Senior AI Engineer

Incredible Health

$135K — $160K *
US-AnywhereRemote in United States
Consumer Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 6+ years of software engineering experience, including 2+ recent years with ML and LLM features from concept to production.
  • Strong skills in Python programming with the ability to run and validate complex ML models independently.
  • Experience evaluating non-deterministic systems using methods like golden datasets and regression tests.
  • Solid understanding of ML principles including classification, ranking, and metrics for user prediction.
  • Production experience with LLM tools and SQL-driven product analysis, plus notebook fluency for model training.

Responsibilities

  • Lead end-to-end development of LLM-powered products, from prototyping to validation and deployment.
  • Own the design and execution of model validation, A/B experiments, and offline validation for search and ranking.
  • Develop and ship production code across the stack, troubleshooting issues like hallucinations and ranking regressions.
  • Translate ambiguous inquiries into actionable data analyses to drive product decisions and enhance code quality through rigorous reviews.

Benefits

  • 100% remote work option available across multiple U.S. states.
  • Opportunity to work on cutting-edge LLM-powered products and features.
  • Collaboration with a diverse team of AI, ML, and product engineers.
  • Engagement in a startup-like environment that values ownership and adaptability.
Full Job Description
Position Summary

This is a hands-on role for an engineer with range: notebook in the morning, eval suite by afternoon, production by Friday - full stack, all the way to the user. You think in systems, tradeoffs, and outcomes, not just models - and you'd rather ship a good feature that helps 10,000 nurses today than a perfect model that stays in a notebook.

This role is 100% remote and can be located in the following states:

Alabama, Arizona, California, Colorado, Connecticut, District of Columbia, Florida, Georgia, Idaho, Illinois, Indiana, Iowa, Maryland, Massachusetts, Michigan, Minnesota, Mississippi, Missouri, Nebraska, Nevada, New Jersey, New York, North Carolina, Ohio, Oregon, Pennsylvania, South Carolina, Tennessee, Texas, Utah, Virginia, Washington, West Virginia

What you'll do
  • Lead LLM-powered products end to end using the foundation models best suited to the job: prototype a capability, validate it with evals, deploy, monitor, and iterate.
  • Take end-to-end ownership of matching, search, and ranking - model design, offline validation, shadow-scoring, and A/B experiments - across embeddings and ElasticSearch.
  • Ship production code across the stack - Python ML services, Rails and React web app - and diagnose whatever breaks: hallucinations, retrieval misses, ranking regressions.
  • Turn ambiguous questions into production-minded data analysis that changes what we build, and raise the bar through rigorous code review with our ML, AI, and product engineers.
What you'll bring
  • 6+ years of software engineering, including 2+ recent years taking ML- and LLM-powered features from concept through deployment and long-term production maintenance.
  • Strong Python software engineering skills - with the depth to independently run, validate, and interpret a ranking-model retrain, an embedding pipeline, or an eval harness.
  • Evaluation for non-deterministic systems: golden datasets, LLM-as-judge, regression tests that gate launches.
  • Core ML foundations - classification, ranking, model evaluation - and the rigor to know what offline metrics predict about real users.
  • Production LLM experience (prompt engineering, tool calling, structured outputs, agent and voice workflows), plus SQL-driven product analysis and the notebook fluency to set up an environment and train or run a model on your own box.
What sets you apart
  • Marketplaces or recommendation systems where both sides say yes.
  • Voice agents in production.
  • Embedding retrieval at scale.
  • Claude Code as a daily force multiplier.
  • Confident full-stack contributor.

Above all, we look for ownership and adaptability: product-driven engineers who start from outcomes, work backwards to impact and step up independently to deliver for the team. If startup ethos is your default setting, you'll fit right in.

Your first 90 days
  • Day 30: Shipped to production across our entire AI product stack. Running our eval suites and core notebooks solo, and already trading substantive PR reviews with our ML engineer.
  • Day 60: Owning a production surface end to end - Lyn's eval loop or the matching retrain cycle - and shipped your first measurable improvement, validated by experiment.
  • Day 90: Led an AI bet from scoping through experiment-validated launch. Operating our production ML services independently and pitching the next quarter's AI roadmap with data behind it.
The stack

Anthropic (Claude) • ElevenLabs • Python/FastAPI + Celery • Rails + React • ElasticSearch + embeddings • Snowflake + dbt + Hex • Claude Code everywhere

Similar Jobs

More Jobs at Incredible Health

More Consumer Technology Jobs

Find similar Senior AI Engineer jobs: