Applied AI Engineer

Soulside AI

$100K — $200K *
Healthcare
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 3+ years in applied ML/AI engineering or a Master's degree in a related field with practical production experience
  • Hands-on expertise in fine-tuning open-source models like Llama or Mistral
  • Familiarity with managed platforms such as Fireworks AI or Baseten
  • Proven ability to construct evaluation frameworks for LLM tasks
  • Solid Python skills and familiarity with ML tooling like PyTorch and Hugging Face
  • Strong understanding of prompt engineering and structured-output validation
  • Ability to work in a fast-paced remote startup environment

Responsibilities

  • Build post-training pipelines for domain-specific clinical tasks
  • Fine-tune and serve models across various inference platforms
  • Design rigorous evaluation sets for critical tasks like clinical reasoning
  • Create fast, trustworthy iteration loops using evaluation results
  • Optimize the entire LLM pipeline including latency and cost
  • Partner with clinical experts to define model behavior and evaluation standards
  • Monitor deployed models for quality, drift, and failures

Benefits

  • Comprehensive health, dental, and vision insurance
  • Flexible remote-first working culture
  • Direct access to founders for influence on technical decisions
  • Professional development budget for learning opportunities
  • Opportunity to create impactful AI in mental health care
Full Job Description
Applied AI Engineer

Soulside AI • US On-Site • Reports to the CTO

The Role

We're looking for an Applied AI Engineer to own the model layer that makes Soulside's documentation trustworthy. In behavioral health, a note isn't just text-it has to be clinically sound, defensible for medical necessity, and safe. Your job is to build the post-training pipelines and evaluation systems that get our models there, and keep them there as we scale.

This is a hands-on role for someone who lives at the intersection of applied ML and product. You'll fine-tune and adapt open-source models, stand up the infrastructure to serve them, and build the rigorous evaluation sets that tell us-objectively-whether a change made the product better or worse.

Why This Role Matters
  • Accuracy Isn't Optional: In behavioral health, a wrong or unsupported note has real clinical and financial consequences. The pipelines and evals you build are what let us ship model changes with confidence.
  • Own the Model Layer: You'll define how we post-train, evaluate, and deploy models end-to-end-not inherit someone else's stack.
  • Direct Clinical Impact: Every improvement in clinical reasoning or note quality directly reduces documentation burden and strengthens the charts clinicians and payers rely on.
What You'll Do
  • Build post-training pipelines on open-source models-supervised fine-tuning, preference optimization (DPO/RLHF), LoRA/adapters, and distillation-for domain-specific clinical tasks.
  • Fine-tune, deploy, and serve models across managed inference and fine-tuning platforms such as Fireworks AI, Baseten, and Together AI, and make pragmatic build-vs-buy calls on where each workload should run.
  • Design and maintain rigorous evaluation sets for high-stakes tasks like clinical reasoning and AI note generation-defining metrics, curating gold-standard data, and building automated and human-in-the-loop eval harnesses.
  • Turn eval results into a fast, trustworthy iteration loop: catch regressions before they ship, and quantify the impact of every model or prompt change.
  • Optimize the full LLM pipeline-prompting, retrieval, structured output validation, latency, and cost.
  • Partner with clinical experts to translate documentation and compliance requirements into model behavior and evaluation criteria.
  • Monitor models in production for quality, drift, and failure modes, and close the loop back into training data and evals.
What We're Looking For
  • 3+ years in applied ML / AI engineering, or a Master's degree in a related field, with hands-on experience taking LLM-based systems into production.
  • Practical experience with post-training / fine-tuning open-source models (e.g., Llama, Qwen, Mistral) using SFT, LoRA/PEFT, or preference-based methods.
  • Experience serving or fine-tuning models on managed platforms such as Fireworks AI, Baseten, or Together AI (or comparable inference/training infra).
  • Demonstrated ability to build evaluation frameworks for LLM tasks-you think in terms of measurable quality, not vibes.
  • Strong Python and familiarity with the modern ML tooling ecosystem (PyTorch, Hugging Face, etc.).
  • Solid grounding in prompt engineering and structured-output validation.
  • Ability to thrive in a fast-paced, remote startup and communicate clearly with technical and clinical teammates.
  • We're willing to sponsor visas, including H-1B and O-1, for the right candidate.
Bonus Points
  • Experience with healthcare, clinical NLP, or other high-stakes / regulated domains.
  • Familiarity with HIPAA and handling sensitive clinical data.
  • RAG systems, retrieval quality tuning, or long-context document workflows.
  • Experience with LLM observability, monitoring, and drift detection in production.
  • Data pipeline and labeling workflow experience for curating high-quality training and eval sets.
  • Open-source contributions in the ML/LLM ecosystem.
What We Offer
  • Salary range of $100,000-$200,000, plus equity with significant upside potential as a founding team member
  • Comprehensive health, dental, and vision insurance
  • Flexible, remote-first culture
  • Direct access to founders and influence on technical direction
  • Professional development budget and conference attendance
  • The chance to build AI that measurably improves mental health care at scale
How to Apply

Send your resume and a short note to [redacted]. Tell us about a model or pipeline you took to production-and how you knew it was actually working.

Similar Jobs

More Jobs at Soulside AI

More Healthcare Jobs

Find similar Applied AI Engineer jobs: