ML Engineer, Retrieval & Grounded Generation

DEFCON AI

• $165K — $200K *
US-AnywhereRemote in United States
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years of experience in building production or near-production retrieval-augmented systems
  • In-depth knowledge of retrieval design, vector storage, and grounding testing
  • Strong Python skills with hands-on experience in embeddings and vector retrieval
  • Clear distinction between prototypes, proposals, and deployed code in past work
  • US Citizenship and active US Secret clearance required

Responsibilities

  • Implement embeddings, vector storage, and retrieval for a provenance-tracked evidence base
  • Integrate language models to ensure generated text is linked to cited sources
  • Design prompts and output schemas for model interactions
  • Manage model packaging, versioning, serving, and rollback processes
  • Instrument telemetry for quality measurement and performance tracking
  • Provide model assistance for complex narrative extraction tied to source passages
  • Maintain a modular serving path to ensure platform reliability

Benefits

  • Fully remote, results-based work environment
  • Comprehensive health insurance for you and your family, fully employer-paid
  • Unlimited PTO with manager approval
  • Flexible work hours to manage your day
  • 14 weeks of fully-paid parental leave
Full Job Description
About the Role

You'll join the analytics and AI engineering team behind a system that genuinely matters: an AI-assisted platform that pulls together records from dozens of external feeds, resolves them to the right person, surfaces what a human reviewer should look at first, and explains every recommendation in plain, defensible terms - running inside a secure government cloud environment. It's the kind of problem where the details you get right are the ones that count, which is exactly what makes it worth doing well.

As ML Engineer, Retrieval & Grounded Generation, you'll build embeddings, vector storage, and retrieval at scale, and integrate language models so that every piece of generated text is bound to cited source records and citation failures are tested for rather than assumed away. You'll also own prompt and output-schema design; model packaging, versioning, serving, and rollback; and the telemetry hooks that make later measurement possible without manual reconstruction - real infrastructure for a real production system, not a demo.

This is a fully remote role, with occasional travel to DEFCON AI HQ, customer sites, and vendor facilities as required.

Key Responsibilities
  • Implement embeddings, vector storage, and retrieval across a large, provenance-tracked evidence base
  • Integrate language models so generated text is bound to cited source records; test for citation failures rather than assuming them away
  • Design prompts and output schemas
  • Own model packaging, versioning, serving, and rollback
  • Instrument telemetry for retrieval and generation quality, recommendation/version attribution, overrides, abstentions, grounding failures, latency, throughput, and measurement events defined with Model Test
  • Provide bounded model assistance for difficult narrative extraction where deterministic processing is insufficient, with every output tied to its source passage
  • Supply the recorded rule context to every model-assisted step, so each output carries the exact rule versions and ordered context it received
  • Maintain a modular in-boundary serving path, self-hosted or managed, alongside the primary managed inference service, so the platform does not depend on one provider's availability or approval

Required Qualifications
  • 5+ years of experience, including a production or near-production retrieval-augmented (RAG) system you built yourself
  • Ability to speak in detail to your retrieval design, which vector store you used and why, how you tested grounding, what citation failures looked like in practice, and how rollback worked
  • Strong Python, with hands-on experience in embeddings and vector retrieval at scale
  • Clarity on what actually shipped in past work - prototype, proposal, or deployed code - since that distinction matters more here than the title on a resume
  • US Citizenship Required
  • Active US Secret clearance required to start.

Preferred Qualifications
  • Experience deploying models into restricted or air-gapped environments
  • Self-hosted or open-weight model operation
  • Fine-tuning, adapters, or custom embeddings
  • Federal DevSecOps, RMF, ATO, or DoW cloud environment experience
  • Active Top Secret clearance

What Success Looks Like
  • Generated explanations that assert no more than the sources support, with the citation path intact and citation failures tested rather than assumed away
  • A retrieval system that performs at scale on a large, provenance-tracked evidence base
  • Model rollback that works when it's needed, with telemetry complete enough that measurement does not require manual reconstruction

What We Offer
  • A fully remote, results-based environment
  • Competitive salary, bonus, and equity package
  • 100% employer paid, comprehensive health insurance including medical, dental, and vision for you and your family
  • Unlimited PTO, with your manager's approval
  • Flexible work environment where you manage your work day
  • 14 weeks of fully-paid parental leave

Salary Range: $165,000-$200,000. This represents the typical salary range for this position based on experience, skills, and other factors.

Similar Jobs

More Jobs at DEFCON AI

More Information Technology Jobs

Find similar ML Engineer, Retrieval & Grounded Generation jobs: