ABOUT THE TEAM Discovery is Anduril's team for taking the newest problems across domains- space, missile systems, air, sensor capability, autonomy, and cyber- and proving out what is worth solving. We build the models, run the tests, and carry what works to the point where a program can pick it up. Discovery works alongside Air Defense, Space, Intelligence, Cyber, GNC, Hardware, and every other group at Anduril to incubate the solutions to the hardest problems.
ABOUT THE JOBDiscovery is building an expeditionary force: a team of engineers who want to solve difficult problems and who navigate unfamiliar territory as an operating standard. As a Discovery engineer, you pick up a concept nobody has proven and build the analysis or prototype that tests it to get to an answer. Whether you spend your time working on hard problems on our autonomy stack, acoustic sensor analysis, or battle space management, we tackle every challenge the same way: by questioning assumptions and building our way to the answer.
We are focused on taking AI models from research into production-onto classified platforms and edge hardware where they have to perform reliably in the real world. As our model portfolio and classified work grow, rigorous, repeatable evaluation of how these models actually perform has become mission-critical.
WHAT YOU'LL DO- Develop Test Scenarios Simulation Environments: Develop comprehensive test scenarios, and simulation environments to assess agentic AI performance in classified simulations.
- Validate AI on Classified & Edge Platforms: Lead the integration and validation of agentic AI systems onto classified platforms and edge hardware.
- Build Reusable Evaluation Pipelines: Build reusable evaluation pipelines, automated test harnesses, and monitoring dashboards for continuous validation.
- Define Actionable Metrics: Figure out what good metrics look like for AI models and traditional models-establishing systematic, historic capture of performance rather than one-shot, deployment-specific measurement.
- Partner Cross-Functionally: Partner with cross-functional teams to define requirements, document test results, and contribute to AI model cards and deployment readiness reviews.
- Own End-to-End Evaluation: Architect, build, and maintain the evaluation and validation infrastructure for our agentic AI and ML models, from unit-level model checks through full-system, scenario-based assessment.
- Troubleshoot and Debug: Analyze and resolve issues uncovered in evaluation and in deployment, ensuring reliability and operational success across every release.
REQUIRED QUALIFICATIONS- Technical Expertise: Bachelor's or Master's degree in Computer Science, Machine Learning, Software Engineering, Computer Engineering, Electrical Engineering, or a related technical field.
- Programming Proficiency: At least 12+ years of hands-on experience writing production-grade code, with strong Python skills for building evaluation pipelines, test harnesses, and tooling.
- Evaluation Ownership Mindset: Demonstrated experience owning evaluation or automated testing for production ML or software systems-test scenarios, metrics, regression suites, and CI infrastructure.
- AI/ML Evaluation Experience: Hands-on experience designing evaluation methodologies for AI or machine learning models-defining metrics, building benchmarks, and assessing model behavior against real-world scenarios.
- Simulation Experience: Experience designing or working extensively with simulation environments to exercise model and system behavior.
- Metrics Discipline: Working knowledge of how to define, capture, and reason about performance metrics for AI models and traditional models, including systematic historic capture over time.
- Systems-Level Thinking: Ability to navigate and contribute to complex systems and established codebases.
- Program Ownership: Comfort operating between technical program management and software engineering-defining requirements, coordinating across teams, and documenting results.
- Real-World Impact: Passion for building the evaluation infrastructure that proves AI works-directly influencing mission-critical outcomes.
- Security Clearance: Must be eligible for a US security clearance.
PREFERRED QUALIFICATIONS- Prior Title Background: Experience as an ML Test & Evaluation Engineer, SDET for ML systems, ML/Evaluation Engineer, Simulation Engineer, or Technical Program Manager for AI/ML.
- Agentic AI Evaluation: Experience designing test methodologies for agentic AI systems-tasking, decision-making, and scenario-based behavior validation.
- Classified / Edge Deployment: Familiarity with validating and deploying AI models onto classified platforms, edge hardware, or resource-constrained environments.
- Model Cards & Readiness Reviews: Experience contributing to AI model cards, deployment readiness reviews, or similar model-governance and release-gating practices.
- Aerospace/Defense T&E: Familiarity with test and evaluation practices in aerospace or defense, including qualification testing, range operations, or operational assessment.
- MLOps & Monitoring: Experience with MLOps tooling, monitoring dashboards, and continuous validation pipelines for ML models in production.
- Programming Skills: Additional experience with Go, C++, or scripting for test automation and tooling.
- Growth into Broader Ownership: Interest in growing from T&E ownership into deeper ownership of the AI evaluation and deployment stack over time.
US Salary Range
$253,000-$336,000 USD
The salary range for this role is an estimate based on a wide range of compensation factors, inclusive of base salary only. Actual salary offer may vary based on (but not limited to) work experience, education and/or training, critical skills, and/or business considerations. Highly competitive equity grants are included in the majority of full time offers; and are considered part of Anduril's total compensation package. Additionally, Anduril offers top-tier benefits for full-time employees, including:
BenefitsAt Anduril, we invest in our people. Our comprehensive, competitive benefits package (available at little to no cost to employees) ensures you're supported in health, recovery, and whatever comes next. For more information, Explore Our Benefits.