Anduril Industries

Senior AI Infrastructure Engineer

Anduril Industries • $191K — $253K *
Technical Services
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years of software engineering experience in production-scale ML infrastructure or distributed systems.
  • Proficiency in Python, Go, or C++ with a solid foundation in software engineering concepts.
  • Hands-on experience with container orchestration tools like Docker and Kubernetes.
  • Experience in creating distributed data pipelines for large-scale unstructured or multi-modal datasets.
  • Proven ability to manage projects from design through to production deployment and monitoring.
  • Eligibility for a U.S. Top Secret security clearance.

Responsibilities

  • Build and optimize scalable ML training, orchestration, and experimentation infrastructure.
  • Identify and resolve ML lifecycle bottlenecks through tooling development.
  • Implement robust data pipelines for processing multi-modal data from various sources.
  • Deploy optimized model serving frameworks for cloud and edge environments.
  • Develop CI/CD pipelines for ML models, ensuring safe rollouts and regression testing.
  • Implement model evaluation and reinforcement learning alignment pipelines.
  • Collaborate with AI Researchers to develop scalable infrastructure in line with modeling needs.

Benefits

  • Comprehensive benefits package available at little to no cost for employees.
  • Focus on health and recovery support for employees.
  • Investments in professional development and training opportunities.
Full Job Description
ABOUT THE TEAM

The Air Dominance & Strike team at Anduril develops aerial and multi-domain robotic systems. The team is responsible for taking products like Fury (unmanned fighter jet) and Barracuda (air-breathing cruise missile) from concept to product. The team also develops Lattice for Mission Autonomy, Anduril's premier software platform that enables masses of Fury, Barracuda, and other first and third party robots to collaborate across various missions. We work in close coordination with specialist teams like Perception, Motion Planning, Hardware, and Test Engineering to solve some of the hardest problems facing our customers. We are looking for software engineers and roboticists excited about creating a powerful autonomy software stack that includes computer vision, motion planning, SLAM, controls, estimation, and secure communications.

ABOUT THE JOB

We are looking for a Senior AI Infrastructure Engineer to build, scale, and optimize the end-to-end machine learning platform that powers Anduril's autonomous systems.

In this role, you will own critical components of our ML platform and MLOps tooling. You will build and operate the infrastructure required to train, evaluate, host, and serve complex AI models (including LLMs, computer vision, and RL agents) across cloud environments and air-gapped, edge-deployed networks. Working closely with AI Research Scientists and Platform Engineers, you will eliminate friction in model development, optimize hardware utilization, and ensure robust delivery of models into safety-critical operational environments.

WHAT YOU'LL DO
  • Build, optimize, and maintain scalable training, orchestration, and experimentation infrastructure to accelerate state-of-the-art model development.
  • Identify and resolve bottlenecks in the ML lifecycle by developing tooling for experiment tracking, automated profiling, and hyperparameter tuning.
  • Implement and scale robust data pipelines (ETL) to process multi-modal data (video feeds, radar, flight telemetry, and simulation logs) captured from physical assets and test sites.
  • Deploy high-throughput, low-latency model serving frameworks optimized for both cloud environments and resource-constrained, air-gapped tactical edge hardware.
  • Develop robust CI/CD pipelines for ML models, including automated regression testing, validation benchmarks, and safe rollout/rollback strategies.
  • Implement pipelines for model evaluation, validation, and reinforcement learning alignment loops (RLHF/DPO) to ensure predictability and safety in mission-critical deployments.
  • Partner with AI Researchers and Computer Vision engineers to translate modeling requirements into scalable, reusable infrastructure.
  • Mentor peers, conduct thorough design and code reviews, and champion engineering best practices across the team.

REQUIRED QUALIFICATIONS
  • 5+ years of software engineering experience with demonstrated success in building and operating production-scale machine learning infrastructure or distributed systems.
  • Proficiency in Python, Go, or C++, with a strong grasp of software engineering fundamentals, systems design, and concurrent programming.
  • Hands-on experience with container orchestration (Docker, Kubernetes) and distributed training frameworks (e.g., PyTorch Distributed, Ray, Slurm, or Megatron-LM).
  • Experience building and maintaining distributed data pipelines handling large-scale unstructured or multi-modal datasets.
  • Track record of owning projects end-to-end-from technical design to production deployment and operational monitoring.
  • Eligible to obtain and maintain an active U.S. Top Secret security clearance.

PREFERRED QUALIFICATIONS
  • Experience deploying ML infrastructure, model serving, or artifacts in secure, air-gapped, or regulated environments (e.g., IL5/IL6, GovCloud).
  • Hands-on experience profiling GPU/accelerator workloads, resolving hardware/network bottlenecks, and optimizing compute utilization.
  • Experience supporting workloads for Large Language Models, Generative AI, or Reinforcement Learning (RL) pipelines.
  • Experience with multi-tenant cluster management, including fair scheduling, GPU slicing, and quota enforcement.
  • Familiarity with production ML observability frameworks, including data drift detection and automated evaluation pipelines.


US Salary Range

$191,000-$253,000 USD

The salary range for this role is an estimate based on a wide range of compensation factors, inclusive of base salary only. Actual salary offer may vary based on (but not limited to) work experience, education and/or training, critical skills, and/or business considerations. Highly competitive equity grants are included in the majority of full time offers; and are considered part of Anduril's total compensation package. Additionally, Anduril offers top-tier benefits for full-time employees, including:

Benefits

At Anduril, we invest in our people. Our comprehensive, competitive benefits package (available at little to no cost to employees) ensures you're supported in health, recovery, and whatever comes next. For more information, Explore Our Benefits.

About Anduril Industries

Anduril Industries is a defense technology company that develops advanced systems for the military. The company was founded in 2017 by Palmer Luckey, Trae Stephens, and Matt Grimm, and has since grown to become a major player in the defense industry. Anduril's products include autonomous drones, surveillance systems, and other advanced technologies that are designed to enhance military capabilities. The company has received significant funding from investors and has partnerships with several major defense contractors. Anduril is headquartered in Mountain View, California.
Learn more about Anduril Industries
Size
200 employees
Industry
Founded
2017

Similar Jobs

More Jobs at Anduril Industries

More Technical Services Jobs

Find similar Senior AI Infrastructure Engineer jobs: