Anduril Industries

Senior AI Infrastructure Engineer

Anduril Industries • $191K — $253K *
Enterprise Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years of software engineering experience in machine learning infrastructure or distributed systems.
  • Proficiency in Python, Go, or C++ with strong software engineering fundamentals.
  • Hands-on experience with container orchestration (Docker, Kubernetes) and distributed training frameworks.
  • Experience maintaining distributed data pipelines for large-scale unstructured or multi-modal datasets.
  • Proven ability to manage projects from design to production deployment.

Responsibilities

  • Build and maintain scalable training and experimentation infrastructure for AI models.
  • Identify and resolve ML lifecycle bottlenecks with new tooling for experiment tracking.
  • Develop robust data pipelines (ETL) to process multi-modal data from various sources.
  • Deploy high-throughput model serving frameworks suitable for cloud and edge environments.
  • Create CI/CD pipelines for ML models with automated testing and validation.
  • Implement evaluation pipelines ensuring model predictability in safety-critical deployments.
  • Collaborate with AI researchers to develop scalable infrastructure for modeling requirements.

Benefits

  • Comprehensive, competitive benefits package available at little to no cost.
  • Support for health and recovery needs.
  • Focus on employee investment and well-being.
Full Job Description
ABOUT THE TEAM

The Air Dominance & Strike team at Anduril develops aerial and multi-domain robotic systems. The team is responsible for taking products like Fury (unmanned fighter jet) and Barracuda (air-breathing cruise missile) from concept to product. The team also develops Lattice for Mission Autonomy, Anduril's premier software platform that enables masses of Fury, Barracuda, and other first and third party robots to collaborate across various missions. We work in close coordination with specialist teams like Perception, Motion Planning, Hardware, and Test Engineering to solve some of the hardest problems facing our customers. We are looking for software engineers and roboticists excited about creating a powerful autonomy software stack that includes computer vision, motion planning, SLAM, controls, estimation, and secure communications.

ABOUT THE JOB

We are looking for a Senior AI Infrastructure Engineer to build, scale, and optimize the end-to-end machine learning platform that powers Anduril's autonomous systems.

In this role, you will own critical components of our ML platform and MLOps tooling. You will build and operate the infrastructure required to train, evaluate, host, and serve complex AI models (including LLMs, computer vision, and RL agents) across cloud environments and air-gapped, edge-deployed networks. Working closely with AI Research Scientists and Platform Engineers, you will eliminate friction in model development, optimize hardware utilization, and ensure robust delivery of models into safety-critical operational environments.

WHAT YOU'LL DO
  • Build, optimize, and maintain scalable training, orchestration, and experimentation infrastructure to accelerate state-of-the-art model development.
  • Identify and resolve bottlenecks in the ML lifecycle by developing tooling for experiment tracking, automated profiling, and hyperparameter tuning.
  • Implement and scale robust data pipelines (ETL) to process multi-modal data (video feeds, radar, flight telemetry, and simulation logs) captured from physical assets and test sites.
  • Deploy high-throughput, low-latency model serving frameworks optimized for both cloud environments and resource-constrained, air-gapped tactical edge hardware.
  • Develop robust CI/CD pipelines for ML models, including automated regression testing, validation benchmarks, and safe rollout/rollback strategies.
  • Implement pipelines for model evaluation, validation, and reinforcement learning alignment loops (RLHF/DPO) to ensure predictability and safety in mission-critical deployments.
  • Partner with AI Researchers and Computer Vision engineers to translate modeling requirements into scalable, reusable infrastructure.
  • Mentor peers, conduct thorough design and code reviews, and champion engineering best practices across the team.

REQUIRED QUALIFICATIONS
  • 5+ years of software engineering experience with demonstrated success in building and operating production-scale machine learning infrastructure or distributed systems.
  • Proficiency in Python, Go, or C++, with a strong grasp of software engineering fundamentals, systems design, and concurrent programming.
  • Hands-on experience with container orchestration (Docker, Kubernetes) and distributed training frameworks (e.g., PyTorch Distributed, Ray, Slurm, or Megatron-LM).
  • Experience building and maintaining distributed data pipelines handling large-scale unstructured or multi-modal datasets.
  • Track record of owning projects end-to-end-from technical design to production deployment and operational monitoring.
  • Eligible to obtain and maintain an active U.S. Top Secret security clearance.

PREFERRED QUALIFICATIONS
  • Experience deploying ML infrastructure, model serving, or artifacts in secure, air-gapped, or regulated environments (e.g., IL5/IL6, GovCloud).
  • Hands-on experience profiling GPU/accelerator workloads, resolving hardware/network bottlenecks, and optimizing compute utilization.
  • Experience supporting workloads for Large Language Models, Generative AI, or Reinforcement Learning (RL) pipelines.
  • Experience with multi-tenant cluster management, including fair scheduling, GPU slicing, and quota enforcement.
  • Familiarity with production ML observability frameworks, including data drift detection and automated evaluation pipelines.


US Salary Range

$191,000-$253,000 USD

The salary range for this role is an estimate based on a wide range of compensation factors, inclusive of base salary only. Actual salary offer may vary based on (but not limited to) work experience, education and/or training, critical skills, and/or business considerations. Highly competitive equity grants are included in the majority of full time offers; and are considered part of Anduril's total compensation package. Additionally, Anduril offers top-tier benefits for full-time employees, including:

Benefits

At Anduril, we invest in our people. Our comprehensive, competitive benefits package (available at little to no cost to employees) ensures you're supported in health, recovery, and whatever comes next. For more information, Explore Our Benefits.

About Anduril Industries

Anduril Industries is a defense technology company that develops advanced systems for the military. The company was founded in 2017 by Palmer Luckey, Trae Stephens, and Matt Grimm, and has since grown to become a major player in the defense industry. Anduril's products include autonomous drones, surveillance systems, and other advanced technologies that are designed to enhance military capabilities. The company has received significant funding from investors and has partnerships with several major defense contractors. Anduril is headquartered in Mountain View, California.
Learn more about Anduril Industries
Size
200 employees
Industry
Founded
2017

Similar Jobs

More Jobs at Anduril Industries

More Enterprise Technology Jobs

Find similar Senior AI Infrastructure Engineer jobs: