ABOUT THE JOBAs a Senior Site Reliability Engineer on the Undersea Dominance team, you will build and operate the infrastructure that keeps our operational and production systems running at full speed. You'll develop and manage CI/CD pipelines, automate infrastructure with code, and deploy applications across cloud and edge environments with security, traceability, and reliability in mind.
You'll work closely with software, data, and operations engineers to turn designs into working systems-streamlining development, improving performance, and keeping production stable as we scale. You'll also collaborate with digital, manufacturing, and corporate technology teams across Anduril in a high-tech, fast-paced culture of innovation focused on solving real problems and delivering results.
If you're driven to build systems that last, thrive on deep technical challenges, and want to see your work directly shape how we design, build, and sustain complex platforms, you'll be helping build the future of digital shipbuilding and the next generation of maritime vehicles.
WHAT YOU'LL DO- Build and Manage CI/CD Pipelines: Develop and maintain CI/CD pipelines using tools like GitHub Actions and Jfrog Artifactory to ensure seamless integration and deployment of machine learning models and applications.
- Infrastructure as Code (IaC): Utilize Terraform and Ansible to automate infrastructure provisioning and management on cloud platforms such as Azure, AWS, or Google Cloud Platform (GCP).
- Containerization and Orchestration: Implement containerization solutions with Docker and manage container orchestration using Kubernetes to ensure reliable deployment and scaling of applications.
- Monitoring and Logging: Establish comprehensive monitoring and logging solutions using tools like ELK Stack (Elasticsearch, Logstash, Kibana), Prometheus, and Grafana to ensure the smooth operation of deployment environments.
- Collaborate with Cross-Functional Teams: Work closely with development, data science, and operations teams to foster collaboration and ensure the efficient and effective deployment of machine learning models.
REQUIRED QUALIFICATIONS- Advanced proficiency in programming languages (Python for scripting and integration).
- Experience with CI/CD tools like GitHub Actions, Jfrog Artifactory, Git, and CircleCI.
- Proficiency with IaC tools (Terraform, Ansible).
- Experience with cloud platforms (Azure, AWS, GCP).
- Proficiency in containerization (Docker) and container orchestration (Kubernetes).
- Knowledge of model registries and feature stores (e.g., MLflow, Kubeflow).
- Experience with logging and monitoring tools ( Prometheus, Grafana).
- Understanding of parallel computing frameworks (CUDA, OpenCL).
- Strong collaboration skills and proficiency with collaborative tools (JIRA, Confluence).
- Eligible to obtain and maintain an active U.S. Secret security clearance.
PREFERRED QUALIFICATIONS- Experience writing cloud-native services using C++, Rust, Python and/or Go
- Familiarity with observability concepts and tools.
- Knowledge of security best practices for DevOps and MLOps.
US Salary Range
$166,000-$250,000 USD
The salary range for this role is an estimate based on a wide range of compensation factors, inclusive of base salary only. Actual salary offer may vary based on (but not limited to) work experience, education and/or training, critical skills, and/or business considerations. Highly competitive equity grants are included in the majority of full time offers; and are considered part of Anduril's total compensation package. Additionally, Anduril offers top-tier benefits for full-time employees, including:
BenefitsAt Anduril, we invest in our people. Our comprehensive, competitive benefits package (available at little to no cost to employees) ensures you're supported in health, recovery, and whatever comes next. For more information, Explore Our Benefits.