Distributed Systems ML Infrastructure Engineer

OpenTeams

$145K — $250K *
Enterprise Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • U.S. citizenship and eligibility for security clearance
  • 6+ years experience in distributed systems or software engineering
  • Production experience with Kubernetes and cloud platforms
  • Experience supporting machine learning workloads
  • Proficient in Python, Go, or similar programming languages
  • Familiar with infrastructure-as-code tools like Terraform or Helm
  • Bachelor's degree in related field or equivalent experience

Responsibilities

  • Design and implement core platform services with documented APIs
  • Operate model gateway and routing services with policy enforcement
  • Maintain provider abstractions for cloud and dedicated deployments
  • Deploy releases into controlled environments and ensure validation
  • Verify environment consistency after each promotion
  • Conduct capacity and performance validation against workloads
  • Constrain dependencies to available services and manage development conditions

Benefits

  • 100% medical, dental, and vision coverage for employees
  • 401(k) match up to 5% with full vesting after 2 years
  • Unlimited PTO with a minimum of 15 days off
  • Remote setup with $3,000 equipment reimbursement
  • Continuous education reimbursement of up to $500
  • 100% employer-paid disability and life insurance
  • HSA & FSA options with contributions from employer
Full Job Description
Distributed Systems ML Infrastructure Engineer

Location: Washington, DC; Denver, CO; or Colorado Springs, CO preferred (hybrid). Highly qualified candidates outside these locations may also be considered for unclassified work.

Work Authorization: U.S. citizenship required

Clearance: An active TS/SCI clearance with CI polygraph is strongly preferred. Candidates without an active clearance may be considered for unclassified work but must be eligible to obtain and maintain a U.S. security clearance.

Salary Range: $145,000-$250,000 USD, dependent on experience level and location
About the Role

We're looking for a Distributed Systems and ML Infrastructure Engineer to build the core services of a containerized, API-first AI platform. This is a role for someone who wants to build the thing itself, not integrate someone else's.

You design and implement the services the platform runs on - workflow orchestration, data ingestion, results management, model serving, policy enforcement, usage accounting, audit logging. Those services have to hold up across cloud, dedicated, isolated, and limited-connectivity deployments, which means portability and operability are design constraints from the first commit rather than problems handed to someone downstream.

Development happens primarily on unrestricted infrastructure with an open-source toolchain. Engineers with the right access also carry releases into controlled production environments, integrate data sources there, and validate the platform in place - so there's a path to seeing your work through to where it actually runs.

This position is contingent upon contract award. Travel of up to 15% may be required, primarily to Government facilities and between company locations. Unclassified work may be performed remotely, while classified promotion and validation activities require onsite work in an accredited facility and the appropriate security clearance.
Key Responsibilities
  • Design and implement platform services for workflow orchestration, data ingest, and results management, exposed through documented APIs with no proprietary front end
  • Implement and operate model gateway and serving services that route invocations to approved managed model services with policy enforcement, usage accounting, and audit logging
  • Maintain a documented provider abstraction so the platform runs on AWS-native managed services where appropriate while remaining deployable across other cloud and dedicated environments
  • Deploy platform releases into classified host environments, perform data source integration, and execute validation procedures on a recurring promotion cadence
  • Verify environment parity after each promotion
  • Size and validate the platform against documented workload models, and verify capacity and performance by load test
  • Constrain platform dependencies to services confirmed available in the target environments, and gate any development-only dependency behind feature flags
  • Reproduce high-side defects on the low side through sanitized feedback paths and fix them where the full toolchain is available
Required Skills & Experience
  • U.S. citizenship and eligibility to obtain and maintain a U.S. security clearance
  • 6+ years of experience in distributed systems, platform engineering, infrastructure engineering, or a related software engineering role
  • Production experience operating Kubernetes and containerized workloads on a major cloud platform
  • Experience with managed Kubernetes services such as Amazon EKS or an equivalent platform
  • Experience supporting machine learning workloads in production, such as model serving, GPU scheduling, or large-scale data and evaluation pipelines
  • Experience designing, building, or operating distributed services that support reliability, scalability, and performance requirements
  • Proficiency in Python, Go, or a comparable programming language
  • Experience with infrastructure-as-code and deployment tools such as Terraform, Helm, or equivalent technologies
  • Experience designing API-first services and implementing documented interface specifications
  • Experience testing platform capacity and performance against expected workload requirements
  • Ability to document technical interfaces, deployment procedures, architectural decisions, and validation results
  • Bachelor's degree in computer science, engineering, or a related field, or equivalent practical experience
Nice to Have
  • Active TS/SCI clearance with CI polygraph
  • Hands-on experience deploying or operating software in classified, air-gapped, or limited-connectivity environments
  • Experience with classified cloud environments, including AWS Secret or Top Secret regions
  • Experience building or operating model gateways, LLM routing layers, inference services, or inference-brokering platforms
  • Familiarity with software promotion into classified environments, cross-domain transfer processes, and release packaging
  • Experience with agentic workflow frameworks or the orchestration of multi-step AI pipelines
  • Experience supporting GPU-accelerated workloads and distributed model inference
  • Experience designing vendor-agnostic platforms that operate across multiple cloud or dedicated environments
  • Experience supporting rapid prototyping programs or defense innovation initiatives
What We Offer
  • Medical, Dental & Vision - 100% paid for employees, 75% for dependents
  • 401(k) Match - Up to 5% with full vesting after 2 years
  • Unlimited PTO - With a required minimum of 15 days off annually
  • Fully Remote Setup - Includes up to $3,000 equipment reimbursement
  • Continuous Education - Includes up to $500 reimbursement
  • Disability & Life Insurance - 100% employer-paid
  • HSA & FSA Options - With monthly HSA contributions from OpenTeams

Similar Jobs

More Jobs at OpenTeams

  • Program Manager
    $145K — $250K *
    Colorado Springs, CO 80918 (El Paso County)
    Aerospace & Defense
    In-Person
  • Program Manager
    $145K — $250K *
    Denver, CO 80219 (Denver County)
    Aerospace & Defense
    In-Person
  • Program Manager
    $145K — $250K *
    Washington, DC 20011 (District Of Columbia County)
    Aerospace & Defense
    In-Person
  • Senior AI/ML Test and Evaluation Engineer
    $145K — $250K *
    Colorado Springs, CO 80918 (El Paso County)
    Aerospace & Defense
    In-Person
  • Technical Delivery Lead
    $145K — $250K *
    Washington, DC 20011 (District Of Columbia County)
    Education, Government & Non-Profit
    Hybrid

More Enterprise Technology Jobs

Find similar Distributed Systems ML Infrastructure Engineer jobs: