AI Training Infrastructure Engineer

Designworks Talent

$120K — $160K *
Enterprise Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of experience in distributed training systems or large-scale machine learning infrastructure.
  • Hands-on support for large AI models and post-training workflows.
  • Strong grasp of challenges related to multi-node GPU training in terms of reliability and scalability.
  • Experience with production machine learning pipeline integration.
  • Solid programming skills in relation to complex distributed systems.
  • Proven ability to independently manage complex projects in fast-paced environments.
  • Comfortable operating with high ownership and minimal oversight.

Responsibilities

  • Build and scale infrastructure for large AI model training across GPU clusters.
  • Design systems to enhance training reliability and optimize resource use.
  • Develop solutions for fault tolerance and recovery in training operations.
  • Integrate AI models into production training pipelines in collaboration with engineering teams.
  • Diagnose and resolve issues affecting training throughput and reliability.
  • Create tools and automation to enhance developer experiences for AI engineers.
  • Establish best practices for training infrastructure and operational protocols.

Benefits

  • Medical, dental, and vision insurance.
  • 401(k) plan with company match.
  • Paid holidays throughout the calendar year.
  • Flexible hybrid work arrangement, requiring three days in the office.
Full Job Description
AI Training Infrastructure Engineer

Location: Hybrid | Bellevue, WA AreaTitles: Senior and Staff (multiple roles available)
Build the Training Infrastructure Powering Next-Generation AI Models
About the Opportunity

We're seeking AI Training Infrastructure Engineers to build and scale the distributed systems that power large-scale AI model training. This team focuses on reliability, efficiency, and operational excellence across GPU clusters, enabling researchers and engineers to train and deploy advanced AI models at scale.

The Opportunity

This is a foundational engineering role focused on building the infrastructure layer behind large-scale AI training workloads. You'll work on distributed training systems, GPU clusters, model pipelines, and the tooling required to make AI development more reliable, efficient, and scalable.

You'll collaborate closely with infrastructure, orchestration, performance, and machine learning teams to solve complex challenges around distributed computing, fault tolerance, training efficiency, and production readiness.

This opportunity is ideal for engineers who enjoy building highly scalable systems and working at the intersection of AI research, infrastructure engineering, and distributed computing.

What You'll Do
  • Build and scale distributed training infrastructure supporting large AI models across large GPU clusters.
  • Design and improve systems that increase training reliability, efficiency, and resource utilization.
  • Develop solutions for fault tolerance, checkpointing, recovery, and large-scale training operations.
  • Integrate AI models into production training pipelines in partnership with platform, orchestration, and performance engineering teams.
  • Diagnose and resolve issues impacting training throughput, stability, reliability, and cost efficiency.
  • Build tools and automation that improve the developer experience for AI researchers and engineers.
  • Establish best practices for training infrastructure, operational processes, and platform reliability.
  • Contribute to the evolution of the AI infrastructure platform as an early member of the engineering team.


What We're Looking For
  • Hands-on experience building and operating distributed training systems or large-scale machine learning infrastructure.
  • Experience supporting large AI models, foundation models, post-training workflows, or similar ML systems.
  • Strong understanding of the reliability, scalability, and efficiency challenges associated with multi-node GPU training.
  • Experience integrating training systems with production machine learning pipelines.
  • Strong programming skills and experience working with complex distributed systems.
  • Ability to independently own technically challenging projects in a fast-moving engineering environment.
  • Comfortable operating with high ownership and limited process overhead.


Preferred Qualifications
  • Experience with distributed training frameworks such as PyTorch Distributed, DeepSpeed, Megatron-LM, Ray, or similar technologies.
  • Experience with supervised fine-tuning (SFT), reinforcement learning from human feedback (RLHF), or other post-training workflows.
  • Background operating AI training infrastructure at scale within a hyperscaler, AI research organization, cloud provider, or GPU cloud environment.
  • Experience optimizing GPU utilization, training performance, or distributed system reliability.
  • Familiarity with Kubernetes, containerized AI workloads, and large-scale infrastructure platforms.


Compensation
  • Competitive base pay for Bellevue market
  • Certain roles are eligible for additional rewards, including merit increases, annual bonus, and stock. These awards are allocated based on individual performance
  • U.S. based employees have access to medical, dental, and vision insurance, a 401(k) plan and company match, employees also receive per calendar year, paid holidays.


Location
  • Hybrid role based in the Bellevue, WA area.
  • Approximately three days per week in the office.
  • Candidates elsewhere in the U.S. who are open to relocation are encouraged to apply.
  • U.S. work authorization is required. Visa sponsorship is not currently available.

Similar Jobs

More Jobs at Designworks Talent

More Enterprise Technology Jobs

Find similar AI Training Infrastructure Engineer jobs: