ML Infrastructure Engineer

SpaceXAI

$180K — $440K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor, Master, Post-graduate, or PhD in computer science, machine learning, or related field, or equivalent work experience.
  • 2+ years in large-scale production environments, distributed systems, or deep learning applications.
  • 2+ years experience with ML platforms and collaboration with modeling engineers and data scientists.
  • Strong proficiency in Python and experience in C++ or Rust.

Responsibilities

  • Design and build GPU compute infrastructure and training frameworks for ML hypotheses.
  • Develop effective data pipelines and integrate large data systems for training and inference.
  • Collaborate with ML teams to integrate models seamlessly across systems.
  • Ensure scalability, reliability, and efficiency of machine learning systems.
  • Independently solve complex problems across the full tech stack.
  • Mentor junior engineers and contribute to team development.

Benefits

  • Comprehensive medical, vision, and dental coverage.
  • Access to a 401(k) retirement plan.
  • Short and long-term disability insurance.
  • Life insurance.
  • Various discounts and perks.
Full Job Description
ABOUT THE ROLE:

As an ML Infrastructure Engineer, you will play a pivotal role in building and optimizing the reliable, high-performance ML platform that powers recommendations on X. We're looking for exceptional engineers who are passionate about our mission and have a strong desire to make a meaningful impact.
RESPONSIBILITIES:
  • Designing, building, and scaling GPU compute infrastructure, training frameworks, and experimentation tools to enable rapid iteration on ML hypotheses
  • Developing data pipelines and integrating large-scale data, training, and inference systems
  • Collaborating with ML teams to productionize models and ensure seamless integration across the stack
  • Ensuring scalability, reliability, and efficiency of large-scale machine learning systems
  • Working across the full stack to solve complex problems independently
  • Mentoring junior engineers and contributing to the growth of the team
BASIC QUALIFICATIONS:
  • Bachelor, Master, Post-graduate or PhD in computer science, machine learning, or other quantitative discipline; or equivalent work experience
  • 2+ years of industry experience working with high traffic or large-scale production environments, distributed systems, GPU infrastructure, and/or deep learning applications
  • 2+ years experience with ML platforms, training infrastructure, or close collaboration with modeling engineers and data scientists
  • Strong proficiency with Python and experience with compiled languages such as C++ or Rust
PREFERRED SKILLS AND EXPERIENCE:
  • Deep familiarity with modern ML frameworks such as JAX or PyTorch
  • Low-level understanding of compute systems, including distributed storage, NVIDIA drivers, CUDA toolkits, and networking
  • Comfortable with Linux systems and orchestration tools
  • Experience with job schedulers (e.g., Slurm), configuration management (Puppet/Ansible), or related infrastructure tooling
COMPENSATION AND BENEFITS:

$180,000 - $440,000 USD

Base salary is just one part of our total rewards package at xAI, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks.

Similar Jobs

More Jobs at SpaceXAI

More Information Technology Jobs

Find similar ML Infrastructure Engineer jobs: