ML Infra Engineer

Humble Robotics

$135K — $160K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of experience building high-availability web services in cloud environments
  • Proficiency with infrastructure-as-code tools like Terraform and Ansible
  • Experience developing and managing CI/CD pipelines
  • Strong understanding of security best practices in Linux and network environments
  • Hands-on experience with large-scale batch computation and cluster scheduling
  • Capable of writing and extending code beyond simple scripting
  • Must be eligible to work in the U.S.

Responsibilities

  • Build reliable data collection infrastructure for sensor data to ML platform
  • Develop efficient batch compute pipelines for high-quality training sets
  • Design and scale distributed ML training operations on GPU clusters
  • Own performance, observability, and security of the entire data pipeline
  • Collaborate with ML team to create infrastructure that enhances their workflow

Benefits

  • Base salary plus benefits and equity compensation
  • Flexible work environment
  • Opportunity to work with innovative technology
  • Impactful role in shaping early-stage projects
  • Small team environment promoting rapid iteration and communication
Full Job Description
Position Overview

We're looking for an ML infrastructure engineer to help design, build, and scale the foundational systems we need to realize our ambitious vision. You'll work on tooling and infrastructure that supports every stage of the ML training flywheel and be an important voice in the technical and organizational decisions that shape our work. From areas spanning vehicle compute to data collection to dataset curation to large-scale model training and deployment, help us build reliable, performant, and secure infrastructure that every team at Humble Robotics can rely on. It's fun here. We are doing cool stuff.

The ideal candidate is a first-principles thinker who is comfortable being a broad generalist. Work on every layer of the stack to help make the software iteration loop as fast and efficient as possible. We're a small team, and your input, experience, and knowledge will play a critical role in shaping every system we build, operate, and depend on to achieve our mission.

Key Responsibilities

  • Work on data collection infrastructure that moves sensor data reliably and efficiently from our vehicles into our ML platform
  • Develop batch compute pipelines for cataloging, exploring, and curating raw data into high-quality training sets
  • Design and scale distributed ML training on our GPU clusters
  • Take ownership of performance, observability, efficiency, and security across the full pipeline
  • Partner with the ML team to understand their workflows and translate them into reliable infrastructure that accelerates their work


Minimum Qualifications

  • Experience building and operating high-availability web services on cloud infrastructure
  • Experience with infrastructure-as-code and configuration management tools (we use Terraform and Ansible)
  • Experience building and maintaining CI/CD pipelines and managing deployments
  • Fluent in security fundamentals including Linux hardening, network security, and cryptographic principles
  • Hands-on experience with cluster scheduling systems for running large-scale batch computation
  • Comfortable reading, writing, and extending non-trivial code (not just scripting)
  • Eligible to work in the United States


Preferred Qualifications

  • Hands-on experience managing large, high-performance ML training clusters
  • Working knowledge of distributed training frameworks and high-performance networking for ML workloads
  • Prior infrastructure experience at an early-stage autonomous vehicle or robotics company
  • Comfort operating as an early team member-high ownership, low ego, fast iteration


Compensation

This role is eligible for base salary + benefits + equity compensation. Salary ranges are determined by role, level, and location. Within the range, individual pay is determined by additional factors, including qualifications, skills, experience, and location.

Additional Information

As part of the interview process, we may use Artificial Intelligence (AI) tools to compare your qualifications and experience to the job description. A human reviews all AI output and makes a final hiring decision. Humble Robotics does not rely on the output to make any employment decisions. Some applicants may have a legal right to opt-out of the use of AI as part of our interview process. Contact **[email protected]** to exercise this right or if you have further questions on the use of AI tools in our hiring process.

Similar Jobs

More Jobs at Humble Robotics

More Information Technology Jobs

Find similar ML Infra Engineer jobs: