Nimble Robotics

Senior/Staff Software Engineer, Infrastructure (ML)

Nimble Robotics$210K — $300K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's, Master's, or PhD in Computer Science or a related field, or equivalent practical experience.
  • 4+ years of experience in infrastructure, distributed systems, ML systems, robotics, or related areas.
  • Proficiency in programming languages such as Rust, Go, Python, or C++.
  • Experience with ML frameworks including PyTorch or JAX.
  • Strong understanding of distributed systems, systems programming, and performance optimization.
  • Familiarity with Kubernetes orchestration and containerized deployment pipelines.
  • Ability to debug and optimize bottlenecks in GPU systems.

Responsibilities

  • Design and maintain ML training infrastructure for efficient job execution and rapid experimentation.
  • Build low-latency inference pipelines to support robotics applications.
  • Develop and optimize CUDA kernels for high performance.
  • Create scalable training systems for large datasets and distributed model training.
  • Lead design reviews to assess technologies and technical tradeoffs.
  • Review code for adherence to best practices in software quality.
  • Mentor junior engineers to enhance team capabilities.

Benefits

  • Unlimited flexible time off to recharge and connect with loved ones.
  • Comprehensive health insurance options covering medical, dental, and vision.
  • Paid parental leave for bonding after the birth of a child.
  • Fully-paid parking to ease commuting stress.
  • Cash bonuses for employee referrals that lead to hires.
  • Retirement savings opportunities through a 401k program.
  • Equity opportunities to foster employee ownership in the company.
Full Job Description
About the Role

We're looking for a Software Engineer to join our ML Infrastructure team. In this role, you'll help build the training and inference systems that power our general-purpose warehouse robots.

You'll own training infrastructure end to end: keeping GPUs highly utilized, making runs reproducible, and ensuring every researcher can launch the next experiment with a single command. You'll work closely with ML and Robotics teams to design, build, and scale the systems that turn our GPU clusters into a reliable, high-throughput platform for model development.

Responsibilities
  • Design, develop, and maintain ML training infrastructure that enables the AI team to run training jobs efficiently, manage and iterate experiments quickly.
  • Build low-latency inference pipelines for production robotics workloads.
  • Develop, tune, and optimize low-level CUDA kernels.
  • Design training-platform systems for scalable model training, including high-throughput data ingestion, dataset sharding and sampling for distributed training.
  • Participate in and lead design reviews with peers and stakeholders to evaluate technical tradeoffs and select appropriate technologies.
  • Review code and provide feedback to uphold best practices around style, correctness, testability, performance, and maintainability.
  • Contribute to documentation and educational materials, adapting content as systems and workflows evolve.
  • Mentor junior engineers and help raise the technical bar across the team.

Qualifications
  • Bachelor's, Master's, or PhD in Computer Science or a related field, or equivalent practical experience.
  • 4+ years of industry experience in infrastructure, distributed systems, ML systems, robotics, or a related area.
  • Experience with programming languages such as Rust, Go, Python, or C++.
  • Experience with ML frameworks such as PyTorch or JAX.
  • Strong understanding of distributed systems, systems programming fundamentals, memory management, and performance optimization.
  • Experience with Kubernetes orchestration, resource scheduling for large distributed jobs, and containerized deployment pipelines.
  • Ability to debug and optimize bottlenecks across GPU memory hierarchy, networking fabric, filesystems, and multi-GPU operations.
  • Ability to reason from first principles and optimize systems for both memory-bound and compute-bound workloads.
  • Strong cross-functional communication skills, ownership, and a growth mindset.

Nice to Have
  • Hands-on experience with distributed training frameworks and techniques such as PyTorch DDP/FSDP, DeepSpeed, Megatron, or NCCL.
  • Hands-on experience with GPU kernel development.
  • Experience with data engineering technologies such as Parquet, Arrow, or similar systems.


Compensation

The pay range for this position at the start of employment is expected to be between $210,000 and $300,000/year. Your exact offer may vary depending on multiple individualized factors, including job-related knowledge, skills, and experience. In addition to cash compensation, this position will also receive generous equity for this position.

Culture:

We embrace challenges and strive to make the impossible possible each day. We're not in this to do what's easy or to be mediocre. We want to create something legendary and leave our mark on the world. We're ambitious, we're gritty, we're humble and we're relentlessly resourceful in pursuit of our goals. If this sounds like you then you might be a great fit!

Nimble's Benefits

Unlimited Flexible Time Off

Enjoy the time you need to travel, rejuvenate, and connect with friends and family.

Health Insurance

Nimble provides medical, dental, and vision insurance through several premier plans and options to support you and your family.

Paid Parental Leave

Enjoy paid bonding time following a birth.

Commuter Benefits

Take the stress out of commuting with access to fully-paid parking spots.

Referral Bonus

Get a cash bonus for any friend or colleagues that you refer to us that we end up hiring.

401k

Contribute towards a 401k for retirement planning.

Equity

Be an owner in Nimble through our equity program

Link: Nimble Closes $106 Million Series C Funding Round, Scales Fully Autonomous Fulfillment with FedEx

Link: FedEx Announces Expansion of FedEx Fulfillment With Nimble Alliance

About Nimble Robotics

Nimble Storage, founded in 2008, produced hardware and software products for data storage, specifically data storage arrays that use the iSCSI and Fibre Channel protocols and includes data backup and data protection features. Nimble is now a subsidiary of Hewlett Packard Enterprise.
Learn more about Nimble Robotics

Similar Jobs

More Jobs at Nimble Robotics

More Information Technology Jobs

Find similar Senior/Staff Software Engineer, Infrastructure (ML) jobs: