Senior Machine Learning Engineer

TensorWave

• $150K — $180K *
US-AnywhereRemote in Las Vegas, NV
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, or a related field, or equivalent experience
  • Expertise in supporting production ML systems using SLURM and Kubernetes
  • Strong understanding of GPU workloads and distributed systems
  • Solid Linux fundamentals with experience in debugging infrastructure issues
  • Proficient in building automation tools using Python, Go, or similar languages

Responsibilities

  • Design, operate, and enhance ML infrastructure for distributed training and inference
  • Build reliable execution and orchestration patterns in shared GPU environments
  • Troubleshoot performance, reliability, and scalability issues in the ML stack
  • Collaborate with ML, systems, and platform teams to enhance developer experience
  • Support operational efficiency
  • Drive improvements in ML infrastructure processes

Benefits

  • Stock Options
  • 100% paid Medical, Dental, and Vision insurance for Employees
  • Company contributions to Health Savings Accounts
  • 100% paid Short Term and Long Term Disability Insurance
  • Life and Voluntary Supplemental Insurance options
  • Miscellaneous insurance options including Pet & Legal Insurance
  • Various Supplementary Health Benefits such as discounted Virtual Healthcare
  • Flexible Spending Account
  • 401(k) plan
  • Employee Assistance Program
  • Flexible PTO
  • Paid Holidays
  • Parental Leave
  • In-Office perks
Full Job Description


About the Role

We're looking for a Senior Machine Learning Engineer to join our team during an exciting phase of growth. In this role, you'll be responsible for building and operating the core systems that power large-scale ML training and inference across TensorWave's GPU platform, working closely with cross-functional partners to support business objectives while upholding our standards for excellence, collaboration, and impact.

What You'll Do
  • Design, operate, and improve ML infrastructure systems supporting distributed training and inference workloads
  • Build reliable, repeatable workload execution and orchestration patterns across shared GPU environments
  • Troubleshoot performance, reliability, and scalability issues across the ML stack
  • Partner with ML, systems, and platform teams to improve developer experience and operational efficiency


Who You Are

Required Qualifications
  • Bachelor of Science in Computer Science, Computer Engineering, or a related technical field, or equivalent practical experience
  • Expertise supporting production ML systems using SLURM and Kubernetes
  • Strong understanding of GPU-accelerated workloads and distributed systems concepts
  • Solid Linux fundamentals and experience debugging infrastructure-level issues
  • Ability to build automation and tooling - Python, Go, etc.

Preferred Qualifications
  • Experience working across schedulers, orchestration platforms, or cluster managers
  • Familiarity with large-scale GPU environments or HPC-style systems
  • Experience improving infrastructure reliability, utilization, or performance at scale


What We Offer
  • Stock Options
  • 100% paid Medical, Dental, and Vision insurance for Employees
  • Company Health Savings Account Contributions
  • 100% paid Short Term and Long Term Disability Insurance for Employees
  • Life and Voluntary Supplemental Insurance Options
  • Other Insurance Options, such as Pet & Legal Insurance
  • Various Supplementary Health Benefits, such as discounted Virtual Healthcare Appointments and Serious Illness Support
  • Flexible Spending Account
  • 401(k)
  • Employee Assistance Program
  • Flexible PTO
  • Paid Holidays
  • Parental Leave
  • Other In-Office Perks


Similar Jobs

More Jobs at TensorWave

More Information Technology Jobs

Find similar Senior Machine Learning Engineer jobs: