ARM

Staff Software Engineer, AI Compute Infrastructure

ARM$209K — $282K *
Enterprise Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years of experience in cloud, compute, HPC, or distributed infrastructure in production settings.
  • Proficiency in programming with Go, Python, or similar systems programming languages.
  • Hands-on knowledge of Kubernetes, Linux, networking, and storage systems.
  • Experience with GPU, accelerator, or distributed machine-learning workloads.
  • Strong troubleshooting skills for complex systems with effective communication across technical teams.

Responsibilities

  • Build and manage Kubernetes clusters while enhancing workload scheduling and resource utilization.
  • Integrate and validate new CPU and GPU systems, focusing on drivers, networking, and health checks.
  • Analyze and resolve performance and reliability issues in applications and hardware.
  • Collaborate with AI teams to tailor automated solutions for cluster management and upgrades.

Benefits

  • Relocation package including visa sponsorship available for eligible candidates.
  • Collaborative work environment fostering innovation and personal contributions.
  • Opportunity to directly impact the success of cutting-edge AI projects.
  • Flexible hybrid working arrangements tailored by teams based on project needs.
Full Job Description
As an engineer on the AI Compute Infra team, you will design, build, and operate large-scale infrastructure for AI training, fine-tuning, evaluation, and inference. You will work across Kubernetes clusters, accelerator enablement, workload scheduling, high-performance networking, storage, and capacity management, partnering with AI researchers and engineers to improve reliability, performance, scalability, and developer productivity.

Responsibilities:

  • Build and operate Kubernetes clusters while improving workload scheduling, topology-aware placement, capacity use, and recovery.
  • Enable new CPU and GPU systems by integrating and validating drivers, networking, storage, monitoring, and health checks.
  • Investigate performance and reliability issues across applications, cloud infrastructure, clusters, and hardware, then turn findings into lasting improvements.
  • Partner with AI teams to understand their workloads and automate cluster provisioning, upgrades, monitoring, and maintenance around their needs.


Necessary Skills and Experience:

  • 5+ years of experience building or operating cloud, compute, HPC, or distributed infrastructure in a production environment.
  • Programming experience in Go, Python, or another systems language, with an interest in developing reliable infrastructure software.
  • Practical knowledge of Kubernetes, containers, Linux, networking, and storage.
  • Experience supporting GPU, accelerator, or distributed machine-learning workloads.
  • An ability to troubleshoot complex systems and communicate clearly with engineers from different technical backgrounds.


"Preferred" Skills and Experience:

  • Familiarity with Kubernetes scheduling, operators, quotas, or resource management.
  • Experience with NVIDIA technologies such as CUDA, NVLink, NVSwitch, NCCL, EFA, or DCGM.
  • Knowledge of AWS EKS, Terraform, Argo CD, Helm, Prometheus, or Grafana.
  • Familiarity with frameworks such as PyTorch, Ray, vLLM, SGLang, or TensorRT-LLM, or experience qualifying accelerators and tuning distributed workloads.


In Return:

You will join a driven group committed to developing world-class AI compute infrastructure. We provide a cooperative setting where your ideas can come to life. Your efforts will directly impact the success of our AI projects, guaranteeing smooth operations and outstanding results. Join us and help build the future of AI compute infrastructure!

Additional Information

Please note that a relocation package (including visa sponsorship support) is available for this role, for candidates who require it.

Salary Range:

$209,100-$282,900 per year
We value people as individuals and our dedication is to reward people competitively and equitably for the work they do and the skills and experience they bring to Arm. Salary is only one component of Arm's offering. The total reward package will be shared with candidates during the recruitment and selection process.

Accommodations at Arm

At Arm, we want to build extraordinary teams. If you need an adjustment or an accommodation during the recruitment process, please email [redacted] . To note, by sending us the requested information, you consent to its use by Arm to arrange for appropriate accommodations. All accommodation or adjustment requests will be treated with confidentiality, and information concerning these requests will only be disclosed as necessary to provide the accommodation. Although this is not an exhaustive list, examples of support include breaks between interviews, having documents read aloud, or office accessibility. Please email us about anything we can do to accommodate you during the recruitment process.

Hybrid Working at Arm

Arm's approach to hybrid working is designed to create a working environment that supports both high performance and personal wellbeing. We believe in bringing people together face to face to enable us to work at pace, whilst recognizing the value of flexibility. Within that framework, we empower groups/teams to determine their own hybrid working patterns, depending on the work and the team's needs. Details of what this means for each role will be shared upon application. In some cases, the flexibility we can offer is limited by local legal, regulatory, tax, or other considerations, and where this is the case, we will collaborate with you to find the best solution. Please talk to us to find out more about what this could look like for you.

About ARM

ARM Holdings is a British multinational semiconductor and software design company, owned by SoftBank Group and its Vision Fund. With its headquarters in Cambridge, England, the company designs microprocessors, physical intellectual property (IP) and related technology and software, and sells development tools to deliver complete solutions for the digital world. ARM's technology is used in a wide range of applications, including automotive, consumer electronics, and Internet of Things (IoT) devices. The company was founded in 1990 and has grown to become one of the world's leading semiconductor IP companies.
Learn more about ARM
Size
6,000 employees
Industry

Similar Jobs

More Jobs at ARM

More Enterprise Technology Jobs

Find similar Staff Software Engineer, AI Compute Infrastructure jobs: