Google

Staff Software Engineer, ML Performance, GPU

Google$207K — $300K *
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree or equivalent practical experience.
  • 8 years of experience in software development.
  • 5 years of experience with ML design and ML infrastructure.
  • Experience with modern GPU architectures and programming techniques.
  • Knowledge of low-level GPU programming (CUDA, Triton).
  • Familiarity with modern LLMs and their deployment on AI accelerators.

Responsibilities

  • Identify and optimize LLM training and serving benchmarks.
  • Collaborate with product teams to enhance ML models on GPU hardware.
  • Conduct architecture-level simulations and performance benchmarking.
  • Analyze performance metrics to identify and resolve bottlenecks.
  • Research and implement efficiency techniques for workloads.

Benefits

  • Comprehensive health, dental, and vision insurance.
  • Generous paid time off and parental leave policies.
  • 401(k) plan with employer contributions.
  • Access to continuous learning and development resources.
  • Flexible working hours and environment.
Full Job Description
info_outline
X In most instances, this position requires in-person interviews as part of the hiring process.

Minimum qualifications:
  • Bachelor's degree or equivalent practical experience.
  • 8 years of experience in software development.
  • 5 years of experience with ML design and ML infrastructure (e.g., model deployment, model evaluation, data processing, debugging, fine tuning).
  • Experience with modern GPU architectures, memory hierarchies, and performance bottlenecks.
  • Experience with low-level GPU programming (CUDA, Triton, CUTLASS, etc.) and performance engineering techniques.
  • Experience with modern LLMs and their deployment on AI accelerators.

Preferred qualifications:
  • Master's degree or PhD in Engineering, Computer Science, or a related technical field.
  • 8 years of experience with data structures and algorithms.
  • 3 years of experience in a technical leadership role leading project teams and setting technical direction.
  • 3 years of experience working in a complex, matrixed organization involving cross-functional, or cross-business projects.
  • Experience in hardware-aware algorithm design and compiler stacks (e.g., OpenXLA), tailoring large-scale ML models and distributed systems for peak performance across accelerator hardware.


About the job

While known for pioneering work with TPUs, GPUs are an equally vital and rapidly expanding frontier within Google's ML infrastructure. GPUs are indispensable to Google's ever-evolving landscape for strategic, pragmatic, and performance-driven reasons - ensuring top performance for our ML models, adapting to ML workloads, achieving results, and influencing next-gen GPU architectures via partnerships.

Core ML's GPU Performance team is responsible for optimizing, modeling, and evaluating GPU systems for comparative analysis and benchmarking for internal and external ML workloads. Our team's focus on performance analysis and optimization identifies opportunities in Google production and research ML workloads and lands optimizations to entire fleet. We evaluate current and future ML workloads and runs performance/total cost of ownership simulations to collect roofline estimates and guide decision-making for the hardware teams.

Individual pay is determined by factors including job-related skills, experience, and relevant education or training.

US: $207000 - $300000 (USD) 20% bonus target equity benefits

Learn more about benefits at Google .

Responsibilities
  • Identify and maintain LLM training and serving benchmarks; use them to identify performance opportunities, drive XLA:GPU/Triton performance and guide XLA releases.
  • Partner with product teams (e.g., Google DeepMind) to onboard, optimize, and scale LLMs and machine learning models on GPU hardware.
  • Conduct architecture-level simulations, performance benchmarking, and roofline analyses using tools like TRT-LLM, vLLM, and SGLang to guide system designs.
  • Analyze fleet-wide performance and efficiency metrics to identify bottlenecks and engineer scalable optimizations across Google's infrastructure.
  • Research and implement model/data efficiency techniques, tooling, and profiling mechanisms to improve workload performance and training efficiency.

About Google

Google is a multinational technology company that specializes in Internet-related services and products. These include online advertising technologies, search engine, cloud computing, software, and hardware. Google was founded in 1998 by Larry Page and Sergey Brin while they were Ph.D. students at Stanford University. The company has grown tremendously since then and has become one of the most valuable companies in the world. Google's mission is to organize the world's information and make it universally accessible and useful.
Learn more about Google
Size
156,500 employees
Market Cap
$1,115.4 billion
Industry
Net Income
$40.2 billion
Founded
1998
5 Year Trend
+23.3%
Revenue
$182.5 billion
NASDAQ

Similar Jobs

More Jobs at Google

More Information Technology Jobs

Find similar Staff Software Engineer, ML Performance, GPU jobs: