Google

Staff Software Engineer, Inference Performance Optimization, GenAI, DeepMind

Google • $207K — $300K *
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in Computer Science, Engineering, Applied Mathematics, or related field, or equivalent experience.
  • 8 years of software development experience.
  • Proficient in Python and C, with skills in navigating and debugging codebases.
  • Knowledge of AI model execution constraints and modern serving architectures.
  • Experience with real-world LLM inference environments or contributions to open-source inference frameworks.

Responsibilities

  • Analyze and optimize AI inference workloads to enhance throughput and reduce latency.
  • Design and implement techniques for inference optimization.
  • Investigate and resolve performance bottlenecks in model inference.
  • Model latency-to-cost impacts and provide actionable insights for production systems.
  • Develop tools and metrics to track compute usage across the fleet.

Benefits

  • Diverse learning opportunities across global teams.
  • Varied career pathways for driven individuals.
  • Access to cutting-edge AI technologies and research.
  • Collaborative work environment focused on mission-driven results.
Full Job Description
Minimum qualifications:
  • Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, Applied Mathematics, or a related technical field, or equivalent practical experience.
  • 8 years of experience in software development.
  • Experience in Python and C , including navigating, debugging, and modifying serving codebases.
  • Experience with AI model execution constraints, throughput-latency tradeoffs, memory bandwidth limitations, and modern serving architectures.

Preferred qualifications:
  • Experience with real world LLM inference serving environments or direct contributions to modern open-source inference frameworks (e.g., vLLM, TensorRT-LLM, SGLang, Dynamo).
  • Experience profiling workloads using standard ML profilers (e.g., PyTorch profiler) and internal trace analysis tools.
  • Experience with observability and reliability for large distributed systems.
  • Familiarity with GPU/TPU/accelerator performance concepts (e.g. memory bandwidth, quantization, collective communication, kernel), and can reason their implications to the overall inference serving performance.


About the job

As a Software Engineer, you will be working with the cutting edge AI agents developed by our exceptional team of Machine Learning and Neuroscience research scientists. Your responsibilities will include everything from creating systems for agent testing using 2D and 3D games to developing test problems within physics simulators. You will create graphical visualization of results, build competitive agent leaderboards and test new algorithms on robots. To succeed in this role you will need to have a strong foundation in software engineering and enjoy working on a wide range of challenging problems within a mission-driven team.

As an Inference Performance Engineer, you will push the boundaries of AI model execution at scale. In this role, you will be at the forefront of making large-scale AI inference faster, cheaper, and more efficient. You will analyze the entire inference stack to identify critical bottlenecks and drive systemic improvements. By combining deep systems profiling, benchmarking, and first-principles problem solving, your work will directly maximize hardware throughput, reduce cost-to-serve, and empower our cross-functional teams to make data-driven capacity and latency tradeoffs.

We are pushing the boundaries across multiple domains. Our global teams offer diverse learning opportunities and varied career pathways for those driven to achieve exceptional results through collective effort.
Individual pay is determined by factors including job-related skills, experience, and relevant education or training.

US: $207000 - $300000 (USD) 20% bonus target equity benefits

Learn more about benefits at Google .

Responsibilities
  • Analyze and optimize AI inference workloads across the application, model, and distributed fleet infrastructure layers to methodically increase throughput-per-GPU and reduce latency.
  • Design and implement inference optimization techniques.
  • Investigate and resolve complex model inference performance bottlenecks across the stack.
  • Model the latency-to-cost impacts of system variables (such as batch-sizing and utilization goals) and translate these insights into actionable signals that drive production systems.
  • Develop investigative tools and metrics (e.g., compute/FLOPs funnels) that track where compute is spent across the fleet.


Information collected and processed as part of your Google Careers profile, and any job applications you choose to submit is subject to Google's Applicant and Candidate Privacy Policy .

About Google

Google is a multinational technology company that specializes in Internet-related services and products. These include online advertising technologies, search engine, cloud computing, software, and hardware. Google was founded in 1998 by Larry Page and Sergey Brin while they were Ph.D. students at Stanford University. The company has grown tremendously since then and has become one of the most valuable companies in the world. Google's mission is to organize the world's information and make it universally accessible and useful.
Learn more about Google
Size
156,500 employees
Market Cap
$1,115.4 billion
Industry
Net Income
$40.2 billion
Founded
1998
5 Year Trend
+23.3%
Revenue
$182.5 billion
NASDAQ

Similar Jobs

More Jobs at Google

More Information Technology Jobs

Find similar Staff Software Engineer, Inference Performance Optimization, GenAI, DeepMind jobs: