Google

Senior Software Engineer, Fleet-level ML Performance

Google$174K — $252K *
Technical Services
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in Computer Science, Electrical Engineering, or related field required.
  • 5 years of experience in systems or computer architecture, focusing on power and performance trade-off analysis.
  • Experience with Reliability, Availability, and Serviceability (RAS) features necessary.
  • Master's degree or PhD in related fields preferred.
  • Knowledge of deep learning workloads and their hardware execution characteristics is a plus.

Responsibilities

  • Conduct fleet-level performance analysis of ML workloads on TPU systems using simulation tools.
  • Collaborate with teams such as model researchers and TPU architects to optimize performance.
  • Design an ML Accelerator Platform Performance Estimation method using C and Python.
  • Work with hardware and software teams to define requirements for AI infrastructure.
  • Evaluate hardware features and software optimizations to inform chip architecture decisions.

Benefits

  • Comprehensive health and wellness programs.
  • Retirement savings plans with company matching.
  • Generous paid time off and holiday schedules.
  • Equity and performance bonus opportunities.
  • Support for continuous learning and development.
Full Job Description
Minimum qualifications:
  • Bachelor's degree in Computer Science, Electrical Engineering, Computer Engineering, or a related field.
  • 5 years of experience in systems architecture or computer architecture, power and performance trade-off analysis, or data center, cloud, infrastructure hardware optimization .
  • Experience with Reliability, Availability, and Serviceability (RAS) features, paradigms, or architecture.

Preferred qualifications:
  • Master's degree or PhD in Electrical Engineering, Computer Engineering or Computer Science, with an emphasis on computer architecture.
  • Knowledge of deep learning workloads, including embedding architectures and their hardware execution characteristics.


About the job

The TPU Chip Architecture and Performance team bridges Google's machine learning workloads and custom silicon architectures. Through hardware/software co-design, we define and shape the future TPU platforms required to meet Google's ambitious AI goals. In this role, you will conduct end-to-end performance analysis of critical ML workloads including Gemini, YouTube, Ads, and key 3P models on next-generation TPU systems. Using advanced simulation and projection methodologies, you will evaluate hardware features and software optimizations to balance performance and cost trade-offs. You will drive high-impact architecture decisions that maximize TPU efficiency while accounting for system-wide constraints such as power, reliability (RAS), and scheduling.
The AI and Infrastructure team is redefining what's possible. We empower Google customers with breakthrough capabilities and insights by delivering AI and Infrastructure at unparalleled scale, efficiency, reliability and velocity. Our customers include Googlers, Google Cloud customers, and billions of Google users worldwide.

We're the driving force behind Google's groundbreaking innovations, empowering the development of our cutting-edge AI models, delivering unparalleled computing power to global services, and providing the essential platforms that enable developers to build the future. From software to hardware our teams are shaping the future of world-leading hyperscale computing, with key teams working on the development of our TPUs, Vertex AI for Google Cloud, Google Global Networking, Data Center operations, systems research, and much more.Individual pay is determined by factors including job-related skills, experience, and relevant education or training.

US: $174000 - $252000 (USD) 15% bonus target equity benefits

Learn more about benefits at Google .

Responsibilities
  • Perform fleet-level performance analysis of key ML workloads (e.g., Gemini) on future TPU systems using advanced simulation tools to evaluate hardware/software trade-offs and guide next-generation chip architecture.
  • Partner with teams across the ML stack including model researchers, compiler developers, systems engineers, and TPU architects to analyze and optimize performance across the design space.
  • Design and implement a unified ML Accelerator Platform Performance Estimation Methodology using C and Python to enable scalable performance and TCO projections across Google.
  • Collaborate cross-functionally with data center, hardware architecture, and framework teams to define key hardware and software requirements for future AI infrastructure.


About Google

Google is a multinational technology company that specializes in Internet-related services and products. These include online advertising technologies, search engine, cloud computing, software, and hardware. Google was founded in 1998 by Larry Page and Sergey Brin while they were Ph.D. students at Stanford University. The company has grown tremendously since then and has become one of the most valuable companies in the world. Google's mission is to organize the world's information and make it universally accessible and useful.
Learn more about Google
Size
156,500 employees
Market Cap
$1,115.4 billion
Industry
Net Income
$40.2 billion
Founded
1998
5 Year Trend
+23.3%
Revenue
$182.5 billion
NASDAQ

Similar Jobs

More Jobs at Google

More Technical Services Jobs

Find similar Senior Software Engineer, Fleet-level ML Performance jobs: