Minimum qualifications:- Bachelor's degree or equivalent practical experience.
- 2 years of experience with software development in one or more programming languages.
- 2 years of experience with performance, large-scale systems data analysis, visualization tools, or debugging.
- 2 years of experience with computer architecture, performance analysis, and performance modeling.
Preferred qualifications:- Master's degree or PhD in Computer Science or related technical fields.
- 2 years of experience with data structures and algorithms.
- Experience developing accessible technologies.
About the jobIn this role, you will bridge the gap between ML workloads and custom Tensor Processing Unit (TPU) hardware. You will analyze and optimize how distributed systems, compiler architectures-such as Accelerated Linear Algebra (XLA)-and emerging software abstractions, such as Compound AI and multi-step agentic systems, execute across Google's AI infrastructure. You will collaborate with various product area architects within Google, such as Google Cloud and YouTube, and external customers to systematically onboard novel workloads with engaged performance. Your optimizations will directly drive TPU adoption, secure pre-sales engagements, and shape our future ML infrastructure roadmap.
Individual pay is determined by factors including job-related skills, experience, and relevant education or training.
US: $147000 - $210000 (USD) 15% bonus target equity benefits
Learn more about benefits at Google .
Responsibilities - Develop and scale benchmarking and workload characterization strategies to enable fast grounding-to-silicon, root-cause performance analysis, and TPU mapping optimization.
- Drive full-stack hardware-software co-design to optimize current and future ML accelerator architectures for business-critical production models (e.g., LLMs and embedding models).
- Partner with Product Areas (e.g., YouTube and Ads) to scale key workload pipelines efficiently (Perf/$/Watts) during TPU Pilot and General Availability (GA) transitions.
- Build and upgrade compiler-aware simulator tools, hardware cost-models, and performance-ladder pathways to baseline and project physical silicon capabilities.
- Distill complex performance analyses and hardware trade-offs into presentations to guide TPU roadmap decision-making in core leadership forums (e.g., ArchForums, NPI, BCR reviews, and TdJs).