info_outline
X In most instances, this position requires in-person interviews as part of the hiring process.
Minimum qualifications: - Bachelor's degree or equivalent practical experience.
- 2 years of experience programming in C or Python.
- 2 years of experience with software design and architecture.
- 2 years of experience testing, and launching software products.
- Experience with ML model optimization.
- Experience with ML frameworks such as TensorFlow, JAX, and PyTorch, or ML compilers (e.g., accelerated linear algebra (XLA)).
Preferred qualifications: - Master's degree or PhD in Computer Science or related technical fields.
- Experience developing accessible technologies.
- Experience with debugging correctness and performance issues at all levels of the ML software stack.
- Experience with ML compilers and their internals, experience writing compiler optimization passes.
- Familiarity with accelerator hardware architectures (TPUs/GPUs).
About the jobWith your technical expertise you will manage project priorities, deadlines, and deliverables. You will design, develop, test, deploy, maintain, and enhance software solutions.
On this team, you will own the optimization of the models powering the YouTube algorithm. Your work will focus on model efficiency optimization, such as low-precision quantization , knowledge distillation, parameter sharing, and designing hardware-friendly model architectures, to reduce training and serving costs while maximizing fleet utilization and value delivered to users.
You will build support and optimize new and existing models in our Recommendation System stack, including new model architectures while adapting to next-generation TPU hardware.
You will engage in model and TPU compiler co-design, with opportunities to work across the stack ranging from end-user ML models down to Hardware/Software architecture.
Individual pay is determined by factors including job-related skills, experience, and relevant education or training.
US: $147000 - $210000 (USD) 15% bonus target equity benefits
Learn more about benefits at Google .
Responsibilities - Profile ML workloads, identify compute and memory bandwidth bottlenecks, and optimize accelerator utilization to maximize compute efficiency.
- Explore, implement, and productionize algorithmic efficiency techniques, including low-precision quantization, knowledge distillation, parameter sharing, and attention optimizations.
- Optimize auxiliary serving and distributed data pipelines, including data ingestion, feature transformation, embedding lookups, and memory caching to support real-time training and inference.
- Partner closely with ML model developers and researchers to co-design hardware-friendly model architectures and deploy universal efficiency libraries.