info_outline
X In most instances, this position requires in-person interviews as part of the hiring process.
Minimum qualifications: - Bachelor's degree or equivalent practical experience.
- 5 years of experience with software development in one or more programming languages.
- 3 years of experience testing, maintaining, or launching software products, and 1 year of experience with software design and architecture.
- 3 years of experience with developing large-scale infrastructure, distributed systems or networks, or experience with compute technologies, storage or hardware architecture.
- Experience programming in Go for software development, including AI/ML applications.
Preferred qualifications: - Master's degree or PhD in Computer Science or related technical field.
- 5 years of experience with data structures and algorithms.
- 1 year of experience in a technical leadership role.
- Experience with container orchestration (e.g., Kubernetes) and cloud-based AI platforms.
About the jobJoin our pioneering team dedicated to democratizing access to Google's AI solutions by integrating and developing them deeply within the Google Distributed Cloud (GDC) Platform. We are passionate about enabling customers to harness the power of AI, no matter their environment. Our work spans the full range of GDC offerings:
If you are excited about shaping the future of AI on the edge and in hybrid clouds, and addressing unique challenges across various deployment models, we want to hear from you!
Individual pay is determined by factors including job-related skills, experience, and relevant education or training.
US: $174000 - $252000 (USD) 15% bonus target equity benefits
Learn more about benefits at Google .
Responsibilities - Lead the technical design, development, and optimization of software components critical for Large Language Model (LLM) inference serving on GDC. This includes areas such as model life-cycle management, efficient data loading, dynamic request routing, and intelligent load balancing.
- Drive horizontal integration across core platform services, including billing, logging, observability, security, and quota management.
- Implement and enhance serving capabilities to support advanced LLM techniques like disaggregated serving, speculative decoding, quantization, and efficient model sharding across distributed hardware.
- Collaborate closely with internal teams developing core LLM frameworks, container orchestration (Kubernetes, Google Kubernetes Engine (GKE)), networking infrastructure, and hardware acceleration to build a cohesive and high-performance serving platform.