info_outline
X In most instances, this position requires in-person interviews as part of the hiring process.
Minimum qualifications: - Bachelor's degree or equivalent practical experience.
- 5 years of experience with software development in one or more programming languages.
- 3 years of experience testing, maintaining, or launching software products, and 1 year of experience with software design and architecture.
- 3 years of experience with developing large-scale infrastructure, distributed systems or networks, or experience with compute technologies, storage or hardware architecture.
Preferred qualifications: - Master's degree or PhD in Computer Science, or a related technical field.
- 5 years of experience with data structures and algorithms.
- 1 year of experience in a technical leadership role.
- Experience in one or more of the following: Java, C/C .
- Experience in Generative AI.
About the jobThe Agent Cloud initiative is evolving Google Cloud to handle the unique challenges of non-deterministic AI agents and workloads. Our team bridges the gap between traditional Cloud networking and the new agentic ecosystem by leveraging traditional infrastructure systems like Bouncer, QuotaServer, QuotaStore to provide quota and rate limiting infrastructure under Agent Gateway, the central conduit securing and governing agent-to-agent and agent-to-tool communications.
You will be handling critical industry challenges: LLM token quota management and backend overload protection. You will design the infrastructure that prevents runaway AI costs, ensures fair share of LLM capacity, and protects enterprise AI agents from DDoS and abuse.
Individual pay is determined by factors including job-related skills, experience, and relevant education or training.
US: $174000 - $252000 (USD) 15% bonus target equity benefits
Learn more about benefits at Google .
Responsibilities - Design and operate globally distributed systems paths that handle millions of queries across multiple regions.
- Build distributed scheduling, dynamic reservation, and fair-sharing tuned to how LLMs actually behave - token-by-token streaming, bursty agent loops, speculative decoding, and prompt caching.
- Build self-healing systems with solid failover, fast coordination, and the telemetry needed to hit five-nines (99.999%) availability.
- Set the architectural direction for token orchestration and distribution across Cloud.