info_outline
X Applicants in San Francisco: Qualified applications with arrest or conviction records will be considered for employment in accordance with the San Francisco Fair Chance Ordinance for Employers and the California Fair Chance Act.Note: By applying to this position you will have an opportunity to share your preferred working location from the following:
Sunnyvale, CA, USA; San Francisco, CA, USA.
Minimum qualifications: - Bachelor's degree or equivalent practical experience.
- 8 years of experience in software development.
- 5 years of experience leading ML design and optimizing ML infrastructure (e.g., model deployment, model evaluation, data processing, debugging, fine tuning).
- 3 years of experience in a technical leadership role.
- 2 years of experience with state of the art GenAI techniques (e.g., LLMs, Multi-Modal, Large Vision Models) or with GenAI-related concepts (e.g., language modeling, computer vision).
- 2 years of experience in a people management or team leadership role.
Preferred qualifications: - Master's degree or PhD in Engineering, Computer Science, or a related technical field.
- 3 years of experience working in a complex, matrixed organization.
- Experience building multi-tenant architectures with strict access control, tenant isolation, and data governance standards.
- Proven experience designing agent memory systems (short-term/working memory, episodic/semantic memory), context caching (e.g., Gemini context caching), context compression/summarization, vector databases/embeddings, and knowledge graphs.
- Deep background in Information Retrieval (IR), hybrid search (semantic lexical), graph data stores, and aggregating complex distributed state (logs, metrics, IAM, infrastructure topology).
About the jobWith your extensive technical expertise you take initiative to independently design and implement new systems, designing, implementing, and testing multiple features with little or no direction from tech lead or manager. You collaborate with key stakeholders to determine future direction of work.
AutoCloud is Google Cloud's autonomous, AI-powered cloud management portfolio. We are transforming how enterprise customers design, deploy, operate, investigate, and optimize their workloads and infrastructure across GCP. Autonomous agents are only as capable as the context and memory they operate on. The AutoCloud Context and Memory team is responsible for the core cognitive backbone that powers AutoCloud agents: managing short-term dynamic context, long-term episodic and semantic memory, cloud topology graphs, context caching, and intelligent retrieval across petabyte-scale cloud logs, metrics, configurations, and historical runbooks.
Individual pay is determined by factors including job-related skills, experience, and relevant education or training.
US: $207000 - $300000 (USD) 20% bonus target equity benefits
Learn more about benefits at Google .
Responsibilities - Define the technical roadmap and architecture for agent memory systems, dynamic context synthesis pipelines, graph-based cloud representations, and hybrid search/RAG platforms.
- Manage, mentor, and grow an exceptional team of software engineers. Drive talent acquisition, foster an engineering culture, conduct performance evaluations, and support career progress.
- Lead the design and implementation of low-latency context caching, token compression/pruning strategies, working memory buffers, and long-term episodic knowledge stores for autonomous agents.
- Build automated evaluation pipelines and benchmarking frameworks to measure and optimize context retrieval precision, memory recall, grounding fidelity, and hallucination reduction.
- Ensure all memory and context subsystems adhere to the highest standards of multi-tenant enterprise data privacy, tenant isolation, compliance, high availability, and operational efficiency. Partner with AutoCloud agent orchestration teams, and GCP service teams to seamlessly integrate cloud state into agent context.