Minimum qualifications:- Bachelor's degree or equivalent practical experience.
- 8 years of experience in software development.
- 5 years of experience with ML design and ML infrastructure (e.g., model deployment, model evaluation, data processing, debugging, fine tuning).
- 5 years of experience in systems architecture, including building or maintaining distributed systems or large-scale storage architectures.
- 5 years of experience in systems programming (e.g., C , Go, or Rust).
Preferred qualifications:- Master's degree or PhD in Engineering, Computer Science, or a related technical field.
- 8 years of experience with data structures and algorithms.
- 3 years of experience in a technical leadership role leading project teams and setting technical direction.
- 3 years of experience working in a complex, matrixed organization involving cross-functional, or cross-business projects.
About the jobThe AI Storage team within Google Distributed Cloud (GDC) builds the foundational data layer that powers next-generation machine learning and generative AI workloads at the edge and in air-gapped environments. AI models require massive throughput and ultra-low latency; our team tackles the complex challenge of delivering high-performance, massively scalable storage systems directly to customer data centers. We focus on optimizing the data pipeline to keep GPUs and accelerators fully saturated, ensuring that enterprise and public-sector customers can run advanced AI on their most sensitive data without compromising on sovereignty, security, or speed.
Individual pay is determined by factors including job-related skills, experience, and relevant education or training.
US: $207000 - $300000 (USD) 20% bonus target equity benefits
Learn more about benefits at Google .
Responsibilities - Design and develop scalable, distributed File and Object storage solutions that serve as the critical foundational backbone for complex AI/ML workloads within the GDC environment.
- Engineer advanced solutions optimized for massive throughput, specifically enabling high-frequency model checkpointing and the ultra-low latency data access required for AI training and inference.
- Drive technical execution with external partners (e.g., VAST Data) to seamlessly integrate industry-leading, high-performance storage hardware with Google's distributed software ecosystem.
- Take full life-cycle responsibility for core storage services, ensuring uncompromising security, data durability, and high availability across disconnected edge and on-premises data center environments.
- Partner closely with AI infrastructure, compute, and networking teams to architect system-wide improvements, eliminate Input/Output bottlenecks, and deliver a unified "cloud-anywhere" experience.