The OpportunityThe AI Biology & Translation (AIBT) department within Genentech's Computational Sciences Center of Excellence (CS-CoE) is building the next generation of AI systems for biology. Our mission is to develop AI models that learn from biological data at unprecedented scale, generating new insights into disease mechanisms, therapeutic opportunities, and human biology. We seek a highly motivated Senior ML Engineer to join our Foundation Models team and help build and scale the next generation of foundation models and agentic systems for therapeutic discovery.
The successful candidate will contribute to the design, development, and scaling of the next generation of large-scale foundation models and AI agents, with the ultimate aim of accelerating target and drug discovery. In this role, the candidate will build, productionize, and operate these systems, with a strong emphasis on the infrastructure, AgentOps, and MLOps that make them robust, reproducible, and efficient at scale. The candidate will join an exciting, multidisciplinary research environment alongside ML scientists, ML engineers, and computational biologists. The ideal candidate will combine strong software and ML engineering skills, a systems mindset, and a 2get-it-done2 attitude. The selected candidate will be a technical leader, driving the delivery of significant technical solutions, raising engineering quality across projects, and partnering closely with researchers to move ideas into impactful applications.
In this role, you will:- Build, scale, and productionize foundation models and AI agents that support target discovery, experimental design, and lab-in-the-loop pipelines.
- Design and operate the AgentOps and MLOps backbone for these systems, including experiment tracking, model and agent evaluation, monitoring, and reproducible training and inference workflows.
- Design, implement, and maintain scalable and reliable ML infrastructure on AWS, and optimize distributed training and inference on high-performance compute.
- Manage and optimize CI/CD pipelines and Git repositories for ML projects, ensuring efficient version control to support collaboration and deployment.
- Automate deployment, monitoring, and operational tasks using infrastructure-as-code and orchestration tooling (e.g., Terraform, Helm, Kubernetes).
- Proactively identify issues and gaps, propose improvements, and champion engineering best practices and code quality across the team.
- Collaborate closely with interdisciplinary and cross-functional teams across gRED and Roche, and help support research output where relevant, including publications.
Who you are- Educational background: BS/MS in Computer Science, Machine Learning, Engineering, or a related quantitative field.
- Experience: 5+ years of industry experience building and delivering ML systems.
- Technical skills:
- Excellent Python programming skills, and proficiency in scripting languages for automation.
- Solid working knowledge of the theory and practice of deep learning, and hands-on experience with ML frameworks such as PyTorch or JAX.
- Practical experience building, finetuning, deploying, and scaling foundation models, LLMs, and/or agentic systems in production.
- Hands-on experience with the MLOps and/or AgentOps lifecycle, including experiment tracking, model and agent evaluation, monitoring, and reproducible training and inference workflows.
- Proven experience designing, deploying, and managing ML infrastructure on Amazon Web Services (AWS), including services such as EC2, S3, EKS, and SageMaker, along with distributed training and inference on high-performance compute.
- Strong software and data engineering fundamentals, with a proven track record of owning production CI/CD pipelines, Git-based workflows, automated testing, and documentation.
- Demonstrated ability to lead technical projects from conception to completion and deliver high-quality, scalable, and reliable software.
- Excellent problem-solving, communication, and collaboration skills, with the ability to thrive in a fast-paced, user-facing environment.
Preferred:- Familiarity with agent orchestration frameworks (e.g., LangGraph, LangChain) and common patterns such as tool use, retrieval, and multi-step workflows.
- Interest or experience in applying ML to scientific discovery (AI for science), such as biology, chemistry, or drug discovery, including working with domain-specific data and models.
Relocation benefits are
NOT available for this job posting
The expected salary range for this position based on the primary location of San Francisco is $147,800 - 274,400 of hiring range. Actual pay will be determined based on experience, qualifications, geographic location, and other job-related factors permitted by law. A discretionary annual bonus may be available based on individual and Company performance. This position also qualifies for the benefits detailed at the link provided below.
Benefits
#ComputationCoE
#tech4lifeComputationalScience