Deep systems thinking and strong distributed systems fundamentals.
Experience designing, building, and operating large-scale production systems.
Strong judgment around tradeoffs involving latency, reliability, capacity, and cost.
Experience taking ownership of complex infrastructure from architecture through production operation.
Excitement about applying systems expertise to AI infrastructure and learning quickly as the underlying technology evolves.
Responsibilities
Partner with frontier labs and providers to supply capacity and infrastructure.
Shape Sierra's inference architecture for traffic flow across models and infrastructure.
Build systems for low latency and high reliability in production workloads.
Run models on GPU infrastructure, managing containers and compute capacity.
Optimize inference performance with techniques like speculative decoding.
Build across a hybrid inference stack with Sierra-managed and third-party platforms.
Push the serving stack forward by tuning engines for Sierra's workloads.
Support the broader model lifecycle in collaboration with Models and Agent Runtime teams.
Benefits
Flexible (unlimited) paid time off
Medical, dental, and vision benefits for you and your family
Life insurance and disability benefits
Retirement plan dependent on country of employment
Parental leave
Fertility and family building benefits through Carrot
Lunch, as well as delicious snacks and coffee
Discretionary benefit stipend for personal spending
Free alphorn lessons
Full Job Description
About the role
Sierra's AI agents depend on foundation models to reason and act in real time. The Inference team builds the systems that make those models fast, reliable, and efficient at scale.
As a Software Engineer on Inference, you'll help define Sierra's inference architecture across both self-hosted models and third-party inference providers. You'll work on the systems responsible for serving and routing inference, managing capacity and quota, and optimizing for latency, reliability, and cost.
This is a systems-first role at the intersection of distributed infrastructure and AI. You don't need to be an ML researcher-we're looking for engineers who love complex systems problems and are excited to apply that expertise to one of the fastest-moving areas of AI infrastructure.
What you'll do
Partner with frontier labs and providers. At our scale, we rely on frontier labs, and inference providers to supply capacity, training and inference infrastructure.
Shape Sierra's inference architecture. Design how inference traffic flows across models, infrastructure, and providers, including new serving and proxy layers as Sierra scales.
Build for low latency and high reliability. Develop systems for routing, failover, capacity management, and quota that keep inference performant and available across large-scale production workloads.
Build and operate self-hosted inference. Run models on GPU infrastructure, from building containers and operating inference engines to managing the underlying compute capacity.
Optimize inference performance. Work with the Applied Research team on techniques such as speculative decoding and serving-engine optimizations that improve latency, throughput, and cost.
Build across a hybrid inference stack. Work with both Sierra-managed infrastructure and leading inference platforms, making architectural decisions about where and how workloads should run.
Push the serving stack forward. Work closely with inference providers to tune engines and infrastructure for Sierra's workloads.
Support the broader model lifecycle. Contribute to infrastructure that enables post-training while partnering closely with our Models and Agent Runtime teams.
What you'll bring
Deep systems thinking and strong distributed systems fundamentals.
Experience designing, building, and operating large-scale production systems.
Strong judgment around tradeoffs involving latency, reliability, capacity, and cost.
Experience taking ownership of complex infrastructure from architecture through production operation.
Excitement about applying systems expertise to AI infrastructure and learning quickly as the underlying technology evolves.
Even better
Experience with ML infrastructure, MLOps, or production inference systems.
Experience serving LLMs or other large models at scale.
Experience operating self-hosted inference and GPU infrastructure.
Familiarity with inference frameworks such as vLLM or SGLang.
Experience with post-training infrastructure or inference-performance optimization.
What we offer
We want our benefits to reflect our values and offer the following to full-time employees:
Flexible (unlimited) paid time off
Medical, dental, and vision benefits for you and your family
Life insurance and disability benefits
Retirement plan dependent on country of employment
Parental leave
Fertility and family building benefits through Carrot
Lunch, as well as delicious snacks and coffee to keep you energized
Discretionary benefit stipend giving people the ability to spend where it matters most
Free alphorn lessons
About Sierra Club
The Sierra Club is a nonprofit organization that promotes environmental conservation. It was founded in 1892 by John Muir and is one of the oldest and largest environmental organizations in the United States. The organization has over 3.8 million members and supporters and is dedicated to protecting the planet's natural resources and wildlife. The Sierra Club engages in a variety of activities, including lobbying for environmental legislation, organizing outdoor activities, and publishing a magazine. The organization is headquartered in Oakland, California.