IonQ

Staff Site Reliability Engineer

IonQ$152K — $228K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 7+ years of production engineering experience with recent hands-on reliability work.
  • Hands-on experience with large-scale, fault-tolerant systems on AWS or GCP.
  • Demonstrated observability ownership in production systems with managed service-level objectives.
  • Experience executing chaos experiments and disaster-recovery exercises with validation.
  • Experience in commanding serious SEV1/SEV2 incidents and leading root cause analysis.
  • Proven track record of reliability outcomes such as availability and recovery metrics.

Responsibilities

  • Own production reliability and represent it in architecture and scaling decisions.
  • Design and operate the observability stack for comprehensive service instrumentation.
  • Define, manage, and review service-level objectives and error budgets with service owners.
  • Execute chaos engineering to validate failure coverages and safeguard effectiveness.
  • Lead high-severity incident response and define incident command processes.
  • Establish and manage escalation paths and on-call rotations for continuous coverage.
  • Drive disaster-recovery testing and improvement based on exercise findings.

Benefits

  • Comprehensive medical, dental, and vision plans.
  • Matching 401(k) program.
  • Unlimited PTO and paid holidays.
  • Parental/adoption leave policy.
  • Legal insurance support.
  • Home technology stipend.
Full Job Description
Location: This role is based at our Santa Clara, CA office, with the option to work a few days a week remotely.
Travel: Up to 25%Job ID: 1874

The Role:

The Platform Engineering team builds, secures, and operates scalable infrastructure for cloud-managed SaaS products with on-premises components deployed at customer sites.

The Site Reliability Engineering discipline keeps the platform stable and reliable, with a strong focus on service continuity and customer experience. It owns production reliability, service-level objectives, observability architecture, backup and disaster recovery, incident response, and resilience, and co-owns cloud security posture and runtime vulnerability management with DevSecOps.

As Staff Site Reliability Engineer, you set the technical direction for reliability across regions and services. You own the reliability strategy, define the standards and mechanisms that guide production operations, and raise the bar through design leadership, operational discipline, and mentorship. You remain deeply hands-on by designing and operating observability platforms, defining and governing SLO programs, leading high-severity incident response, building resilience and disaster-recovery automation, improving reliability of stateful and streaming platforms, and creating AI Ops workflows for triage, remediation, and self-healing.

The work is driven by observability and automation, with a focus on detecting and fixing issues before customers are affected and using every incident to improve the system.
  • Production reliability, SLOs, and error budgets - the reliability of production services end to end, including standards, governance, and escalation for Tier-1 and Tier-2 services.
  • Observability architecture and standards - metrics, logs, distributed tracing, and profiles instrumented across production systems, with consistent platform-wide standards.
  • Chaos engineering and resilience - failure-injection experiments and validation of recovery mechanisms in pre-production and production environments.
  • Backup and disaster recovery - backup validation, disaster-recovery architecture, failover testing, and recovery verification against defined RTO and RPO objectives.
  • Cloud security posture - cloud security posture management, runtime vulnerability detection, and configuration-compliance monitoring, co-owned with DevSecOps.
  • Data and streaming platform reliability - reliability engineering for Postgres, Redis/Valkey, Kafka, OpenSearch, and other critical stateful services.
  • Capacity, efficiency, and AI Ops - resource rightsizing, predictive alerting, autonomous triage, remediation automation, and self-healing workflows.
  • Incident response and command - severity classification, incident command, executive communication, and blameless post-incident review for the highest-severity events.
  • On-call and escalation - rotation design, operational readiness, escalation policy, and clean follow-the-sun handoffs across regions.

Responsibilities:
  • Production reliability - own service-level objectives, error budgets, and production reliability outcomes end to end, and represent reliability in architecture and scaling decisions.
  • Engineer observability - design and operate the observability stack so production services are fully instrumented and define the standards platform and application teams follow.
  • Govern SLOs and error budgets - define and manage service-level objectives, run regular reviews with service owners, and drive corrective action when services consume error budgets unsafely.
  • Drive resilience - design and execute chaos experiments and validate that failure modes are covered by tested safeguards.
  • Lead incident response - define the incident process and serve as incident commander for the highest-severity incidents, including security incidents within the coverage window.
  • Run on-call and escalation - establish and manage rotations and escalation paths that provide continuous coverage with clean follow-the-sun handoffs.
  • Disaster recovery - own disaster-recovery testing and failover validation against defined recovery objectives and turn exercise findings into architectural and operational improvements.
  • Cloud security posture - co-own cloud security posture management, runtime vulnerability detection, and configuration-compliance monitoring with DevSecOps.
  • Data, streaming, and AI Ops - own reliability of stateful and streaming services, capacity planning and rightsizing, and autonomous agents for triage, predictive alerting, remediation, and self-healing.
  • Scale the team and broaden impact- mentor engineers at different seniority levels, set standards adopted across teams, and align Architecture, DevSecOps, Cloud Operations, and Product Development behind a shared reliability roadmap.

Requirements:
  • 7+ years of production engineering experience with recent hands-on reliability work.
  • Hands-on, recent experience operating large-scale, fault-tolerant production systems on AWS or GCP.
  • Observability ownership - has instrumented production systems and governed service-level objectives and error budgets, not only installed dashboards.
  • Resilience practice - has designed and executed failure experiments or disaster-recovery exercises with real failover validation.
  • Incident command - has personally commanded serious SEV1/SEV2 incidents and driven root cause through to a systemic fix.
  • Demonstrated ownership of reliability outcomes with measurable results, such as availability, mean time to recovery, and error-budget adherence.
  • Evidence of multi-team technical leadership through standards, review, coaching, and mechanisms adopted beyond one service or team.

Preferred Qualifications:
  • Proven production experience with cloud security posture management, runtime vulnerability detection, and workload protection across cloud and distributed environments.
  • Strong experience prioritizing risk using identity, workload, and exposure-path context to focus remediation on issues that materially increase attack likelihood and operational impact.
  • Experience with autonomous remediation and self-healing workflows powered by AIOps, including Amazon Bedrock Agent Core or equivalent agentic automation frameworks.
  • Hands-on experience in capacity management, resource rightsizing, efficiency engineering, and practical cost optimization based on FinOps principles.
  • Experience with load-balancing design and operations, including health-based failover, global traffic management, and performance optimization for highly available services.
  • Experience with AI traffic management via an LLM gateway, including request routing, policy enforcement, rate limiting, model fallback, latency optimization, cost controls, and observability for multi-model or multi-provider environments.
  • Ability to connect networking, security, and reliability considerations into cohesive platform design decisions that improve resilience, performance, and operability.


The total compensation package includes base, bonus, equity, and a range of benefit options found on our career site.

If this role has a commission structure, the compensation range below just reflects the base compensation range.

Wage Transparency:

$152,000-$228,000 USD

Compensation will vary based on individual factors such as education, qualifications, and experience of the final candidate(s), specific office location, and calibration against relevant market data and internal team equity. Posted base salary figures are subject to change as new market data becomes available. Our benefits include comprehensive medical, dental, and vision plans, matching 401(k), unlimited PTO and paid holidays, parental/adoption leave, legal insurance, and a home technology stipend. Details of participation in these benefit plans will be provided when a candidate receives an offer of employment.

If you are interested in being a part of our team and mission, we encourage you to apply!

About IonQ

IonQ is a quantum computing company that is developing a general-purpose, full-stack quantum computer based on trapped ion technology. The company was founded in 2015 by Chris Monroe and Jungsang Kim. IonQ's quantum computer is designed to be scalable, error-corrected, and fault-tolerant, and is expected to be able to solve problems that are intractable for classical computers. The company has partnerships with Microsoft, Amazon Web Services, and others. IonQ has raised over $225 million in funding to date.
Learn more about IonQ
Size
100 employees
Market Cap
$633 million
Industry
Founded
2016
NASDAQ

Similar Jobs

More Jobs at IonQ

More Information Technology Jobs

Find similar Staff Site Reliability Engineer jobs: