Site Reliability Engineer III

Jefferson Lab

$118K — $186K *
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • 10+ years in Site Reliability Engineering, DevOps, or similar roles; minimum 2 years of leadership experience
  • Experience with scientific computing, HPC, or research environments preferred
  • Background in establishing operational practices in new facilities is a plus
  • Knowledge of high availability operations and incident management
  • Familiar with containerization and Kubernetes; experience with infrastructure as code tools preferred
  • Solid understanding of storage systems and large-scale data infrastructure

Responsibilities

  • Lead design and operation of monitoring, logging, and alerting systems for HPDF
  • Supervise, mentor, and manage a small team of site reliability engineers
  • Establish and maintain an operational framework for the facility, including on-call structures and documentation
  • Design resilience models incorporating redundancy and disaster recovery
  • Define and report on Service Level Objectives (SLOs) with stakeholders
  • Act as incident commander during significant incidents and lead postmortem processes
  • Drive reliability improvements through automation and process optimization

Benefits

  • Flexible work environment
  • Opportunities for professional development and technical growth
  • Collaboration with leading research physicists and computational scientists
  • Influence on technology choices in a cutting-edge facility
  • Participation in operational planning for a brand-new facility
Full Job Description
The good-faith pay range for this role is $118,400 - $186,850 per year. Actual compensation may vary and may be above the posted range based on factors such as a candidate's skills, experience, education, certifications, and work location.

What your job will be like:

As Lead Site Reliability Engineer on the High Performance Data Facility (HPDF) team, you will play a critical role in establishing and running the reliability practice for the facility's first systems on its path to operations. This is a technical role with manager responsibilities: you will supervise and develop a small team of site reliability engineers, and you will also design and build systems yourself, hands on in the code, the monitoring stack, and the incident response. You will design how the facility stays available and recovers, define and report on the service level objectives that measure how well it serves its users, serve as incident commander for significant incidents, and work day to day with staff at both Jefferson Lab and Berkeley Lab. HPDF is still in design, so there is room in this role to grow into influencing the technology choices the facility is built on. The users you support are research physicists and computational scientists, and helping them succeed is a core measure of this role.

In this job you will:
  • Lead the design, implementation, and operation of monitoring, logging, alerting, and diagnostic tooling for HPDF compute, storage, network, and facility systems, contributing directly to that work as well as directing it.
  • Supervise, mentor, and develop a small team of site reliability engineers: assign and review work, set expectations, give regular feedback, support technical growth, and plan and estimate the multi-person efforts assigned to the team.
  • Establish and maintain the facility's operational framework, including on-call and escalation structure, incident management, change management, and scheduled maintenance, and keep operational records, runbooks, and documentation current.
  • Design the facility's resilience model, including failure domain isolation, redundancy, graceful degradation, and disaster recovery objectives for a geographically distributed facility, and validate that design through testing.
  • Define, implement, and report on Service Level Objectives (SLOs) and Service Level Indicators (SLIs) in collaboration with the architecture team and scientific stakeholders, and hold facility operations to them.
  • Serve as incident commander for significant incidents, own the postmortem process, and drive root cause prevention back into the design and operation of the systems.
  • Drive reliability improvement through automation, process optimization, and the elimination of manual operations, using Python, Go, or shell and standard software development practices.
  • Partner with the architecture team on HPDF technology selection from a reliability standpoint, lead evaluations of vendor and open source technologies, and represent HPDF site reliability engineering in the Berkeley Lab partnership.


Additional Responsibilities
  • Participate in an on-call rotation as the facility moves toward operations.


Lead - Supervisory - Management
  • Supervises a team of site reliability engineers
  • Assigns and reviews work, sets performance expectations, provides regular feedback, conducts performance discussions, and supports the technical development of the team.
  • Participates in hiring for the group. Does not hold fiscal or budget authority.


Experience
  • Required: 10 or more years experience in Site Reliability Engineering, DevOps, systems engineering, or operations engineering, including at least two years leading or supervising engineers. Technical leadership of engineering teams or projects qualifies.
  • Preferred: Supporting scientific computing, HPC, or research environments.
  • Preferred: Establishing operational practice in a new or greenfield facility.
  • Preferred: High availability or around the clock operations.
  • Preferred: Serving as the reliability or availability authority during the design phase of a large system or facility, before it entered operations.
  • Preferred: Evaluating vendor compute, storage, and network solutions against reliability requirements, including acceptance criteria and benchmarking.
  • Preferred: Experience with containers and Kubernetes
  • Preferred: Experience with configuration management and infrastructure as code tools (for example Ansible, Terraform, Puppet).
  • Preferred: Experience with storage systems, data movement, or large scale data infrastructure.
  • Preferred: Experience with IT service management practice and tooling (for example ServiceNow, ITIL).
  • Preferred: Experience with HPC infrastructure and environments.
  • Preferred: Supporting formal project milestone or gate reviews, such as DOE critical decision reviews, and defining KPPs or acceptance criteria.


Education
  • Required: Bachelor's Degree Computer Science or Related Field
  • Preferred: Master's Degree Computer Science or Related Field


Experience and Education Exchange

Education above the minimum may be substituted for experience. Relevant experience may not be substituted for education.

Knowledge, Skills, and Abilities
  • Deep Linux systems expertise, with the ability to troubleshoot across the application, operating system, storage, and network layers and to guide others in doing so.
  • Expertise in designing and operating monitoring and observability stacks (for example Prometheus, Grafana, ELK, OpenTelemetry) and in defining and reporting on SLOs and SLIs.
  • Strong scripting and automation skills (Python, Go, or shell) with standard software development practices, together with the judgment to decide what is worth automating.
  • Demonstrated ability to lead and mentor technical staff: setting expectations, assigning work, giving constructive feedback, and addressing performance issues in a constructive manner.
  • Ability to establish and run incident response and operational process in a production environment.
  • Ability to design for resilience, including failure domain isolation, redundancy, graceful degradation, and recovery objectives, and to validate the design through testing and analysis.
  • Clear written and verbal communication, including the ability to present options, argue persuasively for proposals, and work productively with a community of scientific users and with colleagues across both laboratories.
  • Familiarity with public cloud environments (AWS, Azure, GCP).
  • Networking at scale: IPv4/IPv6, DNS, firewalls and access control lists, high speed interconnects, and data transfer protocols.
  • Ability to review system and vendor designs from a reliability standpoint and to argue a technical position persuasively with architects, vendors, and scientific stakeholders.
  • Load testing, performance analysis, and capacity modeling to validate design assumptions and identify bottlenecks in the data path.
  • Ability to estimate cost and effort for multi-person projects and to plan staffing accordingly.
  • Practical experience developing and deploying AI assisted or autonomous automation for operational work, and routine use of AI tools in day to day engineering.


Similar Jobs

More Jobs at Jefferson Lab

More Information Technology Jobs

Find similar Site Reliability Engineer III jobs: