The good-faith pay range for this role is $91,800 - $145,050 per year. Actual compensation may vary and may be above the posted range based on factors such as a candidate's skills, experience, education, certifications, and work location.What your job will be like:As a Site Reliability Engineer on the High Performance Data Facility (HPDF) team, you will help build and operate the facility's first systems on its path to operations. You will create the monitoring, alerting, and automation that the full facility will eventually run on, participate in incident response, and help define and report on the service level objectives that measure how well the facility serves its users. You will work as part of a small site reliability engineering team, with day to day direction from the team's lead, and alongside staff at both Jefferson Lab and Berkeley Lab. The users you support are research physicists and computational scientists, and helping them succeed is a core measure of this role.
In this job you will:- Build and operate monitoring, logging, and alerting for HPDF systems using modern observability tooling (for example Prometheus, Grafana, OpenTelemetry, ELK), within guidelines set with the lead site reliability engineer and senior staff.
- Develop automation and tooling in Python, Go, or shell that eliminates manual operations and reduces operational risk, following standard software development practices including version control, code review, and testing.
- Respond to incidents and author the runbooks, postmortems, and operational documentation that convert each incident into a lasting improvement to facility reliability.
- Implement and report on Service Level Objectives (SLOs) and Service Level Indicators (SLIs) defined with the architecture team and scientific stakeholders.
- Support expert scientific users by helping research physicists and computational scientists understand system behavior and by translating their needs into reliability requirements.
- Collaborate with colleagues at both Jefferson Lab and Berkeley Lab, sharing tooling, reviews, and operational practice across the HPDF partnership.
- Conduct testing and performance analysis to validate reliability and resilience decisions and to identify bottlenecks.
Additional Responsibilities- Contribute to evaluations of vendor and open source technologies against reliability, performance, and security requirements.
- Participate in an on-call rotation as the facility moves toward operations.
Experience- Required: 3 or more years related experience in Site Reliability Engineering, DevOps, systems engineering, or software engineering with operational responsibility.
- Preferred: 3 or more years experience supporting scientific computing, HPC, or research environments; experience operating systems in support of users outside the engineer's own team
- Preferred: Experience with containers and Kubernetes
- Preferred: Practical experience applying AI assisted or autonomous automation to operational tasks, and routine use of AI tools in day to day engineering work.
- Preferred: Experience with configuration management and infrastructure as code tools (for example Ansible, Terraform, Puppet).
Education- Required: Bachelor's Degree Computer Science or Related Field
- Preferred: Master's Degree Computer Science or Related Field
Experience and Education ExchangeEducation above the minimum may be substituted for experience. Relevant experience may not be substituted for education.
Knowledge, Skills, and Abilities- Solid Linux systems skills and command line fluency, with the ability to troubleshoot across the application, operating system, and network layers.
- Scripting and automation ability in Python, Go, or shell, including familiarity with standard software development practices.
- Working knowledge of monitoring and observability tools (for example Prometheus, Grafana, ELK, OpenTelemetry) and strong motivation to deepen that expertise in a new facility environment.
- Clear written and verbal communication, and the ability to work productively with a community of scientific users and with colleagues across both laboratories.
- A self starter's approach to learning new technologies and to identifying and solving problems within an assigned scope.
- Familiarity with public cloud environments (AWS, Azure, GCP).
- Networking fundamentals (IPv4/IPv6, DNS, firewalls, access control lists) and security conscious operational habits.
- Exposure to HPC or scientific computing environments.