Site Reliability Engineer

LBL$110K — $130K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Experience in a 24/7 onsite team environment supporting large-scale data centers.
  • Proficiency in Linux shell and command-line environments (e.g. SSH).
  • Experience developing tools in C, C++, Perl, Java, Python, or scripting languages.
  • Familiarity with network security protocols and firewalls.
  • Experience collaborating across technical teams to enhance system reliability.

Responsibilities

  • Monitor NERSC HPC Facility during designated Owl shifts.
  • Respond to alerts from various systems and triage issues accordingly.
  • Create solutions to improve operational processes and automate routine responses.
  • Enhance monitoring capabilities and propose automation for problem triage.
  • Develop and maintain monitoring tools in cooperation with the Operations Team.
  • Create alert systems that integrate HPC system APIs into monitoring pipelines.
  • Conduct regular walkthroughs of data center to ensure operational health and efficiency.

Benefits

  • Opportunities for professional development and ongoing education.
  • Access to advanced technology and high-performance computing resources.
  • Collaboration with cutting-edge research and technical teams.
  • Work in a mission-driven environment focused on scientific discovery.
Full Job Description
The National Energy Research Scientific Computing Center (NERSC) is inviting applications for the position of Site Reliability Engineer. NERSC's mission is to accelerate scientific discovery through high performance computing and data analysis for the DOE Office of Science programs. NERSC provides critical HPC and data systems and support for NERSC's 11,000 plus users researching energy, physics, materials science, and chemistry and other DOE mission areas. As a Site Reliability Engineer in the Operations Technology Group, you will be a member of a 24x7 team that helps ensure NERSC is accessible, reliable, and secure for our scientific users. This position leverages advanced data collection and monitoring systems to proactively manage the health of our environment. Ultimately, your work ensures that NERSC's computational power remains an uninterrupted catalyst supporting fundamental scientific research related to energy.

ESSENTIAL DUTIES & RESPONSIBILITIES

Describe the duties/functions essential to performing the job.

  • Works an onsite 5-day weekly schedule consisting of Owl (midnight-8 am) shifts to monitor the NERSC HPC Facility.
  • Review and respond to alerts from computer systems, storage, network, and other data center/facility-related systems by triaging or calling the appropriate on-call staff.
  • Create appropriate solutions to improve processes, prevent issue recurrence, and automate responses to all routine service conditions.
  • Identify issues and propose solutions that will improve monitoring capabilities or provide better automation for triage.
  • Possess expertise in ServiceNow and its usage to develop and implement customized service management solutions.
  • Respond to alerts from multiple systems to ensure that data collection continues 24/7, providing real-time information for diagnoses.
  • Develop and maintain tools within the monitoring pipeline in collaboration with the Operations Team.
  • Create new software to provide alerts and notifications from HPC system APIs into the monitoring pipeline.
  • Builds and maintains application/tool configurations to ensure software runs reliably as data and user demands grow.
  • Collaborate with other groups at NERSC to ensure that communication and workflows are clearly understood.
  • Work closely with other NERSC groups to coordinate center-wide maintenance activities and manage diagnostic and notification software during maintenance periods.
  • Perform regular physical and logical walkthroughs of the data center floor to monitor environmental health, power distribution units, and cooling infrastructure to ensure peak operational efficiency.
  • Provide accurate information in the trouble ticketing system for outages, maintenance updates, and other incidents so that workflows and protocols can be appropriately tracked by others.
  • Work on and resolve problems of diverse scope where data analysis requires the evaluation of identifiable factors.
  • Demonstrate good judgment in selecting methods and techniques for obtaining solutions.
  • Work on and resolve complex issues where the analysis of situations or data requires an in-depth evaluation of variable factors.


POSITION REQUIREMENTS

  • Experience in or willingness to work within a 24/7 onsite team environment to support large-scale data centers or critical installations.
  • Experience on Linux shell and working in a command-line (e.g. SSH) environment.
  • Experience with developing tools using various programming languages such as C, C++, Perl, Java, or Python or a scripting language with knowledge of standard software development practices.
  • Motivated, self-starter who can learn technologies that improve data center management in areas like Kubernetes, Prometheus/VictoriaMetrics, Alertmanager, building management software, evaporative


cooling, and power utilization.

  • Experience with network security: configuring/maintaining ACLs, knowledge of firewalls
  • Experience collaborating across technical teams to resolve operational bottlenecks and ensure system reliability and alignment with service-level objectives.
  • Good to Have : Practical experience in developing and deploying Agentic AI or autonomous automation tools to streamline technical tasks.
  • Experience with ServiceNow implementation is a plus
  • Familiarity with ITSM best practices and an understanding of how to align service lifecycles with business goals is preferred.


Knowledge, Skills & Abilities

  • Strong hands-on knowledge of the Linux shell and working in a command-line (e.g. SSH) environment.
  • Strong hands-on knowledge on various programming languages such as C, C++, Perl, Java, or Python or a scripting language with knowledge of standard software development practices.
  • Knowledge of and ability to work on large data communications networks/ Network Protocols and IT infrastructure supporting highly available systems and applications.
  • Strong communication skills and ability to work effectively across multiple technical teams.
  • Good to Have : Ability to build, and deploy Agentic AI solutions that utilize autonomous agents to automate decision-making, optimize complex workflows, and enhance proactive system monitoring.

About LBL

LBL Careers

Joining LBL offers an unparalleled opportunity to become part of a leading team of professionals dedicated to pioneering innovation and digital transformation. LBL stands as a beacon of excellence, offering a range of job opportunities that cater to various skills and career aspirations.

Explore Career Opportunities

LBL’s dynamic career paths empower professionals to navigate their professional growth with confidence. Whether through full-time positions, internships, or leadership roles, LBL is committed to fostering a culture of growth and learning.

Innovation and Professional Growth

At LBL, innovation isn’t just a buzzword; it's the cornerstone of their mission. The company encourages its team to push the boundaries of technology and strategy, ensuring that every member has the opportunity to contribute to groundbreaking projects.

Diversity and Inclusion

Diversity training and inclusion are at the heart of LBL’s employment strategy. The company believes that a diverse team is a strong team, and actively works to create an environment where all voices are heard and valued.

Benefits and Culture

LBL is renowned for its vibrant culture and comprehensive benefits package designed to support the team in all aspects of life—both professional and personal. From health benefits to flexible work policies, LBL ensures that the team not only excels at work but also enjoys a balanced life.

Networking and Development

Career advancement at LBL is fueled by robust professional networking and development programs. These initiatives are tailored to hone skills, enhance leadership capabilities, and ensure that every team member can achieve their career goals.

Join the LBL Team

LBL is actively hiring and looking for individuals who are passionate, curious, and driven. Explore the open positions that match your skills and interests. Engage with a company that values innovation and offers the tools needed to succeed in a competitive market.

Stay Connected with LBL Jobs

Stay informed about the latest in career opportunities and industry trends by subscribing to LBL job alerts. Tailor your preferences to receive updates that align with your professional interests and career goals.

Prepare for Your Interview

Aspiring candidates can look forward to a transparent interview process that assesses a range of competencies from technical skills to creative thinking. Ensure your resume highlights relevant experiences and skills to stand out in the LBL hiring process.

Career Insights and Tips

Gain insights from industry leaders and get ahead with career tips directly from the professionals at LBL. These resources are invaluable for those looking to make a significant impact in their professional journey.

Explore LBL Careers Today

Discover the exciting and rewarding career opportunities at LBL. Whether you’re seeking an internship or a managerial position, LBL offers a path for everyone. Join a team that’s dedicated to leadership, professional growth, and innovation in the digital era.
Learn more about LBL

Similar Jobs

More Jobs at LBL

More Information Technology Jobs

Find similar Site Reliability Engineer jobs: