Oak Ridge National Laboratory

Infrastructure Services Engineer (Hybrid Eligible)

Oak Ridge National Laboratory$80K — $95K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • BS degree in information technology or a related field with 2 years relevant experience.
  • Experience with enterprise monitoring platforms across various environments.
  • Experience supporting enterprise Windows and Linux server environments.
  • Skilled in automation using PowerShell, Python, or similar tools.
  • Familiarity with cloud infrastructure, container technologies, and virtualization.

Responsibilities

  • Design and maintain monitoring and observability solutions across diverse environments.
  • Develop and optimize alerts, dashboards, and telemetry pipelines for incident response.
  • Evaluate and enhance monitoring coverage and operational visibility tools.
  • Automate deployment and monitoring processes using scripting languages.
  • Collaborate with technical teams to improve observability and system health.
  • Provide metrics and insights for incident management and operational analysis.
  • Support lifecycle activities for monitoring infrastructure, including upgrades and backups.

Benefits

  • Flexible work environment with potential for hybrid/onsite arrangements.
  • Opportunity to work in a respected national laboratory focused on infrastructure and technology innovation.
  • Access to advanced monitoring tools and technologies.
  • Engagement in a mission-driven organization aligned with core values of impact and teamwork.
Full Job Description
Requisition Id 16943

Overview:

We are seeking an Infrastructure Services Engineer who will focus on specializing in monitoring and observability. This position resides in the Infrastructure Operations Center (IOC) in the Digital Services Infrastructure & Operations division of the Information Technology Services Directorate, at Oak Ridge National Laboratory (ORNL).

As part of our team, you will design, operate, and continuously improve monitoring solutions across on-premises, cloud, and containerized environments. The IOC provides 24/7/365 monitoring and operational support for ORNL's enterprise infrastructure and business-essential systems and services.

Major Duties/Responsibilities:
  • Design, implement, administer, and maintain enterprise monitoring and observability solutions across on-premises, cloud, and containerized environments.
  • Develop and optimize alerts, dashboards, reports, synthetic monitors, and telemetry pipelines to identify degradation early and accelerate incident triage and root-cause analysis.
  • Evaluate monitoring coverage, gaps, overlaps, and underused capabilities, and implement tools, integrations, and data sources that improve operational visibility.
  • Automate monitoring deployment, configuration, data collection, and remediation using PowerShell, Python, or other scripting and automation tools.
  • Evaluate and apply AI-assisted capabilities for anomaly detection, predictive analytics, and operational efficiency.
  • Collaborate with infrastructure, network, application, security, and other technical teams to improve system health and observability.
  • Support incident and problem management by providing relevant metrics, logs, performance trends, and historical analysis.
  • Work with vendors and internal subject matter experts to troubleshoot monitoring agents, collectors, integrations, and platform components.
  • Establish and maintain monitoring standards, topology diagrams, technical documentation, runbooks, and team procedures.
  • Support patching, backup, upgrade, and lifecycle activities for monitoring platforms and related infrastructure components.
  • Continuously improve alert thresholds, dashboards, data quality, automated remediations, and monitoring workflows to reduce noise and repetitive operational work.
  • Provide escalated support for monitoring-related issues and participate in an on-call or planned maintenance rotation as required.
  • Deliver ORNL's mission by aligning behaviors, priorities, and interactions with our core values of Impact, Integrity, Teamwork, Safety, and Service. Promote equal opportunity by fostering a respectful workplace - in how we treat one another, work together, and measure success.


Basic Qualifications:
  • BS degree in information technology or a related technical field and 2 years of relevant experience.
  • Experience operating, administering, or engineering enterprise monitoring platforms for infrastructure, applications, networks, or cloud environments.
  • Experience supporting enterprise Windows and Linux server environments, including performance analysis and troubleshooting.
  • Experience developing automated solutions using PowerShell, Python, or similar scripting tools.
  • Working knowledge of cloud infrastructure, container platforms, orchestration technologies, virtualization, and virtual-machine lifecycle operations.
  • Understanding of networking fundamentals, system performance indicators, telemetry, and diagnostic methodologies.


Preferred Qualifications:
  • Strong analytical and problem-solving skills, including the ability to use operational data to identify issues and recommend improvements.
  • Strong written and verbal communication, customer service, collaboration, and technical documentation skills.
  • Ability to prioritize responsibilities and balance project work, operational support, and incident response in a fast-paced environment.
  • Demonstrated experience automating repetitive work or improving technical and operational processes.
  • Experience engineering and administering one or more enterprise-scale monitoring platforms, such as Prometheus, Grafana, Elastic, SolarWinds, or Dynatrace.
  • Experience with observability concepts and technologies, including metrics, logs, traces, baselining, synthetic monitoring, and service-level objectives.
  • Experience with anomaly detection, predictive analytics, AIOps, or automated remediation.
  • Knowledge of automation and infrastructure-as-code frameworks, such as Ansible, Terraform, or Azure Automation.
  • Experience using version-control systems to maintain scripts, configurations, dashboards, or infrastructure code.
  • Experience with virtualized or clustered compute environments, including performance tuning and lifecycle automation.
  • Familiarity with enterprise storage technologies, including direct-attached, SAN, and object storage, and their monitoring requirements.
  • Knowledge of enterprise server, storage, network hardware, and platform-level instrumentation.
  • Experience with enterprise backup, patching, configuration, or lifecycle management practices.
  • Understanding of change management and controlled operational workflows.
  • Experience working in regulated, scientific, government, or similarly complex technical environments.
  • Motivated self-starter with the ability to work independently and participate creatively in collaborative teams across the laboratory.


Special Requirements:
  • Visa sponsorship: Visa sponsorship is not available for this position.
  • Security, Credentialing, and Eligibility Requirements: Q Clearance: This position requires the ability to obtain and maintain a clearance from the Department of Energy. As such, this position is a Workplace Substance Abuse (WSAP) testing designated position. WSAP positions require passing a pre-placement drug test and participation in an ongoing random drug testing program.
  • For Hybrid eligible positions: In addition, we offer a flexible work environment that supports both the organization and the employee. A hybrid/onsite working arrangement may be available with this position.


This position will remain open for a minimum of 5 days after which it will close when a qualified candidate is identified and/or hired.

We accept Word (.doc, .docx), Adobe (unsecured .pdf), Rich Text Format (.rtf), and HTML (.htm, .html) up to 5MB in size. Resumes from third party vendors will not be accepted; these resumes will be deleted and the candidates submitted will not be considered for employment.

If you have trouble applying for a position, please email [email protected].

About Oak Ridge National Laboratory

Oak Ridge National Laboratory (ORNL) is a science and technology national laboratory managed for the United States Department of Energy (DOE) by UT-Battelle. ORNL is the largest science and energy national laboratory in the Department of Energy system by size and by annual budget. ORNL conducts research and development activities in a variety of scientific and technical disciplines. ORNL's scientific programs focus on materials, neutron science, energy, high-performance computing, systems biology and national security. ORNL partners with other national laboratories, universities and industry to solve complex problems and transfer knowledge and technology. ORNL is home to several of the world's most powerful supercomputers, including Summit, the world's most powerful supercomputer as of November 2018.
Learn more about Oak Ridge National Laboratory
Size
5,000 employees
Industry
Founded
1943

Similar Jobs

More Jobs at Oak Ridge National Laboratory

More Information Technology Jobs

Find similar Infrastructure Services Engineer (Hybrid Eligible) jobs: