Requisition Id 16943
Overview: We are seeking an Infrastructure Services Engineer who will focus on specializing in monitoring and observability. This position resides in the Infrastructure Operations Center (IOC) in the Digital Services Infrastructure & Operations division of the Information Technology Services Directorate, at Oak Ridge National Laboratory (ORNL).
As part of our team, you will design, operate, and continuously improve monitoring solutions across on-premises, cloud, and containerized environments. The IOC provides 24/7/365 monitoring and operational support for ORNL's enterprise infrastructure and business-essential systems and services.
Major Duties/Responsibilities: - Design, implement, administer, and maintain enterprise monitoring and observability solutions across on-premises, cloud, and containerized environments.
- Develop and optimize alerts, dashboards, reports, synthetic monitors, and telemetry pipelines to identify degradation early and accelerate incident triage and root-cause analysis.
- Evaluate monitoring coverage, gaps, overlaps, and underused capabilities, and implement tools, integrations, and data sources that improve operational visibility.
- Automate monitoring deployment, configuration, data collection, and remediation using PowerShell, Python, or other scripting and automation tools.
- Evaluate and apply AI-assisted capabilities for anomaly detection, predictive analytics, and operational efficiency.
- Collaborate with infrastructure, network, application, security, and other technical teams to improve system health and observability.
- Support incident and problem management by providing relevant metrics, logs, performance trends, and historical analysis.
- Work with vendors and internal subject matter experts to troubleshoot monitoring agents, collectors, integrations, and platform components.
- Establish and maintain monitoring standards, topology diagrams, technical documentation, runbooks, and team procedures.
- Support patching, backup, upgrade, and lifecycle activities for monitoring platforms and related infrastructure components.
- Continuously improve alert thresholds, dashboards, data quality, automated remediations, and monitoring workflows to reduce noise and repetitive operational work.
- Provide escalated support for monitoring-related issues and participate in an on-call or planned maintenance rotation as required.
- Deliver ORNL's mission by aligning behaviors, priorities, and interactions with our core values of Impact, Integrity, Teamwork, Safety, and Service. Promote equal opportunity by fostering a respectful workplace - in how we treat one another, work together, and measure success.
Basic Qualifications:- BS degree in information technology or a related technical field and 2 years of relevant experience.
- Experience operating, administering, or engineering enterprise monitoring platforms for infrastructure, applications, networks, or cloud environments.
- Experience supporting enterprise Windows and Linux server environments, including performance analysis and troubleshooting.
- Experience developing automated solutions using PowerShell, Python, or similar scripting tools.
- Working knowledge of cloud infrastructure, container platforms, orchestration technologies, virtualization, and virtual-machine lifecycle operations.
- Understanding of networking fundamentals, system performance indicators, telemetry, and diagnostic methodologies.
Preferred Qualifications:- Strong analytical and problem-solving skills, including the ability to use operational data to identify issues and recommend improvements.
- Strong written and verbal communication, customer service, collaboration, and technical documentation skills.
- Ability to prioritize responsibilities and balance project work, operational support, and incident response in a fast-paced environment.
- Demonstrated experience automating repetitive work or improving technical and operational processes.
- Experience engineering and administering one or more enterprise-scale monitoring platforms, such as Prometheus, Grafana, Elastic, SolarWinds, or Dynatrace.
- Experience with observability concepts and technologies, including metrics, logs, traces, baselining, synthetic monitoring, and service-level objectives.
- Experience with anomaly detection, predictive analytics, AIOps, or automated remediation.
- Knowledge of automation and infrastructure-as-code frameworks, such as Ansible, Terraform, or Azure Automation.
- Experience using version-control systems to maintain scripts, configurations, dashboards, or infrastructure code.
- Experience with virtualized or clustered compute environments, including performance tuning and lifecycle automation.
- Familiarity with enterprise storage technologies, including direct-attached, SAN, and object storage, and their monitoring requirements.
- Knowledge of enterprise server, storage, network hardware, and platform-level instrumentation.
- Experience with enterprise backup, patching, configuration, or lifecycle management practices.
- Understanding of change management and controlled operational workflows.
- Experience working in regulated, scientific, government, or similarly complex technical environments.
- Motivated self-starter with the ability to work independently and participate creatively in collaborative teams across the laboratory.
Special Requirements:- Visa sponsorship: Visa sponsorship is not available for this position.
- Security, Credentialing, and Eligibility Requirements: Q Clearance: This position requires the ability to obtain and maintain a clearance from the Department of Energy. As such, this position is a Workplace Substance Abuse (WSAP) testing designated position. WSAP positions require passing a pre-placement drug test and participation in an ongoing random drug testing program.
- For Hybrid eligible positions: In addition, we offer a flexible work environment that supports both the organization and the employee. A hybrid/onsite working arrangement may be available with this position.
This position will remain open for a minimum of 5 days after which it will close when a qualified candidate is identified and/or hired.
We accept Word (.doc, .docx), Adobe (unsecured .pdf), Rich Text Format (.rtf), and HTML (.htm, .html) up to 5MB in size. Resumes from third party vendors will not be accepted; these resumes will be deleted and the candidates submitted will not be considered for employment.
If you have trouble applying for a position, please email
[email protected].