Observability Engineer I, II, III, or Senior

Tri-State Generation and Transmission Association, Inc.

$109K — $139K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in computer science, IT, engineering, or related field, or equivalent experience.
  • Expert-level experience with Splunk and telemetry architecture.
  • Strong knowledge of automation, operational analytics, and monitoring systems administration.
  • Deep understanding of observability principles and service reliability.
  • Experience with diverse data sources and integration into observability workflows.

Responsibilities

  • Set technical direction and establish standards for observability platforms.
  • Lead design and lifecycle of monitoring and observability technologies.
  • Define and govern data onboarding and retention standards for operational effectiveness.
  • Design advanced dashboards and actionable operational analytics.
  • Lead troubleshooting and root-cause investigations for complex observability challenges.

Benefits

  • Medical, Dental, and Vision Insurance
  • Health Savings Account (HSA) and Flexible Spending Accounts (FSA)
  • Tuition Reimbursement Programs
  • Flexible work schedules including remote work options
  • Comprehensive life and disability insurance programs
Full Job Description
Job Description

The Senior Observability Engineer is a technical authority responsible for setting technical direction and defining standards for enterprise observability. This role requires deep knowledge and expert-level experience with Splunk, telemetry architecture, monitoring strategy, monitoring systems administration, automation, and operational analytics. The position leads the design and lifecycle of observability platforms, serves as the senior escalation point for complex issues, mentors engineers, advises leadership, and drives automation and continuous improvement in service visibility, reliability, performance, data quality, and ingest efficiency.

Note:
There is one position available and the position will be filled at one of four job grade levels: Observability Engineer I, job grade 7; Observability Engineer II, job grade 8, Observability Engineer III, job grade 9; or Senior Observability Engineer, job grade 10. This decision will be based on the qualifications and experience of the candidate selected, and Tri-State business needs at the time of hire.

Tri-State recognizes the value of a highly-engaged and committed workforce and provides an excellent benefits program that includes:
Medical Insurance, Dental Insurance, Vision Insurance, Health Savings Account (HSA), Flexible Spending Accounts (FSA), Tuition Reimbursement, Flexible Work Schedules including compressed work week and telecommuting opportunities to work remotely up to 40%, Life Insurance, 401K, Long Term Disability (LTD), Short Term Disability (STD), Employee Assistant Program (EAP) and Paid Leave Benefits.

Senior Observability Engineer
Hiring Salary Range: $109,000-139,000
Observability Engineer III
Hiring Salary Range: $98,000-$124,000
Observability Engineer II
Hiring Salary Range: $88,000-$111,000
Observability Engineer I
Hiring Salary Range: $80,000-$99,000

Actual compensation offer to candidate may vary outside of the posted hiring salary range based upon work experience, education, and/or skill level.

Responsibilities
• Set technical direction and define enterprise standards for the architecture, configuration, security, scalability, availability, and lifecycle of monitoring and observability platforms.
• Serve as the technical authority for Splunk Enterprise, Splunk Cloud, and related observability capabilities, including architecture, data ingestion, indexing, search, platform health, integrations, access, and operational support.
• Design enterprise telemetry strategies spanning logs, metrics, traces, events, service health, application performance, infrastructure performance, networks, cloud services, and business-critical systems.
• Define data-onboarding, parsing, filtering, routing, normalization, retention, quality, lineage, and lifecycle standards that balance operational value, compliance, performance, and cost.
• Lead the design of advanced dashboards, correlation searches, service-health models, executive reporting, alerting strategies, and actionable operational analytics.
• Establish and govern CIM-aligned data models, field standards, naming conventions, reusable knowledge objects, and validation controls across diverse data sources.
• Set strategy for alert quality, event correlation, noise reduction, anomaly detection, and integration with incident and service management processes.
• Lead the most complex troubleshooting and root-cause investigations involving ingestion, duplication, search performance, monitoring-system availability, integrations, data quality, and distributed-system behavior.
• Define automation standards and lead the design of reliable, supportable automation using Python, Bash, PowerShell, Ansible, APIs, version control, testing, and orchestration tooling to eliminate manual work across monitoring operations.
• Evaluate and recommend new observability technologies, products, and architectural approaches; lead technical proofs of concept and enterprise adoption decisions.
• Define monitoring system capacity, performance, resilience, disaster recovery, upgrade, and lifecycle plans; identify and mitigate technical and operational risks.
• Partner with leaders and technical teams across the enterprise to establish monitoring outcomes, service-level indicators, service-level objectives, ownership models, and operational accountability.
• Establish engineering practices, design review standards, testing requirements, documentation expectations, and continuous improvement measures.
• Serve as the senior escalation point for observability engineering challenges and major operational events; provide clear technical guidance during high-impact incidents.
• Mentor, coach, and develop Observability Engineers at all levels; lead knowledge sharing and build deep team capability.
• Lead cross-functional and enterprise-wide initiatives from strategy through implementation, operational transition, measurement, and improvement.
• Advise on compliance and audit strategy, including NERC CIP, and ensure observability systems and processes meet applicable company, legal, regulatory, and security requirements.
• Maintain compliance with all company policies and procedures and remain knowledgeable of regulations, laws, standards, and best practices applicable to the functional area.
• Because Tri-State has an obligation to provide continuous, reliable electric service to its customers, the ability to work overtime at any time of the day or week is considered an essential function of the job.

OTHER DUTIES/RESPONSIBILITIES
• Perform other related duties as assigned.

SUCCESS FACTORS/JOB COMPETENCIES:
• Deep and expert understanding of observability architecture, telemetry pipelines, distributed systems, service reliability, performance engineering, event correlation, and operational analytics.
• Expert-level Splunk knowledge, including search, distributed deployment concepts, data onboarding, indexing, parsing, knowledge management, data models, dashboards, alerting, security, performance, and lifecycle planning; deep knowledge of integrating SolarWinds, Snare, Sunbird DCIM, or comparable systems into enterprise observability workflows.
• Deep expertise in ingest optimization, licensing, retention, data tiering, duplicate-data reduction, monitoring-system capacity, resilience, and cost management.
• Deep technical knowledge of Windows, Linux, networking, virtualization, cloud platforms, containers, APIs, authentication, application architectures, and enterprise integrations.
• Expert automation and engineering capability using Python, Bash, PowerShell, Ansible, APIs, infrastructure as code, version control, and delivery pipelines.
• Demonstrated ability to set technical direction, lead enterprise initiatives, make architecture decisions, mentor engineers, and influence stakeholders without direct authority.
• Excellent written and verbal communication skills, including translating complex technical matters into clear business risks, options, and recommendations for leadership.
• Excellent written and verbal communication skills, with the ability to explain technical information in clear business terms.
• Strong analytical, evaluative, troubleshooting, and problem-solving abilities.
• Keen attention to detail and commitment to accurate, supportable, and well-documented engineering work.

Qualifications

Education and Training
• Bachelor's degree in computer science, information technology, engineering, information systems, business, or related field, or equivalent experience gained through progressively greater responsibilities.

Knowledge, Skills, and Ability:
• Deep and expert understanding of observability architecture, telemetry pipelines, distributed systems, service reliability, performance engineering, event correlation, and operational analytics.
• Expert-level Splunk knowledge, including search, distributed deployment concepts, data onboarding, indexing, parsing, knowledge management, data models, dashboards, alerting, security, performance, and lifecycle planning; deep knowledge of integrating SolarWinds, Snare, Sunbird DCIM, or comparable systems into enterprise observability workflows.
• Deep expertise in ingest optimization, licensing, retention, data tiering, duplicate-data reduction, monitoring-system capacity, resilience, and cost management.
• Deep technical knowledge of Windows, Linux, networking, virtualization, cloud platforms, containers, APIs, authentication, application architectures, and enterprise integrations.
• Expert automation and engineering capability using Python, Bash, PowerShell, Ansible, APIs, infrastructure as code, version control, and delivery pipelines.
• Demonstrated ability to set technical direction, lead enterprise initiatives, make architecture decisions, mentor engineers, and influence stakeholders without direct authority.
• Excellent written and verbal communication skills, including translating complex technical matters into clear business risks, options, and recommendations for leadership.
• Excellent written and verbal communication skills, with the ability to explain technical information in clear business terms.
• Strong analytical, evaluative, troubleshooting, and problem-solving abilities.
• Keen attention to detail and commitment to accurate, supportable, and well-documented engineering work.

Other:
• Willingness to travel as required up to 10% of the time including overnight.
• Willingness to participate in on-call support or planned maintenance activities up to 25% of the time.

DESIRED JOB QUALIFICATIONS
• An MS, MBA, or related advanced degree is desired.
• Splunk certifications, such as Splunk Core Certified User, Splunk Core Certified Power User, or Splunk Enterprise Certified Admin, are beneficial but not required.
• Experience appropriate to the position level with Splunk Cloud, Splunk Enterprise, SolarWinds, Snare, Sunbird DCIM, Splunk IT Service Intelligence, Splunk Enterprise Security, Cribl, or comparable monitoring and data-management systems.
• Experience with metrics, traces, application performance monitoring, network monitoring, digital experience monitoring, or service-health modeling.
• Experience using Python, Bash, PowerShell, Ansible, APIs, version control, infrastructure as code, or orchestration tooling.
• Experience working within regulated environments or regulatory compliance frameworks, such as NERC CIP, SOX, or other applicable industry, legal, security, privacy, or operational requirements.
• ITIL, cloud, Linux, Windows, networking, Kubernetes, or other relevant technical certifications.

Those with less experience will be hired at the Observability Engineer I, II, or III job grade level, as appropriate.

Similar Jobs

More Jobs at Tri-State Generation and Transmission Association, Inc.

More Information Technology Jobs

Find similar Observability Engineer I, II, III, or Senior jobs: