World Wide Technology

Observability & SRE Engineer

World Wide Technology$90K — $112K *
US-AnywhereRemote in United States
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years of experience in IT Operations
  • Proficiency in monitoring tools like Prometheus and Grafana
  • Familiarity with programming in Python, Go, or similar
  • Experience with system frameworks such as Git and GitHub
  • Strong skills in critical thinking and communication
  • Team player with a robust work ethic
  • Knowledge of Linux systems and container technologies like Docker
  • Understanding of SRE principles and practices

Responsibilities

  • Collect and apply metrics for decision-making
  • Provide insights into system health with observability techniques
  • Implement Application Performance Monitoring and synthetic transaction monitoring
  • Enhance reliability by leveraging monitoring and SRE practices
  • Utilize AI methodologies to streamline workflows and improve incident response
  • Drive reliability improvements through SLIs, SLOs, and incident reviews
  • Manage on-call rotation and incident response processes
  • Automate operational tasks to minimize toil and integrate infrastructure as code
  • Conduct capacity planning and performance engineering for scalable systems
  • Collaborate with development on production readiness and resilience testing

Benefits

  • Comprehensive health insurance including medical, dental, and vision coverage
  • 401k plan with company matching and profit sharing
  • Generous paid time off including parental and medical leave
  • Employee assistance programs and wellness initiatives
  • Tuition reimbursement and additional financial benefits
  • Family planning and nursing mothers support
  • Flexible spending accounts and voluntary legal services
Full Job Description
Qualifications:

Minimum Qualifications
  • 5+ years of professional experience in IT Operations
  • Experience in Metrics, Monitoring and Alerting, including tools like Prometheus and Grafana, Big Panda, Splunk, etc.
  • Knowledge of programming languages Python, Go, or equivalent
  • Knowledge of system frameworks including Git and GitHub
  • Critical thinker with excellent written and verbal communication skills
  • Team-oriented individual with very strong work ethic
  • Familiarity with Linux, preferably administrative knowledge
  • Understanding of container technologies (Docker, Podman, etc.)
  • Understanding of SRE concepts such as SLIs, SLOs, incident response, root cause analysis, capacity planning, and reliability automation
  • Practical understanding of AI-enabled tools, automation patterns, and responsible AI use to improve operational efficiency
  • Willingness to participate in an on-call rotation and drive incidents through triage, mitigation, and resolution
  • Experience with Infrastructure as Code (Terraform, Ansible, or similar) and CI/CD pipelines
  • Understanding of distributed systems fundamentals: fault tolerance, redundancy, load balancing, and caching

Preferred Qualifications
  • Kubernetes experience
  • Experience working with Agile methodology
  • Test-First development mindset with functional, end-to-end and regression testing experience
  • Experience with AI-assisted engineering, AIOps, or machine learning techniques for monitoring, alerting, or incident response
  • Hands-on SRE experience with SLOs, error budgets, blameless reviews, toil reduction, and production readiness
  • Experience with chaos engineering or resilience/failure-injection testing
  • Experience authoring and maintaining runbooks, playbooks, and production readiness review checklists
  • Public cloud platform experience (AWS, Azure, or GCP)
  • Familiarity with distributed tracing and log aggregation tooling (e.g., OpenTelemetry, ELK/Splunk)

Certain states and localities require employers to post a reasonable estimate of the salary range. A reasonable estimate of the current base pay range for this position is $90,200 to $112,000 annually. Actual salary will be based on a variety of factors, including shift, location, experience, skill set, performance, licensure and certification, and business needs. The range for this position in other geographic locations may differ. Certain positions may also be eligible for variable incentive compensation, such as bonuses or commissions, that are not included in the base pay.

The well-being of WWT employees is essential. When it comes to our benefits package, WWT has one of the best. We offer the following benefits to all full-time employees:
  • Health and Wellbeing: Health (Medical & Prescription), Dental, and Vision Care, Onsite Health Centers (MO & IL), Employee Assistance Program, Wellness program
  • Financial Benefits: Competitive Pay, Profit Sharing, 401k Plan with Company Matching, Life and Disability Insurance, Flexible Spending Accounts, Tuition Reimbursement
  • Paid Time Off: PTO & Holidays, Parental Leave, Medical Leave, Military Leave, Bereavement, Day of Caring
  • Additional Perks: Family Planning Benefits, Nursing Mothers Benefits, Voluntary Legal, Voluntary Supplemental Accident/Illness/Hospital, Voluntary ID Theft, Pet Insurance, Employee Discount Program


Note: This is not an all-encompassing list and should not be used as a complete description of the plan's benefits. For more information, see our US Benefits Website

What will you be doing?

World Wide Technology (WWT) is seeking an Engineer to join the Observability & Site Reliability Engineering (SRE) team. In this role, the Engineer will design, build, and operate the systems that give WWT visibility into the health of its IT services, while also applying core SRE practices to improve their reliability, performance, and resilience.

The ideal candidate is a passionate technologist who enjoys instrumenting, measuring, and improving systems to strengthen observability, reliability, and operational decision-making. They bring an AI-first mindset, using AI, automation, and data-driven practices responsibly to improve reliability, accelerate response, and reduce manual effort.

The Observability & SRE team brings together infrastructure, operations, automation, and reliability engineering skills to build consumable services and platforms that improve visibility, confidence, and resilience across WWT's IT systems. This is an opportunity for someone looking to grow technically while helping modernize IT through AI-enabled operations and reliability-focused engineering.

Responsibilities:
  • Collection and strategic application of metrics to drive organizational decisions
  • Providing a holistic view of system health using observability practices
  • APM, RUM, and Synthetic Transaction monitoring
  • Driving reliability through monitoring, alerting, observability, and SRE practices
  • Applying AI-first thinking to automate workflows, correlate alerts, surface insights, and speed incident response
  • Improving reliability through SLOs, SLIs, error budgets, incident reviews, and continuous improvement
  • Defining and maintaining on-call rotations, escalation paths, and incident response processes, including participating in on-call coverage
  • Reducing operational toil through automation, self-healing systems, and infrastructure as code
  • Capacity planning and performance engineering to ensure systems scale reliably under load
  • Partnering with development teams on production readiness reviews, architecture reviews, and resilience/chaos testing to prevent incidents before they happen

About World Wide Technology

World Wide Technology (WWT) is a technology solution provider that offers a wide range of services to businesses and organizations. The company was founded in 1990 and is headquartered in Maryland Heights, Missouri. WWT provides a variety of services, including consulting, design, integration, and managed services. The company has a strong focus on innovation and has been recognized for its efforts in this area. WWT has partnerships with many leading technology companies, including Cisco, Dell, and Microsoft. The company has a global presence, with offices in the United States, Europe, and Asia.
Learn more about World Wide Technology
Size
7,000 employees
Industry
Founded
1990

Similar Jobs

More Jobs at World Wide Technology

More Information Technology Jobs

Find similar Observability & SRE Engineer jobs: