Enterprise Site Reliability Engineer

Compunnel

• $125K — $150K *
Enterprise Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • 8+ years in Site Reliability Engineering, Production Engineering, Platform Engineering, or software engineering for large-scale systems.
  • Deep understanding of SRE principles and cloud-native architecture.
  • Extensive hands-on experience with AWS services and large-scale AWS environments.
  • Strong expertise in Kubernetes and Red Hat OpenShift Service on AWS.
  • Proficiency in OpenTelemetry and enterprise observability tools.
  • Strong coding skills in Python, Java, Go, or JavaScript/TypeScript.
  • Experience leading complex incidents and conducting root cause analysis.

Responsibilities

  • Define enterprise SRE standards for SLIs, SLOs, and incident management.
  • Establish observability standards using OpenTelemetry and metrics.
  • Conduct maturity assessments on application observability and resilience.
  • Analyze production systems to identify operational risks and performance issues.
  • Translate technical findings into actionable business impact insights.
  • Develop reliability improvement roadmaps with measurable goals.
  • Mentor engineering teams and facilitate technical decisions.

Benefits

  • Opportunity to lead reliability improvements across enterprise environments.
  • Collaborative work across diverse technical teams and domains.
  • Exposure to cutting-edge technologies in cloud and observability.
  • Opportunities for professional mentoring and influence on technical practices.
  • Chance to define and implement SRE standards within the organization.
Full Job Description
div:has([data-free-thinking-preview-answer=true])+:is(.text-message,.relative:has(>.text-message))]:-mt-2 grow">

Job Summary:
The Enterprise Site Reliability Engineer will advance reliability, resilience, observability, automation, and operational excellence across large-scale enterprise environments. The role will work across SRE, application, architecture, cloud, platform, infrastructure, security, and data teams to identify operational risks, assess application maturity, and develop scalable reliability solutions. The engineer will combine hands-on software development and cloud engineering with architecture reviews, complex incident leadership, observability strategy, and executive-level communication.

Key Responsibilities:
• Define and advance enterprise SRE standards for SLIs, SLOs, error budgets, production readiness, incident management, and operational excellence.
• Establish observability standards covering OpenTelemetry, distributed traces, logs, metrics, events, telemetry correlation, tagging, data quality, and service ownership.
• Conduct application maturity assessments across observability, reliability, resilience, operability, automation, and incident readiness.
• Analyze application architectures and production operations to identify dependencies, failure modes, capacity constraints, performance bottlenecks, and operational risks.
• Translate technical and operational findings into customer impact, business risk, and investment priorities.
• Develop prioritized and measurable reliability improvement roadmaps.
• Recommend and implement reliability patterns including fault isolation, graceful degradation, circuit breakers, retries, rate limiting, load shedding, high availability, and disaster recovery.
• Design end-to-end observability across distributed services, APIs, business transactions, customer journeys, cloud platforms, and cross-domain dependencies.
• Develop production-grade automation, applications, APIs, and integrations that reduce operational toil and improve reliability at scale.
• Automate service onboarding, telemetry validation, SLO reporting, production readiness checks, incident enrichment, and remediation workflows.
• Integrate observability, cloud, CI/CD, IT service management, incident management, and configuration platforms using APIs, SDKs, webhooks, and event-driven patterns.
• Lead complex incident investigations and post-incident reviews, identify contributing factors, and drive corrective actions through completion.
• Analyze operational and observability data to identify patterns, quantify risks, and uncover systemic reliability gaps.
• Proactively identify opportunities to improve application operability, resilience, telemetry, support readiness, and operational processes.
• Define observability requirements for AI agents and AI-enabled applications, including workflows, model and tool interactions, dependencies, latency, failures, quality, token consumption, cost, and reliability.
• Use AI-assisted engineering tools and AI agents to accelerate development, incident analysis, correlation, pattern detection, and operational insights.
• Develop reusable reference architectures, engineering patterns, assessment frameworks, maturity models, scorecards, and implementation guidance.
• Collaborate with SREs and engineering teams to resolve cross-domain reliability issues and promote consistent engineering practices.
• Mentor engineers, facilitate technical decisions, challenge existing approaches constructively, and build alignment across teams.
• Communicate technical risks, business impact, recommendations, and progress to engineering teams and senior leadership.

Required Qualifications:
• 8+ years of experience in Site Reliability Engineering, Production Engineering, Platform Engineering, or software engineering for large-scale production systems.
• Deep knowledge of SRE principles, operational excellence, distributed systems, microservices, APIs, and cloud-native architecture.
• Strong hands-on experience architecting, operating, and troubleshooting large-scale AWS environments across compute, containers, serverless, networking, databases, storage, identity, and cloud observability.
• Experience with AWS services including EC2, EKS, Lambda, VPC, Elastic Load Balancing, Route 53, RDS/Aurora, DynamoDB, S3, IAM, and CloudWatch.
• Strong hands-on experience with Kubernetes and Red Hat OpenShift Service on AWS (ROSA), including architecture, operations, performance, and troubleshooting.
• Strong experience with OpenTelemetry, distributed tracing, logs, metrics, events, and Dynatrace or a comparable enterprise observability platform.
• Experience defining and operationalizing SLIs, SLOs, error budgets, production readiness criteria, and reliability scorecards.
• Experience assessing application architecture and operational maturity from observability, reliability, resilience, and operability perspectives.
• Strong software development skills in Python, Java, Go, JavaScript/TypeScript, or similar programming languages.
• Experience developing production-grade APIs, integrations, automation services, and internal engineering tools.
• Proficiency with REST APIs, SDKs, Git, automated testing, CI/CD, Infrastructure-as-Code, secure development, and software lifecycle practices.
• Experience leading complex incidents, technical investigations, root cause analysis, post-incident reviews, and corrective-action programs.
• Strong analytical and investigative skills with the ability to identify systemic issues beyond immediate symptoms.
• Ability to convert technical findings and operational data into clear, actionable insights and enterprise recommendations.
• Ability to connect reliability risks and technical decisions to customer experience, business impact, and organizational priorities.
• Excellent written, verbal, technical, and executive communication skills.
• Proven ability to mentor experienced engineers, facilitate technical decisions, and influence outcomes without direct authority.
• Ability to collaborate effectively with SREs, architects, application teams, and specialists across multiple IT domains.

Preferred Qualifications:
• Experience integrating enterprise platforms such as Dynatrace, AWS, ServiceNow, GitLab, Kubernetes, and ROSA.
• Experience developing self-service reliability capabilities, internal developer platforms, or enterprise engineering products.
• Experience with performance engineering, capacity planning, resilience testing, and chaos engineering.
• Experience establishing observability and reliability controls for AI agents, LLM-enabled applications, or AI-driven workflows.
• Experience working with large-scale, highly available, business-critical enterprise systems.

Certifications:
• AWS certification, preferably AWS Certified Solutions Architect Professional, AWS Certified DevOps Engineer Professional, or a relevant specialty certification.
• Kubernetes or OpenShift certification, such as CKA, CKS, or Red Hat Certified OpenShift Administrator, is preferred.

Similar Jobs

More Jobs at Compunnel

More Enterprise Technology Jobs

Find similar Enterprise Site Reliability Engineer jobs: