Site Reliability Engineer

Compunnel

$100K — $130K *
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in Computer Science, Information Technology, Engineering, or related field, or equivalent experience.
  • 10+ years in Site Reliability Engineering, DevOps, or Systems Engineering.
  • Proven experience with highly available, cloud-based, or mission-critical platforms.
  • Expert knowledge in monitoring, observability, and logging strategies.
  • Strong scripting skills in Python, Java, Bash, or PowerShell.
  • Familiarity with Infrastructure as Code tools like Terraform and cloud platforms such as AWS, Azure, or GCP.
  • Solid grasp of SRE concepts: SLIs, SLOs, SLAs, and incident management.

Responsibilities

  • Ensure the reliability and availability of critical cybersecurity platforms and infrastructure.
  • Monitor system health, identify operational risks, and respond to incidents.
  • Instrument infrastructure and applications for comprehensive observability.
  • Design and implement monitoring, alerting, and dashboard solutions for distributed systems.
  • Develop actionable alerts to improve incident response efficiency.
  • Define and track service reliability metrics and operational risk parameters.
  • Collaborate with cross-functional teams to enhance platform reliability.

Benefits

  • Support for continuing education and professional development.
  • Opportunities to work on cutting-edge cybersecurity and cloud technology.
  • Collaborative work culture with cross-functional teams.
  • Flexible working arrangements to promote work-life balance.
  • Access to the latest tools and technologies in the field of site reliability engineering.
Full Job Description
Job Summary

We are seeking a Site Reliability Engineer to support Cyber Data Risk & Resilience by ensuring the reliability, availability, performance, and operational visibility of critical cybersecurity platforms and services. This role is responsible for maintaining production systems, building observability solutions, automating operational processes, supporting incident response, enhancing executive dashboards, and driving continuous improvements across cloud-based and distributed environments.

Key Responsibilities

  • Maintain and improve the reliability, availability, scalability, and performance of cybersecurity platforms and supporting infrastructure.
  • Monitor system health, identify operational risks, respond to incidents, and drive timely resolution of service-impacting issues.
  • Instrument infrastructure, applications, APIs, databases, cloud components, and data pipelines for end-to-end observability.
  • Design, build, and enhance monitoring, alerting, logging, tracing, and dashboard solutions across distributed systems.
  • Develop actionable alerts that reduce noise, improve signal quality, and accelerate incident response.
  • Define and monitor SLIs, SLOs, SLAs, error budgets, latency, throughput, availability, and operational risk metrics.
  • Build, maintain, and improve operational dashboards for engineering, operations, cybersecurity, risk, and executive leadership.
  • Continuously enhance executive dashboards with service health, reliability trends, incident reporting, and operational performance metrics.
  • Partner with engineering, cloud, infrastructure, application, and cybersecurity teams to improve platform reliability.
  • Participate in incident response, root cause analysis, post-incident reviews, and problem management activities.
  • Automate operational tasks, health checks, reporting, deployment validation, and recovery procedures.
  • Support CI/CD pipelines, DevOps processes, release readiness, rollback validation, and production support activities.
  • Contribute to resiliency engineering initiatives including capacity planning, performance tuning, disaster recovery, failover testing, and resilience validation.
  • Ensure monitoring, dashboards, alerting, and operational processes comply with enterprise security, governance, and risk standards.
  • Perform other duties as assigned.


Required Qualifications

  • Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field, or equivalent work experience.
  • 10+ years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, Systems Engineering, Software Engineering, or Production Operations.
  • Experience supporting highly available, distributed, cloud-based, or mission-critical technology platforms.
  • Strong experience with monitoring, observability, logging, tracing, dashboards, and service health reporting.
  • Experience instrumenting applications, infrastructure, cloud services, APIs, databases, and distributed systems.
  • Strong understanding of SRE concepts including SLIs, SLOs, SLAs, error budgets, capacity planning, and incident management.
  • Experience designing meaningful monitoring and actionable alerting strategies.
  • Strong scripting or programming skills using Python, Java, Bash, PowerShell, or similar languages.
  • Experience with cloud platforms including AWS, Azure, or GCP.
  • Experience with Infrastructure as Code tools such as Terraform.
  • Experience supporting CI/CD pipelines, DevOps workflows, release management, and production environments.
  • Experience troubleshooting distributed systems, REST APIs, event-driven architectures, messaging platforms, and service integrations.
  • Familiarity with relational and NoSQL databases such as PostgreSQL, Microsoft SQL Server, MongoDB, or similar technologies.
  • Strong analytical, troubleshooting, problem-solving, and communication skills.


Preferred Qualifications

  • Experience supporting cybersecurity, risk, resilience, or enterprise security platforms.
  • Experience with observability platforms such as Splunk, Grafana, Prometheus, Datadog, Dynatrace, New Relic, Azure Monitor, CloudWatch, or OpenTelemetry.
  • Experience building executive-level dashboards, operational scorecards, and reliability reporting.
  • Experience with automated health checks, synthetic monitoring, dependency mapping, and operational runbooks.
  • Experience with Kubernetes, Docker, serverless platforms, or cloud-native technologies.
  • Experience with Apache Kafka or other messaging platforms.
  • Familiarity with cloud security governance tools such as Azure Policy, AWS SCP, Wiz, Prisma, CloudGuard, or similar solutions.
  • Experience with AI cloud platforms such as Azure AI, AWS Bedrock, or Google Vertex AI.
  • Experience supporting Linux and Windows environments through scripting and automation.


Required Skills

  1. Site Reliability Engineering (SRE)
  2. DevOps
  3. Cloud Platforms (AWS, Azure, GCP)
  4. Infrastructure as Code (Terraform)
  5. Monitoring & Observability
  6. Logging & Tracing
  7. Alerting
  8. Dashboard Development
  9. Incident Management
  10. Root Cause Analysis
  11. Service Reliability
  12. SLIs / SLOs / SLAs
  13. Performance Monitoring
  14. Capacity Planning
  15. Disaster Recovery
  16. CI/CD
  17. Python
  18. Java
  19. Bash
  20. PowerShell
  21. REST APIs
  22. Distributed Systems
  23. Microservices
  24. SQL Databases
  25. NoSQL Databases
  26. Automation
  27. Troubleshooting
  28. Cloud Infrastructure
  29. Operational Excellence
  30. Communication Skills
  31. Problem Solving


Preferred Skills

  1. Splunk
  2. Grafana
  3. Prometheus
  4. Datadog
  5. Dynatrace
  6. New Relic
  7. Azure Monitor
  8. AWS CloudWatch
  9. OpenTelemetry
  10. Kubernetes
  11. Docker
  12. Apache Kafka
  13. Wiz
  14. Prisma Cloud
  15. CloudGuard
  16. Azure AI
  17. AWS Bedrock
  18. Google Vertex AI
  19. Linux Administration
  20. Windows Administration
  21. Executive Dashboard Reporting
  22. Synthetic Monitoring
  23. Service Dependency Mapping
  24. Cloud Security
  25. Cybersecurity Platforms


Education

Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field, or equivalent work experience.

Certifications

Relevant cloud, DevOps, SRE, Kubernetes, or cybersecurity certifications are preferred.

Similar Jobs

More Jobs at Compunnel

  • Site Reliability Engineer
    $100K — $130K *
    Toronto, ON M3C 0E3
    Information Technology
    In-Person
  • Systems Analyst III
    $80K — $110K *
    Allentown, PA 18102 (Lehigh County)
    Information Technology
    In-Person
  • Project Manager
    $90K — $120K *
    Aliso Viejo, CA 92656 (Orange County)
    Business Services
    In-Person
  • Full Stack Engineer
    $100K — $130K *
    Durham, NC 27713 (Durham County)
    Information Technology
    In-Person
  • Data Engineer
    $90K — $130K *
    Wayne, PA 19087 (Delaware County)
    Enterprise Technology
    In-Person

More Information Technology Jobs

Find similar Site Reliability Engineer jobs: