Site Reliability Engineer

Compunnel

• $110K — $130K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in Computer Science, IT, Engineering, or related field, or equivalent experience.
  • 5+ years of experience in Site Reliability Engineering, DevOps, or related fields.
  • Experience with business-critical production services in hybrid on-premises and cloud environments.
  • Hands-on with monitoring, observability, alerting, and capacity planning.
  • Skilled in automation using PowerShell, Python, REST APIs, Terraform, Ansible, or similar technologies.
  • Expertise in incident management and root cause analysis processes.
  • Knowledge of ITIL and ITSM practices like Incident and Problem Management.

Responsibilities

  • Define availability targets and reliability metrics for critical services.
  • Partner with Service Owners to document ownership and escalation paths.
  • Identify and mitigate single points of failure and operational risks.
  • Implement monitoring and alerting across all service layers.
  • Drive technical workstreams during P1 and P2 incidents for service restoration.
  • Lead investigations for high-impact incidents and coordinate root cause analysis.
  • Automate manual operational processes to improve operational efficiency.

Benefits

  • Flexible work hours and potential for remote work.
  • Opportunities for professional development and certification.
  • Access to cutting-edge technology and tools.
  • Collaborative work environment with cross-functional teams.
Full Job Description
Job Summary The Site Reliability Engineer will improve the availability, resiliency, recoverability, performance, and operational supportability of critical technology services across on-premises, cloud, network, database, VDI, storage, and vendor-managed environments. The role applies automation, observability, infrastructure engineering, and ITSM practices to reduce recurring incidents, eliminate manual operational work, strengthen recovery readiness, and improve service ownership and accountability. Key Responsibilities • Define availability targets, error budgets, and reliability metrics for critical services. • Partner with Service Owners and technical teams to document ownership, dependencies, escalation paths, runbooks, and recovery requirements. • Identify single points of failure, fragile dependencies, capacity constraints, and operational risks before they cause outages. • Prioritize reliability improvements based on business impact, risk, technical feasibility, and cost. • Implement actionable monitoring, alerting, dashboards, service health indicators, and event correlation across infrastructure, networks, applications, databases, cloud resources, storage, and VDI. • Reduce alert noise through improved thresholds, dependencies, routing, escalation logic, and baselines. • Develop early-warning indicators for storage exhaustion, routing instability, database integrity issues, certificate expiration, capacity saturation, and service availability risks. • Use incident and performance data to identify recurring patterns and reliability improvement opportunities. • Provide technical leadership during P1 and P2 incidents involving multiple technologies, vendors, or support teams. • Drive technical workstreams, validate hypotheses, identify dependencies, and maintain focus on service restoration. • Lead problem investigations for recurring, high-impact, or complex incidents and coordinate root cause analysis. • Translate root cause findings into corrective-action backlogs with owners, target dates, measurable outcomes, and documented risks. • Replace repetitive, manual, error-prone, or inefficient operational processes with sustainable automation. • Build scripts, integrations, workflows, and infrastructure automation using PowerShell, Python, REST APIs, Terraform, Ansible, or comparable technologies. • Automate diagnostics, service validation, recovery checks, configuration reviews, capacity checks, and remediation workflows. • Design automation with appropriate testing, security, logging, rollback, and change-control safeguards. • Define and validate recovery requirements, RTOs, and RPOs for critical business services. • Design and execute failover, fault-tolerance, recovery, and disaster recovery tests. • Validate clustering, replication, backups, high availability, and recovery procedures against business requirements. • Maintain recovery runbooks, dependency maps, test evidence, lessons learned, and remediation actions. • Partner with DBA teams to improve Microsoft SQL Server reliability, integrity, recoverability, and observability. • Validate monitoring for database availability, replication, backups, storage, capacity, clustering, and integrity. • Support testing of SQL Failover Clusters, Availability Groups, secondary-site replication, point-in-time recovery, and database recovery procedures. • Participate in root cause analysis for database corruption, infrastructure failures, recovery delays, and data consistency events. • Partner with Network Engineering to improve WAN, SD-WAN, switching, wireless, VPN, routing, Internet connectivity, and remote-access reliability. • Develop monitoring and validation for BGP route advertisements, convergence, path selection, failover behavior, traffic flow, and asymmetric routing. • Support fault-tolerance and failover testing for critical network paths and services. • Develop network diagrams, routing documentation, dependency maps, troubleshooting procedures, and escalation runbooks. • Apply strong knowledge of TCP/IP, DNS, DHCP, routing, switching, VLANs, VPN, firewalls, wireless, WAN, Internet, and hybrid-cloud connectivity. • Support SD-WAN environments, preferably Cisco Meraki, as well as enterprise networking platforms such as Cisco and Juniper. • Use packet captures, flow data, SNMP, syslog, API telemetry, synthetic testing, and path analysis to troubleshoot and monitor network performance. • Support SASE, ZTNA, and remote-access platforms such as Zscaler and Citrix NetScaler. • Correlate network behavior with application, database, storage, VDI, and cloud-service impact. • Support reliability architecture and operational readiness for Azure, hybrid-cloud services, and Azure Virtual Desktop. • Define monitoring, capacity, availability, security, and recovery requirements for VDI and cloud services. • Analyze VDI compute, memory, profile, session, storage, and application performance. • Ensure cloud migrations include operational acceptance criteria, support documentation, observability, recovery testing, and defined ownership. • Identify unsupported, end-of-life, and approaching-end-of-support technologies across operating systems, virtualization, hardware, databases, network platforms, and supporting services. • Support modernization planning for Windows Server, Red Hat Enterprise Linux, VMware vSphere, VMware vCenter, Cisco UCS, Citrix, and Microsoft SQL Server. • Prioritize vulnerabilities based on exploitability, exposure, service criticality, and business impact. • Support SLA-based vulnerability remediation, exception tracking, and risk acceptance through Ivanti Neurons for ITSM or a comparable ITSM platform. • Participate in architecture, design, change, and production-readiness reviews for critical services. • Define reliability acceptance criteria for production implementations. • Verify that changes include monitoring, testing, rollback, recovery, ownership, documentation, and vendor-support plans. • Support post-implementation validation and confirm that new services meet availability and support requirements. • Coordinate reliability initiatives across internal technology teams and external service providers. • Establish vendor escalation paths, support expectations, communication procedures, and technical accountability. • Lead technical discussions with cloud, infrastructure, networking, virtualization, storage, security, and application vendors. • Support RACI models for service ownership, incident response, vulnerability remediation, lifecycle management, and disaster recovery. Required Qualifications • Bachelor's degree in Computer Science, Information Technology, Engineering, or a related discipline, or equivalent professional experience. • 5+ years of experience in site reliability engineering, infrastructure engineering, DevOps, systems engineering, production engineering, or enterprise operations. • Experience supporting business-critical production services in hybrid on-premises and cloud environments. • Hands-on experience with monitoring, observability, alerting, performance analysis, and capacity planning. • Experience automating operational processes using PowerShell, Python, REST APIs, Terraform, Ansible, or comparable technologies. • Experience with incident management, root cause analysis, problem management, and corrective-action tracking. • Knowledge of availability targets, disaster recovery, high availability, and reliability reporting. • Experience developing technical documentation, dependency maps, operational runbooks, and recovery procedures. • Working knowledge of Windows Server, Linux, virtualization, cloud, networking, storage, or enterprise database platforms. • Knowledge of ITIL and ITSM practices, including Incident, Problem, Change, Configuration, and Knowledge Management. • Strong analytical, troubleshooting, communication, and cross-functional collaboration skills.

Similar Jobs

More Jobs at Compunnel

  • SCCM Administrator
    $90K — $110K *
    Taylor, TX 76574 (Williamson County)
    Information Technology
    In-Person
  • Node Developer
    $100K — $120K *
    Alpharetta, GA 30022 (Fulton County)
    Information Technology
    In-Person
  • Senior AI Engineer
    $120K — $145K *
    Taylor, TX 76574 (Williamson County)
    Enterprise Technology
    In-Person
  • SAP Security Analyst
    $110K — $130K *
    Charlotte, NC 28269 (Mecklenburg County)
    Information Technology
    In-Person
  • Director of AI Engineering
    $150K — $180K *
    Charlotte, NC 28269 (Mecklenburg County)
    Information Technology
    In-Person

More Information Technology Jobs

Find similar Site Reliability Engineer jobs: