The Site Reliability Engineer role focuses on enhancing the reliability and operational excellence of enterprise platforms across hybrid cloud and on-premises environments. This senior technical leader drives improvements in automation, observability, and incident management while collaborating across multiple business and technology teams.
Responsibilities include leading major incident responses, driving problem management, and implementing automation to reduce service downtime.
The role involves standardizing observability practices, mentoring SRE team members, and contributing to enterprise-wide reliability frameworks.
Candidates require 7+ years of experience, expertise in distributed systems, Kubernetes, automation scripting, and strong leadership in incident management.
ESSENTIAL DUTIES AND RESPONSIBILITIES
Following is a summary of the essential functions for this job. Other duties may be performed, both major and minor, which are not mentioned below. Specific activities may change from time to time.
1. Implements software architecture and engineering approaches for complex initiatives within the job area, contributing to technical plans and working to achieve operational targets with major impact on results.
2. Adopts and refines advanced software engineering standards, practices, and governance mechanisms for the job area, influencing how multiple teams improve quality, reliability, and delivery.
3. Collaborates with senior engineers, product partners, and architecture teammates to shape technology approaches for the domain, providing deep technical insight and proposing solution patterns that inform local roadmaps and priorities.
4. Leads the end-to-end technical design and implementation of scalable, secure, and highly available software solutions for the job area, producing patterns and examples that other technical professionals can follow.
5. Independently troubleshoots and resolves complex technical issues in the area of responsibility, designing innovative architectures and performance, reliability, and scalability improvements that advance business objectives.
6. Provides ongoing technical guidance, coaching, and training to other engineers, delegating and reviewing work from lower-level technical professionals and raising the technical bar through design reviews and knowledge sharing.
7. Evaluates emerging technologies and techniques relevant to the job area, building prototypes and solution concepts that contribute measurable input into new features, products, or capabilities.
8. Contributes to the development of long-term technical goals and plans for the area of responsibility through well-reasoned recommendations, design proposals, and implementation experience.
9. Leads large or complex initiatives within the job area, coordinating and delegating technical work that may span outside the immediate team, and ensuring cohesive, high-quality outcomes with limited supervision.
Qualifications
Required Qualifications
The requirements listed below are representative of the knowledge, skill and/or ability required. Reasonable accommodations may be made to enable individuals with disabilities to perform the essential functions.
1. Bachelor’s degree in Computer Science, Software Engineering, or related field.
2. Minimum of 7 years of professional experience in software development.
3. Deep knowledge of multiple programming languages, software architecture, and design principles.
4. Deep understanding of software development lifecycle, testing, deployment, and security practices.
Preferred Qualifications
1. Advanced degree in Computer Science or related technical discipline.
2. Professional certifications such as Certified Software Development Professional (CSDP) or equivalent.
3. Deep expertise in cloud-native architectures, microservices, container orchestration, and DevOps.
4. Strong familiarity with Agile frameworks, continuous integration/continuous deployment (CI/CD), and enterprise innovation management.
5. 7+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Infrastructure Operations.
6. Deep hands‑on experience with distributed systems, container orchestration (Kubernetes), and cloud-native operational tooling.
7. Proficiency with automation and scripting languages (Python, Go, PowerShell, Ansible).
8. Strong understanding of observability platforms (Splunk, Dynatrace) and event-driven monitoring.
9. Proven leadership in major incident management and cross-team technical coordination.
10. Strong grasp of networking, Linux/Unix internals, and modern infrastructure patterns.
11. Excellent communication skills, including executive-level situational awareness during critical incidents.
12. Demonstrated ability to influence technical roadmaps and drive adoption of reliability best practices.
Incident & Problem Management Leadership
Reliability Engineering & Automation
Observability & Operational Excellence
Enhance telemetry coverage across logs, metrics, traces, and events using platforms such as Dynatrace and Splunk.
Cross-Functional Leadership & Influence
Partner with Delivery, Architecture, Security, and Risk teams to embed reliability and resilience into design and execution.
Standardization & Documentation
Develop, maintain, and enforce runbooks, response playbooks, and automated recovery patterns.
Mentorship & Technical Development
Candidate must be willing to work onsite Monday - Friday at either office in Charlotte NC, Raleigh NC, or Atlanta, GA.