Location: Remote (US-based candidates only)
Manager, Site Reliability Engineering (SRE)
Position OverviewWe are seeking a Manager, Site Reliability Engineering (SRE) to lead our US-based SRE team and drive operational excellence across our production platforms.
This is a player-coach leadership role that combines people management with hands-on technical leadership. You will mentor and grow a team of Site Reliability Engineers while actively participating in major incident response, reliability initiatives, and operational reviews. The role is a key part of our global follow-the-sun support model and requires close collaboration with SRE leadership in India.
Key ResponsibilitiesTeam Leadership & Development- Lead, mentor, and develop a US-based team of Site Reliability Engineers.
- Conduct regular 1:1s, performance reviews, and career development discussions.
- Own hiring, onboarding, and retention efforts as the team scales.
- Foster a culture of ownership, blameless postmortems, and continuous improvement.
Operational Excellence & Incident Management- Lead day-to-day production operations and ensure timely incident triage, resolution, and escalation.
- Serve as an escalation point and incident commander for major production incidents.
- Drive problem management and root cause analysis processes.
- Carry PagerDuty on-call escalation responsibilities for critical issues.
- Track and report operational KPIs, SLAs, and SLOs, including availability, MTTR, and incident trends.
Reliability & Automation- Improve system reliability, observability, and resilience using Datadog and related tooling.
- Drive automation, self-healing capabilities, and runbook maturity.
- Partner with Development, DevOps, DevSecOps, and Engineering teams to embed reliability into the SDLC.
- Contribute hands-on to tooling, automation, and technical reviews as needed.
Collaboration & Global Alignment- Coordinate closely with SRE leadership in India to ensure seamless follow-the-sun coverage.
- Represent the US SRE organization in cross-functional planning and operational reviews.
- Communicate effectively with both technical and non-technical stakeholders.
Documentation & Compliance- Maintain high-quality documentation for incidents, postmortems, runbooks, and operational procedures.
- Ensure adherence to healthcare and fintech compliance standards, including HIPAA, PCI DSS, SOC 2, ISO 27001, and HITRUST.
Required Qualifications- 5-8 years of experience in Site Reliability Engineering, DevOps, Production Support, or Platform Engineering.
- 1-2+ years of experience leading, mentoring, or managing engineers.
- Demonstrated success operating in a player-coach leadership model.
- Strong hands-on experience with production incident management and escalation processes.
- Proficiency with Datadog or similar observability platforms.
- Hands-on experience with Kubernetes and Docker in production environments.
- Strong scripting or programming skills in PowerShell, Bash, Python, Java, or C#.
- Experience with Helm, CI/CD pipelines, and deployment automation.
- Working knowledge of ITIL processes and Agile methodologies.
- Experience working with SQL, MySQL, or NoSQL databases.
- Excellent communication and stakeholder management skills.
- Willingness to participate in PagerDuty on-call escalation and work within a global follow-the-sun operating model.
Preferred Qualifications- Experience with cloud platforms such as AWS, Azure, or GCP.
- Experience building or scaling SRE teams and on-call programs.
- Experience defining and managing SLOs, SLIs, and error budgets.
- Prior experience in the healthcare or fintech industry.
- Knowledge of security and compliance frameworks relevant to regulated environments.
Why Join NationsBenefits?- Competitive compensation and comprehensive benefits.
- Unlimited PTO.
- Fully remote work environment (US-based).
- Opportunity to lead and grow a high-impact SRE organization.
- Exposure to modern cloud-native technologies and large-scale reliability challenges.
- Collaborative culture focused on innovation, learning, and continuous improvement.
- Meaningful work that directly impacts healthcare technology and millions of members.
Ideal CandidateWe are looking for a technically strong SRE leader who enjoys building teams, improving operational maturity, and remaining hands-on during critical production events. The ideal candidate combines leadership, systems thinking, and automation expertise to help scale reliability practices across a fast-growing Healthcare FinTech organization.