Full Job Description
We are seeking a Staff Site Reliability Engineer to play a key role in designing, optimizing, and securing the underlying cloud infrastructure that powers our organization's cloud environments. As we continue to transform, consolidate, and evolve our cloud environments, this role will be instrumental in architecting scalable, secure, automated, and resilient infrastructure, ensuring seamless workload migration, and integrating new cloud environments into our ecosystem.
With a focus on core cloud infrastructure, including networking, identity and access management, security, storage, compute, and cross-cloud integrations, you will collaborate closely with peers in security, development, and other areas of the platform engineering team to ensure our cloud environment adheres to best practices, governance frameworks, and automation-first principles.
What You Will Do
Technical Leadership & Architecture
• Own the technical direction of SRE functions for designated teams from architecture definition and test planning through implementation and ongoing operations
• Lead design and development of tooling and processes that improve service delivery at scale: reliability, resiliency, efficiency, visibility, and quality
• Drive adoption of researched, verified, and standardized engineering practices across the teams you support
• Enhance system designs and infrastructure implementations to make them more efficient, standardized, and maintainable
• Take end-to-end ownership of major projects defining architecture, test plans, and implementation and empower others to do the same
• Push for and advocate superior architectural solutions; provide a compelling, data-backed case for change
• Drive transformation and standardization of foundational infrastructure services as the environment evolves and the business grows
M&A Integration & Environment Consolidation
• Architect and lead the integration of newly acquired cloud environments into Datavant's infrastructure, ensuring security, reliability, and consistency across an expanding multi-cloud footprint
• Design and implement hybrid-cloud and cross-cloud connectivity strategies to ensure interoperability as M&A activity adds new environments
• Own and evolve the modular network security edge, ensuring it remains flexible enough to unify additional environments while providing a consistent integration testing surface for development teams
• Drive environment consolidation initiatives bringing order to disparate hybrid and multi-cloud architectures and reducing operational fragmentation across the organization
• Collaborate with security, platform engineering, and development teams to ensure newly integrated environments adhere to Zero Trust principles, governance frameworks, and automation-first practices
• Optimize processes, develop policies, and enable secure, self-service infrastructure for engineering partners operating across consolidated environments
Operational Excellence
• Drive resolution of complex and systemic operational issues without needing guidance
• Lead enhancements and improvements to service delivery practices for all designated teams
• Define and evolve SLO/SLI frameworks and error budget policies across your scope
• Establish and continuously improve on-call practices, incident response processes, and postmortem culture
• Identify opportunities for efficiency gains and actively share that knowledge across the team
Mentorship & Team Enablement
• Serve as a mentor to SREs and developers on designated teams guiding technical growth, reviewing work, and raising the engineering bar
• Responsible for the mentorship and growth of SREs within your scope of influence
• Train and lead others on the team in completing complex tasks; delegate effectively and empower others to own architectures and implementations independently
• Drive best practices and standards for SRE functions within designated teams
Cross-functional Influence
• Partner deeply with development, security, and platform engineering teams to embed reliability, scalability, and operability into the software development lifecycle
• Drive improvements to development processes based on data-driven analysis and team feedback
• Communicate complex engineering challenges and trade-offs clearly to both technical and non-technical stakeholders
• Leverage AI tools and agents strategically to accelerate delivery and improve team workflows
Automation & Platform
• Lead the design and implementation of automation that improves platform reliability, deployment safety, and operational efficiency
• Architect and build scalable observability solutions instrumentation strategy, dashboards, alerting, and on-call tooling
• Own Infrastructure as Code standards and practices (Terraform, Ansible) for designated teams, writing IaC that supports simple, secure, and scalable deployments
• Implement policy-based automation (AWS SCPs, Azure Policy) to maintain governance across cloud environments, including newly integrated M&A environments
• Champion CI/CD pipeline improvements that reduce toil and increase delivery confidence
What You Need to Succeed
• 8+ years of experience in site reliability engineering, platform engineering, or infrastructure architecture
• Strong expertise in cloud infrastructure, including networking, security, compute, storage, and IAM
• Experience migrating workloads and integrating new cloud environments into existing architectures including M&A-driven consolidation scenarios
• Deep understanding of multi-cloud networking (AWS VPCs, Azure VNets, Transit Gateway, ExpressRoute, PrivateLink, DNS, hybrid connectivity)
• Hands-on experience with Infrastructure as Code and automation (Terraform and Ansible)
• Strong security knowledge, including IAM, encryption, network security, PKI, certificate management, and compliance frameworks (SOC2, HITRUST, NIST, FedRAMP)
• Experience managing multi-account cloud governance and security policies (AWS Organizations, Azure Policy, SCPs)
• Proven track record of technical leadership owning large, complex projects end-to-end and driving them to completion
• Demonstrated ability to mentor engineers, lead design reviews, and elevate the technical quality of a team
• Excellent problem-solving skills, with the drive to bring clarity to vague situations and a desire to constantly expand your skills
• Strong communication skills, including the ability to describe complex problems to non-technical staff, paired with the curiosity, adaptability, and flexibility needed to navigate a growing enterprise
• Experience leveraging AI agents to accelerate daily workload
Competency with the following technologies:
• Compute: AWS EC2, Azure VMs, Kubernetes / containerized workloads at scale
• Network security: AWS (VPC, TGW, Peering, SG, ALB, NLB); Azure (VNets, NSG, AGW); NextGen FW, VPN
• Storage: Azure Files & Blob, Amazon S3, AWS FSxN
• Automation: Ansible Automation Platform, Terraform, GitHub
• Core: Linux & Windows, DNS / IPAM, Active Directory
• Observability: Datadog, CloudWatch, Azure Log Analytics advanced instrumentation, SLO tooling, alerting strategy
• IAM: AWS IAM / IAM Identity Center, Azure Entra ID / RBAC
What Helps You Stand Out
• Experience driving environment consolidation and integration bringing order to disparate hybrid and multi-cloud architectures as the business grows
• Experience driving technical initiatives end-to-end across multiple teams or organizational boundaries
• Deep knowledge of cloud services, containerization, and/or Kubernetes
• Knowledge of multi-region architecture and disaster recovery strategies
• Background in cloud cost optimization and FinOps methodologies
• Experience with cloud-native identity management (AWS IAM Identity Center, Entra ID, RBAC)
• Track record of building and scaling an SRE practice within a high-growth engineering organization
• Group Policy, PowerShell, CIFS & NFS storage protocols, Palo Alto, F5
At Datavant our total rewards strategy powers a high-growth, high-performance, health technology company that rewards our employees for transforming health care through creating industry-defining data logistics products and services.
The range posted is for a given job title, which can include multiple levels. Individual rates for the same job title may differ based on their level, responsibilities, skills, and experience for a specific job.
The estimated total cash compensation range for this role is:
$190,000-$235,000 USD