We are seeking a Site Reliability Engineer to play a key role in designing, optimizing, and securing the underlying cloud infrastructure that powers our organization's cloud environments. As we continue to transform, consolidate, and evolve our cloud environments, this role will be instrumental in architecting scalable, secure, automated, and resilient infrastructure, ensuring seamless workload migration, and integrating new cloud environments into our ecosystem.
With a focus on core cloud infrastructure, including networking, identity and access management, security, storage, compute, and cross-cloud integrations, you will collaborate closely with peers in security, development, and other areas of the platform engineering team to ensure our cloud environment adheres to best practices, governance frameworks, and automation-first principles.
What You Will DoReliability & Technical Ownership- Lead the design and implementation of reliability improvements across assigned services, with minimal guidance
- Identify systemic inefficiencies in architecture, implementation, and operational process and drive solutions
- Implement customized solutions to complex operational problems derived from technical requirements
- Review code, systems, and configuration with a focus on efficiency gains, optimization, and best practices and hold peers to those standards
- Own SLO/SLI definitions for assigned services and drive teams toward meeting and improving those targets
- Lead incident response, facilitate postmortems, and ensure action items result in durable reliability improvements
M&A Integration & Environment Consolidation- Support the integration of newly acquired cloud environments into Datavant's existing infrastructure, ensuring reliability, security, and operational consistency from day one
- Contribute to the implementation of hybrid-cloud and cross-cloud connectivity strategies that ensure interoperability across a growing multi-cloud footprint
- Help maintain and expand the modular network security edge, keeping it flexible enough to absorb additional environments as M&A activity requires
- Drive standardization of infrastructure and operational practices across consolidated environments, reducing fragmentation and toil
- Partner with security and platform engineering teams to ensure newly integrated environments adhere to governance frameworks and automation-first principles
Service Delivery & Process Improvement- With limited guidance, develop tools and processes to improve team service delivery including scaling, resiliency, efficiency, visibility, quality, and operations management
- Analyze service delivery data and team feedback to drive meaningful improvements to development processes
- Participate in and help evolve on-call practices, runbooks, and alerting strategy for assigned teams
- Address communication gaps and produce clear documentation of process changes and technical standards
Collaboration & Mentorship- Teach and lead more junior engineers on team processes, technical implementations, and SRE best practices
- Accept and promote sound engineering decisions including those that weren't your own and build alignment around them
- Collaborate cross-functionally with development, security, and platform engineering teams to embed reliability thinking into the software development lifecycle
Automation & Tooling- Build and maintain Infrastructure as Code (Terraform, Ansible) to support scalable, repeatable, and secure deployments
- Develop and improve CI/CD pipelines, automated testing, and deployment tooling
- Implement policy-based automation (AWS SCPs, Azure Policy) to maintain governance across cloud environments
- Extend observability coverage through instrumentation, dashboards, and alerting using Datadog, CloudWatch, and Azure Log Analytics
- Leverage AI tools and agents to accelerate and improve daily engineering workflows
What You Need to Succeed- 5+ years of experience in site reliability engineering, DevOps, or platform/infrastructure engineering
- Strong expertise in cloud infrastructure, including networking, security, compute, storage, and IAM
- Experience supporting workload migrations and integrating new cloud environments into existing architectures
- Hands-on experience with Infrastructure as Code and automation (Terraform and Ansible)
- Strong security knowledge, including IAM, encryption, network security, and compliance frameworks (SOC2, HITRUST, NIST)
- Demonstrated ability to solve complex operational problems independently and drive solutions end-to-end
- Strong proficiency in at least one systems language (Python, Go, or similar) and comfort across multiple languages and configuration formats
- Proven ability to conduct meaningful code and system reviews not just for correctness, but for efficiency and architectural soundness
- Strong communication skills, including the ability to document technical decisions, address gaps in understanding, and build team alignment
- Experience leveraging AI agents to accelerate daily workload
Competency with the following technologies:- Compute: AWS EC2, Azure VMs, Kubernetes / containerized workloads
- Network security: AWS (VPC, TGW, Peering, SG, ALB, NLB); Azure (VNets, NSG, AGW); VPN
- Observability: Datadog, CloudWatch, Azure Log Analytics instrumentation, dashboards, alerting
- Automation: Terraform, Ansible, GitHub Actions or equivalent CI/CD tooling
- Cloud platforms: Hands-on experience in AWS and/or Azure; familiarity with GCP
- IAM: AWS IAM, Azure RBAC, least-privilege access patterns
- Core OS: Linux (required), Windows (helpful); DNS / IPAM
What Helps You Stand Out- Experience with multi-cloud environments (AWS, Azure, GCP) and multi-account cloud governance (AWS Organizations, Azure Policy, SCPs)
- Background in driving environment consolidation and integration bringing order to disparate hybrid and multi-cloud architectures
- Experience defining and operating against SLOs/SLIs and error budgets
- Familiarity with FinOps principles and cloud cost optimization
- Track record of improving on-call health reducing alert fatigue, improving runbook quality, driving down MTTR
- Experience with cloud-native identity management (AWS IAM Identity Center, Entra ID, RBAC)
- Knowledge of multi-region architecture and disaster recovery strategies
At Datavant our total rewards strategy powers a high-growth, high-performance, health technology company that rewards our employees for transforming health care through creating industry-defining data logistics products and services.
The range posted is for a given job title, which can include multiple levels. Individual rates for the same job title may differ based on their level, responsibilities, skills, and experience for a specific job.
The estimated total cash compensation range for this role is:
$100,000-$130,000 USD
This job is not eligible for employment sponsorship.