Datavant

Site Reliability Engineer

Datavant$100K — $130K *
US-AnywhereRemote in United States
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years of infrastructure management or platform engineering experience
  • Deep understanding of multi-cloud networking (e.g., AWS VPCs, Azure VNets)
  • Hands-on experience with Infrastructure as Code (Terraform, Ansible)
  • Excellent problem-solving skills and ability to clarify complex issues
  • Strong communication skills to articulate technical concepts to non-technical personnel
  • Curiosity and adaptability in a growing enterprise
  • Experience with AI agents for workload optimization

Responsibilities

  • Automate cloud infrastructure provisioning using Terraform and Ansible
  • Support platform engineering and app development teams in troubleshooting and maintaining infrastructure
  • Architect and maintain secure network infrastructure across Azure, AWS, and GCP
  • Configure and manage virtual networks, including IP addressing and routing
  • Collaborate with security teams for Zero Trust networking and IAM standards
  • Maintain a flexible network topology for development team integration testing
  • Implement secure connections using Direct Connect, ExpressRoute, and VPN

Benefits

  • Flexible work environment supporting remote collaboration
  • Comprehensive health and wellness programs
  • Opportunities for professional development and career growth
  • Employee recognition programs that highlight achievements
  • Inclusive company culture focused on diversity and teamwork
Full Job Description
We are seeking a Site Reliability Engineer to play a key role in designing, optimizing, and securing the underlying cloud infrastructure that powers our organization's cloud environments. As we continue to transform, consolidate, and evolve our cloud environments, this role will be instrumental in architecting scalable, secure, automated, and resilient infrastructure, ensuring seamless workload migration, and integrating new cloud environments into our ecosystem.

With a focus on core cloud infrastructure, including networking, identity and access management, security, storage, compute, and cross-cloud integrations, you will collaborate closely with peers in security, development, and other areas of the platform engineering team to ensure our cloud environment adheres to best practices, governance frameworks, and automation-first principles.

What You Will Do

Reliability & Technical Ownership
  • Lead the design and implementation of reliability improvements across assigned services, with minimal guidance
  • Identify systemic inefficiencies in architecture, implementation, and operational process and drive solutions
  • Implement customized solutions to complex operational problems derived from technical requirements
  • Review code, systems, and configuration with a focus on efficiency gains, optimization, and best practices and hold peers to those standards
  • Own SLO/SLI definitions for assigned services and drive teams toward meeting and improving those targets
  • Lead incident response, facilitate postmortems, and ensure action items result in durable reliability improvements

M&A Integration & Environment Consolidation
  • Support the integration of newly acquired cloud environments into Datavant's existing infrastructure, ensuring reliability, security, and operational consistency from day one
  • Contribute to the implementation of hybrid-cloud and cross-cloud connectivity strategies that ensure interoperability across a growing multi-cloud footprint
  • Help maintain and expand the modular network security edge, keeping it flexible enough to absorb additional environments as M&A activity requires
  • Drive standardization of infrastructure and operational practices across consolidated environments, reducing fragmentation and toil
  • Partner with security and platform engineering teams to ensure newly integrated environments adhere to governance frameworks and automation-first principles

Service Delivery & Process Improvement
  • With limited guidance, develop tools and processes to improve team service delivery including scaling, resiliency, efficiency, visibility, quality, and operations management
  • Analyze service delivery data and team feedback to drive meaningful improvements to development processes
  • Participate in and help evolve on-call practices, runbooks, and alerting strategy for assigned teams
  • Address communication gaps and produce clear documentation of process changes and technical standards

Collaboration & Mentorship
  • Teach and lead more junior engineers on team processes, technical implementations, and SRE best practices
  • Accept and promote sound engineering decisions including those that weren't your own and build alignment around them
  • Collaborate cross-functionally with development, security, and platform engineering teams to embed reliability thinking into the software development lifecycle

Automation & Tooling
  • Build and maintain Infrastructure as Code (Terraform, Ansible) to support scalable, repeatable, and secure deployments
  • Develop and improve CI/CD pipelines, automated testing, and deployment tooling
  • Implement policy-based automation (AWS SCPs, Azure Policy) to maintain governance across cloud environments
  • Extend observability coverage through instrumentation, dashboards, and alerting using Datadog, CloudWatch, and Azure Log Analytics
  • Leverage AI tools and agents to accelerate and improve daily engineering workflows


What You Need to Succeed
  • 5+ years of experience in site reliability engineering, DevOps, or platform/infrastructure engineering
  • Strong expertise in cloud infrastructure, including networking, security, compute, storage, and IAM
  • Experience supporting workload migrations and integrating new cloud environments into existing architectures
  • Hands-on experience with Infrastructure as Code and automation (Terraform and Ansible)
  • Strong security knowledge, including IAM, encryption, network security, and compliance frameworks (SOC2, HITRUST, NIST)
  • Demonstrated ability to solve complex operational problems independently and drive solutions end-to-end
  • Strong proficiency in at least one systems language (Python, Go, or similar) and comfort across multiple languages and configuration formats
  • Proven ability to conduct meaningful code and system reviews not just for correctness, but for efficiency and architectural soundness
  • Strong communication skills, including the ability to document technical decisions, address gaps in understanding, and build team alignment
  • Experience leveraging AI agents to accelerate daily workload


Competency with the following technologies:
  • Compute: AWS EC2, Azure VMs, Kubernetes / containerized workloads
  • Network security: AWS (VPC, TGW, Peering, SG, ALB, NLB); Azure (VNets, NSG, AGW); VPN
  • Observability: Datadog, CloudWatch, Azure Log Analytics instrumentation, dashboards, alerting
  • Automation: Terraform, Ansible, GitHub Actions or equivalent CI/CD tooling
  • Cloud platforms: Hands-on experience in AWS and/or Azure; familiarity with GCP
  • IAM: AWS IAM, Azure RBAC, least-privilege access patterns
  • Core OS: Linux (required), Windows (helpful); DNS / IPAM


What Helps You Stand Out
  • Experience with multi-cloud environments (AWS, Azure, GCP) and multi-account cloud governance (AWS Organizations, Azure Policy, SCPs)
  • Background in driving environment consolidation and integration bringing order to disparate hybrid and multi-cloud architectures
  • Experience defining and operating against SLOs/SLIs and error budgets
  • Familiarity with FinOps principles and cloud cost optimization
  • Track record of improving on-call health reducing alert fatigue, improving runbook quality, driving down MTTR
  • Experience with cloud-native identity management (AWS IAM Identity Center, Entra ID, RBAC)
  • Knowledge of multi-region architecture and disaster recovery strategies


At Datavant our total rewards strategy powers a high-growth, high-performance, health technology company that rewards our employees for transforming health care through creating industry-defining data logistics products and services.

The range posted is for a given job title, which can include multiple levels. Individual rates for the same job title may differ based on their level, responsibilities, skills, and experience for a specific job.

The estimated total cash compensation range for this role is:

$100,000-$130,000 USD

This job is not eligible for employment sponsorship.

About Datavant

Datavant is a healthcare technology company that specializes in connecting and standardizing healthcare data from various sources. The company's products are used to improve patient care, accelerate drug development, and enhance clinical research. Datavant's platform, called the Datavant Platform, uses artificial intelligence and machine learning to identify and link patient data across different sources while maintaining patient privacy. The company was founded in 2017 by Travis May and is headquartered in San Francisco, California.
Learn more about Datavant
Size
100 employees
Industry
Founded
2017

Similar Jobs

More Jobs at Datavant

More Information Technology Jobs

Find similar Site Reliability Engineer jobs: