Datavant

Senior Site Reliability Engineer

Datavant$168K — $200K *
US-AnywhereRemote in United States
Enterprise Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 6+ years in SRE, platform engineering, or DevOps roles catering to data-intensive or ML-powered applications
  • Proficient with Databricks, including management and integration with CI/CD
  • Solid understanding of cloud-native infrastructure on AWS or similar
  • Proven skills in observability tools, specifically Datadog
  • Strong experience with CI/CD tooling like GitHub Actions and Terraform
  • Familiar with shell scripting and Python
  • Experience in building fault-tolerant systems

Responsibilities

  • Own the lifecycle of Databricks and Snowflake platforms, focusing on automation and cost optimization
  • Architect resilient and secure infrastructure across cloud environments
  • Establish robust monitoring and alerting systems using Datadog
  • Automate the deployment of data pipelines and ML workflows with GitHub Actions and Terraform
  • Support data movement across various cloud systems and tools
  • Utilize cloud-native technologies to create scalable event-driven architectures
  • Act as the SRE partner for analytics, data science, and product teams

Benefits

  • Comprehensive total rewards strategy
  • Opportunities for professional growth and high-performance rewards
  • Collaborative work environment in a health technology space
  • Engagement in impactful projects that shape industry data logistics
  • Supportive of multi-cloud and hybrid cloud environments
Full Job Description
What We're Looking For

We're looking for a Senior Site Reliability Engineer to join our Data & ML Platform team. You'll be at the forefront of building and operating a resilient, observable, and scalable platform that enables mission-critical data and ML workloads across our organization.

This role is ideal for someone who combines a strong SRE mindset with deep cloud infrastructure and data platform experience. You're comfortable operating at scale in a complex, hybrid cloud environment and can architect systems that balance velocity, safety, and cost. You'll work closely with Data & ML Engineers, Data Scientists, Analysts, and App Engineering teams to build a modern data platform that is secure, self-service, and production-grade.

What You Will Do
  • Operate and Improve Databricks and Snowflake: Own Databricks & Snowflake platforms lifecycle-including automation, workspace governance, job orchestration, and cost optimization.
  • Design for Reliability: Architect resilient, scalable, and secure infrastructure across cloud environments. Drive initiatives around failover, autoscaling, chaos testing, and capacity planning.
  • Advance Observability: Build and maintain platform-wide monitoring, alerting, and logging infrastructure using Datadog and other open tooling. Define and enforce SLOs/SLAs for critical services.
  • Drive CI/CD for Data & ML: Automate deployments of data pipelines, ML workflows, and infra components using GitHub Actions, Terraform, and related IaC tooling.
  • Enable Data Flow Across Platforms: Build patterns and tooling to support inter- and intra-cloud data movement across systems like Snowflake, S3, Delta Lake, and Kafka.
  • Champion Event-Driven Architectures: Leverage cloud-native tools like EventBridge, SNS/SQS, and Lambda to build loosely coupled, scalable data systems.
  • Collaborate Across Teams: Serve as the SRE and platform partner for teams across the organization, ensuring the platform meets the needs of analytics, data science, and product use cases.
  • Contribute to Strategy: Influence engineering-wide decisions on data platform architecture, ML enablement, and data product strategy.

What You Need to Succeed
  • 6+ years in SRE, platform engineering, or DevOps roles supporting data-intensive or ML-powered applications.
  • AI-native working style: daily use of Claude Code, Cursor, Copilot, or equivalent, with views on how they make a team faster.
  • Hands-on Databricks experience, including workspace setup, cluster/job management, and integration with CI/CD and data orchestration tools. Experience with Snowflake as well.
  • Deep understanding of cloud-native infrastructure on AWS (or similar), including VPCs, IAM, event-driven patterns, and serverless compute.
  • Proven expertise with observability tools (especially Datadog) and architecting platform-wide logging and monitoring solutions.
  • Strong command of CI/CD tooling, especially GitHub Actions, infrastructure-as-code (Terraform), and deployment automation for data systems.
  • Working knowledge in shell scripting and Python.
  • Experience building and supporting highly available, fault-tolerant systems.
  • Excellent communication and collaboration skills; able to work effectively across teams.

What Helps You Stand Out
  • DevSecOps mindset: Familiarity with implementing security best practices in IaC, CI/CD, secret management, and audit logging.
  • Experience with ML infrastructure tooling such as MLflow, Feature Stores, and GPU workload orchestration.
  • Strong experience in both Databricks and Snowflake in a large scale production lakehouse with cross-warehouse interoperability, e.g. Iceberg v3, Glue, etc.
  • Background in compliance-aware architecture (e.g., HIPAA, SOC 2) or regulated industries.
  • Familiarity with multi-cloud or hybrid cloud data environments; experience with Azure.
  • Contributions to open-source infrastructure, SRE, or observability tools.


At Datavant our total rewards strategy powers a high-growth, high-performance, health technology company that rewards our employees for transforming health care through creating industry-defining data logistics products and services.

The range posted is for a given job title, which can include multiple levels. Individual rates for the same job title may differ based on their level, responsibilities, skills, and experience for a specific job.

The estimated total cash compensation range for this role is:

$168,000-$200,000 USD

About Datavant

Datavant is a healthcare technology company that specializes in connecting and standardizing healthcare data from various sources. The company's products are used to improve patient care, accelerate drug development, and enhance clinical research. Datavant's platform, called the Datavant Platform, uses artificial intelligence and machine learning to identify and link patient data across different sources while maintaining patient privacy. The company was founded in 2017 by Travis May and is headquartered in San Francisco, California.
Learn more about Datavant
Size
100 employees
Industry
Founded
2017

Similar Jobs

More Jobs at Datavant

More Enterprise Technology Jobs

Find similar Senior Site Reliability Engineer jobs: