What We're Looking ForWe're looking for a
Senior Site Reliability Engineer to join our Data & ML Platform team. You'll be at the forefront of building and operating a resilient, observable, and scalable platform that enables mission-critical data and ML workloads across our organization.
This role is ideal for someone who combines a strong
SRE mindset with deep
cloud infrastructure and data platform experience. You're comfortable operating at scale in a complex, hybrid cloud environment and can architect systems that balance velocity, safety, and cost. You'll work closely with Data & ML Engineers, Data Scientists, Analysts, and App Engineering teams to build a modern data platform that is secure, self-service, and production-grade.
What You Will Do- Operate and Improve Databricks and Snowflake: Own Databricks & Snowflake platforms lifecycle-including automation, workspace governance, job orchestration, and cost optimization.
- Design for Reliability: Architect resilient, scalable, and secure infrastructure across cloud environments. Drive initiatives around failover, autoscaling, chaos testing, and capacity planning.
- Advance Observability: Build and maintain platform-wide monitoring, alerting, and logging infrastructure using Datadog and other open tooling. Define and enforce SLOs/SLAs for critical services.
- Drive CI/CD for Data & ML: Automate deployments of data pipelines, ML workflows, and infra components using GitHub Actions, Terraform, and related IaC tooling.
- Enable Data Flow Across Platforms: Build patterns and tooling to support inter- and intra-cloud data movement across systems like Snowflake, S3, Delta Lake, and Kafka.
- Champion Event-Driven Architectures: Leverage cloud-native tools like EventBridge, SNS/SQS, and Lambda to build loosely coupled, scalable data systems.
- Collaborate Across Teams: Serve as the SRE and platform partner for teams across the organization, ensuring the platform meets the needs of analytics, data science, and product use cases.
- Contribute to Strategy: Influence engineering-wide decisions on data platform architecture, ML enablement, and data product strategy.
What You Need to Succeed- 6+ years in SRE, platform engineering, or DevOps roles supporting data-intensive or ML-powered applications.
- AI-native working style: daily use of Claude Code, Cursor, Copilot, or equivalent, with views on how they make a team faster.
- Hands-on Databricks experience, including workspace setup, cluster/job management, and integration with CI/CD and data orchestration tools. Experience with Snowflake as well.
- Deep understanding of cloud-native infrastructure on AWS (or similar), including VPCs, IAM, event-driven patterns, and serverless compute.
- Proven expertise with observability tools (especially Datadog) and architecting platform-wide logging and monitoring solutions.
- Strong command of CI/CD tooling, especially GitHub Actions, infrastructure-as-code (Terraform), and deployment automation for data systems.
- Working knowledge in shell scripting and Python.
- Experience building and supporting highly available, fault-tolerant systems.
- Excellent communication and collaboration skills; able to work effectively across teams.
What Helps You Stand Out- DevSecOps mindset: Familiarity with implementing security best practices in IaC, CI/CD, secret management, and audit logging.
- Experience with ML infrastructure tooling such as MLflow, Feature Stores, and GPU workload orchestration.
- Strong experience in both Databricks and Snowflake in a large scale production lakehouse with cross-warehouse interoperability, e.g. Iceberg v3, Glue, etc.
- Background in compliance-aware architecture (e.g., HIPAA, SOC 2) or regulated industries.
- Familiarity with multi-cloud or hybrid cloud data environments; experience with Azure.
- Contributions to open-source infrastructure, SRE, or observability tools.
At Datavant our total rewards strategy powers a high-growth, high-performance, health technology company that rewards our employees for transforming health care through creating industry-defining data logistics products and services.
The range posted is for a given job title, which can include multiple levels. Individual rates for the same job title may differ based on their level, responsibilities, skills, and experience for a specific job.
The estimated total cash compensation range for this role is:
$168,000-$200,000 USD