Bachelor's degree in a technical field or equivalent experience
5+ years of experience in SRE, DevOps, or platform engineering
Deep hands-on experience with Apache Airflow, including distributed executor configurations
Experience operating enterprise job scheduling platforms like Automic/UC4
Strong knowledge of Linux and Windows systems, with cloud environment experience (AWS preferred)
Proficiency in Python for automation, and familiarity with shell scripting
Experience with container orchestration and CI/CD pipelines
Understanding of observability principles and related tools like ELK and Grafana
Ability to communicate clearly under pressure and drive incidents to resolution
Bias toward automation and minimizing repetitive tasks
Responsibilities
Serve as a primary escalation point for production support involving Airflow and UC4
Own and continuously improve SLOs, SLIs, and error budgets for orchestration platforms
Monitor platform health, capacity, and performance, proactively identifying and remediating issues
Partner with data engineering and application teams to troubleshoot various issues
Manage patching, upgrades, and configuration management for Airflow and UC4 environments
Collaborate with security to harden platform configurations
Contribute to on-call rotations and maintain runbooks and escalation procedures
Design and build tooling to improve developer experience and reduce toil
Lead platform modernization initiatives, such as migrating workloads and improving deployment pipelines
Develop and maintain infrastructure-as-code for platform components
Build observability solutions to enhance visibility into workflows
Participate in design reviews and contribute to platform roadmap
Benefits
Hybrid work model
Opportunity to work on cutting-edge technologies
Focus on continuous improvement and operational excellence
Collaborative work environment
Chance to influence platform modernization initiatives
Full Job Description
Job Description:
Aboutthe Role: We arelooking for a Senior SRE to join our Platform Engineering team whereyoullown the reliability, scalability, and operational excellenceof our workflow orchestration platforms 6primarily Apache Airflow and BroadcomAutomic/UC4.This is a hybrid role:roughly halfyour time will be spent onsteady stateoperationsand incident response, and the other half on engineering projects that meaningfully improve the platforms you support.
This is a role for someone who is genuinely motivated by the pursuit of excellence 6 not justsustaining whatworks butrelentlessly refining it.You care deeply about the craft of reliability and find satisfaction in the distance traveled between good and great, whether that means tightening observability,reducing toil, or building something that makes support activities easier to handle.
WhatYoullWork On: Operations & Reliability ( 50%)
Serve asa primary escalation point for production supportinvolvingAirflow and UC4 6assistingend-userinquiries,incident root cause analysis, andimplementing go-forward solutions
Own and continuously improve SLOs, SLIs, and error budgets for orchestration platforms
Monitor platform health, capacity, and performance; proactivelyidentifyand remediate issues before theyimpactusers
Partner with data engineering and application teamsto troubleshoot DAG failures,job dependencies, and scheduling issues
Managepatching,upgrades, and configuration management for Airflowand UC4 environments
Collaborate with security to harden platform configurations and manage software vulnerabilities
Contribute to on-call rotations andmaintainrunbooks and escalation procedures
Platform Engineering ( 50%)
Design and build tooling and automation to reduce toil and improve developer experience for teams that depend onAirflow and UC4 (among other platforms the teamoffers)
Lead or contribute to platform modernizationinitiatives 6 e.g., migrating workloads, improving deployment pipelines,containerizing components, or adopting managed service offerings
Develop andmaintaininfrastructure-as-code (Terraform, Helm, Ansible, etc.) for platform components
Build observability solutions(e.g.,dashboards, alerting, log aggregation) that give teams better visibility into their workflows
Build and enforce standards aroundplatform usethat help engineering teams adopt best practices at scale
Participate indesignreviews and contribute to the overall platform roadmap
WhatWereLooking For:
Bachelors degree in a technical field or equivalent practical experience
5+ yearsof experience in SRE, DevOps, or platform engineering roles
Deep hands-on experience with Apache Airflow 6ideally including distributed executor configurations (Celery or Kubernetes), DAG authoring best practices,and multi-environment deployments