5+ years in Production Engineering, SRE, or related fields
Experience with holistic monitoring of complex systems
Proficient in PostgreSQL or similar databases
Background in incident management and on-call rotations
Strong programming skills in Python, Go, Java, or Scala
Familiarity with Infrastructure as Code and automation
Experience using AI for operational challenges
Responsibilities
Build advanced monitoring and detection capabilities for early issue warnings
Develop reliable automation for debugging and incident mitigation
Enhance service reliability, scalability, and security
Create dependable solutions for production issues
Lead incident response and perform root-cause analysis
Collaborate with cross-functional teams on follow-up actions
Benefits
Comprehensive benefits and perks tailored to employee needs
Opportunities for professional development and growth
Flexible work arrangements
Supportive work culture focused on innovation
Access to cutting-edge technology and tools
Full Job Description
CSQ127R154
As a Senior Production Engineer working on Databricks' realtime products you will be directly contributing to our customers' success. You will build advanced monitoring and incident mitigation tooling, and drive changes across the stack to proactively make the realtime products like Lakebase, Neon and Model Serving reliable, secure, and scalable in production.
Our production engineers understand the Databricks platform from end to end, and partner across engineering teams, building durable solutions for a platform operating across AWS, Azure, and GCP. The Impact You'll Have
Observability
Build advanced monitoring and detection capabilities that give early warning of issues and deep insights into workload performance and customer experience.
Build reliable, observable automation for debugging and incident mitigation.
Reliability
Improve service reliability, scalability, security, and operational efficiency.
Develop dependable, safe, mitigations for production issues.
On-Call & Incident Response
Participate in a follow-the-sun on-call rotation and lead incident response and mitigation.
Perform root-cause analysis, identify and implement lasting corrective actions.
Partner with Product Engineering, Security, Support, and other infrastructure teams on follow up actions.
What We Look For
Experience
5+ years of experience in Production Engineering, Customer Reliability Engineering (CRE), Site Reliability Engineering (SRE), infrastructure engineering, backend software engineering, or a related field.
Experience in holistic monitoring and alerting of complex stateful systems, applying multiple strategies like workload alerting, anomaly detection and probing.
Experience with PostgreSQL, or related managed databases or distributed systems.
Experience in incident management and participating in on-call rotations for critical infrastructure.
Skillset
A mindset focused on automation, root-cause resolution, and continuous improvement.
Strong programming skills in one or more languages such as Python, Go, Java, Scala, or similar.
Proficiency with infrastructure automation and Infrastructure as Code.
Ability to work across system boundaries and collaborate effectively during complex incidents, and engage with customers' infrastructure teams.
Experience in using AI to address production and operational challenges.
Bonus
Experience with AWS, Azure, or GCP.
Experience with Lakebase or Neon.
Experience with Kubernetes, Terraform.
Experience building internal platforms, operational tooling, or developer productivity systems.
Education
BS degree (or higher) in Computer Science, or a related field.
BenefitsAt Databricks, we strive to provide comprehensive benefits and perks that meet the needs of all of our employees. For specific details on the benefits offered in your region click here.
About Databricks
Databricks is a unified analytics platform that provides data engineering, collaborative data science, and machine learning capabilities. The company was founded in 2013 by the original creators of Apache Spark, a popular open-source big data processing engine. Databricks provides a cloud-based platform that allows data teams to collaborate and build data pipelines, run machine learning models, and perform advanced analytics. The company has raised over $1 billion in funding and is valued at $38 billion as of November 2021.