Vultr is seeking a highly skilled Senior Site Reliability Engineer, Databases to join our Platform Engineering team. You will be the reliability and operational backbone for Vultr's database infrastructure spanning MySQL InnoDB Clusters, PostgreSQL, and other database technologies. Working alongside our Senior Platform Engineer (Databases), you will own database monitoring, backup verification, disaster recovery testing, on-call incident response, replication health, access management, and compliance remediation. This is a peer-level role where you will apply SRE methodology - error budgets, runbooks, automation-first thinking, and toil reduction - specifically to database systems that power a global cloud platform serving millions of customers. You bring deep operational database expertise to complement our existing schema and DDL strengths, and you are comfortable writing PHP, Python, or Go to automate and instrument everything you build.
Key Responsibilities- Own and evolve comprehensive monitoring and alerting for MySQL InnoDB Clusters, and PostgreSQL infrastructure across multiple datacenters
- Establish and execute a quarterly backup verification and disaster recovery testing program across all database systems
- Serve as first-tier on-call responder for database incidents, executing documented runbooks and escalating to senior engineering when architecture-level decisions are required
- Monitor and maintain replication health across databases
- Own database user lifecycle management - provisioning, deprovisioning, access audits, and role-based access control across all database systems
- Execute and track security compliance remediation including pen-test findings, GRC audit requirements, and encryption-at-rest verification
- Manage operational health of the Debezium/Kafka Connect data pipeline in coordination with the Kafka infrastructure team
- Build and maintain Puppet profiles for database infrastructure configuration management and write PHP, Python, or Go automation tooling to reduce operational toil
- Develop and maintain runbooks, operational documentation, and disaster recovery procedures for all database systems
- Partner with the Senior Platform Engineer (Databases) as a peer - reviewing each other's work, sharing on-call, and splitting ownership of database reliability across production systems
Qualifications- 7+ years of experience in Site Reliability Engineering, DevOps, or Database Operations roles in production environments at scale
- Deep operational expertise with MySQL in production - InnoDB Cluster, Group Replication, MySQL Router, and ProxySQL - with strong troubleshooting skills for replication, performance, and reliability issues
- Production experience with PostgreSQL - replication, high availability, performance tuning, and operational management
- Strong proficiency with configuration management tools (Puppet preferred) and infrastructure-as-code practices
- Experience with database backup tools (xtrabackup, pg_dump, mysqldump) and disaster recovery procedures
- Proficiency in PHP and Python or Go for automation, tooling, and integration with existing codebases
- Experience with database observability - Prometheus exporters, Grafana dashboards, alerting frameworks, and SLO/error-budget methodology
- Familiarity with Kafka Connect, Debezium, or similar change-data-capture pipelines
- Strong incident response skills with experience in on-call rotations, including post-incident review and remediation
- Excellent communication skills and ability to collaborate across engineering teams as a senior peer
Compensation$125,000 - $135,000
We are currently accepting applications from candidates residing in the following states: Alabama, Arizona, Colorado, Connecticut, Florida, Georgia, Idaho, Illinois, Indiana, Iowa, Kentucky, Louisiana, Maryland, Massachusetts, Michigan, Minnesota, Missouri, Montana, Nebraska, Nevada, New Jersey, New Mexico, New York, North Carolina, Ohio, Oklahoma, Pennsylvania, Rhode Island, South Carolina, Tennessee, Texas, Utah, Vermont, Virginia, Wisconsin.