About the Data Platform TeamOur mission is to deliver reliable, secure, and scalable data infrastructure as a service across AppDirect. We treat data infrastructure as a product-databases, search, cache, and lakehouse capabilities that engineering teams can depend on to build and ship with confidence.
The Data Platform team supports 80+ data stores across the organization and operates against four quality pillars: resiliency and capacity, security, operational excellence, and performance and scalability. We enable self-serve access and automation, lead large-scale database migrations driven by company growth and acquisitions, and partner closely with Cloud Platform, DevOps, and product engineering so teams get fit-for-purpose data stores that are secure by design.
As data infrastructure enters its AI-native era, we are investing in AI-assisted operations, self-service tooling, and lakehouse modernization-reducing toil, improving time-to-insight, and helping newly acquired teams onboard onto our data platform quickly and safely.
About YouWe are looking for a Senior Database Reliability Engineer to join our Data Platform team. You are a highly technical professional with an entrepreneurial spirit and a data-as-a-service mindset. You thrive on efficiency, automation (including AI-driven approaches), and continuous learning-and you want your work to make a meaningful impact.
This is a unique opportunity to build at the intersection of highly scalable, fault-tolerant data services and modern cloud infrastructure. You will take on complex technical challenges across database reliability, migrations, and platform automation, and play a key role in evolving our cloud-native data environment. You enjoy guiding engineers toward the right data store for the job, building self-healing systems, and leaving behind documentation and patterns that help others move faster.
What you'll do and how you'll have an impact- Lead large-scale database migration projects driven by company acquisitions, client onboarding, and platform growth-planning for minimal downtime, strong rollback strategies, and zero data loss.
- Design and execute deployment strategies for blue/green rollouts, data archiving, and partitioning across relational, key-value, and document data stores.
- Use AI-powered automation to strengthen our data platform across our quality pillars: operational excellence, security and compliance, performance and scalability, and resiliency and capacity planning.
- Apply AI-assisted development tools and spec-driven workflows to design, automate, and ship database platform changes faster-Terraform modules, runbooks, diagnostics, and operational tooling-with clear acceptance criteria and high code quality.
- Partner with the Data Insights team on lakehouse enablement (e.g., Databricks, Apache Iceberg, Unity Catalog)-bridging transactional data stores and analytics infrastructure through reliable access patterns, security/governance alignment, and safe onboarding of workloads onto the lakehouse.
- Guide engineers in delivering fit-for-purpose database solutions across relational (MySQL, PostgreSQL, SQL Server), key-value (Redis), document (MongoDB), and search (Elasticsearch).
- Maintain a strong culture of documentation by developing and sharing high-quality knowledge resources for engineers and AI-assisted workflows.
- Participate in incident response and on-call rotation to maintain data infrastructure health.
What we're looking for- Strong understanding of AI-assisted development workflows, with proven hands-on experience using tools such as Cursor, Claude, to improve efficiency, automation, and code quality.
- 5+ years of hands-on experience running and operating MySQL and/or PostgreSQL in large-scale production environments, with deep knowledge of database internals
- 2+ years of hands-on experience with NoSQL-one of MongoDB, Redis, or Elasticsearch.
- 2+ years of experience with AWS cloud services.
- Strong expertise in Infrastructure as Code (IaC), with a primary focus on Terraform.
- Strong preferred exposure to lakehouse and analytics infrastructure (e.g., Databricks, Apache Iceberg, Unity Catalog) with an interest in collaborating with Data Insights on lakehouse enablement.
- Excellent analytical and problem-solving abilities.
- Strong communication skills for explaining complex technical topics to peers and non-technical stakeholders.
- Experience working in a distributed team.