Director of Platform Engineering

First Due

$240K *
US-AnywhereRemote in United States
Enterprise Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • CI/CD and deployment-pipeline modernization at scale
  • Formal SRE practice including SLOs, observability/alerting tooling, and incident management
  • Experience in designing on-call/incident-management programs and enforcing RCA practices
  • Background in database reliability/operations for high-traffic production databases
  • Experience in vertical SaaS or regulated environments, with a preference for public safety
  • Ability to communicate reliability and operational maturity to executive leadership in business terms
  • Hands-on experience with deployment processes and incident management tools

Responsibilities

  • Assess and centralize CI/CD and deployment practices company-wide
  • Contribute hands-on to containerizing and automating deployments
  • Establish production reliability standards, including SLOs/SLIs, and error budgets
  • Manage on-call strategies and escalation policies in collaboration with multiple teams
  • Define and enforce RCA and post-incident review practices using AI
  • Ensure operational reliability of production databases and define schema-change practices
  • Implement operational metrics for system health and build reporting dashboards

Benefits

  • Comprehensive compensation and benefits package
  • Medical, dental, and vision coverage
  • Flexible PTO and a fully remote workplace
  • Technology stipend and opportunities for advancement
Full Job Description
Job Title: Director of Platform Engineering

Location: Remote - US Only
Country: United States
Department: Engineering
Reports To: SVP, Engineering
Position Type: Full-Time
Salary: $240,000
Job Summary:

This is a new, dedicated leadership role reporting to the SVP of Engineering, accountable for turning First Due's DevOps, SRE, database reliability, incident-management, and developer-experience practices from distributed and ad hoc into centralized and disciplined. The right leader brings genuine hands-on technical depth to build and run this function, and the executive presence to represent platform reliability and operational maturity to the ELT.
Key Responsibilities:
DevOps & CI/CD

  • Assesses and centralizes CI/CD and deployment practices across product lines; owns the rollout
  • Directly contributes where useful - e.g., containerizing and automating deployments, not just directing others
  • Improves deployment frequency, predictability, environment consistency, and release safety

Site Reliability Engineering & Observability
  • Stands up SLOs/SLIs, error budgets, and production reliability standards
  • Owns monitoring, alerting, and observability patterns/dashboards, standardized across teams
  • Partners with development teams on consistent instrumentation rather than owning their code

Incident Management & Operational Discipline
  • Owns on-call tooling, rotation design, and escalation policy company-wide, partnering with Support, Product, QA, and Development to ensure comprehensive escalation resolutions
  • Defines and enforces RCA and post-incident review discipline, using AI to speed investigation and information-gathering
  • Builds these as adopted practices across teams that don't report to this role - requires real influence without authority

Database Reliability Engineering (DBRE)
  • Owns operational reliability of First Due's core production database(s): uptime, query performance, capacity
  • Defines schema-change practice and review discipline for a monolithic, high-traffic database
  • Partners with Infrastructure on backup, continuity, and BCDR runbooks, and is accountable for testing them

Engineering Operations & Metrics
  • Defines and implements the operational metrics set (deployment frequency, MTTR, system health, uptime), distinct from delivery/velocity metrics
  • Builds the dashboards and reporting giving engineering and executive leadership real visibility into system health
  • Formalizes the AI natural-language query layer already in use against observability data into a shared standard for "how we observe system health"

Developer Experience & Shared Frontend Platform
  • Partners with the owner of shared component libraries, frontend build tooling, and design-system/API contracts used by product teams - not feature UI work
  • Sponsors (without managing) a dotted-line Frontend Guild that drives convention adoption across module teams

Cross-Team Architecture Influence
  • Looks across product groups to spot shared structural patterns and drift - e.g., business logic embedded in frontend code that belongs in per-module backend APIs - and drives the fix
  • Influences cross-cutting platform architecture; does not own product/feature architecture or act as a Chief Architect

Team Building
  • Builds this function's structure and hiring plan from the ground up, making the business case for each hire as the need is proven - headcount is not prescribed in advance
  • Establishes the team's operating rhythm (on-call, incident reviews, roadmap planning)
Qualifications:
  • CI/CD and deployment-pipeline modernization at scale
  • Formal SRE practice: SLOs, observability/alerting tooling, incident management
  • On-call/incident-management program design (e.g., PagerDuty or equivalent) and RCA practice enforcement
  • Database reliability/operations experience (schema governance, query optimization, backup/BCDR) for a high-traffic production database, or has directly managed someone who does
  • Vertical SaaS, regulated, or mission-critical software is a plus; public safety experience is an advantage
  • Executive presence: can talk to the ELT about reliability, risk, and operational maturity in business terms
  • Genuinely hands-on - willing to be in the pipeline, the dashboard, or the incident channel


Physical Demands and Work EnvironmentThis role is fully remote with minimal travel expectations at this time. Reasonable accommodation may be made to enable qualified employees and applicants to perform the essential functions as outlined above. If you require an accommodation during the interview process, please reach out to [redacted].

Working at First DueFirst Due offers a comprehensive compensation and benefits package for eligible employees, including competitive pay, medical, dental, and vision coverage, FSA/HSA, 401(k), flexible PTO, a fully remote workplace, a technology stipend, opportunities for advancement, and other benefits and perks that sets our team apart. Visit www.firstdue.com to learn more.

If you are a resident of a state requiring wage transparency, please reach out to [redacted] for a reasonable estimate of annual base compensation and any eligible incentive compensation. The actual compensation offered to successful candidates for roles may be higher or lower, based on non-discriminatory criteria including but not limited to relevant professional experience, geographic location, knowledge, skills, and abilities. This range will be reviewed on a regular basis.

Similar Jobs

More Jobs at First Due

More Enterprise Technology Jobs

Find similar Director of Platform Engineering jobs: