Platform Delivery & Reliability Engineer (Remote) - 29337

Enlighten

$119K — $235K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Must obtain and maintain a U.S. Government Security Clearance, U.S. Citizenship required.
  • 9 years of relevant experience with a Bachelor's or 7 years with a Master's; or 13 years with a High School Diploma.
  • Deep experience with deploying and debugging production Kubernetes clusters.
  • Proven ability to deliver complex distributed systems into production across various environments.
  • Exceptional troubleshooting skills across Linux, networks, containers, and applications.
  • Familiarity with infrastructure as code (e.g., Terraform) and cloud environments (AWS, Azure, GCP).
  • Experience with CI/CD pipelines and programming/scripting skills (e.g., Go, Python, Bash).

Responsibilities

  • Plan and execute deployments and upgrades for the data lakehouse platform across 50 Kubernetes clusters.
  • Debug critical issues throughout the stack and drive them to root cause.
  • Contribute patches and automation improvements to the platform.
  • Catalog and communicate field issues to relevant teams, maintaining a knowledge base.
  • Improve site reliability for deployed environments, ensuring better monitoring and incident response.
  • Identify and eliminate friction in deployment processes through better tooling and communication.
  • Mentor engineers on debugging, deployment, and operational best practices.

Benefits

  • Flexible work location: fully remote or hybrid based on proximity to offices.
  • Opportunity to work on large-scale government projects.
  • Collaboration with cross-functional teams and stakeholders.
  • Focus on continuous improvement and operational excellence in deployments.
Full Job Description
Job Description

Enlighten is looking for a Senior Platform Delivery & Reliability Engineer to own the rollout of our data lakehouse platform across a large, multi-site government enterprise, currently ~50 production Kubernetes clusters and growing. This role is a rare hybrid of platform engineer, SRE, and delivery lead. You can deploy the platform, debug anything you encounter in the field, feed what you learn back to the engineering teams, and help fix underlying issues in the code base. Just as importantly, you can step back from any individual issue and fix the system that produced it by building the processes, tooling, and communication channels inside Enlighten that make every deployment faster and less painful than the one before it.

You will be a full member of the Infrastructure team, working daily with our Ingest, Query, Application and Testing teams as well as government customers and site personnel. Success in this role looks like: rollouts across the enterprise happen predictably and efficiently, issues found in the field are cataloged, communicated, and resolved quickly, and the friction that slows deployments steadily disappears. #LI-DS1 #Senior Level

Essential Job Responsibilities

  • Plan, coordinate, and execute deployments and upgrades of the data lakehouse platform across 50 production Kubernetes clusters in customer environments.
  • Debug and troubleshoot critical issues anywhere in the stack (infrastructure, Kubernetes, platform services, data services, and applications) and drive them to root cause.
  • Contribute patches, configuration changes, and automation improvements directly back to the platform.
  • Catalog and triage issues discovered in the field, communicate them clearly to the Infrastructure, Ingest, Query, and Application teams, and maintain a living knowledge base of failure modes, fixes, and runbooks.
  • Enable and improve site reliability for fielded environments: monitoring, alerting, incident response, and continuous reliability improvement.
  • Identify friction and dysfunction in how deployments happen, such as unclear handoffs, communication gaps, and repeated manual work; then design, implement, and institutionalize the processes that eliminate them (release checklists, readiness reviews, escalation paths, cross-team communication cadences).
  • Continuously improve deployment tooling and automation so rollouts become faster, safer, and more repeatable.
  • Coordinate with a large set of stakeholders (the engineering teams, government programs, security, and site personnel) and keep them informed.
  • Mentor engineers, both junior and senior, on debugging, deployment, and operational excellence.
  • Other duties as assigned.


Minimum Qualifications

  • Clearance Requirement: Must obtain and maintain a U.S. Government Security Clearance, but not required on day one; U.S. Citizenship required.
  • 9 years relevant experience with Bachelors in related field; 7 years relevant experience with Masters in related field; or High School Diploma or equivalent and 13 years relevant experience.
  • Deep, hands-on experience deploying, operating, and debugging production Kubernetes clusters and their ecosystem (networking, storage, service meshes, observability, volume management).
  • Proven record of delivering complex distributed systems into production across many environments or sites: not just building platforms, but landing them with customers.
  • Elite troubleshooting and analytical skills across the full stack: Linux systems, hosts, networks, security, containers, and application services.
  • Experience with infrastructure as code (e.g., Terraform) and modern cloud environments (e.g., AWS, Azure, GCP).
  • Experience with CI/CD pipelines (e.g., GitLab CI) and proficiency in scripting or programming (e.g., Go, Python, Bash).
  • Working knowledge of SRE practices: monitoring and alerting, incident management, blameless postmortems, and runbook development.
  • Demonstrated experience creating or improving engineering and delivery processes that other teams actually adopted; you can point to a workflow that exists because you built it.
  • Excellent verbal and written communication skills; able to translate deep technical issues for engineers, leadership, and customers, and comfortable coordinating a large number of people across organizational boundaries.
  • Work Location: *Remote or Hybrid. This role is fully remote unless you are located near one of our offices in Columbia, MD; San Antonio, TX; Boise, ID; Greenville, SC; or Augusta, GA, where a hybrid schedule applies. Note: Work models are subject to change based on business needs.


Preferred Requirements

  • Experience deploying or operating large-scale data platforms and lakehouse technologies (e.g., Spark, Trino/Presto, Kafka, NiFi, object storage, Iceberg/Delta/Hudi)
  • Experience delivering into DoD, IC, or other federal environments, including STIG-hardened, disconnected, or air-gapped deployments and familiarity with the RMF/ATO process
  • Experience with Kubernetes Operators/Controllers development
  • Prior release management, delivery lead, field engineering, or deployment engineering experience on a multi-team program
  • Understanding of agile software development methodologies and use of standard software development tool suites (e.g., YouTrack, GitLab, Nexus)
  • DoD 8140 / 8570 compliance certifications may be required in this position as directed by the customer


We have many more additional great benefits/perks that you can find on our website at www.enlighten.com .

Similar Jobs

More Jobs at Enlighten

More Information Technology Jobs

Find similar Platform Delivery & Reliability Engineer (Remote) - 29337 jobs: