Legion

Director of Production Engineering

Legion$220K — $265K *
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • 8-12 years in DevOps, SRE, or production infrastructure roles with management experience
  • Hands-on experience with AWS production workloads including EKS and RDS
  • Experience managing security operations including vulnerability and incident response
  • 5+ years using observability tools like Datadog for reliability improvements
  • Proficient in Infrastructure-as-Code (Terraform, CloudFormation) and CI/CD automation
  • Skilled in programming languages like Go, Python, or Bash
  • Hands-on with Linux/Unix production systems
  • Proven ability to collaborate with cross-functional engineering and IT teams

Responsibilities

  • Build and lead a global DevOps/SRE engineering team, fostering a culture of ownership and collaboration
  • Own the reliability and infrastructure roadmap for AWS production, focusing on scalability and efficiency
  • Lead security operations ensuring proactive vulnerability management and incident response
  • Define engineering OKRs for infrastructure reliability and track measurable outcomes
  • Champion observability practices and automate alert triage to enhance response times
  • Leverage AI infrastructure knowledge in DevOps processes for efficiency gains
  • Drive Infrastructure-as-Code and CI/CD practices to boost engineering speed and reduce toil
  • Align on infrastructure standards and compliance requirements with engineering and IT teams

Benefits

  • $0 monthly premium for flexible medical, dental, and vision plans from day one
  • 401k plan offered
  • Discretionary Paid Time Off and Paid Holidays included
  • Parental Leave available
  • Equity options for employees
  • Monthly Wellness Reimbursement provided
  • Company-sponsored lunches
Full Job Description
Director of Production Engineering
Remote, United States

About this Position

Are you passionate about building the reliability, automation, and security foundations that let engineering teams move fast with confidence? At Legion, we are seeking a Director of Engineering, DevOps & SRE to lead the teams responsible for the availability, scalability, and security of our production environment. Our production infrastructure runs on AWS, leveraging services such as EKS, RDS, and a broad set of AWS-native technologies. You will partner closely with engineering and IT to build resilient systems, drive operational excellence, and ensure our platform meets the highest standards of security and compliance.

This is a hands-on leadership role where you'll spend ~20-30% of your time contributing directly to architecture, tooling, and incident response, and the rest driving vision, roadmap, and cross-team execution.

Responsibilities
  • Hire and build a globally-distributed DevOps/SRE engineering team. Recruit, mentor, and manage engineers, and foster a culture of ownership, collaboration, and continuous improvement.
  • Own the reliability and infrastructure roadmap for our AWS-based production environment, including EKS, RDS, and related AWS services, ensuring scalability, high availability, and cost efficiency.
  • Lead the organization's security operations (SecOps) practice, including vulnerability management, threat detection, incident response, and remediation, to proactively identify and resolve security issues before they impact customers.
  • Define and drive engineering OKRs for infrastructure reliability, automation, and security, and track progress against measurable outcomes.
  • Champion observability and alerting best practices (e.g., Datadog), including automating alert triage and response to reduce mean-time-to-resolution.
  • Solid understanding of agentic AI infrastructure and how AI agentic workflows apply to SDLC and DevOps processes (e.g., automated investigation, remediation, and PR-generation pipelines).
  • Drive Infrastructure-as-Code, CI/CD, and automation practices to increase engineering velocity and reduce operational toil.
  • Work closely with engineering and IT teams to align on infrastructure standards, access controls, tooling, and compliance requirements across the organization.
  • Ensure the platform meets the highest standards of security, compliance, and data protection; implement and maintain robust security controls and audit-readiness.
  • Lead and participate in the Incident Management on-call rotation, working with SRE and development teams to meet and exceed availability goals.
  • Stay current on cloud, DevOps, and security best practices, and provide technical guidance and thought leadership to the broader engineering organization.

Required Qualifications
  • 8-12 years of experience in DevOps, Site Reliability Engineering, or production infrastructure roles, including people management experience.
  • Deep hands-on experience running production workloads on AWS, including EKS (Kubernetes), RDS, and other core AWS services (e.g., VPC, IAM, Lambda, S3).
  • Demonstrated experience running security operations (SecOps) - vulnerability management, incident response, and remediation of production security issues.
  • 5+ years of experience leveraging observability platforms (e.g., Datadog, Prometheus, Grafana) to drive reliability, performance, and alerting improvements.
  • Strong experience with Infrastructure-as-Code (e.g., Terraform, CloudFormation) and CI/CD automation.
  • Proficiency in at least one of Go, Python, or Bash, with day-to-day use of Git and test automation pipelines.
  • Hands-on experience operating Linux/Unix production platforms (Amazon Linux, Ubuntu, RHEL/CentOS).
  • Proven track record partnering cross-functionally with engineering and IT teams to align on infrastructure, tooling, and security standards.
  • Demonstrated experience leading incident management and on-call practices for high-availability production systems.
  • Bachelor's degree in Computer Science, Engineering, or related field required; Master's degree preferred.

Preferred Qualifications
  • Experience with major cloud providers beyond AWS, such as Google Cloud Platform or Oracle Cloud Infrastructure (OCI).
  • Relevant security certifications (e.g., CISSP, AWS Security Specialty, CKS).
  • Experience with compliance frameworks such as SOC 2, ISO 27001, or HIPAA.
  • 3+ years of experience with Kubernetes or other containerization/orchestration platforms at scale.
  • Experience with Kubernetes-native delivery tooling, including Argo Workflows and Helm.
  • 5+ years of experience leading teams in an agile/scrum environment.
  • Experience building or scaling automated investigation and remediation pipelines for production error classes.

COMPENSATION & BENEFITS

Salary Range: Base Salary Range $220,000 - $265,000 + Bonus + Stock Equity

At Legion, we offer competitive compensation and benefits packages to all employees. As a fully remote employer, pay for positions is determined using local, national, and industry-specific survey data.

Our posted salary range is done so in good faith based on national data and may be refined for a candidate's region/town/cost of living. We strive to make competitive offers that allow employees room for future growth. Salaries will be based on the applicant's location, level of experience, education, and specialized knowledge and skills. Additionally, we consider the external market rate, the amount we have budgeted internally, and the internal equity for the same position within the company.

Benefits include, but are not limited to:
  • $0 monthly premium and other flexible medical, dental, and vision plans effective on the first day of employment
  • 401k plan
  • Discretionary Paid Time Off and Paid Holidays
  • Parental Leave
  • Equity
  • Monthly Wellness Reimbursement
  • Monthly Lunch on Legion

About Legion

Legion is a software company that provides workforce management solutions for businesses in the retail, hospitality, and healthcare industries. The company's products are designed to help businesses optimize their workforce by predicting demand, scheduling employees, and managing labor costs. Legion's clients include major brands such as Sephora, Bloomingdale's, and Levi's. The company was founded in 2015 and is headquartered in San Francisco, CA.
Learn more about Legion
Size
100 employees
Industry
Founded
2016

Similar Jobs

More Jobs at Legion

More Information Technology Jobs

Find similar Director of Production Engineering jobs: