DraftKings

Lead Site Reliability Engineer

DraftKings$148K — $185K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor’s Degree in Computer Science or related field, or equivalent experience.
  • 7+ years experience in Site Reliability Engineering with hands-on SLO operationalization.
  • Deep experience with observability tools like Datadog for metrics and logging.
  • Experience linking infrastructure metrics to application performance and user experience.
  • Strong understanding of distributed systems and their reliability challenges.
  • Proven track record of influencing engineering teams and aligning cross-functional efforts.
  • Excellent communication skills for presenting reporting to senior stakeholders.
  • Familiarity with cloud infrastructures such as AWS and Kubernetes.

Responsibilities

  • Lead the development of Service Level Objectives (SLOs) for critical services.
  • Collaborate with engineering teams to design SLOs based on user journeys.
  • Refine reliability objectives to match customer impact and business needs.
  • Connect reliability metrics to platform experiences to illustrate dependencies.
  • Create reporting tools for visibility into reliability across components.
  • Influence reliability practices through mentoring and technical guidance.
  • Apply strategic measurement techniques to differentiate service degradation from outages.

Benefits

  • Opportunity to shape reliability practices within a technology-focused company.
  • Guidance provided for necessary gaming licensing related to the role.
  • Access to equity and bonus opportunities as part of total compensation.
  • Full-time position with potential for career growth in a public tech company.
Full Job Description
The Crown Is Yours

As a Lead Site Reliability Engineer, you’ll set the reliability standard across our Infrastructure Engineering organization. You’ll define how we measure reliability for critical services, partnering with engineering teams to build and refine Service Level Objectives that connect infrastructure performance to the experiences our platforms support. You’ll turn complex telemetry into clear, actionable insights that help teams and senior leaders make better decisions about reliability, risk, and priorities. As an individual contributor, you’ll lead through technical expertise and influence, shaping a consistent reliability practice across the organization.

What you’ll do as a Lead Site Reliability Engineer
  • Lead and mature the Service Level Objective development process across Infrastructure Engineering, establishing clear frameworks and standards for setting meaningful reliability targets.

  • Partner with engineering teams to design and implement Service Level Objectives, beginning with the critical user journeys each service supports and translating them into measurable indicators, targets, and error budget policies.

  • Review and refine existing reliability objectives to keep them aligned with changing customer impact, technical dependencies, and business priorities.

  • Connect infrastructure reliability targets to the application and platform experiences they support, making dependencies and their impact on end-user experience clear and measurable.

  • Build reporting processes and tooling that provide a clear view of reliability across critical components, translating technical signals into actionable insights for Senior Managers and Directors.

  • Influence reliability practices across teams through technical guidance, design reviews, mentoring, and collaboration with engineering partners.

  • Help teams distinguish meaningful service degradation from true downtime by applying thoughtful measurement strategies to complex distributed systems.

What you’ll bring
  • A Bachelor’s Degree in Computer Science or a related field, or equivalent relevant education, experience, and training.

  • At least 7 years of experience in Site Reliability Engineering, including hands-on experience defining and operationalizing Service Level Objectives, Service Level Indicators, and error budgets at scale.

  • Deep experience with observability platforms such as Datadog, including building dashboards, monitors, and reporting from metrics and logging pipelines.

  • Experience connecting infrastructure-level reliability objectives to application or platform-level outcomes and evaluating how technical dependencies affect end-user experience.

  • Strong knowledge of distributed systems and the failure modes that can make reliability measurement complex, with the ability to assess what technical signals truly represent.

  • Proven ability to influence across engineering teams, translate reliability concepts for technical and non-technical audiences, and drive alignment without direct authority.

  • Excellent written and verbal communication skills, including experience developing and presenting reliability reporting to Senior Managers, Directors, and cross-functional stakeholders.

  • Working knowledge of cloud and infrastructure environments such as Amazon Web Services, Kubernetes, and on-premise systems, with the technical depth to partner effectively with the teams operating them.

Join Our Team

We’re a publicly traded (NASDAQ: DKNG) technology company headquartered in Boston. As a regulated gaming company, you may be required to obtain a gaming license issued by the appropriate state agency as a condition of employment. Don’t worry, we’ll guide you through the process if this is relevant to your role.

The US base salary range for this full-time position is 148,000.00 USD - 185,000.00 USD, plus bonus, equity, and benefits as applicable. Our ranges are determined by role, level, and location. The compensation information displayed on each job posting reflects the range for new hire pay rates for the position across all US locations. Within the range, individual pay is determined by work location and additional factors, including job-related skills, experience, and relevant education or training. Your recruiter can share more about the specific pay range and how that was determined during the hiring process.

About DraftKings

DraftKings is an American daily fantasy sports contest and sports betting operator. The company allows users to enter daily and weekly fantasy sports–related contests and win money based on individual player performances in five major American sports, Premier League and UEFA Champions League soccer, NASCAR auto racing, Canadian Football League, the XFL, mixed martial arts and Tennis. In August 2018, DraftKings launched DraftKings Sportsbook in New Jersey becoming the first legal mobile sports betting operator in the state. Since launching in New Jersey, DraftKings has opened mobile sports betting operations in Indiana, Pennsylvania, West Virginia and opened in New Hampshire December 30, 2019 after reaching contract with the New Hampshire Lottery. Retail sports betting is available in Iowa, Mississippi and New York. DraftKings Sportsbook mobile and retail sports betting products allow bettors in each state engage in betting for most major U.S. and international sports. As of April 2016, the majority of U.S. states consider fantasy sports a game of skill and not gambling. In November 2016, FanDuel and DraftKings, the two largest companies in the daily fantasy sports industry, reached an agreement to merge. However the merger was terminated in July 2017 due to it being blocked by the Federal Trade Commission as the combined company would have controlled a 90 percent of the market for daily fantasy sports. As of July 2017, DraftKings had eight million users.
Learn more about DraftKings
Size
3,400 employees
Market Cap
$4.9 billion
Industry
Net Income
-$577.9 million
Revenue
$292.3 million
NASDAQ

Similar Jobs

More Jobs at DraftKings

More Information Technology Jobs

Find similar Lead Site Reliability Engineer jobs: