Akamai Technologies

Senior II Site Reliability Engineer

Akamai Technologies$146K — $263K *
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • 8+ years of SRE, infrastructure, or platform engineering experience with large-scale distributed systems.
  • Proven ability to define SLO/SLI frameworks and manage observability platforms.
  • Extensive experience with Kubernetes and container orchestration for compute-intensive workloads.
  • Skilled in building automation and tooling in Python or Go, with knowledge of CI/CD pipelines.
  • Strong capability to lead technical initiatives and mentor other engineers.
  • Familiarity with AI/ML infrastructure, model serving, or GPU workloads.

Responsibilities

  • Lead reliability workstreams for the serverless inference platform.
  • Design and implement observability strategies including telemetry, dashboards, and alerts.
  • Build automation and tooling that reduces operational toil and improves incident response.
  • Own incident management integration for inference workloads, leading incident response during on-call rotations.
  • Define deployment safety practices such as canary analysis and rollback automation.
  • Collaborate with product engineering teams to ensure operational readiness and represent the SRE perspective in design reviews.
  • Mentor and guide junior SREs through code reviews and design discussions.

Benefits

  • Healthcare coverage including mental and financial wellness support.
  • 401K savings plan with company contributions.
  • Generous paid time off (PTO), including sick leave and parental leave.
  • Employee Stock Purchase Plan (ESPP) and equity awards.
  • Family-friendly benefits and employee assistance programs.
Full Job Description
Job Description

Do you want to shape reliability practices for a new AI inference platform?

Are you a senior technical leader who drives solutions across teams?

Join the Akamai Inference Cloud Team!

The Akamai Inference Cloud team is part of Akamai's Cloud Technology Group. We design, implement, deploy and operate AI platforms that enable customers to run inference models and developers to create AI applications.

Partner with the best

In this role, you'll lead reliability workstreams for Akamai's serverless inference platform, design SRE tooling and automation, and drive technical decisions. Opportunities exist to mentor other SREs, influence architecture decisions with product engineering teams, and shape SRE practices for AI inference workloads and GPU infrastructure at scale.

As a Senior II Site Reliability Engineer, you will be responsible for:
  • Taking ownership of observability strategy for the serverless inference platform, designing telemetry, dashboards, and alerts, defining SLO/SLI frameworks, and driving improvements when targets are missed
  • Building production-grade automation and tooling that reduces operational toil, improves incident response, and sets patterns that other SREs adopt
  • Owning incident management integration for inference workloads, designing frameworks, leading incident response during on-call rotations, and driving systemic improvements from post-mortems
  • Defining and implementing deployment safety practices including progressive rollouts, canary analysis, and rollback automation, establishing standards for the team
  • Partnering with product engineering teams to influence architecture decisions, ensure operational readiness, and represent the SRE perspective in design reviews
  • Mentoring Senior and mid-level SREs through code reviews, design discussions, and hands-on problem-solving

Do what you love

To be successful in this role you will:
  • 8+ years of experience in SRE, infrastructure engineering, or platform engineering, working with large-scale distributed systems
  • Possess a proven track record of defining SLO/SLI frameworks, building observability platforms, and running incident management processes at scale
  • Have extensive Kubernetes and containerization experience at scale, including autoscaling, resource scheduling, and container orchestration for compute-intensive workloads
  • Have experience building automation and tooling in Python or Go, with familiarity in CI/CD pipelines, deployment safety, and infrastructure-as-code
  • Possess the ability to lead technical initiatives across teams, mentor other engineers, and drive complex reliability problems to resolution independently
  • Have experience with or exposure to AI/ML infrastructure, model serving, or GPU workloads

Compensation

Akamai is committed to fair and equitable compensation practices. For US based candidates only - the base salary for this position ranges from $146,400 - $263,600/year; a candidate's salary is determined by various factors including, but not limited to, relevant work experience, skills, certifications and location. Compensation for candidates outside the US will vary. The compensation package may also include incentive compensation opportunities in the form of annual bonus or incentives, equity awards and an Employee Stock Purchase Plan (ESPP). Akamai provides industry-leading benefits including healthcare, 401K savings plan, company holidays, vacation (in the form of PTO), sick time, family friendly benefits including parental leave and an employee assistance program including a focus on mental and financial wellness; Eligibility requirements apply.

About Akamai Technologies

Akamai Technologies, Inc. is a global content delivery network (CDN), cybersecurity, and cloud service company. The company provides web and mobile performance solutions, cloud security solutions, enterprise access solutions, and video delivery solutions. Akamai was founded in 1998 and is headquartered in Cambridge, Massachusetts. The company serves a wide range of industries, including media and entertainment, gaming, software, financial services, healthcare, and others. Akamai is publicly traded on the NASDAQ stock exchange under the ticker symbol AKAM.
Learn more about Akamai Technologies
Size
8,700 employees
Market Cap
$13 billion
Industry
Net Income
$557 million
Founded
1998
5 Year Trend
+8.1%
Revenue
$3.1 billion
NASDAQ

Similar Jobs

More Jobs at Akamai Technologies

More Information Technology Jobs

Find similar Senior II Site Reliability Engineer jobs: