Akamai Technologies

Site Reliability Engineer

Akamai Technologies$75K — $136K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Relevant experience supporting large-scale distributed systems
  • Bachelor's degree in Computer Engineering, Computer Science, or equivalent
  • Strong knowledge of Linux and networking (routing, DNS, firewalls, TCP/IP)
  • Proficient in a programming language like Python or Go
  • Experience with observability tools such as Prometheus or Grafana
  • Familiarity with infrastructure automation tools like Terraform or Ansible
  • Knowledge of container technologies such as Docker and orchestration platforms like Kubernetes

Responsibilities

  • Troubleshoot complex issues across Linux systems and distributed services
  • Build software and automation to reduce operational toil and improve efficiency
  • Develop AI-assisted tooling for incident investigation and reliability improvement
  • Use data analysis and debugging tools to enhance performance and reliability
  • Establish and improve monitoring, alerting, SLIs, and SLOs for key services
  • Contribute to root cause analysis and post-incident reviews
  • Partner with engineering teams to enhance system design and operational readiness
  • Participate in an on-call rotation to lead incident response efforts

Benefits

  • Comprehensive healthcare coverage
  • 401K savings plan with company matching
  • Generous paid time off (PTO) and sick leave
  • Family-friendly benefits including parental leave
  • Employee assistance program focusing on mental and financial wellness
  • Equity awards and Employee Stock Purchase Plan (ESPP) options
Full Job Description
Job Description

Our team is responsible for improving the reliability, performance, and scalability of our Compute products and platforms. We solve complex problems, improve how our systems operate, and build automation that makes our platforms more resilient.

In this role, you'll work at the intersection of systems, networking, and software engineering to solve problems across globally distributed services. You'll investigate complex production behavior, turn operational insights into lasting engineering improvements, and help shape how our services are operated.

As a Site Reliability Engineer, you will be responsible for:
  • Troubleshooting complex issues across Linux systems, networking, and distributed services.
  • Building software and automation that reduce operational toil, improve efficiency, and prevent recurring issues.
  • Developing and applying AI-assisted tooling to accelerate incident investigation, identify operational patterns, reduce toil, and improve reliability.
  • Using data analysis, network diagnostics, and debugging tools to identify performance and reliability improvements.
  • Establishing and improve monitoring, alerting, SLIs, and SLOs for critical services.
  • Contributing to root cause analysis, post-incident reviews, and long-term corrective actions.
  • Partnering with Engineering teams to improve system design, deployment safety, and operational readiness.
  • Participating in an on-call rotation and providing leadership during incident response, driving timely service restoration, effective communication, and post-incident improvement efforts.

Do what you love

To be successful in this role you will:
  • Have relevant experience and a Bachelor's degree in Computer Engineering, Computer Science or equivalent
  • Have experience supporting large-scale distributed systems
  • Have Linux and networking knowledge, including routing, DNS, firewalls, TCP/IP, and L7 traffic management
  • Be proficient in a programming language such as Python or Go
  • Have experience with observability tools such as Prometheus, Grafana, Loki, ELK/OpenSearch, or similar
  • Have experience with infrastructure automation or configuration management tools such as Terraform, Ansible, Salt, or similar.
  • Be familiar with container technologies such as Docker or Podman and orchestration platforms such as Kubernetes or Nomad.

Compensation

Akamai is committed to fair and equitable compensation practices. For US based candidates only - the base salary for this position ranges from $75,700 - $136,300/year; a candidate's salary is determined by various factors including, but not limited to, relevant work experience, skills, certifications and location. Compensation for candidates outside the US will vary. The compensation package may also include incentive compensation opportunities in the form of annual bonus or incentives, equity awards and an Employee Stock Purchase Plan (ESPP). Akamai provides industry-leading benefits including healthcare, 401K savings plan, company holidays, vacation (in the form of PTO), sick time, family friendly benefits including parental leave and an employee assistance program including a focus on mental and financial wellness; Eligibility requirements apply.

About Akamai Technologies

Akamai Technologies, Inc. is a global content delivery network (CDN), cybersecurity, and cloud service company. The company provides web and mobile performance solutions, cloud security solutions, enterprise access solutions, and video delivery solutions. Akamai was founded in 1998 and is headquartered in Cambridge, Massachusetts. The company serves a wide range of industries, including media and entertainment, gaming, software, financial services, healthcare, and others. Akamai is publicly traded on the NASDAQ stock exchange under the ticker symbol AKAM.
Learn more about Akamai Technologies
Size
8,700 employees
Market Cap
$13 billion
Industry
Net Income
$557 million
Founded
1998
5 Year Trend
+8.1%
Revenue
$3.1 billion
NASDAQ

Similar Jobs

More Jobs at Akamai Technologies

More Information Technology Jobs

Find similar Site Reliability Engineer jobs: