Job DescriptionOur team is responsible for improving the reliability, performance, and scalability of our Compute products and platforms. We solve complex problems, improve how our systems operate, and build automation that makes our platforms more resilient.
In this role, you'll work at the intersection of systems, networking, and software engineering to solve problems across globally distributed services. You'll investigate complex production behavior, turn operational insights into lasting engineering improvements, and help shape how our services are operated.
As a Site Reliability Engineer, you will be responsible for:
- Troubleshooting complex issues across Linux systems, networking, and distributed services.
- Building software and automation that reduce operational toil, improve efficiency, and prevent recurring issues.
- Developing and applying AI-assisted tooling to accelerate incident investigation, identify operational patterns, reduce toil, and improve reliability.
- Using data analysis, network diagnostics, and debugging tools to identify performance and reliability improvements.
- Establishing and improve monitoring, alerting, SLIs, and SLOs for critical services.
- Contributing to root cause analysis, post-incident reviews, and long-term corrective actions.
- Partnering with Engineering teams to improve system design, deployment safety, and operational readiness.
- Participating in an on-call rotation and providing leadership during incident response, driving timely service restoration, effective communication, and post-incident improvement efforts.
Do what you loveTo be successful in this role you will:
- Have relevant experience and a Bachelor's degree in Computer Engineering, Computer Science or equivalent
- Have experience supporting large-scale distributed systems
- Have Linux and networking knowledge, including routing, DNS, firewalls, TCP/IP, and L7 traffic management
- Be proficient in a programming language such as Python or Go
- Have experience with observability tools such as Prometheus, Grafana, Loki, ELK/OpenSearch, or similar
- Have experience with infrastructure automation or configuration management tools such as Terraform, Ansible, Salt, or similar.
- Be familiar with container technologies such as Docker or Podman and orchestration platforms such as Kubernetes or Nomad.
CompensationAkamai is committed to fair and equitable compensation practices. For US based candidates only - the base salary for this position ranges from $75,700 - $136,300/year; a candidate's salary is determined by various factors including, but not limited to, relevant work experience, skills, certifications and location. Compensation for candidates outside the US will vary. The compensation package may also include incentive compensation opportunities in the form of annual bonus or incentives, equity awards and an Employee Stock Purchase Plan (ESPP). Akamai provides industry-leading benefits including healthcare, 401K savings plan, company holidays, vacation (in the form of PTO), sick time, family friendly benefits including parental leave and an employee assistance program including a focus on mental and financial wellness; Eligibility requirements apply.