Capgemini

Site Reliability Engineer

Capgemini • $86K — $127K *
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • 9+ years of experience in Site Reliability Engineering and product design.
  • Proficient in managing large datasets and creating dashboards with Power BI and Grafana.
  • Strong background in Infrastructure as Code using Terraform and CloudFormation.
  • Hands-on experience with monitoring and logging distributed systems using Datadog and Splunk.
  • Experience in CI/CD automation with Jenkins and Azure DevOps.
  • Skilled in developing automation solutions with Python for application delivery.
  • Knowledge of AWS and Azure scalability and resiliency practices.

Responsibilities

  • Design, build, and operate resilient, scalable systems using best practices.
  • Deliver high-availability services through automation and proactive reliability engineering.
  • Enhance monitoring, logging, and observability for distributed systems.
  • Support CI/CD automation and production tooling to minimize operational toil.
  • Drive incident response and root cause analysis to reduce downtime.
  • Collaborate with engineering teams to integrate reliability into the software lifecycle.
  • Automate provisioning and configuration across cloud and on-prem environments.
  • Validate system resiliency and performance through testing and capacity planning.

Benefits

  • Paid time off including vacation, holidays, personal days, and sick leave.
  • Comprehensive medical, dental, and vision coverage.
  • Retirement savings plans like 401(k) in the U.S. and RRSP in Canada.
  • Life and disability insurance.
  • Employee assistance programs.
  • Additional benefits as per local policy and eligibility.
Full Job Description
Job Role: Site Reliability Engineer

Location: Boston, MA

Duration: Fulltime

Summary:

We are seeking a highly motivated Site Reliability Engineer to help build and operate reliable, scalable, and secure services across our platform. This role is designed for someone who combines strong DevOps practices, modern SRE principles, and software engineering experience to improve system reliability, automate operations, and support high-availability production environments.

The ideal candidate will be passionate about building resilient systems, improving developer productivity, and driving operational excellence through automation, observability, and engineering best practices. This role partners closely with engineering, platform, and product teams to ensure services are built for reliability from the start and remain performant, stable, and supportable at scale

Key Skills - Node.js, Python, DevOps, Jenkins, AWS

Preferred Qualifications
  • 9+ years of demonstrated experience developing and designing products around Site Reliability Engineering principles to improve stability and platform availability for containerized workloads and on-premises services using Kubernetes.
  • Experience managing and interpreting large datasets using query languages and creating dashboards and reports with Power BI and Grafana.
  • Strong background in managing cloud and on-premises infrastructure using Infrastructure as Code tools, including Terraform, and CloudFormation.
  • Hands-on experience building, operating, monitoring, logging, and alerting distributed systems at scale using Datadog and Splunk.
  • Experience supporting DevOps practices for service delivery and operations using Jenkins, Azure DevOps, Team Foundation Version Control, and CI/CD automation.
  • Experience developing software and automation solutions to support application delivery, operations, and repeatable business processes using Python.
  • Knowledge of scalability and resiliency practices for applications deployed on AWS and Azure, including Lambda and API Gateway
  • Strong development experience in scripting, automation, and integration across Linux and Windows-based environments.


Key Responsibilities
  • Design, build and operate resilient, scalable systems using DevOps, SRE, and software development best practices.
  • Deliver high-availability services through automation, infrastructure as code, and proactive reliability engineering.
  • Improve monitoring, logging, alerting, and observability for distributed systems.
  • Support CI/CD automation, deployment workflows, and production tooling to reduce operational toil.
  • Drive incident response, root cause analysis, and recovery improvements to minimize downtime.
  • Partner with engineering teams to embed reliability into the software development lifecycle.
  • Automate provisioning, configuration, and self-healing across cloud and on-prem environments.
  • Validate resiliency and performance through testing, chaos engineering, and capacity planning.'


The base compensation range for this role in the posted location is: $86129 - $127189

Capgemini provides compensation range information in accordance with applicable national, state, provincial, and local pay transparency laws. The base compensation range listed for this position reflects the minimum and maximum target compensation Capgemini, in good faith, believes it may pay for the role at the time of this posting. This range may be subject to change as permitted by law.

The actual compensation offered to any candidate may fall outside of the posted range and will be determined based on multiple factors legally permitted in the applicable jurisdiction.

These may include, but are not limited to: Geographic location, Education and qualifications, Certifications and licenses, Relevant experience and skills, Seniority and performance, Market and business consideration, Internal pay equity.

It is not typical for candidates to be hired at or near the top of the posted compensation range.

In addition to base salary, this role may be eligible for additional compensation such as variable incentives, bonuses, or commissions, depending on the position and applicable laws.

Capgemini offers a comprehensive, non-negotiable benefits package to all regular, full-time employees. In the U.S. and Canada, available benefits are determined by local policy and eligibility and may include:
  • Paid time off based on employee grade (A-F), defined by policy: Vacation: 12-25 days, depending on grade, Company paid holidays, Personal Days, Sick Leave
  • Medical, dental, and vision coverage (or provincial healthcare coordination in Canada)
  • Retirement savings plans (e.g., 401(k) in the U.S., RRSP in Canada)
  • Life and disability insurance
  • Employee assistance programs
  • Other benefits as provided by local policy and eligibility

About Capgemini

Capgemini is a global leader in consulting, digital transformation, technology and engineering services. The company is headquartered in Paris, France and operates in over 50 countries. Capgemini provides a range of services including strategy and transformation, application services, technology services, and engineering services. The company serves clients in a variety of industries including automotive, consumer products, financial services, healthcare, and retail.
Learn more about Capgemini
Industry
Founded
1967
NASDAQ

Similar Jobs

More Jobs at Capgemini

More Information Technology Jobs

Find similar Site Reliability Engineer jobs: