QualiTest Group

#23183 - Site Reliability Engineer

QualiTest Group$110K — $120K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 6-12 years of experience as a Site Reliability Engineer (SRE)
  • Proven AWS Cloud application experience
  • Experience supporting hybrid environments
  • Proficient in Linux/Unix and shell scripting
  • Familiarity with observability and APM tools, especially Datadog
  • Ability to build dashboards with Grafana and Kibana
  • Strong programming skills in Python, Java, or Shell Scripting

Responsibilities

  • Partner with application teams to enhance resiliency and reliability
  • Implement and uphold SLOs, SLIs, and best operational practices
  • Create observability solutions with monitoring and logging tools
  • Design monitoring, alerting, and health checks for applications
  • Automate processes to boost reliability and reduce manual work
  • Enhance capacity planning and performance management
  • Contribute to disaster recovery planning and execution

Benefits

  • Flexible hybrid work model (2-3 days onsite)
  • Access to innovative monitoring tools
  • Engagement in resilience and chaos engineering initiatives
  • Collaborative culture with cross-functional teams
  • Opportunities for hands-on experience with advanced technologies
  • Support for continuous learning and skills growth through diverse projects
Full Job Description
We are looking for a Site Reliability Engineer (SRE)) to join our growing team in the United States!

Location: Riverwoods, IL (Hybrid - 2 to 3 days/week onsite)

Position Overview

We are seeking an experienced Site Reliability Engineer (SRE) with a strong background in AWS Cloud, monitoring/observability platforms, and automation. The ideal candidate will partner closely with application development teams to improve application reliability, resiliency, performance, and operational excellence across hybrid cloud environments.

This role is ideal for engineers with hands-on experience in monitoring engineering using tools such as Datadog, Dynatrace, Grafana, Kibana, and expertise in AWS-based applications.

Must-Have Skills
  • 6-12 years of professional experience as a Site Reliability Engineer (SRE)
  • Strong hands-on experience with AWS Cloud applications and services (mandatory)
  • Experience supporting hybrid environments (AWS Cloud and on-premises deployments)
  • Strong Linux/Unix administration and shell scripting experience
  • Experience with Systems Observability and Application Performance Monitoring (APM) tools, preferably Datadog (Dynatrace experience is also valuable)
  • Experience building dashboards using Grafana and Kibana
  • Experience in performance testing and the ability to translate functional and non-functional requirements into automated non-functional testing (NFT) solutions
  • Strong understanding of application integration, high availability, resilience, and observability
  • Experience with DevOps practices and CI/CD pipelines
  • Strong programming skills in one or more of the following:
    • Python
    • Java
    • Shell Scripting (Unix/Linux)
  • Strong understanding of Software Development Life Cycle (SDLC)


Preferred Qualifications
  • Hands-on experience with ServiceNow (SNOW)
  • Experience with container technologies such as Kubernetes and OpenShift
  • Experience with Jenkins and CI/CD automation
  • Experience with automation tools such as Ansible
  • Strong working knowledge of JIRA
  • Basic understanding of Release Management
  • Understanding of Agile methodologies
  • Experience with AWS Lambda services
  • Knowledge of SQL, MySQL, and database concepts


Key Responsibilities
  • Partner with application development teams to improve application resiliency and reliability.
  • Implement and maintain Service Level Objectives (SLOs), Service Level Indicators (SLIs), and operational best practices.
  • Build end-to-end observability solutions using monitoring, logging, and tracing tools.
  • Design and implement monitoring, alerting, dashboards, and health checks for production applications.
  • Automate operational processes to reduce manual effort and improve system reliability.
  • Develop and enhance capacity planning and performance management capabilities.
  • Support disaster recovery (DR) planning and implementation for critical applications.
  • Develop and support chaos engineering and resilience testing initiatives.
  • Participate in production incident management and on-call support rotation.
  • Collaborate with cross-functional teams to improve system availability, scalability, and operational efficiency.


Technical Skills
  • Cloud: AWS (mandatory), Hybrid Cloud, On-Prem, Linux/Unix, AWS Lambda
  • Monitoring & Observability: Datadog (preferred), Dynatrace, Grafana, Kibana, ELK, APM tools
  • Programming: Python, Java, Shell Scripting, Go (preferred), Ansible
  • DevOps: Jenkins, CI/CD, Kubernetes, OpenShift
  • Tools: JIRA, ServiceNow (SNOW), Agile, Release Management
  • Database: SQL, MySQL


Benefits:

A Competitive pay, the salary range for the role is $110,000 - $120,000.

If you like what you have read, send us your resume and let's start talking!
  • Intrigued to find more about us?
    • Visit our website at https://www.quality-ai.com/
    • LinkedIn: https://www.linkedin.com/company/qualityaigroup

About QualiTest Group

Size
3,500 employees
Industry
Founded
1983

Similar Jobs

More Jobs at QualiTest Group

More Information Technology Jobs

Find similar #23183 - Site Reliability Engineer jobs: