Site Reliability Engineer

Danta Technologies

$110K — $130K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5-10+ years of overall IT experience
  • 5+ years in Site Reliability Engineering, Production Support, DevOps, or Cloud Operations
  • Strong experience with Linux/Unix administration
  • Proficiency in Python and additional scripting languages
  • Hands-on experience with cloud platforms like AWS, Azure, or GCP
  • Experience with containerization technologies such as Docker and Kubernetes
  • Understanding of networking, security, and distributed systems

Responsibilities

  • Monitor and improve the reliability and performance of production systems
  • Design and implement observability solutions
  • Establish and track SLIs, SLOs, and Error Budgets
  • Automate operational tasks using scripting and IaC
  • Lead incident response and post-incident reviews
  • Collaborate with teams to enhance system resilience
  • Perform capacity planning and scalability assessments
  • Support CI/CD pipelines and deployment automation

Benefits

  • Competitive compensation package
  • Healthcare insurance options (Dental, Medical, Vision)
  • Major holidays observed
  • Paid sick leave as per state law
Full Job Description
Role - Site Reliability Engineer
Location - Schaumburg, IL or Secaucus, NJ (Hybrid)

Mandatory Skills: Python/R and ML libraries (scikit-learn, TensorFlow, PyTorch), Data analysis and visualization (Pandas, NumPy, Power BI/Tableau), SQL and database management


Key Responsibilities
• Monitor, maintain, and improve the reliability, availability, and performance of production systems.
• Design and implement monitoring, alerting, logging, and observability solutions.
• Establish and track Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets.
• utomate operational tasks and repetitive processes using scripting and Infrastructure as Code (IaC).
• Lead incident response activities, troubleshooting, root cause analysis (RCA), and post-incident reviews.
• Collaborate with development, infrastructure, and platform teams to improve system reliability and resilience.
• Perform capacity planning, performance tuning, and scalability assessments.
• Support CI/CD pipelines and deployment automation initiatives.
• Implement high-availability, disaster recovery, and failover strategies.

Required Skills
• Strong experience with Linux/Unix administration.
• Proficiency in scripting languages such as Python, Shell, or PowerShell.
• Hands-on experience with cloud platforms (AWS, Azure, or GCP).
• Experience with containerization technologies such as Docker and Kubernetes.
• Knowledge of monitoring and observability tools such as Prometheus, Grafana, ELK, Splunk, Dynatrace, or Datadog.
• Understanding of CI/CD tools such as Jenkins, GitHub Actions, GitLab CI, or Azure DevOps.
• Experience with Infrastructure as Code tools such as Terraform, Ansible, or CloudFormation.
• Strong troubleshooting, debugging, and problem-solving skills.
• Understanding of networking, security, and distributed systems concepts.

Experience
• 5-10+ years of overall IT experience.
• 5+ years of hands-on experience in Site Reliability Engineering, Production Support, DevOps, or Cloud Operations roles.

Benefits: Danta offers a compensation package to all W2 employees that are competitive in the industry. It consists of competitive pay, the option to elect healthcare insurance (Dental, Medical, Vision), Major holidays and Paid sick leave as per state law.

The rate/ Salary range is dependent on numerous factors including Qualification, Experience and Location.

Similar Jobs

More Jobs at Danta Technologies

More Information Technology Jobs

Find similar Site Reliability Engineer jobs: