Site Reliability Engineer II

DAT

$95K — $134K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 2 to 4+ years of industry experience in Site Reliability Engineering or related fields.
  • At least 1 year of software engineering experience with languages such as JavaScript, Python, Go, Java/Kotlin, or C++.
  • Proficiency with modern observability tools, with a preference for Datadog.
  • Experience with cloud platforms, preferably AWS.
  • Strong collaboration and problem-solving skills within SRE or Platform Engineering teams.

Responsibilities

  • Contribute to scalable and reliable system design, implementation, and maintenance.
  • Identify and troubleshoot complex issues in distributed systems for optimal performance.
  • Advocate for best practices in SRE, focusing on automation and monitoring.
  • Participate in capacity planning and performance tuning to mitigate bottlenecks.
  • Leverage AI tools for coding and observability enhancements.
  • Respond to critical engineering incidents as they arise.
  • Drive a culture of continuous improvement within the engineering teams.

Benefits

  • Hybrid work model with flexibility in location.
  • Opportunity for professional growth and skills enhancement in SRE practices.
  • Collaboration with development teams and platform architects for impactful projects.
  • Participation in significant technical initiatives to shape platform reliability.
  • A culture that values sharing expertise and continuous improvement among peers.
Full Job Description
Job Application Deadline: 10/31/2026

The Opportunity

DAT is looking for a Site Reliability Engineer to join our SRE platform team. This position will work hybrid in Denver, CO

Candidate profile

DAT is seeking an experienced Site Reliability Engineer to help grow our SRE practices. In this role, you will be responsible for contributing to technical initiatives and enhancing your skills. You'll work closely with development teams and platform architects to achieve critical reliability goals and help scale our platform.DAT is actively seeking a highly skilled and experienced Site Reliability Engineer (SRE) to play a pivotal role in the expansion and maturation of our SRE practices. In this critical position, the successful candidate will be instrumental in driving key technical initiatives, fostering a culture of continuous improvement, and significantly enhancing their own professional expertise.

This role necessitates close collaboration with various stakeholders, including our dedicated development teams and platform architects. The primary objective of these partnerships is to collectively achieve ambitious reliability goals and strategically scale our platform to meet evolving growth of the company. The SRE will be responsible for ensuring the stability, performance, and scalability of our systems, implementing robust monitoring solutions, automating operational tasks, and proactively identifying and resolving potential issues. This will involve a deep understanding of distributed systems, cloud infrastructure, and a commitment to best practices in site reliability engineering.

What You'll Do
  • Contribute to the design, implementation, and maintenance of scalable and reliable systems. Collaborate with engineering teams to ensure reliability targets are met.
  • Identify and troubleshoot complex issues across distributed systems, ensuring minimal downtime and optimal performance.
  • Advocate for and implement SRE best practices, including automation, monitoring, and incident response, to enhance system resilience.
  • Participate in capacity planning and performance tuning to proactively address potential bottlenecks and support future growth.
  • Leverage new AI tools to assist with coding and observability tasks.
  • Assist and respond to critical engineering incidents.
  • Improve your engineering skills within the SRE team.
  • Provide technical guidance and best practices for use of cloud infrastructure and tooling. Contribute to Infrastructure-as-Code within the platform. We strive to automate all the things!
  • Contribute to reliability-focused initiatives and projects.
  • Help optimize our work to be customer-focused. Continually seek feedback from our customers on how we can improve.
  • Assist in migrating legacy systems to modern, scalable cloud environments.
  • Help develop and drive a culture of continuous improvement with the Platform Engineering and Software Engineering groups.
  • Participate in an on-call rotation.

The Skills and Experience You'll Bring
  • Strong collaboration and problem-solving abilities, especially within SRE or Platform Engineering/Infrastructure teams.
  • Total of 2 to 4+ years industry experience
  • At least 1 year of software engineering experience (JavaScript, Python, Go, Java/Kotlin, C++, etc)
  • Experience with modern observability tools (Datadog preferred).
  • Experience with cloud platforms (preferably AWS).
  • Demonstrated success in contributing to large technical initiatives and acting as a driving force to complete those initiatives.
  • Proven experience assisting in modernizing legacy code and infrastructure.
  • Ability to work closely with peer teams, platform/software architects and management to drive key reliability improvements.
  • Willingness to share your expertise among team members and others within the engineering organization. We value upleveling our peers however we can.
  • Understanding of cloud infrastructure, automation, and best practices for reliability.
  • Experience with our tools (Kubernetes, ArgoCD, Terraform, Github Actions) a plus.


*This position is not eligible for visa sponsorship**

For Colorado-based candidates, in compliance with the Colorado Equal Pay for Equal Work Act, the salary range for this role is $95,000.00 - $134,000.00 + target bonus. DAT considers factors such as scope and responsibilities of the position, candidate's work experience, education and training, core skills, internal equity, and market and business elements when extending an offer.

#LI-RF1

#LI-hybrid

Similar Jobs

More Jobs at DAT

More Information Technology Jobs

Find similar Site Reliability Engineer II jobs: