Site Reliability Engineer

Nscale

$130K — $200K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 3-6 years in SRE, systems engineering, or software engineering, with production experience in cloud or data center environments.
  • Strong programming skills in Python, Go, or similar languages focused on automation.
  • Solid understanding of Linux, networking principles, and distributed systems.
  • Proven ability to troubleshoot live production issues and lead post-incident retrospectives.
  • Experience with monitoring tools and creating observability metrics, logs, and dashboards.

Responsibilities

  • Build and maintain automation and tooling to enhance platform reliability.
  • Define and uphold SLOs and SLIs, creating clear dashboards for service health monitoring.
  • Lead incident management efforts, including troubleshooting, root cause analyses, and retrospectives.
  • Investigate and resolve performance and reliability issues across Linux and networking.
  • Collaborate with engineering teams to enhance overall system reliability and efficiency.
  • Drive improvements in availability and scalability through code-based solutions.

Benefits

  • Competitive base salary plus equity that is reviewed annually.
  • Early responsibility and a clear progression plan tailored to individual skill development.
  • Flexible working arrangements that promote autonomy and trust in managing one's time.
Full Job Description
The Role
This is a career-level SRE role for someone who wants to own systems, not just watch them. You'll take
real surface area: the automation and tooling other engineers depend on, and the reliability of
production services running AI and GPU workloads at scale. You'll sit in the incident rotation, and you'll
be expected to make the systems you touch quieter over time.

What You'll Do
• Build and own the automation and tooling that keeps the platform running; treat operational toil as
a bug to be fixed, not a fact of life.
• Define and maintain SLOs, SLIs, and the dashboards that make service health obvious at a glance.
• Take point during incidents; troubleshoot under pressure, drive root cause analysis, and run post-
incident reviews that actually change the system.
• Investigate performance and reliability problems across Linux, networking, and distributed services,
then fix them at the source.
• Partner with Engineering, Networking, and Infrastructure teams to raise the reliability bar across the
stack.
• Improve availability, scalability, and efficiency through code, not manual effort.
What You'll Bring
• 3-6 years in SRE, systems engineering, or software engineering, including time running production in
a data center or cloud environment.
• Strong programming skills (Python, Go, or similar) and a genuine bias toward automating the work
away.
• Solid command of Linux, networking fundamentals, and distributed systems.
• A track record of troubleshooting live production issues and owning the fix through to the retro.
• Fluency with monitoring and observability; metrics, logs, dashboards, and alerting.
• Comfort in a fast-moving environment where priorities shift and you fill gaps without waiting to be
asked.

Nice to Have
• Experience with AI or GPU workloads, or high-performance computing (HPC).
• Familiarity with high-performance networking (InfiniBand, RDMA).
• Kubernetes, plus virtualized or bare-metal environments.
On-Call and Pace
A quick note on the shape of the job. This role sits close to production, so there is an on-call rotation,
and some weeks are busier than others. We share it fairly, and we treat every page as a signal worth
acting on rather than just an interruption. The goal is to make the systems quieter over time, so each
rotation asks less of the person carrying it. If you take ownership of what you run and like leaving it in
better shape than you found it, you'll do well here.
What We Offer
• Competitive base plus equity, reviewed every 12 months.
• Real scope early, and a progression plan built around the skills you want to sharpen.
• Flexibility that treats you as an adult; we care that the work gets done, and we trust you to shape
your day.

Salary Range
$130,000 - $200,000 USD. Actual compensation varies with skill set, experience, and location, and the
role may be eligible for bonus and equity.

Similar Jobs

More Jobs at Nscale

More Information Technology Jobs

Find similar Site Reliability Engineer jobs: