Site Reliability Architect

Summit

$136K — $175K *
US-AnywhereRemote in United States
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of experience in site reliability engineering or related field
  • Expertise in defining and implementing observability and reliability frameworks
  • Comfortable affecting architecture and strategy across multiple teams
  • Strong mentoring abilities to cultivate team best practices
  • Proficient in automation tools and strategies

Responsibilities

  • Define observability and reliability architecture strategy
  • Implement site reliability concepts across teams
  • Collaborate with teams to ensure resilient and scalable system design
  • Lead design and governance of observability standards
  • Oversee automation of incident monitoring and alerting
  • Serve as escalation point for cross-platform incidents
  • Evaluate and introduce emerging reliability tools

Benefits

  • Flexible Time Off policy
  • Comprehensive Medical, Dental & Vision coverage
  • 401(k) plan with 4% employer match
  • Paid Parental Leave
  • Life & Disability Insurance
  • Wellness Support resources
  • Free access to Colocation & Cloud services
  • Remote Work opportunities
  • Low-ego, results-oriented company culture
Full Job Description
As our Site Reliability Architect, you will own the vision, strategy, and architecture of Summit's observability and reliability platforms. You will ensure that the systems supporting our most critical applications are designed with resilience, scalability, and efficiency in mind. Beyond managing platforms, you will shape the standards and practices that make Summit's infrastructure visible, measurable, and reliable.

You will collaborate across engineering, operations, and product teams to embed observability into the fabric of our technology, mentor engineers on best practices, and influence architecture decisions that improve performance and reliability. The right candidate will thrive in a fast-moving environment, combining deep technical expertise with the ability to see the bigger picture and guide the evolution of Summit's reliability ecosystem.

What You'll Do:

  • Define the observability and reliability architecture strategy across Summit's platforms and services.
  • Implement site reliability concepts based around SLOs, SLIs, and SLAs, across teams and platforms
  • Partner with engineering and operations leadership to ensure system design aligns with resilience and scalability goals.
  • Lead the design, implementation, and governance of observability frameworks and standards across teams.
  • Oversee the automation of monitoring, alerting, and incident response processes to maximize efficiency and reduce human error.
  • Serve as the escalation point and lead for complex, cross-platform incidents, guiding resolution and post-incident analysis.
  • Evaluate and introduce emerging tools, frameworks, and practices to strengthen Summit's reliability and observability posture.
  • Mentor and coach engineering teams, fostering a culture of reliability, automation, and continuous improvement.


What You'll Deliver:

  • A scalable, standards-driven observability framework embedded across Summit's technology stack.
  • Clear alignment of monitoring, alerting, and automation with business-critical definitions of health.
  • Reduced time-to-detect and time-to-resolve for incidents through automation and well-designed response frameworks.
  • Improved reliability practices across teams through coaching, knowledge-sharing, and collaboration.
  • Strategic roadmaps for observability and monitoring platforms that support Summit's growth and evolving client needs.


You'll Thrive in This Role If You:

  • Have a proven track record in designing and implementing observability and reliability platforms at scale.
  • Are comfortable influencing architecture and strategy across engineering, operations, and product teams.
  • Enjoy mentoring engineers and shaping organizational practices, not just managing tools.
  • Think strategically but aren't afraid to get hands-on when solving complex problems.
  • Thrive in environments where you can innovate, set standards, and bring structure to complex ecosystems.


Bonus Points:

  • Expertise with observability stacks such as ELK, Grafana/Graphite/InfluxDB, LogicMonitor, Prometheus, or related platforms.
  • Advanced automation experience using Ansible, Terraform, or equivalent tooling.
  • Experience designing reliability frameworks in multi-cloud environments (Azure, AWS, or hybrid).
  • Knowledge of scripting and development languages (Python, Go, Ruby, or Javascript).
  • Familiarity with compliance-heavy industries where reliability, security, and auditability are paramount.


Benefits:

Summit offers a total rewards package designed to support you at work, at home, and everywhere in between. Here's a snapshot of what that looks like:

  • Flexible Time Off (yes, really) - take what you need, we trust you to manage it
  • Medical, Dental & Vision - comprehensive coverage with HSA/HRA options
  • 401(k) with 4% match - we invest alongside you
  • Parental Leave - paid time for the moments that matter most
  • Life & Disability Insurance - built-in peace of mind
  • Wellness Support - resources for your mental and physical health
  • Free Colocation & Cloud Access - build and experiment in real environments
  • Work From Anywhere - remote-friendly by design
  • A low-ego, get-it-done culture - come as you are, do great work


Compensation:

The salary range for this role is $136,000 - $175,000 annually, depending on your skills, experience, and location. We aim to make offers that feel fair, forward-looking, and reflective of what you bring to the table - and where you want to grow.

Internal candidates may see variations based on current role, compensation, and progression at Summit. Same thoughtful approach, just with additional context.

Similar Jobs

More Jobs at Summit

More Information Technology Jobs

Find similar Site Reliability Architect jobs: