Senior Site Reliability Engineer I

Axon

$134K — $214K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's Degree in Computer Science, Engineering, or equivalent technical field
  • 7+ years of experience in SRE, platform engineering, or infrastructure engineering
  • Experience with agentic AI tooling or building LLM-powered developer tools
  • Strong Linux systems fundamentals and experience in Kubernetes environments
  • Hands-on experience with Loki, Grafana, Tempo/Jaeger, or Mimir/Cortex
  • Experience with Terraform for infrastructure as code, CDK is a plus
  • Proficiency in Golang, Python, or Java

Responsibilities

  • Own and evolve distributed tracing infrastructure with Jaeger and OpenTelemetry
  • Build and operate the log aggregation platform using Grafana Loki and Alloy
  • Maintain and improve the metrics infrastructure with Cortex, Prometheus, and Grafana
  • Develop internal tooling and automation for self-service observability
  • Manage observability infrastructure as code through Terraform and other tools
  • Collaborate with engineering teams to define instrumentation standards and drive tracing adoption

Benefits

  • Competitive salary and 401k with employer match
  • Discretionary time off
  • Paid parental leave for all
  • Medical, Dental, Vision plans
  • Fitness programs
  • Emotional & Development programs
  • Snack availability in offices
Full Job Description
Your Impact

Are you an engineer who gets excited about the challenge of making complex distributed systems observable - not just instrumenting them, but designing the infrastructure that makes traces, metrics, and logs useful at scale? In this role, you will help build and evolve Axon's next-generation observability platform, enabling the entire engineering organization to understand and operate their services with confidence.

You'll work across the full observability stack: from distributed tracing adoption (OpenTelemetry, Jaeger) to log infrastructure (Loki, Alloy) to metrics (Cortex, Prometheus, Grafana). You'll partner directly with Axon's engineering teams to drive adoption of modern observability practices and build the tooling that makes our platform self-service for the teams that depend on it.

You will be part of the Observability team within Axon's Site Reliability organization - a focused team responsible for Axon's metrics, logging, tracing, and alerting infrastructure across dozens of environments globally.

The ideal candidate has a strong infrastructure engineering background, is comfortable working across cloud-native systems, and cares about both the technical depth and the developer experience of observability. You'll thrive here if you have opinions about what good observability looks like, and enjoy the challenge of making it real in a large, fast-moving organization.

Location: This role is based out of our Boston, MA office and follows a hybrid schedule. We rely on in-person collaboration and ask that team members work onsite Tuesdays through Fridays, with the flexibility to work remotely on Mondays, unless there is an approved workplace accommodation. We believe that connection fuels innovation, and our in-office culture is designed to foster meaningful teamwork, mentorship, and shared success.
What You'll Do
  • Own and evolve Axon's distributed tracing infrastructure, including Jaeger and OpenTelemetry-based instrumentation, driving adoption across Axon's service-oriented architecture
  • Build and operate Axon's log aggregation platform (Grafana Loki + Alloy), expanding use cases beyond Kubernetes event logs and reducing organizational dependency on expensive third-party log tooling (including Splunk)
  • Maintain and improve Axon's metrics infrastructure (Cortex, Prometheus, Grafana) - the foundation for alerting, dashboards, and SLO tracking across all of Axon's environments
  • Write internal tooling and automation that makes observability self-service: toolkit commands, agentic on-call helpers, runbook generation, and dashboard scaffolding
  • Manage observability infrastructure as code via Terraform, CDK, ArgoCD, and Helm - including capacity management, cybersecurity requirements and compliance, and on-call rotation participation
  • Work directly with engineering teams across Axon to define instrumentation standards, drive tracing adoption, and help teams build meaningful SLOs for their services
Basic Qualifications
  • Bachelor's Degree in Computer Science, Engineering, or an equivalent highly technical field
  • 7+ years of experience in SRE, platform engineering, or infrastructure engineering
  • Experience with agentic AI tooling or building LLM-powered developer tools
  • Strong Linux systems fundamentals and comfort working in Kubernetes-based environments
  • Hands-on experience with one or more components of the LGTM stack: Loki, Grafana, Tempo/Jaeger, or Mimir/Cortex
  • Experience with infrastructure as code - Terraform strongly preferred, CDK is a plus
  • Experience with any of: Golang, Python, or Java
  • United States Citizen - able to gain CJIS clearance for full US production access
Preferred qualifications
  • Experience deploying and operating distributed tracing systems (OpenTelemetry, Jaeger, Tempo, or similar)
  • Familiarity with OpenTelemetry - instrumentation, collectors, and pipelines
  • Experience with GitOps workflows
  • Exposure to 24/7 high-volume systems with formal SLA requirements
  • Ability to debug complex multi-service distributed systems
Benefits that Benefit You
  • Competitive salary and 401k with employer match
  • Discretionary time off
  • Paid parental leave for all
  • Medical, Dental, Vision plans
  • Fitness Programs
  • Emotional & Development Programs
  • And yes, we have snacks in our offices

Benefits listed herein may vary depending on the nature of your employment and the location where you work.

Axon is a total compensation company, meaning compensation is made up of base pay, bonus, and stock awards. The actual base pay is dependent upon many factors, such as: level, function, training, transferable skills, work experience, business needs, geographic market, and often a combination of all these factors. Our benefits offer an array of options to help support you physically, financially and emotionally through the big milestones and in your everyday life. To see more details on our benefits offerings please visit https://www.axon.com/careers.

Base Pay Range

$134,250-$214,800 USD

Don't meet every single requirement? That's ok. At Axon, we Aim Far. We think big with a long-term view because we want to reinvent the world to be a safer, better place. We are also committed to building diverse teams that reflect the communities we serve.

Studies have shown that women and people of color are less likely to apply to jobs unless they check every box in the job description. If you're excited about this role and our mission to Protect Life but your experience doesn't align perfectly with every qualification listed here, we encourage you to apply anyways. You may be just the right candidate for this or other roles.

Important Notes

The above job description is not intended as, nor should it be construed as, exhaustive of all duties, responsibilities, skills, efforts, or working conditions associated with this job. The job description may change or be supplemented at any time in accordance with business needs and conditions.

Some roles may also require legal eligibility to work in a firearms environment.

Similar Jobs

More Jobs at Axon

More Information Technology Jobs

Find similar Senior Site Reliability Engineer I jobs: