Staff Software Engineer - Observability

Cerebras Systems

• $150K — $180K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Strong experience in backend or systems software engineering
  • Proficiency in Go, C++, Rust, Java, or Python
  • Solid understanding of distributed systems, networking fundamentals, and concurrency trade-offs
  • Hands-on experience with metrics, logs, and distributed tracing
  • Familiarity with observability tools like OpenTelemetry, Prometheus, and Grafana

Responsibilities

  • Design and implement observability instrumentation across services and platforms
  • Build and maintain telemetry pipelines for metrics, logs, and traces at scale
  • Develop internal observability platforms, libraries, and tooling
  • Define and operationalize SLIs, SLOs, and alerting strategies
  • Partner with engineers to enhance system debuggability by design
  • Reduce MTTR by enabling fast root cause analysis during incidents
  • Create clear, actionable dashboards and alerts reflecting system health

Benefits

  • Collaborative work environment across multiple engineering disciplines
  • Opportunity to shape observability infrastructure rather than only dashboards
  • Involvement in high-impact design and operational strategies
  • Focus on improving developer experience around observability and debugging
  • Engagement with cutting-edge tools and technologies for observability
Full Job Description
We\'re looking for a Software Engineer focused on Observability to build and evolve the systems that give us deep visibility into large-scale, performance-critical production systems.

You\'ll design and implement metrics, logging, tracing, and alerting infrastructure that enables fast debugging, high reliability, and confident operation of complex distributed systems. This role sits at the intersection of platform engineering, distributed systems, and reliability.

This is not a dashboards-only role - you\'ll be writing production software, shaping internal platforms, and working closely with engineers across the stack.

Responsibilities
  • Design and implement observability instrumentation across services and platforms
  • Build and maintain telemetry pipelines for metrics, logs, and traces at scale
  • Develop internal observability platforms, libraries, and tooling
  • Define and operationalize SLIs, SLOs, and alerting strategies
  • Partner with engineers to make systems debuggable by design
  • Reduce MTTR by enabling fast root-cause analysis during incidents
  • Create clear, actionable dashboards and alerts that reflect real system health
  • Balance telemetry signal vs cost, noise, and performance impact
  • Improve the developer experience around observability and debugging

Qualifications:

Core Engineering Skills
  • Strong experience in backend or systems software engineering
  • Proficiency in one or more of:
    • Go, C++, Rust, Java, Python
  • Solid understanding of:
    • Distributed systems
    • Networking fundamentals
    • Concurrency and performance tradeoffs

Observability & Reliability Experience
  • Hands-on experience with:
    • Metrics, logs, and distributed tracing
    • Production monitoring and alerting
  • Familiarity with tools such as:
    • OpenTelemetry
    • Prometheus
    • Grafana
    • Datadog / Elastic / Jaeger / Tempo (or similar)
  • Experience designing:
    • High-signal alerts
    • Scalable telemetry pipelines
    • Service-level indicators and objectives

Preferred Qualifications:
  • Experience in high-performance computing, AI/ML systems, or inference platforms
  • Hardware-aware observability (accelerators, GPUs, custom hardware)
  • Prior SRE or platform engineering background
  • Experience debugging large-scale production incidents
  • Building internal developer platforms or shared libraries

Similar Jobs

More Jobs at Cerebras Systems

More Information Technology Jobs

Find similar Staff Software Engineer - Observability jobs: