Mirantis

Senior Site Reliability Engineer (Golang / Kubernetes)

Mirantis$135K — $160K *
US-AnywhereRemote in United States
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years in SRE or related role
  • Proficient in software engineering languages (e.g., Go or Python)
  • Experience defining SLIs/SLOs for production systems
  • Hands-on experience with observability tools (e.g., Prometheus, Grafana)
  • Strong understanding of Kubernetes and its signals

Responsibilities

  • Define SLIs and SLOs from platform signals
  • Develop API for exposing SLIs to Administrators
  • Establish optimized alerting and error-budget practices
  • Collaborate with infrastructure teams to ensure signal collection
  • Diagnose and resolve reliability and performance issues

Benefits

  • Work with a Silicon Valley leader in cloud infrastructure
  • Collaborate with passionate and talented colleagues
  • Engage in cutting-edge, open-source innovation
  • Thrive in a dynamic, collaborative work environment
  • Professional development and training opportunities
  • Attend conferences and industry events
  • Participate in team outings, hackathons, and tech talks
Full Job Description
Define what reliability means for a GPU-accelerated AI platform and make it measurable. You will own the service-level indicators and objectives for the K0rdent Observability Framework (KOF) - deriving meaningful SLIs from the signals the platform already emits, and exposing them to Platform Administrators through a clean API. Work spans hybrid, edge, and air-gapped deployments built on the Mirantis K0rdent stack. We are looking for a Senior SRE who thinks past dashboards to the contract between a platform and its operators. The right candidate can look at raw telemetry from Kubernetes, bare metal, and NVIDIA infrastructure, decide which signals actually predict user-visible reliability, and turn them into SLIs and SLOs that operators can act on. You are equally comfortable writing the service that exposes those SLIs through an API and reasoning about error budgets, alerting quality, and signal-to-noise. You should be self-directed, able to own reliability definitions end to end, and effectively communicate them across teams. Responsibilities • Define SLIs and SLOs based on the signals available across the platform - Kubernetes, bare-metal hosts, and NVIDIA infrastructure (BMC, InfiniBand, NVLink, UFM). • Design and build the API that exposes SLIs and reliability state to Platform Administrators and downstream systems. • Establish alerting and error-budget practices that maximize signal and minimize noise. • Partner with infrastructure, storage, and networking teams to ensure the right signals are instrumented and collected. • Diagnose reliability and performance issues across the observability stack and drive their resolution. Qualifications Required Qualifications • 5+ years in SRE, platform reliability, or a closely related software/infrastructure role. • Strong software engineering skills (e.g., Go or Python) with experience building and operating APIs or services in production. • Demonstrated experience defining SLIs/SLOs and error budgets for real production systems. • Hands-on experience with observability tooling - metrics, logging, and tracing (e.g., Prometheus/VictoriaMetrics, OpenTelemetry, Grafana). • Solid understanding of Kubernetes and the signals it and its workloads emit. • Strong written and verbal communication with technical audiences. Preferred • Experience instrumenting or monitoring bare-metal and NVIDIA infrastructure (BMC/Redfish, InfiniBand, NVLink, UFM). • Experience with the Mirantis K0rdent stack (K0rdent Enterprise, K0rdent AI, KOF) and Cluster API. • Familiarity with VictoriaMetrics/VictoriaLogs at scale. • Proven experience in sovereign or high-security air-gapped environments. Additional Information What does Mirantis offer you? • Work with an established Silicon Valley leader in the cloud infrastructure industry; • Work with exceptionally passionate, talented and engaging colleagues, helping Fortune 500 and Global 2000 customers implement next-generation cloud technologies; • Be a part of cutting-edge, open-source innovation; • Thrive in the high-energy environment of a young company where openness, collaboration, risk-taking, and continuous growth are valued; • Professional development and training; • Attend conferences and working groups; • Company outings, happy hours, hackathons, and tech talks; • Receive a competitive compensation package with a strong benefits plan.

About Mirantis

Mirantis is a software company that provides cloud computing services and solutions. The company was founded in 2011 and is headquartered in Sunnyvale, California. Mirantis offers a range of cloud computing services, including OpenStack, Kubernetes, and Docker. The company's solutions are used by a variety of industries, including telecommunications, finance, and healthcare. Mirantis has over 1,000 employees and offices in the United States, Russia, Ukraine, and the United Kingdom.
Learn more about Mirantis
Size
1,000 employees
Industry
Founded
2011

Similar Jobs

More Jobs at Mirantis

More Information Technology Jobs

Find similar Senior Site Reliability Engineer (Golang / Kubernetes) jobs: