Mirantis

Technical Product Manager, Observability - remote in the US

Mirantis$135K — $160K *
US-AnywhereRemote in United States
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years in product management or technical role focused on observability product
  • Knowledge of Prometheus, OpenTelemetry, and other observability tools
  • Expertise in Kubernetes observability and cloud-native monitoring
  • Experience collaborating with engineering on technical trade-offs and field teams in GPU cloud deals

Responsibilities

  • Own the vision and roadmap for k0rdent AI observability across the stack
  • Translate diverse requirements into clear product direction, collaborating with engineering
  • Manage observability backlog based on feedback from production deployments
  • Track and respond to emerging observability technologies and standards
  • Define strategies for integrating vendor telemetry sources into a unified observability plane
  • Collaborate with product marketing on positioning and represent Mirantis to clients and partners

Benefits

  • Build the observability foundation for the AI cloud era
  • Collaborate with a world-class, distributed team focused on technical excellence
  • Shape product narrative and impact go-to-market strategy
Full Job Description
Job Description

Job Summary

Mirantis is looking for a Technical Product Manager to own observability for k0rdent AI, our control plane for GPU infrastructure and distributed AI workloads. In this role, you will define the observability strategy, roadmap, and feature priorities that determine how operators gain visibility into the health, performance, and resource utilization of GPU clusters running large-scale training and inference. You will shape how k0rdent AI handles everything from GPU-level metrics and distributed tracing across AI workloads, to multi-tenant log aggregation and intelligent alerting - powered by the OpenTelemetry ecosystem, and Prometheus-compatible metrics pipelines.

The ideal candidate brings strong technical fluency in observability tooling and the AI infrastructure stack. You will work directly with engineering to shape requirements, with marketing to define positioning, and with customers to help ensure their success.

Responsibilities
  • Own the vision, roadmap, and priorities for k0rdent AI observability across the full stack: GPU compute, east-west fabric (InfiniBand, RoCE), high-performance storage, DPU/SmartNIC telemetry, workload schedulers, inference serving, and data services
  • Translate requirements from NeoClouds, GPU clouds, telcos, sovereign clouds, and enterprise platform teams into clear product direction; partner with engineering to define requirements and evaluate trade-offs
  • Manage the observability backlog using feedback from production deployments and design partners to refine priorities
  • Track and shape our response to emerging observability standards and technologies, including OpenTelemetry (OTel), DCGM GPU metrics, InfiniBand/RoCE fabric counters, storage platform telemetry APIs, and AI workload profiling
  • Define integration strategies for vendor telemetry sources across the ecosystem - NVIDIA compute and BlueField DPUs, storage platforms (VAST, Weka, DDN), workload managers (SLURM), inference stacks, and vector and relational databases - into a unified, operator-facing observability plane
  • Partner with product marketing and field teams on positioning, technical briefs, and reference architectures; represent Mirantis with customers, analysts, and ecosystem partners


Qualifications

  • 5+ years in product management or a senior technical role owning an observability product or operating large-scale monitoring infrastructure
  • Working knowledge of Prometheus, OpenTelemetry, distributed tracing (Jaeger, Tempo), and log aggregation (Loki, Elasticsearch/OpenSearch)
  • Fluency in Kubernetes observability, cloud-native monitoring, or metrics and alerting pipeline architecture
  • Ability to work directly with engineering on technical trade-offs and with field teams in competitive GPU cloud and NeoCloud deals

Strongly Preferred:
  • Exposure to GPU observability, including DCGM metrics, AI workload profiling and performance analysis
  • Familiarity with east-west fabric telemetry - InfiniBand counters, RoCEv2 congestion metrics (ECN, PFC, DCQCN), or switch-level fabric health
  • Experience with high-performance storage telemetry from platforms such as VAST Data, Weka, or DDN, including IOPS, latency, and throughput instrumentation at scale
  • Familiarity with NVIDIA BlueField DPU telemetry, SR-IOV, or offload pipeline observability
  • Exposure to workload-level visibility for SLURM job scheduling, inference serving stacks (vLLM, Triton, TensorRT-LLM), or data service telemetry from vector databases (Milvus, Qdrant) and relational databases in AI pipelines


Additional Information

Why you'll love Mirantis
  • Build the observability foundation for the AI cloud era, working directly with leading GPU cloud operators, NeoClouds, sovereign clouds, and AI-first enterprises
  • Collaborate with a world-class, distributed team committed to openness and technical excellence
  • Shape the product narrative and influence go-to-market success


#remote

We are a Leader for Container Management in G2 (#2 after AWS)!

About Mirantis

Mirantis is a software company that provides cloud computing services and solutions. The company was founded in 2011 and is headquartered in Sunnyvale, California. Mirantis offers a range of cloud computing services, including OpenStack, Kubernetes, and Docker. The company's solutions are used by a variety of industries, including telecommunications, finance, and healthcare. Mirantis has over 1,000 employees and offices in the United States, Russia, Ukraine, and the United Kingdom.
Learn more about Mirantis
Size
1,000 employees
Industry
Founded
2011

Similar Jobs

More Jobs at Mirantis

More Information Technology Jobs

Find similar Technical Product Manager, Observability - remote in the US jobs: