About the Team
The Data Platform and Observability Engineering (DPOE) team is building Workday's next-generation, multi-petabyte scale Observability Platform. We own the libraries, distributed services, and infrastructure that power ingestion, storage, and query across the observability stack — Iceberg, ClickHouse, Tempo, Mimir, Grafana, S3, Kafka, and Elasticsearch — serving traces, metrics, and logs for every workload at Workday.
How we help our customers: Every engineering, SRE, and product team at Workday depends on us to see inside their systems — from a single service's latency spike to a cross-service cascading failure impacting thousands of customers. By owning the full observability data lifecycle at petabyte scale, we give internal teams the speed and confidence to find and fix issues before they affect Workday's customers. Our roadmap directly shapes how the company detects, diagnoses, and eventually predicts operational issues at scale — turning observability from a reactive debugging tool into a proactive, AI-assisted safety net for every workload running on Workday's platform.
About the Role
We are looking for a hands-on, technical Senior Product Manager (P4) to own and drive our Observability strategy, with a strong emphasis on Distributed Tracing. You will define the product vision for how engineers understand, debug, and optimize complex distributed systems, with a particular focus on buildingAI-powered detection and triage capabilities that reduce time-to-detect and time-to-resolve production issues. This role requires someone comfortable diving deep into technical architecture discussions, reading code/traces, and partnering closely with engineering and applied ML teams to ship technically sound, high-impact products.
What You'll Do- Own the product vision, strategy, and roadmap for Observability, with a primary focus on Distributed Tracing capabilities (trace context propagation, sampling strategies, span analysis, service maps, latency/error analysis, etc.)
- Define and drive the roadmap forAI-enabled anomaly detection, including specifying requirements for statistical and ML-based detection methods (e.g., time-series forecasting, seasonality-aware baselining, change-point detection, multivariate anomaly detection across correlated metrics/traces/logs)
- Partner with ML engineering to define model requirements, evaluation metrics (precision/recall, false-positive rate, alert-to-incident ratio), and feedback loops for continuous model improvement
- Define requirements forLLM-based root cause summarization and triage assistance — e.g., generating human-readable incident summaries from raw trace/log/metric data, ranking probable root causes, suggesting remediation runbooks based on historical incident patterns
- Specify how confidence scores, explainability, and human-in-the-loop review are surfaced in the triage workflow so on-call engineers can trust and act on AI-generated recommendations
- Partner closely with engineering teams to define technical requirements, evaluate architectural trade-offs, and make hands-on contributions to product decisions (e.g., reviewing design docs, understanding OpenTelemetry/OTel standards, tracing protocols, and instrumentation approaches)
- Define and track success metrics (detection precision/recall, MTTD, MTTR, alert noise reduction, triage automation rate, trace coverage) to measure product impact
- Conduct customer and internal stakeholder research to identify pain points in debugging, alert fatigue, and incident triage
- Write clear, detailed product requirements, user stories, and specs; work closely with design, engineering, and data science to bring them to life
- Stay current on industry trends in observability (OpenTelemetry, eBPF-based tracing, service mesh telemetry) and applied AI/ML for anomaly detection, incident correlation, and LLM-based operational tooling
About You
Basic Qualifications
- 8+ years of product management experience, with meaningful time spent in Observability, Monitoring, APM, or related infrastructure/developer tooling domains
- Deep, demonstrable expertise in Distributed Tracing concepts and technologies (e.g., OpenTelemetry, Jaeger, Zipkin, trace sampling, span/context propagation)
- Strong technical background — comfortable reading code, understanding system architecture, APIs, and engaging directly in technical design discussions
- Concrete, hands-on understanding ofAI/ML techniques applied to detection and triage, including:
- Anomaly detection methods (statistical thresholds, time-series forecasting models, change-point/seasonality detection)
- Alert correlation and clustering techniques (embedding/vector similarity, graph-based dependency analysis, unsupervised clustering)
- LLM-based summarization and reasoning for incident root-cause analysis and runbook suggestion
- Model evaluation frameworks (precision/recall, false-positive/false-negative trade-offs, drift monitoring)
- Experience partnering directly with ML engineering teams to translate detection and triage requirements into shippable models and features
- Proven track record of shipping complex, technical products from concept to launch
- Excellent written and verbal communication skills; ability to translate complex technical and ML concepts for varied audiences
- Strong analytical and data-driven decision-making skills
Other Qualifications- Prior experience as a software engineer, SRE, ML engineer, or data scientist before transitioning to product management
- Familiarity with observability platforms (Datadog, New Relic, Honeycomb, Grafana, Splunk, Dynatrace, etc.) and their AI/ML-driven detection features
- Direct experience productionizing ML models (e.g., anomaly detectors, classifiers, LLM-based tools) for operational/incident-management use cases
- Familiarity with vector databases/embedding search, graph-based correlation techniques, or LLM prompt/agent design for triage automation
- Background in large-scale, cloud-native, microservices-based systems
Workday Pay Transparency Statement
The annualized base salary ranges for the primary location and any additional locations are listed below. Workday pay ranges vary based on work location. As a part of the total compensation package, this role may be eligible for the Workday Bonus Plan or a role-specific commission/bonus, as well as annual refresh stock grants. Recruiters can share more detail during the hiring process. Each candidate’s compensation offer will be based on multiple factors including, but not limited to, geography, experience, skills, job duties, and business need, among other things. For more information regarding Workday’s comprehensive benefits, please .
Primary Location: USA.CA.PleasantonPrimary Location Base Pay Range: $168,000 USD - $252,000 USD
Additional US Location(s) Base Pay Range: $140,600 USD - $252,000 USD
Our Approach to Flexible Work
With Flex Work, we’re combining the best of both worlds: in-person time and remote. Our approach enables our teams to deepen connections, maintain a strong community, and do their best work. We know that flexibility can take shape in many ways, so rather than a number of required days in-office each week, we simply spend at least half (50%) of our time each quarter in the office or in the field with our customers, prospects, and partners (depending on role). This means you'll have the freedom to create a flexible schedule that caters to your business, team, and personal needs, while being intentional to make the most of time spent together. Those in our remote "home office" roles also have the opportunity to come together in our offices for important moments that matter.