The application window is expected to close on: 11/27/2026
Meet the team:The Observability Center of Excellence (CoE) exists to ensure the internal incident management process is best-in-class. We run one of the world's largest Kubernetes installations and support critical use cases for some of the most demanding customers, so reliability, operational excellence, and fast, effective response are core to how we operate.
We serve as a demanding Customer Zero for our own products, using them in real production conditions, surfacing gaps, and turning what we learn into product improvement requests and customer-relevant best practices. In parallel, the Developer Experience (DevX) team's goal is to improve engineering velocity while ensuring compliance, quality, and governance requirements are met.
Your Impact:As the Engineering Product Manager for the Observability Center of Excellence and Developer Experience, you will serve as the product leader across engineering, SRE, design, quality, security, sales, marketing, and go-to-market teams to establish scalable observability practices and elevate the organization's AI-driven DevX posture through standards, golden paths, tooling, and measurable outcomes.
What You'll Do:- Own the vision, multi-quarter strategy, roadmap, and operating model for the Observability Center of Excellence.
- Own the organization's Developer Experience vision and roadmap, defining the principles, golden paths, and success measures that reduce developer friction and improve engineering velocity, quality, and satisfaction.
- Own the organization's Developer Experience vision and roadmap, defining the principles, golden paths, and success measures that reduce developer friction and improve engineering velocity, quality, and satisfaction.
- Define clear, outcome-based priorities across product, engineering, and operations teams, linking investments to measurable improvements in time-to-market, availability, performance, digital experience, and incident response.
- Lead prioritization and complex trade-off decisions across multiple products and teams; maintain strategic focus while adapting quickly to operational risks and changing needs.
- Partner closely with engineering, SRE, support, and on-call teams to identify monitoring gaps, remove operational blockers, and improve detection, diagnosis, and resolution of performance and availability issues.
- Serve as a trusted technical product and incident-management-process advisor; partner with sales and marketing teams to translate CoE learnings into compelling customer-facing artifacts.
- Establish feedback loops that turn operational data, incident learnings, and internal product usage into product improvements, reference practices, and reusable guidance.
- Foster candid communication, healthy debate, clear decision-making, and accountability across teams working on shared reliability and platform outcomes.
- Conduct market, competitive, and internal-user validation to identify product gaps, differentiators, adjacencies, and new opportunities for the observability portfolio.
- Influence platform strategy across an 18-24 month horizon and align multiple teams around a cohesive product experience.
- Evangelize the Observability CoE internally, creating a community of practice that enables product teams to adopt, learn from, and contribute to its evolution.
Minimum Requirements:- Bachelor's degree plus 12 years of related experience, Master's degree plus 8 years of related experience, or PhD plus 5 years of related experience.
- Product management experience for complex SaaS, platform, developer, cloud, or observability products.
- Hands-on experience with one or more of: digital experience monitoring, infrastructure monitoring, or application performance monitoring.
- Experience using data, telemetry, and customer or user insights to shape product strategy, priorities, and measurable outcomes.
- Experience leading cross-functional teams through ambiguous, technically complex decisions.
Preferred Qualifications:- Experience establishing or leading an observability, reliability, developer productivity, or operational-excellence center of excellence or community of practice.
- Close working experience with teams responsible for performance and availability incident management, SRE, reliability engineering, or production operations.
- Knowledge of software-development and cloud-deployment tooling, such as CI/CD, infrastructure as code, containers, Kubernetes, public-cloud platforms, and developer workflows. Experience with AI-assisted development workflows.
- Experience defining or working with service-level indicators/objectives, alerting, incident analytics, root-cause analysis, and reliability metrics.
- Ability to influence senior engineering and product leaders, align teams around shared priorities, and drive execution with quality and urgency.
- Experience applying automation or AI tools to improve product, engineering, or operational workflows responsibly.