Distinguished Technologist, Agentic AIOps & Self-Driving Networks
This role has been designed as ''Onsite'' with an expectation that you will primarily work from an HPE office.
Job Description:Key responsibilities- Own the E2E assurance architecture: cross-domain telemetry normalization, correlation, root cause, guarded remediation, and outcome verification.
- Define the shared substrate - a common model spanning intent, topology, config, identity, policy, and state
- Architect the agentic layer: planning and tool-use agents, MCP tool servers, and multi-agent orchestration across domain boundaries.
- Own the safety architecture: approval gates, policy enforcement at tool boundaries, remediation safety tiers, blast-radius limits, and deterministic rollback.
- Set the discipline for when to use deterministic graph reasoning, classical ML, or generative AI - and when not to.
- Establish evaluation as a shipped discipline - golden fault corpora, replayable incident harnesses, and MTTR and false-escalation rate as product metrics.
- Carry the strategy across BUs and influence executive and product stakeholders toward one substrate.
- Mentor engineers and develop the next generation of principal and distinguished technical leaders.
Requirements- 15+ years in engineering and architectural leadership building and operating large distributed systems in production.
- Breadth across networking domains - campus, data center, WAN/SD-WAN, and security - with proven depth in at least three.
- Deep networking fundamentals: EVPN-VXLAN, BGP, SAI-level behavior, overlay/underlay, identity and policy enforcement, and streaming telemetry (gNMI/OpenConfig) at scale.
- Graph systems depth: knowledge or property graph modeling and graph algorithms applied to dependency, impact, and causality.
- Applied AI on both sides: LLM agents, tool calling, MCP, and agent evaluation, plus classical ML and time-series modeling - and the judgment to know which the problem calls for.
- Production AIOps or observability experience: correlation, anomaly detection, noise reduction, incident lifecycle.
- Cloud-native at scale: Kubernetes, microservices, Kafka, time-series/OLAP stores, hybrid and public cloud.
- Strong hands-on coding in Python and Go, sustained through design and architecture. Proficient with AI-first, spec-driven prototyping.
- Demonstrated impact at Principal or Distinguished level spanning multiple organizations, including the ability to run discovery across teams that don't report to you.
- Bachelor's, Master's, or PhD in computer science, engineering, or a related discipline preferred.
What We Can Offer You:Health & WellbeingWe strive to provide our team members and their loved ones with a comprehensive suite of benefits that supports their physical, financial and emotional wellbeing.
Personal & Professional DevelopmentWe also invest in your career because the better you are, the better we all are. We have specific programs catered to helping you reach any career goals you have - whether you want to become a knowledge expert in your field or apply your skills to another division.
Unconditional InclusionWe are unconditionally inclusive in the way we work and celebrate individual uniqueness. We know varied backgrounds are valued and succeed here. We have the flexibility to manage our work and personal needs. We make bold moves, together, and are a force for good.
Let's Stay Connected:#unitedstates
#executive, #networking
Job:Engineering
Job Level:TCP_07
"The expected salary/wage range for this position is provided below. Actual offer may vary from this range based upon geographic location, work experience, education/training, and/or skill level.
- United States of America: Annual Salary USD 194,000 - 412,500 in California
The listed salary range reflects base salary. Variable incentives may also be offered."
Information about employee benefits offered in the US can be found at https://myhperewards.com/main/new-hire-enrollment.html