We are looking for an engineering manager to lead the team responsible for operational processes and platforms across incident management, maintenance, on-call, issue management, and customer-serving readiness.
You will own the roadmap and delivery, from defining how teams work to building the tools they use. You will partner with infrastructure and service owners to improve reliability, reduce manual work, and ensure services are ready to support customers. Your team will use automation, AI, and lessons from operational events to drive improvements.
What you'll be doing:- Lead a team and own the roadmap for operational processes and platforms, from requirements and delivery through adoption and results.
- Set technical direction, prioritize work, and guide execution across engineering and operational disciplines.
- Partner with infrastructure, product, and security teams to establish consistent practices for incident response, maintenance, on-call, issue management, and customer-serving readiness.
- Hire and develop engineers and technical leads, building a team with clear ownership and accountability.
- Align priorities across teams, communicate progress and risks, and provide technical leadership during major incidents.
What we need to see:- BS degree or equivalent experience with 10+ overall years of software engineering or related experience, including 5+ years of engineering leadership managing teams or complex technical programs.
- Knowledge of operational processes and supporting platforms, including roadmap, delivery, adoption, and improvement.
- Strong technical judgment in software architecture, platform integration, and engineering tradeoffs.
- Clear communication with engineers, cross-functional partners, and executive stakeholders.
- A record of developing engineers, growing teams, and delivering results under pressure.
Ways to stand out from the crowd:- Established readiness standards covering service ownership, support coverage, and reliability objectives.
- Built, integrated, and scaled platforms pertaining the incident, maintenance, customer experience management, along with on-call and production readiness
- Applied AI or LLMs to improve triage, knowledge retrieval, incident analysis, or automation.
- Supported EDA, large-scale compute, or hybrid infrastructure with complex dependencies and demanding availability requirements.
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 224,000 USD - 356,500 USD for Level 3, and 272,000 USD - 431,250 USD for Level 4.
You will also be eligible for equity and benefits.
Applications for this job will be accepted at least until September 26, 2026.
This posting is for an existing vacancy.
NVIDIA uses AI tools in its recruiting processes.