NVIDIA Corporation

Senior Software Engineer, AIOps and Observability

NVIDIA Corporation$200K — $322K *
Information Technology
11 - 15 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in computer science, engineering, or related field, or equivalent experience.
  • 12+ years of experience in product development and full stack engineering.
  • 5+ years of experience in developing and operating observability platforms in a cloud-native environment.
  • Strong knowledge of observability tools, such as Prometheus and Grafana.
  • Hands-on knowledge of AIOps tools like BigPanda and Datadog.
  • Experience with Docker, Kubernetes, and microservices architectures.
  • Proficient in programming languages such as Go, Python, Java, or C#.

Responsibilities

  • Lead the design, development, and deployment of AIOps & Observability platforms.
  • Drive the technical vision and roadmap for observability initiatives.
  • Collaborate with teams to understand their observability needs and provide solutions.
  • Establish and implement observability standards and processes across NVIDIA.
  • Provide peer reviews and feedback on engineering performance and security.
  • Work with data scientists to implement machine learning models for anomaly detection.
  • Develop AI agents and observability tools to enhance issue resolution.

Benefits

  • Equity participation.
  • Comprehensive benefits programs.
  • Exposure to cutting-edge technology in AI and observability.
  • Opportunities to mentor and lead other engineers.
Full Job Description
We are looking for a highly skilled Senior Software Engineer to design and develop AIOps & Observability platforms at NVIDIA. The platforms are used by internal teams to monitor, diagnose, and optimize the products, millions of assets and services in cloud, on-prem, data centers, supply chain, and edge. You will work with a team of engineers, product managers, and partners to define the observability strategy, roadmap, and standard methodologies for NVIDIA. You will also mentor and coach other engineers on observability, machine learning, tools and techniques.

What you will be doing:
  • Lead the design, development, and deployment of AIOps & Observability platforms, including metrics, logs, traces, events, alerts, dashboards, and visualizations.
  • Drive the technical vision and roadmap for AIOps and Observability initiatives, aligning with business goals and industry best practices.
  • Collaborate with other teams and customers to understand their observability needs and provide solutions that meet their requirements and expectations.
  • Establish and implement observability standards, guidelines, and processes across NVIDIA. Research, evaluate, and adopt new observability technologies and frameworks that can enhance user experience.
  • Provide peer reviews to other engineers including feedback on performance, scalability, security and correctness.
  • Work with Data scientists to implement machine learning models for anomaly detection, forecasting, and root cause analysis on logs, metrics, and events. Handle large volumes of data and ensure data quality, security, and compliance.
  • Develop and operate scalable, reliable, and distributed systems that can handle high traffic and complex workloads.
  • Develop AI agents and AI-native observability tools that help engineers detect, understand, and resolve production issues faster. Build agentic workflows that reason across logs, metrics, traces, events, alerts, topology, and incident history to support anomaly detection, forecasting, root cause analysis, automated debugging, and remediation recommendations.


What we need to see:
  • Bachelor's degree in computer science and engineering, or related field, or equivalent experience.
  • 12+ years of experience in product development and full stack engineering, with 5+ years of experience in developing and operating observability platforms and solutions, preferably in a cloud-native environment.
  • Strong knowledge and experience with observability tools, such as Prometheus, Victoria Metrics, Vector, Loki, Grafana, Alert Manager, Clickhouse, OpenTelemetry, etc.
  • Hands-on knowledge in AIOps tools such as BigPanda, PagerDuty, Datadog, etc.
  • Experience with Kubernetes, Nomad, Docker, and microservices architectures as well as experience with streaming services to ingest billions of events using NATS, Kafka, etc
  • Proficient in one or more programming languages, such as Go, Python, Java, C#, etc.
  • Passionate about observability and delivering high-quality internal platforms.
  • Experience with developing Observability solutions to monitor On-prem and Public Cloud environments.
  • Experience with running large Observability platforms on BareMetal Infrastructure
  • Establish scalable data pipelines and instrumentation for collecting, aggregating, and visualizing telemetry and operational metrics.


Ways To Stand Out From The Crowd:
  • Deep understanding of implementing Observability solutions to large scale on-prem Infrastructure and Networking.
  • Hands-on experience with managing large scale Observability Platforms with LLMs & ML Models and building custom services to ingest billions of metrics and logs from wide range of assets.
  • Developed unified cloud observability platform to monitor Network, Compute, Power, Storage, Operating Systems, Security, Applications, SaaS Platforms.
  • Demonstrated experience and expertise in using machine learning and Generative AI to develop solutions such as predictive monitoring, incident diagnosis, summarization and correlation.
  • Demonstrate proficiency in AI/ML systems, generative AI, or agentic AI frameworks.


Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 200,000 USD - 322,000 USD.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until August 4, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

About NVIDIA Corporation

Nvidia, a global leader in graphics, gaming, and AI technology, offers Nvidia careers and internship opportunities for those passionate about driving innovation in the tech industry. you'll find a company committed to growth, teamwork, and leadership in computer science and machine learning domains.

About Nvidia

A Pioneer in Technology and Innovation

Nvidia has cemented its reputation as a powerhouse in developing advanced graphics processing units (GPUs) and has significantly contributed to the gaming industry's evolution. Moreover, its foray into AI and machine learning has opened new frontiers in technology, making Nvidia a beacon of innovation and a desirable workplace for ambitious tech professionals.

Job Opportunities

Diverse Positions in a Dynamic Field

Nvidia is continuously on the lookout for talented individuals across various domains, including hardware and software engineering, product design, marketing, and sales. Employment opportunities at Nvidia are vast, catering to a wide range of expertise and career aspirations.

Employment in Hardware and Graphics

For those fascinated by the intricacies of hardware and graphics technology, Nvidia offers positions that sit at the forefront of gaming and computing advancements.

Growth in Machine Learning and AI

Nvidia's leadership in AI and machine learning has created numerous vacancies for specialists eager to contribute to groundbreaking projects.

Recruitment in Computer Science

With the constant demand for innovation, Nvidia's recruitment efforts focus on computer science experts capable of pushing the boundaries of what's possible.

Internship Program

Opening Doors to Future Innovators

Nvidia's internship program is designed to nurture the next generation of technology leaders, offering hands-on experience in a culture that celebrates creativity and teamwork.

Benefits and Culture

Interns at Nvidia enjoy a plethora of benefits, from competitive stipends to mentorship opportunities, all within an environment that values growth and learning.

Opportunities for Students

Whether you're an undergraduate, a master's student, or a Ph.D. candidate, Nvidia's internships provide a real-world glimpse into the tech industry, offering valuable experience in various technology fields.

Pathways to Full-Time Employment

Many interns have transitioned into full-time positions, marking the start of successful careers at Nvidia. The internship program is more than a stepping stone into the company; it’s an investment in the professional development of interns. The goal is to ensure that interns are well-equipped for future challenges.

Nvidia Careers: More Than Just a Job

Nvidia offers more than just a job to its employees; it provides a front-row seat on the journey into the future of technology. Nvidia stands as a pillar of innovation with its vast opportunities in hardware, graphics, gaming, machine learning, and computer science. Nvidia careers serve as a launching pad for talented workers who aim to redefine the technological landscape. Whether through full-time positions or internships, joining Nvidia means contributing to a legacy of breakthroughs and becoming part of a global community dedicated to pushing the boundaries of what's possible.
Learn more about NVIDIA Corporation
Size
22,473 employees
Market Cap
$350.4 billion
Industry
Net Income
$4.3 billion
Founded
1993
5 Year Trend
+31.3%
Revenue
$16.6 billion
NASDAQ

Similar Jobs

More Jobs at NVIDIA Corporation

More Information Technology Jobs

Find similar Senior Software Engineer, AIOps and Observability jobs: