NVIDIA Corporation

Senior Systems Software Engineer - Fleet Debuggability

NVIDIA Corporation$184K — $356K *
Technical Services
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • 10+ years in software with a focus on system software and firmware development.
  • BS, MS, or PhD in CS, CE, EE, or related field, or equivalent experience.
  • Experience in shipping scalable server products or fleet-wide solutions.
  • Proven problem solver with strong communication skills, including for executive reporting.
  • Ability to collaborate effectively across teams and time zones.
  • Proficient in SCM tools (e.g., Git, Perforce) and project management software like Jira.
  • Deep understanding of Linux systems and debugging server platforms.
  • Hands-on experience with out-of-band management protocols and interfaces.

Responsibilities

  • Architect and design fleet-wide log collection and analysis solutions.
  • Develop tools to collect and normalize logs from various sources.
  • Create debug and root-cause tooling to translate fleet logs into actionable insights.
  • Collaborate with developers, QA, and product engineering on end-to-end logging solutions.
  • Manage open-source releases and uphold code quality standards.
  • Write design documentation and oversee project delivery from inception to support.
  • Conduct code reviews and enhance testing practices.
  • Utilize Jira and bug management tools to track and plan engineering efforts.

Benefits

  • Flexible work arrangements to accommodate global team collaboration.
  • Opportunity to work on the cutting edge of technology in a dynamic team environment.
  • Involvement in open-source projects and community engagement.
  • Exposure to advanced tools and methodologies in the field of log analytics.
  • Professional development opportunities through training and mentorship.
Full Job Description
We are the Datacenter System Software team, and we are looking for a highly motivated, creative Senior Engineer o drive Fleet Scale Debuggability end to end. You will design, architect, and build infrastructure, tooling, analytics on how to collect multi-rack scale logs. The solution should normalize, correlate, and reason over logs spanning multiple components, trays, or racks including NVIDIA's GPUs, CPUs, Network products. The logs shall be fetched inband or out of band and should help triage fleet level issues seen by our customers. Your work directly shortens the path from a raw, noisy log stream to an actionable root cause. Join us at the forefront of technological advancement.

Whatyou will be doing:
  • Architect, Design, build, fleet-wide log collection and analysis solutions that aggregate signals across components, trays, and racks.
  • Develop tooling to collect, normalize, and time-align logs from heterogeneous sources - kernel and driver logs, syslog, Redfish event logs, SEL, firmware and BMC logs - over both in-band and out-of-band channels. Build and maintain a log catalog and taxonomy that maps raw log signatures to fault classes, severity, and remediation guidance, so triage is repeatable rather than tribal knowledge.
  • Develop debug and root-cause tooling that turns high-volume fleet logs into ranked, actionable diagnoses for hardware, firmware, and platform faults. Drive the design for collecting and analyzing logs at fleet scale while keeping overhead on production compute nodes low.
  • Partner with all matrixed organizations - developers, SWQA, and product engineering - in a fast-moving environment with end-to-end logging solutions, event schemas, and the contract between log producers and your tooling.
  • Steward the project's open-source release: keep internal and public code paths clean, review community contributions, and represent the tooling in upstream discussions.
  • Write design docs and own end-to-end delivery, working across teams from definition through implementation, debugging, testing, and early customer support.
  • Perform code reviews and partner with development and QA to strengthen unit testing, integration coverage, and test plans.
  • Track work through Jira and bug-management tools and build a realistic end-to-end execution plan in collaboration with other engineers and managers.


What we need to see:
  • 10+ years in the software industry with specialization in system software and/or firmware development.
  • BS, MS, or PhD in CS, CE, EE, or a related technical field - or equivalent experience.
  • Proven track record of shipping scalable server products or fleet-wide experience.
  • A self-starter who loves finding creative solutions to complicated problems, with excellent written and oral communication skills - including executive-level reporting - strong work ethic, and dedication to teamwork.
  • Flexibility to work and communicate effectively across teams, partners, and time zones.
  • Experience with SCM (e.g., Git, Perforce) and project-management tools like Jira. Strong, demonstrable skills in Python or RUST.
  • Deep Linux systems experience: kernel and driver logs, syslog, journald, and the realities of debugging on server platforms.
  • Hands-on experience with out-of-band management and platform interfaces - BMC, Redfish, IPMI, SEL - and an understanding of in-band vs. out-of-band trade-offs.
  • Strong skills in log parsing, normalization, and structured logging, and comfort designing schemas and taxonomies for machine-readable events.


Ways to stand out from the crowd:
  • Experience leading debuggability solutions on sophisticated rack-scale compute architectures like GB200/GB300 NVL72. Familiarity with log and telemetry analytics stacks (e.g., OpenSearch/ELK, Loki, Prometheus, Grafana, PagerDuty) and time-series databases.
  • Hands-on experience with x86/ARM system architecture and coding (C/C++, Python). Experience with SCM (Git, Perforce) and project management tools (Jira).
  • Track record of integrating AI/LLM tooling into engineering workflows - for triage, validation, log analysis, or test generation. Experience standing up follow-the-sun support organizations with measurable response SLAs
  • Experience contributing to or maintaining open-source projects, including managing the boundary between internal and public code.


Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until July 24, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

About NVIDIA Corporation

Nvidia, a global leader in graphics, gaming, and AI technology, offers Nvidia careers and internship opportunities for those passionate about driving innovation in the tech industry. you'll find a company committed to growth, teamwork, and leadership in computer science and machine learning domains.

About Nvidia

A Pioneer in Technology and Innovation

Nvidia has cemented its reputation as a powerhouse in developing advanced graphics processing units (GPUs) and has significantly contributed to the gaming industry's evolution. Moreover, its foray into AI and machine learning has opened new frontiers in technology, making Nvidia a beacon of innovation and a desirable workplace for ambitious tech professionals.

Job Opportunities

Diverse Positions in a Dynamic Field

Nvidia is continuously on the lookout for talented individuals across various domains, including hardware and software engineering, product design, marketing, and sales. Employment opportunities at Nvidia are vast, catering to a wide range of expertise and career aspirations.

Employment in Hardware and Graphics

For those fascinated by the intricacies of hardware and graphics technology, Nvidia offers positions that sit at the forefront of gaming and computing advancements.

Growth in Machine Learning and AI

Nvidia's leadership in AI and machine learning has created numerous vacancies for specialists eager to contribute to groundbreaking projects.

Recruitment in Computer Science

With the constant demand for innovation, Nvidia's recruitment efforts focus on computer science experts capable of pushing the boundaries of what's possible.

Internship Program

Opening Doors to Future Innovators

Nvidia's internship program is designed to nurture the next generation of technology leaders, offering hands-on experience in a culture that celebrates creativity and teamwork.

Benefits and Culture

Interns at Nvidia enjoy a plethora of benefits, from competitive stipends to mentorship opportunities, all within an environment that values growth and learning.

Opportunities for Students

Whether you're an undergraduate, a master's student, or a Ph.D. candidate, Nvidia's internships provide a real-world glimpse into the tech industry, offering valuable experience in various technology fields.

Pathways to Full-Time Employment

Many interns have transitioned into full-time positions, marking the start of successful careers at Nvidia. The internship program is more than a stepping stone into the company; it’s an investment in the professional development of interns. The goal is to ensure that interns are well-equipped for future challenges.

Nvidia Careers: More Than Just a Job

Nvidia offers more than just a job to its employees; it provides a front-row seat on the journey into the future of technology. Nvidia stands as a pillar of innovation with its vast opportunities in hardware, graphics, gaming, machine learning, and computer science. Nvidia careers serve as a launching pad for talented workers who aim to redefine the technological landscape. Whether through full-time positions or internships, joining Nvidia means contributing to a legacy of breakthroughs and becoming part of a global community dedicated to pushing the boundaries of what's possible.
Learn more about NVIDIA Corporation
Size
22,473 employees
Market Cap
$350.4 billion
Industry
Net Income
$4.3 billion
Founded
1993
5 Year Trend
+31.3%
Revenue
$16.6 billion
NASDAQ

Similar Jobs

More Jobs at NVIDIA Corporation

More Technical Services Jobs

Find similar Senior Systems Software Engineer - Fleet Debuggability jobs: