NVIDIA Corporation

Data and Platform Engineer

NVIDIA Corporation$200K — $322K *
US-Anywhere
+ 4 other locationsRemote
Information Technology
11 - 15 years of experience
Job Overview by Ladders

Qualifications

  • BS or MS in Computer Science, Engineering, or a related field (or equivalent experience) with 12+ years in the industry.
  • Proven history of building and managing production software and data platforms or distributed systems.
  • Experience defining technical plans, guiding engineers, and overseeing project delivery across teams.
  • Deep knowledge of Apache Spark or similar distributed processing frameworks, databases, and cloud platforms.
  • Proficiency in Python and SQL, with a focus on writing production-level code and debugging.
  • Strong SQL and data modeling expertise, including query execution and system consistency.
  • Ability to lead complex investigations across teams, applying analytical tools to establish root causes.

Responsibilities

  • Own and define the architecture and roadmap for a significant platform component.
  • Lead technical delivery and coordinate efforts across multiple teams to manage complex projects.
  • Design and implement data pipelines for real-time and batch processing to support operational decision-making.
  • Direct the development of shared platform capabilities and engineering standards.
  • Guide investigations into production issues, driving resolution of critical challenges.
  • Establish data quality and operational standards, working with security teams to enhance system integrity.
  • Deliver user-friendly data models, APIs, and applications to facilitate data-driven decisions.

Benefits

  • Opportunity to work on cutting-edge technology in a leading AI firm.
  • Possibility for continued education and professional development.
  • Flexible work environment supporting work-life balance.
  • Access to health and wellness programs.
  • Potential for career advancement in a growing technology sector.
Full Job Description
NVIDIA's DGX Cloud organization is seeking a Senior Data Engineer to become part of its data team! We develop the reliable data foundation that supports fleet health, capacity, utilization, cost, reliability, and operational decision-making throughout DGX Cloud. Our platform supports engineering, operations, finance, and product teams managing and expanding large GPU fleets across cloud service providers and NVIDIA Cloud Partners. We are looking for a practical engineer and technical lead to take charge of a key part of the Navigator data platform. We develop the systems that transform distributed infrastructure telemetry and operational data into dependable, managed data products that support fleet health, capacity, utilization, cost, and operational decisions.

You will be responsible for the architecture, technical plan, and production results of a major platform area like ingestion and orchestration, data quality and reconciliation, or data serving and consumption. You will clarify requirements with customers, define technical objectives, guide design and development among engineers and partner teams, and stay actively engaged in coding, debugging, and production tasks. Successful candidates have already led complex technical work across team boundaries and delivered improvements that other groups adopted. Our primary implementation environment is Python, SQL, Databricks, and Spark.

What you'll be doing:
  • You will own a major platform component and its roadmap. For example, define its architecture, interfaces, technical goals, and evolution. Anticipate capacity, compatibility, and operational needs over a multi-year horizon, and translate them into achievable breakthroughs that balance immediate delivery with long-term maintainability.
  • Lead technical delivery across teams. Work with customers and interested parties to clarify vague requirements. Break down design and implementation work for contributing engineers. Establish release turning points and manage dependencies and delivery risks. Guide the work process, revise plans when requirements shift, and keep management and partner teams informed and aligned.
  • Build data pipelines and products. Plan and carry out batch and streaming ingestion, transformation, reconciliation, and serving processes for fleet, capacity, utilization, cost, scheduling, and operational telemetry. Establish data models and agreements that remain stable as sources, consumers, and scale progress.
  • Develop shared platform capabilities. Direct the creation and adoption of libraries, workflow and DAG or comparable experience abstractions, deployment tools, and standard implementation approaches. Partner with related teams to solve shared challenges and evaluate progress in onboarding time, engineering effort, reliability, and cost.
  • Lead complex production investigations. Serve as the technical point of accountability for issues spanning pipelines, applications, SQL engines, Spark, storage, networks, and cloud services. Coordinate investigations across owners, drive resolution of release blockers and critical issues from partners, and implement preventive measures.
  • Define quality, security, and operational expectations. Establish and implement testing, data-quality, reconciliation, lineage, SLO, and release-readiness standards for your platform area. Partner with security and infrastructure teams on trust boundaries, service identities, least privilege, secrets, environment isolation, and auditability, and drive adoption across contributing teams.
  • Make trusted data usable. Deliver well-modeled tables, APIs, automation, dashboards, and focused internal applications. Align with consumers on semantics, access patterns, freshness, compatibility, and ownership so that shared capabilities support dependable operational decisions.
  • Provide technical leadership through others. Guide design reviews, mentor engineers taking on larger ownership, and resolve technical disagreements using evidence and clear tradeoffs. Partner with leadership on priorities and explain how technical investments support DGXC objectives.


What we need to see:
  • BS or MS in Computer Science, Engineering, or a related field (or equivalent experience), and at least 12+ years of equivalent experience
  • A sustained record of building and operating production software, data platforms, databases, or distributed systems. This includes owning a major component or complex project from requirements and architecture through release and ongoing operation.
  • Proven ability to outline a component's technical plan, establish objectives for engineers, assign design and implementation tasks, and guide delivery within your team and nearby teams with little supervision.
  • Extensive practical experience in one or more of these areas: distributed processing using Spark or a similar system; relational, distributed, or analytical databases; production ETL, change-data capture, streaming, or event handling; or backend and cloud platforms managing large data volumes. You must grasp the interfaces and failure modes of nearby layers thoroughly to inform solid architectural choices.
  • Strong software-engineering fundamentals and production proficiency in Python or another backend or systems language, with the ability and willingness to work primarily in Python and SQL. Experience designing reusable abstractions, reviewing substantial changes, and personally implementing and debugging critical code paths.
  • Strong SQL and data-modeling skills, with practical depth in query execution, incremental processing, schema evolution, consistency, and analytical consumption. Ability to reason about idempotency, replay, late-arriving data, partial failure, and correctness across system boundaries.
  • Experience leading complex investigations involving multiple components and teams. Ability to use logs, metrics, traces, query plans, profiles, and controlled experiments to establish root cause, coordinate resolution, and prevent recurrence.
  • Demonstrated architectural judgment: evaluating alternatives, anticipating future requirements, and balancing reliability, performance, cost, security, compatibility, and maintainability. Experience leading significant migrations or architectural changes while preserving production service.
  • Experience establishing production quality and operational practices that other engineers adopt, including testing, CI/CD, monitoring, alerting, rollback, incident response, and secure deployment.
  • Proven success in influencing technical decisions without official authority, advising engineers outside your immediate project, and advancing workflow improvements across closely related teams. Ability to simplify complex issues, offer a course of action, and communicate decisions and delivery risks clearly.


Ways to stand out from the crowd:
  • Proven expertise in building, refining, and running Databricks, Apache Spark, PySpark, Spark SQL, Delta Lake, or Unity Catalog workloads along with shared platform features.
  • Experience designing and operating Kafka or comparable streaming systems, including partitioning, consumer behavior, offset management, backpressure, replay, and schema compatibility.
  • Experience scaling, migrating, or tuning relational, distributed, time-series, object-storage, or information retrieval systems, including Elasticsearch or OpenSearch.
  • Background operating compute or GPU clusters, or working with Kubernetes, Slurm, cloud infrastructure, and fleet telemetry across AWS, Azure, GCP, or other providers.
  • Experience building production agentic systems or agent harnesses, including tool integration, context management, evaluation, permissions, observability, and failure recovery. Evidence of measurable improvements in engineering productivity or operational outcomes.


Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 200,000 USD - 322,000 USD.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until September 25, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

About NVIDIA Corporation

Nvidia, a global leader in graphics, gaming, and AI technology, offers Nvidia careers and internship opportunities for those passionate about driving innovation in the tech industry. you'll find a company committed to growth, teamwork, and leadership in computer science and machine learning domains.

About Nvidia

A Pioneer in Technology and Innovation

Nvidia has cemented its reputation as a powerhouse in developing advanced graphics processing units (GPUs) and has significantly contributed to the gaming industry's evolution. Moreover, its foray into AI and machine learning has opened new frontiers in technology, making Nvidia a beacon of innovation and a desirable workplace for ambitious tech professionals.

Job Opportunities

Diverse Positions in a Dynamic Field

Nvidia is continuously on the lookout for talented individuals across various domains, including hardware and software engineering, product design, marketing, and sales. Employment opportunities at Nvidia are vast, catering to a wide range of expertise and career aspirations.

Employment in Hardware and Graphics

For those fascinated by the intricacies of hardware and graphics technology, Nvidia offers positions that sit at the forefront of gaming and computing advancements.

Growth in Machine Learning and AI

Nvidia's leadership in AI and machine learning has created numerous vacancies for specialists eager to contribute to groundbreaking projects.

Recruitment in Computer Science

With the constant demand for innovation, Nvidia's recruitment efforts focus on computer science experts capable of pushing the boundaries of what's possible.

Internship Program

Opening Doors to Future Innovators

Nvidia's internship program is designed to nurture the next generation of technology leaders, offering hands-on experience in a culture that celebrates creativity and teamwork.

Benefits and Culture

Interns at Nvidia enjoy a plethora of benefits, from competitive stipends to mentorship opportunities, all within an environment that values growth and learning.

Opportunities for Students

Whether you're an undergraduate, a master's student, or a Ph.D. candidate, Nvidia's internships provide a real-world glimpse into the tech industry, offering valuable experience in various technology fields.

Pathways to Full-Time Employment

Many interns have transitioned into full-time positions, marking the start of successful careers at Nvidia. The internship program is more than a stepping stone into the company; it’s an investment in the professional development of interns. The goal is to ensure that interns are well-equipped for future challenges.

Nvidia Careers: More Than Just a Job

Nvidia offers more than just a job to its employees; it provides a front-row seat on the journey into the future of technology. Nvidia stands as a pillar of innovation with its vast opportunities in hardware, graphics, gaming, machine learning, and computer science. Nvidia careers serve as a launching pad for talented workers who aim to redefine the technological landscape. Whether through full-time positions or internships, joining Nvidia means contributing to a legacy of breakthroughs and becoming part of a global community dedicated to pushing the boundaries of what's possible.
Learn more about NVIDIA Corporation
Size
22,473 employees
Market Cap
$350.4 billion
Industry
Net Income
$4.3 billion
Founded
1993
5 Year Trend
+31.3%
Revenue
$16.6 billion
NASDAQ

Similar Jobs

More Jobs at NVIDIA Corporation

More Information Technology Jobs

Find similar Data and Platform Engineer jobs: