Responsibilitie
About the Team We are a systems software team building the foundational software for large-scale compute platforms. We work at the hardware/software boundary across the Linux kernel, accelerators, storage, firmware, and platform validation. We value rigorous engineering, clear interfaces, measurable performance and reliability, and upstream collaboration where appropriate. The team partners closely with hardware, architecture, product, validation, and production engineering groups to move new capabilities from design through dependable deployment. About the Role You will lead the technical strategy and cross-functional delivery of the low-level GPU software stack for data-center computing. You will connect accelerator architecture, firmware, kernel drivers, runtimes, communication, telemetry, and fleet reliability into a coherent roadmap. This is a hands-on individual-contributor role: you will make architecture decisions, guide engineers, review critical implementations, and resolve system-level issues without assuming hiring or performance-management responsibilities. Responsibilities - Define the architecture and multi-year technical roadmap for GPU system software across driver, runtime, firmware-interface, management, observability, and reliability layers. - Set priorities and technical standards for accelerator enablement, balancing near-term product commitments with compatibility, performance, serviceability, and long-term maintenance. - Own end-to-end integration from architecture and pre-silicon planning through bring-up, qualification, general availability, fleet monitoring, and sustained operation. - Lead cross-functional execution among silicon, firmware, kernel, compiler, library, machine-learning framework, server, network, storage, validation, and production teams. - Resolve ambiguous system-level tradeoffs involving APIs, resource management, memory and interconnect topology, telemetry, recovery, security, and workload performance. - Remain technically involved through prototypes, critical-path code and design reviews, performance analysis, and leadership during high-severity failure investigations. - Define measurable release and reliability criteria, including qualification coverage, regression thresholds, fault containment, automated repair, and fleet-health indicators. - Mentor engineers, raise the quality of architecture and debugging practices, and communicate technical decisions to both specialist and executive audiences.
Qualification
Minimum Qualifications - Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent practical experience. - 5+ years of hands-on experience in GPU, accelerator, kernel, firmware, runtime, or data-center system software, including technical leadership on complex hardware/software programs. - Demonstrated success setting technical direction and leading a major hardware/software program from concept through production. - Deep understanding of accelerator architecture, memory hierarchy, interconnects, operating systems, device management, and production reliability. - Experience coordinating dependencies and technical decisions across multiple engineering teams and organizational boundaries. - Strong coding, design-review, performance-analysis, and system-debugging capability, with evidence of continued hands-on contribution. Preferred Qualifications - Experience with a major GPU computing stack and its kernel driver, runtime, libraries, tooling, and fleet-management interfaces. - Experience defining platform APIs or hardware/software contracts across multiple accelerator generations. - Knowledge of distributed accelerator workloads, collective communication, topology-aware placement, and large-scale serviceability. - Track record mentoring senior engineers and building alignment where priorities, ownership, or technical evidence initially conflict.
Job Information
【For Pay Transparency】Compensation Description (Annually)
The base salary range for this position in the selected city is $218400 - $480000 annually.
Compensation may vary outside of this range depending on a number of factors, including a candidate's qualifications, skills, competencies and experience, and location. Base pay is one part of the Total Package that is provided to compensate and recognize employees for their work, and this role may be eligible for additional discretionary bonuses/incentives, and restricted stock units.
Benefits may vary depending on the nature of employment and the country work location. Employees have day one access to medical, dental, and vision insurance, a 401(k) savings plan with company match, paid parental leave, short-term and long-term disability coverage, life insurance, wellbeing benefits, among others. Employees also receive 10 paid holidays per year, 10 paid sick days per year and 17 days of Paid Personal Time (prorated upon hire with increasing accruals by tenure).
The Company reserves the right to modify or change these benefits programs at any time, with or without notice.