GPU Application Platform Engineer Graduate (Server Research and Development) - 2027 Start (PhD)

ByteDance

$162K — $316K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • PhD in Electrical Engineering, Computer Engineering, Computer Science, or related fields.
  • Thesis on GPU/AI platform architecture and performance optimization.
  • Strong understanding of computer system architecture, particularly GPU/AI SoC or Platform Architecture.
  • Experience in performance optimization and software/hardware co-design for GPU/AI systems.
  • Proficiency in software development with C/C++ and scripting languages (e.g., Python).
  • Knowledge of GPU/AI virtualization technology and distributed systems.
  • Familiarity with LLM model architecture and accelerator/memory/network training requirements.

Responsibilities

  • Develop benchmarks and optimization tools for GPU/AI systems, focusing on large model training and inference.
  • Identify system bottlenecks through in-depth data analysis and co-design efforts.
  • Create a Total Cost of Ownership model for GPU/AI systems based on performance assessments.
  • Engage with industry groups to stay ahead on emerging standards and technically contribute to research.
  • Collaborate with technology partners for proof-of-concept setups to test and evaluate new innovations.

Benefits

  • Medical, dental, and vision insurance starting from day one.
  • 401(k) plan with company matching.
  • Paid parental leave.
  • Short-term and long-term disability coverage.
  • Life insurance and wellbeing benefits.
  • 10 paid holidays and 10 paid sick days annually.
  • 17 days of Paid Personal Time, increasing with tenure.
Full Job Description
Responsibilitie

About the team: Server Research and Development team is responsible for architecting, designing, and building the best server and storage system to meet the requirements of high-performance, low cost and easy to operate. By joining this team, you will work with the best engineers and talents in this industry and have a broad opportunity to get in touch with the latest AI application system and newly emerged technology in computing , storage and silicon validation. You will gain remarkable hardware architect, development and validation experience in the most advanced hardware infrastructure at a massive scale. We are looking for a self-motivated GPU/AI Application Platform Architect with focus on giant model system optimization. We are looking for talented individuals to join our team. As a graduate, you will get opportunities to pursue bold ideas, tackle complex challenges, and unlock limitless growth. Successful candidates must be able to commit to an onboarding date by the end of the year. Please state your availability and graduation date clearly in your resume. Responsibilities: - Develop application benchmarks, tools and performance optimization methods for GPU/AI systems, including giant model training and inference systems such as LLM. - Identify the system bottleneck/opportunity with deep system-level data-driven study, explore innovative options through SW-HW co-design, and lead them towards implementation to improve training and inference system efficiency. - Develop GPU/AI system TCO model, based on application benchmark and performance optimization. - Work with industry consortiums and open standard committees to investigate the emerging standards or technologies, and to contribute our research results to the industry. - Work with our technology partners and suppliers to setup POC or prototypes to evaluate and test the new technologies or architectural designs.

Qualification

Minimum Qualifications: - Individuals who are completing or have recently completed a PhD degree in Electrical Engineering, Computer Engineering, Computer Science or related majors. - Thesis topics in GPU/AI platform architecture and/or application performance optimization design or software hardware co-design. - Deep understanding of computer system architecture, especially on GPU/AI SoC or Platform Architecture, Interconnect Fabric, and Memory sub-system. - Experienced in GPU/AI system application performance optimization or software hardware co-design. - Strong knowledge and proficiency in software development in C/C++, scripting languages such as Python. - Understand the implementation of GPU/AI virtualization technology, deep learning architecture, and distributed systems. Preferred Qualifications: - Understand LLM model architecture, familiar with training and inference requirements on accelerator/memory/network.

Job Information

【For Pay Transparency】Compensation Description (Annually)

The base salary range for this position in the selected city is $162000 - $316800 annually.

Compensation may vary outside of this range depending on a number of factors, including a candidate's qualifications, skills, competencies and experience, and location. Base pay is one part of the Total Package that is provided to compensate and recognize employees for their work, and this role may be eligible for additional discretionary bonuses/incentives, and restricted stock units.

Benefits may vary depending on the nature of employment and the country work location. Employees have day one access to medical, dental, and vision insurance, a 401(k) savings plan with company match, paid parental leave, short-term and long-term disability coverage, life insurance, wellbeing benefits, among others. Employees also receive 10 paid holidays per year, 10 paid sick days per year and 17 days of Paid Personal Time (prorated upon hire with increasing accruals by tenure).

The Company reserves the right to modify or change these benefits programs at any time, with or without notice.

For Los Angeles County (unincorporated) Candidates:

Qualified applicants with arrest or conviction records will be considered for employment in accordance with all federal, state, and local laws including the Los Angeles County Fair Chance Ordinance for Employers and the California Fair Chance Act. Our company believes that criminal history may have a direct, adverse and negative relationship on the following job duties, potentially resulting in the withdrawal of the conditional offer of employment:

1. Interacting and occasionally having unsupervised contact with internal/external clients and/or colleagues;

2. Appropriately handling and managing confidential information including proprietary and trade secret information and access to information technology systems; and

3. Exercising sound judgment.

Similar Jobs

More Jobs at ByteDance

More Information Technology Jobs

Find similar GPU Application Platform Engineer Graduate (Server Research and Development) - 2027 Start (PhD) jobs: