Full Job Description
Job Description: - Responsible for the computational performance optimization of ByteDance's recommendation mid-platform models, conduct in-depth tuning for inference/training bottlenecks in business scenarios, and improve computing utilization. - Lead the design and development of high-performance kernel libraries, covering general-purpose and business-customized kernels, including general computation and communication parallelism, to ensure the ultimate performance of kernels. - Deeply cultivate model compilation optimization technology, and based on directions such as graph optimization, kernel fusion, computation scheduling, and code generation, build and improve the mid-platform model compilation system. - Collaborate with the business and algorithm teams to identify performance issues, provide full stack performance analysis, bottleneck diagnosis, and optimization solutions, consolidate general-purpose performance optimization components, toolchains, and platform capabilities, and empower multiple internal business.
Qualification
Minimum Qualifications: - Individuals who are completing or have recently completed a Bachelor's/ Master's degree in Software Development, Computer Science, Computer Engineering, or a related technical discipline, or a related discipline. - Familiar with mainstream model compilation stacks (such as TVM, MLIR, XLA, etc.), with relevant experience in development, and optimization; - Proficient in C/C++ development, familiar with assembly, CPU/GPU architecture, and cache mechanism, with practical experience in high-performance kernel development; - Familiar with the underlying principles of deep learning frameworks (such as TensorFlow, PyTorch, OneFlow, etc.), understand the computational graph structure, inference/training execution process, and have experience in implementing model graph optimization and compilation optimization; - Possess strong abilities in independent thinking, problem decomposition, performance troubleshooting, and practical optimization, and be able to independently overcome complex performance bottlenecks; Preferred Qualifications: - Have experience in joint hardware and software design, and possess experience in heterogeneous computing projects; - Have experience in contributing to or developing open-source deep learning kernel libraries, compilers, or inference engines; - Have in-depth research experience on the underlying architecture and mechanisms of at least one machine learning framework (TensorFlow / PyTorch / MxNet or other self-developed frameworks).
Job Information
【For Pay Transparency】Compensation Description (Annually)
The base salary range for this position in the selected city is $128000 - $256000 annually.
Compensation may vary outside of this range depending on a number of factors, including a candidate's qualifications, skills, competencies and experience, and location. Base pay is one part of the Total Package that is provided to compensate and recognize employees for their work, and this role may be eligible for additional discretionary bonuses/incentives, and restricted stock units.
Benefits may vary depending on the nature of employment and the country work location. Employees have day one access to medical, dental, and vision insurance, a 401(k) savings plan with company match, paid parental leave, short-term and long-term disability coverage, life insurance, wellbeing benefits, among others. Employees also receive 10 paid holidays per year, 10 paid sick days per year and 17 days of Paid Personal Time (prorated upon hire with increasing accruals by tenure).
The Company reserves the right to modify or change these benefits programs at any time, with or without notice.