Backend Inference Framework Engineer Graduate (AML Inference) - 2027 Start

ByteDance

$128K — $256K *
Consumer Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's or Master's degree in Software Development, Computer Science, or related technical field.
  • Proficient in C/C++ programming and familiar with data structures and algorithms.
  • Knowledge of multi-threaded concurrency, including threading, synchronization, and basic performance tuning.
  • Experience with high-concurrency distributed services, including service latency optimization.
  • Strong problem analysis and abstraction abilities with good learning and execution skills.
  • Excellent communication, collaboration, and documentation skills.

Responsibilities

  • Design and implement model inference services for large AI models.
  • Optimize core modules of the inference framework, improving performance and stability.
  • Monitor industry trends and innovate technologies relevant to business scenarios.
  • Promote upgrading and standardization of technical systems within the team.
  • Drive complex technical issues resolution and project implementation.

Benefits

  • Medical, dental, and vision insurance from day one.
  • 401(k) savings plan with company match.
  • Paid parental leave and short-term/long-term disability coverage.
  • 10 paid holidays and 10 paid sick days per year.
  • 17 days of Paid Personal Time, increasing with tenure.
Full Job Description
Responsibilitie

Data AML is ByteDance's Machine Learning mid-platform, providing training and inference systems for recommendation/advertising for businesses such as Douyin, Jinri Toutiao, and Xigua Video. It provides powerful Machine Learning computing power for internal business units within the company and conducts research on some general and innovative algorithms for issues in these businesses. We are looking for talented individuals to join our team. As a graduate, you will get opportunities to pursue bold ideas, tackle complex challenges, and unlock limitless growth. Successful candidates must be able to commit to an onboarding date by the end of the year. Please state your availability and graduation date clearly in your resume. Candidates can apply to a maximum of two positions and will be considered for jobs in the order you apply. The application limit is applicable to our Company and its affiliates' jobs globally. Applications will be reviewed on a rolling basis - we encourage you to apply early. Responsibilities: - Responsible for the overall architecture design and implementation of model inference services, building a high-performance, highly available, and scalable enterprise-level inference system for large-parameter, high-complexity AI models, overcoming various architectural challenges in the implementation of complex model inference, and supporting the efficient launch of models across all business scenarios. - Responsible for the R&D and optimization of the core modules of the inference framework, covering core capabilities such as inference engine scheduling, monitoring and alerting, canary release, etc., continuously iterating on the framework performance, and resolving performance bottlenecks, resource bottlenecks, and stability issues in high-concurrency and large-model inference scenarios. - Keep track of the latest inference technologies in the industry, conduct technology selection and innovation in combination with business scenarios, accumulate distributed high-concurrency service architecture solutions, and promote the upgrade and standardization of the team's technical system.

Qualification

Minimum Qualifications: - Individuals who are completing or have recently completed a Bachelor's/ Master's degree in Software Development, Computer Science, Computer Engineering, or a related technical discipline, or a related discipline. - Familiar with basic Linux commands, with solid C/C++ programming skills and knowledge of data structures and algorithms; - Familiar with the basic principles of multi-threaded concurrency, proficient in basic usages such as thread usage, synchronization locks, and thread pools, able to identify common concurrency issues, and possess the ability to perform basic performance tuning in multi-threaded scenarios; - Have experience in R&D projects of high-concurrency distributed services, and be familiar with service latency and resource optimization; - Possess good learning and execution abilities, be willing to proactively understand the model inference service architecture, and have problem analysis and abstraction capabilities; - Possess good cross-team collaboration skills, communication and presentation skills, and document writing skills, have strong sense of responsibility and stress tolerance, and be able to drive the resolution of complex technical issues and the implementation of projects; Preferred Qualifications: - Have practical project experience and understanding of source code in high-concurrency services/frameworks such as Redis, RocksDB, BRPC, GRPC, etc.; - Understand the operating mechanism of GPUs, and have relevant project experience and optimization capabilities in GPU service resource management and control;

Job Information

【For Pay Transparency】Compensation Description (Annually)

The base salary range for this position in the selected city is $128000 - $256000 annually.

Compensation may vary outside of this range depending on a number of factors, including a candidate's qualifications, skills, competencies and experience, and location. Base pay is one part of the Total Package that is provided to compensate and recognize employees for their work, and this role may be eligible for additional discretionary bonuses/incentives, and restricted stock units.

Benefits may vary depending on the nature of employment and the country work location. Employees have day one access to medical, dental, and vision insurance, a 401(k) savings plan with company match, paid parental leave, short-term and long-term disability coverage, life insurance, wellbeing benefits, among others. Employees also receive 10 paid holidays per year, 10 paid sick days per year and 17 days of Paid Personal Time (prorated upon hire with increasing accruals by tenure).

The Company reserves the right to modify or change these benefits programs at any time, with or without notice.

For Los Angeles County (unincorporated) Candidates:

Qualified applicants with arrest or conviction records will be considered for employment in accordance with all federal, state, and local laws including the Los Angeles County Fair Chance Ordinance for Employers and the California Fair Chance Act. Our company believes that criminal history may have a direct, adverse and negative relationship on the following job duties, potentially resulting in the withdrawal of the conditional offer of employment:

1. Interacting and occasionally having unsupervised contact with internal/external clients and/or colleagues;

2. Appropriately handling and managing confidential information including proprietary and trade secret information and access to information technology systems; and

3. Exercising sound judgment.

Similar Jobs

More Jobs at ByteDance

More Consumer Technology Jobs

Find similar Backend Inference Framework Engineer Graduate (AML Inference) - 2027 Start jobs: