Machine Learning Engineer Graduate (AML-Engine-Orchestration) - 2027 Start

ByteDance

$128K — $256K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's or Master's degree in Computer Science, Software Engineering, AI, or related field.
  • Proficient in at least one of Go, C++, or Python; strong grasp of data structures and algorithms.
  • Familiar with Linux; understanding of operating systems, networks, and distributed systems.
  • Hands-on skills for investigating systems through code, metrics, and profiling.
  • Systematic problem-solving ability; skilled in defining measurements and validating improvements.
  • Demonstrated ownership and collaboration through projects or internships.

Responsibilities

  • Design and build orchestration capabilities for ML platforms using Kubernetes and container runtimes.
  • Develop resource and quota systems for multi-tenant environments, enhancing GPU utilization and FinOps processes.
  • Create lifecycle orchestration for online model serving, covering deployment, upgrades, and disaster recovery.
  • Implement serving orchestration and traffic management for distributed serving clusters.
  • Enhance system reliability, serving latency, and availability through efficient orchestration.

Benefits

  • Medical, dental, and vision insurance from day one.
  • 401(k) plan with company match.
  • Paid parental leave and short-term/long-term disability coverage.
  • Life insurance and wellbeing benefits.
  • 10 paid holidays and 10 paid sick days annually, plus 17 days of Paid Personal Time.
Full Job Description
Responsibilitie

The Data-AML-Engine Orchestration team builds large-scale machine learning infrastructure that powers online model serving across ByteDance products, including TikTok. We develop the orchestration, scheduling, and resource management systems that connect heterogeneous compute infrastructure with production ML workloads. You will work on systems that directly affect GPU utilization, serving latency and availability, infrastructure reliability, and MLE productivity. Depending on your background and interests, you may focus on one or more of the following areas. We are looking for talented individuals to join our team. As a graduate, you will get opportunities to pursue bold ideas, tackle complex challenges, and unlock limitless growth. Successful candidates must be able to commit to an onboarding date by the end of the year. Please state your availability and graduation date clearly in your resume. Candidates can apply to a maximum of two positions and will be considered for jobs in the order you apply. The application limit is applicable to our Company and its affiliates' jobs globally. Applications will be reviewed on a rolling basis - we encourage you to apply early. Responsibilities: 1. Design and build foundational orchestration capabilities for machine learning platforms, including Kubernetes Operators, container runtimes, and lifecycle management for jobs, services, and stateful workloads. 2. Build multi-tenant resource and quota systems that support priorities, preemption, fair sharing, elasticity, and cross-cluster scheduling. Improve GPU utilization and cost efficiency through resource pooling and FinOps. 3. Build lifecycle orchestration for online model serving, including model and image distribution, deployment, upgrades, rollback, autoscaling, multi-cluster operation, and disaster recovery. 4. Build serving orchestration and traffic management capabilities for disaggregated serving clusters, including topology-aware scheduling, KV Cache affinity, intelligent request routing, and QoS/SLA management.

Qualification

Minimum Qualifications: - Individuals who are completing or have recently completed a Bachelor's or Master's degree in Computer Science, Software Engineering, Artificial Intelligence, or a related technical field. - Proficiency in at least one of Go, C++, or Python, with a solid foundation in data structures, algorithms, and software engineering principles. - Familiarity with Linux and a foundational understanding of operating systems, computer networks, concurrent programming, and distributed systems. - Strong hands-on and exploratory abilities, with a willingness to investigate systems through source code, metrics, logs, profiling, and experiments. - A systematic and quantitative approach to problem solving, with the ability to define measurements, test hypotheses, and validate system improvements. - Demonstrated ownership and collaboration through coursework, research, internships, open-source contributions, or other engineering projects. Preferred Qualifications: - Experience with Kubernetes, container runtimes, resource scheduling, quota management, multi-tenant systems, or FinOps. - Contributions to open-source infrastructure projects such as Kubernetes, Volcano, Koordinator, or OpenKruise. - Experience with model serving systems such as vLLM, SGLang, Triton, KServe, or Ray Serve, or an understanding of KV Cache, Continuous Batching, Prefill/Decode disaggregation, or model parallelism. - Experience with online services, gateways, traffic management, autoscaling, performance optimization, or highly available distributed systems. - Experience with GPU/NPU programming, heterogeneous resource scheduling, model distribution, or inference performance analysis.

Job Information

【For Pay Transparency】Compensation Description (Annually)

The base salary range for this position in the selected city is $128000 - $256000 annually.

Compensation may vary outside of this range depending on a number of factors, including a candidate's qualifications, skills, competencies and experience, and location. Base pay is one part of the Total Package that is provided to compensate and recognize employees for their work, and this role may be eligible for additional discretionary bonuses/incentives, and restricted stock units.

Benefits may vary depending on the nature of employment and the country work location. Employees have day one access to medical, dental, and vision insurance, a 401(k) savings plan with company match, paid parental leave, short-term and long-term disability coverage, life insurance, wellbeing benefits, among others. Employees also receive 10 paid holidays per year, 10 paid sick days per year and 17 days of Paid Personal Time (prorated upon hire with increasing accruals by tenure).

The Company reserves the right to modify or change these benefits programs at any time, with or without notice.

For Los Angeles County (unincorporated) Candidates:

Qualified applicants with arrest or conviction records will be considered for employment in accordance with all federal, state, and local laws including the Los Angeles County Fair Chance Ordinance for Employers and the California Fair Chance Act. Our company believes that criminal history may have a direct, adverse and negative relationship on the following job duties, potentially resulting in the withdrawal of the conditional offer of employment:

1. Interacting and occasionally having unsupervised contact with internal/external clients and/or colleagues;

2. Appropriately handling and managing confidential information including proprietary and trade secret information and access to information technology systems; and

3. Exercising sound judgment.

Similar Jobs

More Jobs at ByteDance

More Information Technology Jobs

Find similar Machine Learning Engineer Graduate (AML-Engine-Orchestration) - 2027 Start jobs: