Research Scientist, Foundation Model, Speech Understanding

ByteDance

$254K — $480K *
Consumer Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Master's or PhD in computer science, mathematics, engineering, or related field
  • 3+ years of experience in machine learning and deep learning
  • Proven work on Automatic Speech Recognition, Automatic Speech Translation, or related areas
  • Publications in top machine learning/speech venues
  • Strong coding skills in C/C++ and Python

Responsibilities

  • Conduct research and development in speech/audio foundation models
  • Collaborate with teams to identify key research areas
  • Integrate research findings into product applications
  • Work on team-driven projects to solve challenges
  • Enhance the effectiveness of the research team

Benefits

  • Day one access to medical, dental, and vision insurance
  • 401(k) savings plan with company match
  • Paid parental leave
  • Short-term and long-term disability coverage
  • Life insurance
  • Wellbeing benefits
  • 10 paid holidays and 10 paid sick days
  • 17 days of Paid Personal Time, increasing with tenure
Full Job Description
Responsibilities - Conduct research and development in speech/audio foundation models - Collaborate with cross-functional teams to identify key research areas and contribute to the development of innovative speech/audio models. - Work with product development teams to integrate research findings into practical applications for ByteDance and other platforms. - Collaborate on team-driven projects to address complex challenges and enhance the overall effectiveness of the research team. Qualification Minimum Qualifications - Master's or PhD in computer science, mathematics, engineering or related field - Have 3+ years of experience in one or more areas of machine learning and deep learning, including but not limited to: Automatic Speech Recognition, Automatic Speech Translation, Speech/audio self-supervised learning and foundation models, Speaker recognition and verification, Speech emotion recognition, Multimodal foundation models, Large Language Model pre-training and fine-tuning. Preferred Qualifications - Publications in accredited ML/DL venues such as NeurIPS, ICLR, ICML, AAAI and speech venues such as ICASSP, ASRU, Interspeech - Deep understanding of Large Language models - Familiar with distributed computing and large scale model training - Familiar with deep learning frameworks such as Tensorflow and Pytorch. - Familiar with engineering principles and best practices. - Highly competent in algorithms and programming; Strong coding skills in C/C++ and Python. - Ability to work collaboratively in a fast-paced, multi-functional environment Job Information 【For Pay Transparency】Compensation Description (Annually) The base salary range for this position in the selected city is $254400 - $480000 annually. Compensation may vary outside of this range depending on a number of factors, including a candidate's qualifications, skills, competencies and experience, and location. Base pay is one part of the Total Package that is provided to compensate and recognize employees for their work, and this role may be eligible for additional discretionary bonuses/incentives, and restricted stock units. Benefits may vary depending on the nature of employment and the country work location. Employees have day one access to medical, dental, and vision insurance, a 401(k) savings plan with company match, paid parental leave, short-term and long-term disability coverage, life insurance, wellbeing benefits, among others. Employees also receive 10 paid holidays per year, 10 paid sick days per year and 17 days of Paid Personal Time (prorated upon hire with increasing accruals by tenure). The Company reserves the right to modify or change these benefits programs at any time, with or without notice.

Similar Jobs

More Jobs at ByteDance

More Consumer Technology Jobs

Find similar Research Scientist, Foundation Model, Speech Understanding jobs: