Senior Site Reliability Engineer - Applied Machine Learning (Multiple Positions)

TikTok

$206K — $341K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Master's or equivalent degree in Computer Science, Engineering, Information Technology, or a related field, plus 3 years of relevant experience; OR Bachelor's degree with 5 years of experience.
  • 3 years experience in provisioning servers using Linux operating systems.
  • 3 years managing system resources and deployments with configuration management tools.
  • 3 years experience in building and designing tools with programming and scripting languages.
  • 3 years processing and analyzing large data using analytical tools and generating reports, dashboards, and alerts.
  • 3 years deploying and managing applications with containerized tools.

Responsibilities

  • Design and architect large-scale distributed ML and recommendation systems for high-volume traffic.
  • Make architectural decisions for ML platform components, balancing performance, cost, and maintainability.
  • Implement sophisticated MLOps automation, including multi-stage CI/CD pipelines for model deployment.
  • Conduct root cause analysis for distributed system failures and create fault-tolerant, self-healing components.
  • Provide user support, incident response, and conduct postmortems.
  • Participate in technical operations to address performance and reliability issues.
  • Mentor junior SREs and interns.

Benefits

  • Day one access to medical, dental, and vision insurance.
  • 401(k) savings plan with company match.
  • Paid parental leave and short-term/long-term disability coverage.
  • Life insurance and wellbeing benefits.
  • 10 paid holidays and 10 paid sick days per year.
  • 17 days of Paid Personal Time with increasing accruals by tenure.
Full Job Description
Responsibilities Responsibilities Design and architect large-scale distributed machine learning and recommendation systems processing high-volume traffic, including LLM inference infrastructure. Make architectural decisions for ML platform components, considering trade-offs in performance, cost, and maintainability for both traditional ML and LLM workloads. Design and implement sophisticated MLOps automation including multi-stage CI/CD pipelines for ML and LLM model deployment. Conduct root cause analysis for complex distributed system failures and architect fault-tolerant, self-healing components. Provide user support, incident responses and postmortems. Participate in technical operations and rotations in response to performance and reliability issues. Mentor junior SREs and interns. Qualifications Qualifications Must have a Master's degree or foreign equivalent degree in Computer Science, Engineering (any), Information Technology, or a related field, and 3 years of related work experience; OR a Bachelor's degree or foreign equivalent degree in Computer Science, Engineering (any), Information Technology, or a related field, and 5 years of post-bachelor's, progressive related work experience. Of the required experience, must have 3 years of experience in each of the following: Provisioning servers using Linux operating systems; Managing system resources and deployments using configuration management tools; Building and designing tools using programming languages and scripting languages; Processing and analyzing large quantities of logs and data using analytical tools and converting those data into reports, dashboards, and alerts using search processing languages; and Deploying and managing applications using containerized tools. Employer: TikTok USDS Joint Venture LLC Type: Full time, 40 hours/week Location: Bellevue, WA Salary Range: $206086 - $341734 per year To Apply, click the apply button below. Contact [redacted] if you have difficulty submitting resume through the website. Job Information [For Pay Transparency]Compensation Description (Annually) The base salary range for this position in the selected city is $206086 - $341734 annually. Compensation may vary outside of this range depending on a number of factors, including a candidate's qualifications, skills, competencies and experience, and location. Base pay is one part of the Total Package that is provided to compensate and recognize employees for their work, and this role may be eligible for additional discretionary bonuses/incentives, and restricted stock units. Benefits may vary depending on the nature of employment and the country work location. Employees have day one access to medical, dental, and vision insurance, a 401(k) savings plan with company match, paid parental leave, short-term and long-term disability coverage, life insurance, wellbeing benefits, among others. Employees also receive 10 paid holidays per year, 10 paid sick days per year and 17 days of Paid Personal Time (prorated upon hire with increasing accruals by tenure). The Company reserves the right to modify or change these benefits programs at any time, with or without notice. For Los Angeles County (unincorporated) Candidates: Qualified applicants with arrest or conviction records will be considered for employment in accordance with all federal, state, and local laws including the Los Angeles County Fair Chance Ordinance for Employers and the California Fair Chance Act. Our company believes that criminal history may have a direct, adverse and negative relationship on the following job duties, potentially resulting in the withdrawal of the conditional offer of employment: 1. Interacting and occasionally having unsupervised contact with internal/external clients and/or colleagues; 2. Appropriately handling and managing confidential information including proprietary and trade secret information and access to information technology systems; and 3. Exercising sound judgment.

Similar Jobs

More Jobs at TikTok

More Information Technology Jobs

Find similar Senior Site Reliability Engineer - Applied Machine Learning (Multiple Positions) jobs: