Site Reliability Engineer - System Service Global

ByteDance

$162K — $387K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in Electrical Engineering, Computer Engineering, Computer Science, or related field.
  • Experience in large-scale Linux host management including OS deployment and fleet operations.
  • Knowledge of core data center foundational services: DNS, NTP, DHCP, NAT, APT repository management, and Kerberos.
  • Proficiency in DevOps tools (e.g., Ansible, Salt, Puppet) and CI/CD pipelines.
  • Familiarity with SRE practices including SLO/SLI definition and blameless post-mortems.
  • Understanding of high availability designs and disaster recovery strategies.
  • Strong troubleshooting skills within Linux system and network layers.

Responsibilities

  • Manage large-scale host infrastructure in non-China data centers.
  • Ensure reliability and availability of core data center services like DNS and DHCP.
  • Design and implement high availability deployment architectures.
  • Develop and enforce SLOs; lead incident response and post-mortems.
  • Collaborate with network, security, and application teams for service evolution.
  • Identify and drive automation opportunities for operational efficiencies.

Benefits

  • Access to medical, dental, and vision insurance from day one.
  • 401(k) savings plan with company match.
  • Paid parental leave and short-term/long-term disability coverage.
  • Life insurance and wellbeing benefits provided.
  • 10 paid holidays and 10 paid sick days per year.
  • 17 days of Paid Personal Time, with increasing accrual by tenure.
Full Job Description
Responsibilitie

The Global System Service team owns the infrastructure services and management solutions that power ByteDance's data centers outside of China - from day-to-day operations to long-term architecture design and maintenance. The team specializes in composing end-to-end solutions by drawing on both open-source community tools and in-house developed products, tailored to both the business requirements and the operational complexities of large-scale infrastructure across ByteDance's non-China regions. Our mission is to deliver efficient infrastructure solutions and a stable, secure system environment for ByteDance's global business. We are looking for a self-motivated system engineer that is equipped with SRE mindset and DevOps skills. Your responsibilities will include: - Manage and maintain large-scale host infrastructure across ByteDance's non-China data centers, covering OS lifecycle management, configuration standardization, and fleet-wide health monitoring. - Own the reliability and availability of core data center foundational services, including DNS, NTP, DHCP, NAT, APT repository, and Kerberos authentication. - Design and implement deployment architectures for foundational services, ensuring high availability, fault tolerance, and disaster recovery across regions. - Develop and enforce SLOs for managed services; lead incident response, root cause analysis, and post-mortem reviews to drive continuous reliability improvements. - Collaborate with network, security, and application teams to ensure foundational services meet the evolving demands of global business growth. - Identify automation opportunities across host management and service operations; drive tooling and process improvements to reduce toil and increase operational efficiency.

Qualification

Minimum Qualifications: - Bachelor's degree or higher in Electrical Engineering, Computer Engineering, Computer Science or related majors. - Solid experience in large-scale Linux host management, including OS deployment, configuration management, patching, and fleet operations. - Strong hands-on knowledge of core data center foundational services: DNS (BIND/PowerDNS), NTP, DHCP, NAT, APT repository management, and Kerberos. - Proficiency with DevOps tooling, including configuration management tools (e.g., Ansible, Salt, Puppet) and CI/CD pipelines. - Familiarity with SRE principles and practices, including SLO/SLI definition, error budget management, and blameless post-mortems. - Solid understanding of high availability design patterns, active-active/active-passive architectures, and disaster recovery strategies. - Strong troubleshooting skills across the Linux system stack and network layer. Preferred Qualifications: - Experience managing host fleets at scale (thousands of nodes or above) in a production environment. - Scripting or development experience in Python, Go, or Bash for automation and tooling. - Exposure to hybrid or multi-region data center environments.

Job Information

【For Pay Transparency】Compensation Description (Annually)

The base salary range for this position in the selected city is $162000 - $387600 annually.

Compensation may vary outside of this range depending on a number of factors, including a candidate's qualifications, skills, competencies and experience, and location. Base pay is one part of the Total Package that is provided to compensate and recognize employees for their work, and this role may be eligible for additional discretionary bonuses/incentives, and restricted stock units.

Benefits may vary depending on the nature of employment and the country work location. Employees have day one access to medical, dental, and vision insurance, a 401(k) savings plan with company match, paid parental leave, short-term and long-term disability coverage, life insurance, wellbeing benefits, among others. Employees also receive 10 paid holidays per year, 10 paid sick days per year and 17 days of Paid Personal Time (prorated upon hire with increasing accruals by tenure).

The Company reserves the right to modify or change these benefits programs at any time, with or without notice.

For Los Angeles County (unincorporated) Candidates:

Qualified applicants with arrest or conviction records will be considered for employment in accordance with all federal, state, and local laws including the Los Angeles County Fair Chance Ordinance for Employers and the California Fair Chance Act. Our company believes that criminal history may have a direct, adverse and negative relationship on the following job duties, potentially resulting in the withdrawal of the conditional offer of employment:

1. Interacting and occasionally having unsupervised contact with internal/external clients and/or colleagues;

2. Appropriately handling and managing confidential information including proprietary and trade secret information and access to information technology systems; and

3. Exercising sound judgment.

Similar Jobs

More Jobs at ByteDance

More Information Technology Jobs

Find similar Site Reliability Engineer - System Service Global jobs: