Cloud Site Reliability Engineer - DCS Cloud

ByteDance

$153K — $300K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in Computer Science, Software Engineering, Information Security, or related field
  • 2+ years in Linux operations, Site Reliability Engineering (SRE), or DevOps
  • Proficient in programming languages like Go, Python, or C++
  • Strong understanding of Linux OS principles, networks, storage, GPUs, and databases
  • Familiarity with reliability practices like monitoring, capacity management, and incident response
  • Excellent communication and collaboration skills

Responsibilities

  • Design and operate global infrastructure across public and private clouds
  • Develop automation tools and monitoring systems for global infrastructure optimization
  • Standardize and manage cloud AMIs/images for compliance across environments
  • Engage in technical operations and on-call support for cloud and system incidents
  • Drive improvements throughout the entire infrastructure lifecycle from design to user support

Benefits

  • Medical, dental, and vision insurance from day one
  • 401(k) plan with company match
  • Paid parental leave
  • Short-term and long-term disability coverage
  • Life insurance and wellbeing benefits
  • 10 paid holidays and 10 paid sick days per year
  • 17 days of Paid Personal Time with increasing accrual by tenure
Full Job Description
Responsibilitie

Our Infrastructure Engineering team supports the company's fast growth by building and operating hyper-scale datacenters, managing the life cycle of server fleet, providing cloud solutions, and developing various infrastructure services and making sure they are scalable and are reliable. We have three subgroups for this role: - Cloud Host Delivery, Delivery & Standardization - Cloud Host Operation, Operation Efficiency & Reliability - Cloud Management & Security Responsibilities - What You'll Do - Design, build, scale, and operate ByteDance's global infrastructure, including large-scale systems spanning public and private clouds. - Develop tools, automation frameworks, visualizations, and monitoring systems to streamline operations and drive optimization of global infrastructure. - Create, manage, and standardize cloud AMIs/images for use across multiple environments, ensuring strict alignment with the company's global compliance standards. - Thrive in a fast-paced environment, engaging in technical operations and on-call rotations to address incidents related to cloud, OS, network, performance, and reliability. - Drive improvements across the entire infrastructure lifecycle, from ideation and design through development, deployment, user support, and continuous refinement.

Qualification

Minimal Qualifications - Bachelor's degree or above in Computer Science, Software Engineering, Information Security, or a related field. - 2+ years of experience in Linux operations, SRE, or DevOps; - Proficient in at least one programming language such as Go, Python, or C++, with solid engineering capabilities in platform development, system tooling, and automation. - Strong computer science fundamentals, with deep understanding of Linux OS principles, computer networks, storage systems, GPU systems, and databases, along with systematic troubleshooting and root-cause analysis skills. - Familiar with core reliability practices, including monitoring and alerting, capacity management, change management, canary/gray releases, incident response, and postmortem processes. - Strong communication and collaboration skills, with the ability to proactively identify problems, drive cross-team execution, and demonstrate strong ownership and results-oriented mindset. Preferred Qualifications - Hands-on experience operating public cloud platforms, or deep familiarity with major cloud providers such as OCI, AWS, Azure, GCP, etc, including understanding of their underlying mechanisms. - Experience with large-scale cloud host delivery, image/AMI systems, resource scheduling, network adaptation, and virtualization technologies such as KVM/QEMU. - Familiar with containers and cloud-native ecosystems, including Docker, Kubernetes, and containerd, with a solid understanding of isolation mechanisms like cgroups and namespaces. - Experience maintaining GPU clusters, including drivers, CUDA, MIG, topology awareness, troubleshooting, stress testing, and GPU delivery pipelines. - Proven experience in reliability-focused initiatives such as failure drill systems, capacity governance, change governance, observability platforms, and resource cost optimization. - Open-source contributions, technical blogs, patents, or technical sharing experience are highly preferred. - Experience operating large-scale production environments is a strong plus.

Job Information

【For Pay Transparency】Compensation Description (Annually)

The base salary range for this position in the selected city is $153900 - $300960 annually.

Compensation may vary outside of this range depending on a number of factors, including a candidate's qualifications, skills, competencies and experience, and location. Base pay is one part of the Total Package that is provided to compensate and recognize employees for their work, and this role may be eligible for additional discretionary bonuses/incentives, and restricted stock units.

Benefits may vary depending on the nature of employment and the country work location. Employees have day one access to medical, dental, and vision insurance, a 401(k) savings plan with company match, paid parental leave, short-term and long-term disability coverage, life insurance, wellbeing benefits, among others. Employees also receive 10 paid holidays per year, 10 paid sick days per year and 17 days of Paid Personal Time (prorated upon hire with increasing accruals by tenure).

The Company reserves the right to modify or change these benefits programs at any time, with or without notice.

For Los Angeles County (unincorporated) Candidates:

Qualified applicants with arrest or conviction records will be considered for employment in accordance with all federal, state, and local laws including the Los Angeles County Fair Chance Ordinance for Employers and the California Fair Chance Act. Our company believes that criminal history may have a direct, adverse and negative relationship on the following job duties, potentially resulting in the withdrawal of the conditional offer of employment:

1. Interacting and occasionally having unsupervised contact with internal/external clients and/or colleagues;

2. Appropriately handling and managing confidential information including proprietary and trade secret information and access to information technology systems; and

3. Exercising sound judgment.

Similar Jobs

More Jobs at ByteDance

More Information Technology Jobs

Find similar Cloud Site Reliability Engineer - DCS Cloud jobs: