ZfG Operations Engineer

Zoom Video Communications, Inc.

$98K — $228K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Proficiency in at least one programming or scripting language (e.g., Python, Go, Bash) for automation and infrastructure tasks
  • Understanding of distributed systems architecture including scalability and fault tolerance
  • Experience with monitoring tools and observability platforms for production environments
  • Knowledge of incident management protocols and root cause analysis
  • Strong collaborative skills to work with cross-functional teams on technical challenges
  • Ability to manage large-scale projects independently with accountability for results
  • Practical experience in site reliability engineering, DevOps, or equivalent roles

Responsibilities

  • Design and implement automation frameworks and monitoring solutions to enhance reliability
  • Analyze system performance metrics and recommend scalability improvements
  • Lead incident response efforts and conduct root cause analyses
  • Develop and maintain operational documentation and runbooks
  • Mentor team members on troubleshooting, optimization, and best practices

Benefits

  • Comprehensive benefits program supporting physical, mental, emotional, and financial health
  • Promotes work-life balance and community engagement
  • Work structure combines hybrid, remote, and in-office options
Full Job Description

What You Can Expect

You will develop and maintain reliable infrastructure solutions that power distributed systems serving millions of users globally. You will implement automation frameworks, monitoring tools, and performance optimization techniques while collaborating with cross-functional teams to resolve complex technical challenges. You will drive operational excellence and mentor team members, delivering measurable improvements in system reliability, efficiency, and incident response capabilities.


About the Team

Our team ensures Zoom's infrastructure operates reliably at scale. We collaborate across engineering, product, and operations to solve complex distributed systems challenges. We exist to protect customer experience through proactive monitoring, automation, and continuous improvement.


Responsibilities

  • Designing and implementing automation frameworks and monitoring solutions that enhance system reliability, reduce manual interventions, and optimize performance across distributed production environments
  • Analyzing system performance metrics to identify bottlenecks, recommend scalability improvements, and implement proactive solutions that prevent service degradation
  • Leading incident response efforts by coordinating with cross-functional teams, conducting root cause analysis, and implementing preventive measures that reduce future outage risk
  • Developing and maintaining operational documentation, runbooks, and service level objectives that standardize procedures and improve team efficiency
  • Mentoring team members on troubleshooting methodologies, system optimization techniques, and infrastructure best practices while facilitating knowledge sharing across teams

What We're Looking For

  • Demonstrate proficiency in at least one programming or scripting language (Python, Go, Bash, or similar) for building automation tools and infrastructure solutions
  • Apply knowledge of distributed systems architecture, including scalability patterns, fault tolerance mechanisms, and performance optimization principles
  • Utilize monitoring tools, observability platforms, and metrics collection systems to maintain visibility into production environments
  • Execute incident management protocols, including response coordination, root cause analysis, and remediation planning
  • Collaborate effectively with cross-functional teams to solve complex technical problems and drive infrastructure improvements
  • Operate independently on large-scale projects with minimal supervision while maintaining accountability for deliverables and timelines
  • Possess equivalent practical experience in site reliability engineering, DevOps, or infrastructure operations roles
  • Contribute to on-call rotations and demonstrate experience supporting production systems in high-availability environments

Salary Range or On Target Earnings:

Minimum:

$98,900.00

Maximum:

$228,700.00

In addition to the base salary and/or OTE listed Zoom has a Total Direct Compensation philosophy that takes into consideration; base salary, bonus and equity value.

Note: Starting pay will be based on a number of factors and commensurate with qualifications & experience.

We also have a location based compensation structure;  there may be a different range for candidates in this and other locations

At Zoom, we offer a window of at least 5 days for you to apply because we believe in giving you every opportunity. Below is the potential closing date, just in case you want to mark it on your calendar. We look forward to receiving your application!

Anticipated Position Close Date:

08/11/26

Ways of Working
Our structured hybrid approach is centered around our offices and remote work environments. The work style of each role, Hybrid, Remote, or In-Person is indicated in the job description/posting.

Benefits
As part of our award-winning workplace culture and commitment to delivering happiness, our benefits program offers a variety of perks, benefits, and options to help employees maintain their physical, mental, emotional, and financial health; support work-life balance; and contribute to their community in meaningful ways. Click for more information.

Similar Jobs

More Jobs at Zoom Video Communications, Inc.

More Information Technology Jobs

Find similar ZfG Operations Engineer jobs: