Job DescriptionNBCUniversal Operations & Technology is looking for a Staff SRE, Playout Engineering to provide technical leadership to a team of Site Reliability Engineers. This team drives reliability, observability, and operational excellence for cloud-based master control playout systems supporting all NBCUniversal live linear channels - NBC, Telemundo, Peacock Virtual Channels, etc.
In this position, you will shape the reliability strategy for live linear playout-defining service levels, improving resiliency, and strengthening monitoring and incident response-so the platform can meet evolving business needs with predictable performance and availability.
This role requires the ability to operate in a fast-paced environment. For systems in production, you will lead an on-call team and drive L1 and L2 troubleshooting, incident management and continuous improvement to maintain reliable distribution.
Responsibilities- Team lead for SRE engineers on playout Engineering team
- Define and manage reliability targets (SLIs/SLOs) and operational readiness criteria for playout services
- Drive incident response: establish on-call practices, lead major incident management, and ensure post-incident reviews result in measurable improvements
- Partner with engineering, product, and operations teams to improve reliability through capacity planning, performance tuning, and resilience testing
- Provide high-level conceptual drawings and operational runbooks to support architecture reviews, support readiness, and project planning
- Leadership in driving automation to reduce toil and improve reliability and mean time to recovery (MTTR)
- L1 & L2 support to maintain playout infrastructure/services for NBCUniversal including providing after hours on-call support select week
- Leadership in creating monitoring dashboards (Grafana) and proper alerts(teams/slack/ServiceNow)
QualificationsQualifications/Requirements- Bachelor's degree in computer science or related degree / experience
- Eight years' hands-on-keyboard Engineering experience working with broadcast automation playout environments e.g. Snell, Harris, Imagine, Amagi
- Requires on-call 24/7 availability for escalations
- Hands-on-keyboard experience administrating Linux environments
- Experience with monitoring/logging tools e.g. Splunk and Grafana
- Experience with streaming protocols and codecs (e.g. TS, HEVC, H.264, HLS, CMAF, SCTE-35, SCTE-224, ESAM, SRT/RIST)
- Experience with IP networking and interfacing with cloud-based networks
- Experience with containerization (Docker & Kubernetes)
- Excellent communicator and able to clearly articulate complex issues and technologies
- Expert with broadcast playout systems (master control) technologies
- Expert with public cloud environments using AWS services
- Comfortable working in a fast-paced agile environment. Requirements change quickly and our team needs to constantly adapt to meet objectives
- An automate-first and automate everything attitude
Desired Characteristics- Experience with cloud native playout vendor solutions (Amagi, Evertz, GrassValley, Harmonic, Imagine, CoralBay, Veset, etc.)
- Experience building strong operational readiness practices (runbooks, alert tuning, on-call health, incident reviews)
- Ability to create user interface designs based on client workflows
Additional Requirements: - Hybrid: This position currently has a hybrid schedule, which requires contributing from the office a minimum of four days per week. The Company reserves the right to change in-office requirements at any time.