Job Posting Title:
Software Engineer II
Req ID:
Job Description:
Product Engineering is a unified team responsible for the engineering of Disney Entertainment & ESPN digital and streaming products and platforms. This includes product engineering, media engineering, quality assurance, engineering behind personalization, commerce, lifecycle, and identity.
The Observability & Insights group ensures that Disney Streaming's distributed systems are reliable, performant, and transparent. We build ML-powered detection systems, telemetry pipelines, intelligent alerting, and developer experience tooling that enable engineers across the organization to understand system health and take action quickly.
Job Summary: As a Software Engineer II, you will contribute to building and operating machine learning models and AI-driven systems that enhance the reliability of Disney's streaming ecosystem. You will work on production ML models - including autoencoders for anomaly detection, statistical threshold systems, and LLM-powered investigation gates - that transform telemetry and signals into automated detection and proactive insights across Disney+, Hulu, and ESPN.
You will participate in the ML lifecycle: feature engineering on time-series data, model training on GPU clusters, real-time inference pipelines, and model improvement. You will partner with engineering and platform teams to embed intelligence into operational workflows, improving system resilience and customer experience at scale.
As a Software Engineer II, you will deliver features end-to-end, participate in model design and code reviews, and grow into owning components of production ML systems within a fast-paced, AI-native engineering environment.
Responsibilities and Duties of the Role: - Contribute to production ML models for anomaly detection, including autoencoders, statistical threshold models, and ensemble detection systems that monitor thousands of microservices
- Develop and improve ML training pipelines using PyTorch on GPU clusters - including feature engineering, model training, threshold calibration, and deployment through MLflow
- Engineer features from time-series telemetry (error ratios, latency, infrastructure metrics) - implementing windowing, normalization, and data quality safeguards for model consumption
- Support real-time ML inference systems that run prediction cycles in production - including model serving, detection logic, and alert generation
- Build AI-driven capabilities using foundation models (Claude, GPT-4) for automated investigation and reasoning over system health signals
- Create and deploy scalable APIs and services (FastAPI) that deliver ML predictions, health status, and insights to engineering teams and operational tooling
- Partner cross-functionally to embed ML intelligence into workflows such as incident response and release validation
- Contribute to model improvement through evaluation, retraining, and threshold tuning
Basic Qualifications - 3+ years of professional software engineering experience building, scaling, and maintaining ML-powered backends, data-driven microservices, and production RESTful APIs using FastAPI or Flask
- Strong hands-on experience in end-to-end ML engineering using PyTorch or TensorFlow spanning model architecture selection, feature engineering, training, and evaluation (e.g., autoencoders, sequential/time-series models like RNNs/GRUs, anomaly detection, or transformers).
- Proficiency in Python and at least one ML framework (PyTorch preferred)
- Experience managing model experiments, lineage, and hyperparameter tracking using tools like MLflow or Weights & Biases.
- Practical experience processing, transforming, and querying large-scale telemetry, event, or time-series datasets using PySpark, Pandas, or Databricks.
- Experience with modern development practices including version control (GitHub), containerization (Docker), and cloud-native deployments (AWS/EKS)
- Strong analytical and troubleshooting skills with the ability to iterate rapidly
- Strong collaboration and communication skills, with the ability to work cross-functionally
Preferred Qualifications - Experience with foundation model integration, prompt engineering/evaluation, RAG architectures, or orchestration frameworks like LangChain or LangGraph
- Familiarity with observability platforms (e.g., Datadog, Grafana, Conviva) and high-volume telemetry data
- Experience with large-scale data platforms (Databricks, Spark, Snowflake)
Required Education Bachelor's degree in Computer Science, Machine Learning, Statistics, Engineering, or equivalent experience
#disneytech
The hiring range for this position in Glendale, CA is $117,500 - $157,500 per year, and in New York City, NY is $123,000 - $165,000 per year. The base pay actually offered will take into account internal equity and also may vary depending on the candidate's geographic region, job-related knowledge, skills, and experience among other factors. A bonus and/or long-term incentive units may be provided as part of the compensation package, in addition to the full range of medical, financial, and/or other benefits, dependent on the level and position offered.
Job Posting Segment:
PE - Sports, News & Entertainment, Tech Enablement
Job Posting Primary Business:
PE - Sports, News & Entertainment, Enablement - News & Entertainment Engineering
Primary Job Posting Category:
Software Engineer
Employment Type:
Full time
Primary City, State, Region, Postal Code:
Glendale, CA, USA
Alternate City, State, Region, Postal Code:
Date Posted:
2026-08-24