NBCUniversal Media, LLC

Reinforcement Learning Engineer - Ingénieur(e) en apprentissage par renforcement

NBCUniversal Media, LLC$90K — $120K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Graduate degree in Robotics, Computer Science, AI, or related field focused on Reinforcement Learning.
  • Proven experience as an RL Engineer or Research Engineer in a fast-paced environment.
  • Experience in industries with complex multi-disciplinary teams (robotics, smart grids, etc.).
  • Fluency with Python, Git, and Unix shell.
  • Deep familiarity with RL frameworks (Ray Rllib, Stable Baselines3, or CleanRL).
  • Experience with physics engines (MuJoCo, Bullet) or 3D game engines.
  • Strong mathematical background related to Markov Decision Processes.

Responsibilities

  • Coordinate with ML engineers, Annotation teams, and TPMs to define simulation training requirements.
  • Design and maintain high-fidelity 2D/3D simulation environments using Unity, Unreal, or Isaac Sim.
  • Create and optimize reward functions that align agent behavior with product goals.
  • Develop and enhance RL algorithms (e.g., PPO, SAC) for complex 3D observation spaces.
  • Analyze the reality gap and implement techniques for domain adaptation.

Benefits

  • Access to cutting-edge technology and tools in the field of reinforcement learning.
  • Opportunity to work in a cross-functional team with experts from various disciplines.
  • Support for professional development and continuous learning.
  • Flexible working conditions to promote a healthy work-life balance.
Full Job Description
We are seeking a Reinforcement Learning Engineer with experience manipulating virtual environments to train autonomous agents. This role focuses on the design of robust simulation environments, reward structures, and policy architectures that can navigate complex, multi-sensor landscapes. Key Responsibilities - Cross-Functional Coordination: Work with partner ML and Annotation engineers and TPMs to spec out data, simulation, and training requirements. - Environment Design: Build and maintain high-fidelity 2D/3D simulation environments (using tools like Unity, Unreal, or Isaac Sim) that serve as the training ground for RL agents. - Reward Engineering: Design and tune complex reward functions that align agent behavior with product goals and safety constraints. - Algorithm Implementation: Develop and optimize RL algorithms (e.g., PPO, SAC, or Offline RL) capable of handling high-dimensional 3D observation spaces. - Sim-to-Real Strategy: Analyze the "reality gap" and implement domain randomization or adaptation techniques to ensure models perform reliably in real-world scenarios. Nous sommes à la recherche d'un(e) ingénieur(e) en apprentissage par renforcement ayant de l'expérience dans la création et l'exploitation d'environnements virtuels pour l'entraînement d'agents autonomes. Ce rôle consiste à concevoir des environnements de simulation robustes, des structures de récompense et des architectures de politiques capables d'évoluer dans des contextes complexes et multi-capteurs. Vous jouerez un rôle clé dans le rapprochement entre simulation et performance réelle en développant des systèmes RL évolutifs et en garantissant un comportement fiable des agents dans des conditions variées. - Collaboration interfonctionnelle : Travailler avec les ingénieurs ML, les équipes d'annotation et les TPM afin de définir les besoins en données, en simulation et en entraînement. - Conception d'environnements : Développer et maintenir des environnements de simulation 2D/3D à haute fidélité à l'aide d'outils tels que Unity, Unreal ou Isaac Sim. - Ingénierie des récompenses : Concevoir et optimiser des fonctions de récompense afin d'aligner le comportement des agents avec les objectifs produit et les contraintes de sécurité. - Implémentation d'algorithmes : Développer et optimiser des algorithmes d'apprentissage par renforcement (ex. : PPO, SAC, RL hors ligne) adaptés à des espaces d'observation à haute dimension. - Stratégie sim-to-real : Réduire l'écart entre simulation et réalité à l'aide de techniques comme la randomisation de domaine et l'adaptation afin d'assurer des performances fiables en conditions réelles. Qualifications - Education: Graduate degree (Master's or PhD) in Robotics, Computer Science, AI, or a related field with a focus on Reinforcement Learning, Imitation Learning, or other Online Machine Learning fields. - Professional Experience: Proven experience as an RL Engineer or Research Engineer in a fast-paced environment. - Industry Context: Prior experience in industries with complex multi-disciplinary teams such as robotics, smart grids, precision agriculture, game development, or aerospace. Technical Proficiency: - Core Tools: Fluency with Python, Git, and the Unix shell. - RL Frameworks: Deep familiarity with frameworks like Ray Rllib, Stable Baselines3, or CleanRL. - Physics & 3D Engines: Experience with physics engines (MuJoCo, Bullet) or 3D game engines. - Ecosystem: Familiarity with collaborative tools such as Jira/Confluence, Slack, a Git server, and an experiment tracking framework. Attributes: - Strong Mathematical Background: Essential for understanding Markov Decision Processes (MDPs) and gradient-based optimization. - High Attention to Detail: Critical for debugging non-deterministic agent behaviors and ensuring environment parity. - Formation : Maîtrise ou Doctorat en robotique, informatique, intelligence artificielle ou domaine connexe avec une spécialisation en apprentissage par renforcement, imitation ou apprentissage en ligne. - Expérience : Expérience démontrée en tant qu'ingénieur(e) en apprentissage par renforcement ou en recherche dans un environnement dynamique. - Contexte industriel : Une expérience dans des secteurs multidisciplinaires tels que la robotique, les réseaux intelligents, l'agriculture de précision, les jeux vidéo ou l'aérospatiale est fortement valorisée. Compétences techniques - Outils principaux : Excellente maîtrise de Python, Git et des environnements Unix. - Frameworks RL : Expérience avec des frameworks tels que Ray RLlib, Stable Baselines3 ou CleanRL. - Physique et simulation : Expérience avec des moteurs physiques (MuJoCo, Bullet) ou des environnements de simulation 3D. - Écosystème : Familiarité avec des outils collaboratifs tels que Jira, Confluence, Slack, les workflows Git et les plateformes de suivi d'expériences. Qualités recherchées - Solides bases mathématiques : Bonne compréhension des processus de décision de Markov (MDP) et de l'optimisation basée sur le gradient. - Rigueur et précision : Capacité à déboguer des systèmes non déterministes et à assurer la cohérence et la précision des environnements de simulation.

About NBCUniversal Media, LLC

NBCUniversal Media, LLC is a media and entertainment company that operates a variety of businesses, including television networks, film studios, and theme parks. The company was founded in 2004 and is headquartered in New York, New York. NBCUniversal's television networks include NBC, Telemundo, and USA Network, among others. The company's film studios produce and distribute movies under the Universal Pictures brand. NBCUniversal also operates theme parks in the United States and Japan. The company is committed to producing high-quality content and delivering it to audiences around the world.
Learn more about NBCUniversal Media, LLC
Size
35,000 employees
Industry
Founded
1994

Similar Jobs

More Jobs at NBCUniversal Media, LLC

More Information Technology Jobs

Find similar Reinforcement Learning Engineer - Ingénieur(e) en apprentissage par renforcement jobs: