About the roleOur recording rigs are egocentric: a head-mounted stereo or mono camera, sometimes with body-worn sensors, plus a smaller multi-camera ego-exo volume that gives us triangulated ground truth. From the wearer's point of view the hands and forearms are almost always visible, the torso partly, the legs and feet rarely. Today our pipeline fits Meta's Momentum Human Rig (MHR) to the hands, the upper-body keypoints and the metric camera trajectory, and places the result in the world. The lower body is inferred, the feet are an ankle-plus-offset pushed onto an assumed flat floor, and contact is a check that removes frames rather than a property of the estimate.
You will build the model that replaces that. The evidence you can condition on: the metric camera trajectory, hand and wrist poses, upper-body keypoints, 2D detections of whatever leg is in frame, and, in the ego-exo volume, full-body ground truth. The output we need: a full-body pose per frame in one metric world frame, on MHR / SOMA-X, that stands on the floor it is standing on, does not slide or sink, kneels and lies down when the wearer does - floor yoga is a real case - and reports per-frame contact with calibrated confidence. Motion capture, IMU suits and the ego-exo volume are yours to commission when the benchmark needs them.
We think the answer lives where physics-based character control meets human motion generation: a learned full-body motion prior conditioned on partial evidence, and simulation-based tracking or a contact-aware solver that makes the result physically consistent. If you have a better answer, that is the job.
What you'll do- Full-body motion prior. Train generative motion models - diffusion, autoregressive or masked - on mocap-scale data and our own rig data, conditioned on the partial evidence an egocentric rig produces, so they produce the lower body and feet when nothing sees them.
- Physical consistency. Physics-based tracking or imitation of the kinematic estimate in simulation, or contact-aware optimisation on the body model, so the output has ground contact, no penetration, no foot skate and plausible balance - including floor work, where the standing assumption fails.
- Contact as an output. Per-frame foot and body contact and the ground plane, with calibrated confidence, delivered alongside the pose.
- The body model. MHR / SOMA-X native output; pose priors and constraints the production solver can consume.
- Measurement, with the audit team. Define the lower-body, feet and contact metrics and the cohort they run on (ego-exo, mocap, IMU suit); commission captures against named gaps.
- Hand-off. Prototype 1 runnable model on the cohort in a repo we own 1 the pod's production owners integrate. Publish what you learn; we are generous on publication.
What we're looking for- A research record in physics-based character animation, human motion generation, or 3D human pose and shape - a PhD or an equivalent body of work, with publications at SIGGRAPH / SIGGRAPH Asia, CVPR / ICCV / ECCV, NeurIPS / ICLR or CoRL.
- Generative motion models. You have trained diffusion, autoregressive or masked motion models on mocap-scale data and conditioned them on constraints - keyframes, end effectors, paths, text.
- Physics-based control. Motion imitation and tracking for simulated humanoids in the DeepMimic / AMP / MaskedMimic lineage; comfortable in Isaac Gym / Isaac Lab or MuJoCo and with reinforcement learning at scale.
- Body models and kinematics. SMPL-X, MHR or SOMA-X; retargeting; contact and collision modelling.
- Engineering to match. Clean PyTorch research code, and the will to run on noisy, real, large-scale data rather than clean benchmarks.
- Research taste. You can turn "full body from a head camera" into a sequence of measurable steps and make steady progress on it.
Strong plus- Egocentric or partial-observation full-body estimation - head- and hand-tracking to body.
- Human-scene interaction and contact: scene-aware motion synthesis, volumetric body models.
- Motion cleanup or denoising of estimated or corrupted motion.
- Reusable motion priors for control, generative controllers, or large-scale motion tokenisation.
- Having shipped a research model into a product or a game engine; familiarity with Momentum / MHR tooling.
Tech stack- Python / PyTorch for model design and training; reinforcement learning at scale.
- Isaac Lab / MuJoCo; ProtoMotions- or MimicKit-class frameworks for physics-based character control.
- SMPL-X / MHR / SOMA-X body models; AMASS-class motion capture.
- Our data: hundreds of hours of ego-exo multi-view ground truth, thousands of hours of egocentric recordings, mocap and IMU-suit sessions on request.
The exact stack matters less than depth in motion generation or physics-based control and the judgment to make progress on an open problem.
What success looks like- Lower-body and feet pass rate on the body cohort, and the physical-violation rate - penetration, foot skate, ground-plane - are measured from your first quarter and fall every quarter after.
- Floor work - kneeling, sitting, lying, yoga - no longer breaks the estimate.
- The production body solver runs your prior and your contact terms; delivered feet stand on the floor they are on.
- A result on our public benchmark, and a publication, within the first year.
Who this role is not for- Pose-estimation researchers with no generative-model or physics depth - the upper body is already estimated; the problem is what the camera cannot see.
- Graphics researchers who want clean motion capture and no noisy real data.
- Anyone who needs the problem fully specified before starting.