DreamHand

Repurposing Video Diffusion Models for Occlusion-Robust Egocentric 3D Hand Motion Recovery

Yufei Liu1,4, Xixi Wang2, Hao Li3,4, Ganlong Zhao3,4, Kaitong Cai4, Chengkai Jin2,4, Chunxiao Liu4, Jianbo Liu4, Siyuan Huang4, Xingang Pan2, Hongsheng Li3,4 ✉

1Shanghai Jiao Tong University 2Nanyang Technological University 3The Chinese University of Hong Kong 4ACE Robotics

✉ Corresponding author

Shanghai Jiao Tong University logo Nanyang Technological University logo The Chinese University of Hong Kong logo ACE Robotics logo
Paper coming soon Code coming soon
30%MPJPE-p cut · ARCTIC
40%MPJPE-p cut · HOT3D
46–61%Gain out of sight
5Egocentric benchmarks
0External detectors

Abstract

Egocentric video offers scalable manipulation data for embodied AI, yet recovering metric 3D hand trajectories remains challenging due to severe object occlusion and frequent out-of-sight gaps. Existing single-frame and windowed temporal regressors fail when hand shortly leaves the frame, while recent video diffusion models (VDMs) rely on heavy, stochastic multi-step sampling as pixel-space renderers. We instead repurpose VDM into a deterministic geometry encoder. A single forward pass over the clean latent exposes scene content beyond current observations, including occluded and out-of-sight hands. We introduce DreamHand, an offline clip-level framework that extracts features via a Deterministic Clean-Latent Encoder and decodes them with a Bidirectional Spatiotemporal Decoder. DreamHand recovers continuous bimanual trajectories with metric placement and no external detector, while a Ray-Based Camera Solver supports a second configuration that needs no test-time camera intrinsics. Across five egocentric benchmarks, DreamHand sets a new state of the art, cutting MPJPE-p by 30% on occlusion-heavy ARCTIC and 40% on HOT3D. These gains reach 46%–61% once out-of-sight hands are included in the evaluation, offering a scalable path from everyday human video to robot manipulation data.

Method

DreamHand pipeline from the VAE and Wan DiT through the Ray Head and Bidirectional Decoder to 3D hand motion
Overview of DreamHand. Raw video is encoded into latents by the Wan VAE and processed by the Wan DiT (σ=0). Block-15 features (grayed blocks skipped) branch to the Ray Head and Bidirectional Spatiotemporal Decoder. Camera intrinsics supervise the Ray Head during training. At test time they are discarded (K-free) or used only as the bearing source in the translation solve (standard).

Qualitative Results

We present side-by-side qualitative comparisons across ARCTIC, HOT3D and OAKINK2 under a shared protocol, and on unscripted in-the-wild egocentric footage where no ground truth exists. DreamHand remains stable under strong interaction and partial occlusion, which makes the representation suitable for downstream retargeting and manipulation analysis.

ARCTIC HOT3D OAKINK2 IN THE WILD
GT
Ours (DreamHand)
ViDiHand
EgoForce
Dyn-HaMR
HaWoR
WiLoR
HaMeR

Retarget to Dexterous Hand

DreamHand learns a dual mapping between human and robot hands. This makes occlusion-robust retargeting practical for teleoperation, imitation learning, and dexterous deployment.

Results

Result analysis. DreamHand stays ahead most clearly on occlusion-heavy and out-of-sight cases, suggesting that the clean-latent representation preserves hand state even when visual evidence is missing.