UMI-BridgeAction-Anchored Latent Alignment across Human and Robot Manipulation Data
1 Tsinghua University2 Simple AI
Abstract
Real-robot demonstrations are limited, motivating the use of human manipulation data collected without robots, including egocentric videos and handheld Universal Manipulation Interface (UMI) demonstrations. However, differences in viewpoint, embodiment, and available action supervision make it difficult to align representations across these sources according to manipulation motion rather than visual appearance. We introduce UMI-Bridge, which uses UMI as an intermediate domain to align representations according to action equivalence rather than pixel similarity. UMI action supervision anchors the latent representation to end-effector motion and gripper behavior, while synchronized head–wrist observations and paired ego–UMI clips support alignment across views and domains. We train a dual-view latent action model (LAM) on human manipulation data without robot demonstrations, then freeze its wrist teacher and dynamics model to regularize vision-language-action (VLA) post-training on UMI and robot data. The shared wrist interface enables this training-time supervision across both domains while preserving the policy's standard inference architecture. Across three real-robot tasks, UMI-Bridge achieves 91.7% mean success versus 73.3% for Naive Co-training with matched UMI and robot data. On two data-efficiency tasks, it surpasses a full-data Robot-only baseline using 25% of the robot demonstrations together with UMI data. It also achieves 85% and 90% success on two additional tasks learned from UMI demonstrations without task-specific robot demonstrations. These results support action-anchored latent alignment for data-efficient robot learning and UMI-to-robot task transfer.
Method
Learn the representation and guide the policy.
Two training stages connect human manipulation data to a standard robot policy.
Learn action-anchored latents.
UMI actions anchor latent representations to manipulation motion. Latent exchange links synchronized views and paired human–UMI videos.
Supervise policy learning.
A frozen wrist teacher and dynamics model supervise auxiliary latent predictions from the policy’s high-level vision-language model features.
Use the standard policy.
The policy predicts actions from observations, instructions, and proprioceptive state. Auxiliary training components are removed.
See the human–UMI correspondence
Two separately recorded demonstrations, aligned using dynamic time warping (DTW).
01 / Policy Performance
Action grounding matters.
All three methods use the same UMI data and the full robot demonstration set. UMI-Bridge achieves 91.7% mean success across Stain Wiping, Shirt Folding, and Produce Sorting.
02 / Robot Data Efficiency
Learning with fewer robot demonstrations.
With a fixed UMI dataset, UMI-Bridge uses 25% of the robot demonstrations to achieve 77.5% mean success across Stain Wiping and Produce Sorting. Robot-only training with the full robot dataset reaches 62.5%.
03 / UMI Task Transfer
Demonstrate with UMI and execute with a robot.
UMI-Bridge achieves 85% success on Cup Placement on Coaster and 90% on Snack Placement in Tray, without target-task robot demonstrations.
Selected execution with a paper cup absent from our training demonstrations. Demonstration and execution were recorded separately.
Reference
Citation
@misc{liu2026umibridge,
title = {{UMI-Bridge}: Action-Anchored Latent Alignment across Human and Robot Manipulation Data},
author = {Haiyi Liu and Jinming Ma and Ke Rui and Yuteng Wei and Yuan Ma and Yushen Zuo and Honglong Tian and Haoran Jia and Weitao Zhou and Jiawei Wang and Shiyi Chen and Haiyan Mao and Jiaqi Zhang and Chun Zhang and Minglei Li},
year = {2026},
eprint = {2609.18232},
archivePrefix = {arXiv},
primaryClass = {cs.RO},
url = {https://arxiv.org/abs/2609.18232}
}