EgoX treats human, simulated, and heterogeneous robot demonstrations as unified egocentric trajectories, pretrains one shared manipulation representation, and specializes it to each target embodiment.
Abstract
Scaling robot manipulation requires learning from demonstrations collected across humans and heterogeneous robots. These sources cover diverse objects, scenes, tasks, and interaction strategies, but gaps in sensing, kinematics, motion patterns, and action spaces make it difficult to train a shared policy effectively. We introduce EgoX, an egocentric cross-embodiment framework for robot manipulation. Its central design is to share manipulation structure at the level of semantically corresponding modalities, while preserving embodiment-specific sensing and control interfaces. EgoX stores human and robot demonstrations as unified egocentric observation-action trajectories, pretrains a Modularized Cross-Embodiment Transformer with Embodiment Dreaming (MXT+) on the mixture, and fine-tunes it on target-embodiment demonstrations. MXT+ combines embodiment-specific modality tokenizers and action experts with a shared Transformer trunk. An auxiliary embodiment dreaming objective predicts future embodiment latents from egocentric visual and proprioceptive features, using exponential-moving-average copies of the modality encoders; this training-only branch encourages the trunk to encode embodiment-specific temporal structure without adding deployment-time computation. Across 13 task–embodiment pairs spanning five tasks and six robot embodiments, cross-embodiment pretraining consistently improves normalized task progress over target-embodiment training from scratch, with particularly strong gains under out-of-distribution evaluation. Broader pretraining mixtures further strengthen transfer, while embodiment dreaming improves contact-rich tool-use performance and reduces mean cross-embodiment latent-geometry distance by 38.52%.
Method
EgoX has three stages: (i) collect human and robot demonstrations as unified egocentric observation-action trajectories organized by corresponding modality types, (ii) pretrain MXT+ on mixed cross-embodiment data using action prediction and embodiment dreaming, and (iii) finetune the shared representation on target-embodiment tasks. Embodiment-specific modality tokenizers and action experts preserve heterogeneous sensing and control interfaces around a shared Transformer trunk. Dreaming experts predict future egocentric visual and proprioceptive latents during training and are discarded at deployment.
MXT+ Overview. Embodiment-specific modality tokenizers map egocentric vision and proprioception into a shared Transformer trunk. Modular action experts predict embodiment-specific action chunks, while dreaming experts predict future embodiment latents from EMA targets during training only. Bottom: the EgoX embodiments, each with its unified egocentric frame and end-effector-link keypoints.
Embodiments & Tasks
EgoX is evaluated on five manipulation tasks across six robot embodiments (five real robots and one simulated humanoid). Egocentric human demonstrations provide an additional pretraining embodiment.
Embodiment
Task
Sort
G1-Dex5MXT+ · Human pretraining
Sort
G1-Dex5MXT+ · Human + robot pretraining
Sort
G1-Dex5MXT+ · Cross-embodiment pretraining
Insert
G1-Dex5
Sort
G1-Dex5
Scoop
G1-Dex5
Scoop
xArm7-Dex4
Sort
R1Lite-Gripper
Kitchen-breadID
G1-Dex5
Kitchen-breadOOD
G1-Dex5
SortOOD
xArm7-Dex4MXT+ · Cross-embodiment pretraining
Kitchen-bread
R1Lite-Gripper
Kitchen-breadID
R1Lite-GripperMXT+ · Target-only
Kitchen-breadOOD
R1Lite-GripperMXT+ · Target-only
Kitchen-breadID
R1Lite-GripperMXT+ · Cross-embodiment pretraining
Kitchen-breadOOD
R1Lite-GripperMXT+ · Cross-embodiment pretraining
Kitchen-breadID
R1Lite-GripperHIT
Kitchen-breadOOD
R1Lite-GripperHIT
Kitchen-bread
Human-Dex5 ProHuman demonstration
Kitchen-bread
Human-Dex5 ProHuman demonstration
Scoop
LocoMan-Dex2
Kitchen-bread
G1-Sim-Dex3
Kitchen-cleanup
G1-Sim-Dex3
Scoop
Z1-Tool
Insert
xArm7-Dex4
No videos match these filters.Select All in one of the filter groups to explore more footage.
Results
We study three questions: (1) does egocentric cross-embodiment pretraining improve downstream manipulation over target-only training, (2) how do the diversity and composition of pretraining embodiments affect performance, and (3) does embodiment dreaming improve cross-embodiment transfer? Evaluations cover ID and OOD settings using the normalized task progress score.
Main Results. Cross-embodiment-pretrained MXT+ generally improves task progress score across 13 task–embodiment pairs, with substantially larger gains under OOD evaluation.
Pretraining Data Composition. Human-only pretraining improves over no pretraining on several task–embodiment pairs, and adding robot data helps further in some settings. Broader embodiment mixtures tend to improve performance, with full cross-embodiment pretraining showing the strongest overall trend, particularly under OOD evaluation.
Embodiment Dreaming
Embodiment dreaming asks the shared trunk to predict future embodiment latents built from egocentric image features, robot joint configurations, and end-effector joint configurations. EMA copies of MXT+'s own modality encoders provide stable targets. The objective encourages embodiment-aware dynamics modeling during pretraining and finetuning; dreaming heads and EMA teachers are discarded at deployment.
Embodiment Dreaming. The predictive objective improves task progress score across all seven reported task–embodiment settings while adding no deployment-time computation.