Vision-language-action pre-training

UCAG-POne Policy, Many Embodiments: Unified Camera-Centric Action Geometry Pre-training for Heterogeneous Embodied Manipulation

Xiaomi Embodied Intelligence Team University of Macau | CAIR & SKL-IOTSC

Explore the work
11datasets
6,300+hours of demonstrations
9embodiments
1unified checkpoint

Abstract

UCAG-P overview of shared camera-centric action geometry
Camera-observable anchors provide a common geometric interface from demonstrations to execution.

Scaling generalist vision-language-action (VLA) policies is severely bottlenecked by the inherent heterogeneity of embodied data, which spans diverse robot morphologies, camera configurations, and low-level action spaces. Existing paradigms typically address this mismatch through explicit action retargeting, human-to-robot video synthesis, or dataset-specific adaptation branches, fundamentally hindering the joint learning of a unified policy. We introduce UCAG-P, a camera-centric unified action formulation that structurally aligns heterogeneous embodied datasets into a shared geometric action space. Rather than treating robot-specific commands as the shared policy target, UCAG-P represents manipulation through camera-observable anchor motion in image and camera-frame coordinates, treating robot arms, humanoids, and human hands as different embodiments of a common action schema. A geometry-conditioned action translator combines predicted motion with target-embodiment kinematics to produce executable controls. The resulting decoupled architecture allows a shared VLA policy to learn transferable manipulation geometry while retaining embodiment-specific controllability. UCAG-P is trained on 4.03K hours of robot and simulation data and 2.34K hours of human demonstrations. A single checkpoint reaches 98.3% on LIBERO, 88.7% and 89.2% on RoboTwin Easy and Hard, 82.0% zero-shot on LIBERO-Plus, and 62.0% on RoboCasa GR-1, without benchmark-specific fine-tuning.

Paper at a glance

A shared geometric interface,
from data to execution.

SYSTEM OVERVIEWOne policy, many embodiments
UCAG-P unified camera-centric action geometry framework across heterogeneous embodiments
Overview of UCAG-P for unified learning from heterogeneous manipulation data, spanning single-arm robots, bimanual robots, dexterous robotic hands, and human hands. UCAG-P maps their incompatible native action spaces into a shared camera-centric motion space defined by the 2D image and 3D camera-frame trajectories of wrist or end-effector and grasp center. A geometry-conditioned translator then combines the unified motion prediction with camera-to-base transforms, Jacobians, and robot state to produce embodiment-specific executable actions. This unified interface enables joint robot-human training and improves performance across LIBERO, LIBERO-Plus, RoboTwin, and RoboCasa GR-1.

Simulation benchmarks

Success rate (%) · one checkpoint · no benchmark-specific post-training

98.3%LIBEROFour task suites
82.0%LIBERO-PlusOOD robustness
88.7%RoboTwin EasyBimanual
89.2%RoboTwin HardBimanual
62.0%RoboCasa GR-129-DoF humanoid
Shared camera-centric action space across robot and human embodiments
Human and robot motion are represented through visual grounding and spatial information in the camera frame.

The shared space

One geometric schema,
multiple bodies.

Human demonstrations can directly supervise the shared policy because the action target is defined in the camera frame rather than in a robot-specific joint space.

  1. Visual groundingAnchor points identify the manipulation geometry visible in the image.
  2. Spatial informationCamera-frame coordinates preserve the 3D structure needed for translation.
  3. Embodiment translationRobot geometry turns the shared proposal into an executable command.

The idea

Put the shared target where
every embodiment can see it.

We introduce UCAG-P, a camera-centric unified action formulation for heterogeneous embodied manipulation. Instead of using robot-specific commands as the shared target, UCAG-P represents manipulation through camera-observable anchor motion in image and camera-frame coordinates.

This formulation treats robot arms, humanoids, and human hands as different embodiments of a common action schema. A geometry-conditioned action translator then converts the shared prediction into the control representation required by the target robot.

Camera-centric action anchors across heterogeneous embodiments
Camera-observable anchors provide a common geometric target across heterogeneous data.

The method

Share the geometry.
Translate the execution.

UCAG-P separates the representation that should be shared from the execution details that should remain embodiment-specific.

UCAG-P model architecture
Stage 01

Camera-centric specialization

Train the VLM backbone and shared motion head on all robot and human samples with available geometric supervision, establishing p0/p1 anchor motion as the common action interface.

Stage 02

Geometry-conditioned action translation

Train the translator from ground-truth camera-frame trajectories, robot state, calibration, Jacobians, and executable labels, while masking inactive command dimensions.

Stage 03

Joint robot-human training

Jointly optimize the shared motion head and translator on mixed robot-human data, using policy-predicted trajectories so training matches inference-time execution.

Human-to-robot transfer

2× speed

Different viewpoints.
One shared action geometry.

Human demonstrations supervise camera-observable wrist and grasp-center motion. The policy learns this shared geometric target, while the action translator combines it with robot state, kinematics, and camera geometry to produce executable commands.

Bread case

Bread pickup through shared anchor motion

01

Extract shared supervision

Hand landmarks define camera-observable wrist and grasp-center anchor motion.

02

Learn shared motion

The policy predicts a shared action sequence in camera-centric coordinates.

03

Translate to robot control

Robot state, kinematics, and camera geometry turn shared motion into executable commands.

Bowl case

Bowl stacking through shared anchor motion

01

Human demonstration

Track hand wrist and grasp-center anchors during bimanual bowl stacking.

02

Shared motion

Represent bowl manipulation as synchronized camera-centric anchor geometry.

03

Robot execution

Translate shared bowl motion into embodiment-specific robot commands.

Blackboard case

Erasing a blackboard through shared anchor motion

01

Human demonstration

Track left- and right-hand wrist and grasp-center anchors during blackboard interaction.

02

Shared motion

Represent the observed hand trajectories as synchronized camera-centric anchor geometry.

03

Robot execution

Translate the shared motion into embodiment-specific commands for the robot arm.

Qualitative rollouts

One checkpoint,
many embodiments.

Representative simulation and real-world executions across single-arm, bimanual, humanoid, and human-to-robot settings.

LIBEROSingle-arm manipulation
RoboTwin 2.0Bimanual manipulation
RoboCasa GR-1Humanoid control
Pen holderRobot execution
Drawer openingReal-world adaptation
Bowl stackingBimanual adaptation
Bread pickupHuman-to-robot transfer

Evidence

One checkpoint across
simulation and reality.

We evaluate single-arm, bimanual, humanoid, out-of-distribution, cross-embodiment, and real-robot tasks.

BENCHMARKSPerformance across embodiments
UCAG-P benchmark performance panels across heterogeneous embodiments

Real-world Piper tasks

Real-world Piper task success rates
Failure case examples
Representative failure cases from anchor detection, trajectory prediction, and action translation.

Limitations

What remains
to be solved.

Errors in camera calibration, depth estimation, embodiment kinematics, or hand-keypoint localization can propagate into the camera-centric target and downstream controller. Cross-embodiment transfer remains challenging, especially for contact-rich and articulated tasks.

Citation

Cite UCAG-P.

Public arXiv version: arXiv:2608.26058 (2026).

@article{xu2026ucag-p,
  title={One Policy, Many Embodiments: Unified Camera-Centric Action Geometry Pre-training for Heterogeneous Embodied Manipulation},
  author={Shaoqing Xu and Fang Li and Guozhi Zhan and Zhixiang Duan and Yuhan Wang and Yuechen Luo and Shengyin Jiang and Hanbing Li and Zhiying Du and Longlong Wang and Longmei Jiang and Weixiang Liang and Ying Gong and Yong Pan and Ziping Zhao and Zhiyuan Chen and Yangwei You and Kun Ma and Qinyuan Liu and Hangjun Ye and Zhi-xin Yang},
  journal={arXiv preprint arXiv:2608.26058},
  year={2026}
}