Unified Morphological Representation
Standardizes state-action spaces across robotic hands with different degrees of freedom and physical structure.
Overview video for UniBYD and UniManip cross-embodiment manipulation results.
Overview. UniBYD leverages human demonstrations while learning manipulation strategies tailored to diverse robotic hand morphologies.
In embodied intelligence, the embodiment gap between robotic and human hands brings significant challenges for learning from human demonstrations. Existing reinforcement-learning approaches often remain confined to reproducing human manipulation and struggle to support diverse robotic hand configurations.
We propose UniBYD, a unified framework that discovers manipulation policies aligned with each robot's physical characteristics. UniBYD incorporates a unified morphological representation (UMR), dynamic PPO with an annealed reward schedule, and a hybrid Markov-based shadow engine that anchors early-stage training within the expert manifold before transitioning to autonomous exploration.
To evaluate UniBYD, we introduce UniManip, a benchmark for cross-embodiment manipulation spanning diverse robotic morphologies. Experiments demonstrate a 44.08% average improvement in success rate over the current state-of-the-art.
Framework of UniBYD. UniBYD first encodes diverse robotic hands through UMR, then uses dynamic PPO with reward annealing to move from demonstration-guided imitation toward morphology-aligned exploration.
Standardizes state-action spaces across robotic hands with different degrees of freedom and physical structure.
Anneals imitation and goal rewards to bridge offline-informed learning and online morphology-adaptive exploration.
Reduces early training drift by providing fine-grained expert guidance before the learned policy acts autonomously.
UniManip benchmark. Task distribution across hand configurations and manipulation scenarios.
Table 1: Comparative results on UniManip. SR = success rate, PE = position error, OE = orientation error, AS = alignment score. Bold blue indicates the best result. "-" indicates that the method does not support the corresponding hand type.
| Hand Type | Metric | Retarget | ManipTrans | DexMachina* | UniBYD |
|---|---|---|---|---|---|
| 2 fingers 1 hand | SR up (%) | 12.27 | - | - | 78.13 |
| PE down (cm) | 2.81 | - | - | 0.53 | |
| OE down (deg) | 28.77 | - | - | 18.74 | |
| AS up | 4.24 | - | - | 8.93 | |
| 3 fingers 1 hand | SR up (%) | 4.36 | - | - | 71.81 |
| PE down (cm) | 2.65 | - | - | 0.89 | |
| OE down (deg) | 29.19 | - | - | 12.95 | |
| AS up | 2.48 | - | - | 9.13 | |
| 5 fingers 1 hand | SR up (%) | 5.68 | 26.44 | - | 85.67 |
| PE down (cm) | 2.89 | 1.93 | - | 0.77 | |
| OE down (deg) | 29.13 | 22.28 | - | 10.90 | |
| AS up | 3.07 | 5.88 | - | 9.29 | |
| 5 fingers 2 hands | SR up (%) | 2.10 | 28.75 | 25.33 | 57.67 |
| PE down (cm) | 2.96 | 1.44 | 1.66 | 0.72 | |
| OE down (deg) | 29.56 | 16.84 | 14.13 | 13.44 | |
| AS up | 2.58 | 4.34 | 4.81 | 8.16 |
Qualitative comparison. UniBYD learns strategies aligned with robot embodiment, while imitation-centric or goal-only baselines fail in the illustrated case.
Training behavior. Reward and success-rate evolution on a representative task.
Table 2: Ablation study of UniBYD. SE = shadow engine, GR = goal reward, LSC = loss synergy with counterbalancing. Each component contributes to improved success, precision, and embodiment alignment.
| Hand Type | Metric | Base | +SE | +GR | +GR+LSC | UniBYD |
|---|---|---|---|---|---|---|
| 2 fingers 1 hand | SR up (%) | 24.94 | 67.06 | 56.19 | 61.50 | 78.13 |
| PE down (cm) | 2.66 | 0.54 | 1.28 | 1.32 | 0.53 | |
| OE down (deg) | 27.51 | 19.89 | 21.31 | 21.58 | 18.74 | |
| AS up | 5.63 | 7.56 | 7.99 | 8.36 | 8.93 | |
| 3 fingers 1 hand | SR up (%) | 19.56 | 51.06 | 65.13 | 67.63 | 71.81 |
| PE down (cm) | 2.45 | 0.95 | 0.93 | 0.87 | 0.89 | |
| OE down (deg) | 26.09 | 14.67 | 14.06 | 12.85 | 12.95 | |
| AS up | 5.42 | 6.87 | 8.52 | 8.78 | 9.13 | |
| 5 fingers 1 hand | SR up (%) | 13.56 | 55.11 | 63.17 | 65.39 | 85.67 |
| PE down (cm) | 2.46 | 1.24 | 0.87 | 0.79 | 0.77 | |
| OE down (deg) | 27.12 | 17.29 | 13.22 | 13.23 | 10.90 | |
| AS up | 4.87 | 6.33 | 8.29 | 8.64 | 9.29 | |
| 5 fingers 2 hands | SR up (%) | 33.46 | 43.33 | 47.01 | 54.17 | 57.67 |
| PE down (cm) | 1.46 | 1.43 | 1.48 | 1.36 | 0.72 | |
| OE down (deg) | 15.35 | 17.64 | 17.17 | 14.17 | 13.44 | |
| AS up | 3.20 | 3.55 | 6.27 | 6.73 | 8.16 |
Table 3: Component ablation on a representative task. Removing reward annealing causes the largest drop, while boundary and entropy losses improve precision and stability.
| Method | SR up (%) | OE down (deg) | PE down (m) | AS up |
|---|---|---|---|---|
| w/o Boundary Loss | 81.57 | 7.50 | 0.17 | 8.97 |
| w/o Entropy Loss | 84.68 | 5.67 | 0.18 | 8.34 |
| w/o Reward Annealing | 75.44 | 5.80 | 0.20 | 7.65 |
| UniBYD | 93.00 | 2.92 | 0.10 | 9.12 |
Removing the morphology descriptor drops SR from 93.00% to 72.00% and requires about 97.9M steps to reach 80% SR instead of 80.7M steps.
UniBYD achieves 26/50, 32/50, and 35/50 successful trials on 2-fingered, 3-fingered, and 5-fingered real-world platforms.
Base vs. UniBYD. UniBYD discovers more embodiment-aligned manipulation policies.

Policy evolution. Manipulation strategy changes over training epochs.

Real-world platforms. UniBYD transfers to physical robotic hands.
Franka gripper simulation.
Franka gripper simulation.
Franka gripper simulation.
Three-finger hand simulation.
Three-finger hand simulation.
Three-finger hand simulation.
Three-finger hand simulation.
Inspire dexterous hand simulation.
Inspire dexterous hand simulation.
Inspire dexterous hand simulation.
Inspire dexterous hand simulation.
@article{yuan2025unibyd,
title={UniBYD: A Unified Framework for Learning Robotic Manipulation Across Embodiments Beyond Imitation of Human Demonstrations},
author={Yuan, Tingyu and Guan, Biaoliang and Ye, Wen and Tian, Ziyan and Yang, Yi and Zhou, Weijie and Li, Zhaowen and Huang, Yan and Wang, Peng and Zhao, Chaoyang and others},
journal={arXiv preprint arXiv:2512.11609},
year={2025}
}