跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 7 天前AI 评分38

CrossBFM:跨人形机器人形态蒸馏共享潜在行为空间

AI 导读

CrossBFM 将行为基础模型的潜在空间视为可跨形态迁移的资产,用无机器人专属参数的统一编码器同时蒸馏多个训练形态,耗时不到 1 GPU 小时,再经 PPO 训练 10 小时即可实现全身控制。在三个蒸馏人形机器人上,运动跟踪、姿态到达与 41 项奖励提示三种模式均可迁移,编码器仅用四分之一运动语料回归只损失 5% 跟踪性能,在未见机器人上可恢复至已见机器人 89% 的跟踪性能。

正文

View PDF HTML (experimental)

Abstract:Behavior Foundation Models (BFMs) give humanoids a promptable policy over a latent behavior space, enabling one single vector to represent a motion to imitate, a pose to reach, or a reward to maximize. Forward-Backward representations successfully produce such spaces, but at the cost of hundreds of GPU-hours for a single robot. Moreover, when the training process is repeated for a second robot, it produces a second space unrelated to the first, resulting in embodiment-specific latents that do not unify or transfer. We address these problems with CrossBFM, treating the latent space as the transferable asset for various embodiments. As retargeting provides frame-level cross-embodiment correspondence, we propose a unified encoder architecture with no robot-specific parameters for distilling the behavior space to address all training embodiments simultaneously in less than a GPU-hour. Following this encoder, latent-conditioned trackers turn the distilled latent into whole-body control in a conventional PPO training manner in just 10 more GPU-hours. On three distilled humanoids, all three prompting modes transfer: motion tracking with latent-conditioned policy losing only $0.025$ rad to its joint-conditioned counterpart, smooth goal reaching between poses with no falls, and reward optimization for all $41$ reward prompts. Our experiments further reveal that 1) regressing the encoder on a quarter of the motion corpus costs only $5\%$ of tracking performance and 2) training the encoder on a subset of robots and evaluating on an unseen one recovers up to $89\%$ of the tracking performance of seen robots, demonstrating cross-embodiment generalization to morphologically similar robots. We also verify the pipeline on real robots across all three prompting modes and with flow-based generated latents. Project website: this https URL
Comments: Project Website: this https URL
Subjects: Robotics (cs.RO)
Cite as: arXiv:2609.38087 [cs.RO]
  (or arXiv:2609.38087v1 [cs.RO] for this version)
  https://doi.org/10.48550/arXiv.2609.38087

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Tuan Dat Phuong [view email]
[v1] Tue, 29 Sep 2026 17:40:11 UTC (12,471 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org