跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 8 天前AI 评分34

APPL:用智能体先验引导策略学习,打通技能组合与泛化

AI 导读

Agent Priors-guided Policy Learning(APPL)将每个策略的结构先验同时用作训练约束与组合接口:构建智能体把完整演示切分为可复用技能,为每个技能提出多个结构先验并逐一训练、验证策略,运行时智能体再按先验选择并组合这些策略。

正文

View PDF HTML (experimental)

Abstract:Robots that learn from a few demonstrations often require two forms of generalization. Compositional generalization recombines skills to solve new tasks, and skill generalization lets the learned policy behind each skill work in new situations. The two depend on each other, yet information is lost between composition and the skills it calls. Where a skill works is determined by the structure its policy is trained with, while composition sees the skill only through a separate description, such as a name, an instruction, or a symbolic operator, that omits this structure. Our key idea is to use each policy's structural prior as part of the interface between composition and the skill. A structural prior states what a behavior depends on, for example that a grasp depends only on the gripper's pose relative to the object. Built into training, it shapes where the policy generalizes; stated in language, it tells composition where the policy applies. We instantiate this idea in Agent Priors-guided Policy Learning (APPL). A construction agent segments complete demonstrations into reusable skills, proposes several structural priors for each skill, and trains and verifies one policy per prior. A runtime agent then selects among these prior-specific policies and composes them toward new task goals using their interfaces. Across MetaWorld and long-horizon ManiSkill tasks, APPL improves out-of-distribution skill generalization and enables previously unseen skill compositions; ablating the interface information substantially reduces performance. These results support the use of training-time structural assumptions as a bridge between skill learning and skill composition.
Subjects: Robotics (cs.RO)
Cite as: arXiv:2609.35690 [cs.RO]
  (or arXiv:2609.35690v1 [cs.RO] for this version)
  https://doi.org/10.48550/arXiv.2609.35690

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Puming Jiang [view email]
[v1] Mon, 28 Sep 2026 17:37:00 UTC (1,427 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org