HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 7 天前AI 评分39
EVOKE:在智能体中激发世界知识以实现可迁移决策
AI 导读
针对 LLM 智能体在多步决策中难以迁移到未见环境的问题,研究者提出后训练方法 EVOKE,通过在固定状态下引入目标多样性,迫使同一候选动作在不同目标下重新排序,从而激发预训练中已内化的世界知识。在三种骨干模型的多类任务上,EVOKE 提升了任务表现、未见环境泛化能力与数据效率。
正文
Authors:Yuhan Guo, Jinming Liu, Liang Xu, Ziqiang Li, Jianguo Huang, Zhicheng Wang, Hu Zhu, Qiuyu Chen, Yuntao Wei, Xin Jin, Wenjun Zeng
Abstract:Large language models (LLMs) are increasingly deployed as agents for multi-step decision-making, yet transfer poorly to unseen environments. World-model methods address this by training agents to predict future observations, at the cost of additional training and errors that compound when predictions are used for planning. However, for LLM agents operating in digital environments, much of this world knowledge is already internalized during pretraining, which shifts the problem from acquiring it to eliciting it. We argue that typical post-training provides little pressure for such elicitation, since supervision under a single goal at each visited state inadvertently drives policies to rely on superficial contextual habits. We introduce EVOKE, a post-training method that supplies this pressure through goal diversity at fixed states. Motivated by theory showing that an agent competent across diverse goals must encode a world model recoverable from its action preferences, EVOKE holds the environment state and interaction history fixed and ranks the same candidate actions under alternative goals, forcing action preferences to change, so that a policy relying on contextual habits or single-goal correlations cannot order them correctly. This implicitly elicits the policy's pretrained world knowledge to inform decisions. We evaluate EVOKE across diverse tasks in three backbones, demonstrating improved task performance, unseen environment generalization, and data efficiency. We further conduct controlled analyses to better understand what drives these gains. These findings offer a new perspective on eliciting internalized world knowledge for transferable action through direct decision supervision.
| Comments: | 19 pages. Project page: this https URL ; Code: this https URL ; Models: this https URL |
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2609.38334 [cs.CL] |
| (or arXiv:2609.38334v2 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2609.38334 arXiv-issued DOI via DataCite |
Submission history
From: Yuhan Guo [view email]
[v1]
Tue, 29 Sep 2026 18:01:45 UTC (1,436 KB)
[v2]
Thu, 1 Oct 2026 12:36:04 UTC (1,436 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org