HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 6 天前AI 评分40
Box^2-Bench:语言模型能否选择性依赖外部指导?
AI 导读
研究者提出 Box^2-Bench,在固定模型与任务的前提下改变工作流可靠性,以衡量语言模型能否在受益于可靠指导的同时覆盖不可靠指导,即"跳出框框"的能力。测试显示前沿模型虽能从可靠指导中获益,但在指导具有误导性或变得不可靠时仍很脆弱。团队用坏工作流训练两个开放权重模型,发现反事实监督微调可提升鲁棒性,而基于结果的强化学习能提高对有用工作流的使用,该行为还可迁移到同行纠错和受损记忆的鲁棒性上。
正文
Abstract:Agent harnesses often improve language models with human-designed workflows, but as models grow more capable, unreliable guidance can increasingly constrain their execution. We call the ability to benefit from useful guidance while overriding unreliable guidance thinking outside the box. We introduce Box$^2$-Bench, which holds the model and task fixed while varying workflow reliability to isolate how models regulate their reliance on guidance. On Box$^2$-Bench, frontier models often benefit from reliable guidance but remain vulnerable when it is misleading or becomes unreliable. To test whether this capability can be learned, we train two open-weight models using bad workflows, reserving good workflows for evaluation. We explore two complementary training strategies: counterfactual supervised fine-tuning improves robustness, while outcome-based reinforcement learning can shift the balance toward greater use of helpful workflows. We further find that this behavior extends beyond workflows to other forms of external information, improving peer correction and robustness to corrupted memory. Together, our results identify selective reliance on fallible external information as a dimension of agent reliability not captured by task performance alone.
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2609.39578 [cs.CL] |
| (or arXiv:2609.39578v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2609.39578 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Boyuan Wang [view email]
[v1]
Wed, 30 Sep 2026 12:10:24 UTC (16,005 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org