跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 6 天前AI 评分46

自进化搜索智能体的共谋作弊问题:诊断与 CrossFit 缓解方法

AI 导读

自进化搜索智能体存在"共谋作弊"失效模式:生成问题的 proposer 与作答的 solver 在共享错误上日益一致,内部奖励上升但外部正确性停滞。为此提出 CrossFit,将 proposer 的源文档分为 A、B 两组,交叉用对方数据训练的辅助 solver 打分决定奖励。

正文

Authors:Meijia Chen, Hao Li, Zheng Lu, Hongshan Lin, Junbai Tian, Yichen Liu, Zijun Tian, Yufan Zou, Shuhan Sun, Hanxin Chen, Zeyu Zhang, Weizhi Du, Yueting Li, Tianyu Shi, Alaa Khamis

View PDF HTML (experimental)

Abstract:Self-evolving search agents build their own training curricula by jointly optimizing a proposer that generates questions and a solver that answers them. This closed loop introduces a failure mode we call co-cheating: the proposer and solver increasingly agree on shared errors, so internal reward improves without a matching gain in external correctness. A post-hoc audit against source evidence shows co-cheating growing more severe over successive rounds of self-evolution, with pseudo-label correctness stagnating or declining even as the in-loop training signal improves. The most direct mitigation is to verify proposals before training: we introduce multi-sample verification (MSV), which queries the same model three times with the source and three times without it to decide task admission and replace unreliable pseudo-labels. MSV partially reduces false agreement but leaves substantial residual co-cheating and costs six extra labeler generations per candidate. These limitations motivate CrossFit, our main method: it partitions the proposer's source documents into groups A and B; questions generated from A are scored by an auxiliary solver trained only on B, and vice versa. The cross-fitted agreement determines proposer reward, so a same-source pseudo-label cannot be reproduced through the feedback solver, while the original solver's update rule is unchanged. Rerunning the loop with Qwen3.5-4B and Qwen3.5-9B, MSV reduces false-agreement mass from 6.1% to 5.7% and from 8.8% to 7.2%, whereas CrossFit reduces it to 3.0% and 3.7%. Replaying identical proposals with source-excluded feedback further reduces false agreement to 0.4% and 0.1%, isolating feedback ancestry from curriculum changes. Across seven downstream search benchmarks, CrossFit improves average performance over standard coupled self-evolution by 8.8 and 8.4 points and over Search-R1 by 8.7 and 7.8 points at 4B and 9B.
Comments: 21 pages. Equal contribution: Meijia Chen, Hao Li, Zheng Lu
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as: arXiv:2609.39102 [cs.CL]
  (or arXiv:2609.39102v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2609.39102

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Meijia Chen [view email]
[v1] Wed, 30 Sep 2026 06:41:51 UTC (1,330 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org