HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 7 天前AI 评分40
AREX-2:用长时程反思任务推进 LLM 智能体自我改进
AI 导读
AREX-2 通过从机器学习和算法编程任务合成长期改进轨迹,训练出基于 Qwen3.8-27B 的智能体,在 MLE-bench Lite 上得分 81.8、Frontier-CS 上 70.7。
正文
Authors:Hongjin Qian, Chaofan Li, Kun Luo, Wenqing Wei, Jianlyu Chen, Shuqi Lu, Yuyang Hu, Hongwang Xiao, Hui Wang, Chaozhuo Li, Qiwei Ye, Zhicheng Dou, Defu Lian, Zheng Liu
Abstract:We present AREX-2, an effort to advance the self-improving capability of LLM agents, which we define as the ability to iteratively refine a solution at test time. This ability rests on two complementary capabilities: reflection, which produces a solution better than the current one, and long-horizon execution, which keeps the iteration effective over many rounds. We hypothesize that both capabilities are domain-agnostic, and can therefore be learned in scenarios that are well suited for supervision. Accordingly, we synthesize long-horizon improvement trajectories from machine learning and algorithmic programming tasks, two domains that offer verifiable feedback and reward sustained iteration. Trained on this data, our agent, built on Qwen3.8-27B, achieves strong results on MLE-bench Lite (81.8) and Frontier-CS (70.7), transfers to deep research with 84.0 on BrowseComp, 52.6 on HLE, 92.2 on GAIA, and 93.8 on DeepSearchQA, and keeps improving as its budget of rounds grows. These results show that long-horizon reflective data is an effective route toward self-improving agents.
| Comments: | Code will be released at this https URL and models at this https URL |
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2609.38288 [cs.AI] |
| (or arXiv:2609.38288v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2609.38288 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Hongjin Qian [view email]
[v1]
Tue, 29 Sep 2026 16:52:25 UTC (1,413 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org