HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 6 天前AI 评分49
Agent Error Dataset:5 万条错误—诊断对,用于失败分析与错误感知后训练
AI 导读
研究者发布 Agent Error Dataset(AED),包含来自 33 个环境、19 个 harness 家族和 23 个策略模型的 50,228 条错误—诊断对,覆盖 9,961 个源任务。
正文
Abstract:An unsuccessful LLM agent rollout contains more information than its final reward: the observations available to the agent, the actions it chose, and the environment's responses. Reusing this experience for learning requires identifying a decision to revise and testing a concrete alternative. We introduce the Agent Error Dataset (AED), comprising 50,228 error-diagnosis pairs from 9,961 source tasks across 33 environments, 19 harness families, and 23 policy models in text-based agent systems. We retain source traces and execution metadata to support cross-setting failure analysis and re-diagnosis without repeating the original rollout. Our five-stage Agentic Error-to-Training (AET) pipeline collects natural failures, generates diagnoses and proposed corrections, and checks them against recorded evidence. Where replay is supported, we compare corrections with original-action retries from the same checkpoint under matched execution settings. We then construct separate training views for diagnosis and actor recovery. Across 3,062 matched replay pairs, first-proposal corrections raise verifier pass rates from 18.4% to 51.1%, a gain of 32.7 percentage points. Using a separately frozen diagnosis release, full-diagnosis fine-tuning on 1,656 source tasks raises Qwen3-8B's exact-step agreement with internal teacher labels from 47.2% to 63.6%, averaged over three seeds on a 943-case holdout. The strongest prompted reference in this comparison scores 54.7%, and mean agreement improves at each of four increasing training-set sizes. In a single-seed comparison of actor-training recipes, action-only repair training scores 6.67 percentage points higher on WebShop-lite than success-only training.
| Subjects: | Artificial Intelligence (cs.AI); Computation and Language (cs.CL) |
| Cite as: | arXiv:2609.40111 [cs.AI] |
| (or arXiv:2609.40111v2 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2609.40111 arXiv-issued DOI via DataCite |
Submission history
From: Kunlun Zhu [view email]
[v1]
Wed, 30 Sep 2026 16:40:22 UTC (1,423 KB)
[v2]
Thu, 1 Oct 2026 04:54:01 UTC (1,423 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org