跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 6 天前AI 评分39

AutoDataBench:面向数据智能的自动研究测试平台

AI 导读

AutoDataBench 是一个隔离数据因素、专门评估 LLM 数据智能的受控测试平台,覆盖数据诊断、组织与构建三类优化任务,并在工具使用、检索和知识注入场景下评测前沿模型迭代改进训练数据的能力。该平台还通过对比训练前预测与实际结果,检验 LLM 是否真正理解数据干预的效果,而非仅靠试错。复用 AutoDataBench 轨迹进行中期训练可提升下游编码性能。

正文

Authors:Ruifeng Yuan, Yizhi Li, Yaxin Du, Fengyu Cai, Yiqi Liu, Hou Pong Chan, Chenghua Lin, Yun Chen, Jian Yang, Bryan Dai, Pinyan Lu, Chenghao Xiao

View PDF HTML (experimental)

Abstract:Existing auto-research benchmarks often entangle multiple sources of improvement, including training frameworks, hyperparameters, compute budgets, and data, making it difficult to attribute why one frontier agent outperforms another to specific research capabilities. In this work, we isolate and systematically evaluate Data Intelligence: an agent's ability to understand, manipulate, and improve the data that shapes model capabilities. We introduce AutoDataBench, a controlled testbed built on a conceptual framework of data intelligence spanning data diagnosis, data organization, and data construction, instantiated through three highly curated optimization tasks while holding non-data factors fixed. Across tool use, retrieval, and knowledge injection, we evaluate frontier LLMs' ability to improve training data through iterative experimentation under task-specific resource budgets. Beyond optimization performance, we ask: do LLMs understand what their data interventions do? We compare predictions made before training with observed outcomes to seek evidence of data-effect reasoning beyond trial and error, and explore whether iterative feedback helps LLMs better understand how changes to training data affect model performance. Finally, we show that reusing AutoDataBench trajectories for mid-training improves downstream coding performance, highlighting its value in both evaluating data intelligence and generating high-quality training data. Code and resources are available at this https URL.
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2609.40097 [cs.CL]
  (or arXiv:2609.40097v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2609.40097

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Chenghao Xiao [view email]
[v1] Wed, 30 Sep 2026 16:33:24 UTC (448 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org