HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 2026-05-27精选AI 评分70
ResearchMath-14K:通过智能体扩展研究级数学
AI 导读
本文介绍了ResearchMath-14K,这是一个包含14,056个研究级数学问题的数据集,通过多智能体流程从学术资料中策划而成,是目前此类规模最大的集合。研究还生成了ResearchMath-Reasoning(包含220K条教师轨迹),发现语言模型存在回避行为,且新一代模型产生的引用和虚假引用分别是旧模型的5.6倍和5.0倍。经过智能体过滤后,对参数规模为4B到30B的Qwen3模型进行微调,其平均得分比基础模型提高了9.2分,表明过滤后的开放问题尝试能为研究级数学推理提供有效监督。该数据集已公开发布。
推荐理由
这可能是目前数学推理方向最有价值的数据集之一,它暴露了模型编造引用的问题,过滤后微调还能涨点,做数学推理的团队应该立刻拉下来试试。
正文
Abstract:The frontier of mathematics is defined by problems whose solutions are not yet known. However, whether language models can meaningfully engage with such problems without human intervention remains unclear. A major obstacle is the lack of large-scale research-level math datasets. To this end, we introduce ResearchMath-14k, a set of $14{,}056$ problems curated from academic sources via a multi-agent pipeline. ResearchMath-14k spans 11 mathematical domains and ranks above existing math datasets on knowledge, novelty, and procedural difficulty. To our knowledge, it is the largest research-level mathematical problem set available for training. We additionally generate $220$K teacher trajectories through targeted prompting, followed by behavioral filtering. Notably, however, generating correct trajectories is nontrivial at this level, and two LLM judges label only $3.7\%$ and $4.3\%$ of sampled ResearchMath training trajectories as correct. Nevertheless, across three model families, full-parameter training on ResearchMath improves performance on graduate- and research-level mathematics benchmarks by $2.1$ points over the starting checkpoints. In comparison, training on existing datasets such as DASD and Nemotron-SFT-Math-v4 changes performance by $0.0$ and $-0.5$ points, respectively. Notably, mixing DASD with ResearchMath yields higher scores than token-matched DASD alone on benchmarks covering olympiad short-form ($+2.0$), graduate- and research-level short-form ($+0.8$), graduate- and research-level symbolic ($+2.6$), and proof evaluation ($+5.7$). Further analysis suggests that research-level mathematical content and greater reasoning diversity may help explain why ResearchMath provides complementary supervision to contemporary datasets. We make ResearchMath-14k publicly available for future works on research-level mathematical reasoning.
| Comments: | Work in progress. Dataset available at: this https URL |
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2605.28003 [cs.CL] |
| (or arXiv:2605.28003v2 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2605.28003 arXiv-issued DOI via DataCite |
Submission history
From: Guijin Son [view email]
[v1]
Wed, 27 May 2026 05:54:41 UTC (1,912 KB)
[v2]
Sun, 27 Sep 2026 12:55:48 UTC (1,230 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org