跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 8 天前AI 评分30

The Endless Exam:面向超级智能的数学构造基准,覆盖 14 类参数化问题族

AI 导读

研究者推出 The Endless Exam 基准,涵盖 14 个参数化数学构造问题族,用可验证的连续质量分衡量模型在已发表数学前沿内外的进展,且不将改进上限封顶在 1。在 69 个实例上评测 9 个模型,虽无模型超越 30 个已发表前沿参考,连续分数仍能区分性能差异,并给出规模-质量曲线。生成器、验证器、参考、模型回答与分析均已开源。

正文

View PDF HTML (experimental)

Abstract:We introduce the Endless Exam, a benchmark spanning fourteen parameterised families of mathematical construction problems, with verifiable scores that distinguish progress before and beyond published mathematical frontiers. Each submitted object is checked automatically for validity and assigned a relative quality score against a published frontier or construction baseline, without capping improvements at 1. The benchmark draws long-term challenges from open mathematical problems and generates larger instances by varying their parameters. Compact certificates allow large constructions to be verified without listing every element. Across nine models evaluated on 69 distinct instances, continuous quality scores distinguish performance even though none of the 30 published-frontier references is surpassed. Size-quality curves show how construction quality changes as problem size increases. We release the generators, verifiers, references, model responses and analysis to support continued measurement before and beyond human frontiers.
Comments: 54 pages; added Claude Opus 5.5 evaluations, clarified reference baselines and verification limits, and revised presentation. Uses the unchanged bench-v1.0 evaluation suite
Subjects: Artificial Intelligence (cs.AI); History and Overview (math.HO)
Cite as: arXiv:2609.24555 [cs.AI]
  (or arXiv:2609.24555v2 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2609.24555

arXiv-issued DOI via DataCite

Submission history

From: Muhan Zhang [view email]
[v1] Mon, 21 Sep 2026 13:19:36 UTC (413 KB)
[v2] Mon, 28 Sep 2026 12:36:08 UTC (410 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org