跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 2026-06-11精选AI 评分79

MaxProof:面向数学证明的群体级别测试时扩展框架(MiniMax-M3)

AI 导读

MaxProof 是为 MiniMax-M3 系列设计的群体级别测试时扩展框架,用于竞赛级数学证明。M3 模型训练了证明生成、证明验证和基于 critique 的证明修复三种能力,验证器采用低假阳性率的深度防御生成式架构。这些能力合并到单个 M3 模型。测试时,MaxProof 将模型用作生成器、验证器、精炼器和排序器,在候选证明群体中搜索并通过锦标赛选择返回最终证明。M3 模型在 IMO 2025 达 35/42,USAMO 2026 达 36/42,均超过人类金牌阈值。

推荐理由

MiniMax-M3用生成-验证器RL把数学证明推到了人类金牌水平,IMO 2025 35/42,USAMO 2026 36/42。这篇的意义不只分数,而在于验证-修复-群体搜索的技术路线跑通了最难的人类竞赛。

正文

Authors:Jiacheng Chen, Xinyu Zhang, Shunkai Zhang, Yanmohan Wang, Lin Li, Tiancheng Qin, Qin Wang, Zhengmao Zhu, Tianle Li, Jingyang Li, Zehan Li, Binyang Jiang, Jin Zhu, Han Ding, Fei Yu, Chenyu Du, Zijian Song, Jiayuan Song, Zhi Zhang, Yunan Huang, Weiyu Cheng, Pengyu Zhao, Yu Cheng

View PDF HTML (experimental)

Abstract:We present MaxProof, a population-level test-time scaling framework for competition-level mathematical proof in the MiniMax-M3 series. M3 first trains three proof-oriented capabilities -- proof generation, proof verification, and critique-conditioned proof repair -- using a defense-in-depth generative verifier engineered for low false-positive rate. These capabilities are merged into a single released M3 model. At test time, MaxProof treats the model as a generator, verifier, refiner, and ranker, searches over a population of candidate proofs, and returns one final proof through tournament selection. With MaxProof test-time scaling, the M3 model reaches 35/42 on IMO 2025 and 36/42 on USAMO 2026, exceeding the human gold-medal threshold on both.
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
Cite as: arXiv:2606.13473 [cs.LG]
  (or arXiv:2606.13473v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2606.13473

arXiv-issued DOI via DataCite

Submission history

From: Jiacheng Chen [view email]
[v1] Thu, 11 Jun 2026 15:27:06 UTC (2,912 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org