HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 8 天前AI 评分42
AI Night-Scientist:用强化学习训练 LLM 的智能体科研创意
AI 导读
AI Night-Scientist 是一个用强化学习教模型何时以及如何跳出可预测推理的智能体框架,基于行动、过程、结果三个维度用 GRPO 训练。相比基座模型,其科研提案的研究方向范围扩大 27.8%、贡献类型增加 14.9%,预测引用影响最高提升 32.0 个百分点,原创性提升 66.2 分。仅提高解码温度无法复现这些增益,指明追求何种创造力的语义引导才是关键。
正文
Abstract:Large language models (LLMs) excel at structured, verifiable tasks, but their low-entropy bias can produce homogeneous and predictable outputs, limiting their utility for open-ended scientific ideation. Effective discovery, however, spans a broader creative spectrum: from structured day science to loosely structured, serendipitous night science that reaches ideas beyond those typically considered. We introduce AI Night-Scientist, an agentic framework that uses reinforcement learning to teach models when and how to depart from predictable reasoning. Grounded in cognitive science, we model creativity along three axes: action (what to do and how creatively), process (when to explore versus exploit), and outcome (the novelty and usefulness of the resulting idea). We use these axes to train models with GRPO, exposing them to varying degrees and forms of creativity throughout training. This produces substantially more diverse scientific proposals, expanding the range of research directions by 27.8% and contribution types by 14.9% over the base model. It also improves predicted citation impact by up to 32.0 percentage points and originality by 66.2 points. These gains cannot be reproduced by simply increasing decoding temperature; instead, we find that semantic guidance specifying what kind of creativity to pursue is critical. Overall, our results suggest that creativity is a learnable, multi-level ability that can be shaped to help researchers reach ideas beyond those typically explored by LLMs.
| Comments: | Code: this https URL Website: this https URL |
| Subjects: | Artificial Intelligence (cs.AI); Computation and Language (cs.CL) |
| Cite as: | arXiv:2609.35706 [cs.AI] |
| (or arXiv:2609.35706v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2609.35706 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Priyanka Kargupta [view email]
[v1]
Mon, 28 Sep 2026 17:43:52 UTC (225 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org