HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 7 天前AI 评分38
循环语言模型中的循环机制何时有效?HuggingFace 论文给出设计指南
AI 导读
一项受控实验系统考察了循环语言模型(LoopLMs)中循环机制的适用条件、作用位置与条件注入方式。结果显示,循环可提升超出训练范围的推理表现,但会损害知识类任务,且更难的推理样本未必获益更多;非循环输出层能提升欠展开时的鲁棒性。
正文
Abstract:Looped language models (LoopLMs) increase computational depth through parameter sharing, offering a path to scale inference computation without adding parameters. However, it remains unclear when additional recurrence is beneficial and how architectural choices affect its effectiveness. Through controlled experiments, we systematically examine (1) when recurrence helps, (2) where it should be applied, and (3) how its conditioning affects performance. Our evaluation covers inference budgets below, within, and beyond the training horizon under knowledge and reasoning tasks. (1) We find that recurrence can improve reasoning beyond the training horizon while degrading knowledge performance, but harder reasoning instances do not consistently benefit more. (2) Performance also depends on how distinct layers and recurrent iterations are allocated, showing that effective depth alone is insufficient to predict behavior. Non-recurrent output layers improve robustness to under-unrolling, while the preferred placement of input and output layers varies with inference budget. (3) Finally, we find that conventional initial-state injection offers limited robustness to varying recurrence depth. We therefore propose history-state injection as an alternative, and show that channel-wise history-state injection combined with timestep conditioning offers a low-cost and more effective design, better preserving knowledge under extended unrolling while improving robustness across inference budgets. Overall, our results clarify when recurrent computation helps, where it fails, and offer practical guidelines for designing LoopLMs across variable inference budgets.
| Comments: | Preprint, under-review |
| Subjects: | Machine Learning (cs.LG); Computation and Language (cs.CL) |
| Cite as: | arXiv:2609.36636 [cs.LG] |
| (or arXiv:2609.36636v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2609.36636 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Xinlin Zhuang [view email]
[v1]
Tue, 29 Sep 2026 03:43:51 UTC (1,100 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org