HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 9 天前AI 评分41
VisionHOPE:将视觉骨干网络构建为自修改学习系统
AI 导读
VisionHOPE 提出首个将视觉骨干网络形式化为自修改学习系统的方案,让模型在单张图像内"记住什么"与"如何学习"协同演化。它基于 Nested Learning 的自指构造,用五个耦合记忆分别存储内容、生成 key/value 表示并控制学习率与保留率,并推导出稳定性匹配的步长控制方案以保证记忆动态非扩张。
正文
Authors:Siran Peng, Tianshuo Zhang, Tianyu Fu, Weisong Zhao, Haoyuan Zhang, Jiankuo Zhao, Minghui Wu, Ping Jiang, Xiangyu Zhu, Chenxu Zhao, Zhen Lei
Abstract:Visual backbones have evolved from Convolutional Neural Networks (CNNs) with local aggregation to Vision Transformers (ViTs) with global interactions, State-Space Models (SSMs) with input-dependent state transitions, and Test-Time Training (TTT) layers that adapt an inner learner while processing an image. Across this progression, visual computation has become increasingly adaptive to each input, yet the rules governing that adaptation remain largely prescribed by the trained backbone. We introduce VisionHOPE, the first generic visual backbone formulated as a self-modifying learning system, in which what the model remembers and how it learns co-evolve within an image. Building on the self-referential construction of Nested Learning (NL), VisionHOPE realizes this co-evolution through five coupled memories that store content, generate key and value representations, and govern learning rate and retention. These memories evolve jointly as visual context accumulates along each scan. However, directly applying the unconstrained self-referential update to a visual backbone leads to instability. We therefore derive a stability-matched step-size control scheme that combines a soft cap on self-referential injection with a spectral clamp on the retained memory transition, and prove that the resulting memory dynamics are non-expansive along each scan. For two-dimensional feature maps, we adapt NL's chunk formulation by aligning chunks with image rows and columns across four directional scans. The proposed VisionHOPE achieves competitive results on ImageNet-1K, COCO, and ADE20K, establishing self-modifying learning systems as a practical foundation for general-purpose visual backbones. The code is available at this https URL.
| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2609.33325 [cs.CV] |
| (or arXiv:2609.33325v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2609.33325 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Siran Peng [view email]
[v1]
Sun, 27 Sep 2026 07:54:38 UTC (3,689 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org