HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 9 天前AI 评分37
TT-VidT:解耦时间轴的高效运动中心视频预训练
AI 导读
TT-VidT 通过将 DINOv3 初始化的 ViT-B/16 逐帧空间路径与紧凑 Temporal Transfer Layer 结合,用 Diff Compression 从首帧外观锚点和逐帧运动 token 重建目标帧。
正文
Abstract:Comparisons in video self-supervised learning often evaluate complete training recipes rather than isolating the method itself: architecture, objective, data exposure, schedule, scale, and decoder capacity can all vary at once. This makes it hard to identify which choices yield motion-prioritized representations, whose gains concentrate on frame-to-frame change while retaining useful appearance. We address this with a matched $4 \times 6 = 24$ architecture-objective study at roughly 170M ~ 190M encoder scale on $\sim$1.7M OpenVid and Moments-in-Time v2 clips for 8 epochs, and propose TT-VidT. TT-VidT combines a DINOv3-initialized ViT-B/16 per-frame spatial path with a compact Temporal Transfer Layer, trained by Diff Compression to reconstruct target frames from a first-frame appearance anchor and frame-specific motion tokens. The sweep shows that TT3D with Diff Compression, not either component alone, enters the strongest motion-sensitive regime, and decoder ablations favor a compact video-pretrained decoder. In final comparison, TT-VidT leads Jester, Something-Something V2, ARID, and Diving48 fine-tuning simultaneously, improving over the strongest non-TT row by 54% ~ 121%, while using 48% fewer encoder FLOPs than DisMo and 55% fewer than VideoMAE or V-JEPA2. HMDB51, IARD, and EPIC-Kitchens bound the claim.
| Comments: | Accepted by NeurIPS 2026 main track, Project page: this https URL |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2609.33419 [cs.CV] |
| (or arXiv:2609.33419v2 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2609.33419 arXiv-issued DOI via DataCite |
Submission history
From: Shih Ying Yeh [view email]
[v1]
Sun, 27 Sep 2026 10:09:23 UTC (4,992 KB)
[v2]
Tue, 29 Sep 2026 05:59:33 UTC (4,988 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org