跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 9 天前AI 评分37

TT-VidT:解耦时间轴的高效运动中心视频预训练

AI 导读

TT-VidT 通过将 DINOv3 初始化的 ViT-B/16 逐帧空间路径与紧凑 Temporal Transfer Layer 结合,用 Diff Compression 从首帧外观锚点和逐帧运动 token 重建目标帧。

正文

View PDF HTML (experimental)

Abstract:Comparisons in video self-supervised learning often evaluate complete training recipes rather than isolating the method itself: architecture, objective, data exposure, schedule, scale, and decoder capacity can all vary at once. This makes it hard to identify which choices yield motion-prioritized representations, whose gains concentrate on frame-to-frame change while retaining useful appearance. We address this with a matched $4 \times 6 = 24$ architecture-objective study at roughly 170M ~ 190M encoder scale on $\sim$1.7M OpenVid and Moments-in-Time v2 clips for 8 epochs, and propose TT-VidT. TT-VidT combines a DINOv3-initialized ViT-B/16 per-frame spatial path with a compact Temporal Transfer Layer, trained by Diff Compression to reconstruct target frames from a first-frame appearance anchor and frame-specific motion tokens. The sweep shows that TT3D with Diff Compression, not either component alone, enters the strongest motion-sensitive regime, and decoder ablations favor a compact video-pretrained decoder. In final comparison, TT-VidT leads Jester, Something-Something V2, ARID, and Diving48 fine-tuning simultaneously, improving over the strongest non-TT row by 54% ~ 121%, while using 48% fewer encoder FLOPs than DisMo and 55% fewer than VideoMAE or V-JEPA2. HMDB51, IARD, and EPIC-Kitchens bound the claim.
Comments: Accepted by NeurIPS 2026 main track, Project page: this https URL
Subjects: Computer Vision and Pattern Recognition (cs.CV)
Cite as: arXiv:2609.33419 [cs.CV]
  (or arXiv:2609.33419v2 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2609.33419

arXiv-issued DOI via DataCite

Submission history

From: Shih Ying Yeh [view email]
[v1] Sun, 27 Sep 2026 10:09:23 UTC (4,992 KB)
[v2] Tue, 29 Sep 2026 05:59:33 UTC (4,988 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org