HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 6 天前AI 评分36
SemanTok:为高效自回归视频生成提供可预测的语义 token
AI 导读
SemanTok 是一种灵活视频 tokenizer,将冻结的 DINO 特征输入编码器,并加入轻量 head,使每个保留的 token 前缀都能单独重建这些特征。201M 的 SemanTok AR 模型匹配或超越 3.4 倍规模的 VideoFlexTok AR 模型,更大模型还能进一步提升保真度。它在分布外类别上保持语义对齐,并在包括纯噪声在内的各噪声水平上为解码器提供更高语义对齐。
正文
Abstract:Recent video-based world models pair the scalability of autoregressive (AR) prediction with the visual quality of diffusion models. The choice of scene tokenizer is paramount for the optimal performance of each of these, both in terms of fidelity and semantics. Flexible-length, coarse-to-fine tokenizers yield exactly that: the first coarse tokens carry the clip's global semantics while later tokens further specify details. Existing flexible tokenizers only apply a representation-alignment (REPA) loss on early decoder hidden states, a target the decoder can partly meet from its noised input instead. We introduce SemanTok, a flexible video tokenizer that feeds frozen DINO features into its encoder and adds lightweight heads that reconstruct them from each retained token prefix alone. SemanTok achieves high semantic alignment and video fidelity at every AR model size: a 201M SemanTok AR model matches or beats a VideoFlexTok AR model $3.4\times$ its size, and larger SemanTok AR models further improve fidelity. It keeps semantic alignment on out-of-distribution classes and gives the decoder higher semantic alignment at every noise level, including pure noise. It performs well in both reconstruction and generation, and its short token prefixes are cheaper to predict and give better generation fidelity, with pixel detail deferred to later tokens.
| Comments: | 29 pages, 22 figures, including references and appendix; 9 pages of main text |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.00686 [cs.CV] |
| (or arXiv:2610.00686v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00686 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Mikhail Dereviannykh [view email]
[v1]
Wed, 30 Sep 2026 20:29:37 UTC (13,998 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org