跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 7 天前AI 评分37

PixelUMM:无需编码器的统一图像与视频理解生成模型

AI 导读

PixelUMM 是一个无需视觉编码器的统一多模态模型,直接在像素空间完成图像与视频的理解和生成。它将图像表示为空间 patch、视频表示为时空 tubelet,通过单层线性投影接入共享多模态主干,并采用 Mixture-of-Transformers 架构结合共享注意力与任务专属参数,同时支持自回归文本预测和像素空间流匹配。

正文

Authors:Cong Wei, Xuanchi Ren, Bryan Chu, Weiming Ren, Huan Ling, Jiahui Huang, Laura Leal-Taixé, Sanja Fidler, Wenhu Chen, Zian Wang, Jay Zhangjie Wu

View PDF HTML (experimental)

Abstract:Unified Multimodal Models (UMMs) often rely on separate visual representations for understanding and generation, increasing visual context length and complicating integration with established vision-language pretraining pipelines. Recent advances in pixel-space modeling offer an encoder-free alternative, but extending this paradigm from images to videos is non-trivial: video understanding and generation adopt different temporal representations, leaving the design of a unified visual interface an open question. We present PixelUMM, an encoder-free model for unified image and video understanding and generation directly in pixel space. PixelUMM represents images as spatial patches and videos as spatiotemporal tubelets, connecting raw pixels to a shared multimodal backbone through single-layer linear projections. Its Mixture-of-Transformers architecture combines shared attention with task-specific parameters and extends clean-pixel prediction to video generation, jointly supporting autoregressive text prediction and pixel-space flow matching. Experiments show that PixelUMM achieves competitive performance across image and video understanding and generation tasks. We further conduct empirical studies of key design choices, including decoder design and spatial-temporal patch size, providing insights for future pixel-space unified multimodal models.
Comments: 31 pages, 21 figures. Project page: this https URL
Subjects: Computer Vision and Pattern Recognition (cs.CV)
Cite as: arXiv:2609.38597 [cs.CV]
  (or arXiv:2609.38597v1 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2609.38597

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Cong Wei [view email]
[v1] Tue, 29 Sep 2026 21:59:07 UTC (21,766 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org