HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 8 天前AI 评分36
视频-音频联合与跨模态生成及编辑:统一形式化与设计分类体系
AI 导读
一篇综述提出统一形式化框架,将视频-音频联合生成、跨模态生成与联合编辑定义为同一音视频对分布上的三类问题,并按五个设计维度对方法进行分类。该工作称首次系统梳理联合音视频编辑,将其划分为九类编辑、涵盖 28 种编辑类型,并整理了各场景的方法、数据集与评测指标。
正文
Authors:Abhinav Sharma, Sai Karthik Navuluru, Wang Wei, Daksh Dangi, Xiangbo Gao, Li Li, Bo Ni, Vardhan Dongre, Junda Wu, Xiyang Hu, Jiuxiang Gu, Seunghyun Yoon, Tong Yu, Chien Van Nguyen, Mohamed Elmoghany, Nedim Lipka, Hoda Eldardiry, Hongjie Chen, Tyler Derr, Thien Huu Nguyen, Zhengzhong Tu, Nesreen K. Ahmed, Franck Dernoncourt, Ryan A. Rossi
Abstract:Video and audio are perceived together, yet most generative models treat them in isolation. We examine methods that model the two modalities jointly, generate one from the other, or edit them in a coupled manner, organized around a single question: how is the output kept coherent across modalities in time and semantics? A unified formulation casts joint generation, cross-modal generation, and joint editing as three problems defined on a single distribution over audio-visual pairs, and a taxonomy compares methods along five design axes. To our knowledge, this is the first overview to systematically taxonomize joint audio-visual editing, which we map as nine edit categories spanning 28 edit types. We describe methods, datasets, and metrics for each setting and close with the open problems we view as most consequential.
| Comments: | 36 pages, 3 figures, 15 tables |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM); Sound (cs.SD); Audio and Speech Processing (eess.AS) |
| Cite as: | arXiv:2609.34381 [cs.CV] |
| (or arXiv:2609.34381v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2609.34381 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Abhinav Sharma [view email]
[v1]
Mon, 28 Sep 2026 05:57:37 UTC (73 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org