跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 2026-08-04精选AI 评分74

Video-DeepResearch:迈向下一代多模态深度研究智能体

AI 导读

Video-DeepResearch(Video-DR)将多模态智能体从静态图像扩展到连续视频流,提出解耦感知-探索流水线与分阶段工具解锁,以应对模态偏差和参数知识泄漏两大瓶颈。

推荐理由

将深度研究从静态图像扩展到视频流,通过强制视觉定位再检索的两阶段训练,让模型摆脱了对文本搜索的路径依赖,给视频智能体研发提供了可验证的训练范式。

正文

Authors:Zhen Fang, Yu Zeng, Wenxuan Huang, Yiming Zhao, Shiting Huang, Tianfei Ren, Qi Lu, Qingnan Ren, Qisheng Su, Lionel Z. Wang, Qingyu Yin, Shuang Chen, Zehui Chen, Lin Chen, Zhenfei Yin, Yao Hu, Shaohui Lin, Wanli Ouyang, Shaosheng Cao, Feng Zhao

View PDF HTML (experimental)

Abstract:We introduce Video-DeepResearch (Video-DR), extending multimodal agents from static images to continuous video streams, a setting that demands dense spatiotemporal grounding coupled with open-web exploration. Preliminary evaluations reveal two critical bottlenecks in current models: (1) modality bias, where agents bypass visual tools in favor of textual search, and (2) parametric knowledge leakage, where models rely on internal memory rather than genuine tool-augmented execution. To address these challenges, we propose Video-DR, featuring a decoupled perception-exploration pipeline with stage-wise tool unlocking that compels exhaustive cross-frame visual grounding prior to web retrieval. Our framework adopts a two-stage training recipe: supervised fine-tuning followed by Group Relative Policy Optimization (GRPO), enabling autonomous exploration that breaks the imitation-learning ceiling. Furthermore, we curate Video-DR-Bench, a human-AI collaborative benchmark comprising 200 complex, multi-hop VQA instances. Empirical results demonstrate that our Video-DeepResearch-35B-A3B establishes a new state-of-the-art of 64.0% average accuracy, surpassing proprietary Claude-4.5-Sonnet (59.0%) by 5.0 points and significantly outperforming GPT-5 (52.5%) and Gemini 2.5 Pro (57.5%). The 30B-A3B variant achieves 59.3%, competitive with Claude-4.5-Sonnet and demonstrating the effectiveness of our training paradigm even at compact scale. Code: this https URL.
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
Cite as: arXiv:2608.03979 [cs.CV]
  (or arXiv:2608.03979v1 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2608.03979

arXiv-issued DOI via DataCite

Submission history

From: Zhen Fang [view email]
[v1] Tue, 4 Aug 2026 17:45:16 UTC (23,284 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org