HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 6 天前AI 评分36
CoEvoWhen:面向超长视频时序定位的策略-工具协同演化
AI 导读
研究者提出 CoEvoWhen 策略-工具协同演化框架,从 VLM 的智能体推理轨迹中联合演化高层策略与可执行媒体工具,形成可复用技能且无需更新模型参数。在五个 benchmark 和三个 VLM 上,该方法持续提升超长视频时序定位准确率,同时降低推理时的视觉 token 开销,演化出的技能在通用长视频 QA 上也取得显著提升。
正文
Abstract:Ultra-long video temporal grounding requires balancing long-range evidence search with fine-grained event understanding under a limited visual budget, yet existing agentic methods still rely largely on predefined policies and tool capabilities. Motivated by this, we propose a novel policy-tool coevolution framework that jointly evolves high-level policies and executable media tools from the agentic reasoning trajectories of a VLM, forming a reusable skill without updating model parameters. During evolution, an external skill updater distills transferable task experience in long-video temporal grounding, accordingly refining the orchestration of long-range image-based and fine-grained video-based observations. Alongside these policy updates, the updater employs its coding capabilities to upgrade existing tools or create new ones, adapting the tools to long-video evidence acquisition. Equipped with the evolved skill, the VLM autonomously orchestrates tools under the guidance of the evolved policy, coordinating image and video observations for agentic inference without relying on a separate, stronger planning model. Extensive experiments spanning five benchmarks and three VLMs show that policy-tool coevolution consistently improves temporal grounding accuracy in ultra-long videos while reducing visual token cost at inference, and that the evolved skill yields substantial performance gains on general long-video QA without additional task-specific evolution, demonstrating the effectiveness and generalizability of our framework for long-video understanding.
| Comments: | Project page: this https URL |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2609.40048 [cs.CV] |
| (or arXiv:2609.40048v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2609.40048 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Yiduo Jia [view email]
[v1]
Wed, 30 Sep 2026 16:14:14 UTC (15,576 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org