跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 5 天前AI 评分41

OmniSeek:面向多轮音视频推理的原生工具集成框架

AI 导读

OmniSeek 是一个将 Omni-LLM 转变为原生工具调用多轮推理智能体的框架,让模型在推理中动态决定看或听、以及选取哪段时间窗口,把检索到的原始音视频片段回填上下文。

正文

View PDF HTML (experimental)

Abstract:We present OmniSeek, an agentic framework that transforms an Omni Large Language Model (Omni-LLM) into an active, multi-turn reasoning agent with native tool use. Rather than passively processing an entire audio-visual sequence in a single forward pass, OmniSeek makes evidence acquisition part of the reasoning process: it dynamically decides whether to look or listen, and over which temporal window, to retrieve sparse but critical evidence across different modalities within long contexts. Through an iterative multi-turn protocol, the retrieved raw audio or visual segments are appended back into the context to support subsequent reasoning. To cold-start this capability, we build a data engine that synthesizes OmniTraj-170K, a corpus of multi-hop Chain-of-Thought trajectories with interleaved audio and visual evidence. We first supervise the model on these trajectories to instill multi-turn tool-use behavior, and then further optimize the policy via a two-stage reinforcement learning with verifiable rewards. Moreover, we introduce an Audio-Visual Necessity objective that explicitly rewards successful trajectories whose reasoning depends on both modalities, discouraging single-modality shortcuts. Extensive experiments across a wide range of benchmarks demonstrate that OmniSeek learns adaptive cross-modal evidence seeking and consistently improves audio-visual reasoning performance.
Subjects: Computer Vision and Pattern Recognition (cs.CV)
Cite as: arXiv:2610.02181 [cs.CV]
  (or arXiv:2610.02181v1 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2610.02181

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Haibo Wang [view email]
[v1] Thu, 1 Oct 2026 17:58:16 UTC (1,336 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org