HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 8 天前AI 评分37
PlaylistEval:视频语言模型做评判者,在日级长视频上还可靠吗?
AI 导读
PlaylistEval 是一个无需人工标注的智能体框架,可在超过 100 小时的播放列表视频集合上构建视频语言评判基准,自动生成差异由因果退化控制、必须跨集合检索才能作答的问答对。
正文
Abstract:Video-language models are increasingly used as judges of video understanding, both for evaluating model outputs and for training reward models. Whether their judgments remain reliable when the evidence is buried in day-long videos has yet to be established. Existing benchmarks cannot answer this. Their videos are typically only a few minutes long, many answer pairs can be separated from the transcript alone, and collecting human judgments does not scale to ultra-long videos. We introduce PlaylistEval, an agentic framework that builds video-language judge benchmarks over 100-hour playlist collection without human annotation. It automatically generates questions with paired answers whose differences are controlled by causal degradation, so that every pair demands retrieval across the collection. The resulting benchmark contains 630 pairs across seven domains spanning both static and dynamic knowledge, and on a stratified subset of 152 pairs it agrees with human judgments 93.0% of the time (IAA 0.781). Evaluating 17 omnimodal and multimodal models from eight families reveals that frontier judges reach only 75.4% pairwise accuracy, while open-source judge models perform far behind. We further show that both retrieval and final judgment depend on using multiple modalities, and that judge accuracy degrades as the playlist set grows. We release our pipeline, benchmark, and evaluation code at this https URL.
| Comments: | 53 pages, 15 figures, 23 tables. Project page: this https URL |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL) |
| Cite as: | arXiv:2609.34314 [cs.CV] |
| (or arXiv:2609.34314v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2609.34314 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Shayekh Bin Islam [view email]
[v1]
Mon, 28 Sep 2026 04:54:29 UTC (5,924 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org