HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 5 天前AI 评分42
VTR-Bench:评估视频生成中视觉文本渲染的系统性基准
AI 导读
VTR-Bench 是用于评估视频生成模型视觉文本渲染能力的系统性基准,包含覆盖广告、科学视频等五类场景的 300 条提示词,并配套自动化评估流程。在 11 个 SOTA 模型上的实验显示,场景文本渲染普遍存在困难,表现最好的模型整体词错误率(WER)为 0.250。
正文
Authors:Yu Huang, Jungang Li, Zhiyuan Wang, Yonghua Hei, Song Dai, Jiayu Yang, Deyuan Liu, Xiang Zheng, Xiaoshuang Shi, Hao Cheng, Kaidi Xu
Abstract:Recent video generation models can produce highly realistic videos from natural language instructions, with visual quality approaching cinematic standards. Existing evaluation benchmarks, however, predominantly assess visual quality, aesthetic appeal and physical plausibility, while paying limited attention to text, an essential medium for conveying information in everyday scenes. A generated video may appear visually compelling and feature lifelike subjects, yet still render the text within the scene incorrectly. To address this overlooked dimension, we introduce \textbf{VTR-Bench}, a systematic benchmark for evaluating the \textbf{V}isual \textbf{T}ext \textbf{R}endering capabilities of video generation models. VTR-Bench situates text within concrete application scenarios, such as advertisements and scientific videos, with 300 carefully constructed prompts spanning five scenario categories. We develop an automated evaluation pipeline with human alignments that separately assesses text fidelity through carrier-specific transcription and scene and motion requirements through a prompt-specific chain of query. Beyond evaluation, we introduce a \textbf{Keyframe-Guided Agentic Framework} in which a Director agent coordinates image and video generation with visual evaluation, guiding iterative refinement and candidate selection through visual feedback. Experiments on 11 state-of-the-art models reveal widespread difficulties in accurately rendering scene text, with the best-performing model recording an overall word error rate (WER) of 0.250. We further analyze text rendering failures to characterize the challenges faced by current video generation models. These findings highlight visual text rendering as a key challenge for video generation and demonstrate a practical path toward improvement. Code is available at this https URL.
| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2610.01499 [cs.CV] |
| (or arXiv:2610.01499v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2610.01499 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Yu Huang [view email]
[v1]
Thu, 1 Oct 2026 11:40:03 UTC (20,434 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org