跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 8 天前AI 评分35

CineSubBench:用多语言电影字幕评测 LLM 的长篇叙事与文化理解

AI 导读

研究者推出 CineSubBench,一个基于多语言电影字幕评测 LLM 长上下文理解能力的基准,包含 1,012 部电影、六种语言共 6,072 条字幕轨和 813 万条带时间戳字幕。

正文

View PDF HTML (experimental)

Abstract:Large language models are increasingly evaluated in specialized domains such as law, medicine, software engineering, and cybersecurity, yet film remains comparatively underexplored despite requiring long-form narrative integration, multilingual interpretation, and culturally situated audience judgments. We introduce CineSubBench, a benchmark for evaluating long-context film understanding from multilingual movie subtitles. A subtitle track represents a film as thousands of short, temporally ordered utterances from which models must reconstruct characters, relationships, events, causal progression, and themes without explicit scene or event structure. CineSubBench contains 1,012 films with complete subtitle coverage in six languages, yielding 6,072 tracks and 8.13M timestamped subtitle entries. It provides a matched multi-task, multilingual, and multicultural (MultiX) evaluation setting: seven tasks span narrative reconstruction and abstraction, genre prediction, age suitability, country-specific motion-picture ratings across ten national classification systems, and subtitle-grounded language safety. Across nine LLMs, plot premises are recovered more reliably than event-complete synopses; cross-lingual consistency varies substantially across models and languages; national rating systems expose distinct calibration patterns; and strong profanity is far easier to ground than mild obscenity. CineSubBench establishes film as a long-context LLM evaluation domain and provides a unified benchmark for measuring narrative, multilingual, cultural, and evidence-grounding capabilities.
Comments: Preprint
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Emerging Technologies (cs.ET); Machine Learning (cs.LG)
Cite as: arXiv:2609.36218 [cs.CL]
  (or arXiv:2609.36218v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2609.36218

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Mir Tafseer Nayeem [view email]
[v1] Mon, 28 Sep 2026 20:16:09 UTC (4,486 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org