跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 2026-06-08精选AI 评分76

精确性不等于忠实度:完整Oracle下的覆盖感知接地生成评估

AI 导读

无参考忠实度度量仅衡量精确率(陈述是否被支持),鼓励模型少说甚至不说以获得高分。本研究利用F1遥测(确定性完整ground truth)和NOAA天气预报两个完整Oracle领域,证明此盲点:在多语言(EN/ES/PT)共7253个决策实例(覆盖150场比赛)的基准上,最精确的前沿模型仅覆盖不到一半相关事实,按F1排名垫底。引入覆盖度(召回率)后系统排序改变;显式要求详尽也无法弥补差距。作者提出将忠实度与覆盖度合并为单一分数,并给出无参考验证器引导生成方法,同时提升精确率和召回率。相关基准、标注、度量、基线及交互演示已开源。

推荐理由

这个研究戳破了自动评估里 Faithfulness 的泡沫,指标只看模型「说对多少」不看「说全没有」,沉默的模型反而拿高分,以后评测不能只看精确度了,做评估的得补上覆盖度这一环。

正文

View PDF HTML (experimental)

Abstract:Reference-free faithfulness metrics verify each atomic claim a model makes against ground truth, and are increasingly used to evaluate grounded generation. We show they share a blind spot: they measure only precision -- are the stated claims supported? -- and therefore reward abstention, since a model can score near-perfect faithfulness by saying almost nothing. We make this measurable using Formula 1 telemetry, a domain where strategic ground truth is derived deterministically and, crucially, completely: for each decision we know the full set of facts that mattered. This completeness -- absent in open-domain faithfulness benchmarks -- lets us measure recall (coverage of the relevant facts) exactly, alongside precision. On a multilingual (EN/ES/PT) benchmark of 7,253 decision instances spanning 157 races, the most precise frontier model covers under half of the relevant facts and ranks last by F1, so requiring coverage reorders the systems; the same effect reappears in a second complete-oracle domain (NOAA weather forecasts). Fine-tuning small models (1B-7B) on the complete oracle closes the precision-recall gap entirely (F1 ~0.98), beating every zero-shot frontier system regardless of scale. We pair faithfulness with coverage into a single score, validate the metric (controlled perturbation; agreement across a model-free regex extractor and a cross-family LLM extractor, system-level Spearman 1.0), and give a verifier-guided generation method that improves precision and recall without references. We release the benchmark, structured annotations, metric, baselines, and an interactive demo.
Comments: 9 pages. v2: adds Anthropic Claude + 3 additional fine-tuned bases (1B-7B); 6 frontier families x 3 languages. Code this https URL Demo this https URL
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2606.09376 [cs.CL]
  (or arXiv:2606.09376v2 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2606.09376

arXiv-issued DOI via DataCite

Submission history

From: Juan Salas [view email]
[v1] Mon, 8 Jun 2026 11:56:25 UTC (44 KB)
[v2] Tue, 16 Jun 2026 02:01:49 UTC (47 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org