跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 6 天前AI 评分39

JEV 等直接决策模型存在序数尺度利用偏差,BA-LoRA 后训练可将利用率从约 47% 提升至 86%

AI 导读

研究分析 JEV 1.13 与三个开源 KEV 模型,发现直接决策模型存在"序数尺度利用偏差":在 ANLI 上 JEV 准确率 74.95%,却把 38.8% 的预测和 51.3% 的错误都归为 Neutral。

正文

View PDF HTML (experimental)

Abstract:Direct-decision models turn text into low-latency structured labels and scores, making them attractive for classification and automatic evaluation. Yet reliability requires more than accuracy: a model must also use the ordinal decision scale supplied by the user faithfully. We analyze JEV~1.13 and three open KEV models. Our investigation begins with ANLI, where JEV assigns 38.8\% of all predictions and 51.3\% of errors to Neutral despite 74.95\% accuracy, nearly balanced gold labels, and balanced candidate positions. Across 36 ordinal datasets, final decisions use only 67--76\% of the effective gold support, versus 87--102\% on four nominal tasks. Randomizing candidate order weakens but does not remove this compression. Holding items and source scores fixed while balancing gold support and positions, we refine scales from $K=2$ to $14$; utilization falls for every model and reaches 26--75\% at $K=14$, although candidate probabilities remain broad for most models. Targeted BA-LoRA post-training raises gold-relative utilization from roughly 47\% to 86\% on eight supervised scales at both KEV sizes, showing that the compression is learned and modifiable rather than an immutable architectural limit. We call this ordinal scale-utilization bias: decision-stage candidate-space compression distinct from accuracy, gold imbalance, fixed position, and candidate count alone. The code and data are available at this https URL
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2609.38827 [cs.AI]
  (or arXiv:2609.38827v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2609.38827

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Jinzhe Li [view email]
[v1] Wed, 30 Sep 2026 02:50:30 UTC (327 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org