HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 8 天前AI 评分43
如何度量同构面板 LLM 辩论中的崩塌与纠正
AI 导读
研究提出一套可审计的同构 LLM 辩论协议,用 transition ledger 记录崩塌、纠正、起始与带符号干预效用。在 6,925 场 MMLU-Pro 辩论中识别出 253 次崩塌,并发现留一模型探针门控冻结虽能阻止 29 次崩塌,却损失 108 次纠正。
正文
Abstract:Multi-agent large language model (LLM) debate is often evaluated by whether final answers improve, but movement is not necessarily improvement: the same discussion can rescue an initially wrong majority or destroy an initially correct one. Standard final-accuracy evaluations conflate these opposing mechanisms. We introduce an auditable protocol for homogeneous debate on multiple-choice questions (MCQs) that records each run as a transition ledger over collapse, correction, onset, and signed intervention utility. On 6,925 MMLU-Pro debates, the protocol identifies 253 collapses and a parallel correction ledger that changes how interventions should be judged. Replay experiments reveal the central tradeoff: a leave-one-model-out probe-gated freeze prevents 29 collapses but loses 108 corrections under equal weights, so collapse prevention alone can recommend the wrong policy. A compact pre-debate 8-probe screen is a triage signal: its unadjusted family-level association with conditional-collapse risk is high (G=7, Spearman rho=0.893, exact two-sided p=0.0123), but initial-majority accuracy is a close comparator (rho=0.821; family partial rho=0.767, p=0.0877), so we do not treat it as calibrated or capability-adjusted prediction. Round-level traces localize many collapses to the first debate round, where early disagreement can precede both harmful cascades and useful recovery. We release replayable schemas, coders, audits, cost cards, and zero-API rebuild scripts so future model-scaffold rows can be compared under the same denominators and signed utility ledger.
| Comments: | Accepted at NeurIPS 2026 (Evaluations and Datasets Track). Project page: this https URL. Code: this https URL |
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2609.35279 [cs.CL] |
| (or arXiv:2609.35279v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2609.35279 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Xin Li [view email]
[v1]
Mon, 28 Sep 2026 14:31:19 UTC (536 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org