HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 8 天前AI 评分51
Draft-KV:学习语言模型间有用的潜态通信
AI 导读
论文提出 Draft-KV,让共享模型在起草答案时生成 key-value 状态,经线性投影写入侧记忆并由门控注意力分支读取,两模型均冻结,接口仅训练 1.05M 参数。
正文
Abstract:Latent communication passes internal states between language models instead of decoded text, but higher receiver accuracy does not show that the receiver used the message content. Across five method-dataset pairs, replacing each message with one from an unrelated question changes accuracy by at most 0.60 points, even when communication adds 15.44 points over the receiver alone. Thus the interface can supply the gain while making the sharer dispensable. Draft-KV instead sends the key-value states formed while the sharer drafts an answer to the current question. Linear projections place these states in a side memory read through a gated attention branch, and progressive training moves from message reconstruction to answer supervision under a guard on harm from mismatched messages. Both models remain frozen and the interface trains 1.05M parameters, 348x fewer than C2C. With a Qwen3-8B sharer, a frozen Qwen2.5-0.5B-Instruct receiver reaches 78.04% on MMLU-Redux, versus 37.45% alone and 36.40% with reassigned messages. At fixed interface size, scaling the sharer from 0.6B to 8B raises accuracy from 46.11% to 78.04%; communication also transfers to held-out tasks and can exceed both models when each holds different evidence.
| Comments: | 41 pages, 7 figures, 13 tables. Code: this https URL |
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2609.34754 [cs.CL] |
| (or arXiv:2609.34754v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2609.34754 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Linquan Wu [view email]
[v1]
Mon, 28 Sep 2026 09:42:01 UTC (479 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org