跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 7 天前AI 评分42

HeteroFold:面向异构多智能体 LLM 的免预填充跨模型族 KV Cache 传输

AI 导读

HeteroFold 是一种免预填充的跨模型族 KV cache 传输方法,可在发送方与接收方均冻结的前提下,对齐模型结构、将发送方 cache 映射到接收方空间并校准以保持接收方行为。

正文

View PDF HTML (experimental)

Abstract:Recent multi-agent LLM systems increasingly combine heterogeneous models for specialized agent roles. However, text-based communication requires each receiver to prefill shared context already processed by the sender. Reusing the sender's key-value (KV) cache avoids this redundancy, but prefill-free transfer across model families must handle differences in tokenization, model depth, and KV representations. To address these issues, we propose \textit{HeteroFold}, a prefill-free cross-family KV cache transfer method that keeps both the sender and receiver frozen. HeteroFold aligns model structures, maps the sender cache into the receiver space, and calibrates it to preserve receiver behavior. Across six transfer directions, HeteroFold achieves the best cache-transfer performance on all four long-context benchmarks and most short-context settings. It also matches text-based communication on the multi-agent benchmark. At 32K context length, Llama-3.1-8B$\rightarrow$Ministral-3-14B transfer is $10.7\times$ faster than Native Prefill and $1.18$--$1.47\times$ faster than the state-of-the-art prefill-free baselines, Dense Latent and KV Ridge. These results show that HeteroFold enables efficient cross-family KV reuse without receiver prefill.
Subjects: Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Multiagent Systems (cs.MA)
Cite as: arXiv:2609.32259 [cs.AI]
  (or arXiv:2609.32259v2 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2609.32259

arXiv-issued DOI via DataCite

Submission history

From: Vincent-Daniel Yun [view email]
[v1] Sat, 26 Sep 2026 05:30:18 UTC (493 KB)
[v2] Tue, 29 Sep 2026 19:41:30 UTC (737 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org