跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 7 天前AI 评分42

SaveRouter:让 LLM 路由的监督成本自己回本

AI 导读

SaveRouter 是一个稀疏监督的 LLM 路由框架,只选择性获取有信息量的模型反馈并在相关查询间共享能力信息,在四个路由基准上仅用约 33–41% 的可用训练反馈就保持或超过竞争方案的 routing 质量,并将盈亏平衡部署量较最快的传统路由器降低约 1.9–9.5 倍。该研究同时把监督开销与后续服务节省一起纳入评估,发现最小化服务成本的监督水平未必等于最早回本的监督水平。代码已公开。

正文

View PDF HTML (experimental)

Abstract:Large language model (LLM) routing reduces serving cost by assigning each query to an appropriate model while preserving response quality. Learning such a router, however, often requires executing multiple candidate models on historical queries to collect query--model quality feedback, creating a nontrivial supervision cost before deployment. Existing work largely focuses on serving-time efficiency, overlooking whether the resulting savings are sufficient to recover this upfront expenditure. We further observe that routing quality often saturates well before all query--model feedback is collected, suggesting that dense supervision can be economically over-provisioned. We propose SaveRouter, a sparse-supervision routing framework that selectively acquires informative model feedback and shares capability information across related queries, while retaining query-level refinement for fine-grained routing. We evaluate routing by jointly accounting for supervision expenditure and subsequent serving-time savings. Across four routing benchmarks, the main setting uses only about 33--41% of available training feedback while maintaining competitive or better routing quality, and reduces the break-even deployment volume by approximately 1.9--9.5 times compared with the fastest conventional router. Further analysis shows that acquiring more supervision is not always economically preferable: the supervision level that minimizes serving cost can differ from the one that achieves the earliest payback. Our code is publicly available at this https URL.
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2609.37402 [cs.AI]
  (or arXiv:2609.37402v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2609.37402

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Guannan Lai [view email]
[v1] Tue, 29 Sep 2026 12:42:02 UTC (485 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org