HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 8 天前AI 评分42
TokenCast:预测 LLM 智能体执行过程中的 token 消耗
AI 导读
TokenCast 通过学习每个执行片段的可组合成本表示,在 LLM 智能体运行过程中实时预测 token 消耗,无需额外调用 LLM,在 SWE-bench Verified 上平均累计预测耗时 32.8 ms。
正文
Abstract:When a large language model (LLM) agent executes the same task, token consumption can vary by over an order of magnitude across runs. The agent chooses its next steps based on tool feedback and intermediate results, while the growing context steadily inflates the input size of every subsequent call. The total consumption of a task is therefore hard to predict before execution and the prediction must be revised as the run unfolds. In this paper, we propose TokenCast, which learns a composable cost representation for each execution segment, recording its own consumption and the context growth it introduces. Composing adjacent segments yields a cumulative estimate that captures the extra input cost incurred when context from earlier segments is re-read by every later call. As execution unfolds, newly observed evidence refreshes the forecast, requiring no additional LLM calls and incurring a mean cumulative prediction time of 32.8 ms per run on SWE-bench Verified. Across 4 task suites and 6 agent models, TokenCast's mean absolute error reduction against the strongest comparator averages 14.5% over 96 evaluated combinations. In offline budget-control replay, TokenCast uses 21.3% fewer tokens on average than a fixed-budget policy at matched trace completion. The code is available at this https URL.
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Software Engineering (cs.SE) |
| Cite as: | arXiv:2609.35760 [cs.LG] |
| (or arXiv:2609.35760v2 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2609.35760 arXiv-issued DOI via DataCite |
Submission history
From: Chaoqian Ouyang [view email]
[v1]
Mon, 28 Sep 2026 17:59:09 UTC (1,159 KB)
[v2]
Tue, 29 Sep 2026 10:45:52 UTC (1,159 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org