HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 8 天前AI 评分38
SAKIKO:审计工具调用 LLM 内部干预,行为改变不等于修复
AI 导读
SAKIKO 框架用于审计工具调用 LLM 的内部激活干预,在 7 个 LLM 的 When2Call 和 MetaTool 上,通道定向干预让 5 个模型获得方向性净增益,59 个预算匹配的随机方向均无法复现校准目标增益。
正文
Abstract:Before invoking external tools, an agentic LLM must select among a K-way action space: executing a call, seeking clarification, answering directly, or declining. While internal activation steering can alter these pre-execution decisions, conventional aggregate metrics obscure where altered states land and what collateral damage they inflict. We present SAKIKO, an auditing framework that formalizes representation repair via directional error discovery, router-conditioned intervention, destination-resolved verification, and prospectively frozen statistical licensing. Across seven LLMs on When2Call and MetaTool, channel-keyed interventions induce direction-specific net gains in five models; across three sealed evaluations, none of 59 budget-matched random directions matches calibrated target gain. Crucially, destination auditing shows that behavioral movement does not equal repair: an intervention achieving +55 net gain corrupts over half of the baseline-correct decisions it touches, and promising point estimates on Qwen3-4B and Gemma-2-9B are formally declined due to finite-sample uncertainty. SAKIKO establishes the necessity of outcome-resolved adjudication before claiming internal repair. Code: this https URL.
| Comments: | Preprint |
| Subjects: | Computation and Language (cs.CL); Software Engineering (cs.SE) |
| Cite as: | arXiv:2609.36138 [cs.CL] |
| (or arXiv:2609.36138v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2609.36138 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Ruizhe Li [view email]
[v1]
Mon, 28 Sep 2026 19:12:55 UTC (1,660 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org