跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 8 天前AI 评分43

Imprint Reader:从权重更新读出到行为干预

AI 导读

研究者提出 Imprint Reader,用 Semantic Mount-and-Read Tuning(SaRT)训练模型以自然语言描述冻结的权重更新,在留出更新上知识类 Pass@100 为 2%、行为类为 16%。

正文

View PDF HTML (experimental)

Abstract:As language models take a growing role in AI development, a natural aspiration is for them to reflect on their own learning process, as humans do, and use that reflection to improve themselves. At the same time, these models have an advantage that human learners lack, since training leaves parameter-level traces that can, in principle, be inspected directly. However, current models cannot decode these traces into an explicit account of what they have learned. To this end, we introduce the \textit{Imprint Reader}, a model trained with \textit{Semantic Mount-and-Read Tuning} (SaRT) to describe frozen weight updates. SMaRT mounts each update onto the Reader and uses an anchor-free meta-query to elicit a natural-language description, while no-change and random-perturbation controls discourage unsupported claims. On held-out updates, the joint Reader reaches judge-based Pass@100 of $2\%$ for knowledge and $16\%$ for behavior. These results demonstrate the feasibility of natural-language readout while pointing to reliability across updates as the next step. Beyond free-form generation, the Reader provides a differentiable proxy for the gap between a specified target behavior and a candidate weight update. Its coordinate-aligned gradients support intervention through MetaEdit. At a $0.5\%$ pruning rate, Reader-guided selection raises measured harmful-prompt refusal from $57.9\%$ to $64.1\%$ under a safety-maintenance target. Using behavior descriptions without target-task training data, MetaEdit increases the frequency of backtracking and sub-goal expressions in mathematical reasoning traces and raises BFCL Overall from $41.69\%$ to $44.60\%$.
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2609.35261 [cs.AI]
  (or arXiv:2609.35261v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2609.35261

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Guanxu Chen [view email]
[v1] Mon, 28 Sep 2026 14:21:00 UTC (525 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org