HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 5 天前AI 评分50
关键词基准测试对小模型工具调用能力的误判:一套低成本诊断阶梯
AI 导读
一项研究记录了关键词匹配基准的假阳性:一对共享解码器与 tokenizer 的西班牙语安全模型(661.6M 与 1,109M 参数)在宽松工具调用指标上得分几乎相同(B4:0.660 vs. 0.650),但逐字复现检查中 600M 在 6/6 样例上生成有效工具调用,1B 在 0/6 上做到。
正文
Abstract:Keyword-matching benchmarks can credit small models for tool use they never perform. We document such a false positive in a matched-architecture pair of Spanish security language models and propose a ladder of strict, cheap diagnostics. A 661.6M parameter model (approx. 65% code/technical text; no dedicated SFT) and a 1,109M model (web-heavy multi-phase curriculum; 6B-token tool-SFT) share decoder, tokenizer, and special tokens, scoring almost identically on lenient tool-use metrics (B4: 0.660 vs. 0.650).
Verbatim-reproduction checks on training examples separate them completely: the 600M emits valid tool calls with generalized arguments on 6/6 examples; the 1B does so on 0/6 across checkpoints. A first-token probe localizes the 1B's failure to a missing prior (prob. $10^{-4}$--$10^{-5}$ on <|tool_call|>), which was erased by its web-heavy training phase. A targeted SFT recipe (diverse corpus, 5x higher learning rate, 2,202 steps, ~3.3 GPU-hours) repairs the 1B using three orders of magnitude fewer tokens than the failed phase. On all 269 corpus rows, valid emission rises from 0.100 to 0.959 (600M: 0.926). On 238 unseen prompts, the repaired 1B passes 0.536 vs. the 600M's 0.428 ($p = 0.004$). Embedding-drift checks show the repair did not move the trigger token's tied embedding (97.7% of the bf16 table remains bit-identical), meaning changes live in the surrounding network.
Both models over-trigger, rarely answering negative prompts without a call (0.09 for 600M, 0.17 for repaired 1B). Factorial analyses confirm all repair configurations install the format, though suppression benefits from a diverse corpus remain a hypothesis due to seed sensitivity. This cheap diagnostic ladder costs minutes of CPU time and should gate tool-use claims on small models.
| Comments: | 24 pages, 12 tables, preprint |
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2610.02142 [cs.CL] |
| (or arXiv:2610.02142v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.02142 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Juan Salas [view email]
[v1]
Thu, 1 Oct 2026 17:46:38 UTC (107 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org