跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 5 天前AI 评分40

KaliBench:面向 Kali Linux 网络安全工具使用的细粒度基准,支持无运行时可验证奖励

AI 导读

KaliBench 是面向 Kali Linux 自然语言到 CLI 翻译的细粒度基准与数据集,包含 8,504 条 query-command 对,覆盖 1,642 个工具、23 个能力维度和 5 个安全阶段。

正文

View PDF HTML (experimental)

Abstract:LLMs are increasingly applied to cybersecurity workflows, where they are expected to translate analysts' intent into tool invocations. However, existing evaluations focus on knowledge-based assessments or end-to-end agentic tasks, and do not directly measure LLMs' ability to generate executable commands for real-world cybersecurity tools. This gap is critical because cybersecurity operations rely on strict command-line interfaces (CLIs), where minor syntax errors, incorrect flag--value bindings, or argument misordering can invalidate execution. We introduce KaliBench, a fine-grained benchmark and dataset for natural-language--to--CLI translation on Kali Linux, comprising 8,504 query--command pairs spanning 1,642 tools across 23 capability dimensions and 5 security phases. KaliBench is constructed via a manuscript-grounded pipeline with deterministic canonicalization and alias-aware evaluation, enabling precise and reproducible assessment of tool selection and argument construction. To ensure both semantic correctness and practical executability, we develop a multi-stage verification pipeline that combines LLM-based validation, sandboxed terminal execution, and human-in-the-loop refinement. Building on these fine-grained, deterministic signals, KaliBench further enables runtime-free verifiable rewards for training. Across three evaluation modes and 24 configurations of general-purpose and security-focused open-weight models, no open-weight model exceeds 42% exact-command accuracy in the unrestricted setting, highlighting the difficulty of accurate CLI-based cybersecurity tool use without explicit tool hints. We further show that supervised fine-tuning and reinforcement learning with verifiable rewards derived from KaliBench significantly improve an 8B model and achieve performance comparable to a 685B MoE model.
Comments: Accepted at NeurIPS 2026 Evaluations and Datasets Track. Project page: this https URL | Github: this https URL
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR)
Cite as: arXiv:2610.02206 [cs.CL]
  (or arXiv:2610.02206v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.02206

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Naufal Suryanto [view email]
[v1] Thu, 1 Oct 2026 17:59:55 UTC (4,883 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org