跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 8 天前AI 评分60

SEAD 提出基于状态视角的工具调用智能体攻防框架 DART 与 SAGE

AI 导读

SEAD 将工具调用智能体的攻击与防御建模为部分可观测的状态控制,并提出攻击方法 DART 与防御方法 SAGE,代码和数据已在 GitHub 开源。

正文

View PDF HTML (experimental)

Abstract:Language-model agents increasingly use tools to act on external systems. Earlier actions can alter files, permissions, database records, or other state, making a later routine-looking action harmful. Yet the visible interaction may not reveal the underlying state needed to assess that action. We formulate attack and defense as partially observed state control in SEAD, deriving their design requirements from this shared execution process. Because attackers supply instructions while the target chooses concrete actions, DART decomposes harmful goals into locally plausible steps and uses feedback from actual tool execution to guide trajectory search. The defender must decide before execution with incomplete state evidence. SAGE can therefore investigate relevant state through read-only queries before allowing or blocking each action, including those proposed after a block. We construct an environment-verifiable dataset integrating controlled initial states, replayable tool environments, and task-specific executable checks. Across four target models, DART improves semantic attack success by 18.8--35.9 percentage points over the competing baseline, with consistent gains under executable verification. On recorded trajectories, SAGE preserves 95.79% of benign trajectories while intercepting 92.73% of harmful paths by the harm-enabling boundary. In online attack-defense evaluation, it reduces DART's executable attack success from 48.0% to 4.0%. SAGE remains effective across four attack methods and generalizes to out-of-domain environments. Our code and data is available at this https URL.
Comments: Project Website: this https URL
Subjects: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
Cite as: arXiv:2609.34518 [cs.CR]
  (or arXiv:2609.34518v1 [cs.CR] for this version)
  https://doi.org/10.48550/arXiv.2609.34518

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Xinjie Shen [view email]
[v1] Mon, 28 Sep 2026 07:55:56 UTC (256 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org