HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 10 天前AI 评分45
SkillDRE:通过执行前与运行时反馈对 Agent 技能进行双阶段红队演化
AI 导读
SkillDRE 是一个全自动框架,通过执行前扫描与运行时防御反馈的双阶段闭环,演化完整的恶意 Agent 技能包。在 SkillsBench 上对四个受害模型评测,平均攻击成功率达 45.28%,超过最强基线 40.3%,且最终提交的技能未触发任何 SkillScan 告警,并基本保留良性任务性能。结果表明,两阶段防御反馈可作为自适应红队的有效学习信号,单独评估任一防御阶段都会遗漏攻击能力。
正文
Abstract:Agent skills package instructions, executable code, and task-specific resources into reusable artifacts that agents can improve using execution feedback. The same mechanism also enables attackers to evolve malicious skills, making them more effective and less detectable. However, a candidate skill may pass pre-execution scanning yet fail to realize its target under runtime defenses, while a revision that repairs execution may introduce new scanner findings. We introduce SkillDRE, a fully automated framework for evolving complete malicious skill packages through a dual-stage feedback loop. Given a benign task and its associated skills, SkillDRE autonomously constructs and validates a task-conditioned malicious objective and a verifiable judge rule. It then holds both fixed while evolving the skill implementation, with preservation of legitimate task capability. SkillDRE combines scanner-guided evolution with runtime-guided refinement informed by execution outcomes observed under runtime defense. Each runtime-guided revision returns to the pre-execution stage for rescanning and further optimization before re-execution, forming a cross-stage closed loop. Evaluated on SkillsBench across four victim models, SkillDRE achieves an average attack success rate of 45.28%, exceeding the strongest baseline by 40.3%, while its final submitted skills receive no SkillScan findings and largely preserve benign-task performance. These results show that two-stage defense feedback can serve as a useful learning signal for adaptive red teaming and that evaluating either defense stage in isolation can miss the resulting attack capability. Codes is available at this https URL
| Subjects: | Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2609.32400 [cs.CR] |
| (or arXiv:2609.32400v1 [cs.CR] for this version) | |
| https://doi.org/10.48550/arXiv.2609.32400 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Pengyu Zhu [view email]
[v1]
Sat, 26 Sep 2026 09:24:09 UTC (2,428 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org