跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 6 天前AI 评分46

多智能体系统潜在通信的安全性研究

AI 导读

多智能体系统通过潜在通信直接交换内部表示,可降低 token、计算与延迟开销,但研究发现即使良性链路训练也会提高有害请求的顺从率。一种强化学习攻击在三种通信拓扑和四个安全基准上,将平均有害顺从分数从 27.9 提升至 76.9,同时在两个良性效用基准上取得更高平均准确率。调整奖励以偏向更安全行为可修复受损链路,无需更新智能体即可大幅降低有害顺从。

正文

View PDF HTML (experimental)

Abstract:Latent communication enables multi-agent systems to exchange information directly in internal representation space, reducing the token, computation, and latency overhead of text-based communication. To this end, lightweight trainable links are introduced to map the sender's representations into the receiver's input space. In this work, we show that even benign link training can increase harmful compliance relative to text-based communication while the underlying safety-aligned agents remain unchanged. An attacker can amplify this effect by optimizing the links on harmful query--response pairs or poisoning otherwise benign training data. We further develop a reinforcement-learning attack that rewards harmful compliance alongside benign task performance without requiring harmful target responses. Across three communication topologies and four safety benchmarks, this attack raises the mean harmful-compliance score from 27.9 with benignly trained links to 76.9. Compared with direct supervised optimization, it also achieves higher average accuracy on two benign utility benchmarks. Adapting the rewards toward safer behavior also enables repair of compromised links, substantially reducing harmful compliance across all evaluated attacks without updating the agents. Overall, our results show that safety alignment requires considering the multi-agent system as a whole. Code: this https URL
Subjects: Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Multiagent Systems (cs.MA)
Cite as: arXiv:2609.39788 [cs.AI]
  (or arXiv:2609.39788v2 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2609.39788

arXiv-issued DOI via DataCite

Submission history

From: Muhammad Huzaifa [view email]
[v1] Wed, 30 Sep 2026 14:09:57 UTC (105 KB)
[v2] Thu, 1 Oct 2026 09:39:34 UTC (105 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org