HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 8 天前AI 评分45
AdaGuard:支持用户自定义策略的自适应 Guard 模型
AI 导读
研究团队提出 AdaGuard 系列 Guard 模型(0.6B、4B、8B),可在推理时按用户自定义策略评估语言模型智能体的轨迹。
正文
Abstract:Guard models support the safe deployment of language model agents, but fixed risk taxonomies limit their ability to accommodate requirements that vary across applications and tasks. Under user-defined policies, detecting violations requires interpreting both the applicable rules and the agent's behavior, since identical actions can receive different judgments under different policies. To support learning this capability, we introduce AdaptiveSafety, a dataset of 10,939 training examples and 1,000 test examples covering policies with 1--100 rules. The dataset combines trajectories from multiple sources with policy and behavioral counterfactuals, pairing each example with an explanation and the complete set of violated rules. These counterfactuals expose changes that alter compliance, while structural augmentations provide supervision for consistency under rule reordering and identifier remapping. Building on this supervision, we propose SafePO, a reinforcement learning algorithm for refining violation identification while balancing explanatory reasoning and final verdicts. SafePO uses structured rewards to assess prediction correctness, retains group-relative advantages at the response level, and employs a separately trained value model to modulate token weights within explanation and verdict regions. Separate normalization controls their relative contribution to training despite differences in length. Through supervised initialization followed by SafePO, we develop AdaGuard, a family of 0.6B, 4B, and 8B guard models that assess agent trajectories under policies supplied at inference time. Our 4B model achieves binary accuracies of 89.30\% on AdaptiveSafety and 71.82\% on DynaBench. The project repository is available at this https URL
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2609.34241 [cs.AI] |
| (or arXiv:2609.34241v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2609.34241 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Yunhao Feng [view email]
[v1]
Mon, 28 Sep 2026 03:48:12 UTC (484 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org