跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 8 天前AI 评分30

RegLLM:面向受监管智能体 AI 有限自主性的诊断测试框架

AI 导读

研究者提出 RegLLM,一个针对受监管智能体工作流中有限自主性的诊断测试框架,监测引用有效性、来源接地、schema 合规、升级正确性、宪法对齐和不安全动作率六项可信信号。

正文

View PDF HTML (experimental)

Abstract:We propose RegLLM, a diagnostic harness for bounded autonomy in regulated agentic workflows. It instruments six trustworthiness signals: citation validity, source grounding, schema compliance, escalation correctness, constitutional alignment, and unsafe-action rate. Signals are distinguished by their source of supervision: programmatic verifiers, task-level escalation labels, or AI-judge scores. A deterministic runtime supervisor blocks ungrounded answers and forces escalation, logging interventions. The same domain constitution informs evaluation, training rewards, and serving guardrails. Task-level should-escalate labels make the act-versus-defer decision a measurable training signal. We demonstrate the harness at smoke scale. An offline reference run (n=12) lifts escalation recall from 0 to 0.67 and reduces unsafe-action rate from 0.33 to 0.08 when governance is enabled. Two single-GPU Qwen2.5-3B LoRA/DPO pilots (n=8, same seed and evaluation split) expose substantial variation: nominally identical RL-base configurations yield task success of 0.25 versus 0.12 and escalation recall of 1.0 versus 0.5. An answer-quality adapter changes recall from 1.0 to 0.5 in Run A, but from 0.5 to 1.0 in Run B. An escalation-aware variant produces no measurable change in Run B. These small pilots do not establish reliable adapter effects or production readiness. Their contribution is diagnostic: configuration variance can overwhelm apparent tuning effects on bounded-autonomy metrics, motivating larger evaluation sets and repeated runs.
Comments: 9 pages; ancillary evaluation artefacts. Previously submitted to NLLP 2026
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as: arXiv:2609.37501 [cs.CL]
  (or arXiv:2609.37501v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2609.37501

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Dipankar Sarkar [view email]
[v1] Mon, 28 Sep 2026 11:08:38 UTC (37 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org