HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 8 天前AI 评分30
RegLLM:面向受监管智能体 AI 有限自主性的诊断测试框架
AI 导读
研究者提出 RegLLM,一个针对受监管智能体工作流中有限自主性的诊断测试框架,监测引用有效性、来源接地、schema 合规、升级正确性、宪法对齐和不安全动作率六项可信信号。
正文
Abstract:We propose RegLLM, a diagnostic harness for bounded autonomy in regulated agentic workflows. It instruments six trustworthiness signals: citation validity, source grounding, schema compliance, escalation correctness, constitutional alignment, and unsafe-action rate. Signals are distinguished by their source of supervision: programmatic verifiers, task-level escalation labels, or AI-judge scores. A deterministic runtime supervisor blocks ungrounded answers and forces escalation, logging interventions. The same domain constitution informs evaluation, training rewards, and serving guardrails. Task-level should-escalate labels make the act-versus-defer decision a measurable training signal. We demonstrate the harness at smoke scale. An offline reference run (n=12) lifts escalation recall from 0 to 0.67 and reduces unsafe-action rate from 0.33 to 0.08 when governance is enabled. Two single-GPU Qwen2.5-3B LoRA/DPO pilots (n=8, same seed and evaluation split) expose substantial variation: nominally identical RL-base configurations yield task success of 0.25 versus 0.12 and escalation recall of 1.0 versus 0.5. An answer-quality adapter changes recall from 1.0 to 0.5 in Run A, but from 0.5 to 1.0 in Run B. An escalation-aware variant produces no measurable change in Run B. These small pilots do not establish reliable adapter effects or production readiness. Their contribution is diagnostic: configuration variance can overwhelm apparent tuning effects on bounded-autonomy metrics, motivating larger evaluation sets and repeated runs.
| Comments: | 9 pages; ancillary evaluation artefacts. Previously submitted to NLLP 2026 |
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| Cite as: | arXiv:2609.37501 [cs.CL] |
| (or arXiv:2609.37501v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2609.37501 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Dipankar Sarkar [view email]
[v1]
Mon, 28 Sep 2026 11:08:38 UTC (37 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org