跳到正文
原文
PromptArmor:Threat Intelligence· PromptArmor:Threat Intelligence·· 2026-06-10精选AI 评分68

PromptArmor 分析 Claude Dynamic Workflows 的提示注入与权限风险

AI 导读

PromptArmor 分析 Anthropic 于 2026 年 6 月 8 日起对所有租户默认启用的 Claude Dynamic Workflows 的安全风险:单次运行可动态生成最多 1000 个子代理,auto 模式下由分类器代理审批敏感命令,可能被提示注入操纵;首次同意后任何工作流自动执行,Ultracode 开启时则完全跳过确认。

推荐理由

原文梳理了 Claude Dynamic Workflows 默认启用后的审批链路与子代理权限实际行为,并指出文档与实现不符的具体风险点。

正文

Anthropic recommends users operate in auto mode to leverage workflows. This allows an agentic system to determine what commands are safe to run and when to involve a human in the loop. Dynamic workflows release notes state:

"For the best experience, turn on auto mode when using dynamic workflows."

In app, if the user is not in auto mode, and a workflow requires a human approval, the permissions pop-up encourages switching to auto mode:

Permissions modal encourages using auto-mode.
Automatic Command Execution within Workflows

In Auto mode, sensitive actions taken by an agent are evaluated by a second classifier agent, which assesses whether the command is safe. This includes operations such as running MCP tools, making network requests, and editing files outside the active project.

However, this poses a risk as classifier agents can be manipulated by prompt injections to approve malicious commands. This risk is explored in another of our articles, which demonstrates Codex's agent-based approval mode installing malware after ingesting a hidden comment in a GitHub issue.

Kicking off Workflows Without Human Approval

When Claude is in auto mode and attempts to run a workflow, by default, the user is prompted for consent prior to the first workflow they ever run, but after that, any workflow in any project or session executes automatically.

If 'Ultracode' is enabled, workflows execute immediately irrespective whether workflows have been approved before.

Ultracode is a mode that allows Claude to determine when workflows are warranted, and sets the model's reasoning level to 'xhigh'.

Documentation describes the one-time consent:

"First launch only. Any Yes records consent in your user settings, and later launches start without prompting. Skipped entirely when ultracode is on."

Because these workflows execute without a permissions gate, this opens a new avenue for indirect prompt-injection attacks: manipulating Claude to write malicious workflows and run them. Subagents in the malicious workflow, and the classifiers evaluating their command requests are unlikely to have context on the original prompt injection that created the workflow. This increases the likelihood that malicious commands are executed relative to an injection that attempts to manipulate the classifier in the main Claude session.

来源:PromptArmor:Threat Intelligence · promptarmor.com