A subagent's final text is now treated as untrusted and screened for prompt injection.
What's wrong with this entry?
When a subagent finishes, the final text it hands back to the parent agent is now itself reviewed by the auto-mode safeguard, treated as untrusted agent-authored output that could relay a prompt injection. Previously only the subagent's tool calls were reviewed.
- The hand-back text is wrapped in
<subagent_hand_back>tags and submitted as the action to evaluate. - The review now also runs when the subagent made no reviewable tool calls but did produce hand-back text.
- The allowed, blocked, refused and unavailable outcomes were reworked so a policy refusal is no longer reported as unavailable.
<subagent_hand_back>
Strings lifted out of the shipped bundle, so the claim above can be checked against them.
Related
Other releases about the same thing. Found by shared names or similar wording; neither means one caused the other.
-
v2.1.234
Another forged control tag is escaped in subagent output
Both mention subagent
-
v2.1.234
Spawned processes get
--flag=valuewhen the value looks like a flagBoth mention subagent
-
v2.1.235
New error for delegating to a subagent without naming one
Both mention subagent