When a safety check on subagent handoff output is refused, the work passes through with a warning.
What's wrong with this entry?
A sub-agent handoff whose safety classifier request is itself refused by the safeguard is no longer blocked; the output is passed through with a security warning attached.
- The refusal is not treated as a verdict. The sub-agent output is allowed but prefixed with a long warning telling the model the work is unreviewed and untrusted.
- Classifier outcome logging gained a "refused" bucket alongside the existing "unavailable" and "blocked".
Handoff classifier request refused by the safety safeguard, allowing sub-agent output with an unreviewed warning
Strings lifted out of the shipped bundle, so the claim above can be checked against them.
Related
Other releases about the same thing. Found by shared names or similar wording; neither means one caused the other.
-
v2.1.234
Another forged control tag is escaped in subagent output
Both mention subagent
-
v2.1.234
Spawned processes get
--flag=valuewhen the value looks like a flagBoth mention subagent
-
v2.1.235
New error for delegating to a subagent without naming one
Both mention subagent