You'll notice
Blocked-content refusal messages can now name the specific policy category that triggered them
What
When Claude Code's fast safety classifier blocks some content, it now tries to pull a <category> tag out of the classifier's raw response. If that category matches one of a known set, the refusal message now shows a bracketed category name (like [category name]) instead of always showing the same generic fallback message.
Why
This gives a more specific explanation of why content was blocked, when the classifier's response identifies a known category, instead of a one-size-fits-all message.