Safeguard refusals are now a one-liner, with longer explanations only for category-flagged cases.
What's wrong with this entry?
The message shown when a model's safeguards flag your message is shorter and no longer references a model generation, and the longer explanation is now reserved for category-flagged refusals.
- the old paragraph beginning "The safeguards are intentionally broad right now and may flag safe and routine coding, cybersecurity, or biology work" is gone, along with the Mythos-capabilities sentence
- the assembled string
They may flag safe, normal content as well.was removed - the generic case is now a one-liner; the longer explanation appears only for category-flagged refusals
- a cyber-category refusal, when the relevant check passes, gets its own message pointing at the Cyber Verification Program plus a dedicated support link
- the refusal body reads "can't respond to this message with" rather than "this request with"
This sometimes happens with safe, normal conversations., Our intentionally broad safeguards allow us to deliver more capabilities faster, but can sometimes flag legitimate coding, cybersecurity, and biology tasks.
Strings lifted out of the shipped bundle, so the claim above can be checked against them.
Related
Other releases about the same thing. Found by shared names or similar wording; neither means one caused the other.