DisclaimerUnofficial, and not affiliated with Anthropic. Nearly all of this is read straight out of what ships: npm bundles, captured prompts, published docs. Anthropic's own notes go in verbatim, marked as theirs. The rest is my reading, and every entry carries the strings behind it. If one looks wrong, vote it down and say why.
Classifier refusals are tagged separately in blocking telemetry
Under the hood
Useful1Signal0
Telemetry
Safety-classifier refusals are now logged distinctly from ordinary parse failures.
What's wrong with this entry?
Anonymous. No account, no email.
What
The two-stage prompt-injection and safeguard classifier now records when the model stopped with a refusal, which previously looked identical to an ordinary parse failure.
Details
The parse-failure event now carries stopReason plus failureKind: "policy_refusal".
The block result itself carries refusedBySafeguard.
Evidence
refusedBySafeguard
Strings lifted out of the shipped bundle, so the claim above can be checked against them.