Follow Discord
Sweep 09 Oct 2026 · 17:27Z Build v2.1.296 517 read Stable v2.1.287 Latest v2.1.296 Next v2.1.296 Feeds RSS JSON llms.txt llms-full.txt Unofficial
One change · api

mitigate-jailbreaks changedtest-and-evaluate/strengthen-guardrails/mitigate-jailbreaks

Nearest release: v2.1.293, published an hour before this site recorded the change. Shown because the two are within 24 hours of each other. Nothing here says the release caused the edit.

Recorded here
Lines+3added
Lines−3removed
From line 15 where the diff opens
First seen 14 Aug 2026 this site's first read of the page
Recorded edits3to this page, all time

The whole hunk

from line 15, old and new numbered
/
lines
from line 15
1515 
1616In this threat model, a user is deliberately crafting inputs to manipulate your application into producing content or taking actions you don't want it to. These mitigations strengthen your application's guardrails:
1717 
18* **Harmlessness screens:** Use a lightweight model like Claude Haiku 4.5 to pre-screen user input before it reaches your main conversation. Use [structured outputs](https://platform.claude.com/docs/en/build-with-claude/structured-outputs) to constrain the response to a simple classification.
18* **Harmlessness screens:** Use a lightweight model like Claude Haiku 5.5 to pre-screen user input before it reaches your main conversation. Use [structured outputs](https://platform.claude.com/docs/en/build-with-claude/structured-outputs) to constrain the response to a simple classification. Claude Haiku 5.5 runs [safety classifiers](https://platform.claude.com/docs/en/build-with-claude/refusals-and-fallback) that can decline the screening request itself, so treat a response with `stop_reason: "refusal"` as a harmful verdict.
1919 
2020 <Accordion title="Example: Harmlessness screen for content moderation">
2121 ```text User wrap
from line 115
115115 
116116* **Limit Claude's access to sensitive data and actions.** Apply the principle of least privilege so that a successful injection can do minimal damage: don't give Claude access to secrets it doesn't need, run tools in sandboxed environments, and scope permissions as narrowly as possible.
117117 
118* **Screen tool outputs before Claude acts on them.** Apply the same lightweight-model screening pattern you use for user input to the content your tools return. Run each tool, pass its raw output to a small classifier call with Claude Haiku 4.5, and only return the content as a `tool_result` block if the screen reports no injection attempt. Use [structured outputs](https://platform.claude.com/docs/en/build-with-claude/structured-outputs) so the classifier's verdict is a parseable value your application can branch on.
118* **Screen tool outputs before Claude acts on them.** Apply the same lightweight-model screening pattern you use for user input to the content your tools return. Run each tool, pass its raw output to a small classifier call with Claude Haiku 5.5, and only return the content as a `tool_result` block if the screen reports no injection attempt. Use [structured outputs](https://platform.claude.com/docs/en/build-with-claude/structured-outputs) so the classifier's verdict is a parseable value your application can branch on.
119119 
120120 <Accordion title="Example: Injection screen for tool output">
121121 ```text User wrap
from line 147
147147 }
148148 ```
149149 
150 If `injection_suspected` is `true`, return an error or a stripped summary in the `tool_result` block instead of the raw content, and consider surfacing the attempt to the user.
150 If `injection_suspected` is `true`, return an error or a stripped summary in the `tool_result` block instead of the raw content, and consider surfacing the attempt to the user. Treat a `stop_reason: "refusal"` response from Claude Haiku 5.5 the same way: its safety classifiers can decline the screening request itself, which leaves no verdict to read.
151151 </Accordion>
152152 
153153 You can also apply the input-validation patterns from the previous section to tool results before passing them to Claude.
Feedback