develop-tests changedtest-and-evaluate/develop-tests
Nearest release: v2.1.287, published 4 hours before this site recorded the change. Shown because the two are within 24 hours of each other. Nothing here says the release caused the edit.
Recorded here
Lines+8added
Lines−8removed
From line
2,845
where the diff opens
First seen
14 Aug 2026
this site's first read of the page
Recorded edits4to this page, all time
The whole hunk
from line 2845, old and new numbered
/
from line 2845
28452845 A given use case, or even a specific success criteria for that use case, might require several rubrics for holistic evaluation.
28462846 </Note>
28472847* **Empirical or specific:** For example, instruct the LLM to output only 'correct' or 'incorrect', or to judge from a scale of 1–5. Purely qualitative evaluations are hard to assess quickly and at scale.
2848* **Encourage reasoning:** Ask the LLM to reason first before producing an evaluation score, and then discard the reasoning. This increases evaluation performance, particularly for tasks requiring complex judgment.
2848* **Encourage reasoning:** Use a grader model with [thinking](https://platform.claude.com/docs/en/build-with-claude/thinking) on, so that it reasons before it produces an evaluation score. This increases evaluation performance, particularly for tasks requiring complex judgment.
28492849
28502850<Accordion title="Example: LLM-based grading">
28512851 <CodeGroup exclude="shell">
from line 2857
28572857 return f"""Grade this answer based on the rubric:
28582858 <rubric>{rubric}</rubric>
28592859 <answer>{answer}</answer>
2860 Think through your reasoning in <thinking> tags, then output 'correct' or 'incorrect' in <result> tags."""
2860 Output 'correct' or 'incorrect' in <result> tags."""
28612861
28622862
28632863 def grade_completion(output, golden_answer):
from line 2916
29162916 return `Grade this answer based on the rubric:
29172917 <rubric>${rubric}</rubric>
29182918 <answer>${answer}</answer>
2919 Think through your reasoning in <thinking> tags, then output 'correct' or 'incorrect' in <result> tags.`;
2919 Output 'correct' or 'incorrect' in <result> tags.`;
29202920 }
29212921
29222922 async function gradeCompletion(output: string, goldenAnswer: string): Promise<string> {
from line 2972
29722972 Grade this answer based on the rubric:
29732973 <rubric>{rubric}</rubric>
29742974 <answer>{answer}</answer>
2975 Think through your reasoning in <thinking> tags, then output 'correct' or 'incorrect' in <result> tags.
2975 Output 'correct' or 'incorrect' in <result> tags.
29762976 """;
29772977 }
29782978
from line 3053
30533053 return fmt.Sprintf(`Grade this answer based on the rubric:
30543054 <rubric>%s</rubric>
30553055 <answer>%s</answer>
3056 Think through your reasoning in <thinking> tags, then output 'correct' or 'incorrect' in <result> tags.`, rubric, answer)
3056 Output 'correct' or 'incorrect' in <result> tags.`, rubric, answer)
30573057 }
30583058
30593059 func gradeCompletion(output, goldenAnswer string) string {
from line 3134
31343134 Grade this answer based on the rubric:
31353135 <rubric>%s</rubric>
31363136 <answer>%s</answer>
3137 Think through your reasoning in <thinking> tags, then output 'correct' or 'incorrect' in <result> tags.""".formatted(rubric, answer);
3137 Output 'correct' or 'incorrect' in <result> tags.""".formatted(rubric, answer);
31383138 }
31393139
31403140 String gradeCompletion(String output, String goldenAnswer) {
from line 3177
31773177 Grade this answer based on the rubric:
31783178 <rubric>{$rubric}</rubric>
31793179 <answer>{$answer}</answer>
3180 Think through your reasoning in <thinking> tags, then output 'correct' or 'incorrect' in <result> tags.
3180 Output 'correct' or 'incorrect' in <result> tags.
31813181 PROMPT;
31823182 }
31833183
from line 3258
32583258 Grade this answer based on the rubric:
32593259 <rubric>#{rubric}</rubric>
32603260 <answer>#{answer}</answer>
3261 Think through your reasoning in <thinking> tags, then output 'correct' or 'incorrect' in <result> tags.
3261 Output 'correct' or 'incorrect' in <result> tags.
32623262 PROMPT
32633263 end
32643264
No line in this hunk matches that.