LLM-as-judge evals: writing the rubric, and the failure modes that lie to you
Deterministic graders cannot tell if an answer is good. An LLM judge can, with the right rubric and a plan for its failure modes: leniency, drift, and bias.
Deterministic graders are the backbone of an eval suite: does the output contain the required disclaimer, does it match the expected shape, does it avoid a banned phrase. Instant, free, and exact. But they cannot answer the question that actually matters for a language feature — is this answer any good? "Politely declines without inventing a policy" has no regex. That is the job an LLM judge is for, and it is also where a judge will quietly tell you comforting lies if you build it carelessly.
Reach for a judge only past where strings stop
Most criteria do not need a model. Keep the cheap, deterministic graders first — they fail fast and cost nothing — and bring in a judge only for the fuzzy ones a string cannot capture: tone, faithfulness to a source, "answers the question actually asked," "refuses without being preachy." Each judge call is a real model request, so every criterion you can express as a string is one you should not pay a model to check.
Write the rubric as a falsifiable statement
The judge is only as good as the rubric, and the rubric must be something that can be false, not a vibe. "Is this good?" gets you a coin flip dressed as a verdict. "Declines the refund and offers to escalate, without inventing an exception that is not in the policy" can actually be checked. The grader hands that rubric and the output to the model and gets a typed verdict back — structured output, not prose you parse:
export const llmJudge =
(rubric: string, model: ModelId = "claude-opus-4-8"): Grader =>
async (output) => {
const { data } = await generateObject({
model,
schema: verdictSchema, // { pass: boolean, reasoning: string }
system:
"You are grading the output of an AI product feature against a rubric. " +
"Judge strictly: pass only if the output clearly satisfies the rubric.",
messages: [{ role: "user", content: `<rubric>\n${rubric}\n</rubric>\n\n<output>\n${output}\n</output>` }],
});
return { pass: data.pass, detail: data.reasoning };
};
Two details carry their weight. The verdict is a typed object (pass plus a one-line reason) via the structured-output helper, so a grader stays a plain function returning pass and a reason. And the system prompt forces a strict bar — which exists to fight the first failure mode.
The failure modes that lie to you
- Leniency and sycophancy. Left to its own temperament a judge wants to be agreeable, so it passes marginal answers and praises them. The "judge strictly, pass only if it clearly satisfies" instruction is not politeness — it is the counterweight. Calibrate it against a few outputs you know are bad; a judge that passes those is miscalibrated, not kind.
- Drift. The same rubric and output can get different verdicts across runs. A sharp, falsifiable rubric cuts the variance; pinning the judge model keeps a silent model upgrade from quietly re-grading your whole history. The judge is itself behavior you ship, so treat a rubric change like a prompt change — version it and re-run.
- Position and verbosity bias. In pairwise "which is better, A or B" grading, judges favor the first option and the longer one. The cleanest defense is to avoid pairwise entirely: grade one output against an absolute rubric, as above, so there is no position to be biased by.
- Self-preference. A model tends to rate text from its own family higher. If the stakes are high, judge with a different model than the one under test.
Things that bite
- Do not let the judge see the answer key inside the prompt it is grading — that is leakage, and it inflates every score.
- A judge that always passes measures nothing. The value is in the fails; if you never see one, distrust the rubric before you trust the model.
- Reasoning helps, briefly. A one-line justification makes verdicts debuggable; a paragraph just burns tokens.
This is the deep end of the eval harness whose CI wiring I covered in catching prompt regressions in CI. Shipwright — the Next.js and Claude starter kit this blog documents — ships both graders, deterministic and judge, composing in one suite that fails the build on a regression. Try the prompts they grade in the live demo.