· evals  · 11 min read

LLM evaluation LLM-as-a-judge: Evaluation Plan

LLM evaluation LLM-as-a-judge: Evaluation Plan. Comprehensive guide updated for 2026.

LLM evaluation LLM-as-a-judge: Evaluation Plan. Comprehensive guide updated for 2026.

LLM-as-a-Judge: Evaluation Plan

Answer First

A defensible LLM-as-a-judge evaluation plan defines the metric formula, the dataset it runs against, where the pass/fail threshold came from, how offline results differ from what production will show, how the judge itself is calibrated against human labels, and at least one documented case where the judge got it wrong. A plan missing any of these six elements is a demo, not an evaluation system a team can rely on for a release decision.

Scope and Assumptions

This page delivers a template evaluation plan for an LLM-as-a-judge system scoring a RAG-based support assistant for faithfulness to retrieved source documents. The structure generalizes to other tasks (summarization accuracy, code correctness review, agent trajectory quality), but the concrete numbers, dataset composition, and calibration results below are specific to this faithfulness use case and should not be copied as universal benchmarks — no benchmark comparison claims are made here without the underlying data being shown.

Metric Definition and Calculation

Metric name: Faithfulness Score (FS).

Definition: The proportion of atomic factual claims in a generated response that are directly supported by the retrieved source documents provided to the model at generation time.

Calculation method: The judge model receives the response and the source documents, decomposes the response into atomic claims (single-fact statements), and labels each claim as supported, unsupported, or contradicted against the source. FS is computed as:

FS = (count of "supported" claims) / (total claims in response)

A response with zero claims (e.g., a clarifying question back to the user) is excluded from FS calculation and tracked as a separate category, because forcing a faithfulness score onto non-substantive responses produces meaningless denominators.

Why this formula and not a single holistic 1-5 score: a claim-level formula is auditable — a human reviewer can check each claim’s supported/unsupported label individually, whereas a single holistic score gives no way to verify what the judge actually evaluated. This is the difference between a metric a team can trust and a metric a team has to take on faith.

Dataset

Composition: 240 query-response-source triples drawn from three sources: 100 from a curated regression test set covering known historically problematic query types (ambiguous product names, multi-part questions, questions with no answer in the source), 100 randomly sampled from live production traffic over a 14-day window, and 40 adversarially constructed cases designed to have a plausible-sounding but source-contradicting correct-format answer, used specifically to test whether the judge catches subtle contradiction rather than only catching missing information.

Update cadence: The regression subset is append-only — every faithfulness failure found in production that gets fixed is added as a permanent regression case so the fix cannot silently regress. The random production sample is refreshed each evaluation cycle to track current traffic patterns. The adversarial subset is reviewed and expanded quarterly as new failure patterns are discovered.

Known gap: the dataset currently has no non-English query coverage. If the underlying system serves non-English traffic, FS as calculated here should not be assumed to generalize to that traffic without a parallel evaluation.

Threshold and Its Source

Release gate threshold: FS >= 0.92 on the fixed regression + adversarial subset (140 items), computed before any production rollout.

Where this number came from: it is not an industry benchmark — no such standardized benchmark exists for this specific claim-decomposition faithfulness metric. The threshold was derived by running FS against three known-good historical releases (each independently confirmed via manual audit to have no faithfulness incidents in the following 30 days) and three known-bad historical states (each with a confirmed faithfulness incident, defined as a documented case of the assistant stating something not in the source that caused a customer-facing error). The known-good releases scored 0.93-0.97; the known-bad states scored 0.81-0.89. The 0.92 threshold sits above the highest known-bad score with margin, not at the midpoint, because the cost of a false negative (letting a bad release ship) is higher than the cost of a false positive (blocking a good release for extra review).

Limitation to state plainly: three known-good and three known-bad data points is a small basis for a threshold. This threshold should be treated as a working hypothesis, re-validated against every subsequent incident, and tightened or loosened as more incident data accumulates — not treated as a permanently fixed number.

Offline Versus Online Evaluation Differences

Offline evaluation (the regression + adversarial set, run pre-release) uses fixed, curated queries and cannot detect novel failure patterns that only appear in live traffic composition — new product launches, seasonal query shifts, or adversarial user behavior not anticipated in the curated set. It is fast, reproducible across releases, and appropriate for a release gate.

Online evaluation (the 2% continuous production sample) catches drift the offline set misses, but introduces evaluation lag — a faithfulness regression sampled at 2% of traffic with daily aggregation will typically take 24-72 hours to reach statistical confidence for a real drop, depending on traffic volume. Online monitoring is a detection system for gradual or unexpected drift, not a substitute for the pre-release gate, and should trigger a human investigation rather than an automated rollback given the lag and the risk of false alarms from a single noisy sampling window.

Both must run. A system that only has offline evaluation will miss production-only failure modes. A system that only has online monitoring will ship regressions before detecting them, because online monitoring is reactive by construction.

Judge Calibration

Method: 80 items from the evaluation dataset (a stratified sample across the three source categories) were independently claim-labeled by two trained human reviewers using the same supported/unsupported/contradicted schema the judge uses. Reviewer disagreements were resolved by a third reviewer to produce a single human gold label per claim.

Result: judge-human agreement on claim-level labels was 87% raw agreement, Cohen’s kappa 0.79, which falls in the “substantial agreement” range on the standard kappa interpretation scale. Disagreements concentrated in claims requiring numeric precision matching (the judge occasionally marked a claim “supported” when the source stated a similar but not exact number) — this is now a documented known weakness, not an assumed strength.

Recalibration trigger: this human calibration check is re-run whenever the judge model version changes, whenever the rubric or claim-decomposition prompt changes, and at minimum once per quarter regardless of whether either has changed, because judge behavior has been observed to shift even on provider-side model updates that don’t change the model’s public version number.

Counterexample: A Case Where the Judge Was Wrong

One adversarial-set case involved a source document stating a subscription “cannot be paused, only cancelled,” and a response stating “your subscription can be paused for up to 3 months.” The judge initially labeled this claim supported, reasoning (per its justification field) that pausing was “a reasonable interpretation of flexible cancellation options” — an inference not present in the source text. Human review caught this as a contradicted claim during calibration. The fix was tightening the claim-decomposition prompt to explicitly instruct the judge to reject inferential reasoning and require literal textual support, re-running the full calibration set, and confirming agreement improved on this specific failure pattern before considering the fix complete. This case is retained permanently in the adversarial dataset to prevent regression.

Roles and Review Cadence

Who owns what: the ML/AI engineer implementing the judge prompt owns the metric definition and dataset curation. A second engineer or a QA-adjacent reviewer owns the independent human-calibration labeling — this must not be the same person who wrote the judge prompt, because a single person both writing the rubric and labeling the calibration set will unconsciously anchor their labels to match what they expect the judge to say, inflating the reported agreement number. The release manager or team lead owns the threshold decision itself — treating the 0.92 cutoff as an engineering-only decision without a stakeholder informed about the cost of a false negative (a bad release shipping) risks setting a threshold optimized for pipeline convenience rather than actual business risk tolerance.

Review cadence: the full plan (dataset, threshold, calibration) gets a full re-validation quarterly, matching the freshness cadence on this page. Individual pieces trigger an out-of-cycle review sooner: any customer-facing faithfulness incident triggers immediate addition of that case to the regression set and a re-run of the full gate before the next release; any judge model version change triggers a full recalibration before the next release is allowed to use the new judge; any product change that alters what “the source” means (e.g., adding a new document type to the RAG retrieval corpus) triggers a dataset composition review, since the existing 240-item set may no longer represent the current source distribution.

Reporting Format

Every evaluation run against this plan produces a report with four required sections, no exceptions: the headline FS number against the fixed regression set, the pass/fail against the 0.92 threshold, a breakdown of failing items with their specific unsupported/contradicted claim and the judge’s justification text, and a diff against the previous evaluation run’s failing items (new failures versus persistent known failures versus previously-failing items that now pass). The breakdown and diff sections are what make the report actionable — a single headline number tells a release manager whether to ship, but the breakdown is what an engineer needs to actually fix a regression, and without it a failing gate produces a blocked release with no clear next action.

Failure Example From an Earlier Version of This Plan

An earlier version of this evaluation plan used a single holistic 1-5 faithfulness score rather than claim-level decomposition. That version was retired after a post-incident review found the holistic score had passed a release (score 4/5) that a manual audit later found contained two significant unsupported claims buried inside an otherwise well-written response — the judge’s holistic score correctly captured that most of the response was good but gave no visibility into the specific claims that were not. This is the direct justification for the claim-level FS formula used in the current plan: a holistic score can be simultaneously “mostly right” and “release-blocking wrong” in a way a single number cannot distinguish, while a claim-level breakdown surfaces exactly which claims failed and lets a reviewer judge severity rather than relying on the judge’s aggregation to have weighted severity correctly on its own.

Extending This Plan to a New Task Type

This plan is written for RAG faithfulness specifically, but the six-part structure transfers to other judgment-based evaluation needs with different specifics in each section. For a code-generation correctness eval, the metric would be pass rate against a held-out test suite (not an LLM judge at all, since code correctness has a deterministic ground truth once tests exist) with an LLM judge reserved for the secondary dimension of code readability or style adherence, where no deterministic check applies. For an agent-trajectory quality eval, the metric would score the sequence of tool calls against an expected minimal path (did the agent reach a valid answer in a reasonable number of steps without unauthorized actions) rather than scoring a single text response, and the dataset would need full conversation traces rather than single query-response-source triples. In every case, the discipline that transfers is the same: define the metric formula precisely, show where the threshold number came from, separate offline from online evaluation, calibrate the judge against humans, and keep at least one documented counterexample showing the judge failed and how that failure was addressed — a plan missing the counterexample section reads as untested confidence, not evidence-based confidence.

Book CTA

The 0→1 AI Engineer Interview Playbook walks through building an evaluation plan with this exact six-part structure — metric, dataset, threshold provenance, offline/online split, calibration, and counterexamples — across multiple production AI system types.

Get the book: /go/B0H2CML9XD?source=ai-engineers-blog&page=aie-llmjudge-evalplan-001

For the statistical foundations behind threshold-setting and calibration sample sizing referenced in this plan, The 0→1 Machine Learning Engineer Interview Playbook covers the underlying metric theory: /go/B0H256Z1MF?source=ai-engineers-blog&page=aie-llmjudge-evalplan-001

Sources and Freshness

Threshold and calibration figures in this plan are illustrative worked results from a template evaluation, not published benchmark data — treat them as a structural example to replicate with your own dataset, not as a number to cite. Rebuild the calibration step with your own human-labeled sample before using any threshold operationally. Next review due: quarterly, or immediately after any judge model or rubric change.

If you’re actively preparing for this process, the 0→1 AI Engineer Playbook covers the judgment frameworks, real question patterns, and structured answers this article draws on — useful when you want a complete preparation system rather than scattered tips.

    Share:
    Back to Blog

    Related Posts

    View All Posts »