
Book 20 of 50 · Free
Evaluating and Testing LLM Applications
28,641 words · 85 chapters · illustrated

Book 20 of 50 · Free
28,641 words · 85 chapters · illustrated
Book 20 of 50 — AstolixGen Learning Series

Large language models changed what it means to "test" software. A traditional program either passes or fails a test: you give it an input, you compare the output to an expected value, and you get a yes or no. An LLM application gives you a fluent, confident answer that might be right, partially right, beautifully wrong, or subtly dangerous — and asking it the same question twice can produce two different answers. How do you measure that? How do you know whether your system got better or worse after you changed the prompt? How do you report results in a paper that reviewers will trust?
This book answers those questions for students and early-career researchers working in AI. It is practical first: every concept comes with a concrete example, a piece of code you can run, or a template you can copy into your own project. But it is also rigorous enough for publication. By the end, you will be able to design an evaluation plan for an LLM application, execute it with defensible statistics, avoid the mistakes that sink papers at review time, and write up your results so another researcher could reproduce them.
You do not need deep statistics to read this book. You need basic familiarity with language models — what a prompt is, what temperature does, what a benchmark is — and a willingness to think carefully about measurement. Wherever a formula appears, a plain-English explanation comes first.
You don't have to read straight through — though you can. Three recommended paths:
Each chapter ends with a "For your research" box (the actionable point) and key takeaways (the summary). The Learning Dashboard near the end collects every table and checklist in one place for quick reference during a project. The exercises are designed to be done, not just read — Exercise 10, the full eval plan, is the capstone: if you can fill in that template for your own project, you have absorbed the book.
Consider the oldest trick in software testing: the assertion.
assert add(2, 2) == 4
This works because add(2, 2) always returns 4. The function is deterministic: the same input always produces the same output. Most of software engineering — unit tests, integration tests, continuous integration pipelines — rests on this property.
Now consider an LLM application, say a customer-support assistant. You write the analog of a test:
response = assistant.ask("How do I reset my password?")
assert response == "Click Settings, then Security, then Reset Password."
This "test" fails in ways that tell you almost nothing:
The deterministic test collapses because LLM outputs live in an enormous space of possible strings, most of them acceptable, some of them wrong in subtle ways. Evaluating them is not a string comparison problem; it is a judgment problem. And judgment is expensive, noisy, and hard to reproduce.
This chapter explains exactly where the hardness comes from. Once you see the sources of difficulty clearly, the rest of the book — metrics, judges, rubrics, statistics — reads as a set of tools aimed at specific problems.
Language models generate text by sampling from a probability distribution over the next token. A temperature parameter controls how much randomness enters the sampling. At temperature 0 (more precisely, greedy decoding), the model picks the most likely token at each step, which is mostly deterministic — but even greedy decoding is not perfectly stable across runs on real systems. Floating-point differences between GPU kernels, batched inference, and quantized models can all change which token wins a near-tie. At temperature 0.7 or 1.0, two calls with the same prompt routinely produce different text.
Concretely, this means an eval score is not a number — it is a random variable. Run your 200-item eval set today and get 82%; run it tomorrow and get 79%. Which one is "the" score? Neither, really. Both are draws from a distribution, and your job as an evaluator is to estimate that distribution's center and spread. Chapter 8 is entirely about doing this properly.
It also means single-example debugging misleads you. Every LLM developer has had this experience: a prompt change "fixes" the one example you were staring at and breaks ten you were not watching. Without a fixed eval set and repeated runs, you cannot tell whether a change helped. You are steering by noise.
A practical rule follows: never report a score from one run. Run each eval at least three times (more is better), with the same decoding settings you will ship, and report the mean and spread. Record the seed if the API supports one — but remember that a seed only controls your sampling; it does not make a hosted model reproducible across versions (Chapter 7 has more on this).
A classifier chooses among, say, 1,000 labels. An LLM answering an open question chooses among effectively infinite strings. This breaks the two conveniences that make classification evaluation easy:
The response is to decompose quality into dimensions and measure each: factuality, completeness, relevance, fluency, safety. Rubrics (Chapter 3) and structured judge prompts (Chapter 4) do this decomposition. A single aggregate score is fine for a leaderboard; for research and engineering, you want the per-dimension breakdown because it tells you what to fix.
Small prompt changes cause large behavior changes. Adding "Think step by step," reordering examples, or even changing whitespace can shift benchmark scores by double-digit percentages. This is not a bug you can patch; it is a property of how these models work — they are extremely context-sensitive pattern completers.
For evaluation this creates two traps:
Standard benchmarks are public. Public benchmarks end up in training data — accidentally, through web crawls, or deliberately. A model that has memorized MMLU questions will score brilliantly and tell you nothing about its ability to answer new questions. This is contamination, and it is one reason benchmark scores should be read with skepticism (Chapter 5).
Worse, hosted models change under you. The model behind an API endpoint can be updated, re-quantized, or have its safety filters tuned without notice — and your eval scores shift for reasons that have nothing to do with your system. This is why regression testing (Chapter 7) matters even when you change nothing: your baseline can drift.
At the bottom of every automated metric sits a human decision about what "good" means — and humans disagree. Two competent raters grading the same summary will agree perhaps 70–80% of the time on fine-grained rubrics. That disagreement is not noise to be averaged away; it is a ceiling on how meaningful your metric can be. If humans agree with each other at 75%, a metric that agrees with humans at 74% is essentially perfect, and squeezing it to 78% is chasing ghosts (Chapter 4 covers how to validate judges against this ceiling).
This has a deep consequence: there is no ground truth for open-ended generation, only intersubjective agreement. Your eval does not discover the "true quality" of outputs; it operationalizes a particular definition of quality, implemented by particular raters or judges, on a particular sample of inputs. Good papers say this out loud.
These five sources of hardness point to a coherent discipline, which is what this book teaches:
None of this is exotic. It is the ordinary discipline of empirical science — operational definitions, repeated measurement, reported uncertainty — applied to a new kind of artifact. Researchers coming from other empirical fields already know most of it; LLM work just forces you to apply it more carefully, because the artifact under study is unusually good at looking right while being wrong.
For your research: Before you run a single eval, write down in one paragraph: what decision will this eval inform? ("Should I ship prompt v3 or keep v2?" "Is my method better than the baseline on factuality?") Every choice in this book — metric, judge, sample size — flows from that decision. An eval without a decision is a number without a purpose, and numbers without purposes get misread.
To make the five sources of hardness concrete, follow a fictional-but-typical team — two graduate students, Amara and Ben — evaluating a support bot over one week.
Monday — nondeterminism. They run their 200-question eval set at temperature 0.7: 84%. Pleased, they re-run it Tuesday morning before a meeting: 81%. Nothing changed — same code, same data. Three questions flipped from right to wrong and back again across runs. Lesson: the score is a distribution. They switch to three runs per configuration and start reporting means with ranges.
Wednesday — open-endedness. Digging into failures, they find the bot answered "Where do I change my email?" with "Go to Profile → Account settings → Email, enter the new address, and confirm via the link we send." The gold answer was "Update it under Profile settings." Every word of the bot's answer is correct and more helpful — but their exact-match metric marked it wrong. They had been "improving" the prompt for a week against a metric that punished good answers. Lesson: they replace EM with a rubric scored by a validated judge (Chapters 3–4), and the bot's "true" score jumps 12 points without any model change. The metric had been lying, not the model.
Thursday — prompt sensitivity. Amara adds "Be concise." to the system prompt. The headline score holds steady — but on the 30 unanswerable questions, the bot stops saying "I don't know" and starts guessing. Conciseness pressure killed the abstention behavior. Nobody noticed for two days because the aggregate hid it. Lesson: per-slice reporting (Chapter 6) and behavioral assertions (Chapter 7) — "abstention rate ≥ 90%" — become permanent suite members.
Friday — contamination and drift. Comparing against a public benchmark, their bot scores suspiciously high on questions that look memorized. Worse, the provider announces a model update; the following week's baseline shifts 2 points with zero code changes. Lesson: they pin the model version, keep a private held-out set, and stop citing the public benchmark as evidence for their specific claim (Chapter 5).
The following month — judgment. They run a human eval of 150 answers. Two raters agree on 73% of faithfulness judgments. Their LLM judge agrees with the human majority 71% of the time — essentially at the human ceiling. They ship the judge for nightly regression runs and reserve humans for release gates. Lesson: the ceiling was real, the judge was valid for this rubric and task, and knowing the boundary let them spend human effort where it mattered.
Every chapter of this book is a response to one morning of that week. The hardness is not a reason to despair — it is a specification for the discipline. Teams that internalize it stop arguing about single numbers and start discussing distributions, slices, and decision thresholds. That shift, more than any single technique, is what "good at eval" means.

Every metric smuggles in a definition of "good." Accuracy says: getting the exact right answer is all that matters, and all errors are equal. Token-level F1 says: partial credit counts, and overlap with the reference is what we reward. BLEU says: looking like the reference translation, n-gram by n-gram, is quality. None of these is true in general; each is useful in a specific situation. Choosing a metric is therefore not a technical detail — it is the moment you decide what your research values. Reviewers will judge you on it, so this chapter gives you the vocabulary to choose defensibly.
We will walk through the standard metrics one by one: what each computes, when it fits, when it misleads, and how to compute it in a few lines of code.
What it is. The prediction counts as correct only if it is identical to the reference answer (after light normalization — lowercasing, stripping articles and punctuation is the usual convention from the SQuAD evaluation script). Accuracy is the fraction of correct items.
When it fits. Closed tasks with a single canonical answer: multiple-choice questions, yes/no questions, math problems with numeric answers, code generation judged by unit tests (pass/fail per problem). If there is exactly one right answer, exact match is the honest metric.
When it misleads. Any open-ended generation. "Paris," "Paris, France," and "The capital of France is Paris" are all correct answers to "What is the capital of France?" — EM gives credit to exactly one of them. Using EM on open tasks massively understates real performance and, worse, rewards models that learn the format of your references rather than the content.
import re, string
def normalize(text: str) -> str:
text = text.lower()
text = "".join(ch for ch in text if ch not in set(string.punctuation))
text = re.sub(r"\b(a|an|the)\b", " ", text)
return " ".join(text.split())
def exact_match(pred: str, gold: str) -> int:
return int(normalize(pred) == normalize(gold))
A practical tip: extract the answer span before comparing. If your model writes "The answer is 42 because…" and your gold is "42", compare the extracted "42", not the raw strings. Many published "EM" numbers are really "extract-then-EM" numbers — say so in your paper.
What it is. Treat the prediction and reference as bags of tokens. Precision = fraction of predicted tokens that appear in the reference; recall = fraction of reference tokens that appear in the prediction. F1 is their harmonic mean. This gives partial credit: "Paris France" against "Paris" scores well though not perfectly.
When it fits. Short-answer extraction tasks (SQuAD-style reading comprehension) where answers are spans of text and partial overlap is meaningful. It is the standard companion to EM on such benchmarks.
When it misleads. It rewards word overlap, not meaning. "Not guilty" vs. "guilty" share a token and get nonzero F1 while meaning the opposite. On long outputs, F1 degenerates — a rambling answer that happens to contain the reference words scores well.
from collections import Counter
def token_f1(pred: str, gold: str) -> float:
pred_tokens = normalize(pred).split()
gold_tokens = normalize(gold).split()
common = Counter(pred_tokens) & Counter(gold_tokens)
overlap = sum(common.values())
if overlap == 0:
return 0.0
precision = overlap / len(pred_tokens)
recall = overlap / len(gold_tokens)
return 2 * precision * recall / (precision + recall)
What it is. The classic machine-translation metric (Papineni et al., 2002). It measures n-gram precision between the candidate translation and one or more references (for n = 1..4), with a brevity penalty that punishes candidates shorter than the reference. Scores range 0–100 (or 0–1).
When it fits. Machine translation and other tasks with tight paraphrase constraints, where good outputs really do share wording with references, and where you have multiple references per input. In MT, BLEU still correlates reasonably with human judgment at the system level (ranking systems), though poorly at the sentence level.
When it misleads. Almost everywhere else. BLEU punishes legitimate paraphrase, ignores semantics (a fluent wrong translation can outscore a clumsy right one), and is meaningless on creative or open-ended tasks. A BLEU of 35 vs. 37 on summaries tells you essentially nothing about which system humans prefer. If you report BLEU in a paper on summarization or dialogue in 2026, reviewers will ask why.
Use the sacreBLEU implementation, not your own — BLEU has notorious signature variants (tokenization, smoothing) that make scores incomparable across implementations:
from sacrebleu import corpus_bleu
# references: list of reference lists; hypotheses: list of strings
score = corpus_bleu(hypotheses, references)
print(score.score) # 0-100
What it is. The summarization counterpart to BLEU (Lin, 2004). ROUGE-N measures n-gram recall (how much of the reference's content the summary captures); ROUGE-L measures longest common subsequence. Reported as F1 variants in most toolkits.
When it fits. Extractive or highly constrained summarization, as a cheap first-pass signal, and for comparability with older literature that reported it. ROUGE-L is the most commonly reported variant.
When it misleads. ROUGE rewards copying reference phrasing and is blind to factuality — a summary can score high on ROUGE while hallucinating, because hallucinations add n-grams the metric simply ignores (it measures overlap, not truth). It also cannot recognize a good abstractive summary that uses fresh wording. Never use ROUGE as your only summarization metric; pair it with a factuality check (Chapter 9 covers this for RAG, and the same logic applies).
from rouge_score import rouge_scorer
scorer = rouge_scorer.RougeScorer(["rouge1", "rouge2", "rougeL"], use_stemmer=True)
scores = scorer.score(reference_summary, model_summary)
print(scores["rougeL"].fmeasure)
What it is. Generate k candidate solutions per problem; the problem counts as solved if any candidate passes all unit tests (Chen et al., 2021, HumanEval). The unbiased estimator corrects for the fact that you actually sample n ≥ k candidates and compute the probability that at least one of k would pass.
When it fits. Code generation with executable tests — the gold standard setting, because tests check behavior, not text similarity. pass@1 measures the "one shot" a user typically gets; pass@10 or pass@100 measures the model's broader capability.
When it misleads. It depends entirely on test quality: weak tests let wrong code pass. And it says nothing about code quality, efficiency, or readability. Also, generating 100 samples per problem is expensive — budget accordingly.
import math
def pass_at_k(n: int, c: int, k: int) -> float:
"""n samples, c correct, probability >=1 correct in k draws."""
if n - c < k:
return 1.0
return 1.0 - math.comb(n - c, k) / math.comb(n, k)
Modern LLM applications need metrics that check behavior, not text:
The pattern is the same each time: find the observable consequence of a good output and measure that, not the words.
| Task shape | Start with | Add |
|---|---|---|
| Multiple choice / classification | Accuracy | Per-class breakdown, calibration |
| Short extractive answers | EM + token F1 | Human spot-check of failures |
| Machine translation | sacreBLEU / chrF | Human adequacy rating on a sample |
| Summarization | ROUGE-L (legacy comparability) | Factuality check, human rubric |
| Code with tests | pass@k | Test-strength audit |
| Open-ended QA / chat | Human rubric or validated LLM judge | Dimension breakdown |
| RAG | Citation precision/recall, answer EM/F1 | Retrieval hit rate |
| Agents | Task success rate | Step efficiency, tool-call accuracy |
Two rules of thumb. First, never report a single metric on an open task — report a small panel (2–4 metrics) so a weakness in one shows up in another. Second, validate the metric against humans on your data at least once: take 100 outputs, score them with your metric and with a human rubric, and compute the correlation. If they disagree, trust the humans and fix the metric, not the other way round. Chapter 4 shows how to do this validation for LLM judges; the same procedure applies to any automatic metric.
For your research: When you pick metrics for your paper, write one sentence per metric: "We use X because [property of the task it captures] and acknowledge it misses [limitation]." Reviewers rarely object to an imperfect metric; they object to an undefended one. This sentence belongs in your paper's evaluation section, nearly verbatim.
Watch what happens when four metrics score the same model output. Question: "What causes ocean tides?"
Reference: "Tides are caused by the gravitational pull of the Moon and the Sun on Earth's oceans."
Candidate: "The Moon's gravity pulls on ocean water, creating bulges that we experience as tides; the Sun adds a smaller effect."
A human reader calls this an excellent answer — accurate, complete, well-phrased. Now the metrics:
The lesson is not "judges good, strings bad" — it is that the metric must match the answer space. Where the answer space is tight (a number, a name, a code test), string and behavioral metrics are precise and cheap. Where it is open, they systematically under-measure quality, and optimizing against them teaches the model to mimic reference phrasing instead of answering well. This is Goodhart's law in eval clothing: when a flawed measure becomes the target, it stops being a good measure — your prompt iterations will chase ROUGE points while human-perceived quality stalls or drops.
Two more metrics worth knowing for the middle ground:
A practical scoring recipe for open QA that balances cost and honesty: (1) EM/F1 as a cheap first pass to catch clear wins; (2) a semantic metric to rank candidates during development; (3) a validated judge or human rubric on a sample for the reported numbers, focused on factuality. Each layer corrects the blindness of the one below it.
Every automatic metric is a proxy — a cheap, repeatable stand-in for what you actually care about: would a competent person judge this output good? Proxies drift. BLEU stops tracking human preference on new domains; LLM judges inherit biases (Chapter 4). The human evaluation is the measurement everything else is calibrated against. If your human eval is sloppy, your entire results section rests on sand, no matter how fancy the automation above it.
The good news: running a solid human eval is a learnable craft, not a dark art. It has four parts — a rubric, raters, blinding, and agreement measurement — and this chapter walks through each with templates you can reuse.
A rubric turns the vague question "is this good?" into specific, answerable questions. Bad rubric: "Rate the quality of this summary from 1 to 5." Different raters will interpret "quality" differently — one rewards fluency, another punishes a missing fact — and your agreement will be terrible.
Good rubrics share three properties:
Here is a worked example — a rubric for evaluating answers from a RAG question-answering system. Copy and adapt it:
Dimension 1 — Faithfulness (1–3). 3: Every factual claim in the answer is directly supported by the retrieved passages. 2: All central claims supported; minor details unsupported but plausible. 1: At least one central claim contradicted by or unsupported by the passages.
Dimension 2 — Completeness (1–3). 3: Answers every part of the question. 2: Answers the main part; a sub-question is missing or vague. 1: Misses the main point of the question.
Dimension 3 — Usefulness (1–3). 3: A user could act on this answer directly. 2: Useful but requires follow-up. 1: Not useful as written.
Edge rules. (a) If the answer says "I don't know" when the passages contain the answer, Completeness = 1. (b) Score what is written, not what you think the model meant. (c) Ignore fluency unless it blocks understanding.
Notice the scale is 1–3, not 1–5 or 1–10. Fewer points with clear anchors beat many points with vague ones: raters can reliably distinguish three levels, and your agreement statistics will thank you. Reserve wider scales for dimensions where raters genuinely need the granularity.
Pilot the rubric. Before the real study, have two raters score 20 outputs, then sit down (together or over notes) and discuss every disagreement. You will discover ambiguous anchors and missing edge rules. Revise the rubric, re-pilot on a fresh 20, and only then launch. Two pilot rounds is normal; skipping them is the most common cause of failed human evals.
Pay raters fairly and report the arrangement in your paper (reviewers increasingly ask). Report the number of raters, their background ("three CS graduate students"), and approximate effort ("each rated 150 items, ~8 hours").
Raters must not know which system produced the output they are scoring. If they can tell "this fluent one is GPT-4, this clumsy one is my baseline," their scores will reflect brand expectations, not quality. Practical steps:
If your raters do not agree with each other, your scores are noise. Report agreement for every human eval — reviewers expect it, and it is the single strongest signal that your study was run carefully.
from sklearn.metrics import cohen_kappa_score
# rater_a, rater_b: lists of integer labels, one per item
kappa = cohen_kappa_score(rater_a, rater_b)
print(f"Cohen's kappa: {kappa:.2f}")
What to do with disagreements. Decide the rule in advance and state it: majority vote (3 raters), adjudication by a fourth expert, or averaging (for continuous scales). Never silently drop disagreeing items — that biases your sample toward easy cases.
Human rating is expensive, so size the study to the claim. A useful heuristic: to detect a 10-percentage-point difference between two systems at 80% power, you need roughly 150–200 items per system (Chapter 8 explains the statistics). For a paper comparing two systems on one dimension, 200 items × 3 raters is a solid, defensible study. For early development, 50 items × 2 raters is enough to catch big problems — just do not publish strong claims off it.
Stratify your sample: if your eval set has categories (question types, difficulty levels, domains), sample proportionally from each so the human study reflects the full distribution, not just the easy majority.
For your research: Your human eval section should answer five questions a reviewer will ask: (1) What exactly did raters score? — point to the rubric, ideally in an appendix. (2) Who were the raters and how were they trained? (3) Were they blinded to the system? (4) How much did they agree? — report kappa/ICC. (5) How were disagreements resolved? Write your study plan as answers to these five questions before you collect a single rating, and the write-up becomes trivial.
A human eval is a small project. Here is a realistic five-week schedule for a 200-item, 3-rater study comparing two systems:
Week 1 — Rubric and pilot. Draft the rubric (dimensions, anchors, edge rules). Recruit two pilot raters — ideally the people who will do the real rating. They score 20 items; you compute preliminary agreement and hold a 1-hour disagreement review. Expect to rewrite at least a third of your anchors. Re-pilot on 20 fresh items. You are done when kappa clears 0.6 on every dimension.
Week 2 — Rater training and interface. Finalize the rating interface. The layout matters: show the question, the source material (if any), and the answer side by side; one dimension per screen or clearly separated sections; the full anchor text visible next to each scale point (raters should never have to recall what "2" means). Build in attention checks: 5–10% of items with obvious correct ratings (e.g., an answer that is clearly off-topic should score 1 on relevance). Raters who fail attention checks get their work reviewed or discarded. Hold the 30-minute calibration session with worked examples.
Weeks 3–4 — Rating with live monitoring. Do not wait until the end to check quality. After every 50 items, compute running agreement per rater pair and per dimension. Common problems and fixes: - One rater diverges from the other two: usually a misunderstood anchor — re-calibrate that rater with targeted examples, and consider re-rating their completed items on the affected dimension. - Agreement decays over time: fatigue. Cap sessions at ~90 minutes, enforce breaks, and randomize item order so fatigue hits all systems equally. - One dimension has low agreement everywhere: the dimension is ill-defined. Either rewrite its anchors mid-study (and re-rate) or drop the dimension and say so honestly.
Week 5 — Adjudication and analysis. Apply your pre-registered disagreement rule (majority vote, expert adjudication, or averaging). Then analyze: per-system means with CIs, per-dimension breakdowns, and a qualitative pass over the items where all raters agreed the best system still failed — those are your hardest cases and often the seed of your next research idea.
Budgeting. A trained rater scores roughly 15–25 complex items per hour. For 200 items × 3 raters at 20/hour, that is 30 rater-hours plus your ~10 hours of management. At fair graduate-student rates, budget accordingly — and report the cost in your paper's appendix. Reviewers increasingly view "we used 3 graduate raters for ~40 hours total" as a credibility signal, not filler.
One more safeguard: rater independence. Raters should work alone, without discussing items mid-study — discussion converges their judgments artificially and inflates agreement. The disagreement review happens in pilots; during the real study, questions go to you, and your answers go to all raters identically (and get added to the edge-rules document).
Even well-run studies hit situations the rubric didn't anticipate. Plan for these four:
Genuinely ambiguous items. Sometimes two reasonable experts permanently disagree — the question is vague, or two answers are defensible. Don't force false consensus: mark such items with a "disputed" flag during rating, adjudicate what you can, and report the disputed rate. A 5–8% disputed rate is normal on hard tasks; above 15%, your rubric or your items need work, not your raters. In the paper, one sentence covers it: "11 of 200 items (5.5%) were flagged as disputed after adjudication and are analyzed separately in Appendix C."
Rater gaming and speeding. On paid platforms especially, some raters click through as fast as possible. Defenses in layers: attention checks (Week 2), minimum time per item (calibrated from your pilot — e.g., flag items completed in under 20% of the median time), and agreement monitoring (a rater whose kappa with everyone else is near zero is either a genius or not reading). Remove and re-rate affected items; don't just average them in.
Sensitive or disturbing content. If outputs may contain toxic, graphic, or upsetting text, warn raters in advance, let them opt out of categories, limit exposure per session, and provide a way to skip-and-flag. This is both an ethical obligation and a data-quality measure — distressed raters rate badly. State your protocol in the paper; reviewers in 2026 expect it.
Multilingual rating. If your eval spans languages, don't assume one rater pool covers all: recruit per language, and validate that translated rubrics carry the same meaning (back-translation check: translate the rubric to the target language and back, and confirm the anchors still match). Report agreement per language — it often varies, and the variation itself is a finding about where your system is weakest.
The golden rule of study management: every problem you discover mid-study is a problem your pilot was supposed to catch. After each study, spend 30 minutes writing down what surprised you and add one question to your pilot checklist. Three studies in, your pilots will be catching 90% of issues — which is what makes the fourth study cheap.

Human evaluation is slow and expensive. The obvious shortcut: ask a strong language model to do the judging. Give it the rubric, the question, the two candidate answers, and ask which is better — or ask it to score each answer 1–5. This is LLM-as-judge, and it has become one of the most used and most abused tools in AI research. Used well, it lets you evaluate thousands of outputs overnight for dollars. Used badly, it manufactures the result you wanted and dresses it in the authority of automation.
This chapter teaches the disciplined version: how to write a judge prompt that works, how to measure whether your judge is any good, and the known biases that can silently corrupt your results.
A strong instruction-following model can apply a rubric with surprising consistency — often matching the agreement level of trained human raters on well-specified dimensions. The key phrase is "well-specified": LLM judges are excellent at applying clear criteria and terrible at inventing them. Your rubric from Chapter 3 does the hard conceptual work; the judge just executes it at scale.
The standard setups:
A judge prompt has five components. Here is a template for single-answer scoring — adapt the bracketed parts:
You are an impartial evaluator of [TASK, e.g., answers to factual questions].
## Question
{question}
## Reference answer (for guidance; the candidate need not match its wording)
{reference}
## Candidate answer
{candidate}
## Rubric
[Paste your Chapter-3 rubric here, with anchored scales.]
## Instructions
1. Quote the specific part of the candidate answer relevant to each dimension.
2. Score each dimension according to the rubric.
3. Return ONLY a JSON object: {"faithfulness": <1-3>, "completeness": <1-3>,
"usefulness": <1-3>, "rationale": "<one sentence per dimension>"}.
Four details matter more than people expect:
An LLM judge is a measurement instrument, and instruments must be calibrated. The procedure:
The decision rule: if judge–human agreement is close to human–human agreement, the judge is a valid substitute for this task and rubric. If it lags badly, fix the rubric or the prompt — or accept that this dimension needs humans.
Report this validation in your paper. One sentence suffices: "On a 150-item human-labeled subset, our judge agreed with the majority human label on 81% of pairwise preferences, vs. 84% human–human agreement." That sentence is the difference between a credible automated eval and hand-waving.
Validate per dimension. Judges are often good at fluency and bad at factuality — the dimension you care about most. A single aggregate agreement number can hide this. Validate each rubric dimension separately.
LLM judges have documented, repeatable biases. Design around them:
None of these biases is a reason to abandon LLM judges. They are reasons to validate, disclose, and design around them — exactly as you would with any imperfect instrument.
A pairwise judge call on long outputs can cost several cents; 10,000 comparisons adds up. Reduce cost by judging a stratified sample rather than everything, caching judge outputs (keyed by the exact prompt + candidate text hash), and using single-answer scoring (one call per answer) instead of pairwise (one call per pair) during development, switching to pairwise for the final numbers.
Version your judge. The judge model, its prompt, and its decoding settings are part of your method. If the provider updates the judge model mid-project, re-run the validation. A judge whose behavior changed silently is worse than no judge at all.
For your research: Treat "LLM judge" as a claim of the form "this automated procedure approximates human judgment on dimensions D1..Dn for this task." Every element of that claim needs evidence: the rubric (validity), the human-labeled validation set (calibration), the agreement numbers (reliability), and the bias checks (threats). When you can state all four in your paper, the judge stops being a shortcut and becomes a method.
A concrete validation, with numbers, so you can see what "good enough" looks like — and what failure looks like.
Setup. A team evaluates summary quality on three dimensions: faithfulness, coverage, fluency. They sample 150 summaries from their dev set, have two trained raters score each (Chapter 3 rubric), and run their LLM judge on the same items.
Results.
| Dimension | Human–human agreement | Judge–human agreement | Verdict |
|---|---|---|---|
| Fluency | 88% | 87% | Judge valid — use freely |
| Coverage | 81% | 79% | Judge valid — use, spot-check |
| Faithfulness | 82% | 74% | Judge not valid — needs humans |
The judge matches humans on fluency and coverage but lags 8 points on faithfulness — the dimension the paper's claim depends on. Error analysis shows the judge misses subtle contradictions (it checks whether claims sound supported, not whether they are). The team makes a disciplined call: use the judge for fluency and coverage at scale, and pay for human rating on faithfulness for the 200-item test set. Their results section reports all three numbers and the split decision. A reviewer reading this trusts the paper more, not less, because the team demonstrated they know where their instrument is blind.
Bias checks they ran on the same 150 items:
When not to use LLM judges at all. Judges are the wrong tool when (a) humans themselves disagree substantially on the task (the ceiling is too low — the judge is just adding confident noise); (b) the judgment requires expertise the judge lacks and you cannot verify (e.g., novel mathematical proofs); (c) stakes are high and errors are costly (medical or legal deployment decisions need humans); or (d) the setting is adversarial (a judge can be prompt-injected by the very outputs it scores — never let untrusted model output share a context window with judge instructions without strict delimiting and output validation).
The through-line: a judge is a delegated judgment, and delegation without verification is abdication. The validation study is what turns "we used an LLM to score outputs" from a red flag into a method.
A judge prompt is 20% of the work; the harness around it is the other 80%. Here's what production judge harnesses include:
Caching. Judge calls cost money and time. Cache every judge response keyed by a hash of (judge model version + judge prompt + item inputs). Re-running an eval should cost zero judge calls for unchanged items. A simple SQLite or JSONL cache works:
import hashlib, json, sqlite3
def judge_key(model, prompt, item):
h = hashlib.sha256()
h.update(json.dumps([model, prompt, item], sort_keys=True).encode())
return h.hexdigest()
Robust parsing. Judges don't always return clean JSON — they add preamble, truncate, or "helpfully" reformat. Parse defensively: extract the first {...} block, retry once with "return ONLY the JSON object" appended, and log unparseable outputs for inspection rather than silently dropping them. Track your parse-failure rate; above ~2%, fix the prompt.
Retries and rate limits. Wrap calls with exponential backoff, jitter, and a max-retry budget. Run judging jobs idempotently so a crashed run resumes instead of restarting — at 10,000 items, a non-resumable job that dies at 90% is a painful lesson.
Cost tracking. Log tokens in/out per call and maintain a running cost estimate. A pairwise judge on long documents can cost $0.02–0.10 per comparison; 20,000 comparisons is real money. The harness should print projected cost before launching a full run, and support a --sample N flag for stratified pilot runs.
Separation of concerns. Keep four files distinct: the judge prompt (versioned text), the runner (API calls, caching, parsing), the aggregation (means, CIs, bias checks), and the validation (comparison against human labels). When the provider updates the judge model, you change one string (the model version) and re-run validation — not a tangle of edits across a notebook. This separation is also what makes the judge setup describable in a paper: "judge prompt v3, runner v1.2, validated 2026-04-10" is a reproducible method; a notebook with 40 cells is not.
A benchmark is a fixed test set plus a scoring rule, shared across the field so that results are comparable. Benchmarks exist to answer "how capable is this model, in general?" — and they are genuinely useful for that: they compress thousands of judgments into one number, they let you compare against published results without re-running everyone's code, and they reveal broad strengths and weaknesses (a model great at code but poor at multilingual QA tells you something real).
They are also the most misread numbers in AI. This chapter teaches you to read them like a reviewer: what each major benchmark measures, what its score does and does not prove, and the traps — contamination, saturation, prompt-sensitivity — that make leaderboard numbers slippery.
MMLU (Massive Multitask Language Understanding) — Hendrycks et al., 2021. ~16,000 multiple-choice questions across 57 subjects, from high-school math to professional law and medicine. The headline test of "broad knowledge and reasoning." What it proves: the model has absorbed a vast amount of factual and conceptual knowledge and can apply it in a multiple-choice format. What it does not prove: that the model reasons reliably (multiple choice lets models exploit surface cues and elimination strategies), that it would perform similarly on open-ended versions of the same questions, or that its knowledge is current. Caveats: it is the most contamination-suspected benchmark in existence — its questions are all over the training web. Treat small MMLU differences between models (1–3 points) as noise, not signal.
HumanEval — Chen et al., 2021. 164 hand-written Python programming problems, each with a function signature, docstring, and hidden unit tests; scored with pass@k. What it proves: the model can write short, self-contained functions that pass tests — a real, behavioral capability. What it does not prove: ability to work in large codebases, use unfamiliar APIs, debug, or write maintainable code. 164 problems is a small sample; the confidence interval on a pass@1 difference is wide (Chapter 8). Caveats: widely memorized; many models have effectively seen the problems. Also, docstring quality varies — some "failures" are ambiguous specifications, not model errors.
GSM8K — ~8,500 grade-school math word problems with numeric answers. Tests multi-step arithmetic reasoning; chain-of-thought prompting famously lifted scores dramatically (Wei et al., 2022). What it proves: step-by-step quantitative reasoning on well-specified problems. Does not prove: real-world mathematical problem-solving, where the hard part is formalizing the problem. Contamination concerns are serious here too.
SQuAD / TriviaQA — reading-comprehension and factual-QA benchmarks (Rajpurkar et al., 2016; Joshi et al., 2017). What they prove: extracting and recalling factual answers. Caveat: largely saturated — top models score near human performance, so they no longer discriminate between strong models. Useful as sanity checks, not as differentiators.
GLUE / SuperGLUE — collections of natural-language-understanding tasks (Wang et al., 2019). Historically central; now saturated and mostly of historical interest for LLM evaluation, though the idea — a suite of diverse tasks with one aggregate — lives on in HELM.
HELM (Holistic Evaluation of Language Models) — Liang et al., 2022. Not one benchmark but a framework evaluating many models across many scenarios with standardized metrics, emphasizing transparency about what was measured. What it proves: breadth across scenarios with honest reporting. Valuable as a model for how to report evals, not just as a score source.
MT-Bench — 80 multi-turn questions across 8 categories, scored by LLM judges with pairwise comparisons (Zheng et al., 2023). What it proves: conversational and instruction-following quality as judged by a strong model. Does not prove: anything the judge is blind to (see Chapter 4's biases — MT-Bench scores inherit them).
When you see "Model X scores 86.4 on MMLU," run through this checklist:
lm-evaluation-harness, which standardizes prompting.For your research: Before citing a benchmark score as evidence, write the sentence "This score shows ___ about my system" and check that the blank is something the benchmark actually measures (per the descriptions above), not something you wish it measured. If you cannot fill the blank honestly, the score does not belong in your results — it belongs in an appendix, or nowhere.
Contamination deserves a deeper look because it is the quietest way a benchmark score becomes meaningless — and because "the model may have seen the test data" is now a default reviewer question.
How test data leaks into training. The mechanisms are mundane: (1) Benchmarks are published on the web; web-scale crawls ingest them. (2) Researchers discuss benchmark items in papers, blog posts, and GitHub issues — all crawled. (3) Few-shot examples from the test split get pasted into prompts that are logged and later trained on. (4) Fine-tuning datasets assembled from the web include benchmark items verbatim. None of this requires bad intent. A useful rule: if a benchmark was public before a model's training cutoff, assume partial contamination and discount memorization-friendly formats (multiple choice, short factual answers) accordingly.
Detection techniques. None is definitive, but together they build a picture: - Canary strings. Some benchmarks (e.g., BIG-bench) embed a unique "canary" GUID string with instructions not to train on it. Searching model outputs or training-data disclosures for the canary is a direct test — though it only works if the benchmark included one. - Ordering and perturbation tests. If a model scores dramatically higher on the original benchmark items than on paraphrased versions testing the same knowledge, memorization is likely: true understanding survives paraphrase, memorized strings do not. - Cutoff analysis. Compare performance on questions about events before vs. after the model's training cutoff. A model that "knows" 2021 benchmark questions far better than equally difficult 2024 questions it could not have memorized is showing its training data, not its reasoning. - Log-probability probes. For open-weight models, test items the model assigns suspiciously high likelihood to (relative to paraphrases) are candidates for memorized strings.
What to do about it. For high-stakes claims, don't rely on public benchmarks alone: keep a private holdout — items written after the model's cutoff, never published — and report those numbers alongside. Support dynamic benchmarks that regenerate items (new math problems, fresh questions) so memorization has a short half-life. And when you build your own eval set (Chapter 6), treat non-publication as a feature: an unpublished, private test set is contamination-proof by construction, which is a genuine methodological advantage worth stating in your paper.
Running benchmarks reproducibly. If you do report benchmark numbers, produce them yourself with a standard harness rather than copying them. The EleutherAI lm-evaluation-harness is the community standard:
pip install lm-eval
lm_eval --model hf --model_args pretrained=<model-id> \
--tasks mmlu,gsm8k,humaneval --num_fewshot 5 --batch_size 8 \
--output_path results/
Report the harness version, the exact task names, and the few-shot count. Two papers reporting "MMLU 5-shot" from different harness versions can differ by a point from implementation details alone — pinning the harness is what makes your number comparable to the next person's.
Imagine a model release blog post with this table:
| Benchmark | NewModel-X | Competitor-Y |
|---|---|---|
| MMLU (5-shot) | 89.2 | 87.9 |
| HumanEval (pass@1) | 92.1 | 88.4 |
| GSM8K | 95.3 | 94.1 |
Impressive? Run the Chapter 5 checklist on it:
The corrected version you'd want to see: harness name and version, exact prompting per benchmark, decontamination statement, CIs or at least sample sizes, and — most importantly — a custom eval for whatever capability the release actually claims ("our model follows complex instructions better" needs an instruction-following eval, not MMLU). When you write your own paper's related-work or baseline sections, hold others' tables to this standard — and preemptively hold your own to it.
A common confusion: Model A wins on MMLU but loses on HumanEval; which is "better"? The answer is that the question is malformed — the benchmarks measure different capabilities, and models have different strengths. A code-specialized model losing MMLU to a generalist while winning HumanEval is exactly what you'd expect, not a contradiction. When benchmarks disagree, don't average them into a meaningless composite; instead, ask which capability your task needs and weight that benchmark accordingly. And when similar benchmarks disagree (two QA benchmarks ranking models differently), suspect prompt-format differences or contamination asymmetry before concluding anything about the models — re-run both under your harness with identical prompting before theorizing.

Standard benchmarks measure general capability. Your research makes a specific claim — "my retrieval method improves factuality on medical QA," "my prompting technique reduces hallucinations in summarization" — and no public benchmark tests exactly that. Reviewers know this, which is why the strongest papers build custom eval sets tailored to their claim. A well-built custom eval set is also the foundation of regression testing (Chapter 7) and the validation data for your judges (Chapter 4).
Building one is real work: sourcing items, writing gold labels, deciding how many you need. This chapter is the field manual.
An eval item is not just "an input." It is a probe for a specific capability or failure mode. Before collecting anything, list the dimensions your claim touches and make sure your set covers each. Example: you claim your RAG system is more faithful than a baseline. Your dimensions might be:
Aim for a stratified set: each dimension gets its own slice, so you can report per-slice scores. An aggregate score over an unstratified set hides the interesting structure — your method might ace easy items and fail every adversarial one, and the average would look fine.
Write a one-page "eval spec" before collecting: the claim, the dimensions, the target counts per dimension, the input format, the gold-label format, and the metric per dimension. This document is also the first thing to show a reviewer who asks "how did you evaluate?"
Good sources, in rough order of preference:
What to avoid: scraping items from existing public benchmarks (contamination risk, and it adds nothing over citing those benchmarks), and writing all items yourself in one sitting (your blind spots become the set's blind spots — get at least two authors).
The gold label is the reference against which outputs are scored. Its form depends on your metric:
Labeling protocol matters as much as the labels: two independent labelers per item, adjudication of disagreements, and a measured label-agreement rate that you report. Your eval is only as trustworthy as its gold labels — a 5% label error rate puts a hard ceiling on the conclusions you can draw from small score differences.
"How many items?" is the question everyone asks, and the honest answer is: enough to detect the effect you claim, with room for per-slice analysis. Some concrete guidance:
The statistics behind these numbers are in Chapter 8. The intuition: your score is an average over items, and averages over few items are noisy. A 5-point "improvement" on 50 items is about 2–3 items flipping — that is noise, not a finding.
Split your set. Keep a development slice (for iterating on prompts and methods) and a held-out test slice (touched once, at the end, for the reported numbers). The moment you start prompt-tuning against the full set, it becomes a development set and your numbers become optimistic. This is the cheapest insurance against fooling yourself in all of experimental science.
An eval set is a research artifact. Treat it like code:
If you can release the set publicly, do — it is a genuine contribution, and other researchers citing your eval set is how benchmarks are born. If you cannot (privacy, proprietary data), describe the construction process in enough detail that someone could replicate it.
Suppose you are evaluating faithfulness of answers from a customer-support RAG bot over 40 help articles:
Total cost: a few days of agent time plus validation labeling. Total value: an eval you can defend in review and reuse for every future experiment on this system.
For your research: Your eval set is your operational definition of the claim. When a reviewer asks "does your method really improve faithfulness?", they are asking "does your eval set really measure faithfulness?" Spend your skepticism there first — a brilliant method on a sloppy eval convinces no one, while a simple method on a rigorous eval gets cited.
Standard items measure typical performance. Adversarial items measure the failures you most need to prevent — the ones that are rare in logs but catastrophic in production. A faithfulness eval without adversarial items is like crash-testing a car only on straight roads.
Five adversarial patterns, with examples (for a RAG support bot):
Labeling adversarial items. The gold label for an adversarial item is often a behavior ("must abstain," "must cite the 2026 policy") rather than a string. Write the expected behavior explicitly in the gold label, and have your two labelers independently confirm it — adversarial items are exactly where labeler disagreement spikes, because the "right" behavior can itself be debatable. Items your labelers can't agree on are either badly designed (fix them) or genuinely ambiguous (keep them, but score them separately as a "judgment-call" slice rather than mixing them into the main score).
Weighting honestly. Report adversarial slices separately from standard slices — never blended into a single number without disclosure. A system scoring 90% on standard items and 40% on adversarial ones is not a "65% system"; it is a system that works typically and fails dangerously. Both numbers, side by side, tell the truth. If your paper's claim is about robustness, the adversarial slice is your primary evidence; if the claim is about typical-case quality, the standard slice leads and the adversarial slice is the limitations discussion.
An eval set is a dataset, and datasets need schemas. Here's a JSONL schema that has survived several real projects — one object per line, with metadata that pays for itself during analysis:
{"item_id": "medrag_0042",
"slice": "distractor",
"difficulty": "hard",
"input": "Can I get a refund for my missed appointment?",
"context": ["<passage 1 text>", "<passage 2 text>"],
"gold": {"relevant_passages": ["p1"],
"key_points": ["Refunds allowed within 24h", "No-show fee applies after 24h"],
"expected_behavior": "answer"},
"source": "clinic_logs_2026-01",
"author": "labeler_A",
"adjudicated": true,
"version": "1.0"}
Why each field matters: slice enables per-slice reporting (Chapter 6) — without it, stratification exists only in your head. difficulty lets you check whether improvements concentrate on easy items (a common hollow victory). expected_behavior ("answer" vs. "abstain") makes behavioral assertions (Chapter 7) trivially computable. source/author/adjudicated/version are provenance: when someone asks "where did item 42 come from and who labeled it," you answer in seconds instead of excavating chat history.
Item QA checklist — run every item through this before it enters the set: - [ ] The input is self-contained (no "as mentioned above" without the above). - [ ] The gold label is complete (all acceptable answers listed, or behavior specified). - [ ] A second labeler independently produced a compatible label. - [ ] The item is assigned to exactly one slice. - [ ] No PII, no copyrighted text beyond fair use, license noted. - [ ] The item is not in the dev prompt examples or any training data.
Storing it. Keep the eval set in version control as JSONL (diffable, greppable) — not in a spreadsheet, not in a Google Doc. Large binary artifacts (passage corpora) go in object storage with a content hash recorded in the repo. Tag releases: eval-v1.0, eval-v1.1. Your future self, reproducing a paper's numbers two years later, will thank you more for this than for any clever metric.
You ship a support bot that scores 84% on your faithfulness eval. A month later it scores 79% and nobody changed anything — the model provider updated the weights behind the API. Or you change something: you "improve" the system prompt, the headline metric stays flat, but the abstention behavior on unanswerable questions quietly collapses, and users start getting confident fabrications.
Regression testing is the discipline of catching this. The idea is borrowed from software engineering — a suite of tests that must keep passing — but adapted to the nondeterministic, graded nature of LLM outputs (Chapter 1). Instead of asserting exact outputs, you assert statistical properties: the score on each eval slice must stay within a tolerance band of its baseline.
A good regression suite has three layers, from fast to slow:
Layer 1 — Smoke tests (seconds). 10–30 hand-picked items covering your most important behaviors, run on every prompt edit. These are your "canaries": if the new prompt breaks the basic greeting behavior or starts refusing normal questions, you know in seconds, not after a full eval. Score them with cheap checks — EM on closed items, a fast judge on open ones.
Layer 2 — Slice evals (minutes). Your full dev eval set (Chapter 6), run on every significant change: prompt edits, retrieval changes, decoding-parameter changes. Report per-slice scores vs. baseline with a tolerance (e.g., "no slice may drop more than 2 points; the aggregate may not drop at all").
Layer 3 — Full evals (hours). The held-out test set plus human spot-checks, run before releases and after model-provider updates. This is the gate for shipping.
The layers exist because full evals are too slow and expensive to run on every keystroke, and smoke tests are too thin to trust for releases. Match the layer to the decision.
Two techniques make regression tests concrete:
Golden outputs. For a small set of critical items, store a known-good output and compare new outputs against it — not with string equality (Chapter 1: that fails on paraphrase), but with a similarity check: an LLM judge scoring "does the new output preserve the meaning and key facts of the golden output?", or an embedding-similarity threshold. Goldens catch behavioral drift — the model still answers, but its answers changed character.
Behavioral assertions. Assert properties, not strings: "on the 40 unanswerable items, the abstention rate must be ≥ 90%"; "no output may contain a phone number" (regex); "every answer must cite at least one passage" (format check); "average response length must stay within 20% of baseline" (verbosity guard). These are cheap, deterministic, and catch entire classes of regression that aggregate scores miss. A prompt tweak that keeps accuracy flat while doubling average length — and your inference bill — is a regression your score-only suite would never see.
def check_abstention(outputs, must_abstain_ids, min_rate=0.9):
abstained = sum(1 for i, o in enumerate(outputs)
if i in must_abstain_ids and "don't know" in o.lower())
rate = abstained / len(must_abstain_ids)
assert rate >= min_rate, f"Abstention rate {rate:.2f} below {min_rate}"
return rate
Because scores are random variables (Chapter 1), a regression suite needs statistical thinking, not fixed thresholds:
gpt-4o-2026-05-01 rather than gpt-4o). When only an unversioned alias exists, log the model's self-reported version string with every run so you can correlate score shifts with provider updates.Sooner or later you will swap models — a new release, a cost-motivated move to a smaller model, a provider change. Treat it as a first-class experiment:
Document the swap as a one-page decision memo: scores before/after per slice, cost/latency deltas, failure categories, and the final call. Future you — and your reviewers — will be grateful.
Wire Layer 1 and 2 into your continuous integration so every commit touching prompts, retrieval code, or model config triggers them:
# .github/workflows/eval-regression.yml (sketch)
on:
pull_request:
paths: ["prompts/**", "src/retrieval/**", "eval/**"]
jobs:
smoke:
runs-on: ubuntu-latest
steps:
- run: python eval/run_suite.py --layer smoke --runs 3
slice-eval:
needs: smoke
runs-on: ubuntu-latest
steps:
- run: python eval/run_suite.py --layer slice --runs 3 --compare-to baseline.json
Store baseline.json (per-slice means and standard deviations) in version control next to the prompts. A PR that moves a slice outside its tolerance band fails CI — exactly like a broken unit test. The cultural effect matters as much as the technical one: prompt edits become reviewed, tested changes instead of vibes.
For your research: Your paper's "we changed X and the score went from A to B" claim is only credible if A was measured with the same rigor as B — same items, same settings, multiple runs. A regression suite maintained during the project gives you this for free: every experimental result in your paper is then a before/after comparison against a frozen baseline, which is precisely what reviewers want to see.
A mid-size startup's support bot, "prompt v4," shipped on a Friday. The headline faithfulness score went from 82% to 83% — a small win, celebrated briefly. Three weeks later, support tickets about "the bot making things up" tripled. What happened?
The v4 prompt added "be helpful and thorough." The aggregate score held because answers got longer and more detailed — the judge rewarded the extra content (length bias, Chapter 4). But buried in the per-slice numbers, which nobody checked before shipping, the unanswerable slice had collapsed: abstention rate fell from 92% to 64%. "Thorough" had quietly translated into "never say I don't know." The team had a slice eval that would have caught it — they just didn't gate the release on it.
How the three-layer suite would have caught it: - Smoke tests (seconds): one smoke item was "answer an unanswerable question" — it failed on the first run of v4. But smoke tests weren't wired into the prompt-editing workflow; they lived in a notebook someone ran monthly. - Slice evals (minutes): the abstention slice showed the 28-point drop. But the release checklist only required "aggregate score not decreased." - Full evals (hours): never run, because "it was just a prompt tweak."
The fix was procedural, not technical: prompt changes became pull requests, the smoke suite ran on every PR, and the release checklist required every slice to stay within its tolerance band — not just the aggregate. The team estimated the incident cost ~200 support hours; the suite cost an afternoon to build.
Non-functional regression tests. The same machinery catches degradations that aren't about quality at all: - Latency budget: p95 response time ≤ 4s on the smoke set. A prompt that doubles output length can silently double latency and cost. - Cost per 1k queries: track input+output tokens per eval run. "Helpful and thorough" increased cost 40% — a regression the quality score never saw. - Output length guard: mean response length within ±20% of baseline. Length drift is the canary for verbosity regressions, judge gaming, and cost blowups. - Format compliance: % of outputs parsing as the expected schema (JSON, citations format). A model update that changes formatting breaks downstream code without changing "quality" scores.
# Tolerance-band check from measured history (Chapter 7)
def check_regression(current, baseline_mean, baseline_std, k=2):
lo, hi = baseline_mean - k * baseline_std, baseline_mean + k * baseline_std
status = "PASS" if lo <= current <= hi else "FAIL"
return {"current": current, "band": (round(lo,2), round(hi,2)), "status": status}
Run the non-functional checks on the same schedule as the quality suite. A system that gets slightly better answers at 3× the cost and 2× the latency has not unambiguously improved — and your eval suite should say so out loud.
Prompts are code. Version them like code — because every result in your paper depends on the exact bytes of the prompt, and "the prompt we used" from memory is not reproducible.
Store prompts as files. One file per prompt, in prompts/, with a descriptive name and a header comment recording its purpose and history:
# prompts/rag_answer_v3.md
# Purpose: answer generation for MedRAG eval
# History: v1 (2026-03-01) baseline; v2 (2026-03-08) added citation requirement;
# v3 (2026-03-15) added "say I don't know" instruction after abstention failures
---
Answer the question using ONLY the passages below. Cite each claim...
Diff prompts in review. When a prompt change goes through pull-request review, the diff should be readable: keep prompts as plain text/markdown (not embedded in Python strings with escaping), one sentence per line if the team prefers fine-grained diffs. The PR description states the hypothesis ("adding the abstention instruction should raise the unanswerable-slice score without hurting completeness") and links the slice-eval results. This turns prompt engineering from alchemy into an auditable engineering practice.
The prompt review checklist — for the reviewer, not just the author: - [ ] Does the prompt state the task, the constraints, and the output format explicitly? - [ ] Are there instructions that conflict with each other? (Common: "be thorough" + "be concise" — pick one or define the tradeoff.) - [ ] Does it handle the failure modes? (What should the model do when the context lacks the answer?) - [ ] Were few-shot examples chosen from the dev set (never the test set), and do they cover edge cases? - [ ] Is the decoding config (temperature, max tokens) recorded alongside?
Never do this: keep the "real" prompt in a chat window, a playground, or someone's notes app while the repo holds a stale copy. The prompt in version control is the prompt — everything else is a rumor. When your paper says "the full prompt is in Appendix A," generate Appendix A from the version-controlled file, not by retyping.
"Our method scores 84.2% vs. the baseline's 81.7%" — a 2.5-point win. Publishable? It depends entirely on the uncertainty around those numbers. If each is the mean of 1,000 items with small run-to-run variance, the win is real. If each is a single run over 164 HumanEval problems, the win is about four problems — well within noise, and a reviewer who knows statistics will reject the claim.
This chapter gives you the minimum statistical toolkit for eval work: where variance comes from, how to estimate it, how to test whether a difference is real, and how to report all of it honestly. No proofs — just the procedures, the code, and the judgment calls.
An eval score varies for four reasons, and you should know which dominate in your setup:
The practical consequence: your reported score should always come with an uncertainty estimate that accounts for at least (1) and (2). A point estimate alone is an incomplete measurement.
The bootstrap is a beautifully simple way to estimate uncertainty: resample your eval items with replacement, recompute the score on each resample, and look at the spread of the resulting scores. It makes almost no assumptions and works for any metric — accuracy, F1, win rates, even judge scores.
import numpy as np
def bootstrap_ci(scores, n_boot=10000, ci=95):
"""scores: per-item scores (0/1 or continuous). Returns (mean, lo, hi)."""
scores = np.asarray(scores)
n = len(scores)
boot_means = [np.random.choice(scores, size=n, replace=True).mean()
for _ in range(n_boot)]
lo, hi = np.percentile(boot_means, [(100 - ci) / 2, 100 - (100 - ci) / 2])
return scores.mean(), lo, hi
mean, lo, hi = bootstrap_ci(item_scores)
print(f"{mean:.1%} [{lo:.1%}, {hi:.1%}]") # e.g. 84.2% [81.9%, 86.4%]
Report the 95% confidence interval (CI) alongside every score: "84.2% [81.9, 86.4]." If two systems' CIs barely overlap, the difference is suggestive; if they overlap substantially, you do not have evidence of a difference — more items or more runs are needed.
How many bootstrap resamples? 10,000 is cheap for simple metrics and gives stable intervals. For expensive metrics (judge-scored), 1,000–2,000 suffices.
You have system A scoring 84% and system B scoring 81% on the same 500 items. Is A better? The right test for paired data (same items, two systems) is a paired test:
import numpy as np
def paired_permutation_pval(a_scores, b_scores, n_perm=10000):
a, b = np.asarray(a_scores), np.asarray(b_scores)
observed = (a - b).mean()
diffs = a - b
count = sum(
(diffs * np.random.choice([-1, 1], size=len(diffs))).mean() >= observed
for _ in range(n_perm)
)
return (count + 1) / (n_perm + 1)
p = paired_permutation_pval(scores_A, scores_B)
print(f"p = {p:.4f}") # p < 0.05: difference unlikely to be chance
Reading p-values honestly. p < 0.05 means "a difference this large would occur by chance less than 5% of the time if the systems were truly equal." It does not mean the difference is large or important — with 100,000 items, a 0.1-point gap can be "significant." Always report the effect size (the actual gap with its CI) next to the p-value, and ask whether the gap matters in practice. A statistically significant 0.3-point gain that costs 10× more inference is not a win.
Multiple comparisons. If you test 20 slices/metrics and report the one with p < 0.05, you have likely found noise — with 20 tests, you expect one false positive at the 5% level by chance alone. Fixes: decide your primary metric in advance (Chapter 12's template forces this), treat the rest as exploratory, or apply a correction (Bonferroni: divide your threshold by the number of tests; or control the false discovery rate). At minimum, disclose how many comparisons you ran.
Power is the probability your eval will detect a real difference of a given size. An underpowered eval — too few items to detect the effect you care about — wastes everyone's time: a null result tells you nothing, because the eval could not have found the effect anyway.
A handy approximation for comparing two proportions (accuracies): to detect a difference of d percentage points at 80% power with 5% significance, you need roughly this many items per system:
| True difference | Items needed (approx.) |
|---|---|
| 2 points | ~4,800 |
| 5 points | ~800 |
| 10 points | ~200 |
| 20 points | ~50 |
(Assumes baseline accuracy near 80% and paired items; unpaired needs ~2×.)
The sobering lesson: detecting small improvements requires large evals. If your expected gain is 2 points, 200 items cannot show it — and neither can most published eval sets. Either build a bigger set, target bigger effects, or frame the claim around qualitative analysis rather than a small numeric gap. Reviewers increasingly do this math themselves; do it first.
# Quick power check with statsmodels (install: pip install statsmodels)
from statsmodels.stats.power import NormalIndPower
n = NormalIndPower().solve_power(effect_size=0.1, alpha=0.05, power=0.8)
print(f"Items per group (unpaired): {n:.0f}")
Every results table in your paper should carry uncertainty. The pattern:
| System | Accuracy | 95% CI | Δ vs. baseline | p (paired perm.) |
|---|---|---|---|---|
| Baseline | 81.7% | [79.4, 84.0] | — | — |
| Ours | 84.2% | [82.0, 86.3] | +2.5 | 0.03 |
Plus a methods footnote: "Scores are means over 3 runs; CIs from 10,000 bootstrap resamples over items; p-values from paired permutation tests (10,000 permutations) on the primary metric." That footnote is what separates a credible results section from a decorative one. Add error bars to every bar chart — a chart without them is a chart that hides its own uncertainty.
For your research: Before running your final eval, do the power math for your expected effect size. If the required item count exceeds what you have, you have three honest options: collect more items, aim for a larger effect, or pre-register a qualitative/stratified analysis as your primary evidence instead of a headline number. What you cannot do is run an underpowered eval and report the gap as a finding anyway — that is how false "improvements" enter the literature.
A complete statistical workup, start to finish, with the numbers and the interpretation.
The data. Two RAG configurations, A (new) and B (baseline), run on the same 400 test questions. Each answer is scored 0/1 for faithfulness by a validated judge. Three runs each; run-to-run variation is small (±0.8 points), so we pool per-item majority scores.
Step 1 — Point estimates and CIs.
mean_a, lo_a, hi_a = bootstrap_ci(scores_a) # 0.845 [0.810, 0.878]
mean_b, lo_b, hi_b = bootstrap_ci(scores_b) # 0.808 [0.770, 0.845]
A: 84.5% [81.0, 87.8]. B: 80.8% [77.0, 84.5]. The gap is +3.7 points, but the intervals overlap substantially — the overlap alone tells you the evidence is suggestive, not conclusive.
Step 2 — Paired test. Because both systems ran on the same items, use McNemar's test on the disagreement cells: A-right/B-wrong = 38 items; A-wrong/B-right = 23 items.
from statsmodels.stats.contingency_tables import mcnemar
table = [[339 - 0, 38], [23, 0]] # [[both right, A-only], [B-only, both wrong]]
# simplified: mcnemar on [[a, b],[c, d]] uses b vs c
result = mcnemar([[316, 38], [23, 23]], exact=True)
print(result.pvalue) # ≈ 0.07
p ≈ 0.07 — not significant at α = 0.05. The honest conclusion: no statistically significant difference on the full set, despite the 3.7-point gap. This is the moment undisciplined teams publish anyway. The disciplined team looks deeper.
Step 3 — Slice analysis (pre-registered). The eval plan declared two slices: single-passage (250 items) and multi-passage synthesis (150 items).
| Slice | A | B | Δ | p (McNemar) |
|---|---|---|---|---|
| Single-passage (n=250) | 88.0% | 86.8% | +1.2 | 0.62 |
| Multi-passage (n=150) | 78.7% | 70.7% | +8.0 | 0.02 |
The improvement is concentrated in multi-passage synthesis — exactly the capability the method was designed for. The slice p-value (0.02) survives because this comparison was pre-registered as the primary analysis; had the team tested 10 slices and reported the best, a correction would be required.
Step 4 — Effect size and practical meaning. On multi-passage items, the gap is 8 points with CI roughly [+1.5, +14.5]. In practical terms: about 12 more faithful answers per 150 hard questions. Combined with no regression on the single-passage slice and equal cost/latency, this is a real, shippable improvement — on the hard slice, which is where users were complaining.
What goes in the paper. The results table shows all three rows (overall + two slices) with CIs; the text states the primary metric and the pre-registered slice analysis; the p-values are reported exactly (p = 0.07, p = 0.02), not thresholded into "significant/not significant" alone; and the discussion notes that the overall gap is not significant but the targeted slice is. A reviewer can disagree with the interpretation, but they cannot accuse the authors of hiding anything.
Common misreadings this example guards against: (a) "p = 0.07 means no effect" — wrong; it means insufficient evidence, and the CI shows the effect could be as large as +7 points overall. (b) "The slice p = 0.02 proves the method works everywhere" — wrong; it supports the claim for multi-passage synthesis specifically. (c) "We could just run more items until p < 0.05" — that is p-hacking by data peeking; decide the sample size from the power analysis before running (Chapter 8), or pre-register a sequential design.
The p-value tells you whether a difference could be chance. Two richer questions usually matter more: what range of effects is compatible with the data? and is the effect big enough to care about?
Read CIs as compatibility intervals. A 95% CI of [+0.5, +7.0] points doesn't just say "significant" — it says effects from half a point (trivial) to seven points (large) are all reasonably compatible with what you observed. If your CI spans both trivial and large effects, your experiment was too small to pin down the practical question, even if p < 0.05. Report the interval, then discuss both ends: "the improvement could be as small as 0.5 points — below our 2-point practical threshold — or as large as 7."
The region of practical equivalence (ROPE). Before the experiment, define the range of differences you consider practically meaningless — say, ±1.5 points for your task. Afterward, check where the CI falls: entirely outside the ROPE (meaningful effect), entirely inside (practically no difference — a publishable null result!), or overlapping (inconclusive). This turns the vague "is it significant?" into the sharper "is it significant and does it matter?" — and it gives null results a rigorous interpretation instead of leaving them in limbo.
A Bayesian one-liner. If you want a direct statement like "there's a 93% chance the true improvement exceeds 2 points," fit a simple Bayesian model (a Beta-Binomial for binary outcomes takes five lines with any probabilistic programming library). You don't need to become a Bayesian — just know the option exists for when stakeholders ask the direct question that p-values refuse to answer.
The reporting habit that covers all of this: every comparison gets four numbers — the estimated difference, its 95% CI, the p-value, and your pre-declared practical threshold. Readers who care about significance get the p-value; readers who care about decisions get the rest. Nobody has to guess what you think the numbers mean, because the threshold says it out loud.
A RAG (retrieval-augmented generation) system and an AI agent fail in ways that plain LLM evals do not capture. A RAG answer can be fluent, relevant, and unsupported by anything retrieved — the worst failure mode, invisible to BLEU and ROUGE. An agent can produce a beautiful chain of thought while calling tools in the wrong order and never reaching the goal — invisible to any text metric. Evaluating these systems means measuring the pipeline, not just the final string: what was retrieved, what was cited, what tools were called, whether the goal state was reached.
A RAG system has two halves — retrieval and generation — and you must evaluate both, because they fail independently.
Given a question and a corpus, did the retriever find the passages needed to answer?
def recall_at_k(retrieved_ids, gold_ids, k):
retrieved_k = set(retrieved_ids[:k])
gold = set(gold_ids)
return len(retrieved_k & gold) / len(gold) if gold else 0.0
Diagnosing with retrieval slices. Break failures down: is recall low because the query is ambiguous (needs query rewriting), because the chunking split the answer across chunks (needs better chunking or larger windows), or because the embedding model misses domain vocabulary (needs fine-tuning or hybrid sparse+dense retrieval)? Per-failure-category analysis turns "recall@5 = 0.62" into an engineering plan.
The central RAG failure is the unsupported claim — a statement in the answer not backed by retrieved passages. Measure it directly:
A practical claim-decomposition prompt for the judge: "Break the following answer into atomic factual claims, one per line, each verifiable independently." Then verify each claim against the passages. Frameworks like RAGAS (Es et al., 2023) package this decomposition-and-verification loop; whether you use the library or roll your own, validate the judge (Chapter 4) — faithfulness judges have their own error rates, typically 5–15% per claim.
Evaluate in pipeline order: fix retrieval before tuning generation. A common waste of effort is prompt-engineering the generator when recall@5 is 0.4 — the generator is being asked to answer from passages that do not contain the answer, and no prompt fixes that. The disciplined loop:
An agent acts in an environment: calling tools, browsing, writing files. Its "output" is a trajectory of actions plus a final state. Evaluate accordingly:
# Trajectory scoring sketch
def score_trajectory(steps, goal_check, allowed_tools):
violations = [s for s in steps if s.tool not in allowed_tools]
if violations:
return {"success": False, "reason": "constraint violation"}
success = goal_check() # inspect environment state
return {"success": success, "steps": len(steps)}
Evaluating the reasoning trace. Chain-of-thought traces are tempting to score, but beware: a trace can be coherent while the actions are wrong, and models rationalize bad actions fluently. Score traces only for specific, checkable properties ("does the trace mention the tool result before acting on it?") rather than general "reasoning quality." The environment state is the ground truth; the trace is commentary.
Benchmarks for agents (WebArena, SWE-bench, and similar) provide environments with goal-state checks — useful, with the usual contamination and saturation caveats from Chapter 5. For your own agent, build a small suite of 30–100 tasks in a sandboxed environment with programmatic goal checks; this is your agent's equivalent of a unit-test suite, and it pays for itself within weeks.
Agents and RAG systems consume variable amounts of compute per task — retrieval calls, multiple model calls, long trajectories. Always report cost per task (or per 1k tasks) and latency alongside quality metrics. A method that gains 2 points of success rate at 10× the cost is a research result, not a deployment candidate, and reviewers in applied venues will ask. The honest table has four columns: success rate, steps, cost, latency.
For your research: When you claim an improvement to a RAG system or agent, attribute it: show the metric for the component you changed and the end-to-end metric, so the reader can see the causal chain ("better chunking → recall@5 0.62→0.78 → faithfulness 0.81→0.88"). Unattributed end-to-end gains invite the suspicion that something else changed — and in agentic systems, something else always might have.
Multi-hop RAG. When answers require combining facts from several passages, add hop-level diagnostics: hop recall (for each required fact, was a supporting passage retrieved?), and reasoning-chain faithfulness (does each step of a multi-step answer cite its source?). A frequent failure: the system retrieves hop 1 correctly, then generates hop 2 from parametric memory instead of retrieving it — the answer looks complete but half of it is ungrounded. Diagnose by scoring faithfulness per claim and mapping unsupported claims back to hops; if unsupported claims cluster at later hops, your retriever needs iterative (multi-round) retrieval, not a bigger generator.
Long-context eval. When the corpus chunk placed in context grows to tens of thousands of tokens, position effects dominate: models attend best to the start and end of the context and worst to the middle ("lost in the middle"). Test this directly with a needle-in-the-context probe: place the answer-bearing passage at 10 different positions in a long context and plot accuracy by position. If you see the U-shaped curve, mitigations include re-ranking key passages to the top, repeating the question after the context, or simply retrieving less (a short, precise context beats a long, noisy one for most QA tasks).
Agent case study: the ticket-triage agent. A team builds an agent that reads incoming support tickets, queries the order database, and either resolves the ticket or escalates with a summary. Their 40-task sandbox eval:
| Metric | Result | Notes |
|---|---|---|
| Task success rate | 78% (31/40) | Goal state: correct resolution or correct escalation |
| Step efficiency | 14.2 avg steps vs. 6 optimal | Agent re-reads tickets and retries queries excessively |
| Tool-call accuracy | 91% function name, 74% arguments | Date-format and order-ID errors dominate |
| Error recovery | 45% | After a failed DB query, usually spirals instead of reformulating |
| Safety violations | 2/40 | Issued a refund without confirmation twice — automatic failures |
The headline (78% success) looks deployable; the breakdown says otherwise. Argument accuracy of 74% means one in four tool calls has a malformed argument — the agent succeeds despite its tool use, by retrying. Error recovery at 45% means novel failures become infinite loops. And 2 safety violations in 40 tasks is a hard stop: at production volume, that is hundreds of unauthorized refunds. The team's roadmap wrote itself: (1) argument validation layer before tool execution, (2) a recovery policy (max 2 retries, then escalate), (3) a confirmation gate on irreversible actions — re-evaluated against the same 40 tasks. Notice how each metric maps to exactly one engineering fix. That is what good agent evals do: they don't just score, they prescribe.
In agent evals, "tool-call accuracy 74%" is a starting point, not a diagnosis. Break argument errors into types — each has a different fix:
search_orders when it needed get_order_details — usually a sign the tool descriptions are ambiguous. Fix: rewrite the descriptions with contrastive examples ("use X for listing, Y for details of a specific order"), which helps more than adding more tools.How to use the taxonomy: on your next 50 failed tool calls, tag each with one of the six types and count. The distribution tells you exactly where to invest: 60% format errors means fix the schema layer this week; 60% hallucinated IDs means change the prompting strategy. Report the distribution in your paper's failure analysis — "of 83 argument errors, 41% were format errors, 27% hallucinated identifiers…" — and reviewers will see a team that understands its system rather than one that merely scored it.
Reviewers skim the method, but they read the evaluation. It is where they decide whether your numbers mean anything. A surprising number of rejections trace back not to weak ideas but to eval sections that leave basic questions unanswered: What exactly was measured? On what data? With what prompt and settings? How many runs? What is the uncertainty? Each unanswered question is a reason to doubt, and doubt accumulates into "reject."
This chapter gives you the anatomy of a trustworthy eval section plus a reproducibility checklist you can apply before every submission.
1. The claim, stated as a measurable difference. Open with one paragraph: what you claim, and how the eval tests it. "We claim that constrained decoding reduces unsupported claims in RAG answers. We test this by comparing our method against a standard-decoding baseline on faithfulness (claim-level support judged by a validated LLM judge) across 600 questions stratified into four difficulty slices."
2. Data. What eval set(s), how built, how large, which split. Cite public benchmarks with version/harness; describe custom sets with a pointer to the construction details (Chapter 6 — put the datasheet in an appendix). State the test split was held out during development.
3. Systems compared. Every baseline, with enough detail to reproduce: model name and version (e.g., "gpt-4o-2026-05-01", not "GPT-4"), prompt (full text in appendix or supplement — this is non-negotiable for prompt-based methods), decoding settings (temperature, top-p, max tokens, seed), and any retrieval/tooling configuration. "We used GPT-4" is not reproducible; a dated model version plus the exact prompt is.
4. Metrics. One sentence of justification per metric (Chapter 2's rule): what it measures, why it fits the task, what it misses. Name the implementation ("sacreBLEU v2.4", "our judge prompt in Appendix B, validated at 81% agreement with humans on 150 items").
5. Protocol. Number of runs, how randomness was handled, the primary metric (pre-registered — say so), and the statistical analysis: CIs via bootstrap, paired tests, power where relevant (Chapter 8). State how ties and failures were handled ("model outputs that failed to parse as JSON were counted as incorrect; this affected 1.2% of baseline outputs").
6. Results. Tables with uncertainty (Chapter 8's pattern), per-slice breakdowns, and — crucially — analysis, not just numbers. The best eval sections include a failure analysis: sample 50–100 errors from your method, categorize them, and report the distribution. "Of 80 sampled failures, 41% were retrieval misses, 33% were unsupported elaborations, 26% were abstention failures" tells the reader (and you) what to work on next, and reviewers love it because it shows you understand your own system.
7. Limitations and threats. A short, honest paragraph: what the eval does not cover (languages, domains, long contexts), known biases (judge leniency, possible contamination), and what would change your conclusion. This paragraph does not weaken your paper — it strengthens it, because it answers the reviewer's objections before they are raised.
Run through this before submitting. Every "no" is a fix to make:
Tape this list next to your desk. Papers that pass it rarely get rejected on evaluation grounds.
If your eval shows no difference, or your method loses — report it anyway, at least in ablations. Null results with proper statistics ("no significant difference, 95% CI on the gap: [−1.2, +0.8], n=1,000") are informative: they tell the field what does not work, which is knowledge too. Burying null ablations while highlighting the one positive slice is p-hacking by omission, and experienced reviewers can smell it. The field's replication problems come largely from this asymmetry; be part of the fix.
For your research: Write your eval section's methods before you run the final experiments — the claim paragraph, the metric justifications, the primary metric, the statistical plan. This is a lightweight pre-registration, and it protects you from the most common form of self-deception: choosing the analysis after seeing the results. When the numbers come in, you execute the plan you wrote, and whatever it says is what you report.
Below is a mock eval-methods excerpt (280 words) for the medical RAG project from Chapter 12. The bracketed notes map each sentence to this chapter's checklist — imagine them as margin comments from a careful reviewer.
We evaluate on MedRAG-eval v1.0, a set of 600 patient questions sampled from anonymized clinic logs (Jan–Mar 2026), stratified into single-article (250), multi-article (150), distractor (100), and unanswerable (100) slices. [Data: size, source, stratification — ✓] Gold labels (relevant article IDs and key-point lists) were written independently by two medical students (agreement 88% on article IDs) with adjudication of disagreements. [Labeling protocol + agreement — ✓] We split 300/300 into development and test; the test split was frozen on 2026-04-02 before any method tuning. [Held-out test — ✓]
We compare our cite-then-answer pipeline against a standard RAG prompt baseline, both on
med-llm-2026-03-15at temperature 0.2; full prompts are in Appendix A. [Pinned model, settings, prompts published — ✓] Both prompts received equal tuning effort (three iterations each on the development split). [Equal effort — ✓]Our primary metric, pre-registered in our eval plan, is unsupported claims per answer, measured by an LLM judge (
judge-llm-2026-02-01, temperature 0, prompt in Appendix B) that decomposes answers into atomic claims and checks entailment against retrieved articles. [Primary metric pre-registered; judge specified — ✓] On a 150-item human-labeled validation set, the judge agreed with the majority human label on 83% of claims vs. 86% human–human agreement; position-bias checks showed a 4% flip rate. [Judge validation + bias check — ✓] Secondary metrics are completeness (1–3 rubric), abstention rate on the unanswerable slice, and cost per 100 answers. [Panel, not a single metric — ✓]We run each system three times and report means with 95% bootstrap confidence intervals (10,000 resamples); differences on the primary metric are tested with paired permutation tests. [Runs, CIs, paired tests — ✓] Outputs that failed to parse (0.8% of baseline outputs) were counted as incorrect. [Failure handling disclosed — ✓]
A reviewer reading this excerpt can verify every checklist item without hunting through the paper. Notice what is not here: no superlatives ("significantly outperforms" appears nowhere — the numbers will speak in the results), no copied baseline numbers, no unreported settings. The tone throughout is auditable. That tone — more than any individual item — is what makes reviewers trust an eval section.
The results paragraph that follows should mirror this structure: a table with CIs (Chapter 8's pattern), per-slice rows, then a failure analysis paragraph ("Of 60 sampled errors from our method, 55% were retrieval misses on multi-article questions…"), and finally the limitations paragraph ("Our items are English-only and drawn from two clinics; performance on other specialties may differ"). Methods say what you did; results say what happened; the failure analysis says what you learned. Papers that include all three get cited; papers with only the first two get questioned.
Real reviewer comments on eval sections, with how to respond — ideally by fixing the paper, not just the rebuttal:
"The improvement over the baseline is marginal (1.8%)." Response: add CIs and the power analysis. If the CI is [+0.2, +3.4], concede the effect may be small and reframe: argue practical significance (the gain comes with 40% lower cost, or concentrates on the hard slice — show the slice table). If you can't make that case, consider whether the claim belongs in the paper at all. Never respond by adding more test items until p < 0.05.
"The baseline seems weak / undertuned." Response: this one is hard to rebut with words — you usually need to run the stronger baseline. In the revision, tune the baseline with the same budget, report both, and document the effort. One paragraph describing baseline tuning ("we performed 5 prompt iterations on the dev set, matching our method's tuning budget") defuses this better than any argument.
"Why wasn't human evaluation performed?" Response: if you have a validated judge, point to the validation numbers and the per-dimension agreement table — a validated judge is an answer to this question, not an evasion of it. If you don't, run the human study: 100–200 items on the primary dimension is the expected minimum for a strong claim.
"Results are reported on a single dataset." Response: add a second dataset or a robustness slice (paraphrased inputs, a second domain). If truly infeasible, say why and scope the claim down: "our results establish the effect on [dataset]; generalization to [other settings] remains future work." Reviewers accept scoped claims; they reject universal ones supported by one dataset.
"The paper doesn't discuss failure cases." Response: add the failure analysis — sample 50–100 errors, categorize, report. This is the highest-leverage revision per hour of work: it answers the reviewer's underlying question ("do the authors understand their own method?") directly.
The meta-strategy: treat every eval critique as a request for evidence, not as an attack. The rebuttal that works is rarely "the reviewer is wrong" — it's a new table, a new CI, a new slice analysis, or an honest scoping of the claim. Keep your eval artifacts (prompts, splits, scripts) organized precisely so you can produce these in the rebuttal window.
Every mistake below appears regularly in published papers, workshop submissions, and industry eval reports. Each entry follows the same format: the mistake, why it happens, why it matters, and the concrete fix. Read this chapter as a pre-flight checklist for your own work.
Why it happens: One run is fast; the number looks clean; nobody asks for more — until a reviewer does. Why it matters: LLM outputs are stochastic (Chapter 1). A single run's score can easily sit 2–3 points from the true mean, which is larger than many claimed "improvements." Fix: Run everything ≥3 times with fixed decoding settings; report mean ± 95% CI (bootstrap, Chapter 8). Budget for this from the start — it is not optional polish.
Why it happens: The test set is right there, and each tweak visibly moves the number. It feels like progress. Why it matters: You are overfitting the prompt to those specific items. The reported score is optimistic, and the "improvement" often evaporates on fresh data. Fix: Split dev/test before you start iterating (Chapter 6). Touch the test set once, at the end. If you slipped and tuned on it, say so honestly and label the numbers as development results.
Why it happens: Re-running baselines is work; the number is right there in Table 2 of the prior paper. Why it matters: Prompting, harness, model version, and decoding settings differ across papers. You end up comparing your carefully tuned 5-shot number against someone's 0-shot number — or against a different model version entirely. Fix: Re-run every baseline yourself under your exact protocol. If a baseline is truly infeasible to run, flag copied numbers explicitly with the caveat, and never make them the centerpiece comparison.
Why it happens: Judge prompts are easy to write, and the scores look authoritative. Why it matters: An unvalidated judge may reward verbosity, prefer its own model family, or simply disagree with humans on your task — and you would never know. Your results section would then report the judge's biases as findings. Fix: Validate against 100–200 human-labeled items; require judge–human agreement near human–human agreement, per dimension (Chapter 4). Report the validation numbers. No validation, no judge-based claims.
Why it happens: A single number is simple to report and easy to optimize. Why it matters: Every metric is blind to something (Chapter 2) — ROUGE to factuality, EM to paraphrase, judge scores to their own biases. A single metric lets real regressions hide: your "improved" system may have traded factuality for fluency. Fix: Report a small panel (2–4 metrics) covering complementary dimensions, plus a human or validated-judge check on the dimension you care about most.
Why it happens: Error analysis is tedious, and the headline number feels like the deliverable. Why it matters: Without it, you do not know why your method works — or whether it works for the reason you think. Reviewers probe exactly this gap: "Is the gain from better reasoning, or did the baseline just format outputs badly?" Fix: Sample 50–100 failures per system, categorize them with a second reader, and report the distribution (Chapter 10). Budget half a day; it is the highest-insight-per-hour activity in eval work.
Why it happens: The eval set has 150 items, the gap is 3 points, and it feels like a win. Why it matters: On 150 items, a 3-point gap is ~4 items flipping — noise. Publishing it adds a false "improvement" to the literature that someone else will waste months trying to build on. Fix: Do the power math before running (Chapter 8). If you cannot afford the items, shrink the claim: report the result as exploratory, lead with qualitative analysis, or collect more data.
Why it happens: Test items get used in few-shot prompts, discussed in team chats that end up in training data, or included in fine-tuning sets by accident. Why it matters: Once the model has seen the test items, your eval measures memorization, not capability. Scores inflate, and the inflation is invisible from inside. Fix: Keep test items out of prompts, training data, and public repos. Prefer private or newly written items for high-stakes claims. If contamination is possible, say so and discount accordingly.
Why it happens: You give your method the good prompt, the tuned temperature, the careful output parsing — and run the baseline with defaults. Why it matters: The comparison then measures effort, not method quality. This is the "weak baseline" objection (Chapter 10), and it is fatal in review. Fix: Give every system the same tuning budget and the same protocol. Tune the baseline's prompt with the same care as your own. Document the effort per system.
Why it happens: Bare numbers look decisive; intervals look wishy-washy. Why it matters: It is the opposite — bare numbers are the wishy-washy ones, because they hide how much could be noise. Reviewers in 2026 expect uncertainty quantification; its absence reads as either ignorance or concealment. Fix: CIs on every number, error bars on every chart, a limitations paragraph in every paper (Chapters 8, 10). Uncertainty honestly reported is a strength signal, not a weakness.
Notice what these ten share: every one is a way of accidentally fooling yourself — and then, through publication, fooling others. The field's eval crises (contamination, irreproducible gains, judge biases treated as findings) are not caused by dishonesty; they are caused by smart people skipping the boring parts. The fixes are all boring: more runs, held-out splits, re-run baselines, validated judges, reported uncertainty. Boring is what makes science work.
For your research: Pick the three mistakes you are most likely to make — be honest — and write their fixes into your project plan as scheduled tasks with dates ("validate judge against 150 human labels by Oct 20"). Mistakes you plan against do not happen; mistakes you merely intend to avoid do.
Score your current (or planned) project honestly: 0 = not done, 1 = partially done, 2 = done well. Total out of 20.
Interpretation. 18–20: submission-ready on eval grounds. 14–17: solid, with known gaps — fix them before submission and disclose the rest. 10–13: the eval needs a dedicated work cycle; do not submit yet. Below 10: you are exploring, not evaluating — label it as such and keep building.
The "fix this week" planner. Take your three lowest scores and convert each into a calendar task with a concrete deliverable: - Score 0 on #4 → "Validate judge: label 150 items, compute agreement — due Friday." - Score 1 on #2 → "Freeze test split: move 300 items to eval/test/, delete from prompts — due Wednesday." - Score 0 on #6 → "Failure analysis: sample 60 errors, categorize with a partner — due next Monday."
Re-run the audit monthly during an active project. The score should climb; if it doesn't, the eval is rotting while the method advances — the exact situation this book exists to prevent. Tape the worksheet to the wall next to the reproducibility checklist (Chapter 10). Between the two, there is very little a reviewer can ask that you haven't already answered.
Not all ten mistakes are equally likely at every stage. Here's where each one ambushes you — and what to prioritize when time is short:
Course project / hackathon (days). The killers are #1 (single run), #5 (one metric), and #10 (no uncertainty). You're moving fast, but three runs and a bootstrap CI cost minutes and separate a serious project from a demo. Skip #6 (full failure analysis) if you must — but write down three example failures you noticed; that's a mini failure analysis and it counts.
Workshop / short paper (weeks). Add #2 (test-set tuning) and #7 (underpowered claims) to the watch list. With small eval sets, the temptation to tune on test and to overclaim small gaps is strongest. Defenses: freeze the test split on day one, and do the power math before writing the abstract's numbers.
Conference paper (months). All ten are in play, but #3 (copied baselines), #4 (unvalidated judges), and #9 (unequal tuning effort) are the ones reviewers hunt for, because they indicate how seriously you took the comparison. Budget two full weeks for baselines alone — re-running and tuning someone else's method is unglamorous and absolutely decisive.
Production system (ongoing). #2 becomes #8's cousin: test data leaking into training/fine-tuning pipelines over time. And an eleventh mistake appears that this chapter didn't list: stopping evaluation after launch. The regression suite (Chapter 7) is the fix — eval is not a phase, it's a permanent organ of the system.
If you have one hour to de-risk any eval, spend it in this order: (1) 15 min: verify the test split was never tuned on (#2); (2) 15 min: add CIs to the headline numbers (#10, #1); (3) 15 min: check that baselines ran under the same protocol (#3, #9); (4) 15 min: sketch the failure analysis from errors you've already seen (#6). One focused hour catches the mistakes that sink the most papers.
There's one mistake this chapter didn't number, because it's the slow-motion version of Mistake 2: optimizing your method against your own eval set for so long that the eval stops measuring the real task. It happens gradually — each prompt tweak is validated on dev, each dev-validated tweak ships, and after twenty iterations your method is exquisitely fitted to 300 questions that no longer represent the wild. The symptom: dev scores climb steadily while user complaints stay flat. The defenses are structural: refresh a portion of the dev set periodically with new items, keep the test set truly frozen and truly private, and — most honestly — track a live metric (user thumbs-down rate, task completion in production) alongside the eval suite. When the eval and the live metric diverge, believe the live metric and rebuild the eval. An eval set is a model of the task, and like all models, it decays — maintaining it is part of the job, not a one-time cost.
You have the concepts; now you need a document. An eval plan is a short, written plan — made before the final experiments — that says what you will measure, how, and what would count as success. It is the single highest-leverage artifact in eval work: it forces every decision in Chapters 2–9 to be made explicitly, it becomes the methods section of your paper nearly verbatim, and it protects you from post-hoc goalpost-moving (Chapter 11, Mistake 2).
This chapter gives you a fill-in template and then works it through a realistic example so you can see every box filled.
Copy this into your project repo as eval/PLAN.md and fill it in:
# Eval Plan: [PROJECT NAME] — v1.0, [DATE]
## 1. Decision
What decision will these evals inform? (One paragraph.)
## 2. Claim
The measurable claim, in one sentence:
"Method M improves [dimension] on [task] by [expected amount] vs. [baseline]."
## 3. Data
- Eval set(s): [name, version, size, source]
- Construction: [how items were sourced/labeled; agreement rate]
- Splits: [dev N / test N; test held out since DATE]
- Stratification: [dimensions and counts per slice]
## 4. Systems
- Candidate: [model version, prompt location, decoding settings]
- Baselines: [each with same detail; tuning effort per system]
- Blinding: [how outputs are anonymized for raters/judges]
## 5. Metrics
- Primary metric: [ONE metric — pre-registered]
- Secondary metrics: [2-3, labeled exploratory]
- Per metric: implementation + one-sentence justification
## 6. Judges & raters
- Human eval: [N raters, background, training, rubric location,
blinding, agreement target, disagreement rule]
- LLM judge: [model version, prompt location, validation sample size,
judge-human agreement target, bias checks]
## 7. Statistics
- Runs per system: [≥3]
- Uncertainty: [bootstrap 95% CIs, 10k resamples]
- Tests: [paired permutation / McNemar on primary metric]
- Power: [effect size targeted, items needed, items available]
- Multiple comparisons: [how handled]
## 8. Regression & hygiene
- Suite layers: [smoke/slice/full triggers]
- Model pinning: [exact versions]
- Contamination notes: [any risks, mitigations]
## 9. Success criteria
- Ship/claim threshold: [e.g., "primary metric +3 points with p<0.05
AND no secondary metric regresses beyond its CI"]
- Failure plan: [what you conclude if the threshold is not met]
## 10. Timeline & owners
- [Milestone]: [owner] by [date] (judge validation, human eval, final runs...)
Two fields do most of the work: #1 (Decision) keeps the eval honest about its purpose, and #9 (Success criteria) written in advance is what separates science from storytelling. If you cannot fill #9, you are not ready to run the final eval — you are still exploring, which is fine, but label it exploration.
Background. You are building a RAG assistant that answers patient questions from a corpus of 2,000 hospital help articles. Doctors flagged that the bot sometimes adds plausible-but-unsourced medical details. Your proposed fix: a "cite-then-answer" prompt plus a verification pass that drops unsupported sentences. You need to know: does the fix actually reduce unsupported claims, without hurting completeness?
1. Decision. Whether to ship the cite-then-answer pipeline as the default for the patient-facing bot. The deciding factors are faithfulness (must improve) and completeness (must not regress).
2. Claim. "Cite-then-answer reduces unsupported claims per answer on medical RAG questions by at least 30% relative, versus the standard prompt, with no significant completeness regression."
3. Data. Custom set "MedRAG-eval v1.0": 600 real anonymized patient questions sampled from clinic logs (2026-01 to 2026-03), stratified: single-article (250), multi-article synthesis (150), distractor present (100), unanswerable (100). Gold labels: relevant article IDs + key-points lists, written by two medical students independently (agreement: 88% on article IDs, adjudicated). Splits: 300 dev / 300 test, stratified; test frozen 2026-04-02.
4. Systems. Candidate: cite-then-answer pipeline on med-llm-2026-03-15, temperature 0.2, prompt in prompts/cite_then_answer.md. Baseline: standard RAG prompt, same model and settings, prompt in prompts/standard.md, tuned with equal effort (3 prompt iterations each on dev). Outputs stripped of formatting tells; order randomized for judges.
5. Metrics. Primary: unsupported claims per answer (count, lower is better) — validated LLM judge decomposing answers into atomic claims and checking entailment against retrieved articles. Secondary: completeness (1–3 rubric, judge), abstention rate on the unanswerable slice, answer EM on factoid subset. Justifications: unsupported-claim count directly measures the reported failure; completeness guards against the fix making answers uselessly terse.
6. Judges & raters. LLM judge: judge-llm-2026-02-01, temperature 0, prompt in prompts/judge_faithfulness.md. Validation: 150 dev items double-labeled by the two medical students; judge–human agreement 83% per-claim vs. 86% human–human; position bias checked (4% — acceptable, disclosed). Human spot-check: 100 test items × 2 raters on completeness.
7. Statistics. 3 runs per system (temperature 0.2 — small sampling variance; 3 runs enough). 95% CIs via 10k bootstrap over items. Primary test: paired permutation on unsupported-claim counts (items paired across systems). Power: targeting 30% relative reduction from a baseline mean of ~1.1 unsupported claims/answer — with 300 test items, power >90%. Multiple comparisons: primary metric tested at α=0.05; secondary metrics reported with CIs, no significance claims.
8. Regression & hygiene. Smoke suite (20 items) on every prompt edit; slice eval on dev on every pipeline change; full test run once at the end. Model versions pinned; test questions never appear in prompts or training; contamination risk low (private logs, unpublished items).
9. Success criteria. Ship if: unsupported claims drop ≥30% relative with p<0.05 AND completeness mean does not drop beyond its CI AND abstention rate on unanswerables ≥90%. If faithfulness improves but completeness regresses, the conclusion is "trade-off found, needs redesign" — not a ship.
10. Timeline. Judge validation by Apr 10 (you) · human spot-check protocol by Apr 14 (teammate) · final test runs Apr 18 · decision memo Apr 20.
Notice how every chapter of this book appears exactly once: dimensions and stratification (Ch. 6), judge validation (Ch. 4), human spot-check design (Ch. 3), metric justification (Ch. 2), power and paired tests (Ch. 8), regression layers (Ch. 7), contamination hygiene (Ch. 11), and the write-up structure (Ch. 10). The plan is the book, compressed into two pages.
Revisit the plan when reality intrudes — a failed judge validation, a smaller effect than expected — and record the changes with dates. A plan that evolves transparently is science; a plan silently rewritten after seeing results is fiction.
For your research: Your eval plan is also a collaboration tool. Share it with your advisor or teammates before the final runs and ask: "If these results come back positive, will you believe them?" Every objection they raise now is a reviewer objection you get to fix for free. The most expensive sentence in research is "we should have measured X" — spoken after the experiments are done.
The Chapter 12 template scales down as well as up. Two condensed examples:
Plan A — Course project (light). Decision: which of two prompts to use for a class demo classifying movie reviews. Claim: "Prompt B beats Prompt A on accuracy by ≥5 points on our 200-review set." Data: 200 IMDb-style reviews you label yourself (100 dev / 100 test). Systems: same model, temperature 0, prompts in the repo. Metrics: accuracy (primary — single, because the task is closed and the stakes are a demo). Statistics: 3 runs, bootstrap CI; power check: 100 test items can detect ~10-point gaps, so a 5-point claim is exploratory — the plan says so honestly. Success criterion: ship B if the gap clears +5 with non-overlapping CIs; otherwise report "no conclusive difference." Even this tiny plan has the two fields that matter: the decision and the pre-written success criterion.
Plan B — Web agent for form filling (agent-adapted). Decision: whether the agent is safe to pilot with real users. Claim: "The agent completes 80% of 50 form-filling tasks with zero irreversible-action violations." Data: 50 tasks in a sandbox CRM with scripted goal checks (field values correct = success). Systems: agent v2 vs. v1 (frozen baseline from last month's suite run). Metrics (primary + behavioral): task success rate (primary), argument accuracy per tool call, recovery rate on injected tool failures, violation count — where any violation fails the release regardless of success rate (safety as a gate, not a metric). Statistics: 3 runs per task (agents are stochastic); 95% CIs; violations reported as raw counts with task IDs. Hygiene: sandbox reset between tasks (no state leaking across tasks — a classic agent-eval bug), tool-call logs archived per run. Success criterion: success ≥80% AND violations = 0 AND no metric regresses beyond its CI vs. v1. The plan also names the failure conclusion in advance: "If success clears 80% but violations > 0, the outcome is 'not pilotable,' not 'mostly good.'"
Notice what both plans share with the full medical-RAG example: the decision is explicit, the primary metric is singular, the failure conclusion is written before the results exist, and safety/quality gates are stated as rules rather than vibes. The template is the same; only the depth changes. Start every project — a weekend hack or a dissertation — by filling in even the light version. It takes twenty minutes and prevents the two most expensive eval failures: measuring the wrong thing, and deciding what "good enough" means after seeing the numbers.
A template is static; a project moves. Here's how the Chapter 12 template maps onto a realistic six-week timeline for a semester research project:
Week 1 — Decide and specify. Fill in template sections 1–3 (decision, claim, data). Write the eval spec: dimensions, target counts, sourcing plan. The deliverable is a one-page spec your advisor can critique. Most projects should spend more time here, not less — a week of specification saves a month of re-measurement.
Week 2 — Build the dev set. Source and label the development slice (aim for half your target items). Write gold labels with a partner; measure label agreement on the first 30 items before labeling the rest. Set up the JSONL schema (Chapter 6) and version control now, not later.
Week 3 — Baselines and harness. Implement the eval harness (runner, caching, aggregation). Run baselines first — all of them, three runs each — and freeze the baseline numbers. Write the smoke tests and behavioral assertions (Chapter 7). If you're using an LLM judge, draft the judge prompt this week.
Week 4 — Validate. Run the judge validation against human labels (Chapters 3–4): label 100–150 items, compute agreement, run bias checks. If the judge fails validation, you still have time to revise the rubric or fall back to human rating for the test set. Also finalize the test split freeze — no new tuning against it after this week.
Week 5 — Final runs. Execute the full test-set eval: all systems, three runs, all slices. No method changes this week — if you get the urge to tweak the prompt, write it down as future work. Run the failure analysis on sampled errors (Chapter 10) while the runs execute.
Week 6 — Write and audit. Fill in the results using the pre-registered analysis plan. Run the reproducibility checklist (Chapter 10) and the self-audit worksheet (Chapter 11). Write the limitations paragraph before the conclusion — it keeps the conclusion honest.
The two rules that hold the timeline together: (1) never move a deadline by cutting validation — cut scope instead (fewer slices, fewer baselines); (2) the test set stays frozen from Week 4 regardless of what the dev numbers suggest. Teams that follow these two rules produce evals that survive review; teams that don't produce evals that need to be redone.
| Task shape | Primary metric | Why | Watch out for |
|---|---|---|---|
| Multiple choice | Accuracy | Single right answer | Prompt-format sensitivity |
| Short extractive QA | EM + token F1 | Partial credit for spans | "Not guilty" vs "guilty" overlap |
| Machine translation | sacreBLEU / chrF | System-level ranking | Weak sentence-level correlation |
| Summarization | ROUGE-L + factuality check | Legacy comparability | Blind to hallucinations |
| Code with tests | pass@k | Behavioral, not textual | Weak tests inflate scores |
| Open-ended QA/chat | Validated judge or human rubric | Captures graded quality | Judge biases (Ch. 4) |
| RAG answers | Unsupported-claim rate + citation recall | Measures the central failure | Judge error 5–15% per claim |
| RAG retrieval | Recall@k, nDCG@k | Diagnoses pipeline stage | High recall + low precision = noise |
| Agents | Task success rate (goal state) | Behavior, not text | Ignores cost/efficiency |
| Safety evals | Refusal rate, violation rate | Direct policy measurement | Over-refusal hurts usefulness |
| Method | Cost | Speed | Best for | Key risk |
|---|---|---|---|---|
| String metrics (EM, BLEU, ROUGE) | Free | Instant | Closed tasks, legacy comparison | Blind to meaning/paraphrase |
| Human evaluation | High | Days–weeks | Gold-standard quality judgments | Rater disagreement, cost |
| LLM-as-judge | Low–medium | Hours | Scaling rubric judgments | Biases; must validate vs. humans |
| Benchmarks (MMLU, HumanEval…) | Low | Hours | Positioning vs. field | Contamination, saturation |
| Custom eval sets | Medium | Days | Your actual claim | Label error, under-sizing |
| Behavioral checks (tests, goal states) | Low | Minutes | Code, agents, RAG pipelines | Test/goal quality is everything |
| # | Mistake | Fix |
|---|---|---|
| 1 | Single run reported as "the score" | ≥3 runs; mean ± 95% CI |
| 2 | Prompt tuned on the test set | Dev/test split; touch test once |
| 3 | Baseline numbers copied from other papers | Re-run all baselines identically |
| 4 | Unvalidated LLM judge | Validate vs. 100–200 human labels |
| 5 | One metric on open-ended tasks | Panel of 2–4 complementary metrics |
| 6 | No failure analysis | Categorize 50–100 sampled errors |
| 7 | Underpowered significance claims | Power math first; size eval to effect |
| 8 | Contaminated/leaky test data | Keep test items private and unpublished |
| 9 | Unequal tuning effort across systems | Same protocol and budget for all |
| 10 | No uncertainty reported | CIs, error bars, limitations paragraph |
[1] D. Hendrycks et al., "Measuring massive multitask language understanding," in Proc. Int. Conf. Learn. Representations (ICLR), 2021.
[2] M. Chen et al., "Evaluating large language models trained on code," arXiv:2107.03374, 2021.
[3] K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, "BLEU: a method for automatic evaluation of machine translation," in Proc. 40th Annu. Meet. Assoc. Comput. Linguistics (ACL), Philadelphia, PA, USA, 2002, pp. 311–318.
[4] C.-Y. Lin, "ROUGE: A package for automatic evaluation of summaries," in Text Summarization Branches Out: Proc. ACL Workshop, Barcelona, Spain, 2004, pp. 74–81.
[5] L. Zheng et al., "Judging LLM-as-a-judge with MT-Bench and Chatbot Arena," in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 36, 2023.
[6] P. Rajpurkar, J. Zhang, K. Lopyrev, and P. Liang, "SQuAD: 100,000+ questions for machine comprehension of text," in Proc. Conf. Empir. Methods Nat. Lang. Process. (EMNLP), Austin, TX, USA, 2016, pp. 2383–2392.
[7] M. Joshi, E. Choi, D. Weld, and L. Zettlemoyer, "TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension," in Proc. 55th Annu. Meet. Assoc. Comput. Linguistics (ACL), Vancouver, Canada, 2017, pp. 1601–1611.
[8] A. Wang et al., "GLUE: A multi-task benchmark and analysis platform for natural language understanding," in Proc. Int. Conf. Learn. Representations (ICLR) Workshop, 2019.
[9] P. Liang et al., "Holistic evaluation of language models," arXiv:2211.09110, 2022.
[10] J. Wei et al., "Chain-of-thought prompting elicits reasoning in large language models," in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 35, 2022.
[11] T. B. Brown et al., "Language models are few-shot learners," in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 33, 2020, pp. 1877–1901.
[12] S. Es, J. James, L. Espinosa-Anke, and S. Schockaert, "RAGAs: Automated evaluation of retrieval augmented generation," arXiv:2309.15217, 2023.
Exercise 1 — Write a rubric. Pick a task you care about (e.g., answering factual questions, summarizing news articles). Write a 3-dimension rubric with anchored 1–3 scales and at least three edge-case rules, following the Chapter 3 template. Then score 5 real model outputs with it and note every place the rubric was ambiguous — revise accordingly.
Exercise 2 — Metric justification. Choose three metrics from Chapter 2 for a task of your choice. For each, write the one-sentence justification ("We use X because… and acknowledge it misses…"). Then write one paragraph on which metric you would pre-register as primary and why.
Exercise 3 — Compute agreement. Have a friend score 30 items with your Exercise 1 rubric while you score the same 30 independently. Compute Cohen's kappa per dimension with the Chapter 3 code. Where kappa is below 0.6, discuss one disagreement and rewrite the anchor that caused it.
Exercise 4 — Validate a judge. Take the 30 doubly-labeled items from Exercise 3. Write a judge prompt (Chapter 4 template) and run it on the same items. Compute judge–human agreement per dimension and compare with your human–human kappa. Is the judge a valid substitute for any dimension? Write a two-paragraph verdict.
Exercise 5 — Build a 50-item eval set plan. For your own project (or a hypothetical RAG bot), write the eval spec from Chapter 6: claim, 3–4 capability dimensions with target counts summing to 50, sourcing plan, gold-label format, and labeling protocol. Identify the two hardest dimensions to source items for and propose a solution for each.
Exercise 6 — Bootstrap by hand. Take any 100 binary scores (e.g., correct/incorrect on 100 questions). Implement the Chapter 8 bootstrap function from scratch (no copy-paste — type it) and report the mean with 95% CI. Then halve the data to 50 items and re-run: how much wider is the interval? Write down what this implies for your project's eval size.
Exercise 7 — Power analysis. Your baseline scores 78% and you hope your method reaches 83%. Using the Chapter 8 table (or the statsmodels snippet), determine how many items you need for 80% power. If you only have 200 items, what is the smallest effect you could reliably detect? Write the honest claim you could make with 200 items.
Exercise 8 — Regression suite sketch. For a prompt-based application you use or are building, design the three-layer regression suite from Chapter 7: list 10 smoke-test items, define two behavioral assertions with code, and set a tolerance-band procedure. Write it as a eval/PLAN.md section.
Exercise 9 — RAG pipeline diagnosis. Given these numbers — recall@5 = 0.55, faithfulness = 0.90, completeness = 0.62 — where is the bottleneck, and what would you fix first? Justify with the Chapter 9 improvement loop. Then construct a second scenario (different numbers) where the bottleneck is elsewhere, and explain how the fix differs.
Exercise 10 — Full eval plan. Fill in the complete Chapter 12 template for a real or realistic project: decision, claim, data, systems, metrics (one primary), judges/raters, statistics, hygiene, success criteria, timeline. Share it with a peer and ask: "If the results come back positive, would you believe them?" Record their objections and revise the plan.
Exercise 6 sketch. With 100 binary scores at, say, 78% accuracy, the bootstrap 95% CI will be roughly [70%, 86%] — a ±8-point band. Halving to 50 items widens it to roughly ±11–12 points. The implication: on 50 items, even a 10-point "improvement" (5 items flipping) sits inside the noise band. If your project's eval has fewer than 100 items, your honest claims are limited to large effects (15+ points) or qualitative findings — size the claim to the CI, not to your hopes.
Exercise 7 sketch. For 78% vs. 83% (a 5-point gap) at 80% power, α = 0.05, paired items: you need roughly 700–900 items (use the Chapter 8 table or the statsmodels snippet for your exact numbers). With only 200 items, the smallest reliably detectable gap is about 10 points. So the honest claim with 200 items isn't "we improved by 5 points" — it's either "we improved by ≥10 points" (if the data shows it) or "exploratory results suggest a ~5-point gain; a larger eval is needed to confirm." Writing that sentence feels like losing; it's actually what separates publishable work from noise.
Exercise 9 sketch. recall@5 = 0.55 is the bottleneck: the generator can only be faithful to what it was given, and nearly half the needed passages are missing. Fix retrieval first (chunking, hybrid search, query rewriting) before touching the generator — prompt-engineering faithfulness at 0.55 recall is wasted effort. (Second scenario: recall@5 = 0.92 but faithfulness = 0.61 — now retrieval is fine and the generator is adding unsupported claims; fix with citation-forcing prompts or an answer-then-verify pass.) The rule: evaluate in pipeline order, fix the earliest failing stage.
End of Book 20 — Evaluating and Testing LLM Applications. AstolixGen Learning Series.
If this book had to survive as a single paragraph, it would be this: an eval is a decision procedure, not a number. Every choice — the metric, the judge, the items, the runs, the statistics — is justified by the decision it informs, and every number is reported with the uncertainty that keeps it honest. The researchers whose evals you trust aren't the ones with the biggest benchmarks or the fanciest judges; they're the ones whose methods section reads like an audit trail, whose failure analyses show they understand their own system, and whose limitations paragraphs answer your objections before you raise them.
Three habits will carry you further than any technique in these pages. First, write the plan before the experiment — the decision, the primary metric, the success criterion. Thirty minutes of writing saves thirty days of re-running. Second, distrust your own numbers on a schedule — re-validate the judge, refresh the dev set, check the live metric against the eval suite. Evals decay; maintenance is the job. Third, publish the boring parts — the prompts, the seeds, the labeling protocol, the null results. The field's credibility is built from exactly these unglamorous artifacts, and every one you share makes the next researcher's work easier.
The models will keep changing — new architectures, new benchmarks, new failure modes. The discipline in this book doesn't depend on any of them. Operational definitions, repeated measurement, reported uncertainty, validated instruments, honest limitations: that's the scientific method, applied to machines that are very good at sounding right. Learn it once, and every future system — whatever it looks like — becomes evaluable.
Now go measure something, and report the error bars.