← All 50 books
Evaluating and Testing LLM Applications cover

Book 20 of 50 · Free

Evaluating and Testing LLM Applications

28,641 words · 85 chapters · illustrated

Evaluating and Testing LLM Applications

Book 20 of 50 — AstolixGen Learning Series

Book cover: a magnifier glass held over a quality checklist


About This Book

Large language models changed what it means to "test" software. A traditional program either passes or fails a test: you give it an input, you compare the output to an expected value, and you get a yes or no. An LLM application gives you a fluent, confident answer that might be right, partially right, beautifully wrong, or subtly dangerous — and asking it the same question twice can produce two different answers. How do you measure that? How do you know whether your system got better or worse after you changed the prompt? How do you report results in a paper that reviewers will trust?

This book answers those questions for students and early-career researchers working in AI. It is practical first: every concept comes with a concrete example, a piece of code you can run, or a template you can copy into your own project. But it is also rigorous enough for publication. By the end, you will be able to design an evaluation plan for an LLM application, execute it with defensible statistics, avoid the mistakes that sink papers at review time, and write up your results so another researcher could reproduce them.

You do not need deep statistics to read this book. You need basic familiarity with language models — what a prompt is, what temperature does, what a benchmark is — and a willingness to think carefully about measurement. Wherever a formula appears, a plain-English explanation comes first.

What you will be able to do after reading this book

  1. Explain, in plain language, why LLM evaluation is fundamentally harder than testing deterministic software, and what that implies for experimental design.
  2. Choose the right metric for a task — exact match, token F1, BLEU, ROUGE, pass@k, human rubric scores — and justify the choice.
  3. Design a human evaluation study with a written rubric, multiple raters, blinding, and a reported agreement statistic.
  4. Use LLM-as-judge responsibly: write a judge prompt, measure the judge's own accuracy against human labels, and state its limits.
  5. Read a benchmark score (MMLU, HumanEval, and others) critically, understanding what it does and does not prove about a model.
  6. Build your own evaluation set from scratch: source items, write gold labels, and decide how many items you need.
  7. Set up regression tests so prompt edits and model swaps never silently degrade your system.
  8. Apply basic statistical thinking to evals: estimate variance, run a significance test, and report error bars.
  9. Evaluate retrieval-augmented generation (RAG) and agent systems with metrics that match how they fail.
  10. Write a reproducibility checklist for an eval section of a paper that satisfies demanding reviewers.

How to Use This Book

You don't have to read straight through — though you can. Three recommended paths:

  • "I need to evaluate my system this month." Read Chapters 1 (why it's hard), 2 (pick metrics), 6 (build the eval set), 8 (statistics), and 12 (write the plan). Then use Chapters 3–4 as needed for human or judge scoring.
  • "I'm writing a paper and want the eval section to survive review." Read Chapters 10 (reporting), 11 (mistakes), and 8 (statistics) first. Then fill in the Chapter 12 template and run the Chapter 11 self-audit.
  • "I'm building RAG or agents." Read Chapters 1, 9 (RAG/agents), 7 (regression testing), and 6. The statistics in Chapter 8 still apply — agents are stochastic too.

Each chapter ends with a "For your research" box (the actionable point) and key takeaways (the summary). The Learning Dashboard near the end collects every table and checklist in one place for quick reference during a project. The exercises are designed to be done, not just read — Exercise 10, the full eval plan, is the capstone: if you can fill in that template for your own project, you have absorbed the book.


Chapter 1 — Why LLM Evaluation Is Hard

The test that stopped being a test

Consider the oldest trick in software testing: the assertion.

assert add(2, 2) == 4

This works because add(2, 2) always returns 4. The function is deterministic: the same input always produces the same output. Most of software engineering — unit tests, integration tests, continuous integration pipelines — rests on this property.

Now consider an LLM application, say a customer-support assistant. You write the analog of a test:

response = assistant.ask("How do I reset my password?")
assert response == "Click Settings, then Security, then Reset Password."

This "test" fails in ways that tell you almost nothing:

  • The model might reply "Go to Settings → Security → Reset Password." — correct, but not string-identical. Your test fails on a correct answer.
  • Ask the same question tomorrow and get "You can reset your password from the Security page under Settings." — also correct, different wording, different test result.
  • Ask it again with temperature 0.7 and get an answer that mentions the admin password reset flow — wrong for this user, but phrased confidently.

The deterministic test collapses because LLM outputs live in an enormous space of possible strings, most of them acceptable, some of them wrong in subtle ways. Evaluating them is not a string comparison problem; it is a judgment problem. And judgment is expensive, noisy, and hard to reproduce.

This chapter explains exactly where the hardness comes from. Once you see the sources of difficulty clearly, the rest of the book — metrics, judges, rubrics, statistics — reads as a set of tools aimed at specific problems.

Source 1: Nondeterminism and sampling

Language models generate text by sampling from a probability distribution over the next token. A temperature parameter controls how much randomness enters the sampling. At temperature 0 (more precisely, greedy decoding), the model picks the most likely token at each step, which is mostly deterministic — but even greedy decoding is not perfectly stable across runs on real systems. Floating-point differences between GPU kernels, batched inference, and quantized models can all change which token wins a near-tie. At temperature 0.7 or 1.0, two calls with the same prompt routinely produce different text.

Concretely, this means an eval score is not a number — it is a random variable. Run your 200-item eval set today and get 82%; run it tomorrow and get 79%. Which one is "the" score? Neither, really. Both are draws from a distribution, and your job as an evaluator is to estimate that distribution's center and spread. Chapter 8 is entirely about doing this properly.

It also means single-example debugging misleads you. Every LLM developer has had this experience: a prompt change "fixes" the one example you were staring at and breaks ten you were not watching. Without a fixed eval set and repeated runs, you cannot tell whether a change helped. You are steering by noise.

A practical rule follows: never report a score from one run. Run each eval at least three times (more is better), with the same decoding settings you will ship, and report the mean and spread. Record the seed if the API supports one — but remember that a seed only controls your sampling; it does not make a hosted model reproducible across versions (Chapter 7 has more on this).

Source 2: The output space is open-ended

A classifier chooses among, say, 1,000 labels. An LLM answering an open question chooses among effectively infinite strings. This breaks the two conveniences that make classification evaluation easy:

  1. There is no single right answer. For "Summarize this article in two sentences," thousands of summaries are excellent. Any metric that demands closeness to one reference summary will punish good answers that phrase things differently. BLEU and ROUGE try to soften this with n-gram overlap, but they still anchor to references and correlate only weakly with human judgments on open tasks (Chapter 2).
  2. Errors are graded, not binary. A summary can miss a minor detail (small error), miss the main point (large error), or invent a fact (hallucination — potentially catastrophic). A single accuracy number flattens these into "right" and "wrong," destroying exactly the information you need to improve the system.

The response is to decompose quality into dimensions and measure each: factuality, completeness, relevance, fluency, safety. Rubrics (Chapter 3) and structured judge prompts (Chapter 4) do this decomposition. A single aggregate score is fine for a leaderboard; for research and engineering, you want the per-dimension breakdown because it tells you what to fix.

Source 3: Prompt sensitivity

Small prompt changes cause large behavior changes. Adding "Think step by step," reordering examples, or even changing whitespace can shift benchmark scores by double-digit percentages. This is not a bug you can patch; it is a property of how these models work — they are extremely context-sensitive pattern completers.

For evaluation this creates two traps:

  • The prompt you evaluate is part of the system under test. A score like "Model X gets 85% on task Y" is really "Model X with prompt P and decoding settings D gets 85% on task Y as measured by metric M." Change any of P, D, or M and the number moves. Honest reporting names all three (Chapter 10).
  • Optimizing the prompt on the eval set is overfitting. If you iterate your prompt against your 200 test items until the score peaks, you have fit the prompt to those items. The reported score is now optimistic. The fix is the same as in machine learning: keep a held-out set you touch only at the end, or at least report that the number is a development-set number (Chapter 6).

Source 4: Contamination and the moving target

Standard benchmarks are public. Public benchmarks end up in training data — accidentally, through web crawls, or deliberately. A model that has memorized MMLU questions will score brilliantly and tell you nothing about its ability to answer new questions. This is contamination, and it is one reason benchmark scores should be read with skepticism (Chapter 5).

Worse, hosted models change under you. The model behind an API endpoint can be updated, re-quantized, or have its safety filters tuned without notice — and your eval scores shift for reasons that have nothing to do with your system. This is why regression testing (Chapter 7) matters even when you change nothing: your baseline can drift.

Source 5: Evaluation is itself a judgment task

At the bottom of every automated metric sits a human decision about what "good" means — and humans disagree. Two competent raters grading the same summary will agree perhaps 70–80% of the time on fine-grained rubrics. That disagreement is not noise to be averaged away; it is a ceiling on how meaningful your metric can be. If humans agree with each other at 75%, a metric that agrees with humans at 74% is essentially perfect, and squeezing it to 78% is chasing ghosts (Chapter 4 covers how to validate judges against this ceiling).

This has a deep consequence: there is no ground truth for open-ended generation, only intersubjective agreement. Your eval does not discover the "true quality" of outputs; it operationalizes a particular definition of quality, implemented by particular raters or judges, on a particular sample of inputs. Good papers say this out loud.

What this implies for your research practice

These five sources of hardness point to a coherent discipline, which is what this book teaches:

  • Because outputs are nondeterministic, you repeat runs and report spread (Chapter 8).
  • Because outputs are open-ended, you measure quality dimensions with rubrics instead of single string-match scores (Chapters 2–4).
  • Because prompts are sensitive, you version prompts and settings as part of the system and hold out test data (Chapters 6–7).
  • Because benchmarks contaminate and models drift, you build your own eval sets and re-run baselines (Chapters 5–7).
  • Because judgment is the foundation, you validate every automated judge against human labels and report agreement (Chapters 3–4).

None of this is exotic. It is the ordinary discipline of empirical science — operational definitions, repeated measurement, reported uncertainty — applied to a new kind of artifact. Researchers coming from other empirical fields already know most of it; LLM work just forces you to apply it more carefully, because the artifact under study is unusually good at looking right while being wrong.

For your research: Before you run a single eval, write down in one paragraph: what decision will this eval inform? ("Should I ship prompt v3 or keep v2?" "Is my method better than the baseline on factuality?") Every choice in this book — metric, judge, sample size — flows from that decision. An eval without a decision is a number without a purpose, and numbers without purposes get misread.

Seeing all five at once: a worked scenario

To make the five sources of hardness concrete, follow a fictional-but-typical team — two graduate students, Amara and Ben — evaluating a support bot over one week.

Monday — nondeterminism. They run their 200-question eval set at temperature 0.7: 84%. Pleased, they re-run it Tuesday morning before a meeting: 81%. Nothing changed — same code, same data. Three questions flipped from right to wrong and back again across runs. Lesson: the score is a distribution. They switch to three runs per configuration and start reporting means with ranges.

Wednesday — open-endedness. Digging into failures, they find the bot answered "Where do I change my email?" with "Go to Profile → Account settings → Email, enter the new address, and confirm via the link we send." The gold answer was "Update it under Profile settings." Every word of the bot's answer is correct and more helpful — but their exact-match metric marked it wrong. They had been "improving" the prompt for a week against a metric that punished good answers. Lesson: they replace EM with a rubric scored by a validated judge (Chapters 3–4), and the bot's "true" score jumps 12 points without any model change. The metric had been lying, not the model.

Thursday — prompt sensitivity. Amara adds "Be concise." to the system prompt. The headline score holds steady — but on the 30 unanswerable questions, the bot stops saying "I don't know" and starts guessing. Conciseness pressure killed the abstention behavior. Nobody noticed for two days because the aggregate hid it. Lesson: per-slice reporting (Chapter 6) and behavioral assertions (Chapter 7) — "abstention rate ≥ 90%" — become permanent suite members.

Friday — contamination and drift. Comparing against a public benchmark, their bot scores suspiciously high on questions that look memorized. Worse, the provider announces a model update; the following week's baseline shifts 2 points with zero code changes. Lesson: they pin the model version, keep a private held-out set, and stop citing the public benchmark as evidence for their specific claim (Chapter 5).

The following month — judgment. They run a human eval of 150 answers. Two raters agree on 73% of faithfulness judgments. Their LLM judge agrees with the human majority 71% of the time — essentially at the human ceiling. They ship the judge for nightly regression runs and reserve humans for release gates. Lesson: the ceiling was real, the judge was valid for this rubric and task, and knowing the boundary let them spend human effort where it mattered.

Every chapter of this book is a response to one morning of that week. The hardness is not a reason to despair — it is a specification for the discipline. Teams that internalize it stop arguing about single numbers and start discussing distributions, slices, and decision thresholds. That shift, more than any single technique, is what "good at eval" means.

Key takeaways

  • LLM evaluation is hard because outputs are nondeterministic, open-ended, prompt-sensitive, benchmark-contaminated, and ultimately judged by humans who disagree.
  • An eval score is a random variable, not a fact: repeat runs, report mean and spread, never a single run.
  • Quality is multi-dimensional; decompose it (factuality, completeness, relevance…) instead of relying on one aggregate number.
  • The prompt, decoding settings, and metric are part of the system under test — report all three.
  • Human agreement sets the ceiling for any automated metric; validate judges against it.

Metrics dashboard concept: gauges, charts and scorecards showing model accuracy, precision and recall

Chapter 2 — What to Measure: Task Metrics

The metric is a claim about what matters

Every metric smuggles in a definition of "good." Accuracy says: getting the exact right answer is all that matters, and all errors are equal. Token-level F1 says: partial credit counts, and overlap with the reference is what we reward. BLEU says: looking like the reference translation, n-gram by n-gram, is quality. None of these is true in general; each is useful in a specific situation. Choosing a metric is therefore not a technical detail — it is the moment you decide what your research values. Reviewers will judge you on it, so this chapter gives you the vocabulary to choose defensibly.

We will walk through the standard metrics one by one: what each computes, when it fits, when it misleads, and how to compute it in a few lines of code.

Exact match (EM) and accuracy

What it is. The prediction counts as correct only if it is identical to the reference answer (after light normalization — lowercasing, stripping articles and punctuation is the usual convention from the SQuAD evaluation script). Accuracy is the fraction of correct items.

When it fits. Closed tasks with a single canonical answer: multiple-choice questions, yes/no questions, math problems with numeric answers, code generation judged by unit tests (pass/fail per problem). If there is exactly one right answer, exact match is the honest metric.

When it misleads. Any open-ended generation. "Paris," "Paris, France," and "The capital of France is Paris" are all correct answers to "What is the capital of France?" — EM gives credit to exactly one of them. Using EM on open tasks massively understates real performance and, worse, rewards models that learn the format of your references rather than the content.

import re, string

def normalize(text: str) -> str:
    text = text.lower()
    text = "".join(ch for ch in text if ch not in set(string.punctuation))
    text = re.sub(r"\b(a|an|the)\b", " ", text)
    return " ".join(text.split())

def exact_match(pred: str, gold: str) -> int:
    return int(normalize(pred) == normalize(gold))

A practical tip: extract the answer span before comparing. If your model writes "The answer is 42 because…" and your gold is "42", compare the extracted "42", not the raw strings. Many published "EM" numbers are really "extract-then-EM" numbers — say so in your paper.

Token-level F1

What it is. Treat the prediction and reference as bags of tokens. Precision = fraction of predicted tokens that appear in the reference; recall = fraction of reference tokens that appear in the prediction. F1 is their harmonic mean. This gives partial credit: "Paris France" against "Paris" scores well though not perfectly.

When it fits. Short-answer extraction tasks (SQuAD-style reading comprehension) where answers are spans of text and partial overlap is meaningful. It is the standard companion to EM on such benchmarks.

When it misleads. It rewards word overlap, not meaning. "Not guilty" vs. "guilty" share a token and get nonzero F1 while meaning the opposite. On long outputs, F1 degenerates — a rambling answer that happens to contain the reference words scores well.

from collections import Counter

def token_f1(pred: str, gold: str) -> float:
    pred_tokens = normalize(pred).split()
    gold_tokens = normalize(gold).split()
    common = Counter(pred_tokens) & Counter(gold_tokens)
    overlap = sum(common.values())
    if overlap == 0:
        return 0.0
    precision = overlap / len(pred_tokens)
    recall = overlap / len(gold_tokens)
    return 2 * precision * recall / (precision + recall)

BLEU

What it is. The classic machine-translation metric (Papineni et al., 2002). It measures n-gram precision between the candidate translation and one or more references (for n = 1..4), with a brevity penalty that punishes candidates shorter than the reference. Scores range 0–100 (or 0–1).

When it fits. Machine translation and other tasks with tight paraphrase constraints, where good outputs really do share wording with references, and where you have multiple references per input. In MT, BLEU still correlates reasonably with human judgment at the system level (ranking systems), though poorly at the sentence level.

When it misleads. Almost everywhere else. BLEU punishes legitimate paraphrase, ignores semantics (a fluent wrong translation can outscore a clumsy right one), and is meaningless on creative or open-ended tasks. A BLEU of 35 vs. 37 on summaries tells you essentially nothing about which system humans prefer. If you report BLEU in a paper on summarization or dialogue in 2026, reviewers will ask why.

Use the sacreBLEU implementation, not your own — BLEU has notorious signature variants (tokenization, smoothing) that make scores incomparable across implementations:

from sacrebleu import corpus_bleu
# references: list of reference lists; hypotheses: list of strings
score = corpus_bleu(hypotheses, references)
print(score.score)  # 0-100

ROUGE

What it is. The summarization counterpart to BLEU (Lin, 2004). ROUGE-N measures n-gram recall (how much of the reference's content the summary captures); ROUGE-L measures longest common subsequence. Reported as F1 variants in most toolkits.

When it fits. Extractive or highly constrained summarization, as a cheap first-pass signal, and for comparability with older literature that reported it. ROUGE-L is the most commonly reported variant.

When it misleads. ROUGE rewards copying reference phrasing and is blind to factuality — a summary can score high on ROUGE while hallucinating, because hallucinations add n-grams the metric simply ignores (it measures overlap, not truth). It also cannot recognize a good abstractive summary that uses fresh wording. Never use ROUGE as your only summarization metric; pair it with a factuality check (Chapter 9 covers this for RAG, and the same logic applies).

from rouge_score import rouge_scorer
scorer = rouge_scorer.RougeScorer(["rouge1", "rouge2", "rougeL"], use_stemmer=True)
scores = scorer.score(reference_summary, model_summary)
print(scores["rougeL"].fmeasure)

pass@k (for code)

What it is. Generate k candidate solutions per problem; the problem counts as solved if any candidate passes all unit tests (Chen et al., 2021, HumanEval). The unbiased estimator corrects for the fact that you actually sample n ≥ k candidates and compute the probability that at least one of k would pass.

When it fits. Code generation with executable tests — the gold standard setting, because tests check behavior, not text similarity. pass@1 measures the "one shot" a user typically gets; pass@10 or pass@100 measures the model's broader capability.

When it misleads. It depends entirely on test quality: weak tests let wrong code pass. And it says nothing about code quality, efficiency, or readability. Also, generating 100 samples per problem is expensive — budget accordingly.

import math
def pass_at_k(n: int, c: int, k: int) -> float:
    """n samples, c correct, probability >=1 correct in k draws."""
    if n - c < k:
        return 1.0
    return 1.0 - math.comb(n - c, k) / math.comb(n, k)

Beyond strings: structured and behavioral metrics

Modern LLM applications need metrics that check behavior, not text:

  • Tool-call accuracy: did the agent call the right function with the right arguments? Score argument-by-argument (function name EM, then per-argument match).
  • Citation precision/recall (for RAG): of the claims the answer makes, what fraction is supported by retrieved passages (precision)? Of the key facts in the reference answer, what fraction appears (recall)?
  • Task success rate (for agents): did the environment reach the goal state? Binary, but honest — a web agent either booked the flight or did not.
  • Refusal/hedging rates for safety evals; toxicity scores from classifiers for open generation.

The pattern is the same each time: find the observable consequence of a good output and measure that, not the words.

A decision guide

Task shape Start with Add
Multiple choice / classification Accuracy Per-class breakdown, calibration
Short extractive answers EM + token F1 Human spot-check of failures
Machine translation sacreBLEU / chrF Human adequacy rating on a sample
Summarization ROUGE-L (legacy comparability) Factuality check, human rubric
Code with tests pass@k Test-strength audit
Open-ended QA / chat Human rubric or validated LLM judge Dimension breakdown
RAG Citation precision/recall, answer EM/F1 Retrieval hit rate
Agents Task success rate Step efficiency, tool-call accuracy

Two rules of thumb. First, never report a single metric on an open task — report a small panel (2–4 metrics) so a weakness in one shows up in another. Second, validate the metric against humans on your data at least once: take 100 outputs, score them with your metric and with a human rubric, and compute the correlation. If they disagree, trust the humans and fix the metric, not the other way round. Chapter 4 shows how to do this validation for LLM judges; the same procedure applies to any automatic metric.

For your research: When you pick metrics for your paper, write one sentence per metric: "We use X because [property of the task it captures] and acknowledge it misses [limitation]." Reviewers rarely object to an imperfect metric; they object to an undefended one. This sentence belongs in your paper's evaluation section, nearly verbatim.

One answer, four scores: a worked comparison

Watch what happens when four metrics score the same model output. Question: "What causes ocean tides?"

Reference: "Tides are caused by the gravitational pull of the Moon and the Sun on Earth's oceans."

Candidate: "The Moon's gravity pulls on ocean water, creating bulges that we experience as tides; the Sun adds a smaller effect."

A human reader calls this an excellent answer — accurate, complete, well-phrased. Now the metrics:

  • Exact match: 0. The strings differ, so EM reports total failure. On an open question, EM is not strict — it is simply wrong as a measure.
  • Token F1: ~0.45. Shared content words ("moon," "sun," "ocean," "tides") give middling overlap while the phrasing differences ("gravity" vs. "gravitational") score nothing. The number says "half right" about an answer that is fully right.
  • ROUGE-L: ~0.5. The longest common subsequence covers about half the reference. Again: a good answer, a mediocre number.
  • Validated LLM judge (faithfulness 3/3, completeness 3/3): the correct assessment. The judge reads meaning, not strings.

The lesson is not "judges good, strings bad" — it is that the metric must match the answer space. Where the answer space is tight (a number, a name, a code test), string and behavioral metrics are precise and cheap. Where it is open, they systematically under-measure quality, and optimizing against them teaches the model to mimic reference phrasing instead of answering well. This is Goodhart's law in eval clothing: when a flawed measure becomes the target, it stops being a good measure — your prompt iterations will chase ROUGE points while human-perceived quality stalls or drops.

Two more metrics worth knowing for the middle ground:

  • chrF measures character n-gram overlap instead of word overlap. It is far more forgiving of morphological variation ("gravitational" vs. "gravity") and is the recommended default for morphologically rich languages, where word-level BLEU collapses.
  • Semantic similarity metrics (BERTScore, BLEURT) compare contextual embeddings rather than exact tokens, so paraphrases score highly. They correlate better with human judgment than BLEU/ROUGE on open tasks — but they are still similarity-to-reference measures: they cannot detect a fluent, on-topic hallucination that shares vocabulary with the reference, and they inherit biases from their underlying embedding model. Use them as a complement to factuality checks, never as a replacement.

A practical scoring recipe for open QA that balances cost and honesty: (1) EM/F1 as a cheap first pass to catch clear wins; (2) a semantic metric to rank candidates during development; (3) a validated judge or human rubric on a sample for the reported numbers, focused on factuality. Each layer corrects the blindness of the one below it.

Key takeaways

  • A metric is a claim about what "good" means: choose it deliberately and defend it in one sentence.
  • EM/accuracy for closed answers; token F1 for short spans; BLEU for translation (system-level); ROUGE-L for legacy summarization comparability; pass@k for code with tests.
  • N-gram metrics punish paraphrase and are blind to factuality — never let them stand alone on open tasks.
  • For RAG and agents, measure behavior (citations, tool calls, task success), not word overlap.
  • Validate every automatic metric against human judgment on a sample of your own data before trusting it.

Chapter 3 — Human Evaluation Done Right

Why humans are still the gold standard

Every automatic metric is a proxy — a cheap, repeatable stand-in for what you actually care about: would a competent person judge this output good? Proxies drift. BLEU stops tracking human preference on new domains; LLM judges inherit biases (Chapter 4). The human evaluation is the measurement everything else is calibrated against. If your human eval is sloppy, your entire results section rests on sand, no matter how fancy the automation above it.

The good news: running a solid human eval is a learnable craft, not a dark art. It has four parts — a rubric, raters, blinding, and agreement measurement — and this chapter walks through each with templates you can reuse.

The rubric: writing down what "good" means

A rubric turns the vague question "is this good?" into specific, answerable questions. Bad rubric: "Rate the quality of this summary from 1 to 5." Different raters will interpret "quality" differently — one rewards fluency, another punishes a missing fact — and your agreement will be terrible.

Good rubrics share three properties:

  1. Named dimensions. Break quality into 3–5 named aspects. For a summarization study: Factual consistency (does every claim in the summary appear in the source?), Coverage (are the source's main points present?), Conciseness (no filler?), Fluency (grammatical, readable?).
  2. Anchored scales. Each point on the scale gets a concrete description, not just a number. "3 = most main points covered, one minor point missing" beats "3 = average."
  3. Decision rules for edge cases. Write down, in advance, how to score the cases you know will be ambiguous: "If the summary adds background knowledge not in the source but factually true, mark Factual consistency as 2 (unsupported), not 1 (contradicted)."

Here is a worked example — a rubric for evaluating answers from a RAG question-answering system. Copy and adapt it:

Dimension 1 — Faithfulness (1–3). 3: Every factual claim in the answer is directly supported by the retrieved passages. 2: All central claims supported; minor details unsupported but plausible. 1: At least one central claim contradicted by or unsupported by the passages.

Dimension 2 — Completeness (1–3). 3: Answers every part of the question. 2: Answers the main part; a sub-question is missing or vague. 1: Misses the main point of the question.

Dimension 3 — Usefulness (1–3). 3: A user could act on this answer directly. 2: Useful but requires follow-up. 1: Not useful as written.

Edge rules. (a) If the answer says "I don't know" when the passages contain the answer, Completeness = 1. (b) Score what is written, not what you think the model meant. (c) Ignore fluency unless it blocks understanding.

Notice the scale is 1–3, not 1–5 or 1–10. Fewer points with clear anchors beat many points with vague ones: raters can reliably distinguish three levels, and your agreement statistics will thank you. Reserve wider scales for dimensions where raters genuinely need the granularity.

Pilot the rubric. Before the real study, have two raters score 20 outputs, then sit down (together or over notes) and discuss every disagreement. You will discover ambiguous anchors and missing edge rules. Revise the rubric, re-pilot on a fresh 20, and only then launch. Two pilot rounds is normal; skipping them is the most common cause of failed human evals.

Raters: who, how many, how trained

  • Who. For research evals, raters should be competent in the domain — graduate students in the area, not random crowdworkers, for anything technical. Crowd platforms are fine for general-knowledge tasks if you screen with qualification questions and attention checks, but budget for 15–25% of responses being unusable.
  • How many raters per item. Two is the minimum for an agreement statistic; three lets you take majority votes and is the norm in published work. More than three has diminishing returns unless the task is genuinely ambiguous.
  • Training. Walk raters through the rubric with 5–10 worked examples before they start, including at least two examples you expect them to get wrong. A 30-minute calibration session roughly doubles the value of the hours they spend rating.

Pay raters fairly and report the arrangement in your paper (reviewers increasingly ask). Report the number of raters, their background ("three CS graduate students"), and approximate effort ("each rated 150 items, ~8 hours").

Blinding and randomization

Raters must not know which system produced the output they are scoring. If they can tell "this fluent one is GPT-4, this clumsy one is my baseline," their scores will reflect brand expectations, not quality. Practical steps:

  • Strip system identifiers from the rating interface — no model names, no formatting tells (normalize whitespace and markdown).
  • Randomize item order per rater, so position effects (fatigue, contrast with the previous item) average out.
  • Interleave systems: a rater should see outputs from all systems mixed together, never all of system A then all of system B.
  • If the study compares your method against baselines, consider also blinding yourself: have a collaborator hold the system→ID mapping until scoring is done. This is cheap insurance against unconscious bias in how you handle edge cases.

Measuring agreement

If your raters do not agree with each other, your scores are noise. Report agreement for every human eval — reviewers expect it, and it is the single strongest signal that your study was run carefully.

  • For categorical judgments (e.g., correct/incorrect, or 1–3 scales treated as categories): Cohen's kappa (two raters) or Fleiss' kappa (three or more). Kappa corrects for chance agreement. Rough guide: below 0.4 is poor, 0.4–0.6 moderate, 0.6–0.8 substantial, above 0.8 near-perfect. For a 3-point anchored rubric, aim for ≥ 0.6 after piloting.
  • For continuous scores: the intraclass correlation coefficient (ICC) or simply Pearson/Spearman correlation between raters' scores.
  • Always report raw percent agreement alongside kappa — kappa can look bad on skewed data even when raters agree 95% of the time (the "kappa paradox"), and raw agreement keeps the picture honest.
from sklearn.metrics import cohen_kappa_score
# rater_a, rater_b: lists of integer labels, one per item
kappa = cohen_kappa_score(rater_a, rater_b)
print(f"Cohen's kappa: {kappa:.2f}")

What to do with disagreements. Decide the rule in advance and state it: majority vote (3 raters), adjudication by a fourth expert, or averaging (for continuous scales). Never silently drop disagreeing items — that biases your sample toward easy cases.

How many items to rate

Human rating is expensive, so size the study to the claim. A useful heuristic: to detect a 10-percentage-point difference between two systems at 80% power, you need roughly 150–200 items per system (Chapter 8 explains the statistics). For a paper comparing two systems on one dimension, 200 items × 3 raters is a solid, defensible study. For early development, 50 items × 2 raters is enough to catch big problems — just do not publish strong claims off it.

Stratify your sample: if your eval set has categories (question types, difficulty levels, domains), sample proportionally from each so the human study reflects the full distribution, not just the easy majority.

For your research: Your human eval section should answer five questions a reviewer will ask: (1) What exactly did raters score? — point to the rubric, ideally in an appendix. (2) Who were the raters and how were they trained? (3) Were they blinded to the system? (4) How much did they agree? — report kappa/ICC. (5) How were disagreements resolved? Write your study plan as answers to these five questions before you collect a single rating, and the write-up becomes trivial.

Logistics: running the study week by week

A human eval is a small project. Here is a realistic five-week schedule for a 200-item, 3-rater study comparing two systems:

Week 1 — Rubric and pilot. Draft the rubric (dimensions, anchors, edge rules). Recruit two pilot raters — ideally the people who will do the real rating. They score 20 items; you compute preliminary agreement and hold a 1-hour disagreement review. Expect to rewrite at least a third of your anchors. Re-pilot on 20 fresh items. You are done when kappa clears 0.6 on every dimension.

Week 2 — Rater training and interface. Finalize the rating interface. The layout matters: show the question, the source material (if any), and the answer side by side; one dimension per screen or clearly separated sections; the full anchor text visible next to each scale point (raters should never have to recall what "2" means). Build in attention checks: 5–10% of items with obvious correct ratings (e.g., an answer that is clearly off-topic should score 1 on relevance). Raters who fail attention checks get their work reviewed or discarded. Hold the 30-minute calibration session with worked examples.

Weeks 3–4 — Rating with live monitoring. Do not wait until the end to check quality. After every 50 items, compute running agreement per rater pair and per dimension. Common problems and fixes: - One rater diverges from the other two: usually a misunderstood anchor — re-calibrate that rater with targeted examples, and consider re-rating their completed items on the affected dimension. - Agreement decays over time: fatigue. Cap sessions at ~90 minutes, enforce breaks, and randomize item order so fatigue hits all systems equally. - One dimension has low agreement everywhere: the dimension is ill-defined. Either rewrite its anchors mid-study (and re-rate) or drop the dimension and say so honestly.

Week 5 — Adjudication and analysis. Apply your pre-registered disagreement rule (majority vote, expert adjudication, or averaging). Then analyze: per-system means with CIs, per-dimension breakdowns, and a qualitative pass over the items where all raters agreed the best system still failed — those are your hardest cases and often the seed of your next research idea.

Budgeting. A trained rater scores roughly 15–25 complex items per hour. For 200 items × 3 raters at 20/hour, that is 30 rater-hours plus your ~10 hours of management. At fair graduate-student rates, budget accordingly — and report the cost in your paper's appendix. Reviewers increasingly view "we used 3 graduate raters for ~40 hours total" as a credibility signal, not filler.

One more safeguard: rater independence. Raters should work alone, without discussing items mid-study — discussion converges their judgments artificially and inflates agreement. The disagreement review happens in pilots; during the real study, questions go to you, and your answers go to all raters identically (and get added to the edge-rules document).

When raters struggle: difficult cases and how to handle them

Even well-run studies hit situations the rubric didn't anticipate. Plan for these four:

Genuinely ambiguous items. Sometimes two reasonable experts permanently disagree — the question is vague, or two answers are defensible. Don't force false consensus: mark such items with a "disputed" flag during rating, adjudicate what you can, and report the disputed rate. A 5–8% disputed rate is normal on hard tasks; above 15%, your rubric or your items need work, not your raters. In the paper, one sentence covers it: "11 of 200 items (5.5%) were flagged as disputed after adjudication and are analyzed separately in Appendix C."

Rater gaming and speeding. On paid platforms especially, some raters click through as fast as possible. Defenses in layers: attention checks (Week 2), minimum time per item (calibrated from your pilot — e.g., flag items completed in under 20% of the median time), and agreement monitoring (a rater whose kappa with everyone else is near zero is either a genius or not reading). Remove and re-rate affected items; don't just average them in.

Sensitive or disturbing content. If outputs may contain toxic, graphic, or upsetting text, warn raters in advance, let them opt out of categories, limit exposure per session, and provide a way to skip-and-flag. This is both an ethical obligation and a data-quality measure — distressed raters rate badly. State your protocol in the paper; reviewers in 2026 expect it.

Multilingual rating. If your eval spans languages, don't assume one rater pool covers all: recruit per language, and validate that translated rubrics carry the same meaning (back-translation check: translate the rubric to the target language and back, and confirm the anchors still match). Report agreement per language — it often varies, and the variation itself is a finding about where your system is weakest.

The golden rule of study management: every problem you discover mid-study is a problem your pilot was supposed to catch. After each study, spend 30 minutes writing down what surprised you and add one question to your pilot checklist. Three studies in, your pilots will be catching 90% of issues — which is what makes the fourth study cheap.

Key takeaways

  • The rubric is the measurement instrument: named dimensions, anchored 1–3 scales, written edge-case rules, piloted twice.
  • Use 2–3 trained, domain-competent raters per item; calibrate them with worked examples before they start.
  • Blind raters to the system, randomize order, interleave systems — and consider blinding yourself to the mapping.
  • Report agreement (kappa for categories, ICC/correlation for scores) plus raw percent agreement, and pre-register your disagreement-resolution rule.
  • Size the study to the claim: ~200 items × 3 raters for a publishable two-system comparison; 50 × 2 for development signal.

Human evaluation versus AI-as-judge: a human reviewer and a robot reviewer side by side checking test answers

Chapter 4 — LLM-as-Judge: Promise, Pitfalls, and Validation

The tempting shortcut

Human evaluation is slow and expensive. The obvious shortcut: ask a strong language model to do the judging. Give it the rubric, the question, the two candidate answers, and ask which is better — or ask it to score each answer 1–5. This is LLM-as-judge, and it has become one of the most used and most abused tools in AI research. Used well, it lets you evaluate thousands of outputs overnight for dollars. Used badly, it manufactures the result you wanted and dresses it in the authority of automation.

This chapter teaches the disciplined version: how to write a judge prompt that works, how to measure whether your judge is any good, and the known biases that can silently corrupt your results.

Why it works at all

A strong instruction-following model can apply a rubric with surprising consistency — often matching the agreement level of trained human raters on well-specified dimensions. The key phrase is "well-specified": LLM judges are excellent at applying clear criteria and terrible at inventing them. Your rubric from Chapter 3 does the hard conceptual work; the judge just executes it at scale.

The standard setups:

  • Single-answer scoring. The judge sees one answer and the rubric, and returns scores per dimension. Simple, and the scores are directly comparable across systems.
  • Pairwise comparison. The judge sees two answers and picks the winner (or declares a tie). Humans find pairwise judgments easier and more consistent than absolute scoring, and the same seems true of models. Aggregate many pairwise votes into rankings with a Bradley–Terry model or simple win rates. MT-Bench popularized this format (Zheng et al., 2023).
  • Reference-guided grading. The judge sees a reference answer alongside the candidate, and scores the candidate against it. This anchors the judge and substantially improves agreement with humans on tasks with definable correct content (math, factual QA).

Writing a judge prompt that works

A judge prompt has five components. Here is a template for single-answer scoring — adapt the bracketed parts:

You are an impartial evaluator of [TASK, e.g., answers to factual questions].

## Question
{question}

## Reference answer (for guidance; the candidate need not match its wording)
{reference}

## Candidate answer
{candidate}

## Rubric
[Paste your Chapter-3 rubric here, with anchored scales.]

## Instructions
1. Quote the specific part of the candidate answer relevant to each dimension.
2. Score each dimension according to the rubric.
3. Return ONLY a JSON object: {"faithfulness": <1-3>, "completeness": <1-3>,
   "usefulness": <1-3>, "rationale": "<one sentence per dimension>"}.

Four details matter more than people expect:

  1. Demand evidence first. Asking the judge to quote the relevant span before scoring forces it to ground the score in the text. This single trick reduces hallucinatory scoring more than any other.
  2. Require structured output. JSON (or another parseable format) lets you validate and aggregate automatically. Set temperature low (0–0.2) — you want the judge's most likely assessment, not creativity.
  3. Include the rubric verbatim. Do not summarize it or trust the judge to "know what good means." Paste the full anchored scales.
  4. Randomize position in pairwise judging. Present answer A and answer B in random order per item and record which was which, because judges favor the first-presented answer (position bias — see below).

The validation you cannot skip

An LLM judge is a measurement instrument, and instruments must be calibrated. The procedure:

  1. Take a sample of 100–200 outputs from your eval (covering all systems under comparison).
  2. Have humans score them with the same rubric (Chapter 3 — this is the one human-eval cost you cannot avoid).
  3. Run your judge on the same items.
  4. Compute agreement between judge and human: accuracy on pairwise preferences, or correlation/kappa on scores. Compare it to human–human agreement on the same items.

The decision rule: if judge–human agreement is close to human–human agreement, the judge is a valid substitute for this task and rubric. If it lags badly, fix the rubric or the prompt — or accept that this dimension needs humans.

Report this validation in your paper. One sentence suffices: "On a 150-item human-labeled subset, our judge agreed with the majority human label on 81% of pairwise preferences, vs. 84% human–human agreement." That sentence is the difference between a credible automated eval and hand-waving.

Validate per dimension. Judges are often good at fluency and bad at factuality — the dimension you care about most. A single aggregate agreement number can hide this. Validate each rubric dimension separately.

Known biases: the pitfalls

LLM judges have documented, repeatable biases. Design around them:

  • Position bias. In pairwise comparisons, judges favor the answer presented first (or sometimes second — it varies by model). Mitigation: randomize order per item and average both orders on a subset to check.
  • Length bias. Judges prefer longer, more detailed answers even when the extra content adds nothing — or is wrong. Mitigation: instruct the judge explicitly ("do not reward length; penalize unsupported detail"), and check that your winning system is not just the most verbose one.
  • Self-preference. A model judging outputs tends to prefer outputs written in its own style — including outputs from itself. If your judge is Model X and one candidate is Model X, expect inflated scores. Mitigation: use a judge from a different model family than any candidate, or at minimum disclose the overlap and test for it.
  • Style over substance. Judges reward confident tone, bullet points, and formatting polish. A wrong answer in a beautiful layout can outscore a right answer in plain text. Mitigation: reference-guided grading, and a rubric dimension that separates presentation from content.
  • Leniency drift. Some judges rarely use the bottom of the scale, compressing all scores into 4–5. Mitigation: anchored scales with vivid descriptions of poor performance, and few-shot examples of low scores in the prompt.

None of these biases is a reason to abandon LLM judges. They are reasons to validate, disclose, and design around them — exactly as you would with any imperfect instrument.

Cost and practicalities

A pairwise judge call on long outputs can cost several cents; 10,000 comparisons adds up. Reduce cost by judging a stratified sample rather than everything, caching judge outputs (keyed by the exact prompt + candidate text hash), and using single-answer scoring (one call per answer) instead of pairwise (one call per pair) during development, switching to pairwise for the final numbers.

Version your judge. The judge model, its prompt, and its decoding settings are part of your method. If the provider updates the judge model mid-project, re-run the validation. A judge whose behavior changed silently is worse than no judge at all.

For your research: Treat "LLM judge" as a claim of the form "this automated procedure approximates human judgment on dimensions D1..Dn for this task." Every element of that claim needs evidence: the rubric (validity), the human-labeled validation set (calibration), the agreement numbers (reliability), and the bias checks (threats). When you can state all four in your paper, the judge stops being a shortcut and becomes a method.

Case study: validating a judge end-to-end

A concrete validation, with numbers, so you can see what "good enough" looks like — and what failure looks like.

Setup. A team evaluates summary quality on three dimensions: faithfulness, coverage, fluency. They sample 150 summaries from their dev set, have two trained raters score each (Chapter 3 rubric), and run their LLM judge on the same items.

Results.

Dimension Human–human agreement Judge–human agreement Verdict
Fluency 88% 87% Judge valid — use freely
Coverage 81% 79% Judge valid — use, spot-check
Faithfulness 82% 74% Judge not valid — needs humans

The judge matches humans on fluency and coverage but lags 8 points on faithfulness — the dimension the paper's claim depends on. Error analysis shows the judge misses subtle contradictions (it checks whether claims sound supported, not whether they are). The team makes a disciplined call: use the judge for fluency and coverage at scale, and pay for human rating on faithfulness for the 200-item test set. Their results section reports all three numbers and the split decision. A reviewer reading this trusts the paper more, not less, because the team demonstrated they know where their instrument is blind.

Bias checks they ran on the same 150 items:

  • Position bias: on 50 pairwise items, they ran both orders. The winner flipped in 12% of pairs — above their 5% tolerance. Fix: they switched the final eval to single-answer scoring (no positions to bias) and reported the check.
  • Length bias: correlation between judge score and answer length was r = 0.31 on coverage. They added "do not reward length; a concise complete answer outscores a verbose partial one" to the prompt; correlation dropped to r = 0.12. Reported both.
  • Self-preference: the judge was from model family X; one candidate was also family X. On a 40-item subset they re-judged with a family-Y judge: family-X candidates scored 0.2 points higher under the X judge. Small but real — disclosed in a footnote, and the headline comparison used the Y judge.

When not to use LLM judges at all. Judges are the wrong tool when (a) humans themselves disagree substantially on the task (the ceiling is too low — the judge is just adding confident noise); (b) the judgment requires expertise the judge lacks and you cannot verify (e.g., novel mathematical proofs); (c) stakes are high and errors are costly (medical or legal deployment decisions need humans); or (d) the setting is adversarial (a judge can be prompt-injected by the very outputs it scores — never let untrusted model output share a context window with judge instructions without strict delimiting and output validation).

The through-line: a judge is a delegated judgment, and delegation without verification is abdication. The validation study is what turns "we used an LLM to score outputs" from a red flag into a method.

Building a judge harness that won't embarrass you

A judge prompt is 20% of the work; the harness around it is the other 80%. Here's what production judge harnesses include:

Caching. Judge calls cost money and time. Cache every judge response keyed by a hash of (judge model version + judge prompt + item inputs). Re-running an eval should cost zero judge calls for unchanged items. A simple SQLite or JSONL cache works:

import hashlib, json, sqlite3
def judge_key(model, prompt, item):
    h = hashlib.sha256()
    h.update(json.dumps([model, prompt, item], sort_keys=True).encode())
    return h.hexdigest()

Robust parsing. Judges don't always return clean JSON — they add preamble, truncate, or "helpfully" reformat. Parse defensively: extract the first {...} block, retry once with "return ONLY the JSON object" appended, and log unparseable outputs for inspection rather than silently dropping them. Track your parse-failure rate; above ~2%, fix the prompt.

Retries and rate limits. Wrap calls with exponential backoff, jitter, and a max-retry budget. Run judging jobs idempotently so a crashed run resumes instead of restarting — at 10,000 items, a non-resumable job that dies at 90% is a painful lesson.

Cost tracking. Log tokens in/out per call and maintain a running cost estimate. A pairwise judge on long documents can cost $0.02–0.10 per comparison; 20,000 comparisons is real money. The harness should print projected cost before launching a full run, and support a --sample N flag for stratified pilot runs.

Separation of concerns. Keep four files distinct: the judge prompt (versioned text), the runner (API calls, caching, parsing), the aggregation (means, CIs, bias checks), and the validation (comparison against human labels). When the provider updates the judge model, you change one string (the model version) and re-run validation — not a tangle of edits across a notebook. This separation is also what makes the judge setup describable in a paper: "judge prompt v3, runner v1.2, validated 2026-04-10" is a reproducible method; a notebook with 40 cells is not.

Key takeaways

  • LLM judges work by applying a clear rubric at scale — the rubric quality determines everything.
  • Judge prompt anatomy: task, question, reference, candidate, full rubric, evidence-before-score, structured JSON output, low temperature.
  • Validate every judge against 100–200 human-labeled items; require judge–human agreement near human–human agreement, per dimension.
  • Design around known biases: randomize position, penalize verbosity, avoid judging your own model's outputs with itself, separate style from substance.
  • Version the judge model and prompt; re-validate if either changes.

Chapter 5 — Standard Benchmarks: MMLU, HumanEval, and Friends

What benchmarks are for

A benchmark is a fixed test set plus a scoring rule, shared across the field so that results are comparable. Benchmarks exist to answer "how capable is this model, in general?" — and they are genuinely useful for that: they compress thousands of judgments into one number, they let you compare against published results without re-running everyone's code, and they reveal broad strengths and weaknesses (a model great at code but poor at multilingual QA tells you something real).

They are also the most misread numbers in AI. This chapter teaches you to read them like a reviewer: what each major benchmark measures, what its score does and does not prove, and the traps — contamination, saturation, prompt-sensitivity — that make leaderboard numbers slippery.

The major benchmarks, honestly described

MMLU (Massive Multitask Language Understanding) — Hendrycks et al., 2021. ~16,000 multiple-choice questions across 57 subjects, from high-school math to professional law and medicine. The headline test of "broad knowledge and reasoning." What it proves: the model has absorbed a vast amount of factual and conceptual knowledge and can apply it in a multiple-choice format. What it does not prove: that the model reasons reliably (multiple choice lets models exploit surface cues and elimination strategies), that it would perform similarly on open-ended versions of the same questions, or that its knowledge is current. Caveats: it is the most contamination-suspected benchmark in existence — its questions are all over the training web. Treat small MMLU differences between models (1–3 points) as noise, not signal.

HumanEval — Chen et al., 2021. 164 hand-written Python programming problems, each with a function signature, docstring, and hidden unit tests; scored with pass@k. What it proves: the model can write short, self-contained functions that pass tests — a real, behavioral capability. What it does not prove: ability to work in large codebases, use unfamiliar APIs, debug, or write maintainable code. 164 problems is a small sample; the confidence interval on a pass@1 difference is wide (Chapter 8). Caveats: widely memorized; many models have effectively seen the problems. Also, docstring quality varies — some "failures" are ambiguous specifications, not model errors.

GSM8K — ~8,500 grade-school math word problems with numeric answers. Tests multi-step arithmetic reasoning; chain-of-thought prompting famously lifted scores dramatically (Wei et al., 2022). What it proves: step-by-step quantitative reasoning on well-specified problems. Does not prove: real-world mathematical problem-solving, where the hard part is formalizing the problem. Contamination concerns are serious here too.

SQuAD / TriviaQA — reading-comprehension and factual-QA benchmarks (Rajpurkar et al., 2016; Joshi et al., 2017). What they prove: extracting and recalling factual answers. Caveat: largely saturated — top models score near human performance, so they no longer discriminate between strong models. Useful as sanity checks, not as differentiators.

GLUE / SuperGLUE — collections of natural-language-understanding tasks (Wang et al., 2019). Historically central; now saturated and mostly of historical interest for LLM evaluation, though the idea — a suite of diverse tasks with one aggregate — lives on in HELM.

HELM (Holistic Evaluation of Language Models) — Liang et al., 2022. Not one benchmark but a framework evaluating many models across many scenarios with standardized metrics, emphasizing transparency about what was measured. What it proves: breadth across scenarios with honest reporting. Valuable as a model for how to report evals, not just as a score source.

MT-Bench — 80 multi-turn questions across 8 categories, scored by LLM judges with pairwise comparisons (Zheng et al., 2023). What it proves: conversational and instruction-following quality as judged by a strong model. Does not prove: anything the judge is blind to (see Chapter 4's biases — MT-Bench scores inherit them).

Reading a benchmark score critically

When you see "Model X scores 86.4 on MMLU," run through this checklist:

  1. Which split and which prompting? 5-shot? 0-shot? Chain-of-thought? Scores move several points with prompting, and papers do not always use the same setup. Compare only numbers produced under identical protocols — or better, re-run the benchmark yourself with a harness like EleutherAI's lm-evaluation-harness, which standardizes prompting.
  2. Is the difference meaningful? On MMLU's ~16k questions, a 1-point gap is roughly 160 questions — but with prompt sensitivity and possible contamination, gaps under ~2 points between models are weak evidence. On HumanEval's 164 problems, a 5-point gap is about 8 problems — well within noise. Chapter 8 shows how to compute this properly.
  3. Could it be contamination? Ask: was this benchmark public before the model's training cutoff? If yes, discount memorization-prone formats (multiple choice, short factual answers) heavily.
  4. Is the benchmark saturated? If the top 10 models all score 90–92, the benchmark has stopped measuring differences between them. A new state-of-the-art claim on a saturated benchmark is unimpressive; reviewers know it.
  5. Does the benchmark match your task? A high MMLU score does not predict good performance on your RAG system over your documents. Benchmarks measure general capability; your eval set (Chapter 6) measures fitness for your purpose. Papers need both: benchmarks for context, custom evals for the actual claim.

Benchmarks in your paper: how to use them well

  • Use them for positioning, not as your main result. "Our model reaches 71.2 on MMLU (5-shot), comparable to published numbers for X and Y" is good context. "Our method is better because MMLU went from 70.1 to 71.8" is a weak main claim — the gap is small, the benchmark may be contaminated, and it says nothing about your actual task.
  • Re-run baselines yourself rather than copying numbers from other papers, whenever feasible. Prompting and harness differences make copied numbers incomparable, and re-running protects you from accidentally comparing your 5-shot number to someone else's 0-shot number.
  • Report the harness and settings. "Evaluated with lm-evaluation-harness v0.4, 5-shot, temperature 0" lets anyone reproduce you. A bare number does not.
  • Prefer a small panel of diverse benchmarks over one headline number: one knowledge (MMLU), one reasoning (GSM8K or similar), one code (HumanEval), one aligned with your domain. Diversity of evidence beats depth on a single leaderboard.

For your research: Before citing a benchmark score as evidence, write the sentence "This score shows ___ about my system" and check that the blank is something the benchmark actually measures (per the descriptions above), not something you wish it measured. If you cannot fill the blank honestly, the score does not belong in your results — it belongs in an appendix, or nowhere.

Contamination: how it happens, how to detect it

Contamination deserves a deeper look because it is the quietest way a benchmark score becomes meaningless — and because "the model may have seen the test data" is now a default reviewer question.

How test data leaks into training. The mechanisms are mundane: (1) Benchmarks are published on the web; web-scale crawls ingest them. (2) Researchers discuss benchmark items in papers, blog posts, and GitHub issues — all crawled. (3) Few-shot examples from the test split get pasted into prompts that are logged and later trained on. (4) Fine-tuning datasets assembled from the web include benchmark items verbatim. None of this requires bad intent. A useful rule: if a benchmark was public before a model's training cutoff, assume partial contamination and discount memorization-friendly formats (multiple choice, short factual answers) accordingly.

Detection techniques. None is definitive, but together they build a picture: - Canary strings. Some benchmarks (e.g., BIG-bench) embed a unique "canary" GUID string with instructions not to train on it. Searching model outputs or training-data disclosures for the canary is a direct test — though it only works if the benchmark included one. - Ordering and perturbation tests. If a model scores dramatically higher on the original benchmark items than on paraphrased versions testing the same knowledge, memorization is likely: true understanding survives paraphrase, memorized strings do not. - Cutoff analysis. Compare performance on questions about events before vs. after the model's training cutoff. A model that "knows" 2021 benchmark questions far better than equally difficult 2024 questions it could not have memorized is showing its training data, not its reasoning. - Log-probability probes. For open-weight models, test items the model assigns suspiciously high likelihood to (relative to paraphrases) are candidates for memorized strings.

What to do about it. For high-stakes claims, don't rely on public benchmarks alone: keep a private holdout — items written after the model's cutoff, never published — and report those numbers alongside. Support dynamic benchmarks that regenerate items (new math problems, fresh questions) so memorization has a short half-life. And when you build your own eval set (Chapter 6), treat non-publication as a feature: an unpublished, private test set is contamination-proof by construction, which is a genuine methodological advantage worth stating in your paper.

Running benchmarks reproducibly. If you do report benchmark numbers, produce them yourself with a standard harness rather than copying them. The EleutherAI lm-evaluation-harness is the community standard:

pip install lm-eval
lm_eval --model hf --model_args pretrained=<model-id> \
  --tasks mmlu,gsm8k,humaneval --num_fewshot 5 --batch_size 8 \
  --output_path results/

Report the harness version, the exact task names, and the few-shot count. Two papers reporting "MMLU 5-shot" from different harness versions can differ by a point from implementation details alone — pinning the harness is what makes your number comparable to the next person's.

Reading a model release: a benchmark report card, critiqued

Imagine a model release blog post with this table:

Benchmark NewModel-X Competitor-Y
MMLU (5-shot) 89.2 87.9
HumanEval (pass@1) 92.1 88.4
GSM8K 95.3 94.1

Impressive? Run the Chapter 5 checklist on it:

  1. Which prompting? The post says "5-shot" for MMLU but doesn't specify for HumanEval or GSM8K — and doesn't say whether chain-of-thought was used for GSM8K (which moves scores ~10 points). Without that, the numbers are decorative.
  2. Is the difference meaningful? MMLU gap: 1.3 points on ~16k questions ≈ 200 questions — but well within prompt-sensitivity and contamination noise. GSM8K gap: 1.2 points. Neither clears a skepticism threshold. The HumanEval gap (3.7 points on 164 problems ≈ 6 problems) is more interesting but still small-sample.
  3. Contamination? All three benchmarks predate every current model's training data. The vendor doesn't mention decontamination analysis. Discount accordingly.
  4. Saturation? GSM8K at 95.3 is near-ceiling — the benchmark has stopped discriminating. A "win" here is the least informative of the three.
  5. Match to your task? None of these measures your RAG faithfulness, your agent's tool accuracy, or your summarizer's factuality. For a research claim about any of those, this table is context at best.

The corrected version you'd want to see: harness name and version, exact prompting per benchmark, decontamination statement, CIs or at least sample sizes, and — most importantly — a custom eval for whatever capability the release actually claims ("our model follows complex instructions better" needs an instruction-following eval, not MMLU). When you write your own paper's related-work or baseline sections, hold others' tables to this standard — and preemptively hold your own to it.

When benchmarks disagree with each other

A common confusion: Model A wins on MMLU but loses on HumanEval; which is "better"? The answer is that the question is malformed — the benchmarks measure different capabilities, and models have different strengths. A code-specialized model losing MMLU to a generalist while winning HumanEval is exactly what you'd expect, not a contradiction. When benchmarks disagree, don't average them into a meaningless composite; instead, ask which capability your task needs and weight that benchmark accordingly. And when similar benchmarks disagree (two QA benchmarks ranking models differently), suspect prompt-format differences or contamination asymmetry before concluding anything about the models — re-run both under your harness with identical prompting before theorizing.

Key takeaways

  • Benchmarks answer "how capable is this model in general?" — useful for positioning, weak as a main research claim.
  • Know what each measures: MMLU (broad MCQA knowledge), HumanEval (short function synthesis, behavioral), GSM8K (math word problems), MT-Bench (judge-scored conversation).
  • Discount small gaps, suspect contamination on public benchmarks, ignore saturated leaderboards, and match the benchmark to your task.
  • Re-run baselines with a standard harness and identical prompting; report harness version and settings.
  • Your custom eval set (Chapter 6) carries your paper's main claim; benchmarks provide context.

Evaluation pipeline: test dataset feeds into the language model, then a scoring step, then a metrics report

Chapter 6 — Building Your Own Eval Set

Why you need one

Standard benchmarks measure general capability. Your research makes a specific claim — "my retrieval method improves factuality on medical QA," "my prompting technique reduces hallucinations in summarization" — and no public benchmark tests exactly that. Reviewers know this, which is why the strongest papers build custom eval sets tailored to their claim. A well-built custom eval set is also the foundation of regression testing (Chapter 7) and the validation data for your judges (Chapter 4).

Building one is real work: sourcing items, writing gold labels, deciding how many you need. This chapter is the field manual.

Step 1: Define what each item tests

An eval item is not just "an input." It is a probe for a specific capability or failure mode. Before collecting anything, list the dimensions your claim touches and make sure your set covers each. Example: you claim your RAG system is more faithful than a baseline. Your dimensions might be:

  • Questions answerable from a single passage (basic faithfulness)
  • Questions requiring synthesis across two passages (harder)
  • Questions where the retrieved passages contain a plausible-but-wrong distractor (adversarial)
  • Questions unanswerable from the passages (does the system correctly abstain?)

Aim for a stratified set: each dimension gets its own slice, so you can report per-slice scores. An aggregate score over an unstratified set hides the interesting structure — your method might ace easy items and fail every adversarial one, and the average would look fine.

Write a one-page "eval spec" before collecting: the claim, the dimensions, the target counts per dimension, the input format, the gold-label format, and the metric per dimension. This document is also the first thing to show a reviewer who asks "how did you evaluate?"

Step 2: Source the inputs

Good sources, in rough order of preference:

  1. Real user inputs (anonymized). Nothing beats the actual distribution — real queries are messier and more revealing than anything you invent. If you have product logs, sample from them. Scrub PII and get the necessary permissions; say so in your paper.
  2. Domain corpora repurposed. Take documents from your target domain and have annotators write questions about them (the SQuAD recipe). This gives you control over coverage and difficulty.
  3. Expert-written items. For specialized domains (law, medicine), pay domain experts to write items. Expensive per item, but each item is high-value and contamination-proof.
  4. Synthetic generation. Use a strong LLM to generate questions from passages, then have humans verify every item. This scales well, but verification is mandatory — unfiltered synthetic items contain wrong gold labels at rates that will corrupt your eval (studies routinely find 10–20% label error in unverified synthetic sets). Never let synthetic items into your set without a human checking each one.

What to avoid: scraping items from existing public benchmarks (contamination risk, and it adds nothing over citing those benchmarks), and writing all items yourself in one sitting (your blind spots become the set's blind spots — get at least two authors).

Step 3: Write gold labels

The gold label is the reference against which outputs are scored. Its form depends on your metric:

  • For EM/F1: the exact answer string(s). Collect multiple acceptable answers per item where phrasing varies ("Paris", "Paris, France") — this is the cheapest way to make string metrics fair.
  • For rubric scoring: a short description of what a good answer contains (key points list), which human raters or judges use as the reference.
  • For code: hidden unit tests, not a reference solution. Tests define correctness behaviorally; a reference solution is just one way to pass.
  • For unanswerable items: the gold label is the expected behavior ("should abstain" / "should say the passages don't contain the answer"), scored by a classifier or judge.

Labeling protocol matters as much as the labels: two independent labelers per item, adjudication of disagreements, and a measured label-agreement rate that you report. Your eval is only as trustworthy as its gold labels — a 5% label error rate puts a hard ceiling on the conclusions you can draw from small score differences.

Step 4: Decide the size

"How many items?" is the question everyone asks, and the honest answer is: enough to detect the effect you claim, with room for per-slice analysis. Some concrete guidance:

  • Minimum viable: 100 items. Below this, per-slice analysis is meaningless and even aggregate scores are noisy.
  • Development set: 200–500 items. Enough to iterate against with moderate confidence, and to run quick human spot-checks.
  • Test set for a paper claim: 500–1,000+ items for automated metrics; 150–300 for human-evaluated dimensions (Chapter 3's sizing).
  • Per slice: at least 50 items per dimension if you want to report per-dimension numbers with any confidence.

The statistics behind these numbers are in Chapter 8. The intuition: your score is an average over items, and averages over few items are noisy. A 5-point "improvement" on 50 items is about 2–3 items flipping — that is noise, not a finding.

Split your set. Keep a development slice (for iterating on prompts and methods) and a held-out test slice (touched once, at the end, for the reported numbers). The moment you start prompt-tuning against the full set, it becomes a development set and your numbers become optimistic. This is the cheapest insurance against fooling yourself in all of experimental science.

Step 5: Version and document

An eval set is a research artifact. Treat it like code:

  • Version it (v1.0, v1.1…) and record what changed between versions.
  • Write a short "datasheet": where inputs came from, who labeled them, the labeling protocol, agreement rates, known limitations ("our medical items cover cardiology and oncology only"), and the license/privacy status.
  • Freeze the test split. If you must fix a bad item, bump the version and re-run everything — never silently edit items that baselines already ran on.

If you can release the set publicly, do — it is a genuine contribution, and other researchers citing your eval set is how benchmarks are born. If you cannot (privacy, proprietary data), describe the construction process in enough detail that someone could replicate it.

A worked mini-example

Suppose you are evaluating faithfulness of answers from a customer-support RAG bot over 40 help articles:

  1. Spec: claim = "our constrained-decoding method reduces unsupported claims." Dimensions: single-passage questions (100), multi-passage synthesis (60), distractor passages present (40), unanswerable (40). Metric: faithfulness rubric 1–3 (Chapter 3) + unsupported-claim count via judge validated on 150 human-labeled items.
  2. Sourcing: sample 240 real anonymized support queries from logs, stratified by product area; discard ones needing account-specific data.
  3. Gold labels: for each, the set of passage IDs containing the answer + a key-points list written by two support agents independently, adjudicated.
  4. Splits: 120 dev / 120 test, stratified by dimension.
  5. Documentation: datasheet noting the log time window, anonymization steps, and that queries are English-only.

Total cost: a few days of agent time plus validation labeling. Total value: an eval you can defend in review and reuse for every future experiment on this system.

For your research: Your eval set is your operational definition of the claim. When a reviewer asks "does your method really improve faithfulness?", they are asking "does your eval set really measure faithfulness?" Spend your skepticism there first — a brilliant method on a sloppy eval convinces no one, while a simple method on a rigorous eval gets cited.

Designing adversarial items

Standard items measure typical performance. Adversarial items measure the failures you most need to prevent — the ones that are rare in logs but catastrophic in production. A faithfulness eval without adversarial items is like crash-testing a car only on straight roads.

Five adversarial patterns, with examples (for a RAG support bot):

  1. Distractor passages. Retrieve passages that are topically similar but answer a different question, and check the model doesn't blend them. Example: question about the refund policy; distractor passage describes the exchange policy in similar wording. Gold behavior: answer from the refund passage only.
  2. Negation and scope traps. Questions where a small word flips the answer. Example: "Which plans do not include international calls?" Gold: the complement set — models frequently answer with the included set.
  3. Unanswerable variants. Take an answerable question and remove the key passage. Gold behavior: abstain ("the documentation doesn't cover this"). This is the cheapest adversarial slice to build and the most revealing — most RAG systems fail it badly on first measurement.
  4. Paraphrase robustness. Rewrite 20% of questions with different wording, slang, or typos ("how do i get my $$ back" for "what is the refund policy"). Gold: same answer as the clean version. Measures whether your system handles real user language.
  5. Conflicting sources. Include two passages that disagree (e.g., an outdated policy page and a current one). Gold behavior: prefer the newer source, or surface the conflict explicitly — never silently pick one. This tests a genuinely hard capability; even strong systems struggle, which makes it a good diagnostic slice even if you don't gate releases on it.

Labeling adversarial items. The gold label for an adversarial item is often a behavior ("must abstain," "must cite the 2026 policy") rather than a string. Write the expected behavior explicitly in the gold label, and have your two labelers independently confirm it — adversarial items are exactly where labeler disagreement spikes, because the "right" behavior can itself be debatable. Items your labelers can't agree on are either badly designed (fix them) or genuinely ambiguous (keep them, but score them separately as a "judgment-call" slice rather than mixing them into the main score).

Weighting honestly. Report adversarial slices separately from standard slices — never blended into a single number without disclosure. A system scoring 90% on standard items and 40% on adversarial ones is not a "65% system"; it is a system that works typically and fails dangerously. Both numbers, side by side, tell the truth. If your paper's claim is about robustness, the adversarial slice is your primary evidence; if the claim is about typical-case quality, the standard slice leads and the adversarial slice is the limitations discussion.

From spec to spreadsheet: the eval set schema

An eval set is a dataset, and datasets need schemas. Here's a JSONL schema that has survived several real projects — one object per line, with metadata that pays for itself during analysis:

{"item_id": "medrag_0042",
 "slice": "distractor",
 "difficulty": "hard",
 "input": "Can I get a refund for my missed appointment?",
 "context": ["<passage 1 text>", "<passage 2 text>"],
 "gold": {"relevant_passages": ["p1"],
          "key_points": ["Refunds allowed within 24h", "No-show fee applies after 24h"],
          "expected_behavior": "answer"},
 "source": "clinic_logs_2026-01",
 "author": "labeler_A",
 "adjudicated": true,
 "version": "1.0"}

Why each field matters: slice enables per-slice reporting (Chapter 6) — without it, stratification exists only in your head. difficulty lets you check whether improvements concentrate on easy items (a common hollow victory). expected_behavior ("answer" vs. "abstain") makes behavioral assertions (Chapter 7) trivially computable. source/author/adjudicated/version are provenance: when someone asks "where did item 42 come from and who labeled it," you answer in seconds instead of excavating chat history.

Item QA checklist — run every item through this before it enters the set: - [ ] The input is self-contained (no "as mentioned above" without the above). - [ ] The gold label is complete (all acceptable answers listed, or behavior specified). - [ ] A second labeler independently produced a compatible label. - [ ] The item is assigned to exactly one slice. - [ ] No PII, no copyrighted text beyond fair use, license noted. - [ ] The item is not in the dev prompt examples or any training data.

Storing it. Keep the eval set in version control as JSONL (diffable, greppable) — not in a spreadsheet, not in a Google Doc. Large binary artifacts (passage corpora) go in object storage with a content hash recorded in the repo. Tag releases: eval-v1.0, eval-v1.1. Your future self, reproducing a paper's numbers two years later, will thank you more for this than for any clever metric.

Key takeaways

  • Build eval sets around your claim: define capability dimensions first, stratify items across them.
  • Source inputs from real usage when possible; verify every synthetic item with a human.
  • Gold labels need a protocol: two labelers, adjudication, reported agreement — label error caps your conclusions.
  • Size guidance: ≥100 minimum, 200–500 for development, 500–1,000+ for paper test sets, ≥50 per slice.
  • Split dev/test, version the set, write a datasheet, freeze the test split.

Chapter 7 — Regression Testing for Prompts and Model Swaps

The problem: silent degradation

You ship a support bot that scores 84% on your faithfulness eval. A month later it scores 79% and nobody changed anything — the model provider updated the weights behind the API. Or you change something: you "improve" the system prompt, the headline metric stays flat, but the abstention behavior on unanswerable questions quietly collapses, and users start getting confident fabrications.

Regression testing is the discipline of catching this. The idea is borrowed from software engineering — a suite of tests that must keep passing — but adapted to the nondeterministic, graded nature of LLM outputs (Chapter 1). Instead of asserting exact outputs, you assert statistical properties: the score on each eval slice must stay within a tolerance band of its baseline.

What goes into the regression suite

A good regression suite has three layers, from fast to slow:

Layer 1 — Smoke tests (seconds). 10–30 hand-picked items covering your most important behaviors, run on every prompt edit. These are your "canaries": if the new prompt breaks the basic greeting behavior or starts refusing normal questions, you know in seconds, not after a full eval. Score them with cheap checks — EM on closed items, a fast judge on open ones.

Layer 2 — Slice evals (minutes). Your full dev eval set (Chapter 6), run on every significant change: prompt edits, retrieval changes, decoding-parameter changes. Report per-slice scores vs. baseline with a tolerance (e.g., "no slice may drop more than 2 points; the aggregate may not drop at all").

Layer 3 — Full evals (hours). The held-out test set plus human spot-checks, run before releases and after model-provider updates. This is the gate for shipping.

The layers exist because full evals are too slow and expensive to run on every keystroke, and smoke tests are too thin to trust for releases. Match the layer to the decision.

Golden outputs and behavioral assertions

Two techniques make regression tests concrete:

Golden outputs. For a small set of critical items, store a known-good output and compare new outputs against it — not with string equality (Chapter 1: that fails on paraphrase), but with a similarity check: an LLM judge scoring "does the new output preserve the meaning and key facts of the golden output?", or an embedding-similarity threshold. Goldens catch behavioral drift — the model still answers, but its answers changed character.

Behavioral assertions. Assert properties, not strings: "on the 40 unanswerable items, the abstention rate must be ≥ 90%"; "no output may contain a phone number" (regex); "every answer must cite at least one passage" (format check); "average response length must stay within 20% of baseline" (verbosity guard). These are cheap, deterministic, and catch entire classes of regression that aggregate scores miss. A prompt tweak that keeps accuracy flat while doubling average length — and your inference bill — is a regression your score-only suite would never see.

def check_abstention(outputs, must_abstain_ids, min_rate=0.9):
    abstained = sum(1 for i, o in enumerate(outputs)
                    if i in must_abstain_ids and "don't know" in o.lower())
    rate = abstained / len(must_abstain_ids)
    assert rate >= min_rate, f"Abstention rate {rate:.2f} below {min_rate}"
    return rate

Handling nondeterminism in the suite

Because scores are random variables (Chapter 1), a regression suite needs statistical thinking, not fixed thresholds:

  • Run each layer 3+ times and compare distributions, not single numbers. A drop from 84% to 82% on one run is noise; a drop to 80% across three runs is signal (Chapter 8 gives you the test).
  • Fix decoding settings in the suite config (temperature, top-p, max tokens, seed if supported) and record them. A "regression" caused by someone bumping temperature is a config bug, not a model bug.
  • Set tolerance bands from history, not from intuition. Run the suite 10 times on the unchanged system, measure the natural run-to-run spread, and set your alert threshold outside it (e.g., mean ± 2 standard deviations). Otherwise you will chase noise forever.
  • Pin model versions where the provider allows (e.g., gpt-4o-2026-05-01 rather than gpt-4o). When only an unversioned alias exists, log the model's self-reported version string with every run so you can correlate score shifts with provider updates.

The model-swap playbook

Sooner or later you will swap models — a new release, a cost-motivated move to a smaller model, a provider change. Treat it as a first-class experiment:

  1. Freeze the suite first. Run the full Layer 2 + 3 suite on the old model, 3+ runs, and record baselines per slice.
  2. Run the identical suite on the new model — same prompts, same settings, same items. Resist the urge to "tune the prompt for the new model" before measuring; measure the drop (or gain) first, then optimize.
  3. Diff the failures, not just the scores. Pull the items where the new model fails and the old passed. Categorize them: formatting differences (fix with output parsing), capability gaps (real regressions), and behavior changes (different but acceptable — update goldens).
  4. Check the dimensions benchmarks miss: latency, cost per 1k items, refusal behavior, verbosity, and safety properties. A model that scores 2 points higher but costs 5× more or refuses 10% of legitimate queries may be a worse choice.
  5. Canary in production. Route a small fraction of live traffic to the new model and compare real user-facing metrics (thumbs-down rate, task completion) before full cutover.

Document the swap as a one-page decision memo: scores before/after per slice, cost/latency deltas, failure categories, and the final call. Future you — and your reviewers — will be grateful.

CI integration sketch

Wire Layer 1 and 2 into your continuous integration so every commit touching prompts, retrieval code, or model config triggers them:

# .github/workflows/eval-regression.yml (sketch)
on:
  pull_request:
    paths: ["prompts/**", "src/retrieval/**", "eval/**"]
jobs:
  smoke:
    runs-on: ubuntu-latest
    steps:
      - run: python eval/run_suite.py --layer smoke --runs 3
  slice-eval:
    needs: smoke
    runs-on: ubuntu-latest
    steps:
      - run: python eval/run_suite.py --layer slice --runs 3 --compare-to baseline.json

Store baseline.json (per-slice means and standard deviations) in version control next to the prompts. A PR that moves a slice outside its tolerance band fails CI — exactly like a broken unit test. The cultural effect matters as much as the technical one: prompt edits become reviewed, tested changes instead of vibes.

For your research: Your paper's "we changed X and the score went from A to B" claim is only credible if A was measured with the same rigor as B — same items, same settings, multiple runs. A regression suite maintained during the project gives you this for free: every experimental result in your paper is then a before/after comparison against a frozen baseline, which is precisely what reviewers want to see.

Case study: the regression nobody noticed

A mid-size startup's support bot, "prompt v4," shipped on a Friday. The headline faithfulness score went from 82% to 83% — a small win, celebrated briefly. Three weeks later, support tickets about "the bot making things up" tripled. What happened?

The v4 prompt added "be helpful and thorough." The aggregate score held because answers got longer and more detailed — the judge rewarded the extra content (length bias, Chapter 4). But buried in the per-slice numbers, which nobody checked before shipping, the unanswerable slice had collapsed: abstention rate fell from 92% to 64%. "Thorough" had quietly translated into "never say I don't know." The team had a slice eval that would have caught it — they just didn't gate the release on it.

How the three-layer suite would have caught it: - Smoke tests (seconds): one smoke item was "answer an unanswerable question" — it failed on the first run of v4. But smoke tests weren't wired into the prompt-editing workflow; they lived in a notebook someone ran monthly. - Slice evals (minutes): the abstention slice showed the 28-point drop. But the release checklist only required "aggregate score not decreased." - Full evals (hours): never run, because "it was just a prompt tweak."

The fix was procedural, not technical: prompt changes became pull requests, the smoke suite ran on every PR, and the release checklist required every slice to stay within its tolerance band — not just the aggregate. The team estimated the incident cost ~200 support hours; the suite cost an afternoon to build.

Non-functional regression tests. The same machinery catches degradations that aren't about quality at all: - Latency budget: p95 response time ≤ 4s on the smoke set. A prompt that doubles output length can silently double latency and cost. - Cost per 1k queries: track input+output tokens per eval run. "Helpful and thorough" increased cost 40% — a regression the quality score never saw. - Output length guard: mean response length within ±20% of baseline. Length drift is the canary for verbosity regressions, judge gaming, and cost blowups. - Format compliance: % of outputs parsing as the expected schema (JSON, citations format). A model update that changes formatting breaks downstream code without changing "quality" scores.

# Tolerance-band check from measured history (Chapter 7)
def check_regression(current, baseline_mean, baseline_std, k=2):
    lo, hi = baseline_mean - k * baseline_std, baseline_mean + k * baseline_std
    status = "PASS" if lo <= current <= hi else "FAIL"
    return {"current": current, "band": (round(lo,2), round(hi,2)), "status": status}

Run the non-functional checks on the same schedule as the quality suite. A system that gets slightly better answers at 3× the cost and 2× the latency has not unambiguously improved — and your eval suite should say so out loud.

Prompt versioning in practice

Prompts are code. Version them like code — because every result in your paper depends on the exact bytes of the prompt, and "the prompt we used" from memory is not reproducible.

Store prompts as files. One file per prompt, in prompts/, with a descriptive name and a header comment recording its purpose and history:

# prompts/rag_answer_v3.md
# Purpose: answer generation for MedRAG eval
# History: v1 (2026-03-01) baseline; v2 (2026-03-08) added citation requirement;
#          v3 (2026-03-15) added "say I don't know" instruction after abstention failures
---
Answer the question using ONLY the passages below. Cite each claim...

Diff prompts in review. When a prompt change goes through pull-request review, the diff should be readable: keep prompts as plain text/markdown (not embedded in Python strings with escaping), one sentence per line if the team prefers fine-grained diffs. The PR description states the hypothesis ("adding the abstention instruction should raise the unanswerable-slice score without hurting completeness") and links the slice-eval results. This turns prompt engineering from alchemy into an auditable engineering practice.

The prompt review checklist — for the reviewer, not just the author: - [ ] Does the prompt state the task, the constraints, and the output format explicitly? - [ ] Are there instructions that conflict with each other? (Common: "be thorough" + "be concise" — pick one or define the tradeoff.) - [ ] Does it handle the failure modes? (What should the model do when the context lacks the answer?) - [ ] Were few-shot examples chosen from the dev set (never the test set), and do they cover edge cases? - [ ] Is the decoding config (temperature, max tokens) recorded alongside?

Never do this: keep the "real" prompt in a chat window, a playground, or someone's notes app while the repo holds a stale copy. The prompt in version control is the prompt — everything else is a rumor. When your paper says "the full prompt is in Appendix A," generate Appendix A from the version-controlled file, not by retyping.

Key takeaways

  • Regression-test LLM systems in three layers: smoke (seconds), slice evals (minutes), full evals (hours) — match the layer to the decision.
  • Use golden outputs (similarity-checked, not string-matched) and behavioral assertions (abstention rates, format checks, verbosity guards) alongside scores.
  • Compare distributions across 3+ runs, fix decoding settings, set tolerance bands from measured history, pin model versions.
  • For model swaps: freeze baseline, run identical suite, diff failures by category, check cost/latency/safety, canary before cutover.
  • Put prompts and baselines in version control; gate merges on the suite like unit tests.

Chapter 8 — Statistical Thinking: Variance, Significance, Error Bars

Numbers without uncertainty are anecdotes

"Our method scores 84.2% vs. the baseline's 81.7%" — a 2.5-point win. Publishable? It depends entirely on the uncertainty around those numbers. If each is the mean of 1,000 items with small run-to-run variance, the win is real. If each is a single run over 164 HumanEval problems, the win is about four problems — well within noise, and a reviewer who knows statistics will reject the claim.

This chapter gives you the minimum statistical toolkit for eval work: where variance comes from, how to estimate it, how to test whether a difference is real, and how to report all of it honestly. No proofs — just the procedures, the code, and the judgment calls.

Where the variance comes from

An eval score varies for four reasons, and you should know which dominate in your setup:

  1. Item sampling. Your eval set is a sample of all possible inputs. A different 500 questions would give a different score. This is usually the largest source of variance, and it shrinks as your set grows (roughly with the square root of the number of items).
  2. Model sampling. Nondeterministic decoding (Chapter 1) means the same item can score differently across runs. Grows with temperature; present even at temperature 0 on hosted models.
  3. Judge/rater noise. Human raters disagree; LLM judges are stochastic. This adds variance on top, especially for rubric-scored dimensions.
  4. Prompt and setup sensitivity. Small, often undocumented differences — the harness, the few-shot examples, the chat template — shift scores systematically. This is bias, not variance: it does not average out with more runs, and it is why you must hold the setup fixed when comparing systems (Chapter 5).

The practical consequence: your reported score should always come with an uncertainty estimate that accounts for at least (1) and (2). A point estimate alone is an incomplete measurement.

The bootstrap: your main tool

The bootstrap is a beautifully simple way to estimate uncertainty: resample your eval items with replacement, recompute the score on each resample, and look at the spread of the resulting scores. It makes almost no assumptions and works for any metric — accuracy, F1, win rates, even judge scores.

import numpy as np

def bootstrap_ci(scores, n_boot=10000, ci=95):
    """scores: per-item scores (0/1 or continuous). Returns (mean, lo, hi)."""
    scores = np.asarray(scores)
    n = len(scores)
    boot_means = [np.random.choice(scores, size=n, replace=True).mean()
                  for _ in range(n_boot)]
    lo, hi = np.percentile(boot_means, [(100 - ci) / 2, 100 - (100 - ci) / 2])
    return scores.mean(), lo, hi

mean, lo, hi = bootstrap_ci(item_scores)
print(f"{mean:.1%} [{lo:.1%}, {hi:.1%}]")  # e.g. 84.2% [81.9%, 86.4%]

Report the 95% confidence interval (CI) alongside every score: "84.2% [81.9, 86.4]." If two systems' CIs barely overlap, the difference is suggestive; if they overlap substantially, you do not have evidence of a difference — more items or more runs are needed.

How many bootstrap resamples? 10,000 is cheap for simple metrics and gives stable intervals. For expensive metrics (judge-scored), 1,000–2,000 suffices.

Testing whether a difference is real

You have system A scoring 84% and system B scoring 81% on the same 500 items. Is A better? The right test for paired data (same items, two systems) is a paired test:

  • McNemar's test for binary outcomes (correct/incorrect per item): it looks only at the items where the systems disagree and asks whether the disagreements favor one system more than chance would.
  • Paired permutation test for continuous scores: randomly swap A/B labels within each item many times, and see how often the shuffled difference is as large as the observed one. Assumption-free and easy to explain.
import numpy as np

def paired_permutation_pval(a_scores, b_scores, n_perm=10000):
    a, b = np.asarray(a_scores), np.asarray(b_scores)
    observed = (a - b).mean()
    diffs = a - b
    count = sum(
        (diffs * np.random.choice([-1, 1], size=len(diffs))).mean() >= observed
        for _ in range(n_perm)
    )
    return (count + 1) / (n_perm + 1)

p = paired_permutation_pval(scores_A, scores_B)
print(f"p = {p:.4f}")  # p < 0.05: difference unlikely to be chance

Reading p-values honestly. p < 0.05 means "a difference this large would occur by chance less than 5% of the time if the systems were truly equal." It does not mean the difference is large or important — with 100,000 items, a 0.1-point gap can be "significant." Always report the effect size (the actual gap with its CI) next to the p-value, and ask whether the gap matters in practice. A statistically significant 0.3-point gain that costs 10× more inference is not a win.

Multiple comparisons. If you test 20 slices/metrics and report the one with p < 0.05, you have likely found noise — with 20 tests, you expect one false positive at the 5% level by chance alone. Fixes: decide your primary metric in advance (Chapter 12's template forces this), treat the rest as exploratory, or apply a correction (Bonferroni: divide your threshold by the number of tests; or control the false discovery rate). At minimum, disclose how many comparisons you ran.

Power: sizing your eval before you run it

Power is the probability your eval will detect a real difference of a given size. An underpowered eval — too few items to detect the effect you care about — wastes everyone's time: a null result tells you nothing, because the eval could not have found the effect anyway.

A handy approximation for comparing two proportions (accuracies): to detect a difference of d percentage points at 80% power with 5% significance, you need roughly this many items per system:

True difference Items needed (approx.)
2 points ~4,800
5 points ~800
10 points ~200
20 points ~50

(Assumes baseline accuracy near 80% and paired items; unpaired needs ~2×.)

The sobering lesson: detecting small improvements requires large evals. If your expected gain is 2 points, 200 items cannot show it — and neither can most published eval sets. Either build a bigger set, target bigger effects, or frame the claim around qualitative analysis rather than a small numeric gap. Reviewers increasingly do this math themselves; do it first.

# Quick power check with statsmodels (install: pip install statsmodels)
from statsmodels.stats.power import NormalIndPower
n = NormalIndPower().solve_power(effect_size=0.1, alpha=0.05, power=0.8)
print(f"Items per group (unpaired): {n:.0f}")

Reporting: the honest results table

Every results table in your paper should carry uncertainty. The pattern:

System Accuracy 95% CI Δ vs. baseline p (paired perm.)
Baseline 81.7% [79.4, 84.0] — —
Ours 84.2% [82.0, 86.3] +2.5 0.03

Plus a methods footnote: "Scores are means over 3 runs; CIs from 10,000 bootstrap resamples over items; p-values from paired permutation tests (10,000 permutations) on the primary metric." That footnote is what separates a credible results section from a decorative one. Add error bars to every bar chart — a chart without them is a chart that hides its own uncertainty.

For your research: Before running your final eval, do the power math for your expected effect size. If the required item count exceeds what you have, you have three honest options: collect more items, aim for a larger effect, or pre-register a qualitative/stratified analysis as your primary evidence instead of a headline number. What you cannot do is run an underpowered eval and report the gap as a finding anyway — that is how false "improvements" enter the literature.

Worked analysis: System A vs. System B on 400 items

A complete statistical workup, start to finish, with the numbers and the interpretation.

The data. Two RAG configurations, A (new) and B (baseline), run on the same 400 test questions. Each answer is scored 0/1 for faithfulness by a validated judge. Three runs each; run-to-run variation is small (±0.8 points), so we pool per-item majority scores.

Step 1 — Point estimates and CIs.

mean_a, lo_a, hi_a = bootstrap_ci(scores_a)  # 0.845 [0.810, 0.878]
mean_b, lo_b, hi_b = bootstrap_ci(scores_b)  # 0.808 [0.770, 0.845]

A: 84.5% [81.0, 87.8]. B: 80.8% [77.0, 84.5]. The gap is +3.7 points, but the intervals overlap substantially — the overlap alone tells you the evidence is suggestive, not conclusive.

Step 2 — Paired test. Because both systems ran on the same items, use McNemar's test on the disagreement cells: A-right/B-wrong = 38 items; A-wrong/B-right = 23 items.

from statsmodels.stats.contingency_tables import mcnemar
table = [[339 - 0, 38], [23, 0]]  # [[both right, A-only], [B-only, both wrong]]
# simplified: mcnemar on [[a, b],[c, d]] uses b vs c
result = mcnemar([[316, 38], [23, 23]], exact=True)
print(result.pvalue)  # ≈ 0.07

p ≈ 0.07 — not significant at α = 0.05. The honest conclusion: no statistically significant difference on the full set, despite the 3.7-point gap. This is the moment undisciplined teams publish anyway. The disciplined team looks deeper.

Step 3 — Slice analysis (pre-registered). The eval plan declared two slices: single-passage (250 items) and multi-passage synthesis (150 items).

Slice A B Δ p (McNemar)
Single-passage (n=250) 88.0% 86.8% +1.2 0.62
Multi-passage (n=150) 78.7% 70.7% +8.0 0.02

The improvement is concentrated in multi-passage synthesis — exactly the capability the method was designed for. The slice p-value (0.02) survives because this comparison was pre-registered as the primary analysis; had the team tested 10 slices and reported the best, a correction would be required.

Step 4 — Effect size and practical meaning. On multi-passage items, the gap is 8 points with CI roughly [+1.5, +14.5]. In practical terms: about 12 more faithful answers per 150 hard questions. Combined with no regression on the single-passage slice and equal cost/latency, this is a real, shippable improvement — on the hard slice, which is where users were complaining.

What goes in the paper. The results table shows all three rows (overall + two slices) with CIs; the text states the primary metric and the pre-registered slice analysis; the p-values are reported exactly (p = 0.07, p = 0.02), not thresholded into "significant/not significant" alone; and the discussion notes that the overall gap is not significant but the targeted slice is. A reviewer can disagree with the interpretation, but they cannot accuse the authors of hiding anything.

Common misreadings this example guards against: (a) "p = 0.07 means no effect" — wrong; it means insufficient evidence, and the CI shows the effect could be as large as +7 points overall. (b) "The slice p = 0.02 proves the method works everywhere" — wrong; it supports the claim for multi-passage synthesis specifically. (c) "We could just run more items until p < 0.05" — that is p-hacking by data peeking; decide the sample size from the power analysis before running (Chapter 8), or pre-register a sequential design.

Beyond p-values: compatibility intervals and practical equivalence

The p-value tells you whether a difference could be chance. Two richer questions usually matter more: what range of effects is compatible with the data? and is the effect big enough to care about?

Read CIs as compatibility intervals. A 95% CI of [+0.5, +7.0] points doesn't just say "significant" — it says effects from half a point (trivial) to seven points (large) are all reasonably compatible with what you observed. If your CI spans both trivial and large effects, your experiment was too small to pin down the practical question, even if p < 0.05. Report the interval, then discuss both ends: "the improvement could be as small as 0.5 points — below our 2-point practical threshold — or as large as 7."

The region of practical equivalence (ROPE). Before the experiment, define the range of differences you consider practically meaningless — say, ±1.5 points for your task. Afterward, check where the CI falls: entirely outside the ROPE (meaningful effect), entirely inside (practically no difference — a publishable null result!), or overlapping (inconclusive). This turns the vague "is it significant?" into the sharper "is it significant and does it matter?" — and it gives null results a rigorous interpretation instead of leaving them in limbo.

A Bayesian one-liner. If you want a direct statement like "there's a 93% chance the true improvement exceeds 2 points," fit a simple Bayesian model (a Beta-Binomial for binary outcomes takes five lines with any probabilistic programming library). You don't need to become a Bayesian — just know the option exists for when stakeholders ask the direct question that p-values refuse to answer.

The reporting habit that covers all of this: every comparison gets four numbers — the estimated difference, its 95% CI, the p-value, and your pre-declared practical threshold. Readers who care about significance get the p-value; readers who care about decisions get the rest. Nobody has to guess what you think the numbers mean, because the threshold says it out loud.

Key takeaways

  • Report every score with a 95% confidence interval (bootstrap over items, 10k resamples); never a bare point estimate.
  • Use paired tests (McNemar for binary, paired permutation for scores) when both systems ran on the same items.
  • Report effect sizes with CIs, not just p-values; a significant but tiny gap is not a win.
  • Pre-register one primary metric; disclose all comparisons to avoid p-hacking.
  • Do the power math first: small effects need thousands of items — size the eval to the claim or change the claim.

Chapter 9 — Evaluating RAG Systems and Agents Specifically

Why they need their own chapter

A RAG (retrieval-augmented generation) system and an AI agent fail in ways that plain LLM evals do not capture. A RAG answer can be fluent, relevant, and unsupported by anything retrieved — the worst failure mode, invisible to BLEU and ROUGE. An agent can produce a beautiful chain of thought while calling tools in the wrong order and never reaching the goal — invisible to any text metric. Evaluating these systems means measuring the pipeline, not just the final string: what was retrieved, what was cited, what tools were called, whether the goal state was reached.

Evaluating RAG: the two halves

A RAG system has two halves — retrieval and generation — and you must evaluate both, because they fail independently.

Retrieval metrics

Given a question and a corpus, did the retriever find the passages needed to answer?

  • Recall@k: of the gold relevant passages, what fraction appears in the top-k retrieved? The workhorse metric. If recall@k is low, no generator can save you — the facts were never provided.
  • Precision@k / nDCG@k: of what was retrieved, how much was relevant, with nDCG weighting higher ranks more. Matters because generators get distracted by irrelevant context — a retriever with high recall but terrible precision feeds the generator noise.
  • Answer recall (end-to-end retrieval check): does the retrieved set contain the answer string? A cheap, useful diagnostic that skips relevance labels.
def recall_at_k(retrieved_ids, gold_ids, k):
    retrieved_k = set(retrieved_ids[:k])
    gold = set(gold_ids)
    return len(retrieved_k & gold) / len(gold) if gold else 0.0

Diagnosing with retrieval slices. Break failures down: is recall low because the query is ambiguous (needs query rewriting), because the chunking split the answer across chunks (needs better chunking or larger windows), or because the embedding model misses domain vocabulary (needs fine-tuning or hybrid sparse+dense retrieval)? Per-failure-category analysis turns "recall@5 = 0.62" into an engineering plan.

Generation metrics: faithfulness first

The central RAG failure is the unsupported claim — a statement in the answer not backed by retrieved passages. Measure it directly:

  • Faithfulness / citation precision: split the answer into atomic claims; for each, ask (a judge, validated per Chapter 4, or an NLI model) whether it is entailed by the retrieved passages. Report the fraction supported. This is the single most important RAG metric.
  • Citation recall: of the key facts in the reference answer, what fraction appears in the model's answer with support? Catches answers that are faithful but incomplete.
  • Abstention rate on unanswerables: when the retrieved passages do not contain the answer, does the system say so? A RAG system that never abstains will hallucinate on every retrieval failure — measure this on a dedicated slice (Chapter 6).

A practical claim-decomposition prompt for the judge: "Break the following answer into atomic factual claims, one per line, each verifiable independently." Then verify each claim against the passages. Frameworks like RAGAS (Es et al., 2023) package this decomposition-and-verification loop; whether you use the library or roll your own, validate the judge (Chapter 4) — faithfulness judges have their own error rates, typically 5–15% per claim.

The RAG improvement loop

Evaluate in pipeline order: fix retrieval before tuning generation. A common waste of effort is prompt-engineering the generator when recall@5 is 0.4 — the generator is being asked to answer from passages that do not contain the answer, and no prompt fixes that. The disciplined loop:

  1. Measure recall@k and precision@k. If recall is the bottleneck, work on chunking, hybrid retrieval, or query rewriting.
  2. Once the retrieved set usually contains the answer, measure faithfulness. If claims go unsupported, work on citation-forcing prompts, constrained decoding, or answer-then-verify pipelines.
  3. Finally, measure completeness and usefulness with human or validated-judge rubrics.

Evaluating agents: behavior over text

An agent acts in an environment: calling tools, browsing, writing files. Its "output" is a trajectory of actions plus a final state. Evaluate accordingly:

  • Task success rate: did the environment reach the goal state? This is binary and behavioral — the agent either booked the refund or did not. It is the primary metric, full stop.
  • Step efficiency: how many actions did success take vs. an optimal or human reference? An agent that succeeds in 40 steps where 8 suffice is burning money and patience.
  • Tool-call accuracy: per step, was the right tool called with valid arguments? Decompose into function-name accuracy and argument accuracy — argument errors (wrong date format, hallucinated IDs) are the dominant failure mode in practice.
  • Error recovery: when a tool call fails, does the agent recover (retry with fixed args, try an alternative) or spiral? Score recovery rate on a slice of tasks with injected failures — this dimension separates demos from deployable agents.
  • Safety constraints: did the agent respect its action boundaries (no deleting files, no sending emails without confirmation)? Score constraint violations as automatic failures, however "successful" the task.
# Trajectory scoring sketch
def score_trajectory(steps, goal_check, allowed_tools):
    violations = [s for s in steps if s.tool not in allowed_tools]
    if violations:
        return {"success": False, "reason": "constraint violation"}
    success = goal_check()  # inspect environment state
    return {"success": success, "steps": len(steps)}

Evaluating the reasoning trace. Chain-of-thought traces are tempting to score, but beware: a trace can be coherent while the actions are wrong, and models rationalize bad actions fluently. Score traces only for specific, checkable properties ("does the trace mention the tool result before acting on it?") rather than general "reasoning quality." The environment state is the ground truth; the trace is commentary.

Benchmarks for agents (WebArena, SWE-bench, and similar) provide environments with goal-state checks — useful, with the usual contamination and saturation caveats from Chapter 5. For your own agent, build a small suite of 30–100 tasks in a sandboxed environment with programmatic goal checks; this is your agent's equivalent of a unit-test suite, and it pays for itself within weeks.

Cost-aware evaluation

Agents and RAG systems consume variable amounts of compute per task — retrieval calls, multiple model calls, long trajectories. Always report cost per task (or per 1k tasks) and latency alongside quality metrics. A method that gains 2 points of success rate at 10× the cost is a research result, not a deployment candidate, and reviewers in applied venues will ask. The honest table has four columns: success rate, steps, cost, latency.

For your research: When you claim an improvement to a RAG system or agent, attribute it: show the metric for the component you changed and the end-to-end metric, so the reader can see the causal chain ("better chunking → recall@5 0.62→0.78 → faithfulness 0.81→0.88"). Unattributed end-to-end gains invite the suspicion that something else changed — and in agentic systems, something else always might have.

Going further: multi-hop RAG, long context, and an agent case study

Multi-hop RAG. When answers require combining facts from several passages, add hop-level diagnostics: hop recall (for each required fact, was a supporting passage retrieved?), and reasoning-chain faithfulness (does each step of a multi-step answer cite its source?). A frequent failure: the system retrieves hop 1 correctly, then generates hop 2 from parametric memory instead of retrieving it — the answer looks complete but half of it is ungrounded. Diagnose by scoring faithfulness per claim and mapping unsupported claims back to hops; if unsupported claims cluster at later hops, your retriever needs iterative (multi-round) retrieval, not a bigger generator.

Long-context eval. When the corpus chunk placed in context grows to tens of thousands of tokens, position effects dominate: models attend best to the start and end of the context and worst to the middle ("lost in the middle"). Test this directly with a needle-in-the-context probe: place the answer-bearing passage at 10 different positions in a long context and plot accuracy by position. If you see the U-shaped curve, mitigations include re-ranking key passages to the top, repeating the question after the context, or simply retrieving less (a short, precise context beats a long, noisy one for most QA tasks).

Agent case study: the ticket-triage agent. A team builds an agent that reads incoming support tickets, queries the order database, and either resolves the ticket or escalates with a summary. Their 40-task sandbox eval:

Metric Result Notes
Task success rate 78% (31/40) Goal state: correct resolution or correct escalation
Step efficiency 14.2 avg steps vs. 6 optimal Agent re-reads tickets and retries queries excessively
Tool-call accuracy 91% function name, 74% arguments Date-format and order-ID errors dominate
Error recovery 45% After a failed DB query, usually spirals instead of reformulating
Safety violations 2/40 Issued a refund without confirmation twice — automatic failures

The headline (78% success) looks deployable; the breakdown says otherwise. Argument accuracy of 74% means one in four tool calls has a malformed argument — the agent succeeds despite its tool use, by retrying. Error recovery at 45% means novel failures become infinite loops. And 2 safety violations in 40 tasks is a hard stop: at production volume, that is hundreds of unauthorized refunds. The team's roadmap wrote itself: (1) argument validation layer before tool execution, (2) a recovery policy (max 2 retries, then escalate), (3) a confirmation gate on irreversible actions — re-evaluated against the same 40 tasks. Notice how each metric maps to exactly one engineering fix. That is what good agent evals do: they don't just score, they prescribe.

A taxonomy of tool-call errors

In agent evals, "tool-call accuracy 74%" is a starting point, not a diagnosis. Break argument errors into types — each has a different fix:

  1. Format errors. The argument is right in spirit, wrong in syntax: dates as "Oct 8" instead of "2026-10-08", phone numbers with dashes when the API wants digits. Fix: a formatting layer or a stricter schema with examples in the tool description. The cheapest errors to eliminate.
  2. Hallucinated identifiers. The agent invents an order ID, a filename, or a user ID that doesn't exist. Fix: constrain the agent to select IDs from tool outputs (never generate them), and validate IDs against known values before the call executes.
  3. Stale arguments. The agent reuses an argument from three steps ago after the context changed — e.g., querying with last week's date range. Fix: have the agent restate key parameters from the most recent observation before each call ("re-grounding").
  4. Missing required arguments. The call omits a field the schema requires. Fix: client-side validation with an error message that names the missing field — then measure whether the agent recovers (Chapter 9's error-recovery metric).
  5. Wrong-tool errors. The agent calls search_orders when it needed get_order_details — usually a sign the tool descriptions are ambiguous. Fix: rewrite the descriptions with contrastive examples ("use X for listing, Y for details of a specific order"), which helps more than adding more tools.
  6. Premature calls. The agent acts before gathering needed information — booking before confirming the date. Fix: a confirmation gate for irreversible actions, plus a prompt instruction to summarize evidence before acting.

How to use the taxonomy: on your next 50 failed tool calls, tag each with one of the six types and count. The distribution tells you exactly where to invest: 60% format errors means fix the schema layer this week; 60% hallucinated IDs means change the prompting strategy. Report the distribution in your paper's failure analysis — "of 83 argument errors, 41% were format errors, 27% hallucinated identifiers…" — and reviewers will see a team that understands its system rather than one that merely scored it.

Key takeaways

  • Evaluate RAG in pipeline order: retrieval (recall@k, precision@k) first, then generation faithfulness (claim-level support), then completeness and abstention.
  • Faithfulness — the fraction of answer claims entailed by retrieved passages — is the central RAG metric; measure it with a validated claim-decomposition judge.
  • For agents, task success (goal-state reached) is primary; add step efficiency, tool-call/argument accuracy, error recovery, and safety-constraint compliance.
  • Score environment state, not reasoning traces — traces rationalize; states don't lie.
  • Always report cost and latency per task alongside quality; attribute gains to the component you changed.

Chapter 10 — Reporting Evals in Papers (Reproducibility Checklist Reviewers Love)

The eval section is where papers are won and lost

Reviewers skim the method, but they read the evaluation. It is where they decide whether your numbers mean anything. A surprising number of rejections trace back not to weak ideas but to eval sections that leave basic questions unanswered: What exactly was measured? On what data? With what prompt and settings? How many runs? What is the uncertainty? Each unanswered question is a reason to doubt, and doubt accumulates into "reject."

This chapter gives you the anatomy of a trustworthy eval section plus a reproducibility checklist you can apply before every submission.

Anatomy of a strong eval section

1. The claim, stated as a measurable difference. Open with one paragraph: what you claim, and how the eval tests it. "We claim that constrained decoding reduces unsupported claims in RAG answers. We test this by comparing our method against a standard-decoding baseline on faithfulness (claim-level support judged by a validated LLM judge) across 600 questions stratified into four difficulty slices."

2. Data. What eval set(s), how built, how large, which split. Cite public benchmarks with version/harness; describe custom sets with a pointer to the construction details (Chapter 6 — put the datasheet in an appendix). State the test split was held out during development.

3. Systems compared. Every baseline, with enough detail to reproduce: model name and version (e.g., "gpt-4o-2026-05-01", not "GPT-4"), prompt (full text in appendix or supplement — this is non-negotiable for prompt-based methods), decoding settings (temperature, top-p, max tokens, seed), and any retrieval/tooling configuration. "We used GPT-4" is not reproducible; a dated model version plus the exact prompt is.

4. Metrics. One sentence of justification per metric (Chapter 2's rule): what it measures, why it fits the task, what it misses. Name the implementation ("sacreBLEU v2.4", "our judge prompt in Appendix B, validated at 81% agreement with humans on 150 items").

5. Protocol. Number of runs, how randomness was handled, the primary metric (pre-registered — say so), and the statistical analysis: CIs via bootstrap, paired tests, power where relevant (Chapter 8). State how ties and failures were handled ("model outputs that failed to parse as JSON were counted as incorrect; this affected 1.2% of baseline outputs").

6. Results. Tables with uncertainty (Chapter 8's pattern), per-slice breakdowns, and — crucially — analysis, not just numbers. The best eval sections include a failure analysis: sample 50–100 errors from your method, categorize them, and report the distribution. "Of 80 sampled failures, 41% were retrieval misses, 33% were unsupported elaborations, 26% were abstention failures" tells the reader (and you) what to work on next, and reviewers love it because it shows you understand your own system.

7. Limitations and threats. A short, honest paragraph: what the eval does not cover (languages, domains, long contexts), known biases (judge leniency, possible contamination), and what would change your conclusion. This paragraph does not weaken your paper — it strengthens it, because it answers the reviewer's objections before they are raised.

The reproducibility checklist

Run through this before submitting. Every "no" is a fix to make:

  • [ ] Model versions pinned. Every model named with exact version/date; no unversioned aliases.
  • [ ] Prompts published. Full prompt text for every system in appendix or supplement, including judge prompts.
  • [ ] Decoding settings reported. Temperature, top-p, max tokens, seed policy — for every system including baselines.
  • [ ] Data identified. Benchmark name + version + harness, or custom-set datasheet; test split held out during development.
  • [ ] Metrics justified. One sentence per metric: what it captures, what it misses; implementation named.
  • [ ] Runs repeated. ≥3 runs per system; means with 95% CIs reported, not single runs.
  • [ ] Primary metric declared. Stated in advance; other metrics labeled exploratory.
  • [ ] Statistics appropriate. Paired tests for paired data; multiple-comparison disclosure; power considered for the claimed effect size.
  • [ ] Baselines re-run. Not copied from other papers (or, if copied, flagged as such with the caveat).
  • [ ] Failures analyzed. Sampled, categorized, reported — including parse failures and abstention behavior.
  • [ ] Human evals documented. Rubric available, raters described, blinding stated, agreement reported (Chapter 3's five questions).
  • [ ] Judges validated. Agreement vs. humans reported with sample size; biases checked and disclosed.
  • [ ] Code and artifacts shared. Eval scripts, prompts, and (where possible) outputs released; random seeds fixed in code.
  • [ ] Limitations stated. Coverage gaps, contamination risks, and what could overturn the conclusion.

Tape this list next to your desk. Papers that pass it rarely get rejected on evaluation grounds.

Common reviewer objections, pre-answered

  • "The improvement is small." → Show the CI and the power analysis: small but precisely measured effects are real findings; small effects on 100 items are noise. If the effect is small, argue practical significance (cost, latency, a qualitative capability) rather than pretending 1.5 points is a breakthrough.
  • "Why this metric?" → Your one-sentence justification plus the validation-against-humans result.
  • "The baseline is weak." → This is the most damaging objection and the hardest to fix late. Choose the strongest reasonable baseline before running experiments — the current best published method, not a strawman. A win against a weak baseline proves little.
  • "Only tested on one dataset." → At least two datasets, or one dataset plus a robustness slice (paraphrased questions, a second domain). Single-dataset papers survive only when the dataset is the contribution.
  • "Can I reproduce this?" → The checklist above, plus released code. If you used a closed model that may drift, say so and pin the version.

A note on negative and null results

If your eval shows no difference, or your method loses — report it anyway, at least in ablations. Null results with proper statistics ("no significant difference, 95% CI on the gap: [−1.2, +0.8], n=1,000") are informative: they tell the field what does not work, which is knowledge too. Burying null ablations while highlighting the one positive slice is p-hacking by omission, and experienced reviewers can smell it. The field's replication problems come largely from this asymmetry; be part of the fix.

For your research: Write your eval section's methods before you run the final experiments — the claim paragraph, the metric justifications, the primary metric, the statistical plan. This is a lightweight pre-registration, and it protects you from the most common form of self-deception: choosing the analysis after seeing the results. When the numbers come in, you execute the plan you wrote, and whatever it says is what you report.

Annotated example: an eval section with margin notes

Below is a mock eval-methods excerpt (280 words) for the medical RAG project from Chapter 12. The bracketed notes map each sentence to this chapter's checklist — imagine them as margin comments from a careful reviewer.

We evaluate on MedRAG-eval v1.0, a set of 600 patient questions sampled from anonymized clinic logs (Jan–Mar 2026), stratified into single-article (250), multi-article (150), distractor (100), and unanswerable (100) slices. [Data: size, source, stratification — ✓] Gold labels (relevant article IDs and key-point lists) were written independently by two medical students (agreement 88% on article IDs) with adjudication of disagreements. [Labeling protocol + agreement — ✓] We split 300/300 into development and test; the test split was frozen on 2026-04-02 before any method tuning. [Held-out test — ✓]

We compare our cite-then-answer pipeline against a standard RAG prompt baseline, both on med-llm-2026-03-15 at temperature 0.2; full prompts are in Appendix A. [Pinned model, settings, prompts published — ✓] Both prompts received equal tuning effort (three iterations each on the development split). [Equal effort — ✓]

Our primary metric, pre-registered in our eval plan, is unsupported claims per answer, measured by an LLM judge (judge-llm-2026-02-01, temperature 0, prompt in Appendix B) that decomposes answers into atomic claims and checks entailment against retrieved articles. [Primary metric pre-registered; judge specified — ✓] On a 150-item human-labeled validation set, the judge agreed with the majority human label on 83% of claims vs. 86% human–human agreement; position-bias checks showed a 4% flip rate. [Judge validation + bias check — ✓] Secondary metrics are completeness (1–3 rubric), abstention rate on the unanswerable slice, and cost per 100 answers. [Panel, not a single metric — ✓]

We run each system three times and report means with 95% bootstrap confidence intervals (10,000 resamples); differences on the primary metric are tested with paired permutation tests. [Runs, CIs, paired tests — ✓] Outputs that failed to parse (0.8% of baseline outputs) were counted as incorrect. [Failure handling disclosed — ✓]

A reviewer reading this excerpt can verify every checklist item without hunting through the paper. Notice what is not here: no superlatives ("significantly outperforms" appears nowhere — the numbers will speak in the results), no copied baseline numbers, no unreported settings. The tone throughout is auditable. That tone — more than any individual item — is what makes reviewers trust an eval section.

The results paragraph that follows should mirror this structure: a table with CIs (Chapter 8's pattern), per-slice rows, then a failure analysis paragraph ("Of 60 sampled errors from our method, 55% were retrieval misses on multi-article questions…"), and finally the limitations paragraph ("Our items are English-only and drawn from two clinics; performance on other specialties may differ"). Methods say what you did; results say what happened; the failure analysis says what you learned. Papers that include all three get cited; papers with only the first two get questioned.

What reviewers actually write: responding to eval critiques

Real reviewer comments on eval sections, with how to respond — ideally by fixing the paper, not just the rebuttal:

"The improvement over the baseline is marginal (1.8%)." Response: add CIs and the power analysis. If the CI is [+0.2, +3.4], concede the effect may be small and reframe: argue practical significance (the gain comes with 40% lower cost, or concentrates on the hard slice — show the slice table). If you can't make that case, consider whether the claim belongs in the paper at all. Never respond by adding more test items until p < 0.05.

"The baseline seems weak / undertuned." Response: this one is hard to rebut with words — you usually need to run the stronger baseline. In the revision, tune the baseline with the same budget, report both, and document the effort. One paragraph describing baseline tuning ("we performed 5 prompt iterations on the dev set, matching our method's tuning budget") defuses this better than any argument.

"Why wasn't human evaluation performed?" Response: if you have a validated judge, point to the validation numbers and the per-dimension agreement table — a validated judge is an answer to this question, not an evasion of it. If you don't, run the human study: 100–200 items on the primary dimension is the expected minimum for a strong claim.

"Results are reported on a single dataset." Response: add a second dataset or a robustness slice (paraphrased inputs, a second domain). If truly infeasible, say why and scope the claim down: "our results establish the effect on [dataset]; generalization to [other settings] remains future work." Reviewers accept scoped claims; they reject universal ones supported by one dataset.

"The paper doesn't discuss failure cases." Response: add the failure analysis — sample 50–100 errors, categorize, report. This is the highest-leverage revision per hour of work: it answers the reviewer's underlying question ("do the authors understand their own method?") directly.

The meta-strategy: treat every eval critique as a request for evidence, not as an attack. The rebuttal that works is rarely "the reviewer is wrong" — it's a new table, a new CI, a new slice analysis, or an honest scoping of the claim. Keep your eval artifacts (prompts, splits, scripts) organized precisely so you can produce these in the rebuttal window.

Key takeaways

  • Structure the eval section: claim → data → systems → metrics → protocol → results with uncertainty → failure analysis → limitations.
  • Publish full prompts, pin model versions, report decoding settings — these are non-negotiable for reproducibility.
  • Justify each metric in one sentence; declare one primary metric; report CIs and paired tests on every table.
  • Include a failure analysis (categorized sampled errors) and an honest limitations paragraph — they strengthen, not weaken, the paper.
  • Use the 14-point reproducibility checklist before every submission; choose strong baselines early; report null results.

Chapter 11 — The 10 Most Common Eval Mistakes (With Fixes)

Learning from the field's scars

Every mistake below appears regularly in published papers, workshop submissions, and industry eval reports. Each entry follows the same format: the mistake, why it happens, why it matters, and the concrete fix. Read this chapter as a pre-flight checklist for your own work.

Mistake 1: Reporting a single run as "the score"

Why it happens: One run is fast; the number looks clean; nobody asks for more — until a reviewer does. Why it matters: LLM outputs are stochastic (Chapter 1). A single run's score can easily sit 2–3 points from the true mean, which is larger than many claimed "improvements." Fix: Run everything ≥3 times with fixed decoding settings; report mean ± 95% CI (bootstrap, Chapter 8). Budget for this from the start — it is not optional polish.

Mistake 2: Tuning the prompt on the test set

Why it happens: The test set is right there, and each tweak visibly moves the number. It feels like progress. Why it matters: You are overfitting the prompt to those specific items. The reported score is optimistic, and the "improvement" often evaporates on fresh data. Fix: Split dev/test before you start iterating (Chapter 6). Touch the test set once, at the end. If you slipped and tuned on it, say so honestly and label the numbers as development results.

Mistake 3: Copying baseline numbers from other papers

Why it happens: Re-running baselines is work; the number is right there in Table 2 of the prior paper. Why it matters: Prompting, harness, model version, and decoding settings differ across papers. You end up comparing your carefully tuned 5-shot number against someone's 0-shot number — or against a different model version entirely. Fix: Re-run every baseline yourself under your exact protocol. If a baseline is truly infeasible to run, flag copied numbers explicitly with the caveat, and never make them the centerpiece comparison.

Mistake 4: Using an unvalidated LLM judge

Why it happens: Judge prompts are easy to write, and the scores look authoritative. Why it matters: An unvalidated judge may reward verbosity, prefer its own model family, or simply disagree with humans on your task — and you would never know. Your results section would then report the judge's biases as findings. Fix: Validate against 100–200 human-labeled items; require judge–human agreement near human–human agreement, per dimension (Chapter 4). Report the validation numbers. No validation, no judge-based claims.

Mistake 5: One metric on an open-ended task

Why it happens: A single number is simple to report and easy to optimize. Why it matters: Every metric is blind to something (Chapter 2) — ROUGE to factuality, EM to paraphrase, judge scores to their own biases. A single metric lets real regressions hide: your "improved" system may have traded factuality for fluency. Fix: Report a small panel (2–4 metrics) covering complementary dimensions, plus a human or validated-judge check on the dimension you care about most.

Mistake 6: Ignoring the failure analysis

Why it happens: Error analysis is tedious, and the headline number feels like the deliverable. Why it matters: Without it, you do not know why your method works — or whether it works for the reason you think. Reviewers probe exactly this gap: "Is the gain from better reasoning, or did the baseline just format outputs badly?" Fix: Sample 50–100 failures per system, categorize them with a second reader, and report the distribution (Chapter 10). Budget half a day; it is the highest-insight-per-hour activity in eval work.

Mistake 7: Claiming significance from an underpowered eval

Why it happens: The eval set has 150 items, the gap is 3 points, and it feels like a win. Why it matters: On 150 items, a 3-point gap is ~4 items flipping — noise. Publishing it adds a false "improvement" to the literature that someone else will waste months trying to build on. Fix: Do the power math before running (Chapter 8). If you cannot afford the items, shrink the claim: report the result as exploratory, lead with qualitative analysis, or collect more data.

Mistake 8: Contaminated or leaky eval data

Why it happens: Test items get used in few-shot prompts, discussed in team chats that end up in training data, or included in fine-tuning sets by accident. Why it matters: Once the model has seen the test items, your eval measures memorization, not capability. Scores inflate, and the inflation is invisible from inside. Fix: Keep test items out of prompts, training data, and public repos. Prefer private or newly written items for high-stakes claims. If contamination is possible, say so and discount accordingly.

Mistake 9: Moving the goalposts across systems

Why it happens: You give your method the good prompt, the tuned temperature, the careful output parsing — and run the baseline with defaults. Why it matters: The comparison then measures effort, not method quality. This is the "weak baseline" objection (Chapter 10), and it is fatal in review. Fix: Give every system the same tuning budget and the same protocol. Tune the baseline's prompt with the same care as your own. Document the effort per system.

Mistake 10: No error bars, no uncertainty, no humility

Why it happens: Bare numbers look decisive; intervals look wishy-washy. Why it matters: It is the opposite — bare numbers are the wishy-washy ones, because they hide how much could be noise. Reviewers in 2026 expect uncertainty quantification; its absence reads as either ignorance or concealment. Fix: CIs on every number, error bars on every chart, a limitations paragraph in every paper (Chapters 8, 10). Uncertainty honestly reported is a strength signal, not a weakness.

The meta-pattern

Notice what these ten share: every one is a way of accidentally fooling yourself — and then, through publication, fooling others. The field's eval crises (contamination, irreproducible gains, judge biases treated as findings) are not caused by dishonesty; they are caused by smart people skipping the boring parts. The fixes are all boring: more runs, held-out splits, re-run baselines, validated judges, reported uncertainty. Boring is what makes science work.

For your research: Pick the three mistakes you are most likely to make — be honest — and write their fixes into your project plan as scheduled tasks with dates ("validate judge against 150 human labels by Oct 20"). Mistakes you plan against do not happen; mistakes you merely intend to avoid do.

Self-audit worksheet

Score your current (or planned) project honestly: 0 = not done, 1 = partially done, 2 = done well. Total out of 20.

  1. Repeated runs. Have you run every reported number ≥3 times with fixed decoding settings? (Ch. 1, 8)
  2. Dev/test split. Is there a test set you have not tuned against? (Ch. 6, 11)
  3. Re-run baselines. Did every baseline run under your exact protocol, or did any number come from another paper? (Ch. 5, 11)
  4. Judge validation. If you use an LLM judge, is there a human-labeled validation set with reported agreement? (Ch. 4)
  5. Metric panel. Are you reporting ≥2 complementary metrics on every open-ended task? (Ch. 2, 11)
  6. Failure analysis. Have you categorized ≥50 errors from your method? (Ch. 10, 11)
  7. Power. Did the power math support your item count for the claimed effect? (Ch. 8)
  8. Contamination hygiene. Are test items out of prompts, training data, and public repos? (Ch. 6, 11)
  9. Equal effort. Did baselines get the same tuning budget as your method? (Ch. 10, 11)
  10. Uncertainty. Does every number carry a CI and every chart carry error bars? (Ch. 8, 10)

Interpretation. 18–20: submission-ready on eval grounds. 14–17: solid, with known gaps — fix them before submission and disclose the rest. 10–13: the eval needs a dedicated work cycle; do not submit yet. Below 10: you are exploring, not evaluating — label it as such and keep building.

The "fix this week" planner. Take your three lowest scores and convert each into a calendar task with a concrete deliverable: - Score 0 on #4 → "Validate judge: label 150 items, compute agreement — due Friday." - Score 1 on #2 → "Freeze test split: move 300 items to eval/test/, delete from prompts — due Wednesday." - Score 0 on #6 → "Failure analysis: sample 60 errors, categorize with a partner — due next Monday."

Re-run the audit monthly during an active project. The score should climb; if it doesn't, the eval is rotting while the method advances — the exact situation this book exists to prevent. Tape the worksheet to the wall next to the reproducibility checklist (Chapter 10). Between the two, there is very little a reviewer can ask that you haven't already answered.

Mistake patterns by project stage

Not all ten mistakes are equally likely at every stage. Here's where each one ambushes you — and what to prioritize when time is short:

Course project / hackathon (days). The killers are #1 (single run), #5 (one metric), and #10 (no uncertainty). You're moving fast, but three runs and a bootstrap CI cost minutes and separate a serious project from a demo. Skip #6 (full failure analysis) if you must — but write down three example failures you noticed; that's a mini failure analysis and it counts.

Workshop / short paper (weeks). Add #2 (test-set tuning) and #7 (underpowered claims) to the watch list. With small eval sets, the temptation to tune on test and to overclaim small gaps is strongest. Defenses: freeze the test split on day one, and do the power math before writing the abstract's numbers.

Conference paper (months). All ten are in play, but #3 (copied baselines), #4 (unvalidated judges), and #9 (unequal tuning effort) are the ones reviewers hunt for, because they indicate how seriously you took the comparison. Budget two full weeks for baselines alone — re-running and tuning someone else's method is unglamorous and absolutely decisive.

Production system (ongoing). #2 becomes #8's cousin: test data leaking into training/fine-tuning pipelines over time. And an eleventh mistake appears that this chapter didn't list: stopping evaluation after launch. The regression suite (Chapter 7) is the fix — eval is not a phase, it's a permanent organ of the system.

If you have one hour to de-risk any eval, spend it in this order: (1) 15 min: verify the test split was never tuned on (#2); (2) 15 min: add CIs to the headline numbers (#10, #1); (3) 15 min: check that baselines ran under the same protocol (#3, #9); (4) 15 min: sketch the failure analysis from errors you've already seen (#6). One focused hour catches the mistakes that sink the most papers.

Bonus: the eleventh mistake — teaching to the test

There's one mistake this chapter didn't number, because it's the slow-motion version of Mistake 2: optimizing your method against your own eval set for so long that the eval stops measuring the real task. It happens gradually — each prompt tweak is validated on dev, each dev-validated tweak ships, and after twenty iterations your method is exquisitely fitted to 300 questions that no longer represent the wild. The symptom: dev scores climb steadily while user complaints stay flat. The defenses are structural: refresh a portion of the dev set periodically with new items, keep the test set truly frozen and truly private, and — most honestly — track a live metric (user thumbs-down rate, task completion in production) alongside the eval suite. When the eval and the live metric diverge, believe the live metric and rebuild the eval. An eval set is a model of the task, and like all models, it decays — maintaining it is part of the job, not a one-time cost.

Key takeaways

  • The ten mistakes: single-run scores, test-set tuning, copied baselines, unvalidated judges, single metrics, skipped failure analysis, underpowered claims, contamination, unequal tuning effort, missing uncertainty.
  • Every fix is procedural and boring: repeat runs, split data, re-run baselines, validate judges, panel of metrics, error analysis, power math, hygiene, equal effort, error bars.
  • The meta-lesson: eval failures are self-deception failures; process beats intention — schedule the fixes.

Chapter 12 — Designing an Eval Plan for Your Project (Worked Template)

From principles to a plan

You have the concepts; now you need a document. An eval plan is a short, written plan — made before the final experiments — that says what you will measure, how, and what would count as success. It is the single highest-leverage artifact in eval work: it forces every decision in Chapters 2–9 to be made explicitly, it becomes the methods section of your paper nearly verbatim, and it protects you from post-hoc goalpost-moving (Chapter 11, Mistake 2).

This chapter gives you a fill-in template and then works it through a realistic example so you can see every box filled.

The template

Copy this into your project repo as eval/PLAN.md and fill it in:

# Eval Plan: [PROJECT NAME] — v1.0, [DATE]

## 1. Decision
What decision will these evals inform? (One paragraph.)

## 2. Claim
The measurable claim, in one sentence:
"Method M improves [dimension] on [task] by [expected amount] vs. [baseline]."

## 3. Data
- Eval set(s): [name, version, size, source]
- Construction: [how items were sourced/labeled; agreement rate]
- Splits: [dev N / test N; test held out since DATE]
- Stratification: [dimensions and counts per slice]

## 4. Systems
- Candidate: [model version, prompt location, decoding settings]
- Baselines: [each with same detail; tuning effort per system]
- Blinding: [how outputs are anonymized for raters/judges]

## 5. Metrics
- Primary metric: [ONE metric — pre-registered]
- Secondary metrics: [2-3, labeled exploratory]
- Per metric: implementation + one-sentence justification

## 6. Judges & raters
- Human eval: [N raters, background, training, rubric location,
  blinding, agreement target, disagreement rule]
- LLM judge: [model version, prompt location, validation sample size,
  judge-human agreement target, bias checks]

## 7. Statistics
- Runs per system: [≥3]
- Uncertainty: [bootstrap 95% CIs, 10k resamples]
- Tests: [paired permutation / McNemar on primary metric]
- Power: [effect size targeted, items needed, items available]
- Multiple comparisons: [how handled]

## 8. Regression & hygiene
- Suite layers: [smoke/slice/full triggers]
- Model pinning: [exact versions]
- Contamination notes: [any risks, mitigations]

## 9. Success criteria
- Ship/claim threshold: [e.g., "primary metric +3 points with p<0.05
  AND no secondary metric regresses beyond its CI"]
- Failure plan: [what you conclude if the threshold is not met]

## 10. Timeline & owners
- [Milestone]: [owner] by [date] (judge validation, human eval, final runs...)

Two fields do most of the work: #1 (Decision) keeps the eval honest about its purpose, and #9 (Success criteria) written in advance is what separates science from storytelling. If you cannot fill #9, you are not ready to run the final eval — you are still exploring, which is fine, but label it exploration.

Worked example: faithfulness fix for a medical RAG bot

Background. You are building a RAG assistant that answers patient questions from a corpus of 2,000 hospital help articles. Doctors flagged that the bot sometimes adds plausible-but-unsourced medical details. Your proposed fix: a "cite-then-answer" prompt plus a verification pass that drops unsupported sentences. You need to know: does the fix actually reduce unsupported claims, without hurting completeness?

1. Decision. Whether to ship the cite-then-answer pipeline as the default for the patient-facing bot. The deciding factors are faithfulness (must improve) and completeness (must not regress).

2. Claim. "Cite-then-answer reduces unsupported claims per answer on medical RAG questions by at least 30% relative, versus the standard prompt, with no significant completeness regression."

3. Data. Custom set "MedRAG-eval v1.0": 600 real anonymized patient questions sampled from clinic logs (2026-01 to 2026-03), stratified: single-article (250), multi-article synthesis (150), distractor present (100), unanswerable (100). Gold labels: relevant article IDs + key-points lists, written by two medical students independently (agreement: 88% on article IDs, adjudicated). Splits: 300 dev / 300 test, stratified; test frozen 2026-04-02.

4. Systems. Candidate: cite-then-answer pipeline on med-llm-2026-03-15, temperature 0.2, prompt in prompts/cite_then_answer.md. Baseline: standard RAG prompt, same model and settings, prompt in prompts/standard.md, tuned with equal effort (3 prompt iterations each on dev). Outputs stripped of formatting tells; order randomized for judges.

5. Metrics. Primary: unsupported claims per answer (count, lower is better) — validated LLM judge decomposing answers into atomic claims and checking entailment against retrieved articles. Secondary: completeness (1–3 rubric, judge), abstention rate on the unanswerable slice, answer EM on factoid subset. Justifications: unsupported-claim count directly measures the reported failure; completeness guards against the fix making answers uselessly terse.

6. Judges & raters. LLM judge: judge-llm-2026-02-01, temperature 0, prompt in prompts/judge_faithfulness.md. Validation: 150 dev items double-labeled by the two medical students; judge–human agreement 83% per-claim vs. 86% human–human; position bias checked (4% — acceptable, disclosed). Human spot-check: 100 test items × 2 raters on completeness.

7. Statistics. 3 runs per system (temperature 0.2 — small sampling variance; 3 runs enough). 95% CIs via 10k bootstrap over items. Primary test: paired permutation on unsupported-claim counts (items paired across systems). Power: targeting 30% relative reduction from a baseline mean of ~1.1 unsupported claims/answer — with 300 test items, power >90%. Multiple comparisons: primary metric tested at α=0.05; secondary metrics reported with CIs, no significance claims.

8. Regression & hygiene. Smoke suite (20 items) on every prompt edit; slice eval on dev on every pipeline change; full test run once at the end. Model versions pinned; test questions never appear in prompts or training; contamination risk low (private logs, unpublished items).

9. Success criteria. Ship if: unsupported claims drop ≥30% relative with p<0.05 AND completeness mean does not drop beyond its CI AND abstention rate on unanswerables ≥90%. If faithfulness improves but completeness regresses, the conclusion is "trade-off found, needs redesign" — not a ship.

10. Timeline. Judge validation by Apr 10 (you) · human spot-check protocol by Apr 14 (teammate) · final test runs Apr 18 · decision memo Apr 20.

Notice how every chapter of this book appears exactly once: dimensions and stratification (Ch. 6), judge validation (Ch. 4), human spot-check design (Ch. 3), metric justification (Ch. 2), power and paired tests (Ch. 8), regression layers (Ch. 7), contamination hygiene (Ch. 11), and the write-up structure (Ch. 10). The plan is the book, compressed into two pages.

Adapting the template

  • Course project / early exploration: Fill in sections 1–5 and 9 only; mark everything else "exploratory." The value is in the decision and success criteria, even when the statistics are light.
  • Agent evaluation: Add trajectory-level fields to section 5 (task success, step efficiency, tool accuracy — Chapter 9) and a sandbox-reset protocol to section 8.
  • Benchmark-style contribution: If your eval set itself is the contribution, expand section 3 into a full datasheet and add a section on release licensing.

Revisit the plan when reality intrudes — a failed judge validation, a smaller effect than expected — and record the changes with dates. A plan that evolves transparently is science; a plan silently rewritten after seeing results is fiction.

For your research: Your eval plan is also a collaboration tool. Share it with your advisor or teammates before the final runs and ask: "If these results come back positive, will you believe them?" Every objection they raise now is a reviewer objection you get to fix for free. The most expensive sentence in research is "we should have measured X" — spoken after the experiments are done.

Two more quick plans

The Chapter 12 template scales down as well as up. Two condensed examples:

Plan A — Course project (light). Decision: which of two prompts to use for a class demo classifying movie reviews. Claim: "Prompt B beats Prompt A on accuracy by ≥5 points on our 200-review set." Data: 200 IMDb-style reviews you label yourself (100 dev / 100 test). Systems: same model, temperature 0, prompts in the repo. Metrics: accuracy (primary — single, because the task is closed and the stakes are a demo). Statistics: 3 runs, bootstrap CI; power check: 100 test items can detect ~10-point gaps, so a 5-point claim is exploratory — the plan says so honestly. Success criterion: ship B if the gap clears +5 with non-overlapping CIs; otherwise report "no conclusive difference." Even this tiny plan has the two fields that matter: the decision and the pre-written success criterion.

Plan B — Web agent for form filling (agent-adapted). Decision: whether the agent is safe to pilot with real users. Claim: "The agent completes 80% of 50 form-filling tasks with zero irreversible-action violations." Data: 50 tasks in a sandbox CRM with scripted goal checks (field values correct = success). Systems: agent v2 vs. v1 (frozen baseline from last month's suite run). Metrics (primary + behavioral): task success rate (primary), argument accuracy per tool call, recovery rate on injected tool failures, violation count — where any violation fails the release regardless of success rate (safety as a gate, not a metric). Statistics: 3 runs per task (agents are stochastic); 95% CIs; violations reported as raw counts with task IDs. Hygiene: sandbox reset between tasks (no state leaking across tasks — a classic agent-eval bug), tool-call logs archived per run. Success criterion: success ≥80% AND violations = 0 AND no metric regresses beyond its CI vs. v1. The plan also names the failure conclusion in advance: "If success clears 80% but violations > 0, the outcome is 'not pilotable,' not 'mostly good.'"

Notice what both plans share with the full medical-RAG example: the decision is explicit, the primary metric is singular, the failure conclusion is written before the results exist, and safety/quality gates are stated as rules rather than vibes. The template is the same; only the depth changes. Start every project — a weekend hack or a dissertation — by filling in even the light version. It takes twenty minutes and prevents the two most expensive eval failures: measuring the wrong thing, and deciding what "good enough" means after seeing the numbers.

Running the plan: a six-week project timeline

A template is static; a project moves. Here's how the Chapter 12 template maps onto a realistic six-week timeline for a semester research project:

Week 1 — Decide and specify. Fill in template sections 1–3 (decision, claim, data). Write the eval spec: dimensions, target counts, sourcing plan. The deliverable is a one-page spec your advisor can critique. Most projects should spend more time here, not less — a week of specification saves a month of re-measurement.

Week 2 — Build the dev set. Source and label the development slice (aim for half your target items). Write gold labels with a partner; measure label agreement on the first 30 items before labeling the rest. Set up the JSONL schema (Chapter 6) and version control now, not later.

Week 3 — Baselines and harness. Implement the eval harness (runner, caching, aggregation). Run baselines first — all of them, three runs each — and freeze the baseline numbers. Write the smoke tests and behavioral assertions (Chapter 7). If you're using an LLM judge, draft the judge prompt this week.

Week 4 — Validate. Run the judge validation against human labels (Chapters 3–4): label 100–150 items, compute agreement, run bias checks. If the judge fails validation, you still have time to revise the rubric or fall back to human rating for the test set. Also finalize the test split freeze — no new tuning against it after this week.

Week 5 — Final runs. Execute the full test-set eval: all systems, three runs, all slices. No method changes this week — if you get the urge to tweak the prompt, write it down as future work. Run the failure analysis on sampled errors (Chapter 10) while the runs execute.

Week 6 — Write and audit. Fill in the results using the pre-registered analysis plan. Run the reproducibility checklist (Chapter 10) and the self-audit worksheet (Chapter 11). Write the limitations paragraph before the conclusion — it keeps the conclusion honest.

The two rules that hold the timeline together: (1) never move a deadline by cutting validation — cut scope instead (fewer slices, fewer baselines); (2) the test set stays frozen from Week 4 regardless of what the dev numbers suggest. Teams that follow these two rules produce evals that survive review; teams that don't produce evals that need to be redone.

Key takeaways

  • Write the eval plan before final experiments: decision, claim, data, systems, metrics, judges, statistics, hygiene, success criteria, timeline.
  • Pre-register ONE primary metric and write success/failure criteria in advance — this is what makes the result science rather than storytelling.
  • The worked example shows every chapter's ideas landing in one document; adapt the template's depth to your project's stage.
  • Record plan changes with dates; share the plan early and ask collaborators what would convince them.

Learning Dashboard

Metric selection table

Task shape Primary metric Why Watch out for
Multiple choice Accuracy Single right answer Prompt-format sensitivity
Short extractive QA EM + token F1 Partial credit for spans "Not guilty" vs "guilty" overlap
Machine translation sacreBLEU / chrF System-level ranking Weak sentence-level correlation
Summarization ROUGE-L + factuality check Legacy comparability Blind to hallucinations
Code with tests pass@k Behavioral, not textual Weak tests inflate scores
Open-ended QA/chat Validated judge or human rubric Captures graded quality Judge biases (Ch. 4)
RAG answers Unsupported-claim rate + citation recall Measures the central failure Judge error 5–15% per claim
RAG retrieval Recall@k, nDCG@k Diagnoses pipeline stage High recall + low precision = noise
Agents Task success rate (goal state) Behavior, not text Ignores cost/efficiency
Safety evals Refusal rate, violation rate Direct policy measurement Over-refusal hurts usefulness

Eval-method comparison table

Method Cost Speed Best for Key risk
String metrics (EM, BLEU, ROUGE) Free Instant Closed tasks, legacy comparison Blind to meaning/paraphrase
Human evaluation High Days–weeks Gold-standard quality judgments Rater disagreement, cost
LLM-as-judge Low–medium Hours Scaling rubric judgments Biases; must validate vs. humans
Benchmarks (MMLU, HumanEval…) Low Hours Positioning vs. field Contamination, saturation
Custom eval sets Medium Days Your actual claim Label error, under-sizing
Behavioral checks (tests, goal states) Low Minutes Code, agents, RAG pipelines Test/goal quality is everything

Reproducibility checklist (condensed)

  • [ ] Model versions pinned (exact version/date, no aliases)
  • [ ] Full prompts published (systems + judges), decoding settings reported
  • [ ] Data identified (benchmark version/harness or custom datasheet); test split held out
  • [ ] One primary metric pre-registered; each metric justified in one sentence
  • [ ] ≥3 runs per system; means with 95% bootstrap CIs; paired tests on the primary metric
  • [ ] Baselines re-run under identical protocol with equal tuning effort
  • [ ] Human evals: rubric shared, raters described, blinding stated, agreement reported
  • [ ] LLM judges: validated vs. 100–200 human labels; biases checked and disclosed
  • [ ] Failure analysis: sampled errors categorized and reported
  • [ ] Limitations paragraph: coverage gaps, contamination risks, what could overturn the conclusion
  • [ ] Code, prompts, and seeds released

Mistake / fix matrix

# Mistake Fix
1 Single run reported as "the score" ≥3 runs; mean ± 95% CI
2 Prompt tuned on the test set Dev/test split; touch test once
3 Baseline numbers copied from other papers Re-run all baselines identically
4 Unvalidated LLM judge Validate vs. 100–200 human labels
5 One metric on open-ended tasks Panel of 2–4 complementary metrics
6 No failure analysis Categorize 50–100 sampled errors
7 Underpowered significance claims Power math first; size eval to effect
8 Contaminated/leaky test data Keep test items private and unpublished
9 Unequal tuning effort across systems Same protocol and budget for all
10 No uncertainty reported CIs, error bars, limitations paragraph

References

[1] D. Hendrycks et al., "Measuring massive multitask language understanding," in Proc. Int. Conf. Learn. Representations (ICLR), 2021.

[2] M. Chen et al., "Evaluating large language models trained on code," arXiv:2107.03374, 2021.

[3] K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, "BLEU: a method for automatic evaluation of machine translation," in Proc. 40th Annu. Meet. Assoc. Comput. Linguistics (ACL), Philadelphia, PA, USA, 2002, pp. 311–318.

[4] C.-Y. Lin, "ROUGE: A package for automatic evaluation of summaries," in Text Summarization Branches Out: Proc. ACL Workshop, Barcelona, Spain, 2004, pp. 74–81.

[5] L. Zheng et al., "Judging LLM-as-a-judge with MT-Bench and Chatbot Arena," in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 36, 2023.

[6] P. Rajpurkar, J. Zhang, K. Lopyrev, and P. Liang, "SQuAD: 100,000+ questions for machine comprehension of text," in Proc. Conf. Empir. Methods Nat. Lang. Process. (EMNLP), Austin, TX, USA, 2016, pp. 2383–2392.

[7] M. Joshi, E. Choi, D. Weld, and L. Zettlemoyer, "TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension," in Proc. 55th Annu. Meet. Assoc. Comput. Linguistics (ACL), Vancouver, Canada, 2017, pp. 1601–1611.

[8] A. Wang et al., "GLUE: A multi-task benchmark and analysis platform for natural language understanding," in Proc. Int. Conf. Learn. Representations (ICLR) Workshop, 2019.

[9] P. Liang et al., "Holistic evaluation of language models," arXiv:2211.09110, 2022.

[10] J. Wei et al., "Chain-of-thought prompting elicits reasoning in large language models," in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 35, 2022.

[11] T. B. Brown et al., "Language models are few-shot learners," in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 33, 2020, pp. 1877–1901.

[12] S. Es, J. James, L. Espinosa-Anke, and S. Schockaert, "RAGAs: Automated evaluation of retrieval augmented generation," arXiv:2309.15217, 2023.


Glossary

  • Abstention — When a system declines to answer (e.g., "I don't know") because the available information is insufficient. A key behavior to measure in RAG and QA evals.
  • Benchmark — A fixed, shared test set plus scoring rule used to compare models (e.g., MMLU, HumanEval).
  • Blinding — Hiding from raters which system produced each output, so scores reflect quality rather than expectations.
  • Bootstrap — A resampling technique for estimating uncertainty: repeatedly resample your data with replacement and observe the spread of the resulting scores.
  • Bradley–Terry model — A statistical model that converts pairwise win/loss votes into a ranking with ratings, commonly used with pairwise LLM judging.
  • Citation precision/recall — For RAG: precision = fraction of answer claims supported by retrieved passages; recall = fraction of reference key facts present with support.
  • Cohen's kappa — An agreement statistic for two raters on categorical labels, corrected for chance agreement.
  • Confidence interval (CI) — A range of values within which the true score plausibly lies (e.g., 95% CI). Always report with point estimates.
  • Contamination — When benchmark or test items leak into a model's training data, inflating scores via memorization.
  • Decoding settings — Generation parameters such as temperature, top-p, and max tokens. Part of the system under test; must be reported.
  • Effect size — The actual magnitude of a difference (e.g., +2.5 points), as opposed to its statistical significance.
  • Exact match (EM) — A metric giving credit only when the prediction is identical to the reference after normalization.
  • Faithfulness — The degree to which a generated answer's claims are supported by its source material. The central RAG quality dimension.
  • Fleiss' kappa — Chance-corrected agreement for three or more raters on categorical labels.
  • Gold label — The reference answer or expected behavior for an eval item, against which outputs are scored.
  • Golden output — A stored known-good output used in regression testing; new outputs are compared for meaning preservation, not string equality.
  • Hallucination — A fluent, confident model statement that is false or unsupported. The failure mode faithfulness metrics target.
  • HumanEval — A 164-problem Python code-generation benchmark scored with pass@k (Chen et al., 2021).
  • Intraclass correlation coefficient (ICC) — An agreement measure for continuous scores across raters.
  • LLM-as-judge — Using a language model to score or compare outputs according to a rubric. Must be validated against human labels.
  • MMLU — Massive Multitask Language Understanding: ~16k multiple-choice questions across 57 subjects (Hendrycks et al., 2021).
  • Nondeterminism — The property that identical prompts can produce different outputs across runs, due to sampling and system-level variation.
  • p-value — The probability of seeing a difference as large as observed if the systems were truly equal. Not a measure of importance.
  • Pairwise comparison — A judging format where the judge picks the better of two answers (or ties), aggregated into rankings.
  • pass@k — The probability that at least one of k generated samples solves a problem (used with unit tests for code).
  • Position bias — An LLM judge's tendency to favor whichever answer is presented first (or second) in pairwise comparisons.
  • Power (statistical) — The probability that your eval will detect a real difference of a given size. Determines required sample size.
  • Primary metric — The single metric declared in advance as the basis for your main claim; protects against p-hacking.
  • Prompt sensitivity — The tendency of small prompt changes to cause large behavior/score changes.
  • RAG (retrieval-augmented generation) — A system that retrieves relevant passages and conditions generation on them.
  • Recall@k — Of the gold relevant items, the fraction appearing in the top-k retrieved results.
  • Reference-guided grading — A judging setup where the judge compares the candidate against a reference answer.
  • Regression testing — Re-running a fixed eval suite after changes to catch silent degradation in behavior or scores.
  • ROUGE — A family of n-gram-overlap metrics for summarization, emphasizing recall of reference content (Lin, 2004).
  • Rubric — A written scoring guide with named dimensions, anchored scales, and edge-case rules. The foundation of human and judge-based evals.
  • Saturation — When top models all score near a benchmark's ceiling, so it no longer discriminates between them.
  • Self-preference bias — An LLM judge's tendency to favor outputs written in its own style or by itself.
  • Stratification — Dividing an eval set into slices (by difficulty, type, domain) and sampling/reporting per slice.
  • Token-level F1 — Harmonic mean of token precision and recall between prediction and reference bags of words.
  • Trajectory — The sequence of actions (tool calls, observations) an agent takes in an environment, ending in a final state.

Practice Exercises

Exercise 1 — Write a rubric. Pick a task you care about (e.g., answering factual questions, summarizing news articles). Write a 3-dimension rubric with anchored 1–3 scales and at least three edge-case rules, following the Chapter 3 template. Then score 5 real model outputs with it and note every place the rubric was ambiguous — revise accordingly.

Exercise 2 — Metric justification. Choose three metrics from Chapter 2 for a task of your choice. For each, write the one-sentence justification ("We use X because… and acknowledge it misses…"). Then write one paragraph on which metric you would pre-register as primary and why.

Exercise 3 — Compute agreement. Have a friend score 30 items with your Exercise 1 rubric while you score the same 30 independently. Compute Cohen's kappa per dimension with the Chapter 3 code. Where kappa is below 0.6, discuss one disagreement and rewrite the anchor that caused it.

Exercise 4 — Validate a judge. Take the 30 doubly-labeled items from Exercise 3. Write a judge prompt (Chapter 4 template) and run it on the same items. Compute judge–human agreement per dimension and compare with your human–human kappa. Is the judge a valid substitute for any dimension? Write a two-paragraph verdict.

Exercise 5 — Build a 50-item eval set plan. For your own project (or a hypothetical RAG bot), write the eval spec from Chapter 6: claim, 3–4 capability dimensions with target counts summing to 50, sourcing plan, gold-label format, and labeling protocol. Identify the two hardest dimensions to source items for and propose a solution for each.

Exercise 6 — Bootstrap by hand. Take any 100 binary scores (e.g., correct/incorrect on 100 questions). Implement the Chapter 8 bootstrap function from scratch (no copy-paste — type it) and report the mean with 95% CI. Then halve the data to 50 items and re-run: how much wider is the interval? Write down what this implies for your project's eval size.

Exercise 7 — Power analysis. Your baseline scores 78% and you hope your method reaches 83%. Using the Chapter 8 table (or the statsmodels snippet), determine how many items you need for 80% power. If you only have 200 items, what is the smallest effect you could reliably detect? Write the honest claim you could make with 200 items.

Exercise 8 — Regression suite sketch. For a prompt-based application you use or are building, design the three-layer regression suite from Chapter 7: list 10 smoke-test items, define two behavioral assertions with code, and set a tolerance-band procedure. Write it as a eval/PLAN.md section.

Exercise 9 — RAG pipeline diagnosis. Given these numbers — recall@5 = 0.55, faithfulness = 0.90, completeness = 0.62 — where is the bottleneck, and what would you fix first? Justify with the Chapter 9 improvement loop. Then construct a second scenario (different numbers) where the bottleneck is elsewhere, and explain how the fix differs.

Exercise 10 — Full eval plan. Fill in the complete Chapter 12 template for a real or realistic project: decision, claim, data, systems, metrics (one primary), judges/raters, statistics, hygiene, success criteria, timeline. Share it with a peer and ask: "If the results come back positive, would you believe them?" Record their objections and revise the plan.


Selected Exercise Solution Sketches

Exercise 6 sketch. With 100 binary scores at, say, 78% accuracy, the bootstrap 95% CI will be roughly [70%, 86%] — a ±8-point band. Halving to 50 items widens it to roughly ±11–12 points. The implication: on 50 items, even a 10-point "improvement" (5 items flipping) sits inside the noise band. If your project's eval has fewer than 100 items, your honest claims are limited to large effects (15+ points) or qualitative findings — size the claim to the CI, not to your hopes.

Exercise 7 sketch. For 78% vs. 83% (a 5-point gap) at 80% power, α = 0.05, paired items: you need roughly 700–900 items (use the Chapter 8 table or the statsmodels snippet for your exact numbers). With only 200 items, the smallest reliably detectable gap is about 10 points. So the honest claim with 200 items isn't "we improved by 5 points" — it's either "we improved by ≥10 points" (if the data shows it) or "exploratory results suggest a ~5-point gain; a larger eval is needed to confirm." Writing that sentence feels like losing; it's actually what separates publishable work from noise.

Exercise 9 sketch. recall@5 = 0.55 is the bottleneck: the generator can only be faithful to what it was given, and nearly half the needed passages are missing. Fix retrieval first (chunking, hybrid search, query rewriting) before touching the generator — prompt-engineering faithfulness at 0.55 recall is wasted effort. (Second scenario: recall@5 = 0.92 but faithfulness = 0.61 — now retrieval is fine and the generator is adding unsupported claims; fix with citation-forcing prompts or an answer-then-verify pass.) The rule: evaluate in pipeline order, fix the earliest failing stage.


End of Book 20 — Evaluating and Testing LLM Applications. AstolixGen Learning Series.


Afterword: The Eval Mindset

If this book had to survive as a single paragraph, it would be this: an eval is a decision procedure, not a number. Every choice — the metric, the judge, the items, the runs, the statistics — is justified by the decision it informs, and every number is reported with the uncertainty that keeps it honest. The researchers whose evals you trust aren't the ones with the biggest benchmarks or the fanciest judges; they're the ones whose methods section reads like an audit trail, whose failure analyses show they understand their own system, and whose limitations paragraphs answer your objections before you raise them.

Three habits will carry you further than any technique in these pages. First, write the plan before the experiment — the decision, the primary metric, the success criterion. Thirty minutes of writing saves thirty days of re-running. Second, distrust your own numbers on a schedule — re-validate the judge, refresh the dev set, check the live metric against the eval suite. Evals decay; maintenance is the job. Third, publish the boring parts — the prompts, the seeds, the labeling protocol, the null results. The field's credibility is built from exactly these unglamorous artifacts, and every one you share makes the next researcher's work easier.

The models will keep changing — new architectures, new benchmarks, new failure modes. The discipline in this book doesn't depend on any of them. Operational definitions, repeated measurement, reported uncertainty, validated instruments, honest limitations: that's the scientific method, applied to machines that are very good at sounding right. Learn it once, and every future system — whatever it looks like — becomes evaluable.

Now go measure something, and report the error bars.