Model Evaluation: Accuracy, Precision, Recall

Book 5 of 10 — AstolixGen Learning Series (Detailed Edition) · For researcher and publication students

Cover


About This Book

A model that has not been measured properly is not a result — it is a guess. This book teaches you how to measure machine learning models honestly and convincingly. You will learn what accuracy, precision, and recall really mean, why the famous "99% accuracy" claim is often meaningless, how to read a confusion matrix like a diagnostic report, and how to design an evaluation plan that reviewers trust. Everything is explained with small, hand-computable examples so you can follow every number yourself.

Learning objectives: - Explain why evaluation is the core of any ML research claim - Compute accuracy, precision, recall, and F1 score by hand from a confusion matrix - Read a confusion matrix and diagnose what a model gets wrong - Describe what a ROC curve and AUC score mean, and when they mislead - Compute MAE, RMSE, and R-squared for regression problems - Design a k-fold cross-validation setup, including stratified folds - Explain why a single train/test split is not enough evidence - Apply basic statistical significance tests when comparing two models - Choose the right metrics for imbalanced data and rare events - Perform error analysis by slicing results into meaningful groups - Read results sections in papers critically and spot weak evaluation - Write a complete, pre-registered evaluation plan for your own project


Learning Dashboard

Concept Definition (one line) Example Use in research
Accuracy Fraction of all predictions that are correct 95 of 100 predictions right = 95% First-line metric in balanced classification papers
Precision Of predicted positives, fraction truly positive 8 of 10 flagged emails are spam Report when false alarms are costly
Recall (sensitivity) Of actual positives, fraction found Found 8 of 10 sick patients Report when missing a case is costly
F1 score Harmonic mean of precision and recall P=0.8, R=0.6 → F1=0.69 Single summary metric under imbalance
True positive (TP) Correctly predicted positive case Sick patient flagged sick Count in confusion matrix
False positive (FP) Negative case wrongly flagged positive Healthy patient flagged sick Cost of false alarms
True negative (TN) Correctly predicted negative case Healthy patient cleared Count in confusion matrix
False negative (FN) Positive case wrongly missed Sick patient cleared Cost of misses
Confusion matrix Table of TP, FP, TN, FN 2×2 table for a classifier Diagnose exactly where errors happen
Specificity Fraction of actual negatives correctly cleared 95 of 100 healthy cleared Complement to recall in medical papers
ROC curve Plots recall vs false-positive rate across thresholds Curve hugging top-left = good Threshold-independent classifier comparison
AUC Area under the ROC curve (0–1) 0.9 = strong separator Common single-number ranking metric
Decision threshold Score cutoff that turns scores into yes/no Flag if score > 0.5 Tune it; never accept the default blindly
Precision-recall curve Precision vs recall across thresholds Better than ROC under heavy imbalance Fraud/rare-disease evaluation
MAE Mean absolute prediction error Off by 2 units on average Robust regression error in original units
RMSE Root of mean squared error Heavily penalizes big misses Regression where large errors hurt most
R-squared Fraction of target variance explained 0.85 = 85% explained Report alongside error for context
Train/validation/test split Data divided for fitting, tuning, final check 60/20/20 Prevents self-deception about performance
Cross-validation Rotate which part of data tests the model 5-fold: test each fifth once Stable estimate on small datasets
Stratified sampling Keep class ratios in every fold/split Each fold has same sick rate Required for imbalanced classes
Overfitting Great on training data, poor on new data 99% train, 70% test The failure evaluation must catch
Data leakage Test information sneaks into training Scaling before splitting Invalidates results; reviewers check for it
Generalization Performance on unseen data Test-set score The only number that counts as a claim
Baseline Simple reference model to beat "Predict the majority class" Proves your model adds value
Statistical significance Difference unlikely due to chance p < 0.05 on paired test Decides if Model A truly beats Model B
Balanced accuracy Mean of recall on each class 0.5 + 0.9 → 0.7 Fair accuracy under imbalance
Matthews correlation Correlation between truth and prediction 1 = perfect, 0 = random Strong single metric for 2-class imbalance
Error analysis Manual study of the model's mistakes 40% of errors are blurry images Turns a metric into insight and next steps
Ablation study Remove parts of the model, re-measure Dropping feature X costs 3% F1 Shows which ideas actually matter
Pre-registration Fix metrics and plan before running "We will report F1 on the test set" Protects you from cherry-picking

Roadmap of the chapters. Chapters 1–2 set the stakes: evaluation is the evidence behind every claim, and accuracy alone can lie. Chapters 3–5 build the classification toolkit: precision, recall, and F1 (Chapter 3), the confusion matrix (Chapter 4), and ROC/AUC (Chapter 5). Chapter 6 covers regression metrics (MAE, RMSE, R-squared). Chapter 7 explains how to estimate performance reliably with cross-validation, and Chapter 8 adds the statistics needed to compare models fairly. Chapter 9 specializes in imbalanced data and rare events. Chapter 10 goes beyond numbers into error analysis. Chapter 11 teaches you to read other papers' results critically, and Chapter 12 brings everything together into an evaluation plan you can pre-register for your own thesis or paper.


Chapter 1: Why Evaluation Matters More Than Modeling

The model is a claim; evaluation is the evidence

Every machine learning project produces two things: a model and a claim about the model. The claim sounds like "my model detects crop disease with 94% accuracy" or "this predictor cuts forecasting error by 20%". The model itself — the code, the weights, the architecture — is only half the work. The other half is the evidence that the claim is true. That evidence comes from evaluation.

Here is a hard truth about research: reviewers, supervisors, and readers almost never judge your paper by how clever your model is. They judge it by whether they believe your measurements. A simple logistic regression with careful, honest evaluation is a stronger paper than a complex neural network with sloppy evaluation. The model is the idea; evaluation is what makes the idea science.

Think of it like medicine. A pharmaceutical company can invent a promising drug, but nobody prescribes it until clinical trials measure whether it works. The trial design — who is tested, what is compared, how results are reported — is what makes the drug trustworthy. In machine learning, your evaluation plan is your clinical trial. A weak plan means your results cannot be trusted, no matter how interesting the model is.

Training performance is not performance

The most common beginner mistake is reporting how well the model did on the data it learned from. This is called training accuracy or training error, and it tells you almost nothing about whether the model works.

Why? Because a model can simply memorize. Imagine a student who is given the exam questions in advance and memorizes the answers. A perfect score tells you nothing about what the student actually learned. A machine learning model with enough capacity can memorize every training example perfectly — 100% training accuracy — and still fail completely on new data. This is overfitting, and it is the reason evaluation must always happen on data the model has never seen: the test set.

Worked example. You train a classifier on 500 labeled images and test it two ways:

  • On the training images: 490 correct out of 500 → 98% training accuracy.
  • On 200 brand-new images: 142 correct out of 200 → 71% test accuracy.

The 98% is the model remembering. The 71% is the model thinking. Only the 71% is a research result. If you reported 98%, you would be publishing a memorization score, not a capability. Every time you see a suspiciously high accuracy in a paper, ask: was this measured on held-out data?

Evaluation answers the research question

Most MS and PhD projects in applied ML are, at their core, comparison questions: "Is method A better than method B for problem X?" or "Does adding feature Y improve the model?" You cannot answer these questions without a measurement procedure that is fair, repeatable, and appropriate to the problem.

Consider two researchers. Researcher A trains a model, tries ten different settings, picks the best result on the test set, and reports it. Researcher B fixes the plan in advance — the metric, the data split, the baselines — runs it once, and reports whatever comes out. Researcher B's number will usually be lower. Researcher B's paper is the honest one. Researcher A peeked at the test set ten times and effectively tuned the model on the test data. This is a form of data leakage, and it is one of the most common reasons papers fail review.

What good evaluation looks like

A trustworthy evaluation has five properties:

  1. Separation. Training, validation, and test data are strictly separate. The test set is touched once, at the end.
  2. Appropriate metrics. The metric matches the real goal (more on this in Chapters 2–6). Accuracy is not always the right choice.
  3. Baselines. You compare against simple, known methods — not against nothing. Beating "predict the most common class" is the minimum bar.
  4. Honest reporting. You report the metric you planned, on the split you planned, including the runs that looked bad.
  5. Uncertainty. A single number is not enough. Report variation across folds or runs so readers know how stable the result is (Chapters 7–8).

The cost of getting it wrong

Bad evaluation does not just weaken a paper — it misleads everyone who builds on it. If your reported 95% accuracy was actually measured with leakage, the next researcher who tries your method on clean data will get 75% and conclude the method is broken. Worse, in applied fields like medicine or agriculture, a wrongly evaluated model can lead to real-world decisions — missed diagnoses, wasted treatments — based on false confidence. Evaluation is where research meets responsibility.

The three questions every evaluation answers

A complete evaluation answers three questions, in order:

  1. Does it work at all? Compare against the trivial baseline (Chapter 2). If your model cannot beat "always predict the majority class," nothing else matters.
  2. Is it better than the alternatives? Compare against the strongest known methods with significance testing (Chapter 8). This is the core of most papers.
  3. Where does it fail? Error analysis and slicing (Chapter 10). This is where future work — yours or others' — begins.

Notice the order: absolute performance first, comparative performance second, failure modes third. Many student projects jump to question 2, comparing two fancy models, without establishing question 1. If both models are worse than a simple baseline, the comparison is theater. Always climb the ladder from the bottom.

Leakage: the silent killer

Data leakage means information from the test set influences training. It is called "silent" because nothing errors out — your code runs fine, your numbers look great, and your result is meaningless. Common forms:

  • Preprocessing before splitting. Scaling features or selecting features using the entire dataset, then splitting. The scaler "saw" the test data's range. Fix: fit preprocessing on training data only, then apply to test.
  • Duplicate or near-duplicate examples across splits. In medical imaging, two scans from the same patient can land in train and test. The model recognizes the patient, not the disease. Fix: split by patient (or by group), not by image. This is called grouped splitting.
  • Time travel. Training on data from 2024 to "predict" 2023. In any time-ordered problem (stocks, demand, disease spread), the test set must come after the training period. Fix: time-based splits (more in Chapter 7).
  • Test-set peeking during development. Tuning ten architectures and keeping the one with the best test score. The test set became a validation set. Fix: a validation set for decisions, test touched once.

Worked example. A student builds a churn predictor. She scales all features with the full dataset's mean and variance, then splits 80/20 and reports 89% accuracy. Her supervisor asks her to redo it: fit the scaler on the 80% only. Accuracy drops to 84%. The 5-point gap was leakage — the scaler had smuggled test-set statistics into training. The honest number is 84%, and the paper reports 84%. Painful for a week; protective for a career.

Evaluation debt: the cost of skipping ahead

Teams that postpone evaluation design accumulate evaluation debt — and it compounds. The typical story: a student trains models for two months, gets exciting numbers, starts writing, and only then discovers the test set was used for tuning, the baseline was mistuned, and the metric doesn't match the problem. Fixing it means re-running everything: new splits, re-tuned baselines, re-computed statistics. Weeks of work evaporate.

The alternative is the evaluation-first workflow:

  1. Write the one-page plan (Chapter 12) before touching a model.
  2. Build the evaluation pipeline first — data splits, metric functions, baseline scripts — and test it on a dummy model (even random predictions). If the pipeline runs end-to-end and produces sane numbers for a random model, it's ready.
  3. Only then train real models, slotting them into the finished pipeline.

This feels slower for the first week and is dramatically faster overall, because the pipeline never needs rebuilding mid-project. More importantly, numbers produced by a pre-built pipeline are trustworthy by construction — you literally cannot peek at the test set because the pipeline doesn't let you.

A final framing

Modeling asks "what can I build?" Evaluation asks "what do I know?" Research is the second question. Every technique in this book — from a humble accuracy check to nested cross-validation — is machinery for turning "my model got a good score" into "here is reliable evidence that this works, this much better, under these conditions." Readers may forget your architecture. They will remember whether they believed your numbers.

For your research: Before you train your first model, write down your evaluation plan in one page: what metric, what data split, what baselines, and what result would count as success. Show it to your supervisor. This single page will save you months of rework, because most evaluation mistakes are made before training starts, not after.

Key takeaways - Evaluation is the evidence behind every ML claim; reviewers judge it more than the model. - Training accuracy measures memorization; only held-out test performance is a result. - Peeking at the test set repeatedly is data leakage and invalidates the evaluation. - Good evaluation needs separation, appropriate metrics, baselines, honest reporting, and uncertainty. - Write your evaluation plan before you train, not after.


Chapter 2: Accuracy and Its Pitfalls

What accuracy actually measures

Accuracy is the simplest metric in machine learning: the fraction of predictions the model gets right.

Accuracy = (correct predictions) / (total predictions)

If a model classifies 100 emails and gets 93 right, accuracy is 93%. Simple, intuitive, and easy to explain. That is why it is the first metric everyone learns — and the first metric that gets misused.

Accuracy treats every mistake as equal. A wrong prediction on an easy case counts the same as a wrong prediction on a hard case, and a mistake on a common class counts the same as a mistake on a rare class. That equal treatment is fine when classes are balanced and all errors cost the same. It is dangerous when they are not.

The 99% accuracy trap

Now the famous trap. Read this example slowly, because understanding it will change how you read every paper that reports accuracy.

Worked example. A hospital screens 1,000 patients for a rare disease. Only 10 patients actually have the disease; 990 are healthy. A lazy programmer builds a "model" with one line of code: predict "healthy" for everyone, always.

Let's score it:

  • Correct predictions: all 990 healthy patients = 990 correct.
  • Total predictions: 1,000.
  • Accuracy = 990 / 1,000 = 99%.

Ninety-nine percent accuracy! Sounds excellent. But the model found zero sick patients. Every single person with the disease was sent home. As a medical tool, this model is worthless — worse than worthless, because the 99% number makes it look good.

This is the accuracy paradox: on imbalanced data, a useless model can score very high accuracy by always predicting the majority class. The accuracy number hides the failure because the rare class — the one you actually care about — barely affects the total.

How common is this situation? Extremely common. Fraud detection (99.9% of transactions are legitimate), defect detection in manufacturing, rare disease screening, intrusion detection in networks — in all of these, the interesting class is rare, and accuracy is the wrong metric. Chapter 9 is entirely about what to use instead.

When accuracy is fine

Accuracy is not always wrong. It is a perfectly good metric when two conditions hold:

  1. Classes are roughly balanced. If 50% of emails are spam and 50% are not, accuracy reflects real performance.
  2. All errors cost about the same. If a false alarm and a missed case are equally bad, accuracy's equal treatment is fair.

Under these conditions, accuracy is clear, comparable across papers, and easy to interpret. Many benchmark datasets (like balanced image classification sets) satisfy these conditions, which is why accuracy dominates leaderboards. The problem is that real-world problems often do not satisfy them — and researchers carry the leaderboard habit into domains where it misleads.

Accuracy hides the error distribution

Even on balanced data, accuracy hides which mistakes the model makes. Consider a three-class problem — classifying fruit as apple, orange, or banana — with 90% accuracy. That number does not tell you whether the errors are spread evenly or concentrated: maybe the model confuses oranges and apples constantly but never mistakes a banana. If your application cares specifically about oranges, the 90% is misleading. Accuracy compresses all information about error patterns into one number. The confusion matrix (Chapter 4) and error analysis (Chapter 10) recover what accuracy throws away.

A quick test you should always run

Before trusting any accuracy number — yours or someone else's — compute the majority-class baseline: what accuracy would you get by always predicting the most common class? If the dataset is 90% class A, the baseline is 90%. A model reporting 91% accuracy has barely beaten doing nothing. A meaningful result must clear the baseline by a margin that matters, and that margin must be statistically real (Chapter 8).

Worked example. A paper reports 87% accuracy on a customer-churn dataset. You check the dataset description: 85% of customers did not churn. The majority baseline is 85%. The model's real achievement is 2 percentage points above doing nothing — much less impressive than "87% accuracy" sounds. Always ask: accuracy compared to what?

Worked example: accuracy done right

Accuracy shines when its conditions hold. A sentiment classifier for product reviews: 1,000 test reviews, 520 positive, 480 negative (balanced), and a false alarm costs about the same as a miss (a mislabeled review just slightly skews an aggregate dashboard).

Confusion matrix: TP=440, FP=60, FN=80, TN=420.

  • Accuracy = (440+420)/1000 = 0.86
  • Precision = 440/500 = 0.88; Recall = 440/520 ≈ 0.846; F1 ≈ 0.863.

Here accuracy (0.86) agrees with F1 (0.863) — when classes are balanced and costs symmetric, the metrics tell the same story, and accuracy's simplicity is a virtue. Report it confidently, alongside the class balance so readers can verify the conditions hold.

Multi-class accuracy: micro vs macro

With more than two classes, "accuracy" still means correct/total — this is micro-averaged accuracy: every example counts equally. But there's a second way: compute accuracy (or recall) per class, then average the per-class scores — macro-averaging, where every class counts equally.

Worked example. Animal classifier, test set: 90 cats, 90 dogs, 20 rabbits (imbalanced). Correct: 85 cats, 80 dogs, 8 rabbits.

  • Micro accuracy = (85+80+8)/200 = 173/200 = 0.865.
  • Per-class recall: cat 85/90 ≈ 0.944, dog 80/90 ≈ 0.889, rabbit 8/20 = 0.40.
  • Macro recall = (0.944 + 0.889 + 0.40)/3 ≈ 0.744.

Micro says 86.5% — fine. Macro says 74.4% — the model is failing rabbits. Which is honest? Both are arithmetic; macro is honest about the weak class. When classes are imbalanced, always report macro-averaged scores next to micro — reviewers in applied fields specifically look for the macro number.

The accuracy reporting checklist

Before any accuracy number leaves your computer, run this:

  • [ ] Class distribution stated (so readers can compute the majority baseline themselves)
  • [ ] Majority-class baseline computed and beaten by a meaningful, significant margin
  • [ ] Metric defined (accuracy on which set? which classes? micro or macro?)
  • [ ] Test set held out and untouched during tuning
  • [ ] If imbalanced: precision, recall, F1 reported alongside

Five checks, two minutes, and your accuracy claims become bulletproof.

A second trap: the "almost perfect" classifier

The hospital example was dramatic (99% from doing nothing). Here's a subtler, more common version.

Worked example. Online ad clicks: 10,000 ad impressions, 120 clicks (1.2% positive). Model A is a real trained classifier; Model B always predicts "no click."

  • Model B accuracy = 9,880/10,000 = 98.8%.
  • Model A: catches 80 clicks (TP=80), misses 40 (FN=40), with 300 false alarms (FP=300), TN=9,580. Accuracy = (80+9,580)/10,000 = 96.6%.

Model B "wins" on accuracy (98.8% vs 96.6%) while learning nothing. Model A's recall = 80/120 ≈ 0.67, precision = 80/380 ≈ 0.21, F1 ≈ 0.32. Is Model A useful? For the ad business, maybe — catching two-thirds of clicks with a 21% hit rate on flagged impressions might beat showing ads randomly. The point: accuracy ranks the models backwards here. Whenever someone proposes "just use accuracy," run this mental test: would the always-negative model win? If yes, accuracy is disqualified.

When reviewers push back on accuracy

Expect this reviewer comment: "Accuracy is inappropriate given the class imbalance; please report precision, recall, and F1." The professional response isn't to argue — it's to add the metrics and, if accuracy truly fits your problem, a one-sentence justification: "Classes are balanced (48/52) and error costs symmetric, so accuracy is reported as the primary metric; precision/recall/F1 are in Table 3 for completeness." Reviewers accept accuracy when you've shown you considered the alternatives. What they reject is accuracy reported by default, with no evidence you thought about it.

One more habit: sanity-check with a dummy model

Before trusting any accuracy figure, run a dummy classifier — scikit-learn [6] even ships one (DummyClassifier): it predicts the majority class, or randomly, or proportionally to class frequencies. Your real model must beat the dummy by a clear, significant margin on your chosen metric. If it barely does, you don't have a modeling problem — you have a data or problem-framing problem, and no amount of tuning will fix it. The dummy takes thirty seconds to run and has saved countless researchers from months of polishing a model that learned nothing.

For your research: In your thesis or paper, never report accuracy on imbalanced data without also reporting precision, recall, and F1 (Chapter 3). State the class distribution of your dataset explicitly ("the positive class is 4.2% of samples") so reviewers can judge whether your metrics are appropriate. If a reviewer asks "what is the majority-class baseline?", you should already have the answer in your results section.

Key takeaways - Accuracy = correct predictions / total predictions; it treats all errors equally. - On imbalanced data, accuracy can be 99% for a model that finds nothing — the accuracy paradox. - Accuracy is valid when classes are balanced and error costs are symmetric. - Accuracy hides which errors the model makes; it compresses too much into one number. - Always compare accuracy against the majority-class baseline before celebrating it.


Chapter 3: Precision, Recall, and F1

Why we need more than accuracy

Chapter 2 showed that accuracy can lie when classes are imbalanced. The fix is to stop counting "correct" as one lump and instead ask two sharper questions:

  1. Precision: Of all the cases the model flagged as positive, how many were actually positive?
  2. Recall: Of all the cases that were actually positive, how many did the model find?

Precision punishes false alarms. Recall punishes misses. Together, they describe what accuracy hides: the two different ways a classifier can fail.

Figure 1: Precision vs recall visual metaphor

Figure 1. Precision vs recall as a fishing net: precision asks "of everything in the net, how much is fish?", recall asks "of all the fish in the water, how many did the net catch?"

The four outcomes

Every binary prediction falls into one of four boxes. Learn these four terms — they are the vocabulary of the entire book:

  • True Positive (TP): model says positive, truth is positive. A sick patient correctly flagged.
  • False Positive (FP): model says positive, truth is negative. A healthy patient wrongly flagged — a false alarm.
  • True Negative (TN): model says negative, truth is negative. A healthy patient correctly cleared.
  • False Negative (FN): model says negative, truth is positive. A sick patient missed — the dangerous error.

Definitions and formulas

Precision = TP / (TP + FP)

The denominator is everything the model called positive. Precision answers: "When the model raises the alarm, how often is it right?"

Recall = TP / (TP + FN)

The denominator is everything actually positive. Recall answers: "Of all the real cases, how many did the model catch?" (Recall is also called sensitivity or true positive rate.)

F1 score = 2 × (Precision × Recall) / (Precision + Recall)

The F1 score combines both into one number using the harmonic mean. Why the harmonic mean instead of a simple average? Because the harmonic mean punishes extreme imbalance: if precision is 1.0 but recall is 0.0, the simple average is 0.5 (looks okay), but the F1 is 0 (correctly terrible). A model must do reasonably well on both to get a good F1.

Worked example: a spam filter

A spam filter processes 40 emails. The truth: 15 are spam, 25 are not spam. The filter flags 12 emails as spam. Of those 12 flagged, 9 really are spam and 3 are legitimate emails wrongly flagged.

Step by step:

  • TP = 9 (spam correctly flagged)
  • FP = 3 (legitimate emails wrongly flagged)
  • FN = 15 − 9 = 6 (spam emails the filter missed)
  • TN = 25 − 3 = 22 (legitimate emails correctly left alone)

Check: 9 + 3 + 6 + 22 = 40. ✓

  • Precision = 9 / (9 + 3) = 9/12 = 0.75. When the filter says "spam", it is right 75% of the time.
  • Recall = 9 / (9 + 6) = 9/15 = 0.60. The filter catches 60% of all spam.
  • F1 = 2 × (0.75 × 0.60) / (0.75 + 0.60) = 2 × 0.45 / 1.35 = 0.90 / 1.35 = 0.667.
  • Accuracy = (9 + 22) / 40 = 31/40 = 0.775.

Notice the story the numbers tell: accuracy (77.5%) looks decent, but recall (60%) reveals the filter misses 4 out of 10 spam emails. If missing spam is costly, this filter is weaker than accuracy suggests.

The precision–recall trade-off

You cannot maximize both at once. A model that flags everything as positive has perfect recall (it misses nothing) but terrible precision (mostly false alarms). A model that flags only its most confident case has perfect precision but terrible recall. Moving the decision threshold — the score above which the model says "yes" — slides you along this trade-off:

  • Lower threshold → more positives flagged → higher recall, lower precision.
  • Higher threshold → fewer positives flagged → higher precision, lower recall.

Which one matters more? It depends on the cost of each error:

  • Recall matters more when missing a positive is expensive: disease screening, fraud detection, safety defects. A missed cancer is far worse than an extra test.
  • Precision matters more when false alarms are expensive: spam filters that delete real email, recommendation systems, expensive follow-up inspections. Crying wolf too often makes users ignore the system.

A good research paper states this choice explicitly: "we optimize for recall because a missed defect costs far more than a false alarm."

Worked mini-example of the trade-off

A disease model outputs risk scores for 5 patients whose true status is known: scores 0.9 (sick), 0.7 (sick), 0.5 (healthy), 0.4 (sick), 0.2 (healthy).

  • Threshold 0.5: flags the 0.9, 0.7, and 0.5 patients → TP=2, FP=1, FN=1 → Precision = 2/3 ≈ 0.67, Recall = 2/3 ≈ 0.67.
  • Threshold 0.35: flags 0.9, 0.7, 0.5, 0.4 → TP=3, FP=1, FN=0 → Precision = 3/4 = 0.75, Recall = 3/3 = 1.0.

Lowering the threshold caught the missed sick patient (recall 0.67 → 1.0) at the cost of one extra false alarm. In a screening context, that trade is worth it. Notice something interesting: precision went up here too, because the extra caught case was a true positive. The trade-off is a tendency, not a law — which is exactly why you must measure rather than assume.

Reporting the trio

In papers, you will usually see precision, recall, and F1 reported together (often written P/R/F1). Report them per class when classes matter differently, and state which class is the "positive" one. A common beginner error is reporting these numbers without saying which class they refer to — precision of what? Always label: "precision for the disease class = 0.82".

F-beta: weighting what matters

F1 treats precision and recall equally. Sometimes they aren't equal. The F-beta score generalizes F1:

F-beta = (1 + beta²) × (P × R) / (beta² × P + R)

  • beta = 1 → F1, balanced.
  • beta = 2 (F2) → recall counts twice as much as precision. Use for screening: missing a case is worse than a false alarm.
  • beta = 0.5 → precision counts twice as much as recall. Use when false alarms are expensive.

Worked example. Spam-filter numbers from earlier: P=0.75, R=0.60.

  • F2 = 5 × (0.75×0.60) / (4×0.75 + 0.60) = 5×0.45 / 3.60 = 2.25/3.60 = 0.625
  • F0.5 = 1.25 × 0.45 / (0.25×0.75 + 0.60) = 0.5625 / 0.7875 ≈ 0.714

Same model, different verdicts depending on what you value. State your beta and why — "we report F2 because missed spam (phishing) costs more than a false alarm" is a complete justification.

Precision and recall for many classes

For multi-class problems, compute precision/recall/F1 per class (each class vs. all others), then average:

  • Macro average: simple mean of per-class scores. Every class equal — exposes weak rare classes (Chapter 2's rabbit example).
  • Micro average: pool all TP/FP/FN across classes first, then compute. Every example equal — dominated by big classes. (For micro-F1 in multi-class, this equals overall accuracy.)
  • Weighted average: per-class scores weighted by class size. A compromise.

Rule: report macro alongside micro when classes are imbalanced; report per-class scores in an appendix or table when any single class matters (it usually does).

Precision@k and recall@k: ranked retrieval

Search engines and recommender systems don't output yes/no — they output ranked lists. Here the metrics adapt: of the top k results, how many are relevant (precision@k)? Of all relevant items, how many appear in the top k (recall@k)?

Worked example. A search returns 10 documents for a query with 6 relevant documents total in the collection. The top-5 results contain 3 relevant ones: precision@5 = 3/5 = 0.6, recall@5 = 3/6 = 0.5. If your paper evaluates ranking or retrieval, these are your metrics — plain precision/recall don't apply to ranked output.

Reporting checklist for precision, recall, and F1

When these three appear in your paper, make them unambiguous:

  • [ ] The positive class is named ("precision for the spam class").
  • [ ] The data split is stated (test set? which CV fold average?).
  • [ ] The decision threshold is given if one was tuned (default 0.5 otherwise).
  • [ ] Averaging method specified for multi-class (macro/micro/weighted).
  • [ ] The beta stated if using F-beta instead of F1.
  • [ ] Raw confusion counts available (table or appendix) so readers can recompute.

Six checkboxes. A reader should be able to reproduce your P/R/F1 from your paper alone — if they can't, the reporting is incomplete.

For your research: Choose your primary metric before running experiments, and justify it in one sentence tied to error costs ("recall is primary because a missed case costs a life; precision is reported for completeness"). Reviewers respect a metric choice that is argued from the problem, not picked because the number looked biggest. If you report F1, also report its components — a paper that shows only F1 hides whether the model is precise, thorough, or neither.

Key takeaways - Precision = TP/(TP+FP): when the model says yes, how often is it right? - Recall = TP/(TP+FN): of all real positives, how many were found? - F1 is the harmonic mean of the two; it punishes a model that is good at only one. - Lowering the decision threshold raises recall and usually lowers precision — choose based on error costs. - Always label which class your precision/recall/F1 refer to, and justify your primary metric from the problem.


Chapter 4: The Confusion Matrix, Deep Dive

What the matrix is

The confusion matrix is a table that shows, for every predicted class, how the truth was distributed. For a binary classifier it is a 2×2 grid of the four outcomes from Chapter 3:

Predicted Positive Predicted Negative
Actually Positive True Positives (TP) False Negatives (FN)
Actually Negative False Positives (FP) True Negatives (TN)

Rows are the truth; columns are the model's predictions. The diagonal (TP and TN) is where the model was right. Everything off the diagonal is an error, and the matrix shows you which kind of error, in which direction.

Every metric you have learned so far — accuracy, precision, recall, F1 — is just arithmetic on these four numbers. But the matrix itself contains more information than any single metric derived from it, because it preserves the structure of the mistakes.

Reading a matrix like a doctor reads a scan

Worked example. A clinic tests a disease-screening model on 200 patients. The true status: 50 sick, 150 healthy. The model's confusion matrix:

Predicted Sick Predicted Healthy
Actually Sick TP = 42 FN = 8
Actually Healthy FP = 15 TN = 135

From this one table, compute the full diagnostic:

  • Accuracy = (42 + 135) / 200 = 177/200 = 0.885
  • Precision = 42 / (42 + 15) = 42/57 ≈ 0.737 — when the model says "sick", it is right about 74% of the time.
  • Recall = 42 / (42 + 8) = 42/50 = 0.84 — the model catches 84% of sick patients.
  • F1 = 2 × (0.737 × 0.84) / (0.737 + 0.84) = 1.238/1.577 ≈ 0.785
  • Specificity = TN / (TN + FP) = 135 / (135 + 15) = 135/150 = 0.90 — the model correctly clears 90% of healthy patients.
  • False positive rate = FP / (FP + TN) = 15/150 = 0.10 — 10% of healthy patients get a false alarm.
  • False negative rate = FN / (FN + TP) = 8/50 = 0.16 — 16% of sick patients are missed.

Now read it as a story: the model is good at clearing healthy people (specificity 90%) but misses 16% of sick patients and raises false alarms on 10% of healthy ones. For a screening tool, the 8 missed sick patients are the critical problem — recall of 84% may not be good enough, and the next experiment should try lowering the threshold (Chapter 3) or collecting more training examples of the disease class.

That paragraph of insight came from staring at four numbers. No single metric gave it to you. This is why experienced researchers look at the matrix before they look at the metrics.

Two more derived metrics worth knowing

Specificity (true negative rate) is the mirror image of recall: of all actual negatives, how many were correctly cleared? In medical papers you will see "sensitivity and specificity" reported as a pair — sensitivity is just another name for recall. High sensitivity means few missed cases; high specificity means few false alarms.

Negative predictive value (NPV) = TN / (TN + FN): when the model says "healthy", how often is it right? Precision tells you about positive predictions; NPV tells you about negative ones. In screening, a high NPV is what lets a doctor confidently send a patient home.

Multi-class matrices

With three or more classes, the matrix becomes N×N. Rows are true classes, columns are predicted classes. The diagonal holds correct predictions; off-diagonal cells show exactly which classes get confused with which.

Worked example. A fruit classifier (apple / orange / banana) tested on 90 fruits, 30 of each:

Pred. Apple Pred. Orange Pred. Banana
True Apple 25 4 1
True Orange 6 22 2
True Banana 0 1 29

Accuracy = (25 + 22 + 29) / 90 = 76/90 ≈ 0.844. But the matrix reveals the real story: apples and oranges are confused with each other (4 + 6 = 10 cross-errors), while bananas are almost never mistaken (29/30 correct, only 1 error). The research insight writes itself: the model struggles to separate apples from oranges — perhaps they need color or texture features that distinguish them, or more training examples of those two classes. Per-class recall: apple 25/30 ≈ 0.83, orange 22/30 ≈ 0.73, banana 29/30 ≈ 0.97. Orange is the weak class. Your next experiment now has a target.

Normalizing the matrix

Raw counts are good for computing metrics, but for seeing patterns, normalize: divide each row by its row total so every row sums to 1. Each cell then reads as "of the true apples, what fraction were predicted as X?" This makes classes of different sizes visually comparable. Most papers show the normalized matrix as a heatmap figure. When you make yours, label rows "true" and columns "predicted" clearly — a matrix with unlabeled axes is a common figure mistake that confuses reviewers.

Common matrix mistakes

  1. Transposing rows and columns. Conventions differ (some put predictions on rows). Always state your convention in the caption: "rows = true labels, columns = predictions."
  2. Reporting only the diagonal. "Our model got 25, 22, 29 right" without the off-diagonal errors throws away the diagnostic value.
  3. Averaging without weighting. With imbalanced classes, a simple average of per-class scores can mislead; report both macro-average (equal weight per class) and the class distribution.

From counts to rates: seeing the matrix clearly

Raw counts answer "how many"; rates answer "how often," which is what you need when comparing classes of different sizes. Every row of the matrix converts to rates by dividing by the row total:

From the clinic example (Chapter 4's table: TP=42, FN=8, FP=15, TN=135):

Predicted Sick Predicted Healthy Row total
Actually Sick (50) 42/50 = 0.84 8/50 = 0.16 1.00
Actually Healthy (150) 15/150 = 0.10 135/150 = 0.90 1.00

Read across: "84% of sick patients caught, 16% missed; 10% of healthy patients falsely alarmed, 90% cleared." This row-normalized view is what you plot as a heatmap figure — the diagonal shows per-class recall at a glance, and pale off-diagonal cells scream for attention. Column-normalizing instead (dividing by column totals) gives precision-like rates: "of all patients flagged sick, 42/57 ≈ 74% truly were." Rows = recall view; columns = precision view. Know which one your figure shows and say so in the caption.

Cost matrices: when errors have price tags

Sometimes errors aren't just counts — they have costs. A cost matrix assigns a price to each cell, and the model's expected cost becomes the metric to minimize.

Worked example. Disease screening, costs estimated with the clinic:

Predicted Sick Predicted Healthy
Actually Sick 0 (correct) $5,000 (missed treatment, worsened illness)
Actually Healthy $200 (unnecessary follow-up test) 0 (correct)

Using the clinic's matrix (FN=8, FP=15): expected cost = 8 × $5,000 + 15 × $200 = $40,000 + $3,000 = $43,000 per 200 patients.

Now try a lower threshold that catches more: suppose it gives FN=3, FP=40. Cost = 3 × $5,000 + 40 × $200 = $15,000 + $8,000 = $23,000. The lower threshold nearly halves the expected cost — even though accuracy might barely change. When you can estimate costs, expected cost is the most honest metric of all, because it optimizes what the stakeholder actually pays. State your cost assumptions openly; they're debatable, and debatable beats hidden.

The matrix as a paper figure: a checklist

Your confusion-matrix figure will be scrutinized, so get the details right:

  • [ ] Rows labeled "true label," columns labeled "predicted label" (state the convention in the caption).
  • [ ] Normalized (row rates) for pattern-spotting; raw counts in the caption or appendix for recomputation. Many papers show both side by side.
  • [ ] Color scale with a legend — a heatmap without a scale is decoration.
  • [ ] Classes in a sensible order (group similar classes adjacently so confusion blocks are visible).
  • [ ] Readable in grayscale (reviewers print papers; test it).

One well-made matrix figure often communicates more than a page of metric tables. It's also the figure examiners point at during defenses — "tell me about this off-diagonal block" — so know every cell's story before you walk in.

For your research: Put the confusion matrix in your paper or thesis as a figure, not just the metrics in a table. Reviewers read matrices. In your results text, write one sentence per interesting off-diagonal pattern ("the model confuses class A with class B in 12% of cases, suggesting..."). This is the cheapest way to make your results section look like it was written by someone who understands their model — because it proves you looked.

Key takeaways - The confusion matrix (TP, FP, TN, FN) preserves the structure of errors that single metrics destroy. - Specificity = TN/(TN+FP); sensitivity is another name for recall; NPV evaluates negative predictions. - Multi-class matrices reveal which classes are confused — that is your next experiment's target. - Normalize rows to compare classes of different sizes; always label rows as truth and columns as predictions. - Include the matrix as a figure and narrate its off-diagonal patterns in your results.


Chapter 5: ROC Curves and AUC Explained

The threshold problem

Chapter 3 showed that moving the decision threshold trades precision against recall. But a classifier usually outputs scores (probabilities), not hard yes/no answers — and every choice of threshold gives a different (precision, recall) pair. How do you compare two models without arbitrarily picking one threshold? The ROC curve is the answer: it shows performance at every threshold at once.

Building the ROC curve

ROC stands for Receiver Operating Characteristic (a name inherited from radar engineering; the concept is pure statistics). The curve plots:

  • x-axis: False Positive Rate = FP / (FP + TN) — the fraction of actual negatives wrongly flagged.
  • y-axis: True Positive Rate = TP / (TP + FN) — which is just recall: the fraction of actual positives found.

Each point on the curve is one threshold. At a very high threshold, the model flags almost nothing: both rates are near 0 (bottom-left). At a very low threshold, it flags almost everything: both rates are near 1 (top-right). A good model's curve bows toward the top-left corner — high recall with few false alarms, across many thresholds.

Worked example: plotting five points by hand

Six patients, true status and model risk scores:

Patient Truth Score
A Sick 0.95
B Healthy 0.80
C Sick 0.70
D Healthy 0.45
E Sick 0.40
F Healthy 0.10

There are 3 sick and 3 healthy patients. Evaluate thresholds just below each distinct score:

  • Threshold 1.0 (flag nothing): TP=0, FP=0 → TPR = 0/3 = 0, FPR = 0/3 = 0 → point (0, 0).
  • Threshold 0.9 (flag A): TP=1, FP=0 → TPR = 1/3 ≈ 0.33, FPR = 0 → point (0, 0.33).
  • Threshold 0.75 (flag A, B): TP=1, FP=1 → TPR ≈ 0.33, FPR = 1/3 ≈ 0.33 → point (0.33, 0.33).
  • Threshold 0.6 (flag A, B, C): TP=2, FP=1 → TPR ≈ 0.67, FPR ≈ 0.33 → point (0.33, 0.67).
  • Threshold 0.42 (flag A, B, C, D): TP=2, FP=2 → TPR ≈ 0.67, FPR ≈ 0.67 → point (0.67, 0.67).
  • Threshold 0.3 (flag A–E): TP=3, FP=2 → TPR = 1.0, FPR ≈ 0.67 → point (0.67, 1.0).
  • Threshold 0.0 (flag all): TP=3, FP=3 → TPR = 1.0, FPR = 1.0 → point (1, 1).

Plot these points and connect them: the curve climbs steeply at first (good — real cases caught before false alarms pile up), then flattens. You have just drawn a ROC curve by hand, and every point is a threshold you could actually choose.

AUC: one number from the curve

AUC (Area Under the Curve) is the area beneath the ROC curve, between 0 and 1. It has a beautifully intuitive meaning: AUC is the probability that the model ranks a randomly chosen positive case higher than a randomly chosen negative case. An AUC of 0.9 means: pick one random sick patient and one random healthy patient — 90% of the time, the sick one gets the higher risk score.

Reference points:

  • AUC = 1.0: perfect separation. Every positive outscores every negative.
  • AUC = 0.5: random guessing. The curve is the diagonal line.
  • AUC < 0.5: worse than random (usually means labels are flipped somewhere).
  • Rough guide: 0.9+ excellent, 0.8–0.9 good, 0.7–0.8 fair, below 0.7 weak. (Domain-dependent — in hard problems, 0.75 can be a real achievement.)

For our six-patient example, the area under the plotted points (via the trapezoid rule) is approximately 0.78 — a fair classifier on a tiny dataset.

AUC is popular because it is threshold-independent: it summarizes ranking quality without committing to one operating point. Two models can be compared with a single number even if they would be deployed at different thresholds.

But AUC has a blind spot that matters enormously: it can look great on imbalanced data while the model is practically useless. Why? The false positive rate divides by the number of negatives — and when negatives are abundant, even a large absolute number of false alarms produces a small FPR. Example: 10 frauds among 10,000 transactions. A model with 500 false alarms has FPR = 500/9,990 ≈ 0.05 — the ROC curve looks fine. But precision is 10/(10+500) ≈ 0.02: only 2% of flagged transactions are real fraud. The investigators drown in false alarms while the ROC curve smiles.

Rule of thumb: on balanced or moderately imbalanced data, ROC-AUC is fine. On heavily imbalanced data (rare events), use the precision-recall curve and its area (PR-AUC) instead — precision exposes the false-alarm flood that FPR hides. Chapter 9 develops this fully.

Choosing an operating point

The ROC curve does not pick your threshold — you do, using costs. The "best" point is the one where the trade-off matches your problem: in screening, accept a higher FPR to push TPR near 1; in spam filtering, keep FPR tiny even if TPR suffers. A common formal choice is Youden's index (TPR − FPR, the point farthest above the diagonal), but a cost-based choice argued from the domain always beats a formula. State your chosen threshold and why — "we operate at FPR = 0.05 because the clinic can handle that many follow-ups" is a sentence reviewers love.

Comparing two curves: dominance and crossing

Plot two models' ROC curves together and three situations arise:

  1. Dominance: Model A's curve is entirely above B's. A is better at every threshold — the comparison is settled, no debate.
  2. Crossing: A is better at low FPR, B better at high FPR. Neither dominates. The winner depends on your operating region — which is exactly why you must state your operating point. A screening clinic operating at high recall should pick B; a spam filter needing tiny FPR should pick A.
  3. Near-identical: the curves overlap. The models are equivalent for practical purposes; pick the simpler/cheaper one.

Worked example. Two models on the same data: AUC(A) = 0.86, AUC(B) = 0.84. Tempting to declare A the winner — but the curves cross: at FPR = 0.05, A reaches TPR 0.70 while B reaches 0.78. If your application tolerates only 5% false alarms, B is the better model despite the lower AUC. AUC averages over all thresholds, including ones you'd never use. Never let a 0.02 AUC gap overrule the operating region.

Partial AUC: zooming into what matters

When only one region of the curve is relevant (e.g., FPR below 0.1 for a fraud system that can't afford many false alarms), compute the partial AUC — the area under the curve restricted to that FPR range, rescaled to 0–1. It answers "which model ranks best where I'll actually operate?" rather than "which is best on average over thresholds I'll never choose." Report it as "pAUC (FPR ≤ 0.1)" so readers know the window. Few student papers use it; those that do signal unusual maturity.

Building a PR curve by hand (mini-example)

Revisit Chapter 5's six patients (3 sick: A 0.95, C 0.70, E 0.40; 3 healthy: B 0.80, D 0.45, F 0.10). Compute (recall, precision) at each threshold:

  • Thr 0.9 (flag A): R = 1/3 ≈ 0.33, P = 1/1 = 1.0 → (0.33, 1.0)
  • Thr 0.75 (flag A,B): R ≈ 0.33, P = 1/2 = 0.5 → (0.33, 0.5)
  • Thr 0.6 (flag A,B,C): R = 2/3 ≈ 0.67, P = 2/3 ≈ 0.67 → (0.67, 0.67)
  • Thr 0.42 (flag A–D): R ≈ 0.67, P = 2/4 = 0.5 → (0.67, 0.5)
  • Thr 0.3 (flag A–E): R = 1.0, P = 3/5 = 0.6 → (1.0, 0.6)
  • Thr 0.0 (flag all): R = 1.0, P = 3/6 = 0.5 → (1.0, 0.5)

Plot recall (x) vs precision (y): the curve starts at (0.33, 1.0) and sags as recall rises — the visual signature of the precision–recall trade-off. The area under this curve (PR-AUC ≈ 0.72 here) is the number to report on imbalanced data.

For your research: Report ROC-AUC when comparing classifiers, but never report it alone on imbalanced data — pair it with the precision-recall curve or at least precision and recall at your chosen operating threshold. Include the ROC figure with the diagonal (random) line drawn, and mark your operating point on the curve. If a baseline's AUC is 0.52 and yours is 0.55, say so plainly instead of writing "our model achieves strong AUC" — reviewers divide by the baseline in their heads anyway.

Key takeaways - The ROC curve plots recall (TPR) vs false-positive rate across every threshold. - AUC = probability a random positive outranks a random negative; 0.5 is random, 1.0 is perfect. - AUC is threshold-independent, which makes it good for comparing models. - On heavily imbalanced data, ROC-AUC can hide a flood of false alarms — use precision-recall curves there. - The curve doesn't choose your threshold; domain costs do. Mark and justify your operating point.


Chapter 6: Regression Metrics — MAE, RMSE, and R-squared

A different kind of error

So far everything has been classification: right or wrong. Regression predicts a number — a house price, a temperature, a crop yield — where predictions can be close or far, not just right or wrong. Being off by 1 unit and being off by 100 units are both "wrong," but they are very different failures. Regression metrics measure how far off predictions are.

Setup for this chapter: let y be the true value, ŷ (y-hat) the prediction, and the error e = y − ŷ. With n test examples, the three core metrics are:

MAE (Mean Absolute Error) = (1/n) × Σ|y − ŷ|

Average the absolute errors. MAE is in the original units (dollars, degrees, kilograms) and treats every unit of error equally: being off by 4 is twice as bad as being off by 2.

MSE (Mean Squared Error) = (1/n) × Σ(y − ŷ)², and RMSE = √MSE

Squaring punishes large errors disproportionately: an error of 4 contributes 16, while an error of 2 contributes 4 — so one big miss hurts more than several small ones. RMSE takes the square root to return to original units, keeping the big-error penalty.

R-squared (R²) = 1 − (Σ(y − ŷ)²) / (Σ(y − ȳ)²)

R² compares your model against the dumbest possible predictor: always guessing the mean ȳ. The numerator is your model's squared error; the denominator is the mean-guessor's squared error. R² = 0.85 means your model eliminated 85% of the variance that mean-guessing leaves. R² = 0 means you are no better than predicting the average; negative R² means you are worse.

Worked example: predicting apartment rents

Five apartments, true monthly rent (in thousands) and model predictions:

Apt True y Pred ŷ Error (y−ŷ) |Error| Error²
1 50 48 2 2 4
2 60 63 −3 3 9
3 55 55 0 0 0
4 70 64 6 6 36
5 65 70 −5 5 25

Compute step by step:

  • MAE = (2 + 3 + 0 + 6 + 5) / 5 = 16/5 = 3.2 → on average, predictions are off by 3.2 thousand.
  • MSE = (4 + 9 + 0 + 36 + 25) / 5 = 74/5 = 14.8; RMSE = √14.8 ≈ 3.85.
  • For R²: mean ȳ = (50+60+55+70+65)/5 = 300/5 = 60. Denominator = (50−60)² + (60−60)² + (55−60)² + (70−60)² + (65−60)² = 100 + 0 + 25 + 100 + 25 = 250. R² = 1 − 74/250 = 1 − 0.296 = 0.704.

Interpretation: the model explains about 70% of rent variance, with typical errors around 3–4 thousand. Note RMSE (3.85) > MAE (3.2) — RMSE always exceeds MAE when errors vary in size, because squaring inflates the big ones (the 6-thousand miss on apartment 4). The gap between them tells you large errors are present.

Which metric should you use?

  • MAE when you want an honest "typical error" in real units and don't want outliers to dominate. Good default for reporting.
  • RMSE when large errors are genuinely more damaging than small ones — forecasting demand, predicting drug doses, structural load estimates. Also the standard when comparing against literature that uses it.
  • R² when you need context: an MAE of 3.2 is meaningless until you know whether rents vary by 5 or by 500. R² = 0.70 says the model captures most of the signal. Report R² alongside an error metric, not instead of one — R² has no units and can hide systematic bias.

A subtle trap: R² can be high even when the model is systematically biased. If every prediction is exactly 10 thousand too high, the errors are consistent, variance is still "explained," and R² looks fine — but every prediction is wrong by 10. Always plot predicted vs actual values; a straight line off the diagonal reveals bias that R² forgives.

Relative and percentage errors

When targets span orders of magnitude (predicting both small shops' and malls' revenues), absolute errors mislead: being off by 1,000 matters little on a million-scale target but hugely on a thousand-scale one. MAPE (Mean Absolute Percentage Error) = average of |y − ŷ|/|y| × 100% expresses error as a percentage. Use it when relative error matches the real cost — but beware: MAPE explodes when true values are near zero (division by ~0), and it penalizes over-prediction more than under-prediction. State its limits if you use it.

Residual plots: what the metrics don't say

MAE, RMSE, and R² compress all errors into numbers — but patterns in the errors reveal broken assumptions. A residual plot (residual = y − ŷ on the y-axis, predicted value on the x-axis) should look like random scatter around zero. Watch for:

  • Funnel shape (spread growing with prediction): heteroscedasticity — the model is worse on large values. Consider log-transforming the target.
  • Curved pattern: systematic bias — the model misses a nonlinearity. Consider polynomial or interaction features.
  • Clusters: the data has subgroups (e.g., two property types) the model doesn't know about. Consider adding the grouping feature.

Worked example. Our rent model from this chapter: residuals are +2, −3, 0, +6, −5 against predictions 48, 63, 55, 64, 70. Too few points for a real plot, but notice the largest residuals (±5, +6) sit at the largest predictions (64, 70) — a hint of funnel behavior. With 500 points instead of 5, that hint would become a visible fan, telling you the model struggles with expensive apartments specifically. The metric said "RMSE 3.85"; the plot says "and it's worse where it matters most."

Adjusted R-squared: penalizing complexity

R² never decreases when you add features — even useless ones. Add 50 random-noise features to a regression and R² creeps up, because the model exploits chance correlations. Adjusted R² corrects for this:

Adjusted R² = 1 − (1 − R²) × (n − 1) / (n − p − 1)

where n = examples, p = features. Adding a useless feature increases p without reducing error, so adjusted R² falls — exposing the trick.

Worked example. n = 100, R² = 0.70 with p = 5 features: adjusted = 1 − 0.30 × 99/94 = 1 − 0.316 = 0.684. Add 10 noise features (p = 15), R² rises to 0.72: adjusted = 1 − 0.28 × 99/84 = 1 − 0.330 = 0.670. Raw R² celebrated the noise; adjusted R² penalized it. When comparing regressions with different feature counts, adjusted R² is the fair metric.

MAPE: percentage errors, worked

Worked example. True monthly sales (thousands): 100, 200, 50. Predictions: 110, 180, 60.

Percentage errors: |100−110|/100 = 10%, |200−180|/200 = 10%, |50−60|/50 = 20%. MAPE = (10+10+20)/3 ≈ 13.3%. Intuitive for business stakeholders ("we're off by 13% on average"). But note the asymmetry: predicting double (100%) vs half (50%) the truth are equally wrong in business terms, yet MAPE scores them 100% vs 50%. If your errors are symmetric in business impact, prefer MAE/RMSE; if stakeholders think in percentages, report MAPE with this caveat stated.

Choosing between MAE and RMSE: a decision rule

Ask one question: "In my problem, is one error of size 10 as bad as ten errors of size 1?"

  • If yes → MAE. Total harm scales linearly with total error. Example: delivery-time prediction, where each late minute costs roughly the same.
  • If no, and the single big error is much worse → RMSE. Example: flood-level forecasting, where one 10-unit miss floods a town but ten 1-unit misses are rounding noise.
  • If unsure → report both. They disagree interestingly: when RMSE >> MAE, big errors dominate and deserve investigation (Chapter 10's residual slicing).

A related habit: normalize errors when comparing across datasets — RMSE divided by the target's standard deviation, or MAE as a percentage of the mean. "RMSE = 3.85 on rents averaging 60k" travels across papers better than a bare number.

For your research: For regression, report at least two metrics: one in original units (MAE or RMSE) and R² for context. Justify the choice from the problem's cost structure ("we report RMSE because a single large misprediction ruins a harvest plan"). Include a predicted-vs-actual scatter plot — reviewers check it for bias and heteroscedasticity (errors growing with target size) faster than they read your metric table. And never tune on the test set to shave 0.01 off RMSE; Chapter 7 shows the honest way to estimate it.

Key takeaways - MAE = average absolute error (original units, robust); RMSE = root mean squared error (punishes big misses). - R² = fraction of variance explained vs predicting the mean; 1 is perfect, 0 equals mean-guessing. - RMSE ≥ MAE always; the gap between them reveals the presence of large errors. - High R² can hide systematic bias — always plot predicted vs actual. - Report an error metric plus R², and choose the error metric from the real cost of being wrong.


Chapter 7: Cross-Validation — K-Fold, Stratified, and Why

The problem with a single split

Chapters 1–6 assumed you have a test set and measure once. But a single train/test split has a weakness: your result depends on which examples happened to land in the test set. With a lucky split, your model looks brilliant; with an unlucky one, it looks broken — and you would never know which one you got. On small datasets (a few hundred examples, typical for student projects), this luck can swing results by 5–10 percentage points. That is larger than most real improvements researchers claim.

Cross-validation fixes this by rotating: every example gets to be in the test set exactly once, and you average the results. The average is a far more stable estimate of true performance than any single split.

K-fold cross-validation, step by step

  1. Shuffle the dataset and divide it into k equal parts ("folds"). k = 5 and k = 10 are standard.
  2. For each fold i: train on the other k−1 folds, test on fold i, record the metric.
  3. Report the mean of the k scores — and the standard deviation, which tells you how stable the model is.

Figure 2: K-fold cross-validation diagram

Figure 2. Five-fold cross-validation: the data is divided into five parts; each part takes a turn as the test set (highlighted) while the rest trains the model.

Worked example. You have 100 labeled soil samples and run 5-fold cross-validation (each fold: 80 train, 20 test). The accuracy on each test fold:

  • Fold 1: 0.85 (17/20) · Fold 2: 0.90 (18/20) · Fold 3: 0.80 (16/20) · Fold 4: 0.95 (19/20) · Fold 5: 0.85 (17/20)

Mean = (0.85 + 0.90 + 0.80 + 0.95 + 0.85) / 5 = 4.35/5 = 0.87.

Standard deviation: deviations from mean are −0.02, +0.03, −0.07, +0.08, −0.02. Squared: 0.0004, 0.0009, 0.0049, 0.0064, 0.0004. Sum = 0.013; divide by 5 → 0.0026; √0.0026 ≈ 0.051.

Report: accuracy = 0.87 ± 0.05. Now compare honestly: a competing model scoring 0.89 ± 0.06 is not clearly better — the intervals overlap heavily. A model scoring 0.94 ± 0.03 probably is. Without the ±, you would have declared a winner on noise. This is why reviewers ask for error bars.

Stratified k-fold: mandatory for imbalanced data

Plain k-fold splits randomly, which can put almost all rare-class examples in one fold. Imagine 100 samples with only 10 positives: a random fold of 20 might contain zero positives, making recall undefined (division by zero) on that fold. Stratified k-fold preserves the class ratio in every fold — each fold gets ~2 positives out of 20, mirroring the full dataset. Cost: nothing. Benefit: every fold is evaluable. Always use stratified folds for classification, especially with imbalance. (The standard library implementations, e.g., scikit-learn's StratifiedKFold [6], make this a one-word change.)

Leave-one-out and repeated CV

  • Leave-one-out CV (LOOCV): k = n; each single example is its own test set. Maximally uses data, but expensive (n trainings) and the estimate has high variance — the training sets overlap almost completely, so the n scores are highly correlated. Rarely the right choice today.
  • Repeated k-fold: run 5-fold CV several times with different shuffles and average. Reduces the luck-of-the-partition further. Cheap insurance when data is small.

The nested trap: tuning vs evaluating

Here is the subtle failure mode that invalidates many student projects. Suppose you use cross-validation to choose the best hyperparameters (try 20 settings, pick the winner), then report the winner's CV score as your result. That score is optimistic: you selected the setting that got lucky on those particular folds. The selection process itself overfit.

The honest procedure is nested cross-validation:

  • Outer loop: splits data into train/test folds — this measures final performance.
  • Inner loop: within each outer training set, run another CV to select hyperparameters.
  • The outer test folds never influence any choice. The reported score is unbiased.

Nested CV costs k_outer × k_inner trainings (e.g., 5 × 3 = 15), which is fine for classical models. For deep learning, where one training run takes days, researchers instead use a single held-out validation set for tuning and a locked test set for the final number — the same principle (tuning data ≠ evaluation data), cheaper packaging. Either way, the rule from Chapter 1 holds: the data that makes decisions cannot be the data that judges them.

What to actually do

  • Dataset under ~10,000 examples, classical models → stratified 5- or 10-fold CV, report mean ± std.
  • Deep learning / expensive training → one stratified train/validation/test split (e.g., 70/15/15); tune on validation, report once on test.
  • Hyperparameter search involved → nested CV (or locked test set), so the reported number is honest.
  • Always set and report the random seed so others can reproduce your exact folds.

Worked example: building stratified folds by hand

Twelve patients: 4 sick (S), 8 healthy (H). We want 3 folds of 4, stratified — each fold must hold the 1:2 ratio (about 1–2 sick per fold).

Label patients S1–S4, H1–H8. Shuffle within each class, then deal round-robin:

  • Fold 1: S1, H1, H2, H3 → wait, that's only 1 sick in 4 (ratio 1:3). Better: deal sick patients first — S1→fold1, S2→fold2, S3→fold3, S4→fold1 — then deal healthy: H1→fold1, H2→fold2, H3→fold3, H4→fold1, H5→fold2, H6→fold3, H7→fold1, H8→fold2.
  • Fold 1: S1, S4, H1, H4, H7 (2 sick, 3 healthy — 5 members; uneven because 12 doesn't divide evenly with stratification, so folds are 5/4/3... let me redo cleanly.)

Simpler with 12 divisible by 3: aim for 4 per fold with ratio preserved as 1 sick + 3 healthy ≈ matches 4:8 = 1:2 closely enough (1:3 vs 1:2 — stratification is approximate with small numbers, and that's fine):

  • Fold 1: S1, H1, H2, H3
  • Fold 2: S2, H4, H5, H6
  • Fold 3: S3, H7, H8 + S4 → 2 sick, 2 healthy (the leftover S4 must go somewhere)

Final: Fold 1: {S1,H1,H2,H3}, Fold 2: {S2,H4,H5,H6}, Fold 3: {S3,S4,H7,H8}. Every fold contains both classes, so recall is computable on every fold — which is the whole point. With plain random splitting, Fold 3 could easily have been {H1,H2,H3,H4} with zero sick patients. Libraries do this dealing for you; now you know what they're doing.

Time series need special cross-validation

Standard k-fold shuffles data randomly — fatal for time-ordered data, because training on the future to predict the past is leakage (Chapter 1). Forward-chaining (rolling-origin) CV respects time:

  • Fold 1: train on months 1–6, test on month 7.
  • Fold 2: train on months 1–7, test on month 8.
  • Fold 3: train on months 1–8, test on month 9.

The test set always follows the training set in time, mimicking real deployment where you predict the future from the past. Never shuffle time series. If your data has any time component (and most real data does), ask whether random folds are valid before using them.

Five CV mistakes that invalidate results

  1. Preprocessing before splitting — scaling/feature selection on all data leaks test statistics (Chapter 1's example). Pipeline it inside each fold.
  2. Resampling before splitting — oversampling the minority class, then splitting, puts near-duplicates in both train and test. Resample within each fold's training portion only.
  3. Tuning on the test folds — picking the best of 20 settings by CV score, then reporting that score (Chapter 7's nested trap). Nested CV or a locked test set.
  4. Re-shuffling until it looks good — trying seeds until the mean crosses your target is p-hacking. Fix the seed in the plan.
  5. Ignoring grouping — multiple rows per patient/customer/store split across folds lets the model memorize entities. Use grouped (e.g., GroupKFold) splits.

How many folds? Practical guidance

  • k = 5: the workhorse. Each model trains on 80% of data; 5 trainings per model. Good default for datasets up to ~100k examples.
  • k = 10: each model trains on 90%; slightly less biased estimate, double the compute. Standard in classical ML benchmarking.
  • k = 3: when training is expensive (large models) and data plentiful. Less stable — compensate by repeating with different seeds.
  • Single split: when training takes days (deep learning). Accept the noisier estimate; mitigate with a large test set (thousands of examples).

More folds reduce bias (bigger training sets) but increase compute and the correlation between folds. There's no magic number — there's only the trade-off, stated honestly in your methodology.

For your research: Write your CV scheme in the methodology section with full specifics: "stratified 5-fold cross-validation (seed 42); hyperparameters selected by inner 3-fold CV; we report mean ± std of F1 across outer folds." That one sentence answers four reviewer questions at once. And never, ever re-shuffle your folds until the number looks good — that is p-hacking with extra steps, and it is exactly what pre-registration (Chapter 12) protects you from.

Key takeaways - A single split's result depends on luck; k-fold CV averages over k test rotations for a stable estimate. - Always report mean ± standard deviation — the ± is what lets you compare models honestly. - Use stratified folds for classification so every fold contains every class. - Tuning on the same folds you report from is optimistic; use nested CV or a locked test set. - Fix and report your random seed; reproducibility starts with the split.


Chapter 8: Comparing Models — Statistical Significance Basics

"Higher" is not the same as "better"

Model A scores 87.2% accuracy. Model B scores 86.1%. Is A better? Not necessarily — the 1.1-point gap might be noise from the particular data split, random initialization, or fold assignment. Chapter 7 gave you the ±; this chapter tells you how to decide whether the gap is real. This is statistical significance: asking whether an observed difference would likely persist on new data, or whether chance alone could explain it.

Why does this matter for publication? Because "our model beats the baseline by 0.8%" is the most common sentence in ML papers — and without a significance test, it is also the least trustworthy. Reviewers increasingly demand evidence that improvements are real, especially as benchmark gains shrink to fractions of a percent.

The core idea: paired comparison

The right way to compare two classifiers is paired: test both on the same folds (or same test examples) and look at the differences, fold by fold. Pairing removes a huge source of noise — some folds are just harder than others — because both models face the same difficulty. Comparing A's average against B's average from different splits is far weaker.

Worked example: the paired t-test by hand. Two models, same 5 folds, accuracy per fold:

Fold Model A Model B Difference (A−B)
1 0.85 0.83 +0.02
2 0.90 0.86 +0.04
3 0.80 0.79 +0.01
4 0.95 0.90 +0.05
5 0.85 0.84 +0.01

Mean difference d̄ = (0.02+0.04+0.01+0.05+0.01)/5 = 0.13/5 = 0.026. Standard deviation of differences: deviations from 0.026 are −0.006, +0.014, −0.016, +0.024, −0.016; squares: 0.000036, 0.000196, 0.000256, 0.000576, 0.000256; sum = 0.00132; ÷(5−1) = 0.00033; s = √0.00033 ≈ 0.0182.

t-statistic = d̄ / (s/√n) = 0.026 / (0.0182/√5) = 0.026 / 0.00814 ≈ 3.19. With 4 degrees of freedom, the two-sided critical value at the 5% level is 2.78. Since 3.19 > 2.78, the difference is statistically significant (p < 0.05): A is genuinely better than B, not just luckier.

Note what made this work: A beat B on every single fold (all differences positive). Consistency across folds is strong evidence. If the differences had been +0.06, −0.05, +0.04, −0.03, +0.01 (same-ish mean, wild swings), the test would say "not significant" — correctly, because the winner flips with the fold.

The p-value, honestly explained

A p-value is the probability of seeing a difference this large if the models were actually equal and only luck varied the folds. p < 0.05 means: if there were truly no difference, you'd see a gap this big less than 5% of the time. It is not the probability that your conclusion is wrong, and it is not a measure of how big the improvement is. A tiny, meaningless 0.1% gain can be "significant" with enough data; a huge 5% gain can be "non-significant" with 3 folds. Report the effect size (the actual gap, e.g., +2.6 points) alongside the p-value — significance tells you the gap is real, effect size tells you whether it matters.

McNemar's test: comparing on one test set

When you have a single test set (common in deep learning), the paired t-test across folds doesn't apply. McNemar's test compares two classifiers on the same test examples using only the cases where they disagree:

  • b = examples A got right and B got wrong.
  • c = examples A got wrong and B got right.
  • (Cases both got right or both got wrong carry no information about which is better — elegantly ignored.)

Test statistic = (|b − c| − 1)² / (b + c), compared against the chi-squared critical value 3.84 (5% level).

Worked example. On 200 test images: A right/B wrong on 18 (b=18); A wrong/B right on 7 (c=7). Statistic = (|18−7| − 1)² / 25 = 100/25 = 4.0 > 3.84 → significant: A is better. If instead b=14, c=9: (5−1)²/23 = 16/23 ≈ 0.70 < 3.84 → not significant. Same test set, same models — the second result correctly refuses to crown a winner on thin evidence.

Practical guidance and pitfalls

  • Correct for multiple comparisons. Comparing your model against 10 baselines at p < 0.05 gives ~40% chance of at least one fluke "win." Mention this; ideally adjust (Bonferroni: divide 0.05 by number of comparisons) or pre-register one primary comparison.
  • Non-normal differences? The t-test assumes roughly normal differences. With few folds or skewed metrics, use the Wilcoxon signed-rank test (paired, non-parametric) — available in every stats library.
  • Significance ≠ importance. A significant +0.2% on a saturated benchmark may not merit a paper; frame the contribution honestly.
  • Report confidence intervals, not just p-values: "A beats B by 2.6 points (95% CI: 0.3–4.9)" tells the full story in one line.

Confidence intervals: the honest error bar

A p-value says "the difference is real." A confidence interval (CI) says "the difference is real, and here's how big it plausibly is." Readers — and reviewers — increasingly prefer the CI because it carries both significance and size.

Worked example (paired differences, by hand). Reuse Chapter 8's fold differences: d̄ = 0.026, s = 0.0182, n = 5. The 95% CI for the true mean difference:

CI = d̄ ± t × (s/√n), with t = 2.78 (4 degrees of freedom).

Margin = 2.78 × (0.0182/√5) = 2.78 × 0.00814 ≈ 0.0226. CI = 0.026 ± 0.0226 → (0.003, 0.049).

Interpretation: "Model A beats B by 2.6 points (95% CI: 0.3 to 4.9 points)." The interval excludes zero — consistent with p < 0.05 — and tells you the win could be as small as 0.3 points (barely matters) or as large as 4.9 (substantial). One line, total honesty. When the CI includes zero, the difference isn't significant — same conclusion as the p-value, but you also see the range of plausible effects.

Effect size: how big is the win?

Statistical significance asks "is it real?"; effect size asks "does it matter?" A 0.2-point gain can be significant with 10,000 test examples and still be worthless. Report:

  • Absolute difference in the metric's own units ("+2.6 accuracy points") — always.
  • Relative improvement when baselines are strong ("cuts the error rate from 10% to 7.4% — a 26% relative error reduction"). Note the framing: on a 90%-accuracy baseline, +2.6 points sounds small, but eliminating a quarter of remaining errors sounds like what it is — substantial.

A common honest sentence: "The improvement is statistically significant (p = 0.02) but modest in absolute terms (+0.9 points), representing a 9% relative error reduction." Let the reader decide if it matters; your job is to make the size unmissable.

Comparing more than two models

With 3+ models, pairwise t-tests multiply the fluke risk (Chapter 8's multiple-comparison warning). The principled approach:

  1. Rank models per fold/dataset, then run a Friedman test (non-parametric ANOVA for ranks) asking "do any models differ at all?"
  2. If yes, follow with pairwise comparisons corrected for multiplicity (e.g., Nemenyi post-hoc test), often shown as a critical-difference diagram — models joined by a bar are not significantly different.

In practice, many papers simplify honestly: pre-register one primary comparison (your model vs the strongest baseline), test that pair properly, and present the rest of the table as descriptive ("other baselines shown for context, without significance claims"). This is legitimate as long as you don't crown winners among the untested pairs.

Significance in papers: what reviewers actually want

You don't need a statistics chapter in your paper — one careful sentence per comparison suffices. The template:

"Model A outperforms baseline B by 2.6 accuracy points (95% CI: 0.3–4.9; paired t-test across 5 folds, p = 0.03)."

That sentence contains the effect size, the uncertainty, the test, and the verdict. Reviewers who see it relax; reviewers who don't see it reach for the reject button's neighborhood. And if the test fails — report that too: "the 1.1-point gain over baseline C was not significant (p = 0.24)" reads as honesty, not weakness. Papers get rejected for hiding non-significance far more often than for reporting it.

For your research: Pick one primary comparison for your paper (your model vs the strongest baseline, on your primary metric) and test it: paired t-test across folds, or McNemar on a single test set. Report the gap, the p-value, and the confidence interval in one sentence. Pre-register this comparison (Chapter 12) so no one — including you — can suspect you tested ten baselines and reported the one that came out significant. Reviewers trust a paper that says "the improvement over baseline X is significant (p = 0.03)" far more than one that just bolds the biggest number in a table.

Key takeaways - A higher score is not a better model until you rule out luck; that's what significance testing does. - Compare paired: same folds, same test examples — look at per-fold differences. - Paired t-test across folds, or McNemar's test on a single test set using only disagreements. - p < 0.05 means the gap is unlikely to be luck; always report the effect size too. - Correct for multiple comparisons, and pre-register your primary comparison.


Chapter 9: Evaluation for Imbalanced Data and Rare Events

Why this chapter exists

Chapter 2's 99%-accuracy trap was just the opening. Fraud detection, rare-disease screening, network intrusion, manufacturing defects, equipment failure prediction — the most valuable ML applications are exactly the ones where the positive class is rare. Standard evaluation machinery breaks down here in specific, predictable ways, and this chapter gives you the replacement toolkit.

The core problem, restated: when negatives outnumber positives 100-to-1, metrics that average over all examples are dominated by the easy negatives. The model can be perfect on negatives, useless on positives, and still score beautifully. You need metrics that force attention onto the rare class.

What to use instead of accuracy

1. Precision, recall, and F1 on the minority class. Always report these for the rare class explicitly. A fraud model with 99.9% accuracy but recall 0.30 catches less than a third of fraud — the F1 tells the truth.

2. Balanced accuracy = (recall on positives + recall on negatives) / 2. Each class contributes equally regardless of size. The always-predict-majority "model" scores (0 + 1)/2 = 0.5 — correctly exposed as worthless, where plain accuracy gave it 99%.

3. Matthews Correlation Coefficient (MCC) = (TP×TN − FP×FN) / √((TP+FP)(TP+FN)(TN+FP)(TN+FN)). MCC is the correlation between predicted and true labels: +1 perfect, 0 random, −1 totally wrong. It uses all four cells of the confusion matrix, so it stays honest under any imbalance. Many researchers now consider it the best single-number summary for binary imbalanced problems.

4. PR-AUC (area under the precision-recall curve). Chapter 5 showed ROC-AUC hiding false-alarm floods; the PR curve plots precision vs recall across thresholds, and precision has no such hiding place — every false alarm directly lowers it.

Worked example. Fraud detection: 1,000 transactions, 20 fraud (2%), 980 legitimate. Model flags 30 transactions as fraud: 15 are real fraud, 15 are false alarms.

  • TP=15, FP=15, FN=5, TN=965.
  • Accuracy = (15+965)/1000 = 0.98 — looks great, means nothing.
  • Precision = 15/30 = 0.50 — half of all fraud alerts are false.
  • Recall = 15/20 = 0.75 — catches 3 in 4 frauds.
  • F1 = 2×(0.5×0.75)/(0.5+0.75) = 0.75/1.25 = 0.60.
  • Balanced accuracy = (0.75 + 965/980)/2 = (0.75 + 0.9847)/2 ≈ 0.867.
  • MCC = (15×965 − 15×5)/√(30×20×980×970) = (14475−75)/√(570,360,000) = 14400/23881 ≈ 0.603.

Read the story: accuracy 0.98 is a lie; balanced accuracy 0.87 is inflated by the easy negatives; F1 0.60 and MCC 0.60 agree the model is mediocre — decent recall, poor precision. If each false alarm costs an investigator an hour, precision 0.50 is the number the business actually feels.

Threshold tuning is part of the method

On imbalanced data, the default 0.5 threshold is almost never right. Because positives are rare, you typically lower the threshold to catch more of them, accepting more false alarms — then report the full precision-recall curve so readers see the trade-off you chose. Tune the threshold on the validation set (never the test set), optimizing the metric you actually care about (often F1 or a cost-weighted score). Report the chosen threshold in the paper. "We operate at threshold 0.22, maximizing F1 on validation" is a complete, honest sentence.

Sampling fixes change the evaluation, not just the training

Common imbalance remedies — oversampling the minority class, undersampling the majority, SMOTE, class weights — change the training distribution. Critical rule: evaluate on the natural distribution. If you oversample fraud to 50% in training but deploy into a world with 0.1% fraud, your test set must reflect the 0.1% world, or your precision estimate is fantasy. Resample the training folds; keep validation and test folds at natural prevalence. (And with cross-validation: resample inside each fold's training portion, never before splitting — resampling first leaks information across folds.)

Report prevalence, always

Every imbalanced-data paper must state the class ratio plainly: "fraud prevalence 0.18% in the test set." Without it, precision and F1 are uninterpretable — precision 0.5 means something completely different at 10% prevalence vs 0.1%. Reviewers in applied fields check this first.

Walking the PR curve (mini worked example)

Take the fraud example's spirit but smaller, so you can compute by hand. Ten transactions, 3 fraud (F), 7 legitimate (L), model scores:

F 0.95, L 0.85, F 0.70, L 0.60, F 0.55, L 0.40, L 0.30, L 0.20, L 0.10, L 0.05.

Compute (recall, precision) at thresholds:

  • Thr 0.9 (flag top 1): R = 1/3 ≈ 0.33, P = 1/1 = 1.0
  • Thr 0.8 (flag top 2): R ≈ 0.33, P = 1/2 = 0.5
  • Thr 0.65 (flag top 3): R = 2/3 ≈ 0.67, P = 2/3 ≈ 0.67
  • Thr 0.5 (flag top 5): R = 1.0, P = 3/5 = 0.6
  • Thr 0.0 (flag all): R = 1.0, P = 3/10 = 0.3

The PR curve starts at (0.33, 1.0), dips to (0.33, 0.5) — that legitimate transaction scoring 0.85 costs half your precision instantly — recovers, then slides to (1.0, 0.3). That dip is the reality of rare-event detection: one confident false alarm devastates precision because true positives are so few. PR-AUC (≈0.68 here by trapezoid) summarizes the whole trade-off in the currency you actually care about.

Cost-sensitive evaluation: putting a price on errors

When you can estimate real costs, skip abstract metrics and compute expected cost directly (as in Chapter 4's clinic example).

Worked example. Fraud team: each missed fraud costs $2,000 on average; each false alarm costs $25 of investigator time. Model X: 20 frauds in test, catches 15 (FN=5), raises 60 false alarms (FP=60). Model Y (higher threshold): catches 12 (FN=8), only 15 false alarms (FP=15).

  • Cost(X) = 5 × $2,000 + 60 × $25 = $10,000 + $1,500 = $11,500
  • Cost(Y) = 8 × $2,000 + 15 × $25 = $16,000 + $375 = $16,375

Model X wins on cost despite far worse precision (15/75 = 0.20 vs 12/27 ≈ 0.44), because missed fraud dominates the budget. If the business instead capped investigator hours, the ranking could flip — which is precisely why the cost model must come from the stakeholder, not the researcher. Report the cost assumptions; they're the most-argued-over and most-useful line in an applied paper.

A note on anomaly detection

In pure anomaly detection (network intrusion, equipment failure), "positives" may be unlabeled — you only know normal behavior. Standard precision/recall need labels, so evaluation adapts: inject synthetic anomalies with known labels, use time-to-detection on historical incidents, or report precision@k ("of the top-100 flagged events, how many were real incidents per expert review?"). The principle never changes: measure what the deployment actually needs, on data shaped like deployment.

Report both threshold-dependent and threshold-independent metrics

A complete imbalanced-data results table has two layers:

  1. Threshold-independent (ranking quality): PR-AUC, ROC-AUC. Answers "does the model rank positives above negatives?"
  2. Threshold-dependent (operating-point quality): precision, recall, F1 at the chosen, validation-tuned threshold. Answers "how does it perform where we'll actually deploy it?"

Papers that report only layer 1 leave readers guessing about real performance; papers that report only layer 2 hide whether a better threshold was available. Together they tell the full story — and the PR curve figure connects them visually, with your operating point marked.

Calibration: are the probabilities honest?

Imbalanced models often output scores, not just rankings — "87% probability of fraud." Calibration asks whether 87% means 87%: among all cases scored ~0.87, do ~87% turn out positive? Check with a reliability diagram (bin predictions by score, plot mean score vs actual positive rate per bin) or the Brier score (mean squared error of probabilities). Miscalibrated models are dangerous when scores drive decisions (auto-blocking transactions above 0.9) — a model can rank perfectly (AUC 1.0) while its probabilities are nonsense. If your deployment acts on probability values rather than just rankings, report calibration. It's a one-paragraph addition that applied reviewers notice and appreciate. As a quick first check, compare the model's average predicted probability on the test set against the actual positive rate — a large gap means the scores are systematically over- or under-confident, and any threshold you set from them will misbehave in production.

For your research: If your positive class is under ~10%, build your results around PR curves, F1/MCC, and per-class precision/recall — not accuracy or ROC-AUC alone. State prevalence, describe how you handled imbalance in training, confirm the test set kept natural prevalence, and report your tuned threshold. This checklist is what separates a publishable rare-event paper from one that gets rejected with "the evaluation does not address class imbalance."

Key takeaways - On rare events, accuracy lies; use precision/recall/F1 on the minority class, balanced accuracy, MCC, and PR-AUC. - MCC (±1 scale) is the most trustworthy single-number summary for binary imbalance. - Tune the decision threshold on validation data for the metric you care about; report the threshold. - Fix imbalance in training only — test on the natural class distribution, resampling inside CV folds. - Always report class prevalence; without it your metrics cannot be interpreted.


Chapter 10: Beyond Metrics — Error Analysis and Slicing Results

Numbers don't explain themselves

Suppose your model reaches 90% accuracy and beats the baseline significantly. Are you done? No — you still don't know why it fails on the remaining 10%, and that 10% contains your next paper. Error analysis is the systematic study of your model's mistakes: collecting the errors, categorizing them, and finding patterns. Every strong results section does this; weak ones stop at the metric table.

There is a second reason: aggregate metrics can hide that a model works well on average but fails catastrophically for one subgroup. A skin-lesion classifier at 92% overall accuracy that scores 60% on dark skin is not a 92% model — it is a biased model. Slicing results by meaningful subgroups (demographics, image conditions, classes, time periods) is how you find out.

How to do error analysis: a protocol

  1. Collect the errors. Pull every test example the model got wrong (false positives and false negatives separately — they usually have different causes).
  2. Sample and inspect manually. Look at 50–100 errors yourself. This is tedious and irreplaceable; no metric substitutes for looking.
  3. Build an error taxonomy. Group errors into categories with names: "blurry image," "occluded object," "label looks wrong," "rare class confused with similar class." Aim for 4–8 categories covering most errors.
  4. Count and prioritize. What fraction of errors falls in each bucket? Fix the biggest bucket first — that is where the next accuracy points live.
  5. Check the labels. In every error analysis, some "errors" turn out to be mislabeled test data. If 10% of your errors are actually correct predictions with wrong labels, say so — it bounds how much better any model could do, and reviewers respect the honesty.

Worked example. A crop-disease classifier: 200 test images, 24 errors (88% accuracy). Manual review of all 24:

  • 9 errors (37.5%): images taken in low light — disease spots invisible even to the agronomist.
  • 6 errors (25%): early-stage disease, visually ambiguous.
  • 5 errors (20.8%): two similar diseases confused (blight vs. rust).
  • 3 errors (12.5%): mislabeled in the dataset (model was actually right).
  • 1 error (4.2%): corrupted image file.

The research plan writes itself: the biggest bucket (low light) suggests adding low-light augmentation or a "retake photo" prompt in the app; the blight/rust confusion suggests a focused two-class follow-up study; the 3 mislabels mean true accuracy is arguably 91.5%, not 88%. None of this came from the 88% — it came from looking.

Slicing: the subgroup audit

Choose slice dimensions before you look at results (to avoid cherry-picking), based on domain knowledge: patient age group, image lighting, geographic region, class rarity, text length, time of year. Compute your primary metric per slice. Look for two things:

  • Weak slices: performance far below average → targeted data collection or modeling.
  • Disparate impact: systematically worse for a protected or important group → a fairness problem that must be disclosed and addressed, not averaged away.

Worked example. A loan-approval model's overall accuracy is 91%. Sliced by applicant region:

Region n Accuracy Approval recall
North 400 0.94 0.93
South 350 0.92 0.90
East 150 0.89 0.85
West (rural) 100 0.74 0.61

Overall 91% hides that the model fails rural Western applicants — recall 0.61 means nearly 4 in 10 creditworthy applicants there are wrongly rejected. Possible causes: fewer training examples from that region, or features that don't transfer (urban income patterns vs rural). The fix is more data from the weak slice or region-specific calibration — and ethically, this model should not deploy until the gap is addressed. A paper that reports only the 91% is misleading; a paper that reports the slice table and discusses it is doing science.

Ablation: which parts earn their keep?

If your method has multiple components (a new feature set + a new architecture + a new loss function), an ablation study removes them one at a time and re-measures. Result: "removing the texture features costs 4.1 F1 points; removing the new loss costs 0.3." This tells readers — and you — which ideas actually matter. Components that cost nothing when removed should be cut from the paper's claims (or from the model). Reviewers routinely ask "is the gain from X or from Y?"; an ablation table answers before they ask.

Presenting it

Dedicate a real subsection to error analysis ("Error Analysis" or "Qualitative Results"), with: the taxonomy table with counts, 2–4 example error images/cases in a figure, the slice table, and one paragraph of interpretation linking patterns to concrete next steps. This section is often what examiners and reviewers quote — it proves you understand your model rather than just operating it.

Error analysis for regression

Error analysis isn't only for classifiers. For regression, slice residuals by predicted value, by subgroup, or by feature ranges:

Worked example. A house-price model: overall MAE = $12,000. Sliced by price band:

Band n MAE
Under $100k 120 $6,000
$100k–$300k 200 $11,000
Over $300k 60 $28,000

The model is twice as bad (relatively and absolutely) on expensive houses — perhaps because they're rarer in training, or luxury features aren't captured. The fix differs by cause: more luxury examples, or new features (lot size, neighborhood). Also plot residuals vs each input feature: a trend (e.g., residuals growing with house age) means the model mishandles that feature — a modeling insight no aggregate metric provides.

Turning patterns into experiments: the loop

Error analysis earns its keep when it drives the next experiment. The loop:

  1. Pattern: "37% of errors are low-light images" (Chapter 10's crop example).
  2. Hypothesis: "The model never learned low-light features because training images are all daylight."
  3. Intervention: add brightness augmentation + collect 200 low-light training images.
  4. Measurement: re-run the same evaluation; check the low-light slice specifically, not just overall accuracy.

If the slice improves and nothing else regresses, the analysis paid off — and the paper gains a paragraph showing scientific iteration rather than lucky tuning. Document the loop; reviewers love "error analysis revealed X, so we did Y, which improved slice Z by N points."

The qualitative figure

Reserve one figure for qualitative results: 2–4 example predictions with brief captions — one success, two characteristic failures, one interesting edge case. For vision: the images with predicted/true labels. For text: the sentences with model output. This figure does three jobs: it proves you looked at real outputs, it makes failure modes tangible for readers, and it often becomes the most-remembered part of your paper. Caption each example with why it matters ("typical low-light failure from error category 1"), not just what it shows.

Starter taxonomies by domain

Don't start your error taxonomy from a blank page — adapt these:

  • Medical imaging: image quality (blur, artifacts) / anatomical variant / early-stage subtlety / label disagreement between annotators / rare condition.
  • NLP/text: negation handling / sarcasm / domain slang / long-range context / annotation guideline edge cases.
  • Tabular/business: missing-value patterns / outlier inputs / distribution shift (new region, new season) / label noise from heuristic labeling.
  • Vision (general): occlusion / lighting / small objects / unusual viewpoints / background confusion.

Expect 1–2 categories to dominate; that's normal and useful — it focuses your next experiment. And always include a "label looks wrong" bucket: in most real datasets, 2–10% of "errors" are annotation mistakes, and quantifying them sets a ceiling on achievable performance that reviewers find credible.

Knowing when to stop

Error analysis has diminishing returns. Stop when new errors stop forming new categories — typically after 50–150 inspected cases — and when the remaining errors look like irreducible noise (genuinely ambiguous cases even experts disagree on). Document that stopping point: "manual review of 120 errors; categories saturated after ~80." It shows method, not exhaustion, and it keeps the analysis chapter proportionate to the rest of the thesis.

For your research: Budget two full days for error analysis before writing your results chapter — it is the highest-insight-per-hour work in the project. Pre-commit to your slice dimensions in your evaluation plan (Chapter 12) so the audit is principled, not a hunt for flattering subgroups. And when you find an embarrassing weakness (you will), report it: a discussed weakness is a future-work paragraph and a sign of maturity; a hidden one is a rejection reason when the reviewer finds it.

Key takeaways - Aggregate metrics hide error patterns and subgroup failures; error analysis and slicing reveal them. - Protocol: collect errors, inspect manually, build a taxonomy, count, check labels, prioritize the biggest bucket. - Slice results by domain-meaningful subgroups chosen in advance; watch for disparate impact. - Ablation studies show which components of your method actually earn their performance. - Give error analysis its own subsection with tables, examples, and interpretation — reviewers read it closely.


Chapter 11: How Research Papers Report Evaluation (Reading Results Sections Critically)

Why you must read like a reviewer

You will read dozens of papers before writing your own. Most students read results sections passively — "they got 94%, good." Researchers read adversarially: should I believe this number? This chapter is a checklist for that reading. It also secretly teaches you how to write your own results section, because every trap you learn to spot is a trap you will avoid setting.

The critical reading checklist

1. What exactly was measured, and on what? Find the dataset, the split, and the metric definition. Red flags: metric named but not defined ("we report accuracy" — on which class? which split?), test set size missing, or "we split randomly" with no seed or ratio.

2. Is there a real baseline? A result without a baseline is a number without meaning. Strong papers compare against the best known method, tuned fairly. Red flags: comparing only against weak or outdated baselines, against the authors' own reimplementation with no tuning details, or against nothing at all ("our model achieves 92%").

3. Were baselines tuned fairly? The classic trick: spend weeks tuning your own model, run baselines with default settings, then celebrate the gap. Honest papers tune baselines with comparable effort (same validation data, similar search budget) and say so. If tuning effort isn't described, assume it wasn't equal.

4. Is the test set truly untouched? Look for signs of peeking: model selected by test performance ("we chose the architecture with the best test accuracy"), threshold tuned on test, or preprocessing (like feature selection or scaling) fit on the full dataset before splitting. Any of these is leakage (Chapter 1, Chapter 7).

5. Are the gains significant? A table where the proposed method beats five baselines by 0.3–0.8% with no significance test and no error bars is noise presented as victory. Look for ±, confidence intervals, or p-values (Chapter 8). Their absence doesn't prove the result is false — but it means the authors haven't shown it's true.

6. Do the metrics match the problem? Accuracy on a 99-to-1 imbalanced dataset (Chapter 9), RMSE reported without units or context (Chapter 6), AUC without a PR curve on rare events (Chapter 5) — metric choice reveals whether the authors understand their own problem.

7. Can you reproduce it? Seeds, fold counts, hardware-independent details, code or precise hyperparameters. "We used a neural network" is not reproducible. The trend in top venues is mandatory code release; treat missing code as a yellow flag, not a verdict.

Worked example: dissecting a suspicious table

You read: "Our method achieves 96.2% accuracy, outperforming Baseline A (91.0%) and Baseline B (93.5%) on the XYZ dataset."

Apply the checklist:

  • Metric/problem fit: What is XYZ's class balance? You check the dataset paper: 95% of samples are the majority class. The majority baseline is 95% — so "96.2%" is 1.2 points above doing nothing, and Baseline A's 91% is below the do-nothing baseline, suggesting A was broken or mistuned. The headline number collapses on inspection.
  • Baselines: Are A and B the current state of the art, or five-year-old methods? You check citations: both from 2019; two 2024 methods exist and aren't compared. Selective baselining.
  • Significance: No ±, no test. With what test size? Not stated. A 1.2-point edge on 200 test samples is ~2 examples — pure noise territory.
  • Reproduction: No seed, no code link, "hyperparameters chosen empirically."

Verdict: this result is not evidence. A fair version would report balanced metrics on the minority class, compare against 2024 methods tuned on the same validation split, and show significance. Notice you reached this verdict without running a single experiment — critical reading is cheap, and it protects you from building on sand.

Reading figures critically

  • ROC curves without the diagonal random line, or with a zoomed axis that exaggerates tiny gaps.
  • Bar charts with truncated y-axes (bars starting at 80% make a 2-point gap look enormous).
  • Loss curves showing training loss only (hides overfitting) — demand validation curves.
  • Confusion matrices with unlabeled axes (Chapter 4).

Turning criticism into writing guidance

Every item above inverts into a rule for your own paper: define metrics precisely, tune baselines fairly and say so, lock the test set, show uncertainty, match metrics to the problem, release code and seeds. Write the results section you'd want to read as a skeptic. The easiest way: hand your draft to a colleague with this checklist and ask them to attack it.

Red-flag phrases and what they really mean

Learn to translate common results-section phrasing:

  • "Achieves 99% accuracy" (no class balance stated) → probably the majority-class trap. Check prevalence.
  • "Significantly outperforms" (no test named, no p-value) → "significantly" is doing rhetorical work, not statistical work.
  • "Hyperparameters were tuned empirically" → tuned on what data? If the test set, that's leakage.
  • "We compare against state-of-the-art methods" (baselines 5 years old) → selective baselining; check citations.
  • "Results are averaged over multiple runs" (no std reported) → the ± was probably unflattering.
  • "Our method is robust" (one dataset, one split) → robustness untested.
  • "Slightly better" with bold numbers in a table → noise in boldface.

None of these phrases prove misconduct — sometimes they're just sloppy writing. But each one tells you exactly which question to ask next, which is the skill.

Anatomy of a trustworthy results section

Contrast the suspicious table from earlier with this:

"On the XYZ test set (n = 2,000; 6% positive), our model reaches F1 = 0.72 ± 0.03 (5-fold stratified CV, seed 11), vs. 0.65 ± 0.04 for the strongest baseline [14], tuned with equal budget on the same validation folds. The paired difference (+7.0 points, 95% CI 3.1–10.9) is significant (paired t-test, p = 0.008). PR curves (Fig. 3) show the gain holds across thresholds; error analysis (§4.3) attributes it to... Code and seeds: [link]."

Every clause answers a checklist item: prevalence, metric, uncertainty, split scheme, fair baseline tuning, significance with effect size and CI, threshold-robustness, error analysis pointer, reproducibility. Write yours to the same template and reviewers run out of objections.

Questions to ask at a talk or defense

When someone presents results, these five questions cut to the truth fastest:

  1. "What was the majority-class baseline?"
  2. "Was the test set used for any model selection decision?"
  3. "How were the baselines tuned, and with what budget?"
  4. "What's the confidence interval on that improvement?"
  5. "Where does it fail — did you slice the results?"

Ask them of others' work to learn; expect them on your own work to survive.

The rebuttal: defending your evaluation

Sooner or later a reviewer writes: "The evaluation is insufficient." Here's how to respond to the common variants:

  • "Add more baselines" → add the strongest missing one, tuned fairly; if it beats you on some metric, say so and explain the trade-off. Never silently omit a baseline that wins.
  • "Results lack statistical significance" → run the test (Chapter 8). If significant, add the sentence. If not, report it honestly and discuss effect size — sometimes the rebuttal is "the effect is small but consistent across 4 datasets."
  • "Only one dataset" → add a second dataset if at all possible; if truly impossible (rare data), say why and strengthen everything else (more folds, ablations, error analysis).
  • "Metric X would be more appropriate" → compute metric X and add it. If it disagrees with your headline metric, discuss why — metric disagreement is itself an interesting finding.
  • "Potential data leakage" → this is the serious one. Audit your pipeline, document the split procedure precisely, and show the suspicious step is clean. If you find real leakage, fix it and report corrected numbers no matter what they do to your story.

The meta-principle: treat every evaluation critique as a gift. Each one you address makes the paper harder to reject and easier to build on — which is, after all, the point of publishing. Keep the checklist from this chapter taped next to your desk while writing; it's cheaper than a second round of reviews.

For your research: Start a "results autopsy" document: for every paper you cite, fill in one row — dataset/split, metric, baselines, tuning fairness, significance reported?, leakage signs, your verdict. After 15 papers you'll spot weak evaluation in seconds, and your related-work section will practically write itself ("prior work reports accuracy on imbalanced data without significance testing; we address this by..."). That sentence is a contribution claim reviewers understand instantly.

Key takeaways - Read results adversarially: dataset/split, metric definition, baselines, tuning fairness, leakage, significance, reproducibility. - A number without a baseline, without error bars, and without a defined split is not evidence. - Watch for unfair baseline tuning, test-set peeking, metric–problem mismatch, and misleading figures. - Every checklist item inverts into a rule for writing your own results section. - Keep a "results autopsy" table of papers you read — it trains your eye and feeds your related work.


Chapter 12: Designing Your Evaluation Plan (Aligns with Research Methodology)

Evaluation is methodology, not an afterthought

In your thesis proposal or paper's methodology chapter, evaluation deserves the same care as model design. Examiners read methodology asking one question: if I ran this plan, would I trust the conclusion? This chapter turns everything in the book into a concrete plan template — and introduces pre-registration: writing the plan down before seeing results, which protects you from fooling yourself.

The pre-registration principle

Human brains are excellent at finding patterns — including patterns that aren't there. If you try 20 model variants and report the best test score, you've selected for luck (Chapter 7). If you check accuracy, then F1, then AUC, and report whichever looks best, you've done the same with metrics. Pre-registration means fixing in advance: the primary metric, the data splits, the baselines, the statistical test, and what counts as success. Then you run the plan and report the outcome — good or bad.

This doesn't forbid exploration. It separates confirmatory analysis (the pre-registered plan; this is your evidence) from exploratory analysis (everything else; labeled honestly as exploration). Both are valuable; only the first proves the claim. Many theses now include the evaluation plan in the proposal defense — your examiners approve the method, and then the results are what they are.

The evaluation plan template

Fill this in before training. One page. Get your supervisor's sign-off.

1. Research question (one sentence). Example: "Does adding soil-moisture sensor features improve wilt-disease detection over image-only features?"

2. Dataset and splits. Name the dataset, its size, and class distribution. Specify the split scheme: "stratified 5-fold CV (seed 42)" or "70/15/15 stratified split (seed 42)". State that the test set/folds are locked before any tuning.

3. Primary metric (one, with justification). Example: "F1 on the disease class — primary because missing a diseased plant (recall) and false alarms that waste pesticide (precision) are both costly." Name 1–2 secondary metrics.

4. Baselines (at least two). A trivial baseline (majority class / mean predictor) and the strongest known method, tuned with comparable effort on the same validation data.

5. Model selection protocol. How hyperparameters are chosen: "inner 3-fold CV within each outer fold (nested CV)"; or "validation split; test touched once at the end." State the decision threshold policy.

6. Comparison statistics. "Primary comparison (our model vs strongest baseline on F1) tested with a paired t-test across folds at α = 0.05; 95% confidence interval reported."

7. Slices and error analysis. Pre-commit slice dimensions: "metrics reported per disease stage (early/late) and per lighting condition (field/lab)." Error analysis: "100 errors manually categorized."

8. Success criterion. Example: "Success = statistically significant F1 improvement ≥ 3 points over the strongest baseline, with no slice dropping below baseline." Writing this down stops you from redefining victory after seeing the numbers.

9. Reproducibility. Seeds, library versions, hardware notes, code repository link (to be filled at submission).

Worked example: a filled plan

Project: SMS spam detection for Urdu-language messages (a realistic MS thesis topic).

  1. Question: "Does a character-n-gram model beat a word-based baseline for Urdu SMS spam detection?"
  2. Data: 8,000 labeled messages (12% spam). Stratified 5-fold CV, seed 7. Test folds locked.
  3. Primary metric: F1 on the spam class (spam is the minority; false alarms annoy users, misses let fraud through — both matter). Secondary: PR-AUC.
  4. Baselines: (a) majority-class (predict "not spam"); (b) TF-IDF word unigrams + logistic regression, tuned via inner 3-fold CV with the same search budget as our model.
  5. Selection: nested CV; classification threshold tuned on inner validation to maximize F1.
  6. Statistics: paired t-test on per-fold F1, α = 0.05; report mean ± std and 95% CI of the difference.
  7. Slices: F1 per message length (short/long) and per spam type (advertising/fraud). Error analysis on 100 misclassified messages.
  8. Success: significant F1 gain ≥ 2 points over baseline (b), no slice worse than baseline by > 1 point.
  9. Reproducibility: seed 7, scikit-learn [6], code on GitHub at submission.

This plan is a contract. When the results come in — whatever they are — the thesis writes itself, because every methodological decision was already defended.

When results disappoint (they often do)

If the pre-registered test fails, that is a finding, not a failure: "under rigorous evaluation, the proposed features did not significantly improve over the baseline (ΔF1 = +0.8, p = 0.31)." Report it, then explore honestly: the error analysis (Chapter 10) usually reveals why — wrong features, noisy labels, or a baseline stronger than expected. Some of the best theses are careful negative results with a sharp diagnosis. What destroys credibility is not a negative result but a plan quietly rewritten to manufacture a positive one.

Final checklist before you run

  • [ ] Splits defined, seeded, and locked; test data untouched by any decision
  • [ ] Primary metric chosen and justified from error costs
  • [ ] Baselines named, including a trivial one and the strongest known method
  • [ ] Tuning protocol separated from evaluation (nested CV or locked test set)
  • [ ] Statistical test and success criterion written down
  • [ ] Slice dimensions and error-analysis plan fixed
  • [ ] Supervisor has seen and approved the one-page plan

For your research: Put this one-page plan in your thesis proposal and again (condensed) in your paper's methodology section. When examiners ask "why this metric?" or "how do you know the improvement is real?", you point at the plan — approved before results existed. That is the difference between a student who ran experiments and a researcher who tested a hypothesis. Keep the approved plan unedited (version it); any change after seeing results gets documented as a deviation, with reasons.

Key takeaways - Evaluation design belongs in methodology; write the plan before training, not after. - Pre-register: one primary metric, fixed splits, named baselines, a statistical test, and a success criterion. - Separate confirmatory analysis (your evidence) from exploratory analysis (labeled as such). - A rigorous negative result with a sharp diagnosis beats a manufactured positive one. - Get supervisor sign-off on the one-page plan; version it and document any deviations.


References

[1] T. Hastie, R. Tibshirani, and J. Friedman, The Elements of Statistical Learning, 2nd ed. New York, NY, USA: Springer, 2009. [2] C. M. Bishop, Pattern Recognition and Machine Learning. New York, NY, USA: Springer, 2006. [3] T. M. Mitchell, Machine Learning. New York, NY, USA: McGraw-Hill, 1997. [4] K. P. Murphy, Probabilistic Machine Learning: An Introduction. Cambridge, MA, USA: MIT Press, 2022. [5] A. Géron, Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow, 2nd ed. Sebastopol, CA, USA: O'Reilly Media, 2019. [6] F. Pedregosa et al., "Scikit-learn: Machine learning in Python," J. Mach. Learn. Res., vol. 12, pp. 2825–2830, 2011. [7] S. Russell and P. Norvig, Artificial Intelligence: A Modern Approach, 4th ed. Hoboken, NJ, USA: Pearson, 2020.


Glossary

  • Accuracy — Fraction of all predictions that are correct; misleading on imbalanced data.
  • Precision — Of predicted positives, the fraction truly positive; punishes false alarms.
  • Recall — Of actual positives, the fraction found; also called sensitivity or true positive rate.
  • F1 score — Harmonic mean of precision and recall; punishes models good at only one.
  • True positive (TP) — Positive case correctly predicted positive.
  • False positive (FP) — Negative case wrongly predicted positive; a false alarm.
  • True negative (TN) — Negative case correctly predicted negative.
  • False negative (FN) — Positive case wrongly predicted negative; a miss.
  • Confusion matrix — Table of TP, FP, TN, FN (or N×N for multi-class) showing error structure.
  • ROC curve — Plot of recall vs false-positive rate across all decision thresholds.
  • AUC — Area under the ROC curve; probability a random positive outranks a random negative.
  • Precision-recall curve — Precision vs recall across thresholds; preferred under heavy imbalance.
  • MAE — Mean absolute error; average prediction error in original units.
  • RMSE — Root mean squared error; penalizes large errors more than MAE.
  • R-squared — Fraction of target variance explained by the model vs predicting the mean.
  • Train/validation/test split — Data partitioned for fitting, tuning, and final unbiased measurement.
  • Cross-validation — Rotating test folds so every example is tested once; averaged for stability.
  • Stratified sampling — Keeping class ratios constant across folds or splits.
  • Overfitting — High training performance but poor performance on new data; memorization.
  • Data leakage — Test information contaminating training, invalidating the evaluation.
  • Baseline — Simple reference method (e.g., majority class) that a new model must beat.
  • Statistical significance — Evidence a measured difference is unlikely due to chance, typically judged at p < 0.05 (the p-value).
  • Matthews correlation coefficient (MCC) — Correlation between truth and predictions; robust single metric for imbalance.
  • Error analysis — Manual categorization of model mistakes to find patterns and next steps.
  • Pre-registration — Fixing metrics, splits, baselines, and success criteria before running experiments.

Practice Exercises

  1. A classifier tested on 50 emails produces TP=18, FP=4, FN=7, TN=21. Compute accuracy, precision, recall, and F1 by hand. Show each step.
  2. A model predicts "no churn" for all 200 customers; 170 truly did not churn. Compute its accuracy and its recall on the churn class. What does the gap tell you?
  3. From the confusion matrix in Chapter 4's clinic example (TP=42, FP=15, FN=8, TN=135), compute specificity, false positive rate, and negative predictive value.
  4. A 3-class model yields this matrix (rows = true, columns = predicted): [20, 5, 0; 3, 22, 0; 1, 2, 22]. Compute overall accuracy and per-class recall. Which class is weakest, and what would you investigate?
  5. Predictions vs truth: (10, 12), (20, 18), (30, 35), (40, 38), (50, 47). Compute MAE and RMSE by hand. Which is larger, and why?
  6. Using the data in Exercise 5, compute R-squared. (Hint: first find the mean of the true values, then the total sum of squares.)
  7. Six loan applicants have true labels [default, repay, default, repay, repay, default] and model scores [0.9, 0.75, 0.6, 0.5, 0.3, 0.15]. Compute the (FPR, TPR) points for thresholds 1.0, 0.7, 0.4, and 0.0, and sketch the ROC curve. Estimate the AUC.
  8. A 5-fold CV gives F1 scores of 0.72, 0.68, 0.75, 0.70, 0.65. Compute the mean and standard deviation. A baseline scores 0.71 ± 0.04. Is your model clearly better? Why or why not?
  9. Two models tested on the same 300 examples: model A right / B wrong on 22, A wrong / B right on 10. Run McNemar's test at the 5% level. Is A significantly better?
  10. Design a one-page evaluation plan (Chapter 12 template) for this project: "Detecting fake product reviews; 5,000 reviews, 8% fake." Choose and justify a primary metric, name two baselines, specify the split scheme, the statistical test, two slice dimensions, and a success criterion.

End of Book 5. Next: Book 6 — Overfitting and How to Avoid It.