Common AI Mistakes Beginners Make

Book 10 of 10 — AstolixGen Learning Series (Detailed Edition) · For researcher and publication students


Cover

About This Book

Most AI research projects do not fail because the idea was bad. They fail — quietly, and often after weeks of work — because of small, avoidable mistakes: a test set that was accidentally contaminated, a metric that hid the real performance, a baseline that was never run, or a result that cannot be reproduced by anyone, including the original author. This book collects the most common of these mistakes, explains why each one matters, and shows you how to fix it with a concrete worked example in every chapter. By the end, you will have a practical checklist you can apply before every experiment and every submission.

Learning objectives: - Explain why small experimental errors can invalidate an entire research study - Define and detect data leakage in machine learning experiments - Describe test set contamination and how to prevent it - Choose evaluation metrics that match the research question - Justify the use of strong baselines in every experiment - Recognize overclaiming and "state of the art" abuse in papers and drafts - Identify p-hacking and questionable research practices - Make experiments reproducible through seeds, code, and documentation - Read and cite the literature accurately, without strawman arguments - Design experiments that use compute wisely and ethically - Handle privacy, consent, and dual-use concerns in AI research - Apply a pre-submission checklist and survive peer review constructively

Learning Dashboard

Concept Definition (one line) Example Use in research
Experimental validity Whether results actually measure what they claim to measure A 98% accuracy from a leaked test set is invalid Check validity before celebrating any result
Data leakage Test information accidentally used during training Scaling features before splitting the data Split data first, preprocess only on train folds
Test set contamination The model sees test data during development Tuning hyperparameters until test score peaks Lock the test set away; use a validation set
Accuracy theater High accuracy that hides poor real performance 99% accuracy on a 99%-negative dataset Report precision, recall, F1 for the minority class
Baseline The simple method your new method must beat Majority-class classifier scoring 60% Beat a strong baseline, not a strawman
Overclaiming Stating more than the experiments support "State of the art" on one small dataset Scope claims to exactly what was tested
p-hacking Trying many analyses until one is "significant" Testing 20 metrics and reporting only one Pre-register hypotheses and report all results
Reproducibility Another researcher can obtain your results Fixed seeds and shared code Document everything; archive code and data
Strawman citation Misrepresenting prior work to look better "Prior methods fail" without testing them Read primary sources; compare fairly
Experiment budget Planned compute limits per research question 500 GPU hours on one hyperparameter Design small pilots first, then scale
Informed consent Permission from people whose data you use Scraping photos without consent Obtain consent or use licensed datasets
Dual use Research that can help and harm A detector usable for surveillance Discuss risks honestly in the paper

Roadmap: The chapters form a chain. Chapter 1 explains why small errors matter so much — that is the motivation for everything else. Chapters 2 and 3 cover data handling mistakes (leakage and contamination), the most common reason results are wrong. Chapters 4 and 5 cover evaluation mistakes (metric misuse and missing baselines), the most common reason results are misleading. Chapters 6 and 7 cover reporting mistakes (overclaiming and p-hacking), the most common reason papers get rejected. Chapters 8 and 9 cover scholarship mistakes (reproducibility and misreading the literature), the most common reason results don't hold up. Chapter 10 covers efficiency (compute waste), Chapter 11 covers ethics, and Chapter 12 gives you the final checklist that ties every chapter together before you submit.


Chapter 1: Why Mistakes Matter — How Small Errors Invalidate Research

A research paper makes a claim: "Method X works better than method Y." Every step between the raw data and that claim is a link in a chain, and experimental science has an unforgiving rule — a chain is only as strong as its weakest link. One bad link, such as a contaminated test set or a biased metric, does not slightly weaken the claim. It can make the claim entirely meaningless. This chapter explains why AI research is especially vulnerable to this problem and what it costs.

Why AI experiments are fragile. Machine learning models are excellent at finding shortcuts. If there is any way to predict the target that does not reflect the real problem — a timestamp in the file name, a hospital code in an image, a label that leaked into a feature — the model will find it and exploit it. Human researchers do the same thing by accident: they tune settings against the test set, try many variants until one looks good, or report the metric that flatters the result. The danger is that these errors produce numbers that look like success. A 99% accuracy from a leaked experiment feels exactly like a 99% accuracy from a valid one, right up until someone tries to reproduce it or deploy it [1], [5].

The cost of invalid results. A wrong result costs far more than the time spent on it. For you, it can mean a rejected paper, months of rework, or — worse — a published paper that others later cannot reproduce, which damages your reputation as a researcher. For the field, every invalid result adds noise that future researchers must sort through. Some published results become famous and then collapse; whole research directions have been built on contaminated benchmarks. As a beginner, you have one powerful advantage: you can build careful habits now, before bad habits set in. Reviewers and supervisors notice rigor, and a careful experimenter earns trust that compounds over a career [6].

The psychology of self-deception. Most research mistakes are not cheating. They are self-deception, made easier by three biases. Confirmation bias makes you accept the experiment that "worked" and discount the one that didn't. Optimism bias makes you believe your idea is right, so you stop looking for reasons it might be wrong. And publication pressure rewards positive results, which tempts you to present the best number instead of the most honest one. Knowing these biases exist is the first defense; the second is process — checklists, locked test sets, and documented decisions, all covered later in this book [2], [4].

How reviewers spot invalid research. Experienced reviewers have a mental checklist of red flags, and learning it helps you police your own work. They look for numbers that are "too good" — accuracy above 95% on a hard problem with no error analysis invites suspicion. They check whether the data split is described at all; a methods section with no mention of train/test procedure is an immediate concern. They compare your reported baseline numbers against the original papers — if your reproduction of a prior method scores far below its published number, they suspect a strawman (Chapter 5). They look at whether uncertainty is reported; a single point estimate with no standard deviation suggests one lucky run (Chapters 7 and 8). And they read the limitations section — its absence suggests the authors never asked what could be wrong. None of these checks require rerunning your code. They are reading-comprehension checks, which means you can run them on your own draft before submitting [6].

Building a validity mindset. The practical skill underlying this whole book is learning to attack your own results. After every experiment, spend fifteen minutes trying to destroy your own conclusion: Is there leakage? Did I peek at the test set? Would a trivial baseline do as well? Is the metric flattering? Did I try enough seeds? This adversarial self-review feels unnatural because you want your idea to work, but it is the cheapest form of peer review available — and it catches most mistakes while they are still cheap to fix. Keep a short "threats to validity" note in your lab log for each experiment, listing the ways the result could be wrong and what you did about each. When you later write the paper, that note becomes your limitations section almost for free [1], [4].

Mistakes vs misconduct: where the line is. Not all invalid results are equal, and knowing the categories protects you. Honest mistakes — a leaked split you didn't notice — are embarrassing but forgivable, especially when you correct them openly. Questionable research practices — peeking at the test set, reporting only the best seed, hiding a winning baseline — sit in a gray zone: often done without intent to deceive, but they systematically bias the literature and reviewers treat them harshly. Outright misconduct — fabricating data, plagiarizing — ends careers. The uncomfortable truth is that the gray zone does most of the damage to science because it is common and rarely punished directly. The defense is the same process this book teaches: every checklist, locked test set, and preregistered plan moves you from the gray zone into clearly honest territory. When in doubt, disclose — a paper that says "we tuned on the test set in early experiments, then locked it and reran" is infinitely stronger than one that hides it [4], [6].

The economics of being wrong. Put a price on invalid results and rigor starts looking cheap. A flawed experiment that runs for three weeks on a shared GPU costs hundreds of dollars in compute and, more importantly, three weeks of your finite research time — time that could have tested the next idea. A rejected paper costs two to six months of revision and resubmission cycles. A published-but-wrong paper costs reputation, which is unpriced but career-defining: researchers whose results fail to replicate find future papers scrutinized harder and collaborations harder to win. Against these costs, the preventive measures in this book are nearly free: an afternoon for a leakage audit, a day to lock the test set properly, an hour to write a preregistration note. Viewed economically, rigor is not perfectionism — it is the highest-return investment a researcher makes, because it protects every other investment [1], [6].

Keep a mistake journal. Here is a habit that turns every error into tuition: maintain a short journal where you log each methodological mistake you make or catch — what it was, how you found it, what it cost, and the rule that prevents it next time. Review it monthly; patterns emerge (one student discovered all her mistakes came from rushing the last week before deadlines, and fixed it by freezing experiments 72 hours early). Share anonymized entries with labmates — a lab that talks openly about mistakes builds a culture where catching them early is praised, not punished. This book is, in a sense, the field's collective mistake journal. Yours will become the most valuable research document you own, because unlike papers, it records what not to do — knowledge that is rarely published but constantly needed [4], [6].

Worked example — the mistake and the fix. A student team builds a model to predict whether patients will be readmitted to hospital within 30 days, using records from 2020–2023. Their model reaches 93% accuracy — excellent. But the mistake: one of the features is the discharge date, and the dataset was split randomly. Because readmission patterns changed during the COVID years, the model learned that discharge dates in certain months predict readmission — information that would not be available at prediction time in real use, and that leaked across the random split. The fix: the team switched to a time-based split (train on 2020–2021, validate on 2022, test on 2023) and removed the date feature. Accuracy dropped to 74%, but now it measured something real: the ability to predict future readmissions from information available at discharge. The honest 74% became the foundation of a solid paper; the dishonest 93% would have been rejected or, worse, published and wrong.

For your research: Before you run your first experiment, write down exactly what claim you want to make and what evidence would support it. After each experiment, ask: "Is there any way this number could look good without my idea being right?" If you can name one, fix it before writing the paper.

Putting it together. This chapter's message is simple but easy to forget under deadline pressure: validity is not one thing you check at the end, it is a property of every link in your experimental chain. The researchers who internalize this stop asking "is my accuracy high enough?" and start asking "is my accuracy honest?" — and that single shift in questioning prevents most of the mistakes in this book before they happen. When you combine adversarial self-review with a written threats-to-validity note and an honest accounting of costs, you get something more valuable than any single result: a research practice that produces trustworthy numbers by default. Reviewers, supervisors, and future collaborators can sense this practice within pages of reading your work, and it is what turns a beginner into a researcher others want to build on [1], [6].

Key takeaways: - One invalid link in the experimental chain can destroy the whole claim. - ML models and researchers both find shortcuts — rigor is the defense. - Most mistakes are honest self-deception, not fraud; process beats willpower. - An honest, lower number is worth more than a flattering, invalid one. - Build careful habits now; they compound across your research career.


Chapter 2: Data Leakage — The Silent Killer (With Examples)

Data leakage is the most common and most damaging mistake in applied machine learning. It means that information from outside the training set — most often from the test set or from the future — accidentally influences the training process. The result is a model that looks excellent in your experiments but fails in the real world, because in the real world the leaked information will not be available. It is called the silent killer because nothing in the code raises an error; the numbers simply lie [1], [5].

What leakage looks like. There are three main forms. First, preprocessing leakage: you compute something using the whole dataset before splitting — for example, normalizing features with the mean and standard deviation of all data, or selecting the "best" features using all labels. The test set's statistics have now touched the training pipeline. Second, target leakage: a feature contains information that is only known after the outcome — like including the discharge date when predicting readmission, or a "customer cancelled" flag when predicting churn. Third, duplicate or near-duplicate leakage: the same (or nearly the same) records appear in both train and test, common with data augmentation, web-scraped images, or multiple records per patient. The model effectively memorizes the answers [2].

Why preprocessing leakage is so easy to commit. The standard tutorial pattern — load data, clean it, scale it, then split — is backwards. Scaling before splitting means the test set contributed to the scaling parameters. The effect is often small, which makes it tempting to ignore, but it is still wrong, and reviewers increasingly check for it. The correct pattern is: split first, then fit every preprocessing step (scaling, imputation, feature selection, encoding) on the training data only, and apply the fitted transformers to validation and test data. In scikit-learn, a Pipeline object does this automatically inside cross-validation [8].

Target leakage in disguise. The hardest leaks to spot are features that seem legitimate. A classic example: predicting loan default using a feature called "days since last payment" where missing values were filled with 999 for customers who already defaulted — the imputation itself encodes the answer. Another: predicting machine failure using sensor logs that include the maintenance ticket written after the failure. Rule of thumb: for every feature, ask "would this value be known at the moment I must make the prediction?" If not, remove it, no matter how predictive it is [5].

Leakage in time-series and grouped data. Two data shapes make leakage especially likely. In time-series data (stock prices, sensor readings, patient records over years), a random split puts future records in the training set and past records in the test set — the model trains on the future and is tested on the past, the exact reverse of real deployment. The fix is a time-based split: train on the earliest period, validate on the middle, test on the latest, so the model is always evaluated on data from after its training window [1]. In grouped data — multiple scans per patient, multiple images per farm, multiple transactions per customer — a random split scatters one group's records across train and test. The model then learns to recognize the patient or farm rather than the disease or crop condition, and scores brilliantly on "new" scans of patients it already knows. The fix is group-aware splitting: all records from one group go entirely into train or entirely into test (scikit-learn's GroupKFold does this). Whenever your rows are not independent — and in real datasets they rarely are — ask what the natural grouping is and split by it [2], [8].

A leakage audit routine. Make this five-step audit part of every project. (1) For each feature, write down when its value becomes known relative to prediction time; anything known only afterward is removed. (2) Verify the order of operations: split happened before any fitting of preprocessing, feature selection, or imputation. (3) Check for duplicates and near-duplicates across splits — hash the rows or use similarity search on images and text. (4) Run the "impossible model" test: train with only features that should be useless (IDs, timestamps, file names); if it predicts well above chance, something leaked and you have found your prime suspect. (5) Compare a naive random split against a proper time- or group-aware split; a large performance drop is the signature of leakage. This audit takes an afternoon and has saved many theses [5].

Famous leakage patterns and what they teach. Certain leakage stories repeat across the field so often they have become cautionary archetypes. The hospital-marker story: a model trained to detect disease from X-rays actually learned to recognize which hospital's scanner produced the image, because disease prevalence differed by hospital — it was a hospital classifier wearing a disease detector's clothes. The timestamp story: a model predicting equipment failure keyed on log-file timestamps that encoded maintenance schedules rather than sensor patterns. The identity story: a face-verification model tested on photos of people it had seen during training, measuring memorization instead of recognition. Each story teaches the same lesson from a different angle: models exploit the easiest available signal, and the easiest signal is often an artifact of how the data was collected rather than the phenomenon you study. When your results look surprisingly good, your first hypothesis should be "what shortcut did it find?" — and the audit routine above is how you check [1], [5].

Leakage in feature engineering. Some of the subtlest leaks hide inside feature engineering. Target encoding — replacing a categorical value with the mean target for that category — leaks the label into the feature unless computed strictly within training folds. Imputing missing values with the global mean or median leaks test statistics into training data, just like premature scaling. Dimensionality reduction (PCA, embeddings) fitted on the full dataset lets test points shape the projection the model learns from. Even text preprocessing leaks: building a vocabulary, computing IDF weights, or training word embeddings on the whole corpus before splitting. The unifying fix is the pipeline principle from earlier, applied ruthlessly: every transformation that learns anything from data must learn it from training data only. When you use a library function, check whether it "fits" — if it does, it belongs inside the pipeline, after the split [2], [8].

Quick self-test: is my pipeline leak-free? Answer these five questions before any important run. (1) Did the split happen before any fitting, including scaling, imputation, encoding, and feature selection? (2) Could any feature's value be unknown at prediction time? (3) Could the same entity (patient, user, farm, device) appear on both sides of the split? (4) If the data has time structure, does the test set come strictly after the training period? (5) Would the "impossible model" test — training on IDs and timestamps alone — score near chance? Five yes/no answers, two minutes, and you have either confidence or a specific problem to fix. Tape this test next to your monitor; it is the cheapest insurance in this book [5], [8].

Worked example — the mistake and the fix. A beginner classifies emails as spam or not spam. The mistake: they build their TF-IDF vocabulary and remove stopwords using the entire email corpus, then split into train and test. The test set's word statistics have leaked into the vocabulary, and words that only appear in test emails influence the feature space. Reported accuracy: 98.2%. The fix: wrap the vectorizer and classifier in a single scikit-learn Pipeline, and split the raw emails first. Inside cross-validation, the vocabulary is built only from each training fold. Accuracy drops slightly to 97.6% — a small change here, because this leak was mild — but now the evaluation is honest, and more importantly the habit is correct. When the same student later works on a medical dataset where feature selection on the full data was the difference between 91% and 78%, the pipeline habit saves the whole project [1], [8].

Figure 1 — Data leakage as contaminated pipes

For your research: Make "split first, preprocess inside a pipeline" your non-negotiable rule from day one. When you describe your method in a paper, state exactly when the split happened and what was fitted on which data — reviewers look for this sentence, and its absence is a red flag.

Putting it together. Data leakage is the mistake you must fear most precisely because it never announces itself — the code runs, the numbers look great, and only disciplined process stands between you and a false result. The complete defense has four layers: split before you preprocess, keep every fitting step inside training-only pipelines, audit every feature for prediction-time availability, and respect the structure of your data with time- and group-aware splits. Run the five-question self-test before important runs and the full audit before writing up. Students who adopt these habits early find that leakage stops being scary and becomes routine hygiene — like washing your hands. And when a reviewer asks "how did you guard against leakage?", you will have a precise, confident answer instead of a vague hope [1], [5], [8].

Key takeaways: - Data leakage = test or future information influencing training; it produces flattering but meaningless numbers. - The three forms: preprocessing leakage, target leakage, and duplicate leakage. - Split the data first; fit all preprocessing on training data only (use pipelines). - For every feature, ask: "Would this be known at prediction time?" - Small leaks still matter — they teach you the habits that prevent big ones.


Chapter 3: Test Set Contamination and Peeking

If data leakage is the silent killer, test set contamination is its close relative: using the test set, directly or indirectly, to make decisions during development. The test set has one job — to give an unbiased estimate of how your model performs on new data. Every time you peek at it, that estimate becomes biased, and the number you report becomes a story about your tuning process rather than your model's quality [1], [4].

Direct contamination. The obvious form is training on the test set — rare and usually accidental, for instance when a data file contains rows from both sets, or when a web-scraped dataset unknowingly includes benchmark test images. A subtler direct form is using the test set for early stopping or model selection: you train 50 model variants and pick the one with the best test score. The test set has now become a validation set, and the reported score is optimistic because you selected the luckiest variant [2].

Indirect contamination — the peeking problem. This is the form beginners commit most. You run an experiment, check the test score, tweak a hyperparameter, check again, adjust the architecture, check again. Each peek lets test information leak into your decisions. After twenty rounds of this, your "test accuracy" measures how well you tuned to the test set, not how well the model generalizes. The correct workflow uses three splits: train (learn the model), validation (make all decisions — hyperparameters, architecture, early stopping), and test (used once, at the very end, to report the final number). If the test number disappoints you, you do not get to tune and re-test; you go back to the drawing board and, strictly speaking, you need a fresh test set [5].

Contamination from the literature. A newer form of this mistake comes from large pretrained models and benchmarks. If you evaluate a model that was pretrained on data including your benchmark's test set — very common with web-scale pretraining — your "zero-shot" result may be contaminated. Similarly, if a benchmark's test labels have circulated in training data on the internet, fine-tuning on "public" data may include them. Careful researchers now document exactly what data their base models saw and choose benchmarks designed after the model's training cutoff [6], [7].

The validation set done right. A validation set is only useful if it resembles the test set. That means splitting it with the same procedure — stratified by class so rare classes appear in both, grouped or time-ordered if the test set will be (Chapters 2 and 3 go together). Size matters: too small and your model-selection decisions are noise — with 50 validation examples, a 4% "improvement" is within chance. A common rule is 10–20% of the data for validation, more when the dataset is small. When data is scarce, use nested cross-validation: an outer loop estimates performance while an inner loop does model selection, so every example serves both purposes without contamination. And treat your validation score as what it is — an estimate with uncertainty. If two configurations differ by less than the validation noise, prefer the simpler one; chasing noise is how you overfit the validation set itself [1], [2].

Contamination in the age of foundation models. Pretrained models add a new contamination channel. A model trained on web-scale data has likely seen your benchmark's test examples — or discussions of them — during pretraining, so its "zero-shot" score may reflect memorization rather than capability. Researchers now check training cutoff dates against benchmark release dates, and benchmark designers plant canary strings (unique nonsense phrases) in test sets so they can later search model training corpora for them. There is also community-level contamination: when the whole field tunes against the same public benchmark for years, the benchmark stops measuring generalization and starts measuring benchmark-fitting — a version of Goodhart's law. Protect yourself by evaluating on at least one benchmark released after your model's training cutoff, by keeping a private held-out set for important claims, and by stating exactly what data your base model saw [6], [7].

Building a personal no-peek workflow. Good intentions fail under deadline pressure, so design a workflow where peeking is difficult. Keep the test data in a separate directory that your training scripts never read — better yet, keep it out of your working notebook entirely until the final evaluation script runs. Write one script, final_evaluation.py, whose only job is loading the locked test set and the saved best model and printing the score; run it once and save the output. During development, if you catch yourself wanting "just a quick look" at test performance, treat that urge as a signal to improve your validation setup instead — a validation set you don't trust is the real problem. Some labs enforce this socially: the test set lives with the supervisor, and you request the single evaluation run when your validation work is done. Whatever the mechanism, the principle is that discipline should be built into the environment, not left to willpower at 2 a.m. before a deadline [5].

Cross-validation done right. Cross-validation (CV) is the standard tool for model selection on limited data, and it has its own contamination traps. The classic error is preprocessing outside the CV loop: scaling or feature-selecting on the full data, then cross-validating only the classifier — the leakage from Chapter 2, now inside every fold. The fix is the same pipeline principle: the entire preprocessing-plus-model chain goes inside each fold. Use stratified folds for classification so every fold mirrors the class balance, and group-aware folds when rows are grouped (Chapter 2). Repeat the CV with different random seeds when the dataset is small, because a single CV split is itself noisy. And remember what CV gives you: a good estimate for model selection, not a replacement for the final locked test set — after CV picks the winner, that winner still gets exactly one evaluation on untouched test data [1], [8].

When the test score disappoints. Sooner or later you will run the final evaluation and the number will be worse than your validation suggested. The temptation — tune a little more, peek once more — will be strong. Don't. A disappointing test score is information: your validation setup was optimistic, your method overfits, or the problem is harder than you thought. The honest responses are: report the number as is and investigate why (this becomes your limitations and future-work section); improve the method and validate the improvement properly, accepting that you now need fresh test data for a clean claim; or redesign the validation to match the test distribution better. What you must not do is iterate against the test set until the number recovers — that converts a disappointing-but-honest result into a flattering-but-false one, which is strictly worse [1], [5].

Worked example — the mistake and the fix. A student tunes a neural network for flower classification. The mistake: they have only a train/test split. They try learning rates 0.1, 0.01, 0.001, and 0.0001, batch sizes 16, 32, and 64, and two architectures — 24 combinations — each time recording the test accuracy and keeping the best. Final reported accuracy: 94%. But this 94% is the maximum over 24 tries on the test set; the expected performance on truly new data is lower, and the paper's claim is inflated. The fix: split off a validation set (say 15% of the data) before any tuning. All 24 combinations are compared on validation accuracy only; the winner is then evaluated once on the untouched test set, scoring 91%. The paper reports 91% — lower, but honest and defensible. Bonus fix: the student also reports the validation scores of all 24 runs in an appendix, showing the work transparently [1], [5].

For your research: Create the test set first, lock it away (even a separate file you don't open), and make every decision on validation data. Write in your paper: "The test set was evaluated once, after all model selection." That single sentence signals careful methodology to reviewers.

Putting it together. The test set is the one part of your experiment that must remain pristine, because it is the only unbiased witness you have. Everything in this chapter — the three-way split, the no-peek workflow, validation done right, cross-validation inside pipelines, vigilance about foundation-model contamination — serves that single goal. The practical test is behavioral: if you can honestly write "the test set was evaluated once, after all model selection," your results deserve trust. If you cannot, no amount of sophisticated modeling compensates. Build the environment so that honesty is easy — separate directories, a single final-evaluation script, a supervisor-held test set if needed — and treat every urge to peek as a signal to fix your validation setup instead. Guarded this way, your final number means what it says [1], [5].

Key takeaways: - The test set's only job is a final, unbiased estimate — never use it for decisions. - Peeking (repeated tuning against the test set) biases the reported score upward. - Use three splits: train, validation, test — and evaluate the test set exactly once. - Watch for contamination via pretrained models and internet-circulated benchmarks. - A slightly lower honest number always beats an inflated, contaminated one.


Chapter 4: Metric Misuse — Accuracy Theater and Friends

A metric is a lens, and the wrong lens can make a bad model look excellent. Beginners often reach for accuracy because it is simple and familiar, then discover — sometimes only when a reviewer asks — that their 99% accurate model is useless. This chapter teaches you to choose metrics that match your research question and to report them honestly [1], [4].

Why accuracy lies. Accuracy counts the fraction of correct predictions, which is fine when classes are balanced. But many real problems are imbalanced: 99% of transactions are legitimate, 98% of patients are healthy, 95% of emails are not spam. A model that always predicts the majority class scores 99%, 98%, 95% accuracy while learning nothing — this is "accuracy theater." If your paper's headline number is accuracy on an imbalanced problem, reviewers will dismiss the result, and rightly so. The fix is to report metrics that respect the imbalance: precision (of the predicted positives, how many were right), recall (of the actual positives, how many were found), and the F1 score (their harmonic mean) — computed for the class you actually care about, not averaged into meaninglessness [5].

Choosing the right metric for the job. Different questions need different lenses. Medical screening cares about recall — missing a disease is worse than a false alarm. Spam filtering cares about precision — a false positive hides real mail. Ranking problems (search, recommendations) need ranking metrics like mean average precision, not classification accuracy. Regression needs error measures like MAE or RMSE, reported in real units, not just R². Probabilistic predictions need calibration checks, not just the argmax accuracy. Before running experiments, write down which mistake is expensive in your application and pick the metric that penalizes it [2].

The averaging trap and the single-number trap. On multi-class problems, beginners often report a single accuracy or a macro-average that hides catastrophic failure on a rare but important class. Report per-class metrics in a table so readers can see the weak spots. Conversely, never report a metric without its uncertainty: a single test score is one sample. Report confidence intervals or standard deviations across folds or seeds, and use proper statistical tests before claiming one method beats another — a 0.5% gap with overlapping error bars is not a win [1], [4].

Beyond classification: metrics for other tasks. Every task family has its own metric traps. In ranking and retrieval (search engines, recommenders), accuracy is meaningless — use ranking metrics like mean reciprocal rank or normalized discounted cumulative gain, which reward putting the right items near the top. In regression, report error in real units (mean absolute error in dollars, days, or degrees) rather than hiding behind R² alone, and consider whether large errors matter disproportionately — if they do, RMSE penalizes them more than MAE, and you should say why you chose one. In text generation, automatic scores like BLEU or ROUGE correlate weakly with human judgment; treat them as rough guides and budget for human evaluation of a sample of outputs. In clustering, internal scores like silhouette measure compactness, not correctness — validate against known labels or expert review when you can. The principle never changes: the metric must measure what the application actually needs [2], [4].

Reporting uncertainty the right way. A bare number — "accuracy 87.3%" — pretends a precision your experiment does not have. Always accompany headline metrics with uncertainty: the standard deviation across cross-validation folds or random seeds, or a confidence interval. When comparing two methods, use a proper paired statistical test (paired t-test or Wilcoxon signed-rank across folds or seeds) rather than eyeballing two numbers. And distinguish statistical significance from practical significance: with enough data, a 0.1% gain can be "significant" (p < 0.05) yet meaningless for any real application — report the effect size and let readers judge. A results table with means, standard deviations, and significance markers is the mark of a careful experimenter; a table of bare maxima is the mark of a careless one [1], [4].

Choosing thresholds and operating points. Most classifiers output scores, not decisions — and the threshold that turns a score into a "yes" or "no" is a choice, not a given. The default 0.5 threshold is arbitrary; the right threshold depends on the costs of each error type. In cancer screening, you lower the threshold to catch more cases (high recall) at the price of more false alarms. In fraud blocking, you raise it to avoid freezing innocent customers' accounts (high precision). Tune the threshold on validation data using the metric that reflects your application's costs, report the chosen threshold alongside your results, and show how performance varies across thresholds with a curve — readers can then judge whether your operating point suits their use case. A paper that reports metrics at an unexplained threshold is hiding a degree of freedom (Chapter 7); a paper that justifies its operating point is doing applied science [4], [5].

Calibration: when probabilities must mean something. Accuracy, precision, and recall judge hard decisions, but many systems act on predicted probabilities — triage queues, risk scores, abstention rules ("ask a human when unsure"). A model can be 90% accurate yet badly calibrated: predicting 0.9 for cases that occur 60% of the time. If your paper claims the model's scores can guide decisions, you must check calibration — plotting predicted probability against observed frequency, or reporting expected calibration error. Miscalibrated models are a quiet deployment failure: everything looks fine in the results table, but the downstream decision policy built on the scores misbehaves. Simple fixes like temperature scaling or isotonic regression, fitted on validation data, often help. The broader lesson: match your evaluation to how the model will actually be used, not just to what is easy to compute [2], [4].

Metric checklist by task type. Keep this quick reference when designing evaluations. Binary classification with imbalance: precision, recall, F1 on the minority class, plus the confusion matrix. Balanced multi-class: per-class precision/recall and macro-F1, never just accuracy. Ranking/retrieval: precision@k, mean reciprocal rank, or NDCG. Regression: MAE and RMSE in real units, with residual plots. Probabilistic outputs: calibration curves or expected calibration error alongside accuracy. Generation: automatic scores as rough guides plus human evaluation on a sampled subset. Clustering: internal indices plus validation against any available labels or expert judgment. For every task: report uncertainty (std across seeds/folds), use paired tests for comparisons, and justify the operating point or threshold. If your task isn't listed, derive the metric from the application's cost of errors — that derivation itself belongs in your paper [2], [4].

Worked example — the mistake and the fix. A student builds a fraud detector on transactions where only 1% are fraudulent. The mistake: the model predicts "not fraud" for almost everything and the paper's headline is "99.1% accuracy." A reviewer notes that a model predicting nothing would score 99% — the result says nothing about fraud detection. The fix: the student reframes the evaluation around the minority class. They report precision = 62%, recall = 71%, F1 = 0.66 for the fraud class, plus a precision-recall curve and the area under it. They also report the confusion matrix and note the business trade-off: at the chosen threshold, the system catches 71% of fraud while flagging 1.3% of legitimate transactions. The headline number is lower, but the paper now actually evaluates fraud detection, and the discussion of the precision-recall trade-off becomes one of the paper's strongest sections [5].

For your research: Before your first experiment, write one sentence: "A good model for my problem is one that ___," and choose the metric that measures exactly that. Report per-class results and uncertainty (confidence intervals or std across seeds) alongside every headline number — reviewers check for this.

A last check before you finalize. When your results table is done, do one final pass with fresh eyes: cover the metric column and ask whether the experimental setup alone convinces you the comparison is fair — same splits, same budgets, same reporting. Then uncover the numbers and check that every headline claim in your abstract appears in the table with its uncertainty. This two-minute ritual catches the metric mistakes that survive everything else: the forgotten baseline, the unexplained threshold, the missing error bars. Metrics are the language your paper speaks to the field; make sure it's saying what you mean.

Putting it together. Choosing metrics is where your research question meets your evidence, and getting it right is mostly a matter of asking "what does good mean here?" before running anything. Derive the metric from the application's costs, report it for the classes and subgroups that matter rather than hiding behind averages, justify your operating threshold, check calibration when probabilities drive decisions, and always attach uncertainty. A results section built this way reads as a genuine investigation rather than a beauty contest — and it gives reviewers exactly what they look for: evidence that the numbers measure the claimed capability. The headline number may be smaller than accuracy theater would allow, but it will be a number you can defend in front of anyone [1], [4], [5].

Key takeaways: - Accuracy is misleading on imbalanced data — "accuracy theater" hides useless models. - Report precision, recall, and F1 for the class that matters, plus confusion matrices. - Match the metric to the cost of mistakes in your application. - Never hide weak classes behind averages; never report a number without uncertainty. - A smaller, honest metric beats a big, meaningless one.


Chapter 5: Ignoring Baselines — Beating Nothing Is Not a Result

A new method is only an improvement if it beats what came before — and what came before includes simple, strong baselines, not just the weakest prior method you could find. Beginners often compare their sophisticated model against nothing, against a poorly tuned competitor, or against a strawman, and then claim victory. Reviewers have seen this a thousand times, and it is one of the fastest routes to rejection [1], [6].

What counts as a baseline. At minimum, every experiment should include: (1) a trivial baseline — majority-class prediction, random guessing, or predicting the mean — which calibrates your metric (if your model barely beats majority-class, say so); (2) the simplest reasonable method for the task — logistic regression or a small decision tree for tabular data, a standard pretrained model for images or text; and (3) the best prior published method you can run, tuned fairly. If your paper claims to improve on prior work, you must actually run that prior work under the same conditions — same data split, same metric, same compute budget — not copy a number from their paper [5].

Why skipping baselines is so tempting. Strong baselines are humbling. A well-tuned gradient-boosted tree often beats a fancy neural network on tabular data; a simple TF-IDF model sometimes matches a large language model on classification. Beginners skip them because the baseline might win — but a baseline that wins is a finding, not a failure. It tells you your problem doesn't need complexity, which is valuable knowledge and can itself be publishable as a careful empirical study. What is never acceptable is hiding the baseline because it was inconvenient [8].

Fair comparison discipline. Baselines must be treated fairly: tune their hyperparameters with the same effort you spend on your method, give them the same features and data, and report their full results. The classic cheat is to compare your heavily tuned model against a competitor's default settings — reviewers call this a strawman baseline, and it destroys credibility. Document the tuning budget for every method (e.g., "each method received 20 validation-evaluated configurations"), so the comparison is visibly fair [1].

How reviewers see through weak baselines. Reviewers have read the prior literature, and many have run the competing methods themselves. When your baseline numbers are far below the numbers in the original papers, they notice instantly — and they know the usual causes: default hyperparameters, less training time, weaker data augmentation, or an outdated implementation. Some reviewers will even rerun a baseline during review to check. The reputational cost is severe: a paper caught with strawman baselines is not just rejected, it marks the authors as unreliable. By contrast, authors who report a baseline beating their method on some datasets earn unusual trust — the reviewer thinks, "these people tell the truth even when it hurts," and reads the rest of the paper with confidence [6].

When the baseline wins: turning it into a contribution. Suppose the tuned tree beats your neural network. You have three honest options, all publishable. First, characterize where each wins: perhaps your method dominates on large datasets while the tree wins on small ones — that interaction is a genuine finding. Second, ablate your method to find what actually helps; maybe one component matches the tree and the rest is decoration, which simplifies future work. Third, reframe the contribution: "a careful empirical study showing that simple methods remain competitive on this task" is a paper the field needs, and such papers are often highly cited because they save everyone else from the same dead end. What you must never do is hide the winning baseline — someone will run it, and the discovery that you concealed it is far worse than the baseline itself [1], [8].

Baselines for deep learning papers. When your method is a neural network, the baseline bar is higher than beginners expect. The standard strong baselines include: a well-tuned standard architecture for your task (a ResNet for images, a transformer for text) rather than an exotic one; a pretrained model fine-tuned on your data, since pretraining is the dominant source of gains in many domains; and, where applicable, the non-deep method that practitioners actually use (gradient-boosted trees for tabular data). Crucially, match the training budget: if your method trains for 100 epochs with heavy augmentation, the baseline gets the same. Report baseline hyperparameters so readers can verify fairness. And remember that "we compare against the published numbers" is weaker than rerunning — published numbers used different splits, preprocessing, and sometimes different metric definitions. Rerun the baseline in your codebase; the week it costs is the foundation your claim stands on [7], [8].

Ablations as internal baselines. Baselines compare against other methods; ablations compare against simpler versions of your own method — and reviewers expect both. An ablation removes or disables one component at a time (the attention module, the auxiliary loss, the data augmentation) and measures the drop. This answers the question every reviewer asks: "which part of this actually matters?" Often the answer is humbling — the fancy component contributes 0.3% while the data cleaning contributes 5% — but that knowledge is precisely what makes the paper valuable. Report ablations in a dedicated table, and discuss the surprising ones; a component that does nothing is a finding about the problem, not a failure of the paper. A method with thorough ablations reads as understood; a method without them reads as a lucky bundle of tricks [1], [7].

Documenting your baseline protocol. Reviewers judge fairness from what you write, so make the protocol explicit in a short paragraph: which baselines you ran, where their implementations came from, the hyperparameter search space and budget each received, the data splits and metrics used, and the number of seeds. Something like: "We compare against logistic regression, gradient-boosted trees, and the method of Smith et al. [4] (official implementation). Each method received 30 validation-evaluated configurations over the search spaces in Appendix B, trained on identical splits; we report mean ± std over 5 seeds." This paragraph preempts the most common reviewer objection in ML — "the baselines look weak" — because it shows the work. Put the full search spaces and per-configuration results in an appendix; transparency here is cheap and credibility is expensive [1], [8].

Worked example — the mistake and the fix. A student develops a graph neural network to predict house prices from neighborhood data and reports RMSE of 41,000 (currency units), claiming a new state of the art. The mistake: there is no baseline — no comparison at all. The fix: the student adds three baselines under identical splits and tuning budgets. A mean predictor scores RMSE 78,000 (the trivial floor). A linear regression on the same features scores 44,000. A gradient-boosted tree scores 39,500 — better than the new GNN. The honest conclusion: the GNN improves on linear models but does not beat a tuned tree baseline on this data. Instead of a false "state of the art" claim, the paper becomes a careful study showing where graph structure helps (the GNN wins on a subset of dense urban neighborhoods) and where it doesn't — a more honest, more interesting, and more publishable contribution [1], [8].

For your research: Before claiming any improvement, run at least a trivial baseline and the simplest strong method for your task, tuned with equal effort. Put all baselines in one table with identical conditions. If a baseline wins, investigate why — that investigation is often your real contribution.

A final word on intellectual honesty. There is a moment in every honest project when the numbers disappoint — the baseline wins, the ablation shows your idea adds nothing, the test score drops. How you respond in that moment defines you as a researcher. The easy path is to tweak the comparison until your method wins; the right path is to report what happened and think hard about why. Paradoxically, the right path produces better papers: "why the simple method wins" is a more interesting and more cited contribution than another marginal improvement. Reviewers can tell which path you took, usually within minutes. Choose the honest one every time, especially when it costs you — that is precisely when it matters most.

Putting it together. Baselines are how science stays honest about progress: without them, every paper is an island claiming to be a continent. The complete practice is straightforward — run trivial, simple-strong, and prior-method baselines under identical conditions, tune them with equal budgets, ablate your own method's components, document the whole protocol, and report whatever wins with a straight face. When the baseline wins, you have not failed; you have learned something about the problem that is often more publishable than the method you set out to build. Reviewers reward this honesty disproportionately because it is rare. Make fair comparison your signature as a researcher, and your claimed improvements — when they come — will be believed [1], [5], [8].

Key takeaways: - Beating nothing is not a result — every claim of improvement needs baselines. - Include trivial, simple-strong, and prior-method baselines under identical conditions. - Tune baselines fairly; a strawman baseline destroys credibility. - A baseline that wins is a finding, not a failure — report it honestly. - Document equal tuning budgets so the comparison is visibly fair.


Chapter 6: Overclaiming, Hype, and "State of the Art" Abuse

Overclaiming means stating more than your experiments support. It ranges from slightly inflated language ("significantly outperforms") to the full "state of the art" (SOTA) claim based on one small dataset. Hype is overclaiming's louder cousin: framing ordinary results as breakthroughs. Both are epidemic in AI, both are easy for beginners to absorb from reading hyped papers, and both are punished by reviewers [6].

The anatomy of an overclaim. Common forms: claiming SOTA from a single dataset when prior work used several; claiming "significant" improvement for a 0.3% gap with no statistical test; generalizing from one domain ("works on our hospital's X-rays") to all of medicine; implying causation from correlation; and the vague "novel" — every paper is novel in some trivial sense, so the word carries no information. Ask of every sentence in your results section: "Which table or figure proves this exact sentence?" If none does, delete or soften the sentence [4].

Why "state of the art" is usually abused. A legitimate SOTA claim requires: beating the best known prior method, on the standard benchmark(s) the field uses, under comparable conditions, with statistical significance. Beginners typically violate at least two of these — they compare against weak baselines (Chapter 5), use a private or tiny dataset, or test once. Worse, SOTA-chasing distorts research itself: it rewards tiny benchmark gains over genuine insight. Many respected researchers now advise students to aim for clear, honest contributions — a well-understood failure mode, a careful ablation, a new problem formulation — rather than leaderboard positions [6], [7].

Calibrated language. Replace hype with precision. Instead of "our method dramatically outperforms all existing approaches," write "our method improves F1 by 3.2 points (p < 0.05) over the strongest baseline on Dataset X; on Dataset Y the difference is not significant." State limitations explicitly — every method has them, and reviewers trust authors who name theirs. A limitations section is not a confession of failure; it is evidence of understanding, and most top venues now expect one [6].

Reading hyped papers without absorbing the hype. You will read many hyped papers, and you need their knowledge without catching their habits. Practice translating marketing into evidence as you read: "revolutionary framework" becomes "a modified architecture"; "solves X" becomes "improves metric M on dataset D"; "outperforms all prior work" becomes "beats the two baselines they ran." Then check what is missing: no error bars, no ablation, no discussion of failure cases, limitations buried or absent. The real contribution is usually smaller than the abstract and still valuable — a good idea tested on one dataset is worth knowing about. Train yourself to extract the idea and the evidence while discounting the adjectives, and your own writing will naturally become calmer and more credible [6].

The honest abstract template. Write abstracts to this structure and overclaiming becomes nearly impossible: one sentence on the problem, one on what you did, two to three on results with actual numbers and comparisons ("87.3% vs 84.1% baseline, p = 0.02"), and one on limitations or scope ("evaluated on images from two farms; performance drops on a third"). Every number in the abstract must appear in the results section. Notice what this template excludes: the words novel, significant (without a test), state-of-the-art (without meeting its conditions), and any claim about data you did not test. An abstract written this way may feel modest — that modesty is exactly what makes reviewers trust the paper [4], [6].

The incentives behind hype — and how to resist. Hype persists because it is rewarded: bold claims attract citations, media coverage, and funding, while careful caveats feel like they weaken a paper. Lab culture matters enormously here — if your group celebrates leaderboard positions over understanding, you will feel pressure to inflate. Resist structurally, not just morally. Agree with your supervisor at the project's start on what counts as a contribution (a well-understood result, not just a bigger number). Write the limitations section before the introduction, so honesty is baked in rather than bolted on. And remember the long game: researchers known for inflated claims get extra scrutiny on every future paper, while researchers known for careful work get the benefit of the doubt. Your reputation is the only asset that compounds across your entire career; no single paper's headline is worth spending it [6].

Titles and claims that survive contact with reviewers. Your title is a claim, so write it like one you can defend. "A Graph Neural Network for Crop Yield Prediction" survives; "Revolutionary AI Solves Crop Yield Prediction" does not. Apply the same discipline to every section heading and topic sentence: each should be provable from the adjacent tables. A useful exercise is the "claim downgrade ladder": write your result claim, then produce three progressively weaker versions (e.g., "best on all benchmarks" → "best on two of three benchmarks" → "competitive with the best prior method" → "improves over our baseline under these conditions"). Your paper should state the strongest version the evidence actually supports — usually one or two rungs down from your first draft. Reviewers rarely punish modest claims; they routinely punish inflated ones, because an inflated claim forces them to do the downgrading work you skipped [6].

Negotiating claims with coauthors. Overclaiming is often a team sport: one coauthor wants the bold abstract, another wants caution, and the loudest voice wins. Handle this with process, not argument. Make the claim audit (Chapter 6's exercise) a shared pre-submission ritual — sit together, highlight every superlative, and demand the backing table for each. Frame caution as strategy, not timidity: "a reviewer will force us to soften this anyway; let's do it ourselves and keep control of the wording." Agree early on authorship-order implications too, since claim strength affects everyone's reputation equally. Teams that normalize honest drafting produce papers that survive review with fewer bruises — and the junior author who learns to push back politely on hype is learning one of research's most valuable social skills [6].

Worked example — the mistake and the fix. A student's draft abstract reads: "We propose a novel deep learning framework that achieves state-of-the-art results on image classification, significantly outperforming all existing methods." The mistakes: "novel" is empty, "state-of-the-art" is based on one small custom dataset with no standard benchmark, "significantly" has no statistical test, and "all existing methods" were never run. The fix: the student rewrites based on what the tables actually show — "We adapt a convolutional architecture to low-resolution crop images. On our 4,200-image dataset, it reaches 87.3% accuracy versus 84.1% for a fine-tuned ResNet baseline (paired t-test, p = 0.02, n = 5 seeds). Performance drops to 79% on images from a second farm, indicating limited generalization across sites — see Section 5." Every clause is now backed by evidence, the limitation is stated, and the paper reads as careful rather than hyped. This version survives review; the original would not [1], [4].

For your research: Do a final "claim audit" before submitting: highlight every comparative or superlative word in your draft (novel, significant, state-of-the-art, best, first), and for each one, point to the exact table, figure, or test that justifies it. Cut or soften anything you cannot justify — reviewers do this audit too.

Why modesty wins in the long run. Consider two researchers: one publishes bold claims that get attention but collapse under scrutiny, the other publishes careful claims that hold up. After five years, the first is known for work you can't trust and gets extra-hostile reviews; the second is known for work you can build on and gets the benefit of the doubt, more citations, and more collaborations. Hype is a loan with terrible interest — it borrows attention now and repays it in credibility later. Calibrated claiming is the opposite: a slow investment that compounds. As a beginner, you are choosing which researcher to become with every abstract you write. Choose the one whose papers survive contact with reality.

Putting it together. Honest claiming is a skill with two parts: the technical discipline of tying every statement to evidence, and the social discipline of resisting hype when everyone around you is inflating. The tools are concrete — the claim audit, the downgrade ladder, the honest abstract template, stated limitations, negotiated coauthor standards — and they all serve one reputation: the researcher whose papers say exactly what the experiments support, no more. That reputation compounds: reviewers trust your numbers, collaborators trust your judgment, and your own future work rests on a literature you didn't pollute. In a field drowning in superlatives, calibrated language is not modesty — it is a competitive advantage [4], [6].

Key takeaways: - Every claim must be backed by a specific table, figure, or statistical test. - "State of the art" requires standard benchmarks, strong baselines, and significance — rarely all present. - SOTA-chasing rewards leaderboard games over genuine insight; aim for honest contributions. - State limitations explicitly — it builds trust and is expected at top venues. - Calibrated language ("improves X by Y on dataset Z") always beats hype.


Chapter 7: p-Hacking and Researcher Degrees of Freedom

p-hacking means trying many analyses and reporting only the ones that "worked" — the ones with p < 0.05 or the highest score. Researcher degrees of freedom are the many small choices in an experiment (which metric, which split, which preprocessing, which seed) that can each nudge the result. Together, they let an honest researcher accidentally manufacture a false positive. This chapter is about closing those loopholes [4].

How p-hacking happens in ML. You try 5 architectures × 4 learning rates × 3 seeds = 60 runs, and the best run beats the baseline by 2%. You report that run. But with 60 tries, something will look good by chance — the reported number is the maximum of 60 random-ish outcomes, not a typical outcome. Other variants: testing many metrics and reporting the best one, trying several data splits and keeping the flattering one, removing "outlier" runs after seeing results, or stopping data collection the moment significance appears. None of this requires dishonesty — just the very human habit of exploring until something looks publishable [1].

The garden of forking paths. Even without deliberate hacking, researcher degrees of freedom inflate false positives: every choice you make after seeing the data (which features to include, which baseline to compare, which subgroup to highlight) is a fork in the path, and you naturally walk down the path that leads to the nicer result. The defense is to separate exploration from confirmation: explore freely on validation data, but pre-commit to the final evaluation plan — metric, split, baselines, statistical test — before running the decisive experiment [2], [4].

Practical defenses. Pre-register your key hypotheses and analysis plan (even a dated document in your project folder counts). Decide your stopping rules and sample sizes before seeing results. Report all the analyses you tried, not just the winners — a short "we also tried X, which did not help" paragraph is a mark of honesty. Correct for multiple comparisons when you run many tests (Bonferroni correction or false discovery rate control). And always report effect sizes with confidence intervals, not just p-values: a "significant" 0.2% improvement is practically meaningless [1], [4].

The replication crisis and machine learning. Psychology and medicine went through a painful reckoning when large-scale replication attempts found that many famous results did not hold up — driven by p-hacking, publication bias toward positive results, and underpowered studies. Machine learning has its own version brewing: benchmark-chasing rewards the maximum over many tries, negative results go unpublished so the field never learns what failed, and many papers report single-seed numbers that others cannot reproduce. Knowing this history helps you in two ways: it explains why reviewers have become strict about seeds, ablations, and statistical tests, and it shows you that the researchers most respected in the long run are the careful ones whose results survive replication [4], [6].

Preregistration for ML experiments. You do not need a formal registry to get most of preregistration's benefits. Before your final experimental runs, write a dated document — even a text file in your project folder — stating: the hypothesis, the datasets and splits, the exact metric, the baselines, the number of seeds, the statistical test, and the criteria for claiming a win. Then run exactly that. Exploratory work stays free and creative on validation data; the preregistered run is the confirmatory test. When the confirmatory result differs from exploration — and it sometimes will — report both and explain the difference. This practice single-handedly eliminates most p-hacking, and mentioning it in your paper ("our evaluation protocol was fixed before the final runs") is a strong credibility signal [1], [4].

Exploratory vs confirmatory: running both honestly. Good research needs both modes, and the mistake is mixing them up. The exploratory phase is creative and free: try architectures, poke at the data, follow hunches, look at every metric — all on training and validation data. Keep notes on what you tried, because this phase generates your hypotheses. The confirmatory phase is strict: the preregistered protocol, the locked test set, the fixed seeds, the single evaluation. The confirmatory result is what the paper's claims rest on. When exploration and confirmation disagree — the idea that looked great in exploration flops in confirmation — believe the confirmation. That disagreement is the system working: it just saved you from publishing a false positive. Report the exploratory journey briefly in the paper (it shows diligence); rest the conclusions on the confirmatory evidence (it shows rigor) [1], [4].

The file-drawer problem and negative results. Publication bias means positive results get published while negative ones sit in file drawers — so the literature systematically overstates what works. In ML this takes a specific form: the failed architecture, the hyperparameter region that collapses, the baseline that beat the new method — all omitted, leaving future researchers to rediscover the same dead ends. You can push back in small, publishable ways. Include a "negative results" paragraph or appendix section describing what you tried that failed and why you think it failed; reviewers generally respect this. Frame null findings as boundary-setting: "method X does not help when Y" is knowledge. And when your main result is negative — the careful study showing no improvement — consider venues and workshops that welcome such work; a well-executed negative result that saves the field wasted effort is a genuine contribution, and citing it later feels far better than hiding it now [4], [6].

Multiple comparisons: a concrete example. Suppose you test your method against a baseline on 10 datasets, hoping for p < 0.05 on at least one. Even if your method is identical to the baseline, each test has a 5% false-positive chance — across 10 independent tests, the chance of at least one false "win" is about 40%. This is why uncorrected multiple testing manufactures discoveries. The simplest fix is the Bonferroni correction: divide your significance threshold by the number of tests (0.05/10 = 0.005). It's conservative but honest. A more powerful alternative is controlling the false discovery rate, which tolerates a small proportion of false positives among your claimed wins. Either way, the principle is: the more chances you give yourself to get lucky, the stronger the evidence each individual win must provide. Report how many tests you ran, not just the significant ones [1], [4].

Worked example — the mistake and the fix. A student compares their new optimizer against Adam on 6 datasets. The mistake: they run each comparison with 3 seeds, and for each dataset they report the best seed's result. On 4 of 6 datasets their optimizer "wins." But selecting the best of 3 seeds per dataset is 6 chances to get lucky, and the wins are small. The fix: the student pre-commits to reporting the mean and standard deviation across 5 fixed seeds per dataset, with a paired statistical test. Now the optimizer wins on 2 datasets, ties on 3, and loses on 1 — a mixed but honest result. They also report all 6 datasets (including the loss) instead of quietly dropping the unfavorable ones. The paper's conclusion becomes "competitive with Adam, better on datasets with property P" — narrower, true, and more useful to the field than the inflated original [1].

For your research: Write your analysis plan before the final runs: which metric, which seeds, which statistical test, which baselines — and date it. Report means with standard deviations across multiple seeds, correct for multiple comparisons, and include the analyses that didn't work. Reviewers reward this transparency; it is the difference between a result and a rumor.

The mindset shift. The deepest defense against p-hacking isn't any single technique — it's changing what feels like success. If success means "a significant p-value," you will unconsciously steer toward it. If success means "an honest answer, whatever it is," the temptation evaporates. Train yourself to feel good about a well-executed experiment with a null result and uneasy about a sloppy experiment with a flashy one. Supervisors and collaborators take their cues from you: when you present negative results proudly and with clear analysis, you give everyone around you permission to be honest too. Science advances on trustworthy nulls as much as on discoveries — sometimes more.

Putting it together. p-hacking is dangerous precisely because it rarely feels like misconduct — it feels like thorough exploration. The defense is structural: separate free exploration on validation data from strict confirmation under a preregistered protocol, correct for multiple comparisons, report means with uncertainty across seeds instead of maxima, and publish the analyses that failed alongside those that worked. These practices don't just protect you from false positives; they change how you think, replacing "what can I claim?" with "what does the evidence support?" Reviewers increasingly check for exactly these signals, and their presence marks the difference between a result and a rumor. In the long run, the researchers whose findings replicate are the ones who constrained their own degrees of freedom [1], [4].

Key takeaways: - Reporting only the best of many tries manufactures false positives — even unintentionally. - Separate exploration (validation) from confirmation (pre-committed final evaluation). - Pre-register your analysis plan; decide stopping rules and tests before seeing results. - Report all tried analyses, including failures; correct for multiple comparisons. - Report effect sizes with confidence intervals, not p-values alone.


Chapter 8: Reproducibility Failures — Seeds, Code, and Missing Details

A result that cannot be reproduced is a rumor, not a finding. Reproducibility means that another researcher — or you, six months later — can rerun your work and get the same conclusions. Beginners underestimate how quickly their own work becomes unreproducible: the undocumented hyperparameter, the random seed nobody wrote down, the dataset version that changed. This chapter makes reproducibility a habit [5], [8].

The three levels of reproducibility. First, computational reproducibility: the same code and data produce the same numbers — this requires fixed random seeds, recorded library versions, and archived code. Second, method reproducibility: an independent researcher can reimplement your method from the paper and get similar results — this requires complete descriptions of architectures, hyperparameters, preprocessing, and splits. Third, conclusion reproducibility: the claim holds up under reasonable variations — different seeds, slightly different data. Most beginner failures are at levels one and two, which are entirely preventable [6].

Seeds and randomness. Modern ML is full of randomness: weight initialization, data shuffling, dropout, augmentation. Without fixed seeds, every run differs, and "my method beats the baseline" might mean "my method got luckier this time." Set seeds for every random source (Python, NumPy, your framework), record them, and — crucially — report results across multiple seeds with means and standard deviations rather than a single lucky run. Note that some GPU operations are nondeterministic even with seeds; document this and average over runs rather than chasing bit-identical numbers [7].

What to document and share. Your paper and repository should let a stranger rerun everything: the exact dataset version and how to obtain it, the train/validation/test split procedure (or the split files themselves), preprocessing steps in order, model architecture details, all hyperparameters including the ones you tuned and their search ranges, the compute environment (library versions, GPU type), and the random seeds. Share code in a public repository with a README that runs end-to-end, and archive a snapshot (e.g., with a DOI) so it survives link rot. "Code available on request" is widely understood to mean the code does not exist — share it proactively [5], [8].

Nondeterminism across hardware. Even with fixed seeds, results can vary across machines. GPU operations such as certain reductions and convolutions run in nondeterministic order for speed, so the same code on two different GPUs can produce slightly different numbers. Frameworks offer deterministic-mode flags that trade some speed for repeatability — use them when exact reproducibility matters, and otherwise accept small numerical differences while ensuring conclusions are stable across runs. Document the hardware and library versions (a requirements.txt or environment file plus the GPU model), because "it worked on my machine" is not a reproducibility strategy. When reporting, the mean and standard deviation across seeds already absorbs this hardware noise, which is another reason multi-seed reporting is the professional standard [7].

Versioning data and code together. Code without the exact data is half the story, and datasets change — files get corrected, labels get fixed, versions drift. Record a checksum or version identifier for every dataset you use, and keep the split indices so the exact train/test partition can be rebuilt. Version your code with git tags marking which commit produced which paper table. For long-term archiving, deposit a snapshot of code and (when licensing allows) data in a permanent archive that issues a DOI, so the artifacts survive beyond your laptop and your lab's server. Containers or environment files freeze the software stack. None of this is glamorous work, but it is what separates a demo from a scientific result — and it is what lets your future self rebuild a two-year-old project in an afternoon instead of a month [5], [8].

The README that reproduces: a template. A reproduction-friendly README has a fixed anatomy that reviewers and future collaborators recognize instantly. Start with a one-paragraph description of what the code reproduces (which paper, which tables). Then: requirements and setup (exact install commands, tested environment); data (where to download it, expected checksums, how the provided script recreates your splits); quick start (the single command that reproduces the main result, and how long it takes); configuration (where hyperparameters live and which ones you tuned); and a results section showing the expected output numbers so the reproducer can verify success. End with contact information and the license. Write the README as you build, not after submission — a README written from memory weeks later always misses the step that "everybody knows." Test it by having a colleague (or your future self, on a fresh machine) follow it blind [5], [8].

What reproducibility checklists ask for. Major venues now publish reproducibility checklists, and reading one is the fastest way to learn the professional standard. Typical items: a clear description of the data splits and how they were created; the full hyperparameter configuration including search ranges; the compute infrastructure and average runtime; the number of training runs with means and variances; links to code and data (or an explanation when they cannot be shared); and statements of any assumptions or simplifications. Some venues run artifact evaluation, where independent reviewers actually execute your code. You do not need to wait for submission to benefit: run through such a checklist on your own project midway, while there is still time to fix the gaps. Every unchecked box is a reviewer objection waiting to happen — and each one you check is a paragraph of your methods section writing itself [5], [8].

The two-hour reproducibility drill. Schedule this exercise once per project: set a timer for two hours and pretend you are a stranger trying to reproduce your main result from your repository alone. Clone it fresh, follow the README exactly, and note every place you get stuck — the missing dependency, the undocumented download step, the hardcoded path, the "obvious" preprocessing detail that exists only in your head. Fix each blocker as you find it. Most researchers discover three to five blockers in the first drill, which is precisely why the drill exists: every blocker you fix now is a reviewer objection or a failed replication you prevented. Repeat the drill before each submission. Two hours of deliberate awkwardness buys you the confidence that your work stands on its own [5], [8].

Worked example — the mistake and the fix. A student submits a paper with strong results but no code and a methods section that says "we used a CNN with standard hyperparameters." The mistake: reviewers cannot verify anything — what CNN? Which hyperparameters? Which data split? The paper is rejected with "insufficient detail to reproduce." The fix: the student rebuilds the project properly. They create a repository with a README, a requirements file pinning library versions, scripts that download the public dataset and recreate the exact splits (with the split indices saved), a configuration file listing every hyperparameter, fixed seeds, and a single command that reproduces the main table. The methods section now specifies the architecture layer by layer, the optimizer settings, and the tuning ranges. Resubmitted, the paper passes review — and the student's own future self can rerun the experiment in one command a year later [5], [8].

For your research: From your next experiment onward, keep a lab log: date, code version, data version, config, seeds, and results for every run. Before submitting any paper, do the "stranger test" — could a stranger reproduce your main table from your paper plus your repository? If not, you are not done writing.

Reproducibility as a career asset. Beyond papers, reproducibility builds something personal: a portfolio of projects you can actually rerun, extend, and show. When a future employer or collaborator asks what you've built, "here's the repository — one command reproduces the results" is enormously more convincing than a PDF alone. Your old projects become building blocks instead of archaeological sites. Many researchers report that their most-cited work is the paper with the best-maintained code, because usable artifacts attract follow-up work. Invest in reproducibility and you invest in a body of work that keeps paying dividends long after each deadline passes.

Putting it together. Reproducibility is what turns your private experiment into public knowledge, and it is built from unglamorous parts: fixed seeds, multi-seed reporting, versioned data and code, documented environments, complete methods sections, runnable READMEs, and the two-hour drill that finds the gaps. None of it is difficult; all of it is easy to postpone — which is why you should build it into your workflow from the first experiment rather than reconstructing it before submission. The payoff is personal as well as scientific: the researcher with complete records reuses old projects effortlessly, answers reviewer questions in minutes, and never loses a result to a dead laptop. Treat reproducibility as part of doing the experiment, not as paperwork after it [5], [7], [8].

Key takeaways: - Unreproducible results are rumors; aim for all three levels of reproducibility. - Fix and record random seeds; report means and stds across multiple seeds. - Document everything: data version, splits, preprocessing, architecture, hyperparameters, environment. - Share code proactively with a runnable README — "on request" means it doesn't exist. - Keep a dated lab log; apply the "stranger test" before every submission.


Chapter 9: Misreading the Literature — Strawmen and Citation Errors

Your paper's credibility rests on how fairly you treat prior work. Beginners misread the literature in predictable ways: citing papers they haven't read, misrepresenting what prior methods do (the strawman), copying citations from other papers without checking, and overclaiming novelty over work they missed. Reviewers are often the authors of the misrepresented work — this mistake is personal to them, and they punish it [6].

The strawman. A strawman is a weakened version of prior work that is easy to beat: comparing against an old version of a method, using its default settings while tuning yours (see Chapter 5), or describing it inaccurately ("prior methods cannot handle X" when they can). The defense is simple but requires work: read the primary source, run the actual method yourself under fair conditions, and describe it the way its authors would. If you're unsure, email the authors — most researchers are happy to clarify, and the contact often improves your paper [6].

Citation errors. The most common: citing a secondary source for a claim that needs the primary one; citing a paper for a result it doesn't contain (because you copied the citation from another paper's bibliography); citing preprints as established results without noting their status; and the "citation ring" of citing only your group's work. Another frequent error is claiming "to the best of our knowledge, we are the first to..." without a thorough search — reviewers love finding the prior work you missed. Search broadly (multiple databases, preprint servers, related fields with different terminology) before claiming novelty [6].

Reading papers critically. For each key related paper, extract: the exact problem it solves, its method in your own words, the datasets and metrics it used, its stated limitations, and what it does not claim. Keep these in a structured literature table. This does double duty: it prevents misrepresentation and it reveals genuine gaps — the difference between "nobody has done X" (a claim you must verify) and "method A works on images but was never tested on our sensor data" (a defensible gap) [2], [6].

How to search the literature thoroughly. Missing prior work is the most embarrassing citation error, and it is preventable with a systematic search. Use multiple databases — general scholarly search plus preprint servers — because coverage differs. Search with several terminology variants: different subfields name the same idea differently ("domain adaptation" vs "transfer learning" vs "covariate shift"), and a novelty claim that holds under one term may collapse under another. Chain citations in both directions: backward through a key paper's references, and forward through the papers that cite it, which surfaces newer work the original authors could not have known. Set up alerts on your core keywords so new preprints reach you during the project, not after submission. And when you think the search is complete, ask your supervisor or a colleague to try to break your novelty claim — a ten-minute conversation can save a rejection [6].

Citing preprints, negative results, and your own work. Preprints are legitimate citations in fast-moving fields, but mark their status honestly — "a recent preprint reports..." — since they have not been peer-reviewed and their claims may change. Cite negative results and failed approaches when they shaped your thinking; a paper that acknowledges what did not work reads as mature, and it saves other researchers from repeating the dead end. With your own prior work, cite it exactly as you would anyone else's — no more, no less. A related-work section that cites only your own group looks like a sales brochure; reviewers notice the proportion. Finally, keep a citation log as you read: for each paper, one line on what you are citing it for. When you later write "as shown in [4]," you will know it actually shows that [6].

Writing the related-work section. A strong related-work section is organized by ideas, not by papers. Group prior work into two or three approaches or schools of thought, summarize what each approach assumes and achieves, and note their limitations — this structure naturally reveals the gap your work fills. For each group, cite the most representative and most-cited papers, plus the most recent ones, so coverage looks current. Then position your work explicitly: "Unlike approach A, which requires X, our method..." or "We extend approach B to the setting of Y." This positioning paragraph is where many beginners accidentally overclaim, so keep it factual and check every contrast against your literature table (Chapter 9's exercise). End by stating your contributions as a short list, each one verifiable in the paper's experiments. A related-work section written this way does double duty: it proves you understand the field, and it makes your novelty claim checkable [6].

Synthesizing, not summarizing: the literature review as argument. A weak literature review lists papers: "A did X. B did Y. C did Z." A strong one synthesizes: it identifies the two or three core ideas the field keeps rediscovering, points out where results contradict each other and why (different datasets? different metrics? different assumptions?), and builds an argument that ends at your research gap. Contradictions are gold — when paper A claims method M works and paper B claims it fails, the careful reader asks what differed, and that question is often a thesis topic. Write your review as the story of how the field's understanding evolved and where it got stuck; your work then appears as the natural next step rather than a random addition. This is also where honest scholarship pays off visibly: a review that fairly presents work contradicting your approach is far more persuasive than one that pretends it doesn't exist [2], [6].

Emailing authors: how and when. Contacting the authors of prior work is underused by beginners and welcomed by most researchers. Write when: you can't reproduce their result and want to check a detail, you're unsure whether their method handles your setting, or you want to confirm your description of their work is fair. Keep it short and specific: state who you are, cite the exact paper and section, ask one or two concrete questions, and mention what you've already tried. Most authors reply within days — and the exchange sometimes improves your paper beyond the original question, occasionally into a collaboration. Always thank them and cite any private clarification appropriately. The only bad reasons to email are asking for their full codebase to avoid doing the work, or disputing their results aggressively. Used well, this habit prevents strawmen and builds your professional network simultaneously [6].

Worked example — the mistake and the fix. A student's related-work section claims: "Existing methods [3, 7] fail on small datasets, motivating our approach." The mistakes: the student never ran those methods on small datasets, paper [3] actually includes a small-data experiment the student missed, and [7] is cited for a claim it never makes. A reviewer who authored [3] flags all of it. The fix: the student reads both papers fully, runs method [3] on their small dataset under fair conditions (it scores respectably but worse than the new method), and rewrites: "Method [3] reports results on datasets of 50k+ images; on our 2,000-image dataset it reaches 81% versus our 86% (Section 4), suggesting its augmentation strategy is less effective at this scale." The strawman becomes a fair comparison, the false citation is corrected, and the novelty claim is now specific and defensible [6].

For your research: Maintain a literature table with one row per key paper: problem, method (in your words), data, metrics, limitations, and what it doesn't claim. Never cite a paper you haven't at least skimmed, never copy citations secondhand, and verify every "first to" claim with a broad search before writing it.

The scholar's reputation. Here's what fair citation buys you over a career: authors you treated well become reviewers who trust you, collaborators who recommend you, and colleagues who cite you. Scholarship is a repeated game, and the researchers who play it fairly accumulate allies while those who cut corners accumulate enemies. Every careful literature review is also an advertisement for your judgment — it tells the field you can be trusted with its ideas. In a small field like yours, that reputation arrives faster than you think, and it lasts longer than any single paper.

Putting it together. Fair scholarship is a habit of verification: read the primary sources, search the literature systematically across terminologies and databases, maintain a structured literature table, describe competing methods as their authors would, verify every citation supports its claim, and position your work through synthesis rather than strawmen. The practical tools — citation chaining, keyword alerts, emailing authors, the related-work-as-argument structure — turn a dreaded chore into one of the most intellectually rewarding parts of research: genuinely understanding where your work sits in the field's story. Reviewers, who are often the authors you cite, notice fair treatment immediately and reward it with trust. Cite like someone who expects to meet every cited author at a conference — because eventually, you will [2], [6].

Key takeaways: - Never misrepresent prior work — reviewers are often its authors. - Run competing methods yourself under fair conditions; describe them as their authors would. - Cite primary sources; verify every citation actually supports your claim. - Search broadly before claiming novelty; "first to" is a high bar. - A structured literature table prevents strawmen and reveals genuine gaps.


Chapter 10: Compute Waste and Poor Experiment Design

Compute is a research resource — GPU hours cost money, energy, and time — and beginners often spend it badly: giant hyperparameter sweeps before the idea is validated, training huge models when small ones would answer the question, and rerunning everything from scratch after each small change. Poor experiment design wastes compute; good design gets more evidence per GPU hour [7].

The pilot-first principle. Before any large run, validate the idea cheaply: a tiny model on a data subset, a few epochs to check that loss decreases, a single seed to confirm the pipeline works end-to-end. Most ideas fail at the pilot stage, and a failed pilot costs 1% of a full run. Only scale up what survives. Similarly, debug on CPU or a small GPU with tiny data before touching the big machine — an indexing bug discovered after 40 GPU hours is an expensive lesson [7].

Designing informative experiments. Every experiment should answer one question. Change one thing at a time so the result is interpretable — if you change the architecture, the data, and the optimizer at once, a win tells you nothing. Use ablations (removing one component at a time) to show which parts of your method matter; reviewers expect them, and they often reveal that the fancy component contributes nothing — which is worth knowing before you publish. Plan a compute budget per research question in advance, and track spending so one question doesn't eat the whole project [1], [5].

Avoiding the common wastes. Don't grid-search blindly over huge ranges — start coarse, then refine around promising regions, or use random search, which is usually more efficient. Don't retrain from scratch when fine-tuning or warm-starting would do. Cache preprocessed data and features instead of recomputing them every run. And don't chase the last 0.5% with massive sweeps when the paper's contribution is the idea, not the leaderboard position — diminishing returns are real, and your time is also a resource [5], [7].

Green AI: the environmental cost. Training large models consumes real energy — a single big training run can emit as much carbon as several cars do in a year. As a beginner you will rarely train at that scale, but the principle applies at every scale: wasted compute is wasted energy as well as wasted time and money. Report the compute your results required (GPU type and hours) so readers can judge the cost of reproducing your work — many venues now encourage or require this. Treat efficiency as a first-class result: a method that matches the state of the art with one-tenth the training cost is a genuine contribution, sometimes a more important one than a small accuracy gain. And prefer shared or institutional clusters' off-peak scheduling when available; small choices about when and how you compute add up across a research community [7].

Tracking experiments like a lab notebook. You cannot design good experiments if you cannot remember what you ran. Keep a structured log of every experiment: date, code version, config, data version, seeds, compute used, and results — whether you use a dedicated tracking tool or a disciplined spreadsheet matters less than doing it consistently. Name runs descriptively ("resnet50_lr0.01_aug-v2") rather than "run_47_final_FINAL." Record negative results with the same care as positive ones; the experiment that failed teaches you what not to try and, months later, stops you from repeating it. Before any meeting with your supervisor, you should be able to answer "what did you try, what happened, and what will you try next" from the log alone. Good tracking turns a pile of GPU hours into an argument — which is what a paper is [5].

When to stop tuning: knowing "good enough." Tuning has no natural end, so you must impose one. Set a stopping rule before you start: a fixed budget ("20 configurations, then we ship the best"), a target tied to the research question ("within 1% of the baseline is enough to test our hypothesis"), or a plateau rule ("stop when three consecutive rounds improve validation by less than 0.2%"). Remember the opportunity cost: every week spent squeezing 0.3% from a model is a week not spent on the next research question, the ablation that explains why it works, or the writing that communicates it. Ask what the paper needs: if the contribution is the idea, a clean comparison at reasonable tuning beats a leaderboard number. And beware the sunk-cost trap — "we've already spent 300 GPU hours" is a reason to stop, not to continue. Deciding in advance what "done" looks like is what separates systematic researchers from tinkerers [5], [7].

Shared clusters and being a good citizen. Most students train on shared lab or university clusters, where your experiment design affects other people directly. Learn the scheduler basics: request only the GPUs and memory you need, set realistic time limits so jobs don't idle in queues, and use checkpointing so a preempted job resumes instead of restarting. Never run interactive debugging sessions on shared GPUs — debug on a small local setup first. Clean up: kill zombie jobs, delete stale checkpoints filling shared storage, and document your environment so others can reproduce your setup. Communicate big runs to the group in advance ("I'll need 4 GPUs this weekend for the final sweep") rather than silently monopolizing the queue before a deadline. Labs remember who hogs resources; supervisors notice who doesn't. Good cluster citizenship is also good experiment design, because constraints force the piloting and budgeting discipline this chapter recommends [7].

Estimating compute before you run. Before launching any sweep, do back-of-envelope math: time one training run on your setup, multiply by the number of configurations and seeds, and convert to days and GPU-hours. Write the estimate next to your plan — "6 hrs × 24 configs = 144 GPU-hours ≈ 6 days on 1 GPU" — and ask whether the expected information is worth it. This single habit kills most compute waste, because the shocking totals force prioritization: pilots first, coarse-to-fine, fewer seeds for exploration and more for the finalists. Add a 50% buffer for failures and restarts; estimates are always optimistic. And record actual vs estimated afterward — after a few projects your estimates become accurate, which makes you the rare student whose project plans supervisors actually believe [7].

Worked example — the mistake and the fix. A student plans to compare 4 architectures × 6 learning rates × 5 seeds = 120 full training runs, each 6 GPU hours — 720 GPU hours, three weeks on their lab's shared GPU, blocking everyone else. The mistake: no piloting, full factorial design, 5 seeds for every configuration. The fix: first, a pilot — each architecture trained for 2 epochs on 10% of the data (2 GPU hours total) reveals that one architecture diverges and another underfits badly; both are dropped. Second, coarse-to-fine: 2 surviving architectures × 3 learning rates × 1 seed on full data (36 GPU hours) identifies the promising region. Third, only the top 4 configurations get 5 seeds each for the final comparison (120 GPU hours). Total: ~160 GPU hours instead of 720, results in days instead of weeks, and the staged design produces a cleaner story for the paper — pilot findings, coarse search, final validation [7].

For your research: Adopt a written experiment plan for each research question: the pilot, the coarse search, the final validation runs, and a compute budget with a hard cap. Log GPU hours per experiment alongside results — when you can show reviewers you were economical and systematic, it strengthens the paper's credibility.

Efficiency as a research taste. The best researchers develop a feel for which experiments are worth their cost — a kind of taste that this chapter's practices train. They can smell a fishing expedition (huge sweeps with no hypothesis), they kill unpromising directions after a cheap pilot instead of after a month, and they spend their saved compute on the experiments that actually discriminate between ideas. This taste is visible in papers: crisp experiment sections where every table answers a question, versus sprawling appendices of barely-digested runs. Cultivate it deliberately by always asking, before launching anything big: "what will I conclude if this works, and what will I conclude if it doesn't?" If you can't answer both, you're not ready to spend the compute.

Putting it together. Compute is money, energy, time, and other people's queue slots — and experiment design is how you spend it wisely. The full practice: estimate costs before running, pilot cheaply, search coarse-to-fine, stop at a pre-defined "good enough," track every experiment like a lab notebook, treat efficiency as a result worth reporting, and be a good citizen on shared infrastructure. Notice how these habits reinforce each other: pilots make estimates accurate, budgets force prioritization, tracking turns runs into arguments, and constraints breed the creativity that massive sweeps never will. The researcher who gets the most insight per GPU hour isn't the one with the biggest cluster — it's the one with the best-designed experiments [5], [7].

Key takeaways: - Pilot cheaply first — most ideas fail fast, and pilots cost almost nothing. - One experiment, one question; change one variable at a time. - Use ablations to prove each component matters; expect reviewers to ask for them. - Search coarse-to-fine; don't burn compute chasing the last fraction of a percent. - Budget compute per question and track it — your time is a resource too.


AI research increasingly touches real people — their faces, medical records, messages, and behavior. Ethical mistakes here are not just paper rejections; they can harm people and end careers. Beginners often assume "the data was public, so it's fine" or "ethics is someone else's job." Neither is true [6].

Privacy and consent. "Publicly available" does not mean "consented for research." Scraping photos, social media posts, or forum messages and publishing a dataset of them can violate privacy expectations and, in many jurisdictions, data protection law — even if the scraping itself was technically possible. Medical, biometric, and location data need explicit informed consent and usually institutional ethics approval before you collect or use them. De-identification is harder than it looks: "anonymized" datasets have been re-identified repeatedly by linking with public data. Rule: if your data is about people, get ethics approval first, document consent, and treat re-identification risk as your problem [6].

Bias and harm in deployment-facing claims. If your paper claims a model is "ready for" hiring, policing, lending, or medical decisions, you take on responsibility for its failures across groups. Beginners often evaluate on a single population and ignore performance disparities — a model that works well on average but fails for a minority group is a harmful model wearing a good average. Evaluate and report performance across relevant subgroups, discuss who could be harmed by errors, and avoid claims of deployment-readiness you cannot support [4], [6].

Dual use. Some research can be used to help and to harm — a face generator enables both creative tools and deepfakes; a vulnerability detector helps both defenders and attackers. You cannot always prevent misuse, but you must think about it: most venues now expect an honest discussion of potential negative uses and your mitigation choices (what you release, what you withhold, what safeguards you considered). Ignoring dual use doesn't make it go away; it makes reviewers conclude you didn't think [6], [7].

What an ethics review actually asks. Many students fear the institutional ethics board as a bureaucratic obstacle; in practice it is a structured conversation about exactly the questions in this chapter. A typical application asks: what data are you collecting, from whom, with what consent, how will it be stored and who can access it, can participants withdraw, and what are the risks to them? For standard ML projects using licensed or public-consented datasets, approval is often straightforward and fast — the board's main job is catching the risky cases before they happen. Start the process early, because approval can take weeks and you cannot collect data while waiting. Keep the approval letter with your project files; some venues ask for the approval reference number at submission. If your institution has no formal board, write your own ethics memo answering the same questions and have your supervisor sign off — the discipline matters more than the paperwork [6].

Auditing your model for bias. "We evaluated on the full test set" is not enough when errors concentrate in one group. A bias audit means disaggregating your metrics: report accuracy, precision, and recall separately for each relevant subgroup (by gender, age band, ethnicity, region, device type — whatever is relevant to your data). Look at error rates, not just averages: a model with 95% overall accuracy but 70% on one subgroup is a model that fails that subgroup. Examine the worst errors qualitatively — misclassified examples often reveal patterns (lighting conditions, dialects, rare names) that aggregate numbers hide. If you find disparities, report them honestly and investigate causes: unbalanced training data, label bias, or features that proxy for group membership. Fixing bias is a research contribution in its own right, and reporting it is always better than a reviewer discovering it [4], [6].

Writing the ethics and limitations statement. Most venues now expect an explicit statement, and writing it well is a skill. A good statement has four parts: data provenance (where the data came from, what consent or license covers it, ethics approval reference if applicable); subgroup evaluation (which groups you checked and what you found, including disparities); potential misuse (who could be harmed if the work is misapplied, stated plainly); and mitigations (what you are and are not releasing, and why). Keep it concrete — "our dataset was collected with written consent for research use (approval #2024-118); accuracy was 91% overall, 84% on the smallest subgroup; the model could be repurposed for surveillance, so we release evaluation code but not the trained face embeddings." Vague statements ("we considered ethics carefully") signal the opposite. Write this section early: if you cannot fill in the four parts, your project has an ethics gap to fix before submission, not after [6].

Data licensing and terms of service. Beyond consent lies the legal layer beginners often ignore: datasets have licenses, and platforms have terms of service. A dataset labeled "for research use" may forbid commercial application or redistribution — training a startup's product on it, or re-uploading it publicly, can violate the license. Scraping a website against its terms of service can breach contract law even when the data is technically accessible, and some jurisdictions restrict automated collection specifically. Before using any dataset, read its license file (usually a few paragraphs) and record three things: can you use it for your purpose, can you redistribute it or models trained on it, and what attribution is required? When the license is unclear or missing, treat the data as restricted and ask — the dataset authors, your supervisor, or your institution's legal office. "Everyone else uses it" is not a license, and reviewers increasingly ask about data provenance [6].

Sensitive data handling basics. When your project does involve sensitive data — with proper consent and approval — handling discipline matters. Store it encrypted, on institutional systems rather than personal laptops, with access limited to named project members. Never include raw sensitive records in shared repositories, paper appendices, or presentation slides; use aggregated statistics or synthetic examples for illustration. Plan for the project's end: how long will you retain the data, and how will you securely delete it? Document these decisions in your ethics application. And remember the re-identification lesson from earlier: removing names is not anonymization — combinations of quasi-identifiers (age, zip code, dates) can single out individuals, so treat any "anonymized" personal dataset as still sensitive until a proper disclosure analysis says otherwise [6].

Worked example — the mistake and the fix. A student scrapes 50,000 profile photos from a social network to build an "emotion recognition" dataset and plans to publish it. The mistakes: no consent from the photographed people, no ethics approval, faces are inherently identifiable so "anonymization" is impossible, and emotion recognition has documented bias and dual-use concerns (workplace surveillance). The fix: the student stops the scrape, consults their institution's ethics board, and pivots to a licensed dataset collected with explicit consent for research use. The paper now includes an ethics statement describing the data's consent basis, reports accuracy across demographic subgroups, discusses surveillance misuse as a limitation, and releases only the trained evaluation code — not a new scrape of faces. The project takes longer, but it is publishable and defensible; the original plan risked both rejection and real harm [6].

For your research: Before collecting any data about people, answer three questions in writing: (1) Do I have consent or a licensed basis for this data? (2) Have I obtained ethics approval if my institution requires it? (3) Who could be harmed if this works as claimed, and what am I doing about it? Put the answers in your paper's ethics statement — reviewers and readers increasingly expect it.

Ethics as a differentiator. As AI ethics expectations rise across venues, the researchers who already practice them have an edge: their papers pass ethics review smoothly, their datasets get reused because they're properly licensed, and their work is deployable in real institutions that require compliance. Meanwhile, projects built on scraped data and vague consent increasingly hit walls at submission or, worse, after publication. Treat ethics as part of research quality, not as overhead — the careful researcher this book describes is careful here too, and the field is moving decisively in that direction.

Putting it together. Ethics in AI research is not a separate chapter of your project — it is woven through every decision about data and deployment. Consent before collection, approval before experiments, licenses checked before use, sensitive data handled with discipline, subgroup performance audited, dual use considered honestly, and all of it written into a concrete ethics statement. Students sometimes experience this as friction slowing down exciting work; experienced researchers experience it as the thing that lets them sleep at night and defend their work anywhere. The projects you're proudest of years from now will be the ones that were both technically strong and ethically clean — and the habits that produce them are the ones this chapter just gave you [6].

Key takeaways: - Public data ≠ consented data; get ethics approval before collecting data about people. - "Anonymized" data is often re-identifiable — treat that risk as your responsibility. - Report performance across subgroups; don't hide disparities behind averages. - Consider dual use honestly and discuss mitigations in the paper. - An ethics statement is now expected — write it before you need it, not after.


Chapter 12: The Careful Researcher's Checklist + Surviving Peer Review

Everything in this book converges on one practice: checking your work systematically before anyone else does. This final chapter gives you a pre-experiment checklist, a pre-submission checklist, and a guide to surviving peer review — because even careful researchers get critical reviews, and handling them well is a skill [5], [6].

The pre-experiment checklist. Before running the decisive experiments, confirm: the data split is done and the test set is locked away; preprocessing is inside a pipeline fitted on training data only (Chapter 2); every feature would be known at prediction time; baselines are chosen and will get equal tuning budgets (Chapter 5); the metric matches the research question and includes uncertainty reporting (Chapter 4); the analysis plan — metric, seeds, tests — is written and dated (Chapter 7); seeds are fixed and multiple runs are planned (Chapter 8); compute budget is set with a pilot first (Chapter 10); and ethics approval and consent are in place if the data involves people (Chapter 11) [1], [5].

The pre-submission checklist. Before submitting, verify: the test set was evaluated once, after all decisions; every claim in the abstract and conclusion points to a specific table or figure (Chapter 6); all baselines were run fairly and reported fully; limitations are stated explicitly; every citation was actually read and supports its claim (Chapter 9); code and data documentation pass the "stranger test" (Chapter 8); the ethics statement is written; and a claim audit removed or softened every unjustified superlative. Then put the paper away for two days and reread it as a hostile reviewer — you will find problems your tired eyes missed [6].

Surviving peer review. Rejection is normal — most papers are rejected at least once, including good ones. When reviews arrive, wait a day before responding; read them as free expert consulting, not personal attacks. Sort comments into three piles: (1) valid and fixable — do the work and say exactly what changed; (2) valid but out of scope — acknowledge and add as a limitation or future work; (3) mistaken — rebut politely with evidence, never with irritation, because the reviewer may be right in a way you haven't seen yet. Write a response letter that quotes each comment and answers it point by point. Reviewers remember gracious, thorough responses — the same reviewer often sees your resubmission [6].

Anatomy of a good rebuttal. The response letter is a genre with conventions, and following them helps. Open by thanking the reviewers — sincerely; they donated hours to improve your work. Then address every numbered comment in order, quoting or paraphrasing each one so the reviewer sees nothing was skipped. For each: state what you did (new experiment, rewritten paragraph, added limitation), point to exactly where in the revised manuscript it appears ("Section 4.2, Table 3"), and include the key new numbers directly in the letter so the reviewer does not have to hunt. When you disagree, concede what is true first ("the reviewer is right that our dataset is small"), then present your evidence calmly. Never be sarcastic, dismissive, or personal — the reviewer is anonymous but the area chair is not, and graciousness is remembered. Keep the letter self-contained: a tired reviewer should be able to judge your response without reopening the manuscript [6].

After acceptance: the record you leave. Publication is not the end of the careful researcher's job. Release the code and data artifacts you promised, and keep them working — a repository whose README rots within months undermines the paper. If you later find an error, publish a correction or an updated preprint promptly; the community respects authors who correct themselves far more than those caught hiding mistakes. Respond to questions about your work from other researchers — today's confused reader is tomorrow's citer. And archive everything: the final code snapshot, the datasets or their identifiers, the experiment logs. Years later, when someone builds on your work or a controversy arises, the researcher with complete records is the researcher who is trusted [5], [8].

Dealing with harsh or unfair reviews. Sometimes reviews are not just critical but harsh, dismissive, or mistaken — and it stings. First, separate the tone from the content: a rude reviewer can still have a valid point, and a polite one can be wrong. Extract every actionable criticism regardless of how it was phrased. For genuinely mistaken reviews, rebut with evidence and without heat; area chairs can spot reviewer error when you document it clearly. If a review is abusive rather than critical, most venues let you flag it confidentially to the chairs — use that channel instead of responding in kind. Know when to appeal (clear factual errors by reviewers, not mere disagreement) and when to move on: a rejection with useful feedback, revised and resubmitted elsewhere, beats a months-long appeal that rarely succeeds. And protect your morale deliberately — every established researcher has a drawer of rejections; the ones who last are those who learned to mine reviews for gold without internalizing the gravel [6].

Choosing where to submit. Careful work still fails at the wrong venue, so choose deliberately. Read the last two years of proceedings of your candidate venues: do they publish your kind of contribution (a careful empirical study fits some venues better than others)? Check the acceptance rates and the review criteria they publish — some venues explicitly reward rigor and negative results. Match the paper's maturity to the venue: early, exploratory work often fits workshops, where feedback is constructive and expectations are lighter; mature, complete stories fit conferences; long, definitive studies fit journals. Ask your supervisor where similar papers landed, and read those papers' reviews if they're public. A well-chosen venue means reviewers who understand your contribution — which is half the battle this chapter describes [6].

Your first submission timeline. Work backward from the deadline and protect each phase. Six weeks out: experiments frozen, main tables final. Four weeks out: full draft complete, figures polished. Three weeks out: the claim audit and both checklists done, paper rested for two days then reread as a hostile reviewer. Two weeks out: supervisor and at least one colleague have read it; code repository passes the two-hour reproducibility drill. One week out: response to internal feedback incorporated, camera-ready formatting checked against the venue template, supplementary material and code links verified live. The most common beginner failure is compressing writing and checking into the final 72 hours — exactly when fatigue makes you blind to the mistakes this book covers. A paper submitted a day early after a calm final read beats a paper submitted at 11:58 p.m. after a frantic one, every time [5], [6].

Worked example — the mistake and the fix. A student submits a paper and receives three reviews: "test set used for model selection," "no strong baseline," and "overclaimed SOTA." The mistake response would be to argue or to quietly resubmit elsewhere unchanged. The fix: the student treats the reviews as a checklist. They redo the experiments with a locked test set and validation-based selection (Chapter 3), add a tuned gradient-boosted baseline that beats their method on one dataset — reported honestly (Chapter 5) — and rewrite the claims to match the evidence (Chapter 6). The response letter addresses each point with the new tables. The paper is accepted on resubmission, and the student notes that the reviews made the paper genuinely better — the baseline that "embarrassingly" won became the paper's most-cited insight [5], [6].

Figure 2 — The careful researcher's checklist

For your research: Print the two checklists from this chapter and tape them where you work. Use the pre-experiment one before every decisive run and the pre-submission one before every deadline. When reviews come back critical, remember: a reviewer who found your mistakes before publication did you a favor — fixing them now is infinitely cheaper than retracting later.

One last thought. If this book had to fit on an index card, it would say: split before you preprocess, lock the test set, tune baselines fairly, measure what matters, preregister your plan, document everything, cite honestly, budget your compute, respect the people behind the data — and check it all twice. None of these steps is heroic; together they are what separates research that lasts from research that collapses. You now know every item on the card and, more importantly, why each one matters. Go do careful work. The field needs it, and your future self will thank you.

Putting it together. This final chapter distills the entire book into something you can actually use under pressure: two checklists, a review-handling method, and the knowledge that rejection is normal. The deeper message is that careful research is not a personality trait — it is a set of checkable procedures, and procedures can be learned, taped to walls, and followed when you're tired. Every chapter gave you one: the leakage audit, the locked test set, the metric derivation, the baseline protocol, the claim audit, the preregistration note, the reproducibility drill, the literature table, the compute budget, the ethics statement. Run them, and you join the researchers whose work others can trust and build on. That is the whole game — and you now know how to play it [5], [6].

Key takeaways: - Systematic checklists beat memory — use them before experiments and before submission. - Reread your paper as a hostile reviewer after a two-day break. - Treat reviews as free expert consulting; sort comments into fixable, out-of-scope, and mistaken. - Answer every reviewer point-by-point with evidence, never irritation. - Rejection is normal; a thorough, gracious revision usually wins on resubmission.


References

[1] T. Hastie, R. Tibshirani, and J. Friedman, The Elements of Statistical Learning, 2nd ed. New York, NY, USA: Springer, 2009. [2] C. M. Bishop, Pattern Recognition and Machine Learning. New York, NY, USA: Springer, 2006. [3] T. M. Mitchell, Machine Learning. New York, NY, USA: McGraw-Hill, 1997. [4] K. P. Murphy, Probabilistic Machine Learning: An Introduction. Cambridge, MA, USA: MIT Press, 2022. [5] A. Géron, Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow, 2nd ed. Sebastopol, CA, USA: O'Reilly Media, 2019. [6] S. Russell and P. Norvig, Artificial Intelligence: A Modern Approach, 4th ed. Hoboken, NJ, USA: Pearson, 2020. [7] I. Goodfellow, Y. Bengio, and A. Courville, Deep Learning. Cambridge, MA, USA: MIT Press, 2016. [8] F. Pedregosa et al., "Scikit-learn: Machine learning in Python," J. Mach. Learn. Res., vol. 12, pp. 2825–2830, 2011.

Glossary

  • Ablation — An experiment that removes one component of a method to measure its contribution.
  • Baseline — A simple or prior method your new method must beat for the comparison to be meaningful.
  • Calibration — How well a model's predicted probabilities match real observed frequencies.
  • Claim audit — A review pass checking that every comparative or superlative statement is backed by evidence.
  • Confusion matrix — A table showing predicted versus actual classes, revealing exactly where a classifier errs.
  • Cross-validation — Splitting data into folds and rotating which fold validates, to estimate performance robustly.
  • Data leakage — Test or future information accidentally influencing training, inflating reported performance.
  • Dual use — Research outputs that can be used for both beneficial and harmful purposes.
  • F1 score — The harmonic mean of precision and recall, summarizing both in one number.
  • Informed consent — Explicit permission from people whose data is used, given with understanding of the purpose.
  • Overclaiming — Stating conclusions stronger than the experiments actually support.
  • p-hacking — Trying many analyses and reporting only the favorable ones.
  • Pipeline — A chained sequence of preprocessing and modeling steps fitted only on training data.
  • Precision — Of all positive predictions, the fraction that were correct.
  • Pre-registration — Recording hypotheses and analysis plans before seeing results, to prevent p-hacking.
  • Recall — Of all actual positives, the fraction the model found.
  • Reproducibility — The ability of others (or your future self) to obtain the same results from your work.
  • Seed — A fixed starting value for random number generators, making runs repeatable.
  • State of the art (SOTA) — The best known performance on a standard benchmark; frequently abused as a claim.
  • Statistical significance — Evidence that an observed difference is unlikely due to chance alone.
  • Strawman — A weakened or misrepresented version of prior work, set up to be easily beaten.
  • Target leakage — A feature that contains information only available after the outcome being predicted.
  • Test set contamination — Using the test set during development, biasing the final performance estimate.
  • Validation set — Data used for tuning and model selection, kept separate from the final test set.

Practice Exercises

  1. A classmate reports 99% accuracy on a fraud dataset with 1% fraud. Explain in two sentences why this number is meaningless, and name the metrics they should report instead.
  2. Spot the mistake: A student normalizes all features to zero mean and unit variance, then splits into train and test. Name the mistake, explain why it matters, and describe the correct order of operations.
  3. Spot the mistake: A researcher tries 30 hyperparameter combinations, each evaluated on the test set, and reports the best test score as the final result. Identify the error and rewrite their workflow using train/validation/test splits.
  4. You want to claim your method is "state of the art" on a benchmark. List the four conditions a legitimate SOTA claim must satisfy.
  5. Spot the mistake: A paper's related-work section says "Prior method X cannot handle missing data [2]," but you check reference [2] and it describes a method for handling missing data. Name two distinct mistakes here and explain how to fix each.
  6. Design a fair baseline comparison for a new tabular-data classifier: name three baselines you would run and state what "equal tuning effort" means concretely.
  7. A student runs an experiment once with an unfixed random seed and beats the baseline by 1.5%. Write a short paragraph explaining why this is not yet a result, and what they should do next.
  8. Spot the mistake: A team scrapes 20,000 public profile photos to train a face-recognition model for a paper, with no consent process. List three ethical problems and describe a compliant alternative.
  9. You have 200 GPU hours for a project comparing 3 model architectures. Write a staged experiment plan (pilot, coarse search, final validation) with an hour budget for each stage.
  10. Take the pre-submission checklist from Chapter 12 and apply it to a draft paper of your own (or a published paper you admire): write down three checklist items it fails and how you would fix each.

End of Book 10 — the final book of the AstolixGen Learning Series (Detailed Edition).