
Book 3 of 50 · Free
Python for AI: From Zero to First Model
19,837 words · 25 chapters · illustrated

Book 3 of 50 · Free
19,837 words · 25 chapters · illustrated
Book 3 of 10 — AstolixGen Learning Series (Detailed Edition) · For researcher and publication students

Python is the language nearly all modern AI research is written in. Every machine learning paper you read — every model you train, every dataset you clean, every graph you put in your publication — runs on Python and its ecosystem of libraries. This book takes you from a fresh Python install to training and evaluating your first machine learning model, with the habits that make your code reproducible and publication-ready.
Unlike generic Python tutorials, this book is written for you: a researcher or publication student (MS/PhD, early researcher). Every chapter ends with practical advice on how the topic connects to your thesis, your experiments, and your papers. You will write real code in every chapter, and you will finish with a complete mini-project you can document as a research artifact.
Learning objectives: - Install and configure Python, pip, virtual environments, and Jupyter notebooks for AI work - Write clean, working Python code using types, control flow, functions, and comprehensions - Use NumPy arrays, broadcasting, and vectorization for fast numerical computing - Load, inspect, and clean real datasets with Pandas - Create exploration and publication-quality plots with Matplotlib and Seaborn - Train, evaluate, and tune a classifier end to end with scikit-learn - Debug and profile machine learning code efficiently - Make experiments reproducible with random seeds, environment management, and dependency pinning - Structure a project so it grows from a notebook to a reusable codebase - Document code and results to the standard a paper or thesis expects
| Concept | Definition (one line) | Example | Use in research |
|---|---|---|---|
| Python | General-purpose, readable programming language used across AI research | model.fit(X_train, y_train) |
Writing every experiment, script, and tool |
| pip | Python's package installer for adding libraries | pip install numpy |
Installing research libraries (scikit-learn, pandas) |
| Virtual environment | An isolated Python setup so each project has its own packages | python -m venv ai-env |
Keeping experiments reproducible and conflict-free |
| Jupyter notebook | Interactive document mixing code, output, and text | notebook.ipynb cell running analysis |
Exploring data and drafting experiment narratives |
| Variable / type | Named storage with a data kind (int, float, str, list) | accuracy = 0.94 |
Storing metrics, parameters, results |
| List | Ordered, changeable collection of items | scores = [0.91, 0.93, 0.89] |
Holding fold scores, predictions |
| Dictionary | Key-value mapping for named lookups | {"model": "SVM", "acc": 0.93} |
Configs, results logs, metadata |
| Function | Reusable named block of code with inputs and outputs | def train_model(X, y): |
Repeating experiment steps reliably |
| List comprehension | One-line expression building a list | [x**2 for x in range(10)] |
Clean, fast data transformations |
| NumPy array (ndarray) | Fast, fixed-type N-dimensional array for math | np.array([[1, 2], [3, 4]]) |
Storing features, images, weights |
| Broadcasting | NumPy rule that aligns different-shaped arrays in math | arr + 5 adds 5 to every element |
Normalizing data without loops |
| Vectorization | Replacing Python loops with array operations | np.mean(arr, axis=0) vs a loop |
10–100x faster experiment code |
| Pandas DataFrame | Table with labeled rows and columns | pd.read_csv("data.csv") |
Loading and cleaning real datasets |
| Pandas Series | One labeled column of a DataFrame | df["label"] |
Feature/target extraction |
| Missing values | Gaps in data (NaN) that must be handled | df.fillna(df.mean()) |
Cleaning datasets before training |
| Matplotlib | Core plotting library for Python | plt.plot(x, y) |
Line plots, figure control |
| Seaborn | High-level statistical plots on top of Matplotlib | sns.heatmap(cm) |
Confusion matrices, distributions |
| Scikit-learn | The standard ML library: models, metrics, pipelines | LogisticRegression().fit(X, y) |
Training and evaluating classifiers |
| Feature | Input variable the model learns from | df[["age", "bp"]] |
Model inputs in every experiment |
| Label (target) | The answer the model must predict | df["disease"] |
Supervised learning ground truth |
| Train/test split | Dividing data into learning and held-out evaluation sets | train_test_split(X, y, test_size=0.2) |
Honest, credible model evaluation |
| Classifier | Model that predicts a category label | SVC(), RandomForestClassifier() |
Baseline models for papers |
| Cross-validation | Rotating train/test splits to estimate performance robustly | cross_val_score(model, X, y, cv=5) |
Reliable results for publication tables |
| Hyperparameter | Setting chosen by you, not learned by the model | max_depth=5 |
Tuning reported in methodology sections |
| Overfitting | Model memorizes training data, fails on new data | Train acc 99%, test acc 61% | The main failure mode to diagnose |
| Random seed | Fixed starting point for randomness | random_state=42 |
Reproducible splits and training |
| Reproducibility | Anyone can rerun your code and get the same result | Seed + pinned versions + documented env | Required standard for published research |
| Dependency pinning | Recording exact library versions used | numpy==1.26.4 in requirements.txt |
Rerunning code months later |
| Debugging | Finding and fixing why code behaves wrongly | Traceback reading, prints, breakpoints | Fixing shape mismatches, NaNs, crashes |
| Profiling | Measuring where code spends time | %timeit, cProfile |
Speeding up slow experiments |
Roadmap — how the chapters connect: Chapter 1 explains why Python is the right bet, then Chapters 2–3 give you a working setup and the language basics. Chapters 4–6 build your three daily tools — NumPy (numbers), Pandas (tables), Matplotlib/Seaborn (plots) — the same stack almost every published ML experiment uses [1], [2]. Chapter 7 is the payoff: your first trained model with scikit-learn, with Chapters 8–9 covering the real-world skills of dataset handling, debugging, and profiling. Chapters 10–11 teach you the professional habits (reproducibility, project structure) that separate a student script from publishable research code. Chapter 12 ties it all together in a complete mini-project documented like a paper artifact.

If you look at the code behind any recent machine learning paper, it is almost certainly Python. This is not an accident, and it is not just fashion. Python won AI research for concrete, practical reasons that matter directly to you as a student: it makes experiments fast to write, fast to run, and fast to share.
1. Readability speeds up research. Research code is read far more often than it is written — by you next month, by your supervisor, by reviewers. Python reads close to plain English. Compare what it takes to compute the average of a column: in Python, np.mean(scores). The clarity means fewer bugs and faster iteration. When you are testing ten ideas a week, the language that gets out of your way wins.
2. The library ecosystem does the heavy lifting. You never write a neural network or a classifier from scratch. NumPy handles arrays, Pandas handles tables, scikit-learn handles classical models [1], and PyTorch [3] and TensorFlow [4] handle deep learning. These libraries are written by large communities, tested by millions of users, and optimized in fast C/C++ code underneath. You write short, readable Python; the heavy math runs at machine speed. This is the single biggest reason a student can reproduce a paper's experiment in an afternoon.
3. Reproducibility is built into the culture. The Python AI community expects code, data, and exact environments to be shared. Standard tools — requirements.txt, virtual environments, random seeds, notebooks — are exactly what journals now ask for when they say "make your work reproducible." Learning Python is learning the reproducibility workflow at the same time.
4. Notebooks match how research actually happens. Research is exploratory: try something, look at the result, adjust. Jupyter notebooks let you run code in small steps, see tables and plots right next to the code, and write explanations between cells. The exploratory loop in a notebook is the natural home of early-stage research.
5. Everyone you need is already here. Your supervisor's students, the paper authors whose code you download, Stack Overflow answers, dataset loaders, university courses — all speak Python. Choosing Python means every question you ask has an answer somewhere, and every tool you need already exists.
Python is slow at raw number crunching in pure form — a Python loop over a million numbers can be a hundred times slower than the same loop in C. The ecosystem solves this: NumPy, Pandas, and scikit-learn push the heavy work into compiled code, so your Python is just the conductor, not the orchestra. You will learn this pattern (called vectorization) in Chapter 4, and it is one of the most valuable skills in this book.
Python also has quirks — dynamic typing means a variable can silently change type and cause a confusing error later, and the famous "dependency hell" (library versions conflicting) is real. Chapters 2 and 10 teach you the standard defenses: virtual environments, pinning versions, and writing small testable functions.
You can think of the ecosystem as layers, each built on the one below:
Books 5 and 6 of this series will take you into deep learning. This book makes you strong in Layers 1–3, which is exactly where most student publications live: data loading, cleaning, classical models, evaluation, and clear plots.
Here is a complete machine learning experiment in Python. Do not worry if the details are unfamiliar — every piece is explained in later chapters. The point is how little code it takes:
import pandas as pd # Layer 2: tables
from sklearn.model_selection import train_test_split
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import accuracy_score
df = pd.read_csv("patients.csv") # load data
X = df[["age", "blood_pressure", "cholesterol"]]
y = df["disease"] # features and label
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, random_state=42)
model = RandomForestClassifier(random_state=42)
model.fit(X_train, y_train) # train
preds = model.predict(X_test) # predict
print("Accuracy:", accuracy_score(y_test, preds))
Fifteen lines: load data, split it, train a model, evaluate it. In older languages this same experiment could take hundreds of lines. That compression is why Python dominates — and why you can realistically go from zero to a publishable experiment in weeks, not months.
You may hear about other languages. Here is the short, honest version:
None of these is "bad." But research is a team sport played with shared tools, and Python is where the team is. When a paper releases code (increasingly expected), it is Python. When a dataset publishes a loader, it is Python. When you get stuck at 2 a.m., the answer exists for Python.
It does not mean memorizing syntax or passing a quiz on language trivia. It means you can do six things reliably:
Notice that only a small part of that list is "the Python language." Most of it is workflow — and workflow is exactly what this book teaches. A student who can do those six things is more useful in a lab than one who knows every language feature but cannot finish an experiment.
Tools that write code for you are now everywhere, and you should use them — but as an accelerator, not a replacement. They are excellent at boilerplate ("write a function that loads a CSV and prints summary stats") and dangerous at the parts that require judgment (choosing metrics, spotting leakage). The rule: never run generated code you could not explain line by line. This book gives you the understanding that makes AI assistants safe to use.
Python releases a new minor version yearly (3.10, 3.11, 3.12...), and AI libraries need months to catch up — compiled packages like NumPy must be rebuilt for each version. That is why this book recommends 3.11: new enough for modern features and speed improvements (3.11 is roughly 25% faster than 3.10), old enough that every library you need supports it. Two rules: never chase the newest release for research work, and always record the exact version (Chapter 10). A one-line version difference has broken more student experiments than any bug.
You will get stuck — everyone does. When you ask for help (supervisor, forum, AI assistant), include a minimal reproducible example: the smallest complete code that shows the problem, plus the full error message:
# Good help request: 6 lines, runs anywhere, error included
import numpy as np
from sklearn.linear_model import LogisticRegression
X = np.array([1, 2, 3]) # 1-D by mistake
model = LogisticRegression().fit(X, [0, 1, 0])
# ValueError: Expected 2D array, got 1D array instead
Preparing the example solves half of all problems by itself — shrinking the code isolates the bug. And responders answer good examples in minutes while ignoring "my code doesn't work" with no details.
For your research: When your supervisor or a reviewer asks "why Python?", the honest answer is: it is the language of the field's shared infrastructure. Choosing it means you inherit the libraries, the datasets loaders, the pretrained models, and the community that produced the papers you cite. Note this choice in your thesis methodology section — reviewers expect it, and one sentence ("All experiments were implemented in Python 3.11 using scikit-learn 1.3 [1]...") signals you follow standard practice.
Key takeaways - Python dominates AI because of readability, the library ecosystem, reproducibility culture, notebooks, and community size. - The stack is layered: Python → NumPy/Pandas → scikit-learn → PyTorch/TensorFlow. - Python is slow in raw loops; the ecosystem fixes this with vectorized compiled code underneath. - Most student papers live in Layers 1–3: data, classical models, evaluation, plots — exactly this book's territory.

A correct setup saves you from the most common beginner disaster: code that works on your laptop today and breaks everywhere else tomorrow. Spend one careful hour here and it pays back for your entire degree.
Download Python 3.11 or newer from python.org (avoid 3.13+ for now — some AI libraries lag behind new releases). During installation on Windows, check the box "Add python.exe to PATH" — this is the step most beginners miss, and without it the terminal cannot find Python.
Verify with:
python --version
# Python 3.11.9
If that prints a version, you are set. On some systems the command is python3 instead of python; pick whichever works and use it consistently.
pip is Python's package installer, and it comes with Python. Installing a library is one command:
pip install numpy pandas matplotlib seaborn scikit-learn jupyter
This installs the core AI stack: NumPy (arrays), Pandas (tables), Matplotlib and Seaborn (plots), scikit-learn (models) [1], and Jupyter (notebooks). Check an install worked:
python -c "import sklearn; print(sklearn.__version__)"
# 1.3.2
The -c flag runs a one-line Python program — handy for quick checks.
Practical tip: keep a requirements.txt file listing what you installed. Chapter 10 will show you how to generate one automatically (pip freeze), but even a hand-written list today saves confusion later.
Here is the problem virtual environments solve. Project A needs scikit-learn 1.3. Project B needs scikit-learn 1.1 (an older paper's code). Installed globally, they conflict and one project breaks. A virtual environment is a private Python sandbox per project — each has its own packages, and they never fight.
Create and activate one:
# Create (do this once per project)
python -m venv ai-env
# Activate — Windows
ai-env\Scripts\activate
# Activate — macOS/Linux
source ai-env/bin/activate
Your terminal prompt changes to show (ai-env) at the front — that is your confirmation. Now pip install goes into the sandbox only. When you are done, deactivate exits it.
Rule of thumb: one virtual environment per project, created in the project folder. Name it .venv or ai-env and never commit it to Git (add it to .gitignore). This single habit prevents the majority of "it worked on my machine" failures.
Install Jupyter (inside your virtual environment), then launch:
pip install notebook
jupyter notebook
A browser tab opens showing your files. Create a new notebook (.ipynb) and you get cells — boxes where you type code and press Shift+Enter to run. Output, including tables and plots, appears directly under the cell. Between code cells you can add Markdown cells for headings and explanations (double-click a cell and choose Markdown from the dropdown, or press Esc then M).
The workflow that makes notebooks powerful for research:
Two honest warnings about notebooks. First, cells can be run in any order, which means the notebook on screen may not match a top-to-bottom run — always do Kernel → Restart & Run All before trusting results. Second, notebooks are for exploration, not for final reusable code — Chapter 11 shows when and how to graduate working code into plain .py scripts.
Create a notebook and run these cells in order. If every cell runs without errors, your environment is ready:
# Cell 1: check the core libraries
import numpy as np
import pandas as pd
import matplotlib.pyplot as plt
import sklearn
print("numpy", np.__version__)
print("pandas", pd.__version__)
print("sklearn", sklearn.__version__)
# Cell 2: make a tiny dataset and plot it
data = {"x": [1, 2, 3, 4, 5], "y": [2, 4, 5, 4, 5]}
df = pd.DataFrame(data)
df.plot(x="x", y="y", kind="scatter")
plt.title("Setup check: my first plot")
plt.show()
# Cell 3: train a toy model
from sklearn.linear_model import LinearRegression
model = LinearRegression()
model.fit(df[["x"]], df["y"])
print("Prediction for x=6:", model.predict([[6]])[0])
Three cells: libraries verified, a plot rendered, a model trained. Save this notebook as setup_check.ipynb in your project folder — rerunning it is the fastest way to check a new machine or a fresh environment.
You will see two recommended setups and wonder which to pick:
conda). One big installer, fewer compilation headaches on Windows, and conda install sometimes succeeds where pip struggles with tricky packages.Either works. Start with plain Python; if a package repeatedly fails to install (common with some scientific libraries on Windows), try Miniconda for that project. Do not mix both managers in one environment — pick one per project and stick with it.
Almost every beginner hits one of these. None of them means you did something wrong:
| Symptom | Likely cause | Fix |
|---|---|---|
'python' is not recognized |
Missed the "Add to PATH" checkbox | Reinstall with the box checked, or use the py launcher (py --version) |
pip install fails with C++ / compiler errors |
Package needs compilation | On Windows, install Microsoft C++ Build Tools; or use conda for that package |
| Jupyter kernel dies immediately | Notebook using a different Python than your venv | In the venv: pip install ipykernel, then select the venv kernel in the notebook |
import sklearn works in terminal but not notebook |
Same as above — kernel/environment mismatch | Check sys.executable in both; they must match |
Two Pythons fighting (python vs python3) |
Multiple installs | where python (Windows) / which -a python3 (macOS/Linux) to see them; use the full path or virtual environments to disambiguate |
The golden diagnostic: when imports behave strangely, run this in the exact place the code runs (terminal or notebook cell):
import sys
print(sys.executable) # which Python is this, exactly?
Nine times out of ten, the answer is "not the one you thought," and the fix is activating the right environment.
pip install --upgrade pip.rm -rf ai-env after deactivating). They are cheap to recreate from requirements.txt (Chapter 10).Setup is done — now build the habit. A routine that works for busy students:
random_state, fit the scaler on all data — and watch the numbers change. Understanding why they change is the lesson.Thirty focused minutes daily beats a monthly marathon. Research coding is a practice, like lab technique: regular, deliberate, and recorded in your notebook.
pip install notebook gives the classic interface; JupyterLab (pip install jupyterlab, run jupyter lab) is its modern successor — tabs, a file browser, a visual debugger, and a better text editor in one window. Either is fine for this book; JupyterLab is the better long-term home as your projects grow. VS Code with the Python and Jupyter extensions is a popular third option that many labs standardize on.
For your research: Your thesis or paper methodology should name your exact environment: Python version, key library versions, and operating system. Create this record the day you start, not the day you submit. A simple
environment.txtwith the output ofpython --versionandpip freezeis enough. Reviewers increasingly ask for it, and reconstructing it months later from memory is unreliable.
Key takeaways
- Install Python 3.11+, verify with python --version, install libraries with pip.
- One virtual environment per project (python -m venv ai-env) — never install research packages globally.
- Jupyter notebooks are your exploratory lab: run cells, see output, write explanations alongside.
- Always Restart & Run All before trusting a notebook; notebooks are for exploration, scripts are for reuse.
- Record your environment from day one — future-you (and your reviewers) will thank you.
This chapter covers the 20% of Python you will use 80% of the time in AI work. If you have programmed before, skim fast and do the worked example. If Python is new to you, type every snippet into a notebook and run it — reading code teaches little; running it teaches a lot.
A variable is a named box holding a value. Python figures out the type automatically:
name = "Asif" # str — text
epoch = 10 # int — whole number
accuracy = 0.942 # float — decimal number
done = True # bool — True/False
Check a type with type(accuracy) → <class 'float'>. The key types in AI code are int, float, str, list, dict, and later ndarray (NumPy) and DataFrame (Pandas).
Dynamic typing gotcha: a variable can be reassigned to a different type (x = 5 then x = "five"), and Python will not warn you. In long experiment scripts this causes confusing errors far from the mistake. The defense: keep variable names meaningful (test_accuracy, not x) and re-run cells top to bottom.
Lists hold ordered items and can be changed:
scores = [0.91, 0.87, 0.93] # three cross-validation scores
scores.append(0.89) # add one
print(scores[0]) # first item: 0.91 (indexing starts at 0)
print(scores[-1]) # last item: 0.89
print(scores[0:2]) # slice: [0.91, 0.87]
Tuples are like lists but unchangeable — used for fixed records:
image_shape = (256, 256, 3) # height, width, color channels
Dictionaries map keys to values — the workhorse for configs and results:
config = {"model": "random_forest", "max_depth": 5, "n_estimators": 100}
print(config["model"]) # random_forest
config["accuracy"] = 0.93 # add a new key
You will log experiment results as dictionaries and configs as dictionaries throughout your research career.
# if: make decisions
if accuracy >= 0.90:
print("Good enough to report")
elif accuracy >= 0.80:
print("Needs tuning")
else:
print("Try a different approach")
# for: repeat over a collection
for depth in [3, 5, 10]:
model = RandomForestClassifier(max_depth=depth)
model.fit(X_train, y_train)
print(depth, model.score(X_test, y_test))
# while: repeat until a condition changes (use sparingly)
epoch = 0
while epoch < 3:
print("training epoch", epoch)
epoch += 1
The for loop over a list of hyperparameter values is the simplest form of hyperparameter search — you will write this pattern constantly.
A function packages logic with a name, inputs, and an output:
def evaluate(model, X_test, y_test):
"""Return accuracy of a fitted model on test data."""
preds = model.predict(X_test)
return (preds == y_test).mean()
acc = evaluate(model, X_test, y_test)
print(f"Test accuracy: {acc:.3f}") # f-string formatting: 0.933
Three habits that make functions research-grade:
load_data(), train(), evaluate(), not one giant do_everything()."""...""" line) saying what it returns. Future-you will forget; the docstring remembers.A list comprehension builds a list in one readable line:
# The long way
cleaned = []
for s in scores:
cleaned.append(round(s, 2))
# The comprehension way
cleaned = [round(s, 2) for s in scores]
# With a condition
good = [s for s in scores if s >= 0.90]
Dictionary comprehensions work the same way: {m: evaluate(m, X_test, y_test) for m in models} builds a name→accuracy map in one line. Comprehensions are not just style — they are faster than explicit loops and they are the idiom reviewers and collaborators expect.
This example combines everything — it trains three models, evaluates each, and reports the winner, the pattern you will reuse in every comparative study:
from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LogisticRegression
from sklearn.svm import SVC
from sklearn.ensemble import RandomForestClassifier
def run_comparison(models, X_train, X_test, y_train, y_test):
"""Train each model and return {name: accuracy}."""
results = {}
for name, model in models.items():
model.fit(X_train, y_train)
results[name] = round(model.score(X_test, y_test), 3)
return results
X, y = load_iris(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, random_state=42)
models = {
"logistic_regression": LogisticRegression(max_iter=200),
"svm": SVC(),
"random_forest": RandomForestClassifier(random_state=42),
}
results = run_comparison(models, X_train, X_test, y_train, y_test)
for name, acc in sorted(results.items(), key=lambda kv: kv[1], reverse=True):
print(f"{name:20s} {acc:.3f}")
best = max(results, key=results.get)
print("Winner:", best)
Expected output (your numbers may differ slightly by version):
random_forest 1.000
logistic_regression 1.000
svm 0.967
Winner: random_forest
On the small Iris dataset, simple models often tie — and that is itself a research lesson: always report all compared models, not just the winner, because the comparison table is what a paper's results section is built from.
You will print metrics hundreds of times. f-strings (formatted string literals) are the clean way:
acc = 0.93347
print(f"Accuracy: {acc:.3f}") # Accuracy: 0.933 (3 decimals)
print(f"Accuracy: {acc:.1%}") # Accuracy: 93.3% (percentage)
print(f"Samples: {12500:,}") # Samples: 12,500 (thousands separator)
print(f"Model: {name:20s} {acc:.3f}")# padded column — neat tables in the terminal
Format specs worth memorizing: :.3f (decimals), :.1% (percent), :, (thousands), :20s (padded text). Clean console output becomes your lab notebook's first draft.
1. Mutable default arguments. This is Python's most famous trap:
def add_result(score, results=[]): # DANGER: the list persists between calls!
results.append(score)
return results
Call it twice and the second call sees the first call's scores. The fix:
def add_result(score, results=None): # safe pattern
if results is None:
results = []
results.append(score)
return results
2. Float comparison. Never compare floats with == — tiny representation errors make 0.1 + 0.2 == 0.3 false. Use a tolerance:
abs(a - b) < 1e-9 # or
import math; math.isclose(a, b)
This matters when you assert that two runs produced "the same" metric.
3. Notebook variable ghosts. In notebooks, a variable set in a cell you later deleted still exists in the kernel. If results look impossibly good, restart the kernel and run all cells — the ghost disappears and the truth appears.
4. is vs ==. Use == for value comparison (if acc == best:). Use is only for None (if x is None:) — a convention the whole ecosystem follows.
You will see three error types constantly — learn their meaning once:
NameError: name 'df' is not defined — you used a variable before creating it (or the cell that creates it did not run).TypeError: unsupported operand — wrong types mixed, e.g., adding a string to a number; check type() of each piece.KeyError: 'age' — dictionary/DataFrame has no such key; check spelling and df.columns.Each error names the line and the values involved. Read it literally — it is almost always telling the truth.
Three small tools that make loops cleaner — you will see them in every codebase you read:
models = ["lr", "svm", "rf"]
scores = [0.91, 0.87, 0.93]
for i, name in enumerate(models): # index + value together
print(i, name)
for name, score in zip(models, scores): # walk two lists in parallel
print(f"{name}: {score:.2f}")
best_name, best_score = "rf", 0.93 # unpacking: one line, two variables
first, *rest = scores # first=0.91, rest=[0.87, 0.93]
zip is how you pair model names with their scores for reporting; enumerate replaces manual counter variables. Small, but they remove entire categories of off-by-one mistakes.
Older code uses map and filter; modern Python prefers comprehensions — they read better and the community standard is clear:
# Equivalent, but the comprehension is the idiom reviewers expect
squared = [x**2 for x in range(10)] # prefer this
squared = list(map(lambda x: x**2, range(10))) # not this
Nested comprehensions flatten grids and matrices: [v for row in matrix for v in row] turns a 2-D list into 1-D. Read it as nested loops, left to right.
For your research: Every paper's results section is a comparison table: rows are methods, columns are metrics. The
run_comparisonpattern above is that table in code form. Build this habit now: never evaluate one model in isolation — always evaluate a small set including at least one simple baseline (logistic regression is the classic). Reviewers distrust a new method that was never compared against a simple alternative.
Key takeaways
- Master variables, lists, dicts, if/for, functions, and comprehensions — that is most of daily AI coding.
- Dictionaries are your config and results-log format; functions should do one job and return values.
- The train-several-models-and-compare pattern is the code behind every results table you will publish.
NumPy gives Python its number-crunching muscle. Its central object, the ndarray (N-dimensional array), stores data in a compact, fixed-type block of memory and performs math in compiled C code. Datasets, images, model weights, and predictions are all ndarrays underneath. If Pandas is the table you look at, NumPy is the engine room below it.
import numpy as np
a = np.array([1, 2, 3, 4]) # 1-D array (a vector)
b = np.array([[1, 2], [3, 4]]) # 2-D array (a matrix)
c = np.zeros((3, 4)) # 3x4 array of zeros
d = np.ones((2, 2)) # 2x2 array of ones
e = np.arange(0, 10, 2) # [0 2 4 6 8]
f = np.linspace(0, 1, 5) # 5 evenly spaced values 0..1
g = np.random.rand(3, 3) # 3x3 random values in [0, 1)
Every array has three vital attributes — learn to print them when confused:
print(b.shape) # (2, 2) — dimensions
print(b.ndim) # 2 — number of dimensions
print(b.dtype) # int64 — element type (all elements share one type)
The shape is the single most important debugging fact in ML code. Most errors you will meet are shape mismatches ("expected (100, 5), got (100,)"). Print shapes early and often.
m = np.array([[1, 2, 3],
[4, 5, 6],
[7, 8, 9]])
print(m[0, 1]) # 2 — row 0, column 1
print(m[:, 1]) # [2 5 8] — all rows, column 1
print(m[1, :]) # [4 5 6] — row 1, all columns
print(m[0:2, 1:3]) # [[2 3] [5 6]] — sub-block
print(m[m > 5]) # [6 7 8 9] — boolean mask: elements where condition holds
Boolean masks (m[m > 5]) are how you filter data without loops — "give me all samples where the label is 1" is X[y == 1]. You will use this constantly when analyzing errors ("show me the images the model got wrong").
A vectorized operation applies to the whole array at once, in compiled code:
# Slow: Python loop (avoid for large data)
result = []
for x in a:
result.append(x ** 2)
# Fast: vectorized (do this)
result = a ** 2
The difference is not subtle — on a million elements, the loop can take a second while the vectorized version takes milliseconds. Rule: if you find yourself writing for over array elements to do math, stop and look for the NumPy function (np.mean, np.sum, np.max, np.dot, ...). Aggregation functions take an axis argument: np.mean(m, axis=0) averages down each column, axis=1 across each row.
Broadcasting is NumPy's rule for combining arrays of different shapes: the smaller array is conceptually "stretched" to match. The classic use is normalization — subtract the mean of each column from every row:
data = np.array([[10., 20.],
[30., 40.],
[50., 60.]])
col_mean = data.mean(axis=0) # [30. 40.] — shape (2,)
centered = data - col_mean # broadcasts: subtracts from each row
print(centered)
# [[-20. -20.]
# [ 0. 0.]
# [ 20. 20.]]
Broadcasting rules in one sentence: dimensions are aligned from the right, and a dimension of size 1 (or a missing dimension) stretches to match. When broadcasting fails, NumPy raises a clear error — read it, print the shapes, and you will see the fix.
v = np.arange(6) # [0 1 2 3 4 5]
print(v.reshape(2, 3)) # 2 rows, 3 cols
print(v.reshape(-1, 2)) # -1 means "figure it out": 3 rows, 2 cols
flat = m.flatten() # 2-D -> 1-D (images become feature vectors this way)
combo = np.vstack([a, a]) # stack vertically
reshape(-1, ...) is the idiom for "I know the columns, compute the rows." Flattening a 28×28 image to 784 numbers is exactly img.reshape(-1) — the first step of feeding images to a classical classifier.
Feature scaling (making columns comparable) is required by many models [2], and this example shows vectorization, broadcasting, and a timing comparison in one script:
import numpy as np
import time
rng = np.random.default_rng(42)
X = rng.normal(loc=50, scale=15, size=(20000, 10)) # 20k samples, 10 features
# Vectorized standardization: (x - mean) / std, per column
means = X.mean(axis=0)
stds = X.std(axis=0)
X_scaled = (X - means) / stds # broadcasting does the work
print("Column means after scaling:", X_scaled.mean(axis=0).round(6))
# all ~0.0 — each feature now centered at zero
# Speed comparison: vectorized vs Python loop
t0 = time.perf_counter()
_ = (X - means) / stds
t_vec = time.perf_counter() - t0
t0 = time.perf_counter()
out = np.empty_like(X)
for i in range(X.shape[0]):
for j in range(X.shape[1]):
out[i, j] = (X[i, j] - means[j]) / stds[j]
t_loop = time.perf_counter() - t0
print(f"Vectorized: {t_vec*1000:.1f} ms | Loop: {t_loop*1000:.1f} ms "
f"| Speedup: {t_loop/t_vec:.0f}x")
Typical output: Vectorized: 1.2 ms | Loop: 380.4 ms | Speedup: 317x. That gap is why vectorization is non-negotiable: on real datasets, loop-based preprocessing turns minutes into hours.
A Python list is a collection of objects — each number carries type information and overhead (about 28 bytes per float). A NumPy array stores raw numbers packed together (8 bytes per float64) with one shared type. Consequences:
data[data > threshold] has no list equivalent in one line.The rule is simple: the moment data becomes numeric and sizable, it becomes an array. Lists remain for mixed or small collections (file names, model names, config values).
Old tutorials use np.random.seed() + np.random.rand(). The modern, recommended API uses a Generator object, which gives independent, reproducible streams:
rng = np.random.default_rng(42) # seed once
a = rng.normal(0, 1, size=(100, 5))# 100x5 from a standard normal distribution
b = rng.integers(0, 10, size=50) # 50 random integers in [0, 10)
c = rng.choice(["a", "b", "c"], size=20) # random sampling from options
Use default_rng(SEED) in your own code (you saw it in the worked example). It is also what makes synthetic datasets for testing — like Chapter 12's capstone data — reproducible.
Keep this list handy; each replaces a loop you might otherwise write:
np.mean(X, axis=0) # column means (feature averages)
np.std(X, axis=0) # column standard deviations
np.min(X), np.max(X) # range checks — catch bad data fast
np.argmin(errors) # index of the best model in a list of scores
np.unique(y) # sorted unique values — instant class inventory
np.bincount(y) # count of each integer label — class balance in one call
np.concatenate([a, b]) # join arrays
np.dot(A, B) # or: A @ B # matrix multiplication — the heart of ML math
np.linalg.norm(v) # vector length — used in distance-based methods
np.percentile(X, 95) # 95th percentile — robust outlier thresholds
np.where(X > 0, 1, 0) # vectorized if/else — binarize in one line
np.isnan(X).any() # the NaN check from Chapter 9
np.unique and np.bincount deserve emphasis: "what classes exist and how many of each?" is a question you ask of every dataset, and these answer it in one line.
Slicing an array does not copy it — it creates a view sharing the same memory. Modifying the view modifies the original:
a = np.arange(5) # [0 1 2 3 4]
b = a[1:4] # view, not a copy!
b[0] = 99
print(a) # [0 99 2 3 4] — the original changed!
This is a feature (views are free — no memory copied, which matters for large datasets), but it surprises everyone once. When you need independence, copy explicitly: b = a[1:4].copy(). Boolean-mask indexing (a[a > 2]) does return a copy — the inconsistency is historical, so when in doubt, .copy() and move on.
axis=0 means "down the rows" (collapsing rows, one result per column); axis=1 means "across the columns" (one result per row). The mnemonic: the axis number is the dimension that disappears:
m = np.array([[1, 2, 3],
[4, 5, 6]])
print(m.mean(axis=0)) # [2.5 3.5 4.5] — one mean per column (rows collapsed)
print(m.mean(axis=1)) # [2. 5.] — one mean per row (columns collapsed)
For a dataset shaped (samples, features), axis=0 gives per-feature statistics — the direction you aggregate in 90% of preprocessing. Getting axes wrong silently produces plausible-looking wrong numbers, so verify with a tiny example like the one above whenever you are unsure.
For your research: Reviewers will not ask "did you vectorize?" — but they will notice if your experiments take implausibly long or if your code cannot scale to the full dataset. More importantly, NumPy fluency is what lets you implement a method section faithfully: when a paper says "we standardized features to zero mean and unit variance," the code is the three lines above. Practice translating one formula from a paper into NumPy each week; it is the fastest path from reading papers to reproducing them.
Key takeaways
- The ndarray is the foundation: fixed type, N-dimensional, with a shape you should always know.
- Vectorize instead of looping — expect 10–100x+ speedups on real data.
- Broadcasting aligns mismatched shapes; it is how normalization and many paper formulas are written.
- Boolean masks filter data without loops; reshape moves between images and feature vectors.
Real research data arrives messy: spreadsheets with merged headers, missing values, inconsistent labels, stray text in numeric columns. Pandas is the toolkit for turning that mess into a clean table a model can learn from. Researchers routinely spend more time in Pandas than in any model library — data work is most of the work.
A DataFrame is a table: rows are samples, columns are variables, both labeled. A Series is a single column.
import pandas as pd
df = pd.read_csv("patients.csv") # the most common line in research code
print(df.shape) # (rows, columns)
print(df.head()) # first 5 rows — always look at your data
print(df.dtypes) # type of each column
print(df.describe()) # count, mean, std, min, max per numeric column
df.head(), df.info(), and df.describe() are your first three moves on any new dataset. Run them before anything else — they reveal missing values, wrong types, and nonsense values in seconds.
df["age"] # one column (a Series)
df[["age", "bp"]] # several columns (a DataFrame)
df.iloc[0] # first row by position
df.iloc[0:5, 0:2] # rows 0-4, columns 0-1 by position
df.loc[df["age"] > 60] # rows by condition (boolean filter)
The pattern df.loc[condition] is how you answer research questions directly in code: "how many patients over 60 have the disease?" is df.loc[(df["age"] > 60) & (df["disease"] == 1)].shape[0].
Missing data is the norm, not the exception. First, find it:
print(df.isna().sum()) # missing count per column
Then choose a strategy deliberately — your choice belongs in the paper's methodology:
df_clean = df.dropna() # drop rows with any missing value
df["bp"] = df["bp"].fillna(df["bp"].mean()) # fill with column mean
df["gender"] = df["gender"].fillna(df["gender"].mode()[0]) # fill with most common
Guidance: dropping is honest but shrinks small datasets; mean-filling is standard for numeric features; never fill the label column — if the answer is missing, the row cannot be used for supervised learning. Whatever you choose, report it: "3.2% of blood-pressure values were missing and were imputed with the column mean."
df["age"] = pd.to_numeric(df["age"], errors="coerce") # text -> numbers; bad values become NaN
df["date"] = pd.to_datetime(df["date"]) # parse dates properly
df = df.drop_duplicates() # remove exact duplicate rows
df["disease"] = df["disease"].astype(int) # ensure label is integer
CSV files love to smuggle numbers in as text ("1,200" or "45 years"). pd.to_numeric(..., errors="coerce") converts what it can and turns the rest into NaN for you to handle — far safer than code that crashes on row 9,412.
groupby splits the table into groups and computes statistics per group — it is how you produce the descriptive statistics tables in papers:
summary = df.groupby("disease")[["age", "bp", "cholesterol"]].mean()
print(summary)
# age bp cholesterol
# disease
# 0 45.2 120.1 198.4
# 1 58.7 142.3 231.9
One line tells you the diseased group is older with higher blood pressure — the kind of observation that motivates features and appears in your paper's data description section.
merged = pd.merge(patients, lab_results, on="patient_id") # join two tables
df_clean.to_csv("patients_clean.csv", index=False) # save your cleaned data
Golden rule: never overwrite the raw file. Save cleaned data under a new name (*_clean.csv), and keep the cleaning script. When a reviewer asks "how did you handle the missing values?", you point at the script, not your memory.
import pandas as pd
import numpy as np
# Simulate a messy real-world CSV
raw = pd.DataFrame({
"age": ["45", "52", "unknown", "61", "45", "38"],
"bp": [120, np.nan, 140, 155, 120, 118],
"gender": ["M", "F", "M", None, "M", "F"],
"disease": [0, 1, 1, 1, 0, 0],
})
raw.to_csv("messy.csv", index=False)
# --- Cleaning pipeline ---
df = pd.read_csv("messy.csv")
print("Before:\n", df.isna().sum(), "\nduplicates:", df.duplicated().sum())
df["age"] = pd.to_numeric(df["age"], errors="coerce") # "unknown" -> NaN
df["age"] = df["age"].fillna(df["age"].median()) # robust fill
df["bp"] = df["bp"].fillna(df["bp"].mean())
df["gender"] = df["gender"].fillna(df["gender"].mode()[0])
df = df.drop_duplicates().reset_index(drop=True)
print("\nAfter:\n", df.isna().sum())
print(df.describe().round(1))
df.to_csv("messy_clean.csv", index=False)
print("\nSaved cleaned data:", df.shape)
Output shows the transformation: 6 rows in, text ages coerced, NaNs filled, one duplicate removed, clean file saved. Every line maps to a sentence you can write in a methodology section.
Real datasets can exceed your computer's memory. Three techniques handle that:
# 1. Peek at 1000 rows before committing to the full file
peek = pd.read_csv("huge.csv", nrows=1000)
# 2. Load only the columns you need
df = pd.read_csv("huge.csv", usecols=["age", "bp", "disease"])
# 3. Process in chunks (e.g., compute the mean without loading everything)
total, n = 0, 0
for chunk in pd.read_csv("huge.csv", chunksize=100_000, usecols=["bp"]):
total += chunk["bp"].sum()
n += len(chunk)
print("Mean bp:", total / n)
For this book's projects you will rarely need chunks — but knowing they exist means a large dataset never stops you.
Models need numbers. A column like gender with values "M"/"F" must be encoded. The standard approach is one-hot encoding — one binary column per category:
df = pd.get_dummies(df, columns=["gender"], dtype=int)
# gender_M | gender_F
# 1 | 0
Two cautions: one-hot encoding a column with hundreds of categories (like city names) explodes your feature count — consider grouping rare categories first. And when you split data, encode after splitting or ensure train and test get identical columns (a category missing from one split causes shape mismatches — Chapter 9's bug #1).
Pandas string operations are vectorized — no loops needed:
df["disease"] = df["disease"].str.strip().str.lower() # " Healthy " -> "healthy"
df["has_fever"] = df["notes"].str.contains("fever", na=False)
When logic is too custom for vectorized ops, apply runs a function per row — flexible but slow (it is a loop in disguise):
def risk_band(row):
if row["age"] > 60 and row["bp"] > 140:
return "high"
return "low"
df["risk"] = df.apply(risk_band, axis=1) # works, but avoid on millions of rows
Prefer vectorized operations; reserve apply for genuinely row-wise logic on modest data.
pd.merge joins on shared keys, and the how argument decides which rows survive:
# patients: id, age | lab: id, cholesterol
inner = pd.merge(patients, lab, on="id", how="inner") # only ids in BOTH (default)
left = pd.merge(patients, lab, on="id", how="left") # all patients, NaN where lab missing
Use inner when you need complete records (typical for modeling), left when you must keep every original row and will handle the gaps. After any merge, check the row count — an unexpected jump means duplicate keys multiplied your rows (a patient id appearing twice in lab doubles those rows), and an unexpected drop means keys did not match (check for "P001" vs "p001" mismatches with .str.strip().str.lower()).
df.sort_values("bp", ascending=False).head(10) # highest blood pressures — sanity check
df["bp_rank"] = df["bp"].rank(pct=True) # percentile rank per row
df.nlargest(5, "cholesterol") # top 5 without full sort
Sorting is an underrated debugging tool: the largest and smallest values of every column reveal data-entry errors (age 999, negative blood pressure) faster than any statistic. Make "sort each column both ways and eyeball the extremes" part of your inspection ritual.
Datasets arrive in two shapes. Wide format has one column per measurement (one row per patient, columns bp_jan, bp_feb, ...). Long format has one row per measurement (columns patient, month, bp). Models usually want wide; plotting and grouping usually want long. Pandas converts both ways:
wide = pd.DataFrame({"patient": [1, 2],
"bp_jan": [120, 140], "bp_feb": [122, 138]})
long = wide.melt(id_vars="patient", var_name="month", value_name="bp")
# patient month bp
# 0 1 bp_jan 120 ...
back = long.pivot(index="patient", columns="month", values="bp")
melt (wide→long) and pivot (long→wide) are the reshaping pair. Recognizing which shape your data is in — and which shape your next step needs — removes a whole class of "why won't this plot/model accept my data" confusion.
For your research: Your paper needs a "Dataset" subsection, and this chapter is its source material: where the data came from, how many rows/columns, what was missing and how you handled it, and what the class balance is (
df["label"].value_counts()— always check; imbalanced classes change which metrics are honest). Keep the raw file, the cleaning script, and the cleaned file together. If your data cannot be shared, the cleaning script plus a small synthetic sample still lets others understand exactly what you did.
Key takeaways
- First moves on any dataset: head(), info(), describe(), isna().sum().
- Handle missing values deliberately (drop, mean/median/mode fill) and report the choice.
- Never overwrite raw data; save cleaned versions and keep the cleaning script.
- groupby produces the descriptive statistics your paper's dataset section needs.
You will make two kinds of plots in your research life: exploration plots (quick, ugly, for your eyes only — "is there a pattern here?") and publication plots (clean, labeled, for reviewers — "here is the evidence"). Matplotlib gives you full control; Seaborn gives you beautiful statistical plots with one line. Learn both, in that order.
Every Matplotlib figure follows the same pattern — create, plot, label, show:
import matplotlib.pyplot as plt
x = [1, 2, 3, 4, 5]
y = [2, 4, 5, 4, 5]
plt.figure(figsize=(6, 4)) # figure size in inches
plt.plot(x, y, marker="o") # line plot with markers
plt.xlabel("Study hours") # always label axes
plt.ylabel("Exam score")
plt.title("Score vs study hours")
plt.grid(True, alpha=0.3)
plt.savefig("plot.png", dpi=150) # save BEFORE show()
plt.show()
The three plot types you will use most:
plt.scatter(df["age"], df["bp"]) # scatter: relationship between two variables
plt.hist(df["age"], bins=20) # histogram: distribution of one variable
plt.bar(["A", "B", "C"], [0.91, 0.87, 0.93]) # bar: compare values across groups
Subplots put several plots in one figure — essential for paper figures:
fig, axes = plt.subplots(1, 2, figsize=(10, 4))
axes[0].hist(df["age"], bins=20); axes[0].set_title("Age distribution")
axes[1].scatter(df["age"], df["bp"]); axes[1].set_title("Age vs blood pressure")
plt.tight_layout() # prevents overlapping labels
plt.show()
Seaborn sits on top of Matplotlib and understands DataFrames directly:
import seaborn as sns
sns.histplot(data=df, x="age", hue="disease", kde=True) # distributions by class
sns.boxplot(data=df, x="disease", y="bp") # spread per class
sns.heatmap(df.corr(), annot=True, cmap="coolwarm") # correlation matrix
sns.pairplot(df, hue="disease") # every variable vs every variable
Two Seaborn plots deserve special attention because they appear in almost every ML paper:
Set a clean style once per notebook: sns.set_theme(style="whitegrid").
Before training any model, plot. Ten minutes of plots routinely reveals what hours of modeling cannot:
df["label"].value_counts().plot(kind="bar") — if one class is 95% of the data, accuracy is a misleading metric.hue="disease") — features whose distributions separate by class are the ones the model will use.A plot in a paper must be self-contained: a reader who skips the text should still understand it. Checklist:
fontsize=12+; default is often too small)sns.despine(), keep grids subtleplt.savefig("fig1.png", dpi=300, bbox_inches="tight")Common beginner mistake: screenshots of plots. Never screenshot — always savefig to PNG/PDF. Vector PDF (plt.savefig("fig.pdf")) stays razor-sharp at any zoom and is what journals prefer.
import pandas as pd
import matplotlib.pyplot as plt
import seaborn as sns
from sklearn.datasets import load_breast_cancer
sns.set_theme(style="whitegrid")
data = load_breast_cancer(as_frame=True)
df = data.frame # features + 'target' column
fig, axes = plt.subplots(2, 2, figsize=(11, 8))
# 1. Class balance
df["target"].value_counts().plot(kind="bar", ax=axes[0, 0], color=["#4c78a8", "#f58518"])
axes[0, 0].set_title("Class balance (0=benign, 1=malignant)")
axes[0, 0].set_xlabel("Class"); axes[0, 0].set_ylabel("Count")
# 2. Distribution of a key feature by class
sns.histplot(data=df, x="mean radius", hue="target", kde=True, ax=axes[0, 1])
axes[0, 1].set_title("Mean radius by class")
# 3. Boxplot: outlier check
sns.boxplot(data=df, x="target", y="mean area", ax=axes[1, 0])
axes[1, 0].set_title("Mean area by class (outlier check)")
# 4. Correlation heatmap of top features
top = ["mean radius", "mean texture", "mean area", "mean smoothness"]
sns.heatmap(df[top].corr(), annot=True, fmt=".2f", cmap="coolwarm", ax=axes[1, 1])
axes[1, 1].set_title("Feature correlations")
plt.tight_layout()
plt.savefig("exploration.png", dpi=150, bbox_inches="tight")
plt.show()
Four panels, one figure: balance, separation, outliers, redundancy. This exact figure — or a cleaned-up version — belongs in your paper's dataset description. Notice the workflow: explore first, model second.
| Question | Plot | Code |
|---|---|---|
| How is one variable distributed? | Histogram / KDE | sns.histplot(df["age"], kde=True) |
| Does the distribution differ by class? | Grouped histogram | sns.histplot(data=df, x="age", hue="label") |
| How do two variables relate? | Scatter | plt.scatter(df["age"], df["bp"]) |
| How does something change over time/epochs? | Line | plt.plot(epochs, accuracies, marker="o") |
| Which group is biggest / which model wins? | Bar | df["label"].value_counts().plot(kind="bar") |
| Where are the outliers? | Boxplot | sns.boxplot(data=df, x="label", y="bp") |
| Which features move together? | Heatmap | sns.heatmap(df.corr(), annot=True) |
| What errors does my model make? | Confusion matrix | sns.heatmap(confusion_matrix(...), annot=True) |
| How sure am I? (variation across runs) | Bar with error bars | plt.bar(models, means, yerr=stds) |
When in doubt, start with a histogram (one variable) or scatter (two variables). They answer "what does my data look like?" — the question behind every other question.
Exploration plots can be rough; paper plots need polish. The upgrades that matter most:
plt.figure(figsize=(6, 4))
plt.plot(epochs, acc, marker="o", linewidth=2)
plt.xlabel("Epoch", fontsize=12)
plt.ylabel("Validation accuracy", fontsize=12)
plt.title("Training curve: random forest", fontsize=13)
sns.despine() # remove top/right borders — cleaner look
plt.tight_layout()
plt.savefig("fig2_training_curve.pdf", bbox_inches="tight") # vector PDF for journals
Three rules for multi-figure papers: one color palette everywhere (define it once, reuse it), consistent font sizes (a figure shrunk to column width must stay readable), and the same axis ranges when figures invite comparison (different y-axis scales on two accuracy plots mislead even when honest).
Annotating: point the reader at what matters with plt.annotate("overfitting begins", xy=(30, 0.81), xytext=(15, 0.7), arrowprops=dict(arrowstyle="->")). A well-placed annotation turns a plot from "data" into "evidence."
Two y-axes on one plot (e.g., accuracy on the left, loss on the right) are tempting but easy to manipulate — rescaling either axis changes the visual story without changing the data. Prefer two aligned subplots sharing the x-axis:
fig, (ax1, ax2) = plt.subplots(2, 1, figsize=(7, 6), sharex=True)
ax1.plot(epochs, acc, marker="o"); ax1.set_ylabel("Accuracy")
ax2.plot(epochs, loss, marker="s", color="tab:red"); ax2.set_ylabel("Loss")
ax2.set_xlabel("Epoch")
General honesty rules for research plots: never truncate a y-axis to exaggerate small differences (or if you must, say so in the caption); use the same scale when inviting comparison; and show uncertainty (error bars, shaded std bands) whenever you averaged multiple runs — a line without its wobble overstates confidence.
Static plots go in papers; interactive plots speed up exploration. Plotly Express creates zoomable, hoverable figures in one line:
import plotly.express as px
fig = px.scatter(df, x="age", y="bp", color="disease",
hover_data=["cholesterol"], title="Explore: hover any point")
fig.show()
Hovering reveals the exact values behind outliers — invaluable when a strange point turns out to be your most interesting sample. Use interactive plots for yourself, static Matplotlib/Seaborn for the paper.
Journals print figures at fixed column widths, and a figure designed for your laptop screen turns into unreadable micro-text when shrunk. Design at the final size:
# Single column (~3.5 inches wide) vs double column (~7 inches)
plt.figure(figsize=(3.5, 2.6)) # single-column figure
plt.figure(figsize=(7.0, 4.0)) # double-column figure
Set fonts with the shrink in mind: axis labels at 9–10 pt look right in a single-column figure (they render at roughly that size in print). Save with dpi=300 for raster formats (PNG) or as PDF for vector sharpness, always with bbox_inches="tight" so labels are not clipped. One practical test: print the figure at its final width — if you cannot read it on paper, the reviewer cannot either.
For your research: Every figure in your paper needs a number, a caption, and a sentence in the text that references it ("Figure 3 shows..."). Write the caption to state the conclusion, not just the content: weak — "Accuracy by model"; strong — "Random forest achieves the highest accuracy (93.1%) across all feature sets." Reviewers skim figures first; make each one argue for you. And keep the script that generated every figure — when a reviewer asks for a different visualization, you regenerate in minutes instead of rebuilding from memory.
Key takeaways - Matplotlib = control; Seaborn = fast statistical plots. Learn the basic pattern: create, plot, label, save, show. - Plot before modeling: balance, distributions, outliers, correlations. - Publication plots must be self-contained, high-resolution (savefig, never screenshot), and captioned with the conclusion.
This is the chapter the whole book has been building toward. Scikit-learn [1] is the standard library for classical machine learning in Python, and its brilliance is a consistent interface: every model is trained with .fit(), predicts with .predict(), and scores with .score(). Learn the workflow once, and you can try dozens of models.
Memorize this — it is the skeleton of every experiment in Books 3 through 6 and of most published ML papers [2].
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score, classification_report, confusion_matrix
# 1. Load: 569 samples, 30 numeric features, benign/malignant labels
X, y = load_breast_cancer(return_X_y=True)
# 2. Split: 80% train, 20% test, stratified keeps class ratios equal
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, random_state=42, stratify=y)
# 3-4. Scale features, then train (many models need scaling [2])
scaler = StandardScaler()
X_train_s = scaler.fit_transform(X_train) # fit scaler on TRAIN only
X_test_s = scaler.transform(X_test) # apply same transform to test
model = LogisticRegression(max_iter=1000)
model.fit(X_train_s, y_train)
# 5. Evaluate
preds = model.predict(X_test_s)
print("Accuracy:", accuracy_score(y_test, preds))
print(classification_report(y_test, preds,
target_names=["benign", "malignant"]))
The classification report shows precision (of predicted positives, how many were right), recall (of actual positives, how many were found), and F1 (their balance) per class. On medical or imbalanced data, these matter more than accuracy — a model that calls everything "benign" scores 63% accuracy here while finding zero cancers.
Critical rule — no leakage: fit the scaler (or any preprocessing) on the training data only, then transform the test data. Fitting on all data leaks test information into training and inflates your score. Reviewers check for this; it is one of the most common reasons student results get questioned.
from sklearn.svm import SVC
from sklearn.ensemble import RandomForestClassifier
models = {
"logistic_regression": LogisticRegression(max_iter=1000),
"svm": SVC(),
"random_forest": RandomForestClassifier(random_state=42),
}
for name, m in models.items():
m.fit(X_train_s, y_train)
print(f"{name:20s} accuracy={m.score(X_test_s, y_test):.3f}")
Typical output:
logistic_regression accuracy=0.982
svm accuracy=0.982
random_forest accuracy=0.965
Report all three, not just the winner. A paper that says "we tried X, Y, Z; X won" is more credible than one that presents a single model with no comparison [5].
A single train/test split can be lucky or unlucky. Cross-validation rotates the split (typically 5 times) and averages the scores:
from sklearn.model_selection import cross_val_score
scores = cross_val_score(model, X_train_s, y_train, cv=5)
print(f"CV accuracy: {scores.mean():.3f} ± {scores.std():.3f}")
# CV accuracy: 0.978 ± 0.012
"97.8% ± 1.2%" is a publishable number — it tells the reader both the performance and its stability. Use cross-validation for model selection, then report final performance on the held-out test set.
A hyperparameter is a setting you choose (like max_depth), as opposed to parameters the model learns. The simplest tuning is a loop (Chapter 3's pattern), but scikit-learn's GridSearchCV does it properly with cross-validation built in:
from sklearn.model_selection import GridSearchCV
grid = GridSearchCV(
RandomForestClassifier(random_state=42),
param_grid={"n_estimators": [50, 100], "max_depth": [5, 10, None]},
cv=5)
grid.fit(X_train_s, y_train)
print("Best:", grid.best_params_, "CV score:", round(grid.best_score_, 3))
print("Test accuracy:", round(grid.score(X_test_s, y_test), 3))
Report the searched grid and the winner in your methodology — "tuned via 5-fold grid search over ..." is a sentence reviewers expect.
import matplotlib.pyplot as plt
import seaborn as sns
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import confusion_matrix, classification_report
X, y = load_breast_cancer(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, random_state=42, stratify=y)
X_train_s = StandardScaler().fit_transform(X_train)
X_test_s = StandardScaler().fit(X_train).transform(X_test)
model = RandomForestClassifier(n_estimators=100, random_state=42)
model.fit(X_train_s, y_train)
preds = model.predict(X_test_s)
print(classification_report(y_test, preds,
target_names=["benign", "malignant"]))
cm = confusion_matrix(y_test, preds)
plt.figure(figsize=(5, 4))
sns.heatmap(cm, annot=True, fmt="d", cmap="Blues",
xticklabels=["benign", "malignant"],
yticklabels=["benign", "malignant"])
plt.xlabel("Predicted label"); plt.ylabel("True label")
plt.title("Confusion matrix — Random Forest")
plt.savefig("confusion_matrix.png", dpi=150, bbox_inches="tight")
plt.show()
The confusion matrix shows which errors the model makes (e.g., malignant cases missed as benign — the dangerous error). Discussing the error pattern, not just the accuracy number, is what lifts a results section from student-level to publication-level.
A practical starter set, in the order to try them:
| Model | Code | When it shines | Watch out for |
|---|---|---|---|
| Logistic regression | LogisticRegression(max_iter=1000) |
Fast, interpretable baseline; surprisingly strong on clean tabular data | Needs scaled features; linear boundaries only |
| Random forest | RandomForestClassifier(random_state=42) |
Strong default; handles unscaled data and mixed types well | Can overfit tiny datasets; slower to train |
| SVM | SVC() |
Excellent on small, clean datasets | Needs scaling; slow on large data |
| k-NN | KNeighborsClassifier() |
Intuitive baseline; no training assumptions | Needs scaling; slow prediction on big data |
Start with logistic regression as your baseline — it trains in seconds and its coefficients tell you which features matter. Then try random forest, which wins outright on many tabular problems with almost no tuning. Report both; the comparison is the point [5].
Everything in this chapter applies to predicting numbers (house prices, temperatures) instead of classes — only the models and metrics change:
from sklearn.linear_model import LinearRegression
from sklearn.metrics import mean_squared_error, r2_score
reg = LinearRegression().fit(X_train, y_train)
preds = reg.predict(X_test)
print("RMSE:", mean_squared_error(y_test, preds) ** 0.5) # error in original units
print("R²:", r2_score(y_test, preds)) # 1.0 = perfect, 0 = as good as guessing the mean
RMSE (root mean squared error) is in the target's units ("off by 4.2 degrees on average") — interpretable. R² tells you how much better than a naive mean-guess you are. Report both.
Retraining every time you want predictions is wasteful and unreproducible. Save the fitted pipeline:
import joblib
joblib.dump(pipe, "models/random_forest.joblib") # save
pipe = joblib.load("models/random_forest.joblib") # load later — no retraining
Save the pipeline (preprocessing + model), not just the model — otherwise you must perfectly recreate the preprocessing at prediction time, and any mismatch is a silent bug. Version your model files (random_forest_v1.joblib) and note which one produced each result.
Before celebrating any accuracy, beat the dumbest possible model — scikit-learn's DummyClassifier makes this one line:
from sklearn.dummy import DummyClassifier
dummy = DummyClassifier(strategy="most_frequent") # always predicts the majority class
dummy.fit(X_train, y_train)
print("Dummy accuracy:", dummy.score(X_test, y_test))
If your tuned model barely beats the dummy, the features carry little signal — a finding worth knowing before you write a paper around the result. Other strategies ("stratified" for random guessing at class rates, "uniform") give you the full floor-to-ceiling picture. Every results table should implicitly answer "better than what?" — the dummy is the "what."
Accuracy needs a threshold; ROC curves show performance across all thresholds, and AUC summarizes the curve in one number (0.5 = random, 1.0 = perfect):
from sklearn.metrics import RocCurveDisplay
RocCurveDisplay.from_estimator(pipe, X_test_s, y_test)
plt.savefig("roc_curve.png", dpi=150, bbox_inches="tight")
AUC is the standard metric in medical and imbalanced problems because it is threshold-independent — report it alongside precision/recall whenever false positives and false negatives have different costs. Comparing models? Overlay their ROC curves in one plot; the differences are visible at a glance.
For your research: Your methodology section should read like a recipe anyone can follow: dataset and split ratio, preprocessing (with the no-leakage note), models compared with hyperparameter grids, evaluation metrics, and cross-validation scheme. Each maps to code you already wrote in this chapter. And cite scikit-learn properly — it is a published paper [1], and citing your tools is both honest and expected. A reviewer who sees "implemented with scikit-learn [1]" knows exactly what
LogisticRegressionmeans without you explaining it.
Key takeaways
- The scikit-learn workflow: prepare → split → choose → fit → predict → measure. Same interface for every model.
- Always split train/test (use stratify), never fit preprocessing on test data, and report precision/recall/F1 — not just accuracy.
- Cross-validation gives honest, stable numbers; grid search tunes hyperparameters reproducibly.
- Compare multiple models including a simple baseline, and analyze your errors with a confusion matrix.
Tutorials use clean built-in datasets. Research uses the messy real thing: a CSV from a hospital, sensor logs from a field deployment, a downloaded archive with a README written in 2014. This chapter is the bridge from toy data to data you can actually publish on.
Good starting points: UCI Machine Learning Repository (classic tabular datasets), Kaggle (competitions with real-world messiness), Hugging Face Datasets (modern, well-documented), and domain archives (e.g., PhysioNet for medical signals). For your first paper, prefer a dataset that is public, labeled, cited by other papers (so you have related work and baselines), and small enough to iterate quickly.
Before committing to a dataset, check: is there a clear label column? How many samples? Is it imbalanced? What license governs it (can you publish results on it)? Has anyone published on it — those papers are your baselines and your citations.
import pandas as pd
df = pd.read_csv("data.csv") # comma-separated
df = pd.read_csv("data.tsv", sep="\t") # tab-separated
df = pd.read_excel("data.xlsx", sheet_name=0) # Excel (needs openpyxl)
df = pd.read_json("data.json") # JSON records
# Peek inside archives without unpacking everything
import zipfile
with zipfile.ZipFile("archive.zip") as z:
print(z.namelist())
Encoding errors (UnicodeDecodeError) are the classic CSV headache — try pd.read_csv("data.csv", encoding="latin1"). Wrong separators (semicolons in European exports) need sep=";". When read_csv produces one giant column, the separator is almost always the culprit.
Run this on every new dataset before any modeling — it takes five minutes and prevents five hours of confusion:
print(df.shape) # how much data?
print(df.head()) # what does it look like?
print(df.dtypes) # are types correct?
print(df.isna().sum()) # what's missing?
print(df.duplicated().sum()) # any duplicate rows?
print(df["label"].value_counts(normalize=True)) # class balance?
print(df.describe()) # ranges, outliers?
Write the answers down — they become your paper's dataset description paragraph. "The dataset contains 4,177 samples with 12 features; 6.1% of cholesterol values were missing; the positive class is 34% of samples."
The basic split you know:
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, random_state=42, stratify=y)
Three refinements for real research:
python
X_tr, X_tmp, y_tr, y_tmp = train_test_split(X, y, test_size=0.3, random_state=42, stratify=y)
X_val, X_test, y_val, y_test = train_test_split(X_tmp, y_tmp, test_size=0.5, random_state=42, stratify=y_tmp)GroupKFold with patient IDs as groups.This pulls together Chapters 5 and 8 — a realistic pipeline you can adapt to your own data:
import pandas as pd
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
# 1. Load and inspect
df = pd.read_csv("crop_sensors.csv", encoding="latin1")
print("Raw shape:", df.shape)
print(df.isna().sum())
# 2. Clean
df = df.drop_duplicates()
df["label"] = df["label"].str.strip().str.lower() # "Healthy " -> "healthy"
for col in ["temperature", "humidity", "soil_moisture"]:
df[col] = pd.to_numeric(df[col], errors="coerce")
df[col] = df[col].fillna(df[col].median())
# 3. Encode the label as integers
label_map = {"healthy": 0, "diseased": 1}
df = df[df["label"].isin(label_map)] # drop unexpected labels
df["label"] = df["label"].map(label_map)
# 4. Split (stratified) and scale (fit on train only)
X = df[["temperature", "humidity", "soil_moisture"]].to_numpy()
y = df["label"].to_numpy()
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, random_state=42, stratify=y)
scaler = StandardScaler()
X_train = scaler.fit_transform(X_train)
X_test = scaler.transform(X_test)
print("Train:", X_train.shape, "Test:", X_test.shape)
print("Class balance (train):", pd.Series(y_train).value_counts(normalize=True).round(2).to_dict())
Every step here answers a reviewer question in advance: duplicates handled, text normalized, bad values coerced, unexpected labels dropped, split stratified, scaling leak-free.
Many real problems are imbalanced — 95% healthy, 5% diseased. Three things change:
class_weight="balanced", which penalizes mistakes on the rare class more heavily — often the single most effective one-line fix.Deeper techniques (resampling, SMOTE, threshold tuning) belong to Book 4 — but recognizing imbalance on day one already puts you ahead of most beginners.
"Datasets evolve" — a collaborator sends a corrected CSV, you add 200 samples, and suddenly last month's numbers are unreproducible. Minimum viable discipline:
sensors_v2.csv), and the cleaning script documents what changed.DATA_README.md.The principle mirrors code versioning: any result must be traceable to an exact, recoverable data snapshot.
Print this checklist and tape it to your wall until it is muscle memory. Every item prevents a bug or a reviewer question.
If your data has a time dimension (sensor readings by date, yearly records), random splits are wrong: training on 2025 data to "predict" 2023 is leakage through time. Split chronologically — train on the past, test on the future:
df = df.sort_values("date")
split = int(len(df) * 0.8)
train, test = df.iloc[:split], df.iloc[split:]
For cross-validation on time series, use TimeSeriesSplit, which only ever trains on earlier folds. Ask yourself for every dataset: "could a real deployment see the future?" If the answer involves time, split by time.
The subtlest leakage: a feature derived from the label. Examples: a "total_cost" column that includes the treatment you are trying to predict, or a "days_in_hospital" feature recorded after the diagnosis. The model scores spectacularly — because it is reading the answer key. Defense: for every feature, ask "would this be known at prediction time in the real world?" If not, drop it. Spectacular accuracy on a first attempt should make you suspicious, not happy — verify the features before celebrating.
Every serious dataset ships with documentation — a README, a data dictionary, or a "dataset card" describing each column, its units, and how it was collected. Read it before the inspection ritual, not after. It answers questions no code can: what does a 0 in the label column mean? Are blood pressures in mmHg? Were samples collected in one hospital or five? Was the data anonymized, and what was removed?
Documentation also carries the license — the legal terms of use. Common ones: CC-BY (use freely, credit the creators), non-commercial variants (no commercial use), and custom academic-use agreements requiring you to request access. Your paper's dataset section should name the license: it tells readers they can legally build on your work, and it protects you. When documentation is missing, say so in your paper and describe your best understanding — honesty about ambiguity beats silent guessing.
For your research: Document your dataset like it will be audited — because it might be. Keep a
DATA_README.mdnext to your data noting: source and download date, license, number of samples/features, the cleaning decisions above, and the exact split (includingrandom_state). If the data is private, describe it precisely anyway and share the pipeline on synthetic or sample data. Reproducibility reviewers forgive private data; they do not forgive undocumented data.
Key takeaways
- Choose datasets that are public, labeled, cited by others, and legally usable.
- The inspection ritual (shape, head, dtypes, isna, duplicates, class balance) comes before any modeling.
- Use stratified splits; go three-way (train/val/test) when tuning; use group-aware splits when samples are not independent.
- Clean deliberately, encode labels, scale without leakage — and write it all down.
Machine learning code fails in specific, recurring ways. This chapter is a field guide to the failures you will meet — shape mismatches, silent NaNs, data leakage, and slow code — and the systematic way to fix each one. Debugging is not a talent; it is a procedure.
When Python crashes, it prints a traceback — read it bottom-up. The last line names the error; the lines above show the call chain with file names and line numbers. Your code is usually the lowest frame that mentions your file:
Traceback (most recent call last):
File "experiment.py", line 42, in <module>
model.fit(X_train, y_train)
...
ValueError: X has 100 features, but LogisticRegression is expecting 30 features
Translation: the model was trained on 30 features but is now seeing 100 — you probably fit the scaler or the model on differently-processed data. The fix is at line 42's neighborhood: check what X_train looked like at fit time versus now.
1. Shape mismatches. The most common error in ML code. A model expecting (n_samples, n_features) gets (n_samples,) (forgot to keep 2-D: use X[:, np.newaxis] or df[["col"]]) or features counts differ between train and test.
# Defense: assert shapes at pipeline boundaries
assert X_train.shape[1] == X_test.shape[1], "feature count changed!"
assert X_train.shape[0] == y_train.shape[0], "samples != labels!"
2. NaN poisoning. One NaN in your features, and scikit-learn refuses to fit (Input contains NaN). Worse, NaNs in metrics silently produce nan scores that look like code ran fine. Defense: assert not np.isnan(X).any() after cleaning, and check df.isna().sum() before training.
3. Data leakage. Test information sneaks into training — scaling on all data, imputing with full-data statistics, or (the sneakiest) duplicate rows across train and test. Symptom: suspiciously high test accuracy that collapses on truly new data. Defense: do all preprocessing inside the train split, and consider Pipeline (below).
4. The random-but-not-random bug. Forgetting random_state means every run gives different results — you cannot tell whether a change helped or you just got lucky. Defense: set random_state everywhere (splits, models, CV shuffles) until final reporting.
5. Silent wrongness. Code runs, numbers come out, but they are meaningless — e.g., evaluating on training data, or labels accidentally shuffled relative to features. Defense: sanity checks. A model should beat a dumb baseline (predict the majority class); if it does not, something is wrong. Train accuracy far above test accuracy means overfitting [5].
A Pipeline chains preprocessing and modeling into one object, so transformations are fit on train and applied to test automatically — leakage becomes structurally impossible:
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.impute import SimpleImputer
from sklearn.linear_model import LogisticRegression
pipe = Pipeline([
("impute", SimpleImputer(strategy="mean")), # fill missing values
("scale", StandardScaler()), # standardize
("model", LogisticRegression(max_iter=1000)),
])
pipe.fit(X_train, y_train) # impute+scale fit on train only
print("Test accuracy:", pipe.score(X_test, y_test))
Pipelines also work with GridSearchCV (param_grid={"model__C": [0.1, 1, 10]}) and can be saved as one object. For any paper experiment, prefer pipelines over hand-rolled preprocessing sequences.
When an experiment is slow, do not guess — measure. In notebooks, the %timeit magic times a line; %%timeit times a whole cell:
%timeit model.fit(X_train, y_train)
# 1.24 s ± 38 ms per loop
For whole scripts, cProfile shows where time goes:
python -m cProfile -s cumulative slow_script.py | head -30
The output ranks functions by total time. In practice, 90% of slowness in student code comes from three places:
Optimization order: make it correct, make it clear, then make it fast — and only the part cProfile points at. Premature optimization of clear code is how readable experiments become undebuggable ones.
Here is a deliberately broken script. Read the errors, apply the fixes — this is the debugging procedure in miniature:
import numpy as np
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LogisticRegression
rng = np.random.default_rng(0)
X = rng.normal(size=(200, 5))
X[::20, 0] = np.nan # BUG 1: NaNs injected
y = (X[:, 0] + X[:, 1] > 0).astype(int)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2)
model = LogisticRegression()
model.fit(X_train, y_train) # ValueError: Input X contains NaN
Fix 1 — handle NaNs with an imputer (or drop the rows):
from sklearn.impute import SimpleImputer
X_train = SimpleImputer(strategy="mean").fit_transform(X_train)
X_test = SimpleImputer(strategy="mean").fit_transform(X_test) # BUG 2: leakage!
Fix 2 — no leakage: the imputer must be fit on train, then applied to test (or better, use a Pipeline as above):
imp = SimpleImputer(strategy="mean")
X_train = imp.fit_transform(X_train)
X_test = imp.transform(X_test) # transform only — correct
Fix 3 — sanity check the result:
model.fit(X_train, y_train)
train_acc = model.score(X_train, y_train)
test_acc = model.score(X_test, y_test)
print(f"train={train_acc:.3f} test={test_acc:.3f}")
baseline = max(np.bincount(y_test)) / len(y_test)
print(f"majority-class baseline={baseline:.3f}")
assert test_acc > baseline, "model worse than dumb baseline — investigate!"
The pattern: reproduce the error, read the traceback bottom-up, fix at the source (not the symptom), add an assertion so it can never silently return, and sanity-check the numbers against a baseline.
Scattering print statements works, but Python's built-in debugger is faster for tricky bugs. Put breakpoint() anywhere in your code — when execution reaches it, you get an interactive prompt:
def train_fold(X_train, y_train):
breakpoint() # execution pauses here
model.fit(X_train, y_train)
At the (Pdb) prompt: p X_train.shape prints a value, n runs the next line, c continues running, q quits. Inspecting the actual shapes and values at the failure point beats guessing from the outside. In notebooks, the same idea works with the graphical debugger (the bug icon in JupyterLab), which lets you step through cells visually.
| Error message | Meaning | Fix |
|---|---|---|
Expected 2D array, got 1D array |
Model got shape (n,) instead of (n, 1) |
X.reshape(-1, 1) or select with double brackets df[["col"]] |
Input X contains NaN |
Missing values reached the model | Impute or drop before fitting (Chapter 5) |
X has N features, but model is expecting M |
Train/test feature mismatch | Check one-hot columns match; use a Pipeline |
y contains previously unseen labels |
Test has a class the encoder never saw | Fit encoders on train only; handle unknowns explicitly |
n_splits=5 cannot be greater than the number of members in each class |
A class has fewer than 5 samples | Reduce cv, or collect more data for the rare class |
ConvergenceWarning: lbfgs failed to converge |
Optimizer did not finish (logistic regression) | Scale features; increase max_iter |
Copy the exact message into a search engine when stuck — for scikit-learn, someone has asked your question before, usually with the accepted answer explaining why.
Experiment scripts outgrow print. Python's logging module writes timestamped messages to both console and file — a permanent experiment record:
import logging
logging.basicConfig(filename="results/experiment.log",
level=logging.INFO,
format="%(asctime)s %(message)s")
logging.info("Starting training: n_estimators=200, seed=42")
logging.info(f"Test accuracy: {acc:.4f}")
The log file becomes part of your results folder: six months later, experiment.log tells you exactly what ran and what it produced, no memory required.
When an experiment fails on 100,000 rows, do not debug on 100,000 rows. Slice the first 10–100 rows and reproduce the failure there — it runs instantly and the data fits on your screen:
X_small, y_small = X_train[:50], y_train[:50] # debug here first
If the bug reproduces on 50 rows, you iterate in seconds. If it does not, the bug is scale-dependent (memory, a rare value) — and knowing that narrows the search enormously. Professional debugging is mostly about shrinking the problem until it is obvious.
Chapter 9 introduced shape assertions; take the idea further — assert your assumptions at every pipeline stage:
assert X_train.shape[0] > 100, "suspiciously little training data"
assert set(np.unique(y_train)) == {0, 1}, "labels changed unexpectedly"
assert not np.isnan(X_train).any(), "NaNs leaked into training"
assert (X_train.std(axis=0) > 0).all(), "a feature has zero variance"
Each assertion is a contract: "if this breaks, stop immediately and tell me." They cost nothing at runtime, document your expectations better than comments, and turn silent corruption into loud, early failures — exactly what you want before numbers reach a paper.
For your research: Keep a debugging log during experiments — one line per bug: symptom, cause, fix. It feels like overhead, but it becomes gold twice: first, when the same bug returns three weeks later and you have the fix written down; second, when you write the paper's "limitations and threats to validity" section, because you will actually know what you checked. Reviewers respect authors who can say "we verified no leakage by..." — and now you can, with the assertions to prove it.
Key takeaways
- Read tracebacks bottom-up; your bug is in the lowest frame mentioning your file.
- The five classic ML bugs: shape mismatches, NaN poisoning, data leakage, missing random states, silent wrongness.
- Use Pipeline to make leakage structurally impossible; assert shapes and NaN-freedom at boundaries.
- Profile before optimizing (%timeit, cProfile); vectorize loops and hoist repeated work out of loops.
- Sanity-check every result against a dumb baseline.
Reproducibility means someone else — or you, six months later — can rerun your code and get the same numbers. Journals increasingly require it, supervisors expect it, and it is the difference between "a result" and "a scientific result." The good news: reproducibility in Python is mostly four habits, and this chapter makes them automatic.
ML is full of randomness: train/test splits shuffle data, models like random forests sample randomly, cross-validation shuffles folds. Without fixed seeds, every run differs and you cannot tell improvement from luck.
import random
import numpy as np
SEED = 42
random.seed(SEED)
np.random.seed(SEED)
# then pass random_state=SEED to every sklearn splitter and model
Set one SEED constant at the top of every experiment script and thread it through everything: train_test_split(..., random_state=SEED), RandomForestClassifier(random_state=SEED), KFold(shuffle=True, random_state=SEED). One constant, changed in one place, controls the whole experiment.
Honest caveat: seeds make runs repeatable, not truths universal. A result that only holds for one seed is fragile — strong papers check key results across 3–5 seeds and report mean ± std. Seed 42 for development; multiple seeds for the final numbers.
Chapter 2's virtual environments are a reproducibility tool, not just convenience: they guarantee the experiment runs against a known set of packages, not whatever happens to be installed globally. One environment per project, recorded in the project README.
Libraries change. Code written against scikit-learn 1.1 can behave differently — or break — under 1.5. Pinning records exact versions so the environment can be rebuilt identically:
pip freeze > requirements.txt
produces lines like:
numpy==1.26.4
pandas==2.2.1
scikit-learn==1.4.2
Anyone rebuilds your environment with pip install -r requirements.txt. Commit requirements.txt to version control (Chapter 11) and regenerate it when you add packages. For publication, this file is part of the artifact — some venues ask for it explicitly.
Lighter alternative: hand-write only direct dependencies (scikit-learn==1.4.2) and let pip resolve the rest. Full pip freeze output is more exact but noisier; either is infinitely better than nothing.
A README.md (or experiment log) per project should answer: what Python version, what OS, how to recreate the environment, what command runs the experiment, and what the expected output is:
## Environment
- Python 3.11.9, Ubuntu 22.04
- Create env: python -m venv .venv && source .venv/bin/activate
- Install: pip install -r requirements.txt
## Reproduce results
python experiment.py --seed 42
# Expected: test accuracy 0.930 ± 0.008 (see results/summary.csv)
Add python --version and the pip freeze output to an environment.txt at project start (Chapter 2's advice) and you have covered nearly everything a reproducibility checklist asks for.
| It guarantees | It does not guarantee |
|---|---|
| Same code + same data + same environment → same numbers | Same numbers on a different OS or library version |
| You can rerun your own work months later | Your result generalizes to new data (that is validity, a separate question) |
| Reviewers can verify your claims | Bit-identical results on GPUs (floating-point nondeterminism) |
For GPU deep learning (Books 5–6), add torch.manual_seed(SEED) and accept that exact bit-reproducibility needs extra flags [3], [4]. For this book's CPU-based classical ML, seeds plus pinned environments give you full rerun reproducibility.
Run this at the end of every experiment — it captures the environment facts your paper's reproducibility statement needs:
import sys, platform, random
import numpy as np
import sklearn, pandas
SEED = 42
random.seed(SEED); np.random.seed(SEED)
report = {
"python": sys.version.split()[0],
"os": f"{platform.system()} {platform.release()}",
"numpy": np.__version__,
"pandas": pd.__version__ if False else pandas.__version__,
"sklearn": sklearn.__version__,
"seed": SEED,
}
for k, v in report.items():
print(f"{k:10s}: {v}")
with open("reproducibility.txt", "w") as f:
for k, v in report.items():
f.write(f"{k}={v}\n")
print("\nSaved to reproducibility.txt — attach it to your results.")
Pair this with pip freeze > requirements.txt and your random_state=SEED discipline, and any experiment in this book becomes rerunnable by anyone, including future-you.
It will happen: you rerun code and get different numbers. Work through this list in order — the culprit is almost always one of these:
random_state/seed set in every random operation, including any new code you added? One unseeded shuffle is enough.git diff / git status — did you change something and forget? Uncommitted edits are the classic cause.pip freeze diffed against the recorded requirements.txt.Fix the cause, re-record the environment, and note the incident in your experiment log. Each investigation makes the next one faster.
A checksum is a fingerprint of a file — if one byte changes, the fingerprint changes:
sha256sum data/raw/sensors.csv
# 9f2c...a41b data/raw/sensors.csv
Record checksums in your DATA_README.md. When results shift mysteriously, recompute: same checksum means the data is innocent, and you look at code and environment instead. In Python: hashlib.sha256(open("file.csv","rb").read()).hexdigest().
This book's CPU-based classical ML is fully reproducible with seeds plus pinned environments. Deep learning on GPUs (Books 5–6) is harder: GPU operations can be nondeterministic, and libraries need extra flags (torch.use_deterministic_algorithms(True)) plus seeded data-loader workers [3], [4]. Know this now so you are not surprised later: for classical ML, demand exact reruns; for GPU deep learning, expect tiny floating-point wobbles and report means across seeds.
pip freeze > requirements.txt (pip/venv workflow) and conda env export > environment.yml (conda workflow) solve the same problem — pick the one matching your setup. A hand-maintained requirements.txt listing only direct dependencies is more readable:
numpy==1.26.4
pandas==2.2.1
scikit-learn==1.4.2
matplotlib==3.8.2
seaborn==0.13.2
Five lines a human can audit beats fifty lines of transitive dependencies. Either way, the file is committed to version control and regenerated whenever dependencies change — it is as much a part of the experiment as the code.
When a paper says "code available at github.com/...", links rot — repositories get renamed, accounts lapse. Zenodo (zenodo.org) archives a snapshot of your GitHub release and issues a permanent DOI, the same identifier system journals use. Many venues now encourage or require it. The workflow: tag a release on GitHub (v1.0-paper), connect the repo to Zenodo once, and cite the DOI in your paper. Your future self, trying to find "the code from that 2026 paper," will be grateful.
Software versions are half the environment; hardware is the other half. Runtime, memory limits, and CPU/GPU models explain why your experiment took 20 minutes and whether someone else can rerun it at all:
import platform, multiprocessing
print("CPU:", platform.processor() or platform.machine())
print("Cores:", multiprocessing.cpu_count())
# GPU (if torch installed): torch.cuda.get_device_name(0)
Add these to reproducibility.txt alongside the software versions, plus the wall-clock runtime of the experiment ("training took ~14 min on 8 CPU cores"). A result that needs 64 GB of RAM is not reproducible on a laptop — saying so upfront is part of honest reporting, and it is exactly the detail reviewers ask about when they cannot rerun your code.
For your research: Write the reproducibility paragraph of your paper while you run the experiments, not after. One honest paragraph — "Experiments used Python 3.11 with scikit-learn 1.4 [1]; random seeds were fixed (seed 42) for all splits and models; the environment is specified in requirements.txt; code and the train/test split indices are available at [repository link]" — preempts the most common reviewer objection to student papers. If your institution requires thesis submission of code, this chapter's four habits are exactly what the examiners will look for.
Key takeaways
- Four habits: seed all randomness, isolate environments, pin dependencies, document the run.
- One SEED constant threaded through every splitter and model; verify final results across multiple seeds.
- pip freeze > requirements.txt + a README with the exact rerun command is the publication standard.
- Reproducibility means rerunnable and verifiable — it is necessary for science, distinct from generalization.
Notebooks are where experiments are born; they are not where finished research code lives. At some point — when an experiment works and you need to rerun it reliably, share it, or build on it — working code must graduate into scripts and a project layout. This chapter shows the graduation path without the enterprise-software ceremony.
Stay in the notebook while you are exploring: trying plots, testing ideas, understanding data. Move to scripts when code is working and will be rerun: the final data-cleaning pipeline, the training script, the evaluation that produces the paper's tables. The rule: notebooks explore, scripts reproduce.
The move itself is mechanical: copy working cells into a .py file, in top-to-bottom order, replacing the last cell's informal prints with saved outputs:
# train.py
import pandas as pd
from sklearn.model_selection import train_test_split
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import accuracy_score
SEED = 42
def main():
df = pd.read_csv("data/crop_sensors_clean.csv")
X = df[["temperature", "humidity", "soil_moisture"]].to_numpy()
y = df["label"].to_numpy()
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, random_state=SEED, stratify=y)
model = RandomForestClassifier(random_state=SEED)
model.fit(X_train, y_train)
acc = accuracy_score(y_test, model.predict(X_test))
print(f"Test accuracy: {acc:.3f}")
pd.DataFrame({"metric": ["accuracy"], "value": [acc]}).to_csv(
"results/summary.csv", index=False)
if __name__ == "__main__":
main()
Run with python train.py — no kernel state, no out-of-order cells, same result every time. The if __name__ == "__main__": guard lets you also import train elsewhere without rerunning.
You do not need a complex structure. This one handles everything from a course project to a thesis:
my-research/
├── README.md # what this is, how to reproduce it
├── requirements.txt # pinned dependencies
├── data/
│ ├── raw/ # originals — never modified
│ └── processed/ # cleaned outputs of clean_data.py
├── notebooks/
│ └── 01-exploration.ipynb
├── src/
│ ├── clean_data.py # raw -> processed
│ ├── train.py # processed -> model + metrics
│ └── evaluate.py # model -> figures and tables
├── results/
│ ├── summary.csv
│ └── figures/
└── reproducibility.txt # environment facts (Chapter 10)
Principles behind it: raw data is read-only; every transformation is a script that can be rerun; notebooks stay in notebooks/ for exploration only; results are files, not screenshots. When your supervisor asks "how did you get this number?", the answer is a script path and a command, not a notebook you hope was run in order.
When two scripts need the same logic (say, load_and_clean()), put it in one module and import it:
# src/data_utils.py
import pandas as pd
def load_and_clean(path):
"""Load sensor CSV, clean it, return (X, y)."""
df = pd.read_csv(path)
...
return X, y
# src/train.py
from data_utils import load_and_clean # run from src/ directory
X, y = load_and_clean("../data/processed/sensors.csv")
One definition, used everywhere — fix a bug once, fixed everywhere. This is the moment copy-paste coding ends and maintainable research code begins.
Git records every change to your code with a message, so you can always answer "what did I change last Tuesday?" and undo mistakes:
git init
git add src/ README.md requirements.txt
git commit -m "Add training script with random forest baseline"
Commit code, not data or environments: add data/, .venv/, and *.ipynb outputs to .gitignore (notebooks' outputs bloat repositories; commit the code cells, or clear outputs before committing). Commit message habit: what changed and why, in one line. Push to GitHub/GitLab for backup and sharing — a repository link is what goes in your paper.
You do not need advanced Git. add, commit, log, and one remote are enough for 95% of research work. The goal is not Git mastery; it is never losing working code again.
Take the classifier experiment from Chapter 7 and graduate it:
# src/train_classifier.py — the reproducible version of Chapter 7
import pandas as pd
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import train_test_split, cross_val_score
from sklearn.preprocessing import StandardScaler
from sklearn.pipeline import Pipeline
from sklearn.ensemble import RandomForestClassifier
SEED = 42
def main():
X, y = load_breast_cancer(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, random_state=SEED, stratify=y)
pipe = Pipeline([
("scale", StandardScaler()),
("model", RandomForestClassifier(n_estimators=100, random_state=SEED)),
])
cv = cross_val_score(pipe, X_train, y_train, cv=5)
pipe.fit(X_train, y_train)
test_acc = pipe.score(X_test, y_test)
summary = pd.DataFrame({
"metric": ["cv_mean", "cv_std", "test_accuracy"],
"value": [round(cv.mean(), 4), round(cv.std(), 4), round(test_acc, 4)],
})
summary.to_csv("results/summary.csv", index=False)
print(summary.to_string(index=False))
print("\nRerun with: python src/train_classifier.py")
if __name__ == "__main__":
main()
Compare with the notebook version: same logic, but now it is a single command, writes its results to a file, uses a Pipeline (no leakage), and carries its seed. This is the artifact a paper's "code available at..." link points to.
You do not need to memorize PEP 8 (Python's style guide). Follow these five habits and your code will look professional:
test_accuracy, not ta; load_sensor_data(), not f().acc = 0.93, not acc=0.93.pip install black, then black src/) — it fixes style automatically, ending all debates.Style is not vanity: consistent code is faster to debug and signals to collaborators (and examiners) that you are careful.
A good README fits on one screen and answers four questions:
# Project title — one line on what it does and why
## Quickstart
<the exact commands to reproduce results>
## Results
<the headline numbers, in a small table>
## Environment
Python X.Y, key packages (full list in requirements.txt)
## Contact / notes
<anything unusual: data access, expected runtime>
Write it for a tired reviewer at midnight: commands they can copy-paste, numbers they can verify, no prose they must wade through.
Push your repository to GitHub or GitLab — it backs up your work and gives you the link your paper needs. Two final rules:
LICENSE file with its text.git status before every commit.Chapter 3 introduced docstrings; here is the format worth standardizing on — what goes in, what comes out, and one example:
def train_model(X_train, y_train, model_name="rf", seed=42):
"""Train a classifier and return the fitted pipeline.
Args:
X_train: Feature array, shape (n_samples, n_features).
y_train: Integer labels, shape (n_samples,).
model_name: One of "lr", "svm", "rf".
seed: Random seed for reproducibility.
Returns:
Fitted sklearn Pipeline (scaler + classifier).
Example:
>>> pipe = train_model(X_train, y_train, model_name="svm")
"""
Six months later, this docstring — not your memory — is how you reuse the function. Tools can even test the Example lines automatically. Write docstrings for every function in src/; skip them only for throwaway notebook cells.
When graduating notebook code to a script, run through this list:
logging callsThe last item is the most skipped and most valuable: # Trains 3 classifiers on cleaned sensor data; writes results/summary.csv.
For your research: Examiners and reviewers increasingly look at your repository, not just your PDF. A repository with a README, requirements.txt, a sane layout, and scripts that regenerate the paper's tables signals a serious researcher — it is often the difference between "accept with minor revisions" and a painful round of "please clarify how results were obtained." Build the layout in Chapter 12's capstone and reuse it for every project after; it becomes your personal research template.
Key takeaways
- Notebooks explore; scripts reproduce. Graduate working code into .py files with a main() and saved outputs.
- Use the simple layout: README, requirements, data/raw (read-only), src/, notebooks/, results/.
- Share logic via imported modules, not copy-paste; track code with Git (commit code, not data or environments).
- Every number in your paper should be regenerable by one script command.
Everything in this book converges here: a complete mini-project, from raw data to documented results, built the way a publishable experiment is built. Follow it end to end, then adapt the template to your own research problem.
Problem statement (one sentence, Chapter 5's checklist): Can low-cost temperature, humidity, and soil-moisture readings distinguish healthy from diseased crop plots well enough to be useful for early warning?
Plan: 1. Generate a realistic synthetic sensor dataset (in real work: your collected data). 2. Explore it with plots (Chapter 6). 3. Clean it with a script (Chapters 5, 8, 11). 4. Train and compare three classifiers with cross-validation (Chapter 7). 5. Analyze errors with a confusion matrix. 6. Save every result and document the project like a paper artifact (Chapters 10, 11).
# notebooks/01-exploration.ipynb
import numpy as np, pandas as pd
import matplotlib.pyplot as plt, seaborn as sns
rng = np.random.default_rng(7)
n = 1200
temp = rng.normal(28, 4, n)
humidity = rng.normal(65, 12, n)
moisture = rng.normal(40, 10, n)
# Diseased plots run hotter and drier — a learnable pattern with noise
disease_score = 0.5*(temp-28) - 0.3*(moisture-40) + rng.normal(0, 3, n)
label = (disease_score > 2).astype(int)
df = pd.DataFrame({"temperature": temp, "humidity": humidity,
"soil_moisture": moisture, "label": label})
df.to_csv("data/raw/sensors.csv", index=False)
print(df["label"].value_counts(normalize=True).round(3).to_dict())
sns.pairplot(df, hue="label", diag_kind="kde")
plt.savefig("results/figures/exploration.png", dpi=150, bbox_inches="tight")
The pairplot immediately shows temperature and soil moisture separating the classes — your first evidence, and the figure for the dataset section.
# src/clean_data.py
import pandas as pd
def main():
df = pd.read_csv("data/raw/sensors.csv")
assert df.isna().sum().sum() == 0, "unexpected missing values"
df = df.drop_duplicates().reset_index(drop=True)
df.to_csv("data/processed/sensors_clean.csv", index=False)
print("Cleaned shape:", df.shape)
if __name__ == "__main__":
main()
# src/train.py
import pandas as pd
from sklearn.model_selection import train_test_split, cross_val_score
from sklearn.preprocessing import StandardScaler
from sklearn.pipeline import Pipeline
from sklearn.linear_model import LogisticRegression
from sklearn.svm import SVC
from sklearn.ensemble import RandomForestClassifier
SEED = 42
def main():
df = pd.read_csv("data/processed/sensors_clean.csv")
X = df[["temperature", "humidity", "soil_moisture"]].to_numpy()
y = df["label"].to_numpy()
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, random_state=SEED, stratify=y)
models = {
"logistic_regression": LogisticRegression(max_iter=1000),
"svm": SVC(),
"random_forest": RandomForestClassifier(n_estimators=200, random_state=SEED),
}
rows = []
for name, m in models.items():
pipe = Pipeline([("scale", StandardScaler()), ("model", m)])
cv = cross_val_score(pipe, X_train, y_train, cv=5)
pipe.fit(X_train, y_train)
rows.append({"model": name,
"cv_mean": round(cv.mean(), 4),
"cv_std": round(cv.std(), 4),
"test_accuracy": round(pipe.score(X_test, y_test), 4)})
summary = pd.DataFrame(rows).sort_values("test_accuracy", ascending=False)
summary.to_csv("results/summary.csv", index=False)
print(summary.to_string(index=False))
if __name__ == "__main__":
main()
Run order: python src/clean_data.py → python src/train.py. Typical output:
model cv_mean cv_std test_accuracy
svm 0.9012 0.0183 0.9125
logistic_regression 0.8979 0.0161 0.9083
random_forest 0.8938 0.0214 0.9000
All three models land near 90% — and the honest conclusion is that on these three features, simple models perform comparably. That is a result, and it is exactly the kind of careful, well-documented comparison that makes a solid first paper.
Your README.md turns the project into an artifact a reviewer can actually use:
# Crop health classification from low-cost sensors
Classifies crop plots as healthy/diseased from temperature, humidity,
and soil-moisture readings.
## Reproduce
1. python -m venv .venv && source .venv/bin/activate
2. pip install -r requirements.txt
3. python src/clean_data.py
4. python src/train.py # writes results/summary.csv
## Results (seed 42, 5-fold CV, 80/20 stratified split)
| model | cv_mean | cv_std | test_accuracy |
|---|---|---|---|
| svm | 0.9012 | 0.0183 | 0.9125 |
| logistic_regression | 0.8979 | 0.0161 | 0.9083 |
| random_forest | 0.8938 | 0.0214 | 0.9000 |
Environment: Python 3.11, scikit-learn 1.4 (see reproducibility.txt).
How this maps to a paper: the README's reproduce steps become your methodology's implementation paragraph; results/summary.csv becomes the results table; the exploration figure becomes Figure 1; the problem statement becomes the abstract's first line. You did not just run an experiment — you produced a paper-ready package.
requirements.txt pinned; reproducibility.txt recordedThe capstone as written is complete — but research is iterative. Each upgrade below maps directly to a stronger results section:
GridSearchCV (Chapter 7) over max_depth and n_estimators. Report the grid and the winner — methodology sections love this sentence.summary.csv with precision, recall, and F1 from classification_report. On imbalanced data this is not optional.| Artifact you produced | Paper section | What to write |
|---|---|---|
| One-sentence problem statement | Abstract + Introduction | The gap: what is unknown, and why it matters |
| Exploration figure + dataset description | Data / Dataset section | Source, size, balance, cleaning decisions |
| README reproduce steps + scripts | Methodology | Exact pipeline: split, preprocessing, models, tuning, metrics |
summary.csv comparison table |
Results | Table + which model won and by how much |
| Confusion matrix + error discussion | Results / Discussion | Where the model fails and what that implies |
| Checklist + reproducibility.txt | Reproducibility statement | Seeds, environment, code availability |
Notice the pattern: nothing in the paper is invented at writing time. Every section points at an artifact that already exists. Writing becomes assembly, not creation — which is why researchers who document as they go finish papers faster.
When you present this work — to your supervisor, in a seminar, in the paper — lead with three things: the problem (one sentence), the headline number (best model, with uncertainty), and the limitation (synthetic data, three features, one domain). Leading with limitations sounds counterintuitive, but it is what credible researchers do: it shows you understand your result's boundaries, and it hands your audience the exact next experiment — which is often where your second paper comes from.
Students running this capstone pattern on their own data hit the same five walls — here they are, pre-solved:
encoding="latin1"), separator (sep=";"), and open the file in a text editor to see what it actually looks like. The file is always right; your assumption about it is wrong.value_counts), then try class_weight="balanced", then question whether your features carry any signal at all.pip freeze, Restart & Run All — in that order.If this mini-project grows into a thesis chapter, the structure writes itself:
results/ folder)Every section has a corresponding artifact from the capstone. That is the whole secret of productive research writing: the writing is done when the experiments are documented, not when you "start writing."
For your research: Your first paper does not need a novel algorithm — it needs a complete, honest, reproducible experiment, which is exactly what you just built. Swap the synthetic data for your real dataset, write the four paper sections this project already contains (problem, data, method, results), and you have a draft. Supervisors say yes to students who arrive with running code and documented numbers; this capstone is how you become that student. When you are ready, Book 4 (Data Preprocessing and Feature Engineering) will deepen the data skills this project relied on.
Key takeaways - A complete project: problem statement → data → exploration → cleaning script → training script → documented results. - Every paper section maps to an artifact you already produced: README → methodology, summary.csv → results table, figures → figures. - The publication-ready checklist is your pre-submission gate — run it before any draft goes to your supervisor.
[1] F. Pedregosa et al., "Scikit-learn: Machine learning in Python," J. Mach. Learn. Res., vol. 12, pp. 2825–2830, 2011.
[2] A. Géron, Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow, 2nd ed. Sebastopol, CA, USA: O'Reilly Media, 2019.
[3] A. Paszke et al., "PyTorch: An imperative style, high-performance deep learning library," in Proc. Adv. Neural Inf. Process. Syst., vol. 32, 2019.
[4] M. Abadi et al., "TensorFlow: Large-scale machine learning on heterogeneous systems," arXiv:1603.04467, 2016.
[5] T. M. Mitchell, Machine Learning. New York, NY, USA: McGraw-Hill, 1997.
[6] F. Chollet, Deep Learning with Python, 2nd ed. Shelter Island, NY, USA: Manning, 2021.
[7] C. M. Bishop, Pattern Recognition and Machine Learning. New York, NY, USA: Springer, 2006.
ai-env, and install numpy, pandas, matplotlib, seaborn, scikit-learn, and jupyter. Verify each with a one-line import check.describe_numbers(nums) that takes a list of floats and returns a dictionary with keys count, mean, min, max — using a list comprehension somewhere in your solution..mean(axis=0) and .std(axis=0).load_breast_cancer(as_frame=True)). Run the full inspection ritual and write a five-line dataset description paragraph from the output.Pipeline per model and rerun with 5-fold cross-validation. Do the rankings change compared to the single split? Write two sentences on what you conclude.src/train.py script with a main() function, a SEED constant, saved CSV results, and a README with rerun instructions. Confirm python src/train.py reproduces the notebook's numbers.End of Book 3. Next: Book 4 — Data Preprocessing and Feature Engineering.