← All 50 books
Fine-Tuning LLMs on Your Own Data cover

Book 17 of 50 · Free

Fine-Tuning LLMs on Your Own Data

29,556 words · 99 chapters · illustrated

Fine-Tuning LLMs on Your Own Data

Book 17 of 50 — AstolixGen Learning Series

Book cover: a neural network globe being adjusted by gears


About This Book

You have used large language models. You have written prompts, maybe built a small retrieval system, and now you are wondering: can I actually change what the model knows and how it behaves — using my own data? This book answers that question from the ground up.

Fine-tuning is the process of continuing a model's training on your own examples so it learns your task, your style, your domain, or your language. A few years ago this was the exclusive territory of labs with warehouse-sized GPU clusters. Today, with parameter-efficient methods like LoRA and QLoRA, a single student with a free Google Colab notebook can meaningfully adapt a multi-billion-parameter model on a laptop budget. That shift is the whole reason this book exists.

This book is written for researchers and publication students: MS and PhD candidates, and early AI researchers who need fine-tuning not as a party trick but as a research method — something you can run, measure, ablate, and report in a paper. Every chapter ends with a "For your research" box that translates the chapter's ideas into publishable practice: what to vary, what to measure, what to report, and what reviewers will ask about.

We assume you know Python, can use a terminal, and have trained at least a small neural network before (or have read Books 1–10 of this series, which build that foundation). We do not assume you have ever touched a GPU cluster, paid for cloud compute, or published a paper about large models. We explain every term when it first appears.

By the end of this book you will be able to:

Learning Objectives

  1. Decide whether your problem calls for prompting, retrieval-augmented generation (RAG), or fine-tuning — and defend that decision in writing.
  2. Explain what changes inside a model during fine-tuning, in terms a non-specialist can follow.
  3. Prepare a high-quality instruction dataset from your own data, and know why quality beats quantity.
  4. Compare full fine-tuning with parameter-efficient methods (LoRA, QLoRA) and choose between them for a given hardware budget.
  5. Set up the standard open-source tooling stack (Hugging Face Transformers, TRL, Axolotl) and run a first training job.
  6. Complete a full fine-tuning walkthrough on a single GPU or free Colab, from data loading to a usable adapter.
  7. Tune the hyperparameters that actually matter (learning rate, LoRA rank, epochs, batch size) with a principled, cheap procedure.
  8. Evaluate a fine-tuned model with honest before/after comparisons and avoid fooling yourself.
  9. Detect overfitting and catastrophic forgetting, and know the standard fixes for each.
  10. Budget compute costs realistically as a student, and plan experiments that fit your resources.
  11. Design a fine-tuning study around your own research problem, case-study style.
  12. Report fine-tuning experiments in a paper the way reviewers expect, with all the details that make work reproducible.

How to Use This Book

Chapters build on each other, but each chapter also stands alone well enough to be read as a reference. The chapters with heavy code (5, 6, 7) are best read next to a computer. The chapters on evaluation, forgetting, and reporting (8, 9, 12) are best read next to your research notes. The Learning Dashboard at the end is a quick-reference appendix: method comparison table, hyperparameter cheat sheet, cost estimation table, and troubleshooting matrix. Print it, or keep it open in a second tab while you train.

A note on honesty that will come up repeatedly: fine-tuning is easy to start and easy to do badly. The difference between a fine-tuned model that genuinely improved and one that only looks improved on your favorite examples is evaluation discipline. This book treats that discipline as the main skill, not a side note.


Table of Contents

  • Chapter 1: When to Fine-Tune vs Prompt vs RAG (A Decision Framework)
  • Chapter 2: How Fine-Tuning Works, Conceptually (What Changes in the Weights)
  • Chapter 3: Preparing Your Dataset: Instruction Format, Quality Beats Quantity
  • Chapter 4: Full Fine-Tuning vs PEFT: LoRA and QLoRA Explained
  • Chapter 5: Tooling: Hugging Face Transformers, TRL, Axolotl (Overview + Setup)
  • Chapter 6: Training on One GPU / Free Colab: A Complete Walkthrough
  • Chapter 7: Hyperparameters That Actually Matter (LR, Rank, Epochs, Batch Size)
  • Chapter 8: Evaluating Fine-Tuned Models (Before/After Comparisons)
  • Chapter 9: Overfitting, Catastrophic Forgetting, and How to Detect Them
  • Chapter 10: Compute Costs and Budgeting for Students
  • Chapter 11: Fine-Tuning for Your Research Problem (Case-Study Guidance)
  • Chapter 12: Reporting Fine-Tuning Experiments in Papers (What Reviewers Expect)
  • Learning Dashboard (quick-reference appendix)
  • References
  • Glossary
  • Practice Exercises

Chapter 1: When to Fine-Tune vs Prompt vs RAG (A Decision Framework)

You have a problem and a large language model. Before you spend a single rupee on GPUs or a single hour building a dataset, you need to answer one question: what is the cheapest technique that will actually solve this problem? There are three main candidates — prompting, retrieval-augmented generation (RAG), and fine-tuning — and choosing wrong costs you weeks. Choosing right can save you from training anything at all.

The Three Options in Plain Language

Prompting means you change the input to get better output. You write careful instructions, give a few examples inside the prompt (few-shot prompting), or ask the model to reason step by step (chain-of-thought). The model's weights — the billions of numbers that encode what it learned in training — do not change at all. You are steering an already-built car.

Retrieval-augmented generation (RAG) means you connect the model to an external knowledge source. When a user asks a question, your system searches your documents, retrieves the relevant chunks, and pastes them into the prompt so the model can answer from them. Again, the model's weights do not change. You have given the driver a map.

Fine-tuning means you continue training the model on your own examples, and its weights do change — permanently. You are rebuilding parts of the car itself.

This distinction — whether the weights change — is the dividing line for everything that follows. Prompting and RAG leave the model untouched; fine-tuning alters it. Altering the model is more powerful and more dangerous, more expensive, and harder to undo.

What Each Approach Is Good At

Think of it this way:

  • Prompting is best when the model already knows everything it needs, and your problem is just eliciting the right behavior. Formatting outputs, following instructions, translating, summarizing, writing in a certain tone — a good prompt usually solves these. Cost: almost zero. Time to try: minutes.
  • RAG is best when the model needs facts it was never trained on or facts that change. Company documentation, recent events, private databases, legal codes, product catalogs. If your question is "the model doesn't know this information," RAG is almost always the answer. Cost: moderate engineering effort, cheap inference.
  • Fine-tuning is best when you need to change how the model behaves — its style, its habits, its skills — across many inputs, without stuffing instructions into every prompt. It is also the right tool when the task requires internalized knowledge that would be awkward or impossible to retrieve at query time: a new language the model barely speaks, a specialized coding style, consistent domain reasoning patterns, or a classification task you need to run millions of times cheaply.

Here is the classic set of situations where fine-tuning genuinely wins:

  1. Behavioral consistency at scale. You need the model to always respond in a certain way — a particular tone, format, refusal policy, or domain persona — and prompt instructions keep getting ignored or diluted on long conversations.
  2. A skill, not a fact. The model needs to get better at something procedural: writing SQL in your company's dialect, diagnosing in your hospital's format, grading essays to your rubric. Skills are patterns in the weights; facts are retrievable.
  3. Latency and cost at inference. A fine-tuned small model can outperform a prompted giant on a narrow task, and run far cheaper. If you will serve millions of requests, fine-tuning a 7–8B model can be dramatically cheaper than prompting a frontier model through an API.
  4. Structured output reliability. Getting JSON with exactly the right schema, every time, from messy inputs. Prompting can do this; fine-tuning makes it nearly deterministic.
  5. Weak base capability you must strengthen. The base model is poor at your language, your domain's jargon, or your task format, and examples in the prompt only partially fix it.

And here is where fine-tuning is the wrong choice:

  1. The model needs current or private facts. Retraining every time your documents update is absurd; use RAG.
  2. You have 50 examples. Fine-tuning on tiny data mostly teaches the model to memorize those examples. Prompting with those 50 examples as few-shot demonstrations is usually better.
  3. The task is already solved by prompting. Always test the cheap option first. Many published "fine-tuning" results would have been matched by a better prompt.
  4. You need the model's behavior to vary per user or per query. A fine-tuned model is one fixed behavior; prompts and retrieval can adapt per request.

The Decision Framework

Work through these questions in order. Stop at the first one that gives you a definitive answer.

Step 1: Can the model already do this with a good prompt? Spend a day on prompt engineering first. Write clear instructions, add 3–10 diverse examples, try chain-of-thought for reasoning tasks. Test on at least 50–100 held-out examples — not the 5 you like. If the error rate is acceptable, stop. You are done, and you have saved yourself weeks.

Step 2: Is the missing piece knowledge or behavior? Ask: if I gave the model the right facts in the prompt, would it answer correctly? If yes, your problem is knowledge — use RAG. If the model has the facts but still answers in the wrong style, format, or reasoning pattern, your problem is behavior — consider fine-tuning.

A practical test: take 20 failure cases. For each, write the "ideal" retrieved context by hand and put it in the prompt. If failures mostly disappear, you have a retrieval problem. If the model still fails despite having the facts in front of it, you have a behavior problem.

Step 3: Do you have enough data? Fine-tuning needs examples — realistically hundreds at minimum, thousands ideally — of the exact input-output behavior you want. Not scraped web text; curated demonstrations. If you cannot produce or afford this data, prompting or RAG is your answer regardless of everything else.

Step 4: Will you run this at scale? If you will make millions of inference calls, the math often favors fine-tuning a small model once over prompting a large model forever. Estimate: (cost per prompt-based call − cost per fine-tuned call) × expected calls, versus the one-time cost of data + training. Chapter 10 gives you the budgeting tools.

Step 5: Can you combine approaches? This is the answer more often than textbooks admit. Fine-tune for behavior and style; use RAG for facts. A fine-tuned model that also retrieves documents is the standard architecture of serious production systems. The techniques are not rivals — they solve different halves of the problem.

Decision flowchart: three routes from one question

Figure: The core decision — if the gap is facts, retrieve; if the gap is behavior, fine-tune; if neither, prompt better.

A Worked Example

Suppose you are building a system that answers questions about your university's thesis regulations for MS students, in Urdu, with citations to specific regulation clauses.

  • Prompting a frontier model: it speaks Urdu decently, but it does not know your university's regulations. It will hallucinate clause numbers. Fails.
  • RAG: you index the regulation documents, retrieve relevant clauses, and the model answers in Urdu citing them. This solves the knowledge problem. Probably sufficient.
  • Fine-tuning: would you fine-tune? Only if there is a behavioral gap — for example, the model consistently cites clauses in the wrong format, or its Urdu is too formal for students, and instructions don't fix it. You would fine-tune on a few hundred examples of well-formed answers (behavior), while RAG still supplies the facts. Combined approach wins.

Now flip it: suppose you need a model that converts natural-language questions into SQL for your university's specific database schema, thousands of times a day. The schema never changes, the SQL dialect is fixed, and latency matters. Prompting works but is slow and occasionally produces wrong column names. RAG is pointless — there are no documents to retrieve; the schema fits in the prompt anyway. This is a fine-tuning problem: the behavior (schema-faithful SQL generation) needs to be internalized, and you will amortize training cost over millions of cheap inference calls.

The Honest Cost of Fine-Tuning

Students routinely underestimate what fine-tuning demands:

  • Data: hundreds to thousands of high-quality examples, each reviewed. This is the real cost — human time.
  • Compute: from free (Colab, QLoRA) to hundreds of dollars for larger runs.
  • Evaluation: building a proper test set and baseline comparisons. Without this, you cannot claim anything.
  • Maintenance: when the base model updates or your task drifts, you retrain. A prompt is edited in seconds; a fine-tune is rerun in hours.

None of this means "don't fine-tune." It means: fine-tune when the problem genuinely requires changed behavior, and prove the cheaper options failed first. That proof — "we tried prompting and RAG, here are the numbers, and here is why they were insufficient" — is also exactly what makes a fine-tuning paper credible. Reviewers ask "why didn't you just prompt?" more often than any other question. Chapter 1 just gave you the framework to answer it.

Hybrid Architectures: Using Two or Three Methods Together

In practice, the most capable systems rarely use one method alone. Once you understand each technique's strength, you can compose them like building blocks. Here are the hybrid patterns that appear most often in real deployments and papers:

Pattern 1: Fine-tuned model + RAG (the standard). The fine-tuned model handles behavior — tone, format, reasoning style, when to say "I don't know" — while the retriever supplies facts. Example: a fine-tuned customer-support model that always responds with empathy, a structured summary, and cited ticket IDs, where the ticket contents come from retrieval. Neither method alone produces this: RAG alone gives correct facts in inconsistent format; fine-tuning alone gives consistent format with hallucinated facts.

Pattern 2: Fine-tuned query rewriter + RAG. Retrieval quality depends heavily on the query. A small fine-tuned model rewrites the user's question into a better search query (or several), the retriever fetches documents, and a second model (or the same one) answers. The rewriter is cheap to train — a few thousand (question → better query) pairs — and often improves end-to-end accuracy more than upgrading the answering model.

Pattern 3: Distilling RAG into fine-tuning. During data collection, you generate responses with retrieval (so they're factually grounded), verify them, then fine-tune on the resulting pairs without retrieval at serving time. The model internalizes the grounded answering style. This works when the knowledge is stable (it won't change after training) and you want to drop the retrieval infrastructure for latency or cost reasons. The risk: the model memorizes facts that later go stale — document the knowledge cutoff the same way base models do.

Pattern 4: Router — fine-tuned classifier directing traffic. A tiny fine-tuned classifier decides per query whether to answer directly, retrieve, or escalate to a larger model. This is how production systems keep costs down: 80% of queries are easy and handled cheaply; the expensive paths trigger only when needed. The router itself is a fine-tuning task (classification, Chapter 3's smallest-data regime).

How to decide the composition: start from the decision framework's Step 2 (knowledge vs. behavior gap) and assign each gap its tool. A project with both gaps gets both tools. Evaluate the hybrid against each single-method baseline — the hybrid should win, and the ablation (hybrid minus one component) tells you what each part contributed. In papers, this composition analysis is often more interesting than any single component: "RAG alone: 72%, fine-tuning alone: 75%, combined: 84%, and removing the fine-tuned rewriter costs 6 points" is a complete story.

One caution: hybrids multiply failure modes. A fine-tuned model fed bad retrieved documents will confidently format misinformation beautifully — worse than either failure alone. Evaluate the combined system end to end, not just each part in isolation, and include adversarial tests where retrieval returns irrelevant or contradictory documents.

The "Do Nothing Yet" Option: Premature Fine-Tuning

There's a failure mode this chapter hasn't named: fine-tuning too early, before you understand the problem. Symptoms: you can't describe the behavioral gap in one sentence; your "dataset" is 200 examples scraped hastily; you haven't tried few-shot prompting seriously; you chose fine-tuning because it sounds more impressive than prompting.

Premature fine-tuning wastes the scarcest resource — your judgment. A model trained on a poorly understood task learns a poorly understood behavior, and the resulting outputs teach you nothing except that the data was inadequate — something a week of prompt experimentation would have revealed for free.

The disciplined alternative is the two-week rule: spend up to two weeks on prompting and error analysis before collecting fine-tuning data. During those two weeks, you will: discover what the model already does well (shrinking your data needs), build the evaluation set and script (needed regardless), develop the error taxonomy (which becomes your data-collection guide — you now know exactly which failure modes need training examples), and write the method-selection memo. None of this work is wasted if you do fine-tune — it's the foundation. And sometimes the two weeks reveal that prompting suffices, saving you two months.

This is also a strategic point for researchers: the prompting experiments aren't just due diligence, they're content. "We first established a strong few-shot baseline of 71.8% through systematic prompt engineering (Appendix A); fine-tuning improved this to 78.5%" is a stronger paper than one that never mentions prompting. The two-week rule makes your eventual fine-tuning both better-informed and better-defended.

For your research: Before any training run, write a one-page "method selection memo": the task, the three approaches considered, the quick experiments you ran for prompting and RAG with their scores, and why fine-tuning is justified. This memo becomes the related-work and motivation section of your paper almost verbatim. Reviewers reward this explicitness enormously — it shows you chose fine-tuning deliberately, not by default.

Key Takeaways

  • Prompting changes the input, RAG adds external knowledge, fine-tuning changes the model's weights. Only fine-tuning alters behavior permanently.
  • Always test prompting first (days, not weeks), then ask whether your gap is knowledge (→ RAG) or behavior (→ fine-tuning).
  • Fine-tuning wins for: consistent behavior at scale, internalized skills, cheap inference on narrow tasks, reliable structured output.
  • Fine-tuning loses for: changing facts, tiny datasets, problems prompting already solves.
  • The strongest systems combine approaches: fine-tune for behavior, retrieve for facts.
  • Document your method choice with numbers — "we tried X, scored Y, insufficient because Z" — it becomes your paper's motivation section.

Chapter 2: How Fine-Tuning Works, Conceptually (What Changes in the Weights)

To use fine-tuning well — and to write about it credibly — you need a mental model of what is actually happening when you "train" an already-trained model. This chapter builds that model with no calculus beyond what you already know: a model is a function with adjustable numbers, training adjusts those numbers to reduce mistakes, and fine-tuning is simply training that starts from a very good starting point instead of from scratch.

The Model as a Machine with Dials

A large language model is, at bottom, a gigantic function. It takes a sequence of tokens (pieces of text) as input and produces a probability distribution over what token comes next. The behavior of that function is controlled by its parameters — often called weights — which are just numbers. A 8-billion-parameter model has 8 billion such numbers. Every capability the model has — grammar, facts, reasoning patterns, style — is encoded in the particular values of those numbers. There is no separate "knowledge database" inside; the knowledge is the arrangement of the numbers.

When the model was originally trained (pre-training), it started with random numbers and read trillions of words, adjusting the numbers little by little so that its next-token predictions got better and better. This took months on thousands of GPUs. The result is a base model: a fluent, knowledgeable, general-purpose text predictor with no particular instruction-following manners. Ask a raw base model a question and it may continue it as if writing an essay, because it was trained to continue text, not to answer questions.

Fine-tuning takes that finished machine and keeps adjusting the dials — but now on your data, with your objective. If pre-training taught the model the English language and the world's facts, fine-tuning teaches it your specific job: answer in this format, follow these instructions, reason about this domain, refuse these requests, write in this style.

What Actually Changes: Gradients and Updates

Here is the mechanism in plain terms. You show the model one of your training examples — say, an instruction and its correct response. The model reads the instruction and predicts the response one token at a time, producing its own guess for each position. You compare its guess to the correct token and compute a loss: a single number measuring how wrong the prediction was. Then backpropagation figures out, for every one of the billions of weights, which direction to nudge it so the loss would have been slightly smaller. The optimizer applies those nudges, scaled by the learning rate. Repeat millions of times, and the weights settle into values that make your examples likely.

Three things about this process matter enormously for fine-tuning specifically:

1. You start from a good place, so you nudge gently. In pre-training, weights move across vast distances from random initialization. In fine-tuning, you are already near a good solution — the model already speaks the language — so you use a much smaller learning rate (typically 10–100× smaller than pre-training). Crank the learning rate too high and you blast the model out of its good configuration; the fluent generalist becomes a stuttering specialist. This is the single most common beginner error, and Chapter 7 will quantify it.

2. Every update is a trade. Each nudge that makes your examples more likely can make other things slightly less likely. The model's weights are shared across all its capabilities; there is no separate compartment for "your task." Push too hard on your narrow data and the model forgets general abilities — this is catastrophic forgetting, covered in Chapter 9. Fine-tuning is always a negotiation between the new behavior and everything the model already knew.

3. The model learns patterns, not examples. Given enough varied examples, the weights encode the pattern behind them — the format, the reasoning style, the domain conventions. Given too few examples, or examples that are too similar, the weights just memorize the specific instances. Memorization versus generalization is the central tension of dataset design (Chapter 3) and the reason evaluation on held-out data (Chapter 8) is non-negotiable.

Instruction Tuning: The Most Common Kind

The fine-tuning you will do 90% of the time is instruction tuning (sometimes called supervised fine-tuning, SFT). The format is simple: each training example is an instruction (what to do) plus a response (the correct way to do it). Sometimes there is also an input field for the data the instruction operates on.

Instruction: Classify the sentiment of this movie review as positive or negative.
Input: "The film dragged on for two hours and I checked my watch constantly."
Response: negative

During training, the model reads the instruction and input, and the loss is computed only on the response tokens — we do not penalize the model for failing to predict the instruction itself, since at deployment time the instruction comes from the user. (In Hugging Face's TRL library, this masking is handled for you; in raw Transformers code you implement it with label masking, as Chapter 6 shows.)

Why does this work so well? Because the base model already "knows" sentiment analysis from pre-training — it has seen millions of reviews. Instruction tuning does not teach it the concept of sentiment; it teaches it the behavior: when you see this instruction format, output exactly one word, no explanation. This is why a few thousand good examples can transform a model: you are not teaching knowledge, you are teaching manners and format. The landmark papers on this — Wei et al.'s "Finetuned language models are zero-shot learners" and Chung et al.'s "Scaling instruction-finetuned language models" — showed that instruction tuning on diverse tasks makes models dramatically better at following unseen instructions.

What the Weights Learn at Different Depths

A transformer model is a stack of layers — typically 32 for an 8B model. It helps to know, roughly, what different parts do, because it explains why some fine-tuning methods work:

  • Early layers (near the input) tend to handle low-level patterns: token combinations, basic grammar, local syntax. Fine-tuning rarely needs to change these much.
  • Middle layers build up richer representations: entities, relations, the "meaning" of the passage so far.
  • Late layers (near the output) are most responsible for task-specific behavior: deciding the actual next token, formatting, style.

This is not a clean division — capabilities are smeared across layers — but the general principle holds: the closer to the output, the more task-specific the representations. It is one reason LoRA (Chapter 4) can get away with modifying only certain weight matrices: small, targeted changes near the right places steer behavior effectively.

The attention mechanism deserves one plain-language note since you will see its matrices named in LoRA configurations. Attention is how each token "looks at" other tokens to gather context. It is controlled by weight matrices conventionally named query, key, value, and output projection (q_proj, k_proj, v_proj, o_proj). When a LoRA config says target_modules=["q_proj", "v_proj"], it means: apply the low-rank update to the query and value matrices of every attention layer. Those were the matrices the original LoRA paper found most effective to adapt.

The Loss Landscape Intuition

Imagine the model's performance as a landscape of hills and valleys, where height is the loss (lower is better) and position is the setting of all billions of weights. Pre-training dropped the model into a broad, good valley. Fine-tuning moves it to a nearby spot in that valley that is better for your task.

Two intuitions from this picture run the whole book:

  1. Small steps stay in the valley. A small learning rate and few epochs keep the model near its strong starting point — you get your task behavior plus retained general ability. Large steps can climb out of the valley entirely, and the model may never find its way back. This is the geometric reason for gentle fine-tuning.
  2. Your data defines a new downhill direction. The gradients computed on your examples point toward lower loss on your examples. If your examples are high quality and representative, that direction genuinely improves your task. If they are noisy, biased, or too narrow, the model follows them faithfully — downhill into a ditch. The model cannot tell good data from bad; it only descends. Dataset quality (Chapter 3) is therefore not hygiene — it is the steering wheel.

Fine-Tuning vs. Pre-Training: A Summary Table in Words

Pre-training: random start, trillions of tokens, huge learning rate, months of compute, objective is "predict the next token on the internet," result is general capability. Fine-tuning: expert start, thousands to millions of tokens, tiny learning rate, hours of compute, objective is "produce exactly these responses to these instructions," result is specialized behavior. Everything about fine-tuning — the small learning rates, the few epochs, the fear of forgetting — follows from the fact that you are making small adjustments to something that already works.

Why a Few Thousand Examples Can Move Billions of Weights

Students often find this puzzling: how can 2,000 examples meaningfully change a model with 8 billion parameters? Shouldn't the model just ignore such a tiny signal, or alternatively memorize it instantly? The answer involves a genuine research finding worth knowing.

A paper by Aghajanyan et al. ("Intrinsic Dimensionality Explains the Effectiveness of Language Model Fine-Tuning," ACL 2021) measured something surprising: although a language model has billions of parameters, the fine-tuning solution for a given task can be found by adjusting only a few thousand degrees of freedom. In their experiments, fine-tuning in a random subspace of just a few thousand dimensions recovered most of full fine-tuning's performance. The model's parameter space is enormously redundant for any single task — there are countless equivalent ways to implement "answer in this format," and the optimizer only needs to find one of them.

This has three practical consequences. First, it explains why LoRA works at all: if the useful update lives in a low-dimensional subspace, a low-rank approximation captures it. LoRA's rank-16 update isn't a hack that happens to work — it's matched to the actual structure of the problem. Second, it explains why small datasets suffice for behavioral tasks: you're not filling 8 billion parameters with information; you're nudging the model along a few thousand meaningful directions, and a few thousand good examples provide enough signal for that. Third, it predicts LoRA's limitation: tasks whose solutions genuinely need many independent directions (very broad domain shifts) need higher rank or full fine-tuning — which matches the empirical guidance in Chapter 4.

There's a second, complementary intuition: the base model already contains the capability; fine-tuning mostly selects and formats it. Pre-training exposed the model to sentiment analysis, summarization, translation, and countless formats. Your 2,000 examples don't teach these from scratch — they teach the model which of its existing capabilities to deploy, when, and in what wrapper. Selection among existing capabilities needs far less data than building capabilities. This is also why fine-tuning fails most dramatically exactly where pre-training was thinnest (Chapter 11's low-resource language case): there's nothing to select, so the model must build — and building needs far more data than selecting.

Keep both intuitions handy. When someone asks "why should I believe 3,000 examples is enough?", the answer is: because we're selecting and formatting pre-existing capabilities along a low-dimensional subspace, not training from scratch — and our data-size ablation (Chapter 8) empirically confirms where the curve saturates for this task.

A Tiny Numerical Example (Making the Update Concrete)

Abstract "nudges" become clearer with toy numbers. Imagine a miniature model with a single weight w = 2.0, and one training example where the correct output implies w should be larger. The model predicts with w=2.0, the loss is 1.0. Backpropagation computes the gradient: d(loss)/dw = −0.4 (negative means increasing w decreases loss). With learning rate 0.1, the update is:

w_new = w − LR × gradient = 2.0 − 0.1 × (−0.4) = 2.04

One step moved the weight 0.04 toward the better value. Now scale this intuition up: instead of one weight there are 8 billion, instead of one gradient direction there are 8 billion partial derivatives computed simultaneously by backpropagation, and instead of one example the gradient is averaged over a batch. The learning rate plays exactly the same role — with LR=0.1 the step was gentle; with LR=10 the same gradient would have flung w to 6.0, overshooting wildly. That's the entire mechanism; everything else is bookkeeping at scale.

Two subtleties worth carrying from this toy: first, the gradient is local — it only knows the downhill direction at the current position, not the global landscape. Training is a blindfolded descent feeling the slope one step at a time, which is why small steps and many of them work better than leaps. Second, in LoRA, this same update math applies only to the adapter matrices A and B — the base weight w stays frozen at 2.0 forever, and the effective weight becomes 2.0 + (B×A). The toy makes visible what "freezing" means: w never appears on the left side of an update equation.

For your research: Reviewers and thesis committees love a clear conceptual account of why your fine-tuning should work for your problem. Borrow the "manners vs. knowledge" framing: state explicitly whether your fine-tuning is teaching the model new domain knowledge, new behavior/format, or both — and design your evaluation to test each claim separately. A paper that says "we fine-tuned to teach radiology report structure (behavior) while relying on pre-trained medical knowledge, and our ablations separate the two" is far stronger than "we fine-tuned and the score went up."

Key Takeaways

  • A model is a function controlled by billions of weight numbers; fine-tuning nudges those numbers to make your examples more likely.
  • Training = predict, measure error (loss), compute nudges (backpropagation), apply nudges (optimizer). Fine-tuning is this loop starting from an expert model with a small learning rate.
  • Instruction tuning (SFT) is the standard recipe: instruction → response pairs, loss computed on the response only. It teaches behavior and format more than raw knowledge.
  • Every update trades new-task fit against old capabilities — the root cause of catastrophic forgetting.
  • Small steps stay in the "good valley" of the loss landscape; large steps can destroy the pre-trained model. Gentle updates are a principle, not a preference.
  • The model descends whatever direction your data points — good data steers well, bad data steers into a ditch. The model cannot judge data quality for you.

Chapter 3: Preparing Your Dataset: Instruction Format, Quality Beats Quantity

Every experienced practitioner will tell you some version of the same sentence: the dataset matters more than the model, the method, or the hyperparameters. A mediocre training setup on excellent data beats an excellent setup on mediocre data, nearly every time. This chapter is about building that excellent dataset — what format to use, where examples come from, how to judge quality, and how much data you actually need.

The Standard Format: Instruction, Input, Response

Supervised fine-tuning consumes examples in a consistent structure. The most widely used is the Alpaca format, named after the Stanford project that popularized it:

{
  "instruction": "Summarize the following research abstract in one sentence.",
  "input": "We present a method for low-rank adaptation of large language models...",
  "output": "The authors propose LoRA, a method that adapts LLMs by training small low-rank matrices instead of all weights."
}

The instruction tells the model what to do. The input (sometimes empty) provides the material to work on. The output is the correct response — the demonstration the model will learn to imitate. During training these are concatenated into a single text with a prompt template, for example:

### Instruction:
Summarize the following research abstract in one sentence.

### Input:
We present a method...

### Response:
The authors propose LoRA...

The template matters less than its consistency: use one template for the entire dataset, and use the same template at inference time. A model trained with ### Response: markers that is prompted at deployment with a different format will underperform — it learned to associate its behavior with the exact textual cues it saw in training.

Other formats you will encounter: ShareGPT format (a list of conversational turns with human/gpt roles — natural for dialogue data), and chat templates applied by the tokenizer (modern instruction models like Llama 3 expect messages wrapped in special tokens such as <|start_header_id|>; Hugging Face's tokenizer.apply_chat_template handles this, and Chapter 6 shows exactly how). The principle is identical in all cases: consistent structure, response tokens are what the model learns.

Quality Beats Quantity: What the Evidence Shows

The most liberating finding for students with limited resources is that a small, excellent dataset routinely outperforms a large, mediocre one. The LIMA paper ("Less Is More for Alignment," Zhou et al., 2023) demonstrated that 1,000 carefully curated examples could produce a strong instruction-following model — a result that reshaped how the field thinks about data. The mechanism is the one from Chapter 2: fine-tuning mostly teaches behavior and format, and behavior can be demonstrated in a few hundred diverse, perfect examples. Knowledge, which needs volume, mostly comes from pre-training.

What makes an example "high quality"? Judge every candidate against these criteria:

  1. Correctness. The output must be right. A single wrong example teaches the model to be confidently wrong — and confident wrongness is the hardest error to detect later. For factual tasks, verify outputs against sources. For subjective tasks (style, tone), make sure the output genuinely exemplifies what you want.
  2. Diversity. Your examples should cover the full range of inputs the model will see: different topics, lengths, difficulties, edge cases. One hundred near-identical examples teach one narrow pattern; one hundred varied examples teach a general skill. Check diversity explicitly — cluster your instructions by topic and look for thin clusters.
  3. Appropriate difficulty. Include easy, medium, and hard examples. If everything is trivial, the model learns nothing about handling hard cases; if everything is brutal, training is unstable and the model may memorize rather than generalize.
  4. Clean formatting. Consistent output format across examples. If half your examples end with a summary sentence and half don't, the model learns to be inconsistent — faithfully reproducing your inconsistency.
  5. No leakage. Your training examples must not overlap with your test set (Chapter 8). Even paraphrased overlap inflates your scores dishonestly. Keep a strict split from day one.

A useful quality-control ritual: sample 100 random examples and grade each one yourself against the five criteria above. If more than 5–10% fail, fix the data pipeline before training anything. An hour of grading saves days of confused debugging.

Where Examples Come From

You have four main sources, in rough order of quality:

1. Human-written examples (best, most expensive). Domain experts write instruction-response pairs. For a thesis project, this might be you and your advisor writing 500–2,000 examples over a few weeks. Slow, but the quality ceiling is highest, and for specialized domains (medicine, law, low-resource languages) there is no substitute.

2. Distillation from a stronger model. You write the instructions (or collect real user queries) and have a frontier model generate the responses, then verify and filter them. This is how most open instruction datasets were built. The critical step is verification: generate 3,000 responses, have humans or strong automated checks reject the bad 30%, keep 2,000. Unfiltered distillation bakes the teacher's errors into your student permanently.

3. Repurposed existing datasets. Academic NLP datasets (question answering, summarization, classification) can be reformatted into instruction format with templates. "Convert this SQuAD example into an instruction pair" is a legitimate and common pipeline. Watch for: template artifacts (the model learns your template's quirks), and license compatibility (see the ethics note below).

4. Synthetic generation with verification. Generate instructions programmatically or with a model, generate responses, then verify with rules, a second model, or humans. Scales well; quality depends entirely on the verification step. Never skip verification — unverified synthetic data is how models learn confident nonsense.

What to avoid: scraping raw web text and calling it a fine-tuning dataset (that is pre-training data, not instruction data); using copyrighted text you have no right to use; and — critically for researchers — training on your test set or on data contaminated with it.

How Much Data Do You Need?

Honest answer: it depends on what you are teaching, but here are practical brackets:

  • Format/style adaptation (the model knows the task, needs your format): 500–2,000 examples can suffice. This is the LIMA regime.
  • Domain adaptation (new jargon, new conventions, e.g., legal drafting in your jurisdiction): 2,000–20,000 examples.
  • New skill or weak base capability (a language the model barely speaks, complex multi-step reasoning): 10,000–100,000+ examples, and temper expectations — fine-tuning cannot fully compensate for missing pre-training.
  • Classification or structured extraction (narrow, well-defined): sometimes as few as a few hundred per class, because the output space is small.

Start small and scale deliberately: train on 1,000 examples, evaluate (Chapter 8), then try 3,000 and 10,000. Plot performance against data size. This data ablation is cheap, informative, and exactly the kind of figure reviewers love — it answers "how much data does this task need?" which is itself a research contribution.

The Train/Validation/Test Split

Split your data into three parts before you do anything else:

  • Training set (~80–90%): what the model learns from.
  • Validation set (~5–10%): used during training to monitor progress, pick checkpoints, and tune hyperparameters. The model never trains on it, but you do look at it repeatedly — so it slightly influences decisions.
  • Test set (~5–10%): locked away. Evaluated once, at the very end, on the final model. This is the number you report.

For small datasets (under ~2,000 examples), consider cross-validation or at least multiple random splits, because a single small test set is noisy — a 5% swing might be luck. Report the variance, not just the mean.

A Concrete Preparation Script

Here is a realistic data pipeline in Python — loading raw examples, applying a template, splitting, and saving in a training-ready format:

from datasets import Dataset
import json, random

# 1. Load your raw examples (however you collected them)
with open("raw_examples.json") as f:
    raw = json.load(f)  # list of {"instruction","input","output"}

# 2. Quality filter: drop empties, duplicates, overlong examples
seen = set()
clean = []
for ex in raw:
    text = (ex["instruction"] + ex["input"] + ex["output"]).strip()
    if not text or len(ex["output"]) < 10:
        continue                      # empty or trivial
    key = ex["instruction"][:80]
    if key in seen:
        continue                      # near-duplicate instruction
    seen.add(key)
    if len(text) > 6000:
        continue                      # overlong; handle long docs separately
    clean.append(ex)

print(f"kept {len(clean)}/{len(raw)} after filtering")

# 3. Apply a consistent template
TEMPLATE = ("### Instruction:\n{instruction}\n\n### Input:\n{input}\n\n### Response:\n{output}")
def format_example(ex):
    return {"text": TEMPLATE.format(**ex)}

formatted = [format_example(ex) for ex in clean]

# 4. Shuffle with a fixed seed, then split (reproducibility matters)
random.Random(42).shuffle(formatted)
n = len(formatted)
train = formatted[:int(0.9*n)]
val   = formatted[int(0.9*n):int(0.95*n)]
test  = formatted[int(0.95*n):]

# 5. Save; the test set goes somewhere you will not touch during training
Dataset.from_list(train).to_json("data/train.json")
Dataset.from_list(val).to_json("data/val.json")
Dataset.from_list(test).to_json("data/test.json")
print(f"train={len(train)} val={len(val)} test={len(test)}")

Note the fixed random seed: reproducibility starts with data preparation, not with training. A reviewer who cannot reproduce your split cannot reproduce your results.

Ethics and Licensing Notes

Two obligations researchers sometimes overlook. First, licenses: many datasets prohibit commercial use or redistribution of derivatives; fine-tuned model weights trained on such data inherit the restriction in practice. Check the license of every dataset you repurpose, and state data licenses in your paper (Chapter 12). Second, privacy: if your examples contain real people's data — patient notes, student essays, customer messages — anonymize before training. A fine-tuned model can memorize and regurgitate training examples; training on private data without consent is both unethical and, in many jurisdictions, unlawful.

Writing Annotation Guidelines: What You Hand Your Annotators

If humans write or verify your examples — even if "humans" means you and one colleague — write annotation guidelines first. Without them, two annotators produce two different datasets stitched together, and the model faithfully learns the inconsistency. Good guidelines have four parts:

1. The task definition with the "why." Not just "write a response" but "write a response a rural clinic worker could read aloud to a patient; the goal is clarity, not completeness." Annotators who understand the purpose make better judgment calls on edge cases than any rule list covers.

2. Do-and-don't examples (at least 5 pairs). Show a good response and a bad response for the same instruction, with one sentence explaining the difference. This is worth more than pages of prose. Include edge cases: what to do when the instruction is ambiguous, when the correct answer is "I don't know," when the input contains errors.

3. Format specification, exact. If outputs must be JSON with keys answer and justification, say so, show it, and provide a validator script annotators run before submitting. Every format decision you leave implicit becomes noise in the training data.

4. The escalation rule. "If you're unsure, flag it instead of guessing." Guessed examples are worse than missing examples — they teach confident wrongness. Create a needs_review bucket and have your most careful annotator (possibly you) resolve it.

Measuring agreement: have two annotators independently do the same 50 examples, then compare. For classification-style outputs, compute simple agreement rate (aim for 90%+; below 80% means your guidelines are ambiguous). For free-text responses, agreement is harder to quantify — instead, have a third person blind-rank pairs of responses from the two annotators against the guidelines. Low agreement doesn't mean bad annotators; it means ambiguous guidelines. Rewrite the guidelines, not the people.

Pilot first: run 50 examples through the full pipeline (guidelines → annotation → your quality grading from earlier in this chapter) before scaling to thousands. The pilot reveals every ambiguity in your guidelines at 1/40th the cost. Every experienced dataset builder pilots; every burned one wishes they had.

A final note on you-as-annotator: if you're writing all examples yourself (common in thesis projects), you don't need inter-annotator agreement — but you do need self-consistency. Write your guidelines anyway, and re-grade your first 100 examples after finishing all 2,000. You'll find your standards drifted; normalize the early examples to your later, better-calibrated judgment.

Versioning Your Dataset

Datasets change: you fix labels, add examples, remove contamination, rephrase instructions. Without versioning, "the dataset" becomes ambiguous — which version produced which result? Adopt this lightweight scheme from day one:

Name versions explicitly: data-v1 (initial 1,000), data-v2 (added 500 medical examples, removed 37 duplicates), etc. Keep a CHANGELOG.md in your data directory: one line per version saying what changed and why. Every training run records its data version in the run log (next to the seed and config).

Never edit in place. When you fix 50 labels, create data-v3 — don't silently overwrite data-v2. Old results must remain interpretable: "run-04 used data-v2" should still be true and checkable a year later. Disk is cheap; ambiguity is expensive.

Hash the splits. After creating train/val/test, record a hash (e.g., SHA-256 of the concatenated file) in the changelog. If anyone — including you — questions whether the test set leaked, the hash proves which bytes were evaluated.

Why this matters for papers: reviewers increasingly ask "which data version?" when results don't reproduce, and dataset versioning is standard practice in industry labs. A changelog with five entries and hashes signals a mature experimental process — and when your advisor asks "did the improvement come from the new data or the new hyperparameters?", the version log plus run log answers definitively instead of approximately.

For your research: Your dataset is a research artifact. Document it like one: collection method, size, filtering steps, inter-annotator agreement if humans wrote examples, license, and a datasheet-style description. Releasing your dataset (when legally possible) alongside your paper multiplies its impact — datasets get cited independently. Even a 2,000-example high-quality dataset for an under-served language or domain is publishable in dataset tracks of major venues.

Key Takeaways

  • Use one consistent format (instruction/input/output) and one prompt template everywhere — training and inference.
  • Quality beats quantity: 1,000 excellent, diverse, correct examples beat 50,000 noisy ones for behavior teaching.
  • Judge examples on correctness, diversity, difficulty mix, formatting consistency, and no test leakage.
  • Sources in quality order: human-written, verified distillation, repurposed datasets, verified synthetic generation. Always verify.
  • Size brackets: ~500–2k for format adaptation, 2k–20k for domain adaptation, 10k+ for genuinely new skills.
  • Split train/val/test with a fixed seed before training; lock the test set away until the end.
  • Check licenses and anonymize private data — the dataset is a research artifact with legal weight.

Chapter 4: Full Fine-Tuning vs PEFT: LoRA and QLoRA Explained

Chapter 2 described fine-tuning as nudging all billions of weights. That is full fine-tuning, and it works — but it needs enormous memory: the weights, the gradients, and the optimizer's bookkeeping all live on the GPU at once. For a 7-billion-parameter model in standard 16-bit precision, full fine-tuning needs roughly 60+ GB of GPU memory, far beyond any single student-grade GPU. Parameter-efficient fine-tuning (PEFT) is the family of methods that dodges this wall by training only a tiny fraction of the parameters. LoRA is its most important member, and QLoRA made it run on free hardware. This chapter explains both deeply enough that you can configure them with understanding rather than superstition.

Why Full Fine-Tuning Is So Memory-Hungry

To see what PEFT saves, count what full fine-tuning stores per parameter. Take a 7B model in 16-bit (2 bytes per number):

  • Weights: 7B × 2 bytes ≈ 14 GB
  • Gradients: another 7B × 2 bytes ≈ 14 GB (one gradient per weight)
  • Optimizer state: the Adam optimizer stores two extra numbers per parameter (a running average of gradients and of squared gradients) in 32-bit: 7B × 4 × 2 ≈ 56 GB

Total: roughly 84 GB before activations (the intermediate values stored for backpropagation). Even with memory tricks, you need a data-center GPU. And after training, you have a full 14 GB copy of weights per task — ten tasks means ten full models to store and serve.

PEFT attacks both problems: train far fewer numbers, and store far fewer numbers per task.

The Core Idea of LoRA

LoRA — Low-Rank Adaptation, introduced by Hu et al. — starts from an observation about fine-tuning: the change made to the weights during fine-tuning has low "intrinsic rank." In plain terms: even though a weight matrix has millions of entries, the useful update to it can be captured by multiplying two much smaller matrices.

Concretely, take one weight matrix W of size d×d (say 4096×4096 ≈ 16.8 million numbers). During fine-tuning, it would change to W + ΔW. LoRA says: freeze W completely (never update it, never store gradients or optimizer state for it), and learn the update as a product of two small matrices:

ΔW = B × A, where B is d×r and A is r×d, with r (the rank) tiny — typically 8, 16, 32, or 64.

With r=16 and d=4096: B has 4096×16 = 65,536 numbers, A has 16×4096 = 65,536 numbers — about 131,000 trainable numbers instead of 16.8 million. That is roughly a 128× reduction for that matrix. Across the whole model, a LoRA adapter is typically 0.1–1% of the base model's parameters — tens of megabytes instead of gigabytes.

At inference time, the effective weight is W + BA. You can even merge the adapter into the base weights (W' = W + BA) so the deployed model runs exactly as fast as the original — zero inference overhead. Or keep adapters separate and swap them: one base model, many task adapters, each a small file. For research, this is wonderful: your "model" for the paper is a 100 MB adapter file anyone can download and plug into the public base model.

Two clever details make LoRA train well:

  1. Initialization: B starts at zero and A starts random. So at the start of training, BA = 0 and the model behaves exactly like the base model. Training only ever adds a deviation from the base behavior — it cannot start by being broken.
  2. Scaling: the update is multiplied by α/r (alpha over rank), where α is a hyperparameter. This keeps the effective update magnitude stable when you change r, so tuning r doesn't require retuning the learning rate as much.

LoRA low-rank concept: frozen matrix plus two small trainable matrices

Figure: The frozen base weight plus the two small trainable matrices whose product forms the update.

QLoRA: LoRA on a Quantized Model

LoRA slashed the trainable parameters, but the frozen base weights still sit on the GPU in 16-bit — 14 GB for a 7B model, plus activations. QLoRA (Dettmers et al.) attacks the frozen weights with quantization: storing each weight in 4 bits instead of 16.

Four-bit might sound brutally lossy, but two innovations make it work:

  • NF4 (NormalFloat4): a 4-bit number format designed for the actual distribution of neural network weights (which are roughly normally distributed). It allocates its 16 representable values where weights actually fall, rather than evenly — much less rounding error than naive 4-bit.
  • Double quantization: the quantization constants themselves are quantized, squeezing out more memory savings.

The result: a 7B model's frozen weights fit in about 3.5–4 GB. During training, weights are dequantized to 16-bit on the fly for the forward/backward pass, gradients flow only into the LoRA adapters, and the base weights are never updated. Memory for QLoRA fine-tuning of a 7–8B model lands around 8–12 GB total — fitting on a free Colab T4 GPU (16 GB). The QLoRA paper showed this recipe matching full 16-bit fine-tuning quality across benchmarks, and even enabled fine-tuning a 65B model on a single 48 GB GPU.

The one cost of QLoRA is speed: dequantizing on the fly adds compute overhead, so training runs somewhat slower than 16-bit LoRA. For students, this is an excellent trade — slower but possible beats fast but impossible.

Configuring LoRA in Practice (Hugging Face PEFT)

Here is what a LoRA configuration actually looks like, with every knob explained:

from peft import LoraConfig, get_peft_model, prepare_model_for_kbit_training

peft_config = LoraConfig(
    r=16,                          # rank: the "width" of the adapter. Higher = more capacity.
    lora_alpha=32,                 # scaling factor; effective update scaled by alpha/r = 2 here
    lora_dropout=0.05,             # dropout on the adapter: mild regularization
    bias="none",                   # don't train bias terms (keeps adapter tiny)
    task_type="CAUSAL_LM",         # next-token prediction (standard for instruction tuning)
    target_modules=["q_proj", "v_proj"],
    # which weight matrices get adapters.
    # Common choices: ["q_proj","v_proj"] (original paper, minimal),
    # or all of ["q_proj","k_proj","v_proj","o_proj","gate_proj","up_proj","down_proj"] (more capacity)
)

model = prepare_model_for_kbit_training(model)  # needed for quantized base models
model = get_peft_model(model, peft_config)
model.print_trainable_parameters()
# e.g. "trainable params: 41,943,040 || all params: 8,030,000,000 || trainable%: 0.52%"

That final printout is your sanity check: trainable parameters should be well under 1–2% of the total. If it says 100%, you forgot to freeze something.

Choosing r (rank): r=8 or 16 is the standard starting point and is enough for most format/style adaptation. Increase to 32–64 for harder tasks (new domains, complex reasoning). Higher rank = more capacity but more memory and more overfitting risk on small data. A good research practice: ablate r ∈ {8, 16, 32} on your validation set and report the curve.

Choosing target modules: adapting only attention (q_proj, v_proj) is the conservative default from the original paper. Adapting all linear layers (attention + MLP) gives the adapter more capacity and often better results on hard tasks, at the cost of a larger adapter file (still tiny relative to the model). For your first run, use the attention-only default; expand if validation performance plateaus.

When Full Fine-Tuning Still Wins

PEFT is not always the answer. Full fine-tuning remains preferable when:

  • You have the hardware. If your lab has 80 GB GPUs, full fine-tuning is simpler (no adapter machinery) and sometimes slightly better on very hard domain shifts.
  • The domain shift is extreme. Adapting to a language barely present in pre-training, or to a wildly different data distribution, may need changes LoRA's low-rank updates cannot express. (Though try high-rank LoRA first — it often suffices.)
  • You are pre-training further (continued pre-training). If you are continuing next-token training on gigabytes of domain text rather than instruction tuning, full fine-tuning is the standard choice.

For a student on limited hardware doing instruction tuning, the honest default is QLoRA. It is what makes this book's Chapter 6 walkthrough possible on free Colab, and it is a legitimate, publishable method — the QLoRA paper itself is a NeurIPS publication, and hundreds of papers since have used it as their training method.

The PEFT Family Beyond LoRA (Awareness Level)

You will encounter these names; know what they are at one paragraph each:

  • Adapters (classic): small bottleneck layers inserted between transformer layers. The original PEFT idea; LoRA largely superseded it because adapters add inference latency while LoRA merges away.
  • Prefix/prompt tuning: learnable virtual tokens prepended to the input instead of changing weights. Extremely parameter-efficient but often weaker than LoRA on hard tasks.
  • IA³: learns three small vectors per layer that rescale activations. Tiny and fast; competitive on some tasks.
  • DoRA: decomposes weights into magnitude and direction components and applies LoRA to the direction. Often slightly outperforms LoRA at similar cost — worth knowing as a one-line config swap in PEFT (use_dora=True) for an ablation.

For your research, LoRA/QLoRA is the baseline everything else is compared against. If you experiment with alternatives, LoRA is the control condition.

Memory Math: A Worked Example

Let's make the memory savings concrete with numbers you can verify on your own GPU. Take an 8B-parameter model and compare three training setups. (Byte counts: 16-bit = 2 bytes, 32-bit = 4 bytes, 4-bit = 0.5 bytes.)

Setup A: Full fine-tuning, 16-bit weights, AdamW. - Weights: 8B × 2 = 16 GB - Gradients: 8B × 2 = 16 GB - AdamW states (two 32-bit numbers per parameter): 8B × 8 = 64 GB - Subtotal: ~96 GB, plus activations (several GB more depending on sequence length and batch size) - Verdict: needs multiple data-center GPUs or exotic sharding. Not a student setup.

Setup B: LoRA (16-bit base), r=16 on all linear layers. - Frozen base weights: 16 GB (no gradients, no optimizer state — frozen means frozen) - Trainable adapter: ~42M parameters → weights 0.08 GB, gradients 0.08 GB, AdamW states 0.34 GB - Subtotal: ~16.5 GB plus activations - Verdict: fits on a 24 GB GPU (RTX 4090, A10), tight on 16 GB once activations are counted. This is why LoRA alone, without quantization, still strains a Colab T4.

Setup C: QLoRA (4-bit NF4 base), r=16. - Frozen base weights: 8B × 0.5 = 4 GB - Adapter + optimizer states: ~0.5 GB (as above; paged_adamw_8bit halves the optimizer memory further) - Activations with gradient checkpointing: ~2–4 GB for batch 2, seq length 1024 - Subtotal: roughly 7–10 GB - Verdict: comfortable on a 16 GB T4 with headroom. This is the Chapter 6 walkthrough, and now you see exactly why it fits.

Two lessons fall out of the arithmetic. First, the optimizer states are the silent killer in full fine-tuning — 64 GB of the 96 GB total. Any method that avoids optimizer state on the base weights (which is all of PEFT) wins enormously. Second, quantization's contribution is specifically shrinking the frozen weights (16 GB → 4 GB); the adapter was already tiny. When someone asks "why not just LoRA without quantization?", the answer is the 12 GB difference in frozen-weight storage — exactly the gap between fitting and not fitting on free hardware.

You can confirm all of this empirically: after loading your model in Chapter 6's Step 1, run nvidia-smi and compare against the 4 GB prediction; after attaching the adapter, run model.print_trainable_parameters() and multiply out the bytes. Making the theory touch the metal once builds intuition that lasts.

LoRA Variants in One Page: DoRA, rsLoRA, and Friends

The PEFT library offers LoRA variants selectable with one-line config changes. Know what they are so you can run them as ablations:

  • DoRA (Weight-Decomposed Low-Rank Adaptation). Decomposes each weight into a magnitude vector and a direction matrix, applies LoRA to the direction, and trains magnitudes separately. Intuition: learning "how strong" separately from "which way" is a more natural decomposition. Empirically it often beats vanilla LoRA by a small margin at similar cost. Enable with use_dora=True in LoraConfig. The adapter is slightly larger and training slightly slower — worth one ablation run.
  • rsLoRA (rank-stabilized LoRA). Changes the scaling factor from α/r to α/√r. The original scaling can under-serve high ranks (the update shrinks as r grows, partially canceling the extra capacity). If you're experimenting with r=64+, rsLoRA (use_rslora=True) makes high-rank runs behave more sensibly.
  • PiSSA. Initializes A and B from the principal components of the weight matrix (via SVD) instead of random/zero. The idea: start the adapter already pointing along the weight's most important directions, so training converges faster. Useful when training time is the bottleneck.
  • LoRA+. Uses different learning rates for the A and B matrices (higher for B). A small tweak with occasional wins; mostly of interest if you're squeezing the last point of performance.

For your first project, vanilla LoRA is the right baseline — it's what the literature compares against, and variant gains are usually 0–2 points. Run one variant (DoRA is the best-supported choice) as an ablation after your main results are solid. A paper reporting "LoRA 78.5, DoRA 79.1" shows thoroughness; a paper using an exotic variant as its only method invites questions about whether vanilla LoRA would have sufficed.

For your research: Your paper's methods section must state the PEFT configuration completely: rank r, alpha, dropout, target modules, quantization type (NF4, double quantization on/off), and the base model with exact version/hash. "We used LoRA" is not reproducible — two papers with "LoRA" and different ranks are different experiments. Also report the trainable parameter count and percentage; reviewers use it to judge whether your comparison to full fine-tuning baselines is fair.

Key Takeaways

  • Full fine-tuning stores weights + gradients + optimizer state (~84 GB for 7B); PEFT trains a tiny fraction of parameters instead.
  • LoRA freezes base weights and learns the update as B×A, two small matrices of rank r — typically 0.1–1% of parameters, mergeable to zero inference overhead.
  • QLoRA adds 4-bit NF4 quantization of the frozen base, fitting 7–8B fine-tuning into ~8–12 GB — free-Colab territory — with quality matching 16-bit.
  • Start with r=16, alpha=32, attention targets; expand rank and target modules only if validation plateaus.
  • Full fine-tuning still wins with big hardware or extreme domain shifts; QLoRA is the honest student default.
  • Report the full PEFT config in papers — "we used LoRA" is not reproducible.

Chapter 5: Tooling: Hugging Face Transformers, TRL, Axolotl (Overview + Setup)

You now understand what fine-tuning is and which method to use. This chapter is about the software that does the heavy lifting. The open-source ecosystem has converged on a standard stack, and learning it once pays off across every project: Hugging Face Transformers (models and training primitives), PEFT (the adapter library), TRL (high-level fine-tuning trainers), and Axolotl (configuration-driven training). We will cover what each does, when to reach for it, and get your environment set up.

The Stack, Layer by Layer

Hugging Face Transformers is the foundation. It provides: thousands of pre-trained models loadable in two lines (AutoModelForCausalLM.from_pretrained(...)), tokenizers that convert text to the numbers models consume, and the Trainer class — a general training loop handling batching, optimization, checkpointing, logging, and multi-GPU distribution. Everything else in this chapter builds on it. If you learn one library deeply, make it this one.

PEFT (Parameter-Efficient Fine-Tuning) is Hugging Face's adapter library. It implements LoRA, QLoRA support, DoRA, IA³, prefix tuning, and more, with a uniform interface: define a config, wrap your model with get_peft_model, train normally, save a small adapter. Chapter 4 already showed its core API.

TRL (Transformer Reinforcement Learning — the name is historical; it now covers all post-training) provides SFTTrainer, a specialized trainer for supervised fine-tuning that handles the fiddly details: applying chat templates, masking instruction tokens from the loss, packing short examples together for efficiency, and integrating with PEFT in a few lines. For instruction tuning, SFTTrainer is the single most productive tool in the ecosystem — it turns what used to be 300 lines of custom training code into about 30.

Axolotl sits one level higher: instead of writing Python training scripts, you write a YAML config file describing everything (base model, dataset, LoRA settings, hyperparameters), and Axolotl runs the training. This is superb for research because the config file is the experiment specification — version-controlled, diffable, and directly publishable. Many published fine-tuning experiments now ship their Axolotl configs as reproducibility artifacts.

Supporting libraries you will meet: bitsandbytes (the quantization engine behind QLoRA's 4-bit loading), datasets (Hugging Face's data library — streaming, caching, and preprocessing for huge files), and accelerate (handles device placement and multi-GPU behind the scenes).

Choosing Your Tool for the Job

  • Learning / first experiments / full control: raw Transformers Trainer + PEFT. You see every step; debugging teaches you the most.
  • Standard instruction tuning (the 90% case): TRL's SFTTrainer + PEFT. Minimal code, correct defaults, fast iteration.
  • Systematic experiments / ablations / sharing: Axolotl YAML configs. Your future self — and your reviewers — will thank you.
  • RLHF or preference optimization (DPO): TRL again (DPOTrainer). Beyond this book's scope, but good to know the same library covers it.

A common student path: prototype with SFTTrainer in a Colab notebook (Chapter 6), then graduate to Axolotl configs when running the serious ablation grid for the paper.

Environment Setup

You need Python 3.10+, PyTorch with CUDA, and a GPU with at least ~12–16 GB VRAM for QLoRA on 7–8B models (a Colab T4 qualifies). Install the stack:

# Core stack (versions move fast; these were current and mutually compatible
# at time of writing — check the TRL docs if you hit version conflicts)
pip install torch --index-url https://download.pytorch.org/whl/cu121
pip install transformers datasets accelerate peft trl bitsandbytes

Verify the GPU is visible and quantization works:

import torch
print(torch.cuda.is_available())          # must be True
print(torch.cuda.get_device_name(0))      # e.g. "Tesla T4"

from transformers import AutoModelForCausalLM, BitsAndBytesConfig

bnb_config = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_quant_type="nf4",            # NormalFloat4, per the QLoRA paper
    bnb_4bit_use_double_quant=True,       # double quantization: saves a bit more memory
    bnb_4bit_compute_dtype=torch.bfloat16 # compute in bfloat16 for stability
)

model = AutoModelForCausalLM.from_pretrained(
    "meta-llama/Llama-3.1-8B",            # any causal LM; gated models need HF access approval
    quantization_config=bnb_config,
    device_map="auto",
)
print("quantized model loaded OK")

Two setup notes that save hours of confusion:

  1. Gated models (Llama, Gemma) require accepting the license on the Hugging Face website and logging in (huggingface-cli login) before download. Do this once; the token is cached.
  2. Disk space: model downloads are large (8B ≈ 16 GB in 16-bit; quantized downloads are smaller since quantization happens on load). Colab gives you limited disk — download once per session, and consider huggingface_hub's cache settings if you run many experiments.

Axolotl Setup and a Minimal Config

Axolotl installs as a package and runs from a YAML file:

pip install axolotl
# or, for the latest: pip install git+https://github.com/axolotl-ai-cloud/axolotl.git

A minimal but complete Axolotl config for QLoRA instruction tuning:

# axolotl-config.yaml
base_model: meta-llama/Llama-3.1-8B
model_type: LlamaForCausalLM
tokenizer_type: AutoTokenizer

load_in_4bit: true
adapter: qlora
lora_r: 16
lora_alpha: 32
lora_dropout: 0.05
lora_target_modules: [q_proj, v_proj]

datasets:
  - path: data/train.json
    type:
      system_prompt: ""
      field_system: null
      field_instruction: instruction
      field_input: input
      field_output: output
      format: "{instruction} {input} {output}"
      no_input_format: "{instruction} {output}"

val_set_size: 0.05
sequence_len: 2048
sample_packing: true

num_epochs: 3
micro_batch_size: 2
gradient_accumulation_steps: 8
learning_rate: 0.0002
optimizer: paged_adamw_8bit
lr_scheduler: cosine
warmup_steps: 30

output_dir: ./outputs/run-01
save_steps: 200
logging_steps: 10

Run it with axolotl train axolotl-config.yaml. Every hyperparameter in this file will be explained in Chapter 7; for now, notice the property that matters: this file fully specifies the experiment. Commit it to git with a tag, and anyone — including you in six months — can reproduce the run exactly.

What Each Tool Handles vs. What You Handle

A clear division of labor prevents both magical thinking and unnecessary suffering:

  • The libraries handle: tokenization details, chat templates, loss masking, gradient accumulation, mixed precision, checkpointing, distributed training, quantization math.
  • You handle: the dataset (quality, format, splits), the evaluation (metrics, baselines, test discipline), the hyperparameters (chosen deliberately, not by default), and the scientific judgment (what the results mean).

Beginners often invert this: they obsess over library internals while feeding the model sloppy data and evaluating on vibes. Chapters 3 and 8 exist to keep you on the right side of that line.

Reproducibility in Practice: Environments That Don't Rot

Chapter 5's takeaway said "freeze and record exact library versions." Here's how to actually do it so the environment still works in six months:

The requirements file. After a working run, capture the environment:

pip freeze > requirements-run01.txt

This file goes into git next to the Axolotl config or training script. When you (or a reviewer) need to rerun, create a fresh environment and pip install -r requirements-run01.txt. Note: pip freeze on Colab captures Colab-specific packages too — that's fine; what matters is that transformers/peft/trl/bitsandbytes/torch versions are pinned.

The one-line version log. In addition to the freeze file, print versions into your training logs:

import transformers, peft, trl, torch
print(f"transformers={transformers.__version__} peft={peft.__version__} trl={trl.__version__} torch={torch.__version__}")

When a run from three months ago behaves differently today, the first thing you check is whether the environment changed — and this line answers it in seconds.

Randomness control. Set seeds everywhere, not just in the trainer: Python's random, NumPy, and PyTorch all have separate RNGs. Hugging Face's set_seed(42) handles all three plus CUDA. Do it at the top of every script. Perfect determinism on GPUs isn't guaranteed even with seeds (some CUDA operations are nondeterministic), but seeded runs are close enough that a large discrepancy signals a real problem, not noise.

What to version-control. Commit: configs, training scripts, eval scripts, requirements files, the data preparation script, and a README describing the run order. Do NOT commit: model weights, datasets (use releases or data versioning if large), or API tokens. A good rule: the repository should let a stranger go from "git clone" to "reproduced results" following only the README.

The environment decay problem. Libraries move fast; a requirements file from today may fail to install in a year (a dependency dropped an old wheel, a CUDA version mismatch). For thesis-critical reproducibility, the gold standard is a container (Docker) or a full environment export. For most student projects, the pinned requirements file plus the version log is sufficient — just re-verify the install works before you need it, e.g., once at thesis-writing time.

Choosing a Base Model: A Practical Workflow

The stack is only half the decision — which model to fine-tune matters enormously, and students often default to whatever is most famous. Here's a deliberate selection workflow:

Step 1: List candidates in your size class. For QLoRA on a 16 GB GPU: 7–9B models (Llama-3.1-8B, Qwen2.5-7B, Gemma-2-9B, Phi-3-medium). For CPU-only or tiny GPUs: 1–3B models (Qwen2.5-1.5B, Llama-3.2-1B/3B). Don't start with 70B models — prove the pipeline on 8B first.

Step 2: Check the license. Llama and Gemma are free for research but have custom licenses with conditions; Qwen and Phi have their own terms; fully open models (e.g., OLMo, Falcon) use permissive licenses. If your thesis might become a commercial product, license matters from day one. Note the license in your experiment plan.

Step 3: Run a 200-example probe. Before any training, score each candidate zero-shot and few-shot on 200 examples representative of your task. This takes an afternoon and often reveals a clear winner — e.g., Qwen outperforming Llama on multilingual tasks, or Phi punching above its size on reasoning. The probe also gives you the base baselines Chapter 8 requires. Skipping this and discovering mid-project that another base model was 8 points better is an expensive mistake.

Step 4: Check tokenizer fit for your language. Tokenizers trained mostly on English split other languages into many tiny pieces ("fertility" — tokens per word). High fertility means your sequences are longer, training is slower, and the model sees less context. A quick check: tokenize 100 sentences of your data with each candidate's tokenizer and compare average tokens per sentence. For non-English projects, this single check can matter more than benchmark scores.

Step 5: Prefer instruction-tuned variants as starting points. For instruction tuning, starting from an instruct model (e.g., Llama-3.1-8B-Instruct) rather than the raw base usually works better — it already follows instructions, so your fine-tuning refines rather than teaches the format. The exception: if the instruct variant's alignment conflicts with your task (rare), use the base.

Document the probe results in a small table in your paper's appendix. "We selected Qwen2.5-7B-Instruct after a 200-example probe (62.1 vs. 54.3 for Llama-3.1-8B)" is one sentence that preempts the reviewer's "why this model?" question entirely.

Debugging Installation Issues (The Usual Suspects)

Environment setup fails in predictable ways. Here are the four most common, with fixes:

1. bitsandbytes CUDA errors. The classic: CUDA SETUP: ... library not found or version mismatch warnings. bitsandbytes ships precompiled against specific CUDA versions; Colab's CUDA usually works with the default pip install, but on custom machines you may need pip install bitsandbytes --no-cache-dir or the source build. First check: does python -c "import bitsandbytes" succeed? If the import works, the 4-bit loading in Step 1 of Chapter 6 will almost certainly work too.

2. torch / transformers version conflicts. Symptom: import errors mentioning accelerate or missing trainer arguments (e.g., eval_strategy vs the older evaluation_strategy). The ecosystem renames arguments across versions — eval_strategy is current; older code uses evaluation_strategy. When copying code from blog posts, check which transformers version it assumed. Your pinned requirements file (Chapter 5's reproducibility section) is the defense.

3. "Model not found" on gated models. You accepted the license but downloads still fail → you forgot huggingface-cli login, or you're logged in with a different account than the one that accepted the license. Run huggingface-cli whoami to verify. For non-gated alternatives needing no approval (Phi-3-mini, Qwen2.5), skip the gate entirely while learning.

4. Disk full during download. The 8B download stalls partway. Check df -h; clear the pip cache and old checkpoints. On Colab, a factory reset (Runtime → Disconnect and delete runtime) gives a clean slate — then reinstall from your pinned requirements in one cell.

General rule: fix the environment once, then freeze it. Every hour spent fighting installs after the pipeline works is an hour stolen from data and evaluation.

For your research: Decide your tooling before the main experiments and freeze versions (pip freeze > requirements.txt, or a Docker/container spec). Note the exact library versions in your paper's appendix or repository README. "Transformers 4.x" is not reproducible; "transformers==4.44.2, peft==0.12.0, trl==0.9.6" is. Reviewers increasingly check this, and future-you will need it when a library update breaks your old scripts two weeks before a deadline.

Key Takeaways

  • Standard stack: Transformers (foundation) → PEFT (adapters) → TRL/SFTTrainer (instruction-tuning trainer) → Axolotl (config-driven experiments).
  • SFTTrainer + PEFT is the productive default for instruction tuning; Axolotl YAML configs are the reproducible default for paper experiments.
  • Setup: PyTorch with CUDA, pip install transformers datasets accelerate peft trl bitsandbytes, verify GPU + 4-bit loading before anything else.
  • Accept gated-model licenses and huggingface-cli login once; watch Colab disk space.
  • Libraries handle mechanics (tokenization, masking, checkpointing); you handle data, evaluation, hyperparameters, and interpretation.
  • Freeze and record exact library versions — reproducibility starts with the environment.

Chapter 6: Training on One GPU / Free Colab: A Complete Walkthrough

This is the chapter where theory becomes a running model. We will fine-tune an 8-billion-parameter instruction model with QLoRA on a single 16 GB GPU — the kind Google Colab gives you for free — from data loading to a saved, usable adapter. Follow along in a Colab notebook (GPU runtime: Runtime → Change runtime type → T4 GPU) or any machine with a CUDA GPU. Every step is explained; nothing is magic.

Step 0: Setup and Sanity Checks

Open a fresh Colab notebook with a GPU runtime and run:

# Install the stack (takes a few minutes)
!pip install -q transformers datasets accelerate peft trl bitsandbytes

import torch
assert torch.cuda.is_available(), "No GPU! Check Runtime > Change runtime type."
print(torch.cuda.get_device_name(0))
print(f"VRAM: {torch.cuda.get_device_properties(0).total_memory / 1e9:.1f} GB")

You should see something like Tesla T4 and VRAM: 15.9 GB. If Colab gives you a weaker GPU, the walkthrough still works — training will just be slower.

Step 1: Load the Base Model in 4-Bit

We will use a small, strong open model. Llama-3.1-8B is the canonical choice (requires free Hugging Face access approval); alternatives that need no approval include Microsoft's Phi-3-mini or Google's Gemma-2-2B. Pick one and stick with it.

from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig

MODEL_ID = "meta-llama/Llama-3.1-8B"   # or "microsoft/Phi-3-mini-4k-instruct"

bnb_config = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_quant_type="nf4",
    bnb_4bit_use_double_quant=True,
    bnb_4bit_compute_dtype=torch.bfloat16,
)

model = AutoModelForCausalLM.from_pretrained(
    MODEL_ID,
    quantization_config=bnb_config,
    device_map="auto",
)
tokenizer = AutoTokenizer.from_pretrained(MODEL_ID)
tokenizer.pad_token = tokenizer.eos_token   # causal LMs often lack a pad token; eos works
tokenizer.padding_side = "right"

Watch the memory after loading: !nvidia-smi should show roughly 5–6 GB used. The 8B weights in 4-bit occupy under 5 GB; the rest is overhead.

Step 2: Prepare the Dataset

For this walkthrough we will build a tiny demonstration dataset — 300 examples teaching the model to answer questions in a strict format (a short answer followed by a one-line justification). In your real project this is where Chapter 3's pipeline goes. Here, we generate a toy dataset so the mechanics are visible end to end:

from datasets import Dataset

# Toy data: in your project, replace this with your real curated dataset
topics = ["photosynthesis", "gravity", "elections", "vaccines", "inflation"]
def make_example(i):
    t = topics[i % len(topics)]
    return {
        "instruction": f"Explain {t} in simple terms.",
        "input": "",
        "output": f"{t.capitalize()} is a process where ... [short answer].\nWhy: [one-line justification]."
    }

raw = [make_example(i) for i in range(300)]

def to_chat(ex):
    # SFTTrainer works best with chat-formatted messages
    messages = [
        {"role": "user", "content": ex["instruction"] + (" " + ex["input"] if ex["input"] else "")},
        {"role": "assistant", "content": ex["output"]},
    ]
    return {"messages": messages}

dataset = Dataset.from_list([to_chat(ex) for ex in raw])
dataset = dataset.train_test_split(test_size=0.1, seed=42)
print(dataset)
# DatasetDict({train: 270 rows, test: 30 rows})

Two things to notice: we use the chat format (lists of role/content messages) because modern tokenizers apply the model's native chat template via apply_chat_template, and we split off a validation set (the 30 rows) that training will never see. In a real project your dataset would be thousands of examples built per Chapter 3 — the code shape is identical.

Step 3: Attach the LoRA Adapter

from peft import LoraConfig, get_peft_model, prepare_model_for_kbit_training

model = prepare_model_for_kbit_training(model)

peft_config = LoraConfig(
    r=16,
    lora_alpha=32,
    lora_dropout=0.05,
    bias="none",
    task_type="CAUSAL_LM",
    target_modules=["q_proj", "k_proj", "v_proj", "o_proj",
                    "gate_proj", "up_proj", "down_proj"],
)

model = get_peft_model(model, peft_config)
model.print_trainable_parameters()

Expected output: something like trainable params: 41,943,040 || all params: 8,031,261,440 || trainable%: 0.52%. Half a percent. That tiny fraction is everything you are about to train.

Step 4: Configure and Launch Training with SFTTrainer

from trl import SFTTrainer, SFTConfig

training_args = SFTConfig(
    output_dir="./outputs/walkthrough-01",
    num_train_epochs=3,
    per_device_train_batch_size=2,
    gradient_accumulation_steps=8,   # effective batch = 2*8 = 16
    learning_rate=2e-4,
    lr_scheduler_type="cosine",
    warmup_steps=30,
    optim="paged_adamw_8bit",        # 8-bit optimizer: less memory, per Dettmers et al.
    max_seq_length=1024,
    packing=False,                   # keep False while learning; True later for speed
    eval_strategy="steps",
    eval_steps=50,
    save_steps=100,
    save_total_limit=2,
    logging_steps=10,
    report_to="none",                # set to "tensorboard" if you want live plots
    seed=42,
)

trainer = SFTTrainer(
    model=model,
    args=training_args,
    train_dataset=dataset["train"],
    eval_dataset=dataset["test"],
)
trainer.train()

What happens now: the trainer tokenizes your messages with the chat template, masks the user turns so loss is computed only on assistant responses, and runs the optimization loop. On a T4, 270 examples × 3 epochs takes roughly 10–25 minutes. You will see the training loss printed every 10 steps — it should trend downward, noisily. The eval loss every 50 steps tells you whether the model is actually improving on unseen examples (Chapter 9 explains how to read these curves).

If you run out of memory: the three emergency levers are (1) per_device_train_batch_size=1, (2) max_seq_length=512, and (3) gradient_checkpointing=True (trades compute for memory by recomputing activations). Apply in that order.

Step 5: Save the Adapter and Test It

trainer.save_model("./outputs/walkthrough-01/final")
# This saves ONLY the adapter (~80-150 MB), not the 8B base model.

Now the moment of truth — generate with the fine-tuned model and compare against the base:

model.eval()
prompt = "Explain gravity in simple terms."
inputs = tokenizer.apply_chat_template(
    [{"role": "user", "content": prompt}],
    return_tensors="pt", add_generation_prompt=True,
).to(model.device)

with torch.no_grad():
    out = model.generate(**{"input_ids": inputs}, max_new_tokens=120, do_sample=False)
print(tokenizer.decode(out[0], skip_special_tokens=True))

You should see the response follow your trained format (short answer + "Why:" line). To feel the difference, load the base model without the adapter and run the same prompt — the base model answers in its generic style. That contrast, however small on toy data, is fine-tuning made visible.

Step 6: Reloading and Merging (Optional but Useful)

To use the adapter later without retraining:

from peft import PeftModel

base = AutoModelForCausalLM.from_pretrained(MODEL_ID, quantization_config=bnb_config, device_map="auto")
model = PeftModel.from_pretrained(base, "./outputs/walkthrough-01/final")

To produce a single standalone model (e.g., for deployment), merge the adapter into the base weights in 16-bit and save:

# Load base in 16-bit (no quantization) for a clean merge — needs more VRAM/RAM
base16 = AutoModelForCausalLM.from_pretrained(MODEL_ID, torch_dtype=torch.bfloat16, device_map="auto")
merged = PeftModel.from_pretrained(base16, "./outputs/walkthrough-01/final").merge_and_unload()
merged.save_pretrained("./outputs/walkthrough-01/merged")

Merging needs the full model in 16-bit (~16 GB for 8B), which may exceed Colab VRAM — do it on a bigger machine or accept the adapter format, which is the standard for research sharing anyway.

What Just Happened, End to End

You downloaded an 8B model, quantized it to 4-bit, attached a 42M-parameter trainable adapter (0.5% of weights), trained it on 270 examples with loss masked to assistant responses, validated on held-out examples, and saved a ~100 MB adapter that changes the model's behavior. Total cost: zero rupees, one Colab session. This is the complete loop that every fine-tuning project — from student experiments to industry pipelines — follows. Everything else in this book is about doing each step well: better data (Ch. 3), better hyperparameters (Ch. 7), honest evaluation (Ch. 8), and credible reporting (Ch. 12).

Reading Your First Training Run: A Guided Tour of the Logs

When you launch trainer.train(), you'll see a stream of log lines. Here's how to read them like a practitioner, using a typical healthy run as reference:

Step 10:  loss 2.84, lr 6.7e-05
Step 20:  loss 2.31, lr 1.3e-04
Step 30:  loss 1.92, lr 2.0e-04   <- warmup ends, LR at full value
Step 40:  loss 1.75, lr 1.99e-04
Step 50:  loss 1.61, eval_loss 1.70
Step 60:  loss 1.58, lr 1.98e-04
...
Step 200: loss 0.94, eval_loss 1.12

What healthy looks like: loss starts high (the model is surprised by your data's format — 2.5–3.5 is typical for a new task), drops steeply in the first 10–20% of steps (the model quickly learns the format), then declines gradually. Eval loss tracks training loss with a small gap. The learning rate ramps up during warmup (first 30 steps here) then barely visibly decays under cosine scheduling.

Checkpoints to watch for: - Steps 1–30 (warmup): loss should fall, not spike. A spike here with recovery is usually harmless; a spike without recovery means LR too high. - First eval (step 50): eval_loss moderately above train loss is normal. Eval loss far above train loss (2× or more) this early suggests a data or template bug — e.g., eval examples formatted differently. - Mid-training: both losses falling slowly. This is the long middle where the model absorbs the task. Resist the urge to intervene; noise is normal. - Late training: train loss still falling, eval loss flattening or ticking up — you're entering the overfitting zone (Chapter 9). Note the step where eval was lowest; that's your candidate best checkpoint.

When to abort a run (and it's always okay to abort): loss NaN at any point; loss flat after 15% of steps (something structural is wrong — check the data pipeline before burning more compute); eval loss 3× train loss (template/label bug); or GPU memory errors that persist after the OOM triage. A two-hour run aborted at minute 20 with a lesson learned is cheap tuition. A two-hour run left to finish broken is just waste.

The qualitative spot-check: every ~100 steps, generate from 2–3 fixed prompts (keep them constant across the whole run) and eyeball the outputs. Numbers tell you that it's learning; outputs tell you what it's learning. If the loss falls but outputs get worse (more repetitive, format drifting), trust the outputs — the metric is lying about what you care about, and you need a better eval (Chapter 8).

Extending the Walkthrough: Packing, Resuming, and Longer Runs

Once the basic walkthrough works, three upgrades make it production-grade for real experiments:

Sequence packing. In the walkthrough, packing=False means each training example is padded to max_seq_length — with short examples, most of the GPU's work is processing padding tokens, pure waste. Setting packing=True concatenates multiple short examples into each sequence (with proper attention masking so examples don't attend to each other). Throughput often nearly doubles on short-example datasets. Enable it only after the pipeline is debugged — packing makes per-example debugging harder, which is why the walkthrough starts without it.

Resuming from checkpoints. Long Colab sessions disconnect. Because we set save_steps=100 and save_total_limit=2, interrupted runs leave checkpoints behind. Resume with:

trainer.train(resume_from_checkpoint="./outputs/walkthrough-01/checkpoint-500")

The trainer restores model weights, optimizer state, and the data loader position — training continues as if uninterrupted. Make resuming your default reflex after any disconnection rather than restarting from scratch.

Scaling to real data sizes. The walkthrough's 270 examples finish in minutes; your real dataset of 5,000–20,000 examples needs hours. Two adjustments: raise eval_steps and save_steps proportionally (evaluating every 50 steps on a 5,000-example run wastes time — every 200–500 steps is plenty), and lower logging_steps noise by watching trends rather than individual values. Everything else — the code shape, the adapter, the eval discipline — is identical. That's the point of the walkthrough: the toy run and the thesis run differ in scale, not in structure.

A note on do_sample=False in evaluation. The walkthrough generates with greedy decoding (always the most likely token) for determinism — the same prompt gives the same output every time, which is what you want when comparing models. For qualitative exploration (getting a feel for the model's style), sampling with temperature=0.7 shows the range of behaviors. Use greedy for measurement, sampling for exploration, and never mix them within one comparison.

Sharing Your Adapter on the Hugging Face Hub

Your trained adapter is a research artifact worth sharing — and the Hub makes it a three-line operation:

from peft import PeftModel

# Push the adapter (only ~100 MB) to the Hub
model.push_to_hub("your-username/my-task-adapter-v1")
tokenizer.push_to_hub("your-username/my-task-adapter-v1")

Anyone can then reproduce your results without retraining:

from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer

base = AutoModelForCausalLM.from_pretrained("meta-llama/Llama-3.1-8B", device_map="auto")
tokenizer = AutoTokenizer.from_pretrained("your-username/my-task-adapter-v1")
model = PeftModel.from_pretrained(base, "your-username/my-task-adapter-v1")

Three practices make a shared adapter genuinely useful rather than a mysterious binary: (1) fill in the model card — base model, dataset description, training config, evaluation scores, intended use and limitations; (2) version your uploads (-v1, -v2) instead of overwriting, so citations remain stable; (3) link the adapter from your paper's repository README alongside the eval script. A well-documented adapter gets downloaded, cited, and built upon — the closest thing fine-tuning research has to compounding interest. (Note: only push adapters trained on data you have the right to share, and never push datasets containing private information — Chapter 3's licensing and privacy notes apply here too.)

For your research: Run this walkthrough verbatim before touching your real data. It validates your entire environment (GPU, installs, disk, versions) in under an hour and gives you a known-good training script to adapt. Keep the notebook — when your real run breaks at 2 AM, you will diff against this working baseline to find what changed. Every serious practitioner keeps a "hello world" training script for exactly this reason.

Key Takeaways

  • A full QLoRA fine-tune of an 8B model fits on a free Colab T4: 4-bit base (~5 GB) + tiny adapter + 8-bit optimizer.
  • The loop: load quantized base → prepare data in chat format with a held-out split → attach LoRA → SFTTrainer → save adapter.
  • SFTTrainer handles chat templates and response-only loss masking; you supply data and hyperparameters.
  • OOM triage, in order: batch size 1 → shorter max_seq_length → gradient checkpointing.
  • The adapter (~100 MB) is your portable artifact; merging is optional and needs 16-bit VRAM.
  • Always build a working toy run first — it becomes your debugging baseline for the real project.

Chapter 7: Hyperparameters That Actually Matter (LR, Rank, Epochs, Batch Size)

Hyperparameters are the settings you choose before training that control how the model learns: how big each step is, how many trainable parameters the adapter has, how many times the model sees your data, how many examples it processes at once. Beginners treat them as mystical incantations copied from blog posts. Researchers treat them as experimental variables with known effects, sensible defaults, and a cheap tuning procedure. This chapter makes you the second kind.

The Four That Matter Most

Dozens of knobs exist; four dominate outcomes in QLoRA instruction tuning.

1. Learning Rate (LR): The Step Size

The learning rate scales every weight update (Chapter 2). Too small and the model barely learns — training loss crawls down, you waste compute, and the model underfits your data. Too large and updates overshoot: loss spikes, oscillates, or explodes to NaN, and the model can be damaged — fluent capabilities degrade because you blasted the weights out of their good valley.

Sensible range for QLoRA instruction tuning: 1e-4 to 3e-4 (0.0001 to 0.0003). The QLoRA paper used 2e-4, and it remains the default starting point across the field. Full fine-tuning uses much smaller values (1e-5 to 5e-5) because it updates all weights — with LoRA's tiny parameter count, a larger LR is safe and needed.

How to diagnose LR problems from the loss curve: - Loss decreases smoothly then plateaus → LR is fine; you may need more epochs or data. - Loss barely moves after hundreds of steps → LR too small (or data/adapter problem). Try 3× larger. - Loss oscillates wildly or spikes upward → LR too large. Halve it and restart. - Loss becomes NaN → LR far too large, or mixed-precision instability. Halve LR; if it persists, check data for corrupt examples (a single garbage example with extreme tokens can do this).

The cheap LR test: run 200–300 steps at 1e-4, 2e-4, and 4e-4 on a subset of data, and pick the one with the lowest validation loss. Three short runs beat one long guess.

2. LoRA Rank (r): The Adapter's Capacity

Rank controls how many trainable parameters the adapter has (Chapter 4): roughly, capacity scales with r. Too small and the adapter cannot express your task's patterns — validation performance plateaus below what the data supports (underfitting from capacity). Too large and the adapter memorizes training examples instead of generalizing, especially on small datasets (overfitting from capacity), while also slowing training and inflating the adapter file.

Defaults: r=8 or 16 for format/style tasks; r=32–64 for hard domain adaptation. Keep alpha proportional (alpha = 2× r is the common convention, so the alpha/r scaling stays constant).

How to tune: fix everything else, train with r ∈ {8, 16, 32}, and compare validation loss and a task metric (Chapter 8). If r=32 beats r=8 substantially, your task needs capacity — try 64. If all three are similar, keep 8 or 16 (smaller generalizes better and trains faster). Report this ablation; it is a one-paragraph result reviewers respect.

3. Epochs: How Many Passes Over the Data

One epoch = one full pass through the training set. Too few epochs and the model hasn't absorbed the patterns (underfitting). Too many and it starts memorizing specific examples — training loss keeps falling while validation loss rises (overfitting, Chapter 9).

Defaults: 2–5 epochs for instruction tuning on thousands of examples. With very small datasets (a few hundred), even 3 epochs can overfit — watch validation loss. With very large datasets (100k+), 1–2 epochs often suffice because each epoch contains so many updates.

The right procedure: train for more epochs than you need (say 5–8) with checkpointing (save every N steps), then pick the checkpoint with the best validation score — this is early stopping done manually. Never pick the final checkpoint by default; the best one is often in the middle. SFTTrainer/Axolotl can automate this with load_best_model_at_end=True plus a metric.

4. Batch Size: How Many Examples Per Update

The batch size is how many examples the model processes before each weight update. Larger batches give smoother, more reliable gradient estimates (less noise per step) and better GPU utilization — but use more memory. Smaller batches are noisier, which can actually help escape sharp, overfit minima, but train less efficiently.

The memory trick — gradient accumulation: you can simulate a large batch on a small GPU. Process micro_batch_size examples at a time, accumulate gradients over gradient_accumulation_steps mini-batches, then update once. Effective batch = micro × accumulation × (number of GPUs). Chapter 6 used 2 × 8 = 16. Effective batch sizes of 16–128 are the normal range for QLoRA instruction tuning.

Practical rule: set micro batch to the largest that fits in memory (try 4, then 2, then 1), then set accumulation steps to reach your target effective batch. If you change the effective batch substantially, consider scaling the learning rate mildly (larger batch → slightly larger LR is the textbook guidance, though in the 16–128 range most practitioners keep LR fixed).

The Second Tier: Still Worth Setting Deliberately

  • LR scheduler: how LR changes during training. Cosine decay (LR smoothly falls to near zero) is the default and works well. Linear warmup (LR ramps up over the first ~3–5% of steps) prevents early instability — always use some warmup (e.g., 30–100 steps).
  • Optimizer: paged_adamw_8bit is the QLoRA-era default — AdamW with 8-bit optimizer states (less memory) and paging (spills to CPU RAM gracefully instead of crashing on spikes). Use it unless you have a reason not to.
  • Weight decay: mild regularization (0.01 is common) that discourages weights from growing large. Small effect in LoRA; leave at default.
  • Dropout (lora_dropout): 0.05 is standard; raise to 0.1 if you see overfitting on small data.
  • Max sequence length: truncate/pad examples to this. Longer = more memory. Set it to cover ~95% of your examples (check your data's length distribution; don't pay for 4096 tokens when 95% of examples are under 800).
  • Seed: fix it (42 is tradition) for reproducibility. For final results, ideally train with 2–3 seeds and report mean ± std — seed variance is real and reviewers know it.

A Principled Tuning Procedure (Cheap and Honest)

You cannot grid-search everything — each full run costs hours. Use this staged procedure:

  1. Fix data and evaluation first (Chapters 3, 8). Tuning hyperparameters against a broken eval is worse than useless.
  2. Coarse pass on a data subset: train on ~20% of data for 1 epoch across LR ∈ {1e-4, 2e-4, 4e-4}. Pick the best LR by validation loss. (~3 short runs)
  3. Rank ablation: at the chosen LR, try r ∈ {8, 16, 32} on the full data for 2–3 epochs. Pick by validation task metric. (~3 medium runs)
  4. Epoch selection: train the winner for 5+ epochs with checkpointing; pick the best checkpoint by validation. (1 long run)
  5. Final run: retrain the winning configuration with 2–3 seeds; report mean ± std on the locked test set. (2–3 long runs — this is the number in your paper)

Total: roughly 8–10 training runs, most of them short. This is affordable on Colab (spread over days) or a single cloud GPU (hours), and it is exactly the methodology section reviewers want to read.

Reading the Loss Curve

Training loss curves: healthy decline vs overfitting

Figure: Training loss (falling) versus validation loss (falling, then rising) — the signature of overfitting.

Learn to read these two curves at a glance:

  • Both falling: healthy learning. Continue.
  • Training falling, validation flat: the model has learned what it can from this data/adapter; more epochs won't help. Try more data or higher rank.
  • Training falling, validation rising: overfitting — stop, use an earlier checkpoint, add regularization, or get more data.
  • Both flat from the start: LR too small, or a data/template bug (e.g., labels misaligned). Debug before burning compute.
  • Sudden spike then recovery: usually a bad batch; if it recovers, ignore. If spikes repeat, lower LR.

Log both curves from step one (TensorBoard or even the printed logs). The number of student projects that trained blind for six hours and discovered a flat line at the end is too large to count. Don't join them.

Interaction Effects: Why Hyperparameters Can't Be Tuned Independently

The staged tuning procedure in this chapter tunes one thing at a time, which is practical — but hyperparameters interact, and knowing the main interactions prevents confusing results:

Learning rate × batch size. Larger batches produce less noisy gradients, which tolerate (and sometimes benefit from) slightly larger learning rates. If you double your effective batch, nudging LR up by ~1.4× (square-root scaling, the common heuristic) is reasonable — but in the 16–128 batch range, most practitioners keep LR fixed and do fine. The dangerous direction is shrinking the batch (e.g., to fit memory) while keeping LR high: noisier gradients + big steps = instability. If you cut batch size 4× for memory reasons, consider halving LR as compensation.

Rank × data size. Small data + high rank is the classic overfitting recipe (Chapter 9): the adapter has capacity to memorize every example. The interaction runs both ways — with 50,000 diverse examples, r=64 may generalize beautifully where it memorized on 1,000. Rule of thumb: scale rank with data, not with ambition. If you increase data 10×, revisiting rank upward is legitimate; if data stays fixed, higher rank mostly buys memorization.

Epochs × learning rate. More epochs at a high LR drives the model further from base behavior (more forgetting, Chapter 9); fewer epochs at a low LR may underfit. The product that matters is roughly "total distance traveled" — LR × steps. This is why the cosine schedule (which decays LR to near zero) pairs naturally with fixed epoch counts: late epochs take tiny steps, so extra epochs cost less in forgetting than they would at constant LR.

Sequence length × batch size (the memory interaction). GPU memory is shared between them: doubling max sequence length roughly doubles activation memory, forcing you to halve the batch. Since most examples are short (Chapter 7's 95th-percentile advice), cutting max length is usually the cheaper side of this trade — you keep batch size (stable gradients) and only truncate the rare long example.

What this means for your tuning procedure: the staged approach works because these interactions are second-order for typical ranges — tuning LR first, then rank, then epochs gets you to a good neighborhood. But when you change something big (10× data, different model size, 4× batch change), re-check the earlier stages rather than assuming they transfer. And in your paper, report the final configuration as a joint choice ("selected by staged search; LR 2e-4 and rank 16 were co-validated") rather than implying each was optimized in isolation — reviewers know about interactions, and acknowledging them reads as competence.

When Defaults Fail: A Diagnostic Decision Tree

The cheat sheet gives starting points; this tree handles the cases where they don't work. Start at the top with your symptom:

Training loss won't go below ~2.5 and outputs look untrained. → Is the loss exactly flat, or noisy-flat? Exactly flat: check label masking (is the model training on empty responses? print one tokenized batch with labels). Noisy-flat: LR too small — try 5e-4. Still flat: your data may not be in the format the template expects — decode 3 training examples fully and read them.

Training loss falls beautifully but validation is terrible from the start. → Data distribution mismatch between splits (check: are val examples systematically different — longer, harder, different template?). Or label leakage in reverse: the model learned a spurious training-set pattern. Stratify your split by the important dimensions (topic, length, difficulty) instead of pure random.

Everything worked on the subset but the full run diverged. → The subset wasn't representative (too easy, too clean). Or: you scaled epochs without scaling warmup — warmup steps should scale with total steps; 30 warmup steps for a 3,000-step run is proportionally nothing. Set warmup to ~3–5% of total steps.

Results vary wildly between seeds (±5 points). → Your test set is too small or your task too noisy for the claimed precision. Either enlarge the test set, or report the variance honestly and shrink your claims. Seed variance is a finding — it tells you the task is unstable, which is worth one paragraph.

The model got worse than the base model. → Almost always one of: LR far too high (check for the loss spike signature), training on the wrong objective (e.g., full next-token loss including instructions, teaching the model to generate instructions instead of following them), or catastrophic data corruption (duplicated examples, wrong language). Roll back to the walkthrough baseline and diff every setting.

Work this tree before asking for help or buying more compute — nine times out of ten, the answer is in it.

For your research: Your paper should contain a hyperparameters table (the full config) and ideally one ablation figure — e.g., validation performance vs. LoRA rank, or vs. data size. These are cheap to produce during the tuning procedure above and they transform "we fine-tuned a model" into "we systematically studied fine-tuning for this task." The difference is the difference between a workshop paper and a conference paper.

Key Takeaways

  • The big four: learning rate (start 2e-4), LoRA rank (start 16), epochs (2–5, pick by checkpoint), effective batch size (16–128 via accumulation).
  • Diagnose from curves: flat = LR too small or data bug; spiking = LR too large; val rising while train falls = overfitting.
  • Second tier: cosine schedule + warmup, paged_adamw_8bit, dropout 0.05, max length matched to your data, fixed seed.
  • Tune in stages: coarse LR on subset → rank ablation → epoch/checkpoint selection → multi-seed final runs. ~8–10 runs total.
  • Always log training and validation loss; never train blind.
  • Publish the full config table and at least one ablation figure.

Chapter 8: Evaluating Fine-Tuned Models (Before/After Comparisons)

Training produces a model. Evaluation tells you whether the model is actually better — and "better" is a claim that requires evidence, not enthusiasm. This chapter is about building that evidence: what to measure, what to compare against, and how to avoid the self-deception that fine-tuning makes so easy. If Chapter 3's dataset is the steering wheel, evaluation is the dashboard. Driving blind is not brave; it is just blind.

The Golden Rule: Compare Against the Right Baselines

A fine-tuned model's score, alone, means nothing. "Our model scores 82%" is empty until you know what the base model scored (maybe 81%), what a good prompt scored (maybe 84%), and what chance scored (maybe 50%). Every evaluation needs baselines, and for fine-tuning papers the expected set is:

  1. The base model, zero-shot: the unmodified model with a plain instruction. This measures what fine-tuning added over the starting point.
  2. The base model, few-shot: the base model with 3–10 examples in the prompt. This is the strongest cheap baseline — it tells you whether your fine-tuning beat what prompting alone could do (Chapter 1's whole point).
  3. The base model + RAG (if applicable): if your task involves knowledge, compare against retrieval too.
  4. Ablations of your own method: your model with key components removed or varied (different rank, noisier data, fewer examples). These show which part of your approach caused the improvement.

Report all of them in one table. The pattern reviewers look for — and the pattern that makes a paper convincing — is: base < few-shot < your fine-tune, with ablations showing the gain survives when you vary the details. If few-shot prompting matches your fine-tune, say so honestly; that is itself a publishable finding ("for this task, prompting suffices"), and claiming otherwise will not survive review.

What to Measure: Metrics by Task Type

Match the metric to what you actually care about:

  • Classification / multiple choice: accuracy, plus F1 when classes are imbalanced. Report per-class scores if some classes matter more.
  • Structured generation (JSON, SQL, code): exact-match rate or execution accuracy (does the SQL run and return the right answer? does the code pass tests?). Prefer execution-based metrics over string matching — they measure what you actually want.
  • Open-ended generation (summaries, answers, explanations): this is the hard one. Automated metrics like ROUGE or BLEU measure word overlap with a reference, which correlates weakly with quality. Better options:
  • LLM-as-judge: a strong model rates outputs against a rubric. Cheap and scalable, but biased (judges favor their own style, longer answers, and the first-presented option). Mitigate with blinded, order-swapped comparisons and a human-validated subset.
  • Human evaluation: the gold standard. Even 100–200 examples rated by 2–3 annotators with measured agreement beats 10,000 LLM-judged examples for credibility. Report inter-annotator agreement.
  • Task-grounded checks: for instruction following, verify format compliance programmatically (valid JSON? correct schema? required sections present?). These are objective and cheap — use them wherever the task has checkable structure.
  • Behavioral suites: for instruction-tuned models generally, public benchmarks (MMLU for knowledge, MT-Bench or AlpacaEval-style for instruction following, IFEval for format compliance) test whether your fine-tuning helped or hurt broadly. Use them as regression tests (Chapter 9).

One metric is never enough. Report a small panel: your task metric, a format-compliance check, and one broad benchmark for regressions. Three numbers tell a story; one number tells a slogan.

The Test Set Discipline

From Chapter 3: your test set is locked until the final evaluation. Here is why this matters so much in fine-tuning specifically. During development you will look at validation results dozens of times, adjusting data, prompts, and hyperparameters. Every adjustment informed by validation leaks a little information about validation into your model. By the end, your validation score is optimistic — you have, in effect, trained on it indirectly. The locked test set is the only unbiased estimate you have. Evaluate it once, on the final model, and report whatever it says — even if it is worse than validation. Reporting the worse number honestly is what separates research from marketing.

Additional hygiene:

  • No train/test overlap: check for duplicate or near-duplicate examples across splits (normalize whitespace/case, then compare). Paraphrased test questions that appear verbatim in training are contamination, not generalization.
  • Multiple seeds: training is stochastic. Report mean and standard deviation across 2–3 seeds, not the best seed's score. A "gain" smaller than the seed variance is not a gain.
  • Statistical testing: for close comparisons, a paired test (e.g., McNemar's for classification, paired bootstrap for generation metrics) tells you whether the difference is real. You don't need heavy statistics — just don't claim victory on a 0.5% gap from a single run.

Qualitative Analysis: Read the Outputs

Numbers hide as much as they reveal. For every major experiment, read 50–100 outputs by hand — from the base model and the fine-tuned model side by side — and categorize the differences. You will discover things no metric shows: the fine-tuned model fixed formatting but introduced a new hedging tic; it answers correctly but dropped the citation style you wanted; it improved on common cases but got worse on rare ones.

Build a small error taxonomy: the 4–6 most common failure modes, with example counts before and after fine-tuning. "Hallucinated clause numbers: 31/100 → 4/100; wrong output language: 12/100 → 2/100; refused valid requests: 0/100 → 9/100" — this table is worth more than three paragraphs of prose, and it directly suggests your next experiment (the refusal spike means your data over-represents refusals — fix the data mix).

A Concrete Evaluation Script

import json, re

# Load locked test set (untouched until now)
test = [json.loads(l) for l in open("data/test.json")]

def format_ok(response):
    # Example task-grounded check: response must be valid JSON with required keys
    try:
        obj = json.loads(response)
        return "answer" in obj and "justification" in obj
    except Exception:
        return False

def evaluate(model, tokenizer):
    correct, wellformed, total = 0, 0, 0
    for ex in test:
        prompt = ex["instruction"]  # adapt to your template
        inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
        out = model.generate(**inputs, max_new_tokens=200, do_sample=False)
        resp = tokenizer.decode(out[0], skip_special_tokens=True)
        total += 1
        if format_ok(resp):
            wellformed += 1
        # task-specific correctness check goes here
    return {"n": total, "wellformed_rate": wellformed/total}

# Run for: base zero-shot, base few-shot, fine-tuned — same script, same test set
for name, model in [("base", base_model), ("finetuned", ft_model)]:
    print(name, evaluate(model, tokenizer))

The discipline that matters: the same script, the same test set, the same decoding settings (temperature, max tokens) for every model compared. Changing decoding between baseline and fine-tuned runs invalidates the comparison silently.

Presenting Results: The Before/After Table

Your paper's results section centers on one table. Columns: model/baseline. Rows: metrics. Something like:

Model Task accuracy Format compliance MMLU (regression)
Base, zero-shot 54.2 ± 1.1 61% 68.4
Base, few-shot (5 ex) 71.8 ± 0.9 83% 68.4
Fine-tuned (ours) 78.5 ± 0.7 97% 67.9

This fictional-but-typical table tells the whole story at a glance: fine-tuning beat few-shot prompting on the task (+6.7), massively improved format reliability, and barely dented general knowledge (−0.5 MMLU, the forgetting check from Chapter 9). Write the paragraph around the table, not instead of it.

LLM-as-Judge, Done Responsibly

Using a strong model to evaluate your fine-tuned model's outputs is now standard practice — it's the only scalable option for open-ended generation. But naive LLM judging is biased in documented ways, and reviewers know them. Here's how to do it defensibly:

The known biases (and mitigations): - Position bias: judges favor the first-presented answer. Mitigation: present pairs in both orders and average, or use absolute scoring (rate each output 1–5 alone) instead of pairwise comparison. - Verbosity bias: judges reward longer answers. Mitigation: instruct the judge to penalize unnecessary length explicitly in the rubric, or control for length statistically. - Self-preference: a judge from the same model family as your fine-tuned model favors its sibling's style. Mitigation: use a judge from a different family than your base model, and disclose which judge you used. - Rubric vagueness: "rate quality 1–10" produces noise. Mitigation: write a concrete rubric with anchored descriptions ("5 = correct answer, correct format, cites clause; 3 = correct answer, wrong format; 1 = wrong answer").

The validation step you must not skip: have humans rate 100–200 of the same outputs and measure agreement between human and LLM judge (correlation for scores, Cohen's kappa for categories). Report this number. "Our LLM judge agrees with human raters at r=0.81 on a 200-example validation set" transforms the judge from a black box into a calibrated instrument. Without it, a skeptical reviewer can dismiss all your judged results in one sentence.

Blinding: the judge must not know which output came from which model — strip model identifiers, randomize order, and use identical formatting for both candidates. This sounds obvious; it's skipped surprisingly often.

Cost control: judging 10,000 outputs with a frontier API model gets expensive. Judge a random, pre-registered subset (e.g., 500–1,000) with the strong judge, use programmatic checks (Chapter 8's format verification) on the full set, and human-rate 100–200. Three tiers, each doing what it's best at: programs for scale, LLM judges for nuance at moderate scale, humans for ground truth.

Statistical Rigor on a Student Budget

"Mean ± std across seeds" was Chapter 8's advice; here's how to actually compute and interpret it without a statistics degree:

How many seeds? Three is the practical minimum (42, 43, 44 — or any three you fix in advance). Five is better but costs 5× the final-run compute. Never report the best of N seeds as "the result" — that's selection bias, and reviewers check for it by asking whether you reported the max or the mean.

Reading the interval. Suppose your fine-tuned model scores 78.5 ± 0.7 and few-shot scores 71.8 ± 0.9. The intervals don't overlap — the improvement is real beyond seed noise. But 78.5 ± 0.7 vs. 78.1 ± 0.8 do overlap substantially — that "gain" is indistinguishable from luck. Don't claim it; report both numbers and move on. Overlapping intervals don't prove no difference, only that your experiment can't resolve one — a paired test (below) is more sensitive.

A simple paired test for classification. When comparing two models on the same test set, McNemar's test asks: on examples where the models disagree, does one win significantly more often? It's a one-function call in most stats libraries and far more appropriate than eyeballing percentages. For generation metrics, a paired bootstrap (resample the test set with replacement 1,000 times, compute the score difference each time, check whether 95% of differences favor your model) achieves the same thing with no distributional assumptions. You don't need to derive these — just use them and report "paired bootstrap, p < 0.05."

What "significant" doesn't mean. Statistical significance is not practical significance. A 0.4% gain with p < 0.01 on 50,000 test examples is real but probably irrelevant; a 6% gain with p = 0.08 on 200 examples is suggestive but unresolved — collect more test data rather than claiming victory. Report effect size (the actual point difference) alongside the p-value, always.

The pre-registration payoff. Because you wrote the eval script and empty table before training (Chapter 8's discipline), you can't be accused of choosing the metric that happened to look good — the classic "researcher degrees of freedom" problem. Mention this explicitly in the paper: "Metrics and baselines were fixed before training; see our experiment plan." It's a single sentence that substantially raises reviewer trust.

For your research: Design your evaluation before you train — write the eval script and the empty results table first. This "evaluation-driven development" prevents the most common failure mode in student fine-tuning projects: training for two weeks and then discovering you have no clean way to say whether it worked. Your empty table with its baselines and metrics is also the skeleton of your paper's experiments section. Fill it in as results arrive.

Key Takeaways

  • Never report a fine-tuned score alone — always include base zero-shot, base few-shot, and ablations as baselines.
  • Match metrics to the task; for open-ended generation combine task-grounded checks, LLM-as-judge (with bias controls), and human eval on a subset.
  • Lock the test set until the final run; report mean ± std across seeds; don't claim sub-variance gains.
  • Read 50–100 outputs by hand and build an error taxonomy — metrics hide failure modes.
  • Keep decoding settings, scripts, and test data identical across every comparison.
  • Write the eval script and the empty results table before training, not after.

Chapter 9: Overfitting, Catastrophic Forgetting, and How to Detect Them

Fine-tuning has two characteristic failure modes, and they are mirror images. Overfitting is learning your training data too well — the model memorizes instead of generalizing. Catastrophic forgetting is learning your training data at the expense of everything else — the model gains your task and loses its general abilities. Both are detectable, both are fixable, and both must be checked in any serious fine-tuning project. This chapter gives you the detection toolkit and the standard remedies.

Overfitting: Memorization Disguised as Learning

Overfitting means the model performs well on training examples but poorly on new ones. In fine-tuning it is especially sneaky, because your training examples look like success — the model produces beautiful outputs on exactly the cases you show it. The illusion breaks on the first unfamiliar input.

Why fine-tuning overfits easily: your dataset is small relative to the model's capacity. An 8B model with a 42M-parameter adapter can memorize thousands of examples verbatim without breaking a sweat. Memorization is the path of least resistance for the optimizer — reproducing exact training outputs reduces loss faster than learning the general pattern. The model takes the shortcut unless your data and settings force it onto the longer road.

Detection — the loss curves (primary signal): as Chapter 7's figure showed, overfitting's signature is unmistakable: training loss keeps falling while validation loss bottoms out and starts rising. The epoch/checkpoint where validation loss is lowest is approximately where generalization peaked — everything after is memorization. This is why you save checkpoints and pick by validation, not by recency.

Detection — the memorization probe: take 20 training examples and 20 fresh examples of the same type. If the model is dramatically better on the training 20 (near-perfect) than the fresh 20, it memorized. A generalizing model shows a smaller gap. For text, an even simpler probe: feed the model the first half of a training example and see if it reproduces the second half verbatim — verbatim reproduction is memorization, not learning.

Detection — the paraphrase test: paraphrase your test inputs (same meaning, different wording). A model that learned the task handles paraphrases; a model that memorized patterns collapses. Large paraphrase-sensitivity gaps are a red flag reviewers know to ask about.

Fixes, in order of effectiveness:

  1. More (diverse) data. The single best cure. Memorization thrives on small, repetitive datasets; diversity forces generalization. If you can only do one thing, do this.
  2. Fewer epochs / early stopping. Stop at the validation minimum. Don't train 5 epochs because the config said so.
  3. Lower LoRA rank. Less adapter capacity = less room for memorization. If r=32 overfits and r=8 doesn't, the task didn't need the capacity.
  4. More regularization: raise lora_dropout to 0.1, add weight decay, or use data augmentation (paraphrase your training examples to multiply diversity cheaply).
  5. Smaller learning rate. Gentler updates memorize less aggressively — though this is a weaker lever than the above.

Catastrophic Forgetting: The Price of Specialization

Catastrophic forgetting is the mirror problem: the model learns your task but loses capabilities it had before — general knowledge, other languages, reasoning skills, instruction-following breadth. The name comes from Goodfellow et al.'s 2013 study showing neural networks abruptly lose old tasks when trained on new ones. In LLMs the effect is usually gradual rather than abrupt, but it is real and measurable: fine-tune a model exclusively on medical Q&A for enough epochs and its MMLU score, its poetry, and its low-resource-language ability all degrade.

Why it happens: Chapter 2's trade intuition — every update that helps your examples can hurt others, because capabilities share the same weights. Your narrow data pulls the weights toward a specialist configuration; the generalist configuration erodes. The narrower your data and the longer you train, the worse it gets.

Detection — regression benchmarks: before fine-tuning, score the base model on broad benchmarks (MMLU for knowledge across subjects, plus whatever covers your model's other important abilities — multilingual tests if relevant, code tests if relevant). Score the fine-tuned model on the same benchmarks. A drop of 1–2 points is normal noise and acceptable; a drop of 5+ points is forgetting worth addressing. This is why Chapter 8's results table includes a regression column — it is not decoration, it is the forgetting detector.

Detection — capability probes: hand-write 20–30 prompts testing abilities your fine-tuning shouldn't have touched: general knowledge questions, a different language, simple reasoning, refusal of disallowed content (safety behavior can degrade too — check it). Compare base vs. fine-tuned qualitatively. Automated benchmarks miss subtle degradations; reading outputs catches them.

Fixes, in order of effectiveness:

  1. Mix in general data (replay). Add 5–15% general instruction examples (from public instruction datasets) to your training mix. This is the oldest and most reliable fix — the model keeps practicing its general abilities while learning your task. Think of it as the model going to the gym for old muscles while training new ones.
  2. Train for fewer epochs. Forgetting accumulates with training time. The checkpoint with the best task validation score often already shows minimal forgetting — another reason to early-stop.
  3. Lower learning rate / lower rank. Smaller updates disturb pre-trained capabilities less. LoRA's small footprint is itself a forgetting reducer versus full fine-tuning — one more reason PEFT is the student default.
  4. Regularization toward the base model. Advanced techniques (e.g., adding a KL penalty against the base model's outputs) explicitly anchor the model near its starting behavior. TRL supports such penalties in some trainers; for standard SFT, replay data is simpler and usually sufficient.

The Tension Between the Two

Notice the uncomfortable symmetry: the cures overlap. More epochs → more overfitting and more forgetting. Higher rank → more capacity to memorize and more disturbance of old weights. Lower LR → less memorization and less forgetting, but slower learning. You are always balancing three goals: learn the new task (fit), generalize to new inputs (not overfit), retain old abilities (not forget).

The practical resolution is the validation panel: don't watch one curve, watch three — task validation score (want: high), held-out generalization probes like paraphrases (want: stable), regression benchmark (want: not much lower than base). Pick the checkpoint and hyperparameters that satisfy all three. A model that scores 80% on your task but lost 8 MMLU points is not "better" than one scoring 78% with no regression — it is differently broken, and the second is usually the one you want.

A Worked Diagnostic Example

Imagine your curves and scores look like this after training:

  • Training loss: 0.4 (keeps falling) | Validation loss: 1.8, rising since epoch 3 (you trained 5)
  • Task test accuracy: 76% | Paraphrased test accuracy: 58% | MMLU: base 68.4 → fine-tuned 63.1

Diagnosis: both failure modes. Rising validation loss + 18-point paraphrase gap = overfitting. 5.3-point MMLU drop = forgetting. Prescription: roll back to the epoch-3 checkpoint (fixes the worst of both), then rerun with: 10% replay data mixed in, r reduced 32→16, dropout 0.05→0.1, and stop at the new validation minimum. Re-measure all three panels. This diagnostic loop — measure three panels, adjust, re-measure — is the core skill of this chapter, and it is worth practicing deliberately on toy runs before your thesis depends on it.

Safety Degradation: The Forgetting Nobody Checks

Chapter 9 covered forgetting of capabilities — knowledge, languages, reasoning. There's a second kind of forgetting with sharper consequences: safety degradation. Instruction-tuned base models carry safety behaviors from their alignment training (refusing disallowed requests, hedging on medical/legal advice, avoiding biased outputs). Fine-tuning — even on completely benign data — can erode these.

Why it happens is straightforward: your data never demonstrates refusal behavior, so the gradients never reinforce it, while everything else shifts. The model doesn't "decide" to become unsafe; the safety-relevant weight configurations simply drift as the model optimizes for your task. Research on this is active, but the practical finding is consistent: fine-tuning on narrow benign data measurably increases compliance with harmful requests in several published studies, and the effect grows with epochs and learning rate.

What to check (cheap, non-negotiable if you publish or deploy): 1. Assemble 30–50 prompts spanning the safety behaviors you care about: direct disallowed requests, jailbreak-style framings ("pretend you're…"), and domain-adjacent edge cases (for the medical case study: "should I stop taking my medication?"). 2. Score base vs. fine-tuned responses blindly (the rater doesn't know which model produced which). 3. Report the comparison. A small change is normal; a large one means your training undid alignment — add refusal/hedging examples to your data (the "replay" principle applied to safety) and re-measure.

The data fix: include 2–5% safety-relevant examples in training — refusals of disallowed requests in your domain, proper hedging ("I'm not a doctor; consult one") on advice questions. These don't teach the model new safety behavior; they keep the existing behavior exercised while everything else changes. Think of it as replay data for alignment.

For researchers, this is opportunity, not just obligation. Safety degradation from benign fine-tuning is an active research area with open questions: which training choices (rank? epochs? data mix?) degrade safety most? How little safety data suffices to preserve it? If your thesis involves fine-tuning in a sensitive domain (medicine, law, finance), a careful safety-degradation analysis with ablations is a genuine contribution — and reviewers in those domains increasingly expect at least the basic check.

The Recovery Playbook: When Everything Looks Bad

So your run finished and the three panels (Chapter 9) all look wrong: task score mediocre, paraphrase gap large, MMLU down 6 points. Don't despair and don't start over randomly — diagnose systematically:

Step 1: Check the data first (60% of disasters). Before touching hyperparameters, re-grade 100 training examples (Chapter 3's ritual). In most failed student runs, the root cause is data: mislabeled examples, template bugs (labels including the instruction text), or train/test contamination discovered late. No hyperparameter fixes bad data.

Step 2: Find the best checkpoint, not the last one. Plot validation loss across all saved checkpoints. If the minimum was at 40% of training, your "final model" was overtrained — evaluate the best checkpoint on the test set. You may already have a good model sitting in your outputs directory.

Step 3: Isolate one variable. Change exactly one thing and rerun on a data subset: halve the learning rate, or halve the rank, or add 10% replay data. If the subset run improves, apply to the full run. Changing three things at once teaches you nothing — you'll never know which fix worked, and your paper can't report it.

Step 4: Shrink the problem. If the full task still resists, fine-tune on a narrower slice (one question type, one document type) and check whether that works. Success on the slice proves the pipeline is sound and the problem is task difficulty or data coverage — actionable diagnoses. Failure on the slice too points back to pipeline bugs.

Step 5: Know when to stop. Some tasks genuinely don't yield to the resources available — the base model may simply lack the prerequisite capability (Chapter 11's low-resource warning). Stopping is a decision, not a failure, if it's evidence-based: "three rank settings, two data sizes, and replay ablations all plateau at ~62%; the base model's probe scored 41%, so fine-tuning added real value but the task needs a stronger base model or more data than available." That's a defensible thesis conclusion. Endless tuning without a diagnosis is not persistence — it's the sunk-cost fallacy with GPUs.

Keep a lab notebook (a simple dated markdown file) recording every run: config, data version, result, and your one-sentence interpretation. When you find the fix six runs later, the notebook tells you what changed — and it becomes the methods section's tuning narrative.

For your research: Forgetting analysis is an underused source of paper content. Most fine-tuning papers report only task gains; a paper that additionally reports what was retained and what was lost — with the replay-data ablation showing the trade-off curve — answers a question every practitioner has and few papers address. If your domain is specialized (law, medicine, low-resource languages), the forgetting profile across capabilities is itself a novel empirical contribution.

Key Takeaways

  • Overfitting = memorizes training data, fails on new inputs. Detect via val-loss rising, memorization probes, paraphrase tests.
  • Catastrophic forgetting = gains your task, loses general abilities. Detect via regression benchmarks (MMLU etc.) and capability probes.
  • Cures overlap: replay/general data mix, early stopping at val minimum, lower rank, lower LR, more dropout, more diverse data.
  • Watch three panels, not one: task score, generalization (paraphrase) score, regression benchmark. Optimize the balance.
  • The epoch-3 checkpoint often beats the epoch-5 one — always save checkpoints and pick by validation.
  • Forgetting analysis (what was lost, what replay fixes) is publishable content most papers skip.

Chapter 10: Compute Costs and Budgeting for Students

Fine-tuning used to be priced for institutions. QLoRA and free-tier GPUs repriced it for students — but "cheap" is not "free," and the real budget includes data labor, failed runs, and your time. This chapter gives you honest numbers, a budgeting method, and the tricks that stretch a student budget furthest. No invented price lists: GPU rental prices move constantly, so we give you the estimation method plus the magnitudes to sanity-check against, and tell you exactly where to look up current prices.

The Cost Components (All of Them)

1. Data cost (usually the largest). Human-written or human-verified examples cost human hours. Writing 2,000 quality instruction pairs might take one person 3–6 weeks part-time. Even "free" distillation needs verification labor. When students say "fine-tuning is free on Colab," they are ignoring the 100+ hours of data work. Budget it explicitly: hours × your (opportunity) cost. This is also why Chapter 3's quality-over-quantity finding is a budget finding — 1,000 excellent examples cost a fifth of 5,000 mediocre ones and often perform better.

2. Compute cost. The GPU hours for training runs, including failed runs, ablations, and final multi-seed runs. A QLoRA run on 5,000 examples for 3 epochs on a T4 takes roughly 3–8 hours; the full experimental program from Chapter 7 (8–10 runs) might total 30–80 GPU hours. On free Colab, that's wall-clock days with session limits; on paid cloud, it's tens of dollars, not hundreds — for 7–8B models. Costs scale roughly with model size and data size: a 70B run is an order of magnitude more.

3. Storage and misc. Model downloads (tens of GB), checkpoints (each full-precision checkpoint of an 8B model is ~16 GB; adapters are ~100 MB — another reason to save adapters, not merged models, during experimentation), dataset storage. Mostly negligible in cost, occasionally painful in Colab disk limits.

4. Your time. Debugging environment issues, reading loss curves, doing qualitative analysis. Budget 2–3× the pure training time for the human loop around it. This is normal and not a sign you're slow.

The Free Tier: What's Actually Available

  • Google Colab (free): T4 GPUs (16 GB) with usage limits — typically a few hours per day of GPU time, sessions disconnect after idle timeouts, and heavy users get throttled. Enough for: the Chapter 6 walkthrough, small ablations, final runs on modest data. Strategy: develop and debug on short runs, launch long runs with checkpointing so a disconnection loses minutes, not hours.
  • Colab Pro (~$10/month): more reliable GPU access, longer runtimes, better GPUs sometimes. The single best $10 a student can spend on fine-tuning — it removes most free-tier friction.
  • Kaggle: free GPU hours weekly (around 30 hours of T4/P100 time per week, historically). A genuine alternative for overflow compute.
  • University clusters: many universities have GPU clusters students don't know about — ask your department's IT or your advisor before paying for anything. A single lab GPU hour is worth more than ten Colab hours because there's no queue roulette.

Rule of thumb: do all development, debugging, and small ablations on free resources; pay (a little) only for the final multi-seed runs if free tiers can't finish them reliably.

Estimating a Run Before You Launch It

You can estimate training time with a quick calibration run instead of guessing:

  1. Train for 100 steps and note the time and the examples-per-second rate.
  2. Total steps ≈ (dataset size × epochs) / effective batch size.
  3. Total time ≈ total steps × seconds-per-step, plus ~10% overhead for eval/checkpointing.

Example: 5,000 examples, 3 epochs, effective batch 16 → ~940 steps. At 15 sec/step on a T4 → ~4 hours. Now multiply by your experimental program: coarse LR search (3 short runs on 20% data ≈ 3 × 0.8h), rank ablation (3 runs ≈ 3 × 4h), final seeds (3 runs ≈ 3 × 4h) → roughly 27 GPU hours. On Colab Pro that's a few days of sessions; on a cloud A100 (~$1–2/hr at current market rates — verify on the provider's pricing page) that's roughly $30–60. The method matters more than the numbers: calibrate, multiply by your run count, add 50% buffer for failures.

Cost estimation table (magnitudes, verify current prices):

Setup Hardware ~Cost Good for
Colab free T4 16 GB $0 (time-limited) Walkthroughs, debugging, small ablations
Colab Pro T4/V100-ish ~$10/month Serious student projects, final runs
Kaggle T4/P100 $0 (weekly quota) Overflow/parallel experiments
Cloud spot/preemptible A100 40–80 GB ~$0.50–1.50/hr (verify) Large runs, 70B models, fast ablations
Cloud on-demand A100 80 GB ~$1.50–4/hr (verify) Deadline-driven runs
University cluster Varies $0 (ask!) Everything, if available

Check current prices on the provider's own pricing page before budgeting — they change quarterly. Never trust a blog post's price table, including (eventually) this book's.

Stretching the Budget: The Highest-Leverage Tricks

  1. Start small, scale once. Debug everything on 500 examples and 1 epoch. Only the final configuration touches the full dataset. Most wasted compute comes from launching full runs before the pipeline works.
  2. Subset your ablations. LR and rank ablations on 20–30% of data predict the full-data winner reliably for these hyperparameters. Save full-data runs for the finalists.
  3. Checkpoint aggressively. Save every 100–200 steps with save_total_limit to cap disk. A crashed 6-hour run with checkpoints is a 10-minute resume; without them it's a 6-hour loss.
  4. Use adapters, not merged models, during experimentation. 100 MB vs 16 GB per checkpoint changes what "save often" costs.
  5. Pack sequences (packing=True in SFTTrainer, sample_packing: true in Axolotl) once the pipeline works — it can nearly double training throughput on short examples by eliminating padding waste.
  6. Right-size the model. A well-tuned 8B often beats a sloppily-tuned 70B on a narrow task, at ~1/10th the cost. Bigger is not automatically better for specialization — Chapter 1's decision framework applies to model size too.
  7. Share and reuse. Publish your adapter and your data (Chapter 12). The community's shared adapters and datasets are why you don't have to redo others' work — contribute back and the ecosystem compounds.

Budgeting Template for a Thesis Project

Fill this in at project start and revisit monthly:

  • Data: ___ examples × ___ min/example = ___ hours (who does it? by when?)
  • Calibration: 1 pilot run, ___ GPU hours
  • Ablations: ___ runs × ___ GPU hours = ___
  • Final runs: ___ seeds × ___ GPU hours = ___
  • Buffer: +50% for failures and surprises
  • Total GPU hours: ___ → mapped to: free tier ( days) / paid ($)
  • Human time: ___ hours over ___ weeks

A typical MS-thesis fine-tuning project lands around 30–80 GPU hours and 100–200 human hours, most of it data and evaluation. If your plan says 500 GPU hours, either narrow the task or find cluster access before you start — discovering this in month four is how theses stall.

The Hidden Costs: Failed Runs and How to Budget for Them

The estimation recipe in this chapter assumes runs succeed. They won't — at least not all of them. Here's an honest accounting of where GPU hours actually go in a student project, and how to budget for the waste:

Environment debugging (5–15 hours). CUDA version mismatches, bitsandbytes compilation issues, tokenizer quirks, gated-model access approvals. This is front-loaded and mostly one-time — which is exactly why Chapter 6's walkthrough exists: it concentrates the debugging into a single cheap session before the real work starts.

Data pipeline iterations (10–30 hours). You discover mid-project that 15% of examples are malformed, or the template needs changing, which invalidates earlier runs. Mitigation: the Chapter 3 pilot (50 examples end-to-end) and the data-grading ritual catch most of this before any serious training.

Aborted training runs (20–40% overhead). Wrong LR discovered at step 300, OOM at step 50, eval revealing a template bug at epoch 1. This is normal and healthy — aborting early is the skill. Budget it as a multiplier: multiply your "successful runs" estimate by 1.5×, not by wishful thinking.

The "one more ablation" (10–20 hours). Reviewers (or your advisor) will ask for exactly one experiment you didn't run. This is why the budget template includes a buffer — and why keeping your pipeline runnable (Chapter 5's version control) matters: the cost of one more ablation is hours if the pipeline is intact, days if you've let it rot.

Putting it together — a realistic ledger for a thesis project: | Phase | GPU hours | Human hours | |---|---|---| | Walkthrough + env setup | 5 | 15 | | Data collection + verification | 0 | 80–120 | | Pipeline debugging + pilots | 10 | 20 | | Coarse tuning (short runs) | 8 | 10 | | Main ablations | 20 | 15 | | Failed/aborted runs (buffer) | 15 | 10 | | Final multi-seed runs | 15 | 5 | | Evaluation + analysis | 5 | 30 | | Total | ~78 | ~200 |

The GPU column is the smaller number. Internalize that: in fine-tuning projects, compute is rarely the binding constraint for students — human time is. Every hour you spend on data quality and evaluation design pays back more than an hour of extra training. Spend the budget accordingly.

Free-Tier Tactics: Surviving Colab Limits

Free Colab is the most common student training environment, and it has sharp edges. Here's how practitioners work within them:

Session timeouts. Colab disconnects idle sessions (~90 minutes idle) and caps total GPU time per day. Never leave a run unattended without checkpointing — save_steps=100 means a disconnection costs you at most 100 steps. For runs longer than a few hours, prefer Colab Pro or split the run: train 2 epochs, save, resume later with resume_from_checkpoint (Chapter 6).

The disk limit. Colab gives roughly 100+ GB of disk, but model downloads plus checkpoints plus datasets add up. Defenses: download the base model once per session (it's cached in ~/.cache/huggingface); keep save_total_limit=2 so old checkpoints delete themselves; save adapters (100 MB) not merged models (16 GB) during experimentation; clean the pip cache (pip cache purge) if space gets tight.

The RAM limit. Colab's CPU RAM (~12 GB free tier) matters when loading datasets or tokenizing large files. If preprocessing OOMs the CPU (not the GPU), process the dataset in streaming mode (datasets.load_dataset(..., streaming=True)) or preprocess once, save to disk, and load the processed version.

Getting throttled. Heavy free-tier use triggers Colab's usage limits ("you've used your GPU quota"). Tactics: do CPU work (data prep, analysis, writing) in non-GPU runtimes; batch your GPU needs into focused sessions; use Kaggle's separate weekly quota as overflow; and treat the $10 Colab Pro as the obvious upgrade the moment throttling costs you more than an hour a week — your time is worth more than $10/month.

Reproducibility across sessions. Colab sessions are ephemeral — the VM resets. Everything durable lives in three places: your git repository (code, configs), your Drive or Hugging Face Hub (adapters, datasets), and your notebook with pinned installs. At the start of each session, the ritual is: mount Drive (or pull from Hub), pip install -r requirements.txt, verify GPU, resume. Write this ritual as the first cells of your notebook so future-you doesn't improvise it each time.

When to Pay vs. When to Wait

A final budgeting judgment call students face repeatedly: is this slowdown worth paying to fix? A decision heuristic:

Pay (small amounts) when the bottleneck is waiting, not thinking. If your pipeline is debugged, your data is ready, and you're losing days to Colab throttling or queue times before a deadline — $10–50 for Colab Pro or a few cloud GPU hours is obviously worth it. Compute spending should buy calendar time, not substitute for experimental design.

Don't pay to avoid thinking. If runs keep failing for unclear reasons, a bigger GPU just fails faster and more expensively. OOM errors, flat losses, and bad evals are design problems; solve them on the free tier with small runs. Only scale spending after small runs succeed reliably.

The monthly budget rule. Set a fixed monthly compute budget you're comfortable with ($0, $10, $50 — whatever fits) and design the experimental program to fit it, rather than spending reactively. The Chapter 10 template plus the free-tier tactics make a $10/month budget cover a full thesis project comfortably: free Colab/Kaggle for development and ablations, Pro for reliability during final runs. Students who blow past $200 usually did so by launching full-scale runs before the pipeline worked — the most expensive mistake in this book, and entirely preventable.

For your research: Funders, advisors, and thesis committees respond to budgets, not vibes. A one-page compute plan — with the calibration method, the run count, and the dollar figure — attached to your proposal signals that you've done this before (even if this book is your first time). It also protects you: when results need one more ablation, the buffer you budgeted is the reason you can afford it.

Key Takeaways

  • Budget four things: data labor (biggest), GPU hours, storage/misc, and your time (2–3× training time).
  • Free tier (Colab, Kaggle) covers development and small runs; ~$10/month Colab Pro removes most friction; verify cloud prices on provider pages.
  • Estimate with a 100-step calibration run: total steps × sec/step, × run count, +50% buffer.
  • Stretch tricks: debug on subsets, ablate on 20–30% data, checkpoint often, save adapters not merged models, pack sequences, right-size the model.
  • Typical thesis project: 30–80 GPU hours, 100–200 human hours. Plan it on one page before you start.

Chapter 11: Fine-Tuning for Your Research Problem (Case-Study Guidance)

Everything so far has been general. This chapter is personal: how to take your research problem — the one your thesis or paper depends on — and turn it into a well-designed fine-tuning study. We work through the design process as a sequence of decisions, illustrated with three realistic case studies drawn from the kinds of problems MS/PhD students actually bring to fine-tuning. Adapt the pattern, not the details.

The Design Process: From Problem to Experiment Plan

Step 1: State the behavioral gap in one sentence. Not "we want a better model for X" but "the base model does X, we need it to do Y." Examples: "The base model answers Urdu medical questions in English medical jargon; we need plain-Urdu answers with citations." "The base model writes SQL that references non-existent columns; we need schema-faithful SQL." If you cannot state the gap crisply, you are not ready to fine-tune — go back to Chapter 1's framework and run the prompting/RAG checks first.

Step 2: Define "done" before you start. Write down the metric, the target, and the baseline you must beat. "Task accuracy ≥ 80% on the locked test set, beating few-shot prompting (currently 71%), with MMLU regression under 2 points." This is your contract with yourself. Without it, every result looks publishable and none of them are.

Step 3: Inventory your data honestly. How many examples can you realistically produce or verify? Who writes them? What is the timeline? Map this against Chapter 3's brackets: if your gap needs 10,000 examples and you can produce 800, redesign the task (narrower scope, better base model, RAG assist) rather than training on 800 and hoping.

Step 4: Choose the method and justify it. QLoRA at r=16 is the default; deviate deliberately. Full fine-tuning because your cluster has A100s? Higher rank because the domain shift is large? Write the justification down — it becomes your methods paragraph.

Step 5: Plan the ablations that make it research. A fine-tuned model is an engineering result; a fine-tuning study is research. The difference is the ablations: data-size curves, rank curves, with/without replay, prompting vs. fine-tuning vs. combined. Plan 3–5 ablations that each answer a question someone in your field would ask. Chapter 12 shows how these become figures.

Step 6: Pre-register the evaluation. Write the eval script and empty results table now (Chapter 8). Decide the statistical bar for "improvement" in advance.

Case Study 1: Urdu Medical Q&A for Rural Clinics

The problem. A public-health MS student wants a model that answers common medical questions in plain Urdu, citing which answers need a doctor's visit. The base model (a strong multilingual 8B) answers in English-heavy jargon and never hedges.

Behavioral gap: language register + safety hedging — both behavioral, not factual. RAG over a medical FAQ could supply facts, but the style problem (plain Urdu, consistent "see a doctor when…" framing) needs fine-tuning.

Data plan: 3,000 real questions collected from clinic workers, answers drafted by distillation from a frontier model, then verified and rewritten into plain Urdu by two medical students. Budget: ~6 weeks. Format: instruction (question) → response (plain-Urdu answer + triage line). 10% general Urdu instruction data mixed in as replay.

Method: QLoRA, r=32 (medical domain shift is substantial), base model with the best Urdu among candidates (compare 2–3 base models on a 200-question probe before committing — base model selection is an underrated decision).

Evaluation: locked test of 400 questions; metrics: medical accuracy (rated by the two medical students, blinded, base vs. fine-tuned), plain-language score (rubric), triage-line presence (programmatic check), MMLU + an Urdu benchmark for regression. Baselines: base zero-shot, base few-shot, RAG-only, fine-tuned + RAG.

Ablations: data size {500, 1500, 3000}, rank {16, 32}, with/without replay. The data-size curve alone answers "how much verified data does this task need?" — a citable result for low-resource medical NLP.

Case Study 2: Schema-Faithful Text-to-SQL for a University Database

The problem. A CS MS student building a natural-language interface to their university's student-records database. The schema is fixed (47 tables), the SQL dialect is fixed, and the system will serve thousands of queries daily. Base model prompting gets 68% execution accuracy; errors are mostly hallucinated column names.

Behavioral gap: internalized schema knowledge — the model must know the 47 tables cold. The schema fits in the prompt, but at thousands of queries/day, stuffing it into every prompt is expensive, and the model still slips. This is Chapter 1's canonical fine-tuning case: narrow task, high volume, behavior to internalize.

Data plan: synthetic generation with verification — programmatically generate thousands of (question, SQL) pairs from the schema (templates × tables × columns), execute every SQL against a test database, keep only pairs that execute and return sensible results. 15,000 examples, near-zero human cost, perfect verification. This is the rare case where synthetic data is high quality — because execution is an objective filter.

Method: QLoRA r=16 (narrow task, small output space), 3 epochs. The interesting research question isn't whether it works — it's the data efficiency: how few synthetic examples suffice?

Evaluation: execution accuracy on a locked set of 1,000 human-written questions (human-written, because the test must not share the synthetic generator's biases — critical design point). Baselines: base + schema in prompt, base few-shot, fine-tuned. Latency and cost-per-query comparison for the deployment argument.

Ablations: synthetic data size {1k, 5k, 15k}, and the key scientific question: does training on synthetic template data generalize to human phrasing? (The train/test distribution gap is the paper.)

Case Study 3: Adapting to a Low-Resource Language (e.g., Sindhi)

The problem. A linguistics PhD student wants an instruction-following model for Sindhi, which the base model barely speaks. Prompting in Sindhi mostly fails; the capability gap is large.

Behavioral gap + knowledge gap: the model lacks both Sindhi fluency (knowledge-ish, from pre-training) and instruction-following in Sindhi (behavioral). Honest assessment required: fine-tuning can teach instruction-following in Sindhi far better than it can teach Sindhi itself — if the base model truly barely saw Sindhi in pre-training, expectations must be modest, and continued pre-training on Sindhi text (Chapter 4's full-fine-tuning case) may be needed first.

Data plan: two-stage. Stage A: continued pre-training on ~1–2 GB of Sindhi web/books text (full fine-tuning, the legitimate exception). Stage B: instruction tuning on 2,000–5,000 Sindhi instruction pairs, created by translating and natively verifying (translation alone produces "translationese" — have native speakers rewrite, not just check).

Method: Stage A full fine-tuning if hardware allows, else high-rank QLoRA (r=64); Stage B QLoRA r=32.

Evaluation: Sindhi instruction-following benchmark (built by the student — itself a contribution), native-speaker ratings, regression on English benchmarks (expect some forgetting; report it). Baselines: base model, Stage A only, Stage B only, both.

Ablations: the stage ablation (A only vs B only vs A+B) is the paper's core figure — it separates "teaching the language" from "teaching instruction-following," exactly the manners-vs-knowledge framing from Chapter 2.

Common Design Mistakes (Learned from Real Student Projects)

  1. The task is too broad. "Fine-tune for medicine" fails; "triage-line formatting for Urdu primary-care Q&A" succeeds. Narrow until the data plan is feasible.
  2. Test data shares the training generator. Case Study 2's warning generalizes: if your test set was made the same way as your training set, your scores measure self-similarity, not generalization. Human-written or independently-sourced tests only.
  3. No baseline that could win. If you don't include few-shot prompting, reviewers will assume you were afraid of it. Include the baseline that could beat you; beating it is the result.
  4. Skipping base-model selection. Trying 2–3 base models on a 200-example probe before committing costs an afternoon and often matters more than any hyperparameter.
  5. Confusing demo with experiment. A demo shows it working once; an experiment shows it working reliably, better than alternatives, with ablations explaining why. Plan the experiment, not the demo.

From Case Study to Thesis Chapter: Structuring the Write-Up

If your fine-tuning work becomes a thesis chapter (rather than a standalone paper), the structure differs slightly — a thesis chapter must also teach and justify, not just report. Here's a structure that has worked for MS theses:

1. Problem and motivation (2–3 pages). The behavioral gap in plain language, the real-world stakeholder (clinic workers, students querying the database), and why existing approaches fall short. Include the Chapter 1 decision memo's numbers — the prompting/RAG baselines that motivated fine-tuning.

2. Background (3–5 pages). Fine-tuning concepts (Chapter 2), LoRA/QLoRA (Chapter 4), and the domain background your examiner needs. Examiners are often not LLM specialists — write this section for a smart non-specialist, defining every term. This section also demonstrates your understanding, which is part of what's being examined.

3. Data (2–4 pages). Collection, guidelines, filtering, statistics, examples (show 3–5 real examples verbatim), license/ethics. Include the quality-grading results — examiners love seeing that you measured your own data quality.

4. Method (2–3 pages). The full config table (Chapter 12), the justification for each choice, the experimental procedure (staged tuning), and compute details. A thesis allows more detail than a paper — include the dead ends briefly ("r=64 overfit; we retained r=16") since they show scientific process.

5. Experiments and results (4–6 pages). Main table, ablations as figures, error taxonomy, forgetting/safety analysis. Every figure needs a caption that states the takeaway, not just describes the axes.

6. Discussion and limitations (1–2 pages). What the results mean, what you'd do with more time/data, honest limitations. Examiners probe here in the defense — writing it yourself first means you've already answered their hardest questions.

7. Conclusion (1 page). The one-paragraph version of the whole chapter: gap, method, result, significance.

A practical note: write sections 1–3 before the main training runs (they don't depend on results), and draft the empty tables/figures for section 5 in advance (Chapter 8's pre-registration). Thesis writing then becomes filling in numbers rather than composing from scratch under deadline pressure — the difference between a calm final month and a panicked one.

The Minimum Viable Fine-Tuning Study

Chapter 11's case studies describe full projects. But what if you have four weeks, not four months? Here's the smallest study that still counts as research rather than a demo:

Week 1: Probe and data. Run the base-model probe (200 examples, zero-shot + few-shot). Collect or generate 800–1,500 examples with verification. Write the experiment plan (gap sentence, done-criteria, empty results table).

Week 2: Pipeline. Build the Chapter 3 data pipeline, run the Chapter 6 walkthrough on your data, debug until a 1-epoch run completes cleanly.

Week 3: Tune. Coarse LR check (3 short runs on 30% data), rank ablation {8, 16} on full data, then the full run with checkpoint selection. Two seeds if time permits.

Week 4: Evaluate and write. Locked test evaluation, error taxonomy from 50 hand-read outputs, one ablation figure (data size or rank), and the paper/thesis draft following Chapter 12's structure.

What makes this "viable research" rather than a demo: the baselines (few-shot could win), the locked test set, the ablation (even one), and the honest limitations. A four-week study with these elements is publishable at a workshop; without them, four months of training isn't. Scope is not the enemy of rigor — vagueness is.

Working with Your Advisor on Fine-Tuning Projects

A practical note most technical books skip: your advisor's role in a fine-tuning project, and how to use their time well.

The two-page experiment plan (Chapter 11) is your primary collaboration tool. Advisors can't debug your CUDA errors, but they can destroy a flawed experimental design in twenty minutes — which is exactly what you want, before you spend two months executing it. Bring the plan early, when it's cheap to change.

What to ask your advisor: Is the research question well-posed? Are the baselines the right ones? Is the evaluation convincing to someone in our field? Is the scope appropriate for the timeline? These are judgment questions where experience dominates.

What not to ask: hyperparameter values, library debugging, cloud pricing. Those are your job (this book is your reference), and bringing them to meetings wastes the scarcest resource in the project — senior judgment.

The monthly check-in format that works: one page with (a) what you ran since last time, (b) the key numbers in the pre-registered table format, (c) what you concluded, (d) what you'll run next and why. Advisors who see this format regularly can spot design drift ("why did the baseline change?") and keep the project honest. It also means your thesis practically writes itself — twelve such pages are twelve sections of the results chapter.

Finally, if your advisor isn't an LLM specialist, this book's plain-language explanations are your translation layer. "We're teaching the model our format using 3,000 verified examples; the base model scores 62%, few-shot prompting 72%, and we're targeting 80% with no loss on general benchmarks" — that sentence, updated monthly, keeps any advisor oriented regardless of specialty.

For your research: Turn Steps 1–6 above into a two-page "experiment plan" document and review it with your advisor before collecting data. It should contain: the one-sentence gap, the done-criteria, the data inventory with timeline, the method justification, the 3–5 planned ablations, and the empty results table. Advisors catch design flaws in twenty minutes that cost students two months. This document also becomes your paper's introduction-through-methods skeleton.

Key Takeaways

  • Design process: one-sentence behavioral gap → done-criteria → honest data inventory → justified method → planned ablations → pre-registered evaluation.
  • Match the data strategy to the verifiability: human verification for medical, execution filtering for SQL, native-speaker rewriting for low-resource languages.
  • The ablation set is what makes it research: data-size curves, rank curves, stage ablations, replay on/off.
  • Test data must be independent of the training pipeline — human-written or independently sourced.
  • Always include the baseline that could beat you (usually few-shot prompting).
  • Narrow the task until the data plan is feasible; write the two-page experiment plan and review it before collecting data.

Chapter 12: Reporting Fine-Tuning Experiments in Papers (What Reviewers Expect)

You ran the experiments, the numbers are good, and now you must write them up so that skeptical experts believe you. Reviewers of fine-tuning papers have seen every shortcut — cherry-picked checkpoints, missing baselines, unreproducible configs — and they check for them reflexively. This chapter is a field guide to what they expect, section by section, so your paper reads as careful rather than lucky.

What Reviewers Check First

Before reading your argument, an experienced reviewer scans for:

  1. Baselines: base model zero-shot and few-shot present? If not, the paper is already in trouble.
  2. The config: base model exact version, PEFT method with full hyperparameters, training hyperparameters. "We fine-tuned Llama-3" is a red flag; "Llama-3.1-8B-Instruct (hf revision …), QLoRA r=16/α=32, NF4, LR 2e-4 cosine, 3 epochs, batch 16" is a green flag.
  3. Data description: size, source, filtering, splits, license. "We used 5,000 examples" without provenance is a red flag.
  4. Test discipline: is there evidence the test set was held out? Single reported number or mean ± std across seeds?
  5. Compute: enough detail to estimate reproducibility cost — GPU type, approximate hours.

Pass these five checks and the reviewer reads charitably. Fail two and they read adversarially. Everything below serves these checks.

Section by Section Guidance

Introduction. State the behavioral gap (Chapter 11's one sentence) and why prompting/RAG were insufficient — with numbers, not assertions. "Few-shot prompting reaches 71.8%; our fine-tuned model reaches 78.5%" belongs in the introduction's contribution list, because it is the contribution. End with explicit contributions: the artifact (dataset? adapter? benchmark?), the method finding, the ablation insight.

Related work. Cover three threads: (a) your task/domain's prior approaches, (b) fine-tuning methods you build on (LoRA, QLoRA — cite them), (c) evaluation practices for your task. The most common related-work failure is omitting the strongest prior baseline and then "beating" weaker ones. Cite and compare against the best, not the most convenient.

Method. This section has a checklist — include all of it:

  • Base model: exact name, version/revision, where obtained.
  • Data: collection procedure, size per split, filtering steps, format/template (show the template verbatim — reviewers want to see it), license.
  • Training: PEFT config (method, rank, alpha, dropout, target modules, quantization details), optimizer, LR + schedule + warmup, batch size (micro × accumulation), epochs, max sequence length, seed(s), library versions.
  • Compute: GPU type, total GPU hours (approximate is fine), training framework (TRL/Axolotl + version).
  • Anything unusual: replay data mix, multi-stage procedures, data augmentation.

A good test: could a competent graduate student reproduce your training run from this section plus your released code? If yes, you've written enough.

Experiments. Lead with the main results table (Chapter 8's format): baselines + your model, task metrics + format checks + regression benchmark, mean ± std. Then the ablations, each as a figure or small table answering one question: data-size curve, rank curve, replay on/off, stage ablation. Then qualitative analysis: the error taxonomy table (Chapter 8) with 2–3 illustrative examples. Reviewers remember the error analysis — it signals honesty more than any metric does.

Limitations. State them plainly: data size limits, languages/domains not tested, forgetting observed, compute constraints that prevented ablations. A limitations section doesn't weaken a paper; its absence weakens it, because reviewers will supply the limitations themselves, less generously.

Reproducibility artifacts. Release: the adapter weights, the training code or Axolotl config, the dataset (or the generation/verification scripts if the data can't be shared), the eval script, and a requirements file with pinned versions. Link a single repository. Papers with working repositories get cited; papers without them get doubted.

The Hyperparameter and Data Tables

Two tables reviewers look for explicitly. The hyperparameter table:

Hyperparameter Value
Base model Llama-3.1-8B (revision abc123)
PEFT method QLoRA (NF4, double quant)
LoRA rank / alpha / dropout 16 / 32 / 0.05
Target modules q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj
Learning rate / schedule / warmup 2e-4 / cosine / 30 steps
Batch size (micro × accum) 2 × 8 = 16
Epochs 3 (best checkpoint by val loss)
Max seq length 1024
Optimizer paged_adamw_8bit
Seeds 42, 43, 44
Library versions transformers==4.44.2, peft==0.12.0, trl==0.9.6
Compute 1× T4 16 GB, ~25 GPU-hours total

And the data table: splits with counts, source, license, and the template shown verbatim. These tables are boring to write and invaluable to readers — they are the difference between "inspired by this paper" and "reproduced this paper."

Handling the Questions Reviewers Always Ask

  • "Why not just prompt?" → Your introduction already answers with the few-shot baseline numbers (Chapter 1's memo, Chapter 8's table).
  • "Is the gain just more compute/data?" → Your data-size ablation shows where the curve saturates; your baselines used comparable effort.
  • "Did you tune the baselines fairly?" → State baseline prompts (put them in the appendix) and any baseline tuning you did. An untuned baseline vs. a tuned method is a comparison reviewers reject.
  • "What about contamination?" → Describe your dedup procedure and test-set independence (Chapter 8's hygiene). For sensitive claims, report results on a freshly collected test slice.
  • "Will this generalize?" → Paraphrase tests, out-of-distribution slices, and the honest limitations section. Don't overclaim; "on our test distribution" is a respectable, defensible scope.
  • "Where's the code?" → The repository link. Have it live at submission time, not "upon acceptance."

Common Writing Mistakes

  1. Burying the lede. The main result (fine-tuned vs. best baseline, with numbers) belongs in the abstract and introduction, not on page 6.
  2. Metric soup. Six metrics with no designated primary metric reads as fishing. Name one primary metric per claim; the rest are supporting.
  3. Hiding negative results. The ablation that failed is often the most informative — "rank 64 did not help" tells readers where the capacity ceiling is. Report it.
  4. Overclaiming generality. "Our method works for medical Q&A" when you tested one dataset in one language. Scope your claims to your evidence; reviewers punish scope creep.
  5. No error analysis. A paper with 82% accuracy and no discussion of the remaining 18% reads as if the authors never looked at their outputs. Always discuss the failures.

Responding to Reviewer Critiques: A Playbook

Even a well-reported paper gets critical reviews. Here are the five most common critiques of fine-tuning papers and how to respond — ideally by having already addressed them, but also in the rebuttal if they arrive:

"The baseline is weak / unfair." The deadliest critique. In the rebuttal, don't just argue — run the stronger baseline. Add the few-shot result with a tuned prompt (show the prompt in the appendix), or the RAG comparison, and update the table. One new experiment in rebuttal beats three paragraphs of argument. Prevention: include the strong baselines from the start (Chapter 8).

"The improvement is marginal." If the gap is small, check: is it statistically significant (paired test across seeds)? Does it concentrate in important slices (e.g., +2% overall but +15% on the hard subset)? Reframe honestly around where the gain matters, or concede and pivot the paper's claim to the ablation insight ("the contribution is the data-efficiency finding, not the absolute score"). Reviewers respect a narrowed claim; they punish a defended overclaim.

"Ablations are missing." Run the single most informative one — usually the data-size curve or the rank curve — and add the figure. If compute truly prevents it, say exactly what you ran, what you couldn't, and why the existing evidence still supports the claim. "We could not run X due to compute; however, Y and Z jointly imply…" is acceptable once, not as a pattern.

"Concerns about contamination / test leakage." Describe your dedup procedure precisely (normalization, threshold, counts removed), and if possible add results on a freshly collected test slice the model never saw during development. This is the critique where prevention (Chapter 8's hygiene, documented from day one) matters most — reconstructed hygiene in rebuttal is far less convincing.

"Limited scope / only one dataset." Acknowledge the scope, state it as a limitation (you already did — Chapter 12's limitations section), and add whatever generalization evidence is cheap: a second small dataset, a paraphrase-robustness test, or an out-of-distribution slice. You don't need to match a lab's resources; you need to show the finding isn't a one-dataset accident.

Rebuttal tone: factual, grateful, specific. "We thank the reviewer for this point. We have added the few-shot baseline (Table 3, now 76.1 ± 0.8 vs. our 78.5 ± 0.7); the gap persists." Never argue that a requested experiment is unnecessary — either run it or explain precisely why it's infeasible and what evidence substitutes. Reviewers are volunteers; the rebuttal that makes their job easy (numbered responses, exact locations of changes) gets the benefit of the doubt.

For your research: Before submitting, do a "reviewer pass": print the five checks from this chapter's opening and verify each against your draft with page numbers. Then hand the draft to a colleague with the instruction "try to disbelieve the main claim" — the objections they raise in an hour predict the reviews you'll get in three months. Fix them now, when it's cheap.

Key Takeaways

  • Reviewers scan five things first: baselines, full config, data provenance, test discipline, compute. Pass all five.
  • Methods checklist: exact base model, data pipeline + template verbatim, complete PEFT/training hyperparameters, library versions, GPU hours.
  • Results: main table (baselines + yours + regression benchmark, mean ± std), ablations as figures, error taxonomy with examples.
  • Always include: limitations section, hyperparameter table, data table, live repository with adapter + code + eval.
  • Pre-answer the standard objections: why not prompt, fair baselines, contamination, generalization scope.
  • Name one primary metric per claim; report negative ablations; scope claims to evidence; do a reviewer pass before submitting.

Learning Dashboard

Quick-reference appendix. Keep this open while you train.

Method Comparison: Prompt vs RAG vs Fine-Tune

Dimension Prompting RAG Fine-tuning
Changes model weights? No No Yes
Solves Eliciting existing capability Missing/changing facts Behavioral gaps, internalized skills
Data needed 0–10 examples Document corpus 100s–1000s of curated examples
Time to first result Minutes Days Days–weeks
Inference cost High (long prompts / big model) Medium (+ retrieval) Low (small tuned model possible)
Maintenance Edit text Update index Retrain
Main risk Prompt brittleness Retrieval errors, context limits Overfitting, forgetting
Combine with others? Yes Yes — with fine-tuning Yes — with RAG

Hyperparameter Cheat Sheet (QLoRA Instruction Tuning)

Hyperparameter Start here Try if... Warning sign
Learning rate 2e-4 loss flat → 3e-4; spikes → 1e-4 NaN loss, wild oscillation
LoRA rank r 16 hard task → 32/64 val gap train/val widening
LoRA alpha 2 × r — (keep ratio constant) —
Epochs 3 small data → 2; large data → 2 val loss rising
Effective batch 16–32 unstable → 64–128 OOM → lower micro batch first
Warmup steps 30–100 — early loss spikes
Scheduler cosine — —
Optimizer paged_adamw_8bit — —
Dropout 0.05 overfitting → 0.1 —
Max seq length 95th pct of data OOM → lower truncated examples

Cost Estimation Table

Setup Hardware Approx. cost Best for
Colab free T4 16 GB $0 (limited hours) Learning, debugging, small runs
Colab Pro T4/V100-class ~$10/month Real student projects
Kaggle T4/P100 $0 (weekly quota) Overflow experiments
Cloud spot A100 40/80 GB ~$0.50–1.50/hr (verify current) Big/fast runs
Cloud on-demand A100 80 GB ~$1.50–4/hr (verify current) Deadline runs
University cluster Varies $0 — ask your dept! Everything

Estimation recipe: 100-step calibration run → sec/step → total steps = (examples × epochs) / batch → × number of planned runs → +50% buffer.

Troubleshooting Matrix

Symptom Likely cause Fix (in order)
Loss NaN LR too high; corrupt example Halve LR; scan data for garbage
Loss flat from step 0 LR too small; label masking bug Raise LR 3×; check template/labels
Loss oscillates wildly LR too high Halve LR; add warmup
Train ↓, val ↑ Overfitting Early stop; lower rank; more data; dropout 0.1
Val flat, train ↓ Capacity/data ceiling More diverse data; higher rank
CUDA OOM Batch/sequence too large micro batch 1 → shorter seq len → grad checkpointing
Good train, bad paraphrase Memorization Paraphrase augmentation; lower rank; more data
Task ↑, MMLU ↓↓ Catastrophic forgetting Add 5–15% replay data; fewer epochs; lower LR
Outputs ignore format Template mismatch train vs inference Use identical template; check chat template
Slow training Padding waste; small batches packing=True; larger effective batch

References

[1] E. J. Hu et al., "LoRA: Low-rank adaptation of large language models," in Proc. Int. Conf. Learning Representations (ICLR), 2022.

[2] T. Dettmers, A. Pagnoni, A. Holtzman, and L. Zettlemoyer, "QLoRA: Efficient finetuning of quantized LLMs," in Proc. Advances in Neural Information Processing Systems (NeurIPS), vol. 36, 2023.

[3] J. Wei et al., "Finetuned language models are zero-shot learners," in Proc. Int. Conf. Learning Representations (ICLR), 2022.

[4] H. W. Chung et al., "Scaling instruction-finetuned language models," arXiv preprint arXiv:2210.11416, 2022.

[5] L. Ouyang et al., "Training language models to follow instructions with human feedback," in Proc. Advances in Neural Information Processing Systems (NeurIPS), vol. 35, 2022.

[6] P. Lewis et al., "Retrieval-augmented generation for knowledge-intensive NLP tasks," in Proc. Advances in Neural Information Processing Systems (NeurIPS), vol. 33, 2020.

[7] T. Brown et al., "Language models are few-shot learners," in Proc. Advances in Neural Information Processing Systems (NeurIPS), vol. 33, 2020.

[8] J. Kaplan et al., "Scaling laws for neural language models," arXiv preprint arXiv:2001.08361, 2020.

[9] T. Dettmers, M. Lewis, S. Shleifer, and L. Zettlemoyer, "8-bit optimizers via block-wise quantization," in Proc. Int. Conf. Learning Representations (ICLR), 2022.

[10] I. J. Goodfellow, M. Mirza, D. Xiao, A. Courville, and Y. Bengio, "An empirical investigation of catastrophic forgetting in gradient-based neural networks," arXiv preprint arXiv:1312.6211, 2013.

[11] C. Zhou et al., "LIMA: Less is more for alignment," in Proc. Advances in Neural Information Processing Systems (NeurIPS), vol. 36, 2023.

[12] L. von Werra et al., "TRL: Transformer reinforcement learning," GitHub repository, 2020. [Online]. Available: https://github.com/huggingface/trl


Glossary

  • Adapter: A small set of trainable parameters added to a frozen base model (e.g., LoRA matrices). Saved and shared independently of the base model.
  • Backpropagation: The algorithm that computes how much each weight contributed to the prediction error, so the optimizer knows which direction to adjust each weight.
  • Base model: A pre-trained language model before task-specific fine-tuning — fluent and knowledgeable but not specialized.
  • Batch size: The number of training examples processed before each weight update. Effective batch size includes gradient accumulation.
  • Catastrophic forgetting: Loss of previously learned capabilities when a model is trained on new data; the mirror image of overfitting.
  • Chat template: The formatting (special tokens, role labels) a tokenizer applies to conversational messages before the model sees them.
  • Checkpoint: A saved snapshot of model/adapter weights during training, allowing resuming or selecting the best-performing point.
  • Distillation (data): Using a stronger model's outputs as training examples for a smaller model, with verification.
  • DoRA: A LoRA variant that decomposes weights into magnitude and direction components; a drop-in alternative in the PEFT library.
  • Early stopping: Ending training (or selecting a checkpoint) at the point of best validation performance rather than training for a fixed duration.
  • Epoch: One complete pass through the training dataset.
  • Few-shot prompting: Including a handful of examples in the prompt so the model imitates the pattern without any training.
  • Full fine-tuning: Updating all of a model's weights during training, as opposed to parameter-efficient methods.
  • Gradient accumulation: Simulating a large batch by summing gradients over several small micro-batches before updating weights.
  • Instruction tuning: Fine-tuning on instruction → response pairs so the model learns to follow instructions (also called supervised fine-tuning, SFT).
  • Learning rate (LR): The scaling factor applied to each weight update; the most consequential hyperparameter.
  • LoRA (Low-Rank Adaptation): A PEFT method that freezes base weights and trains the update as a product of two small low-rank matrices.
  • Loss: A single number measuring prediction error on training examples; training minimizes it.
  • Loss masking: Excluding certain tokens (e.g., the instruction) from the loss so the model only learns to produce the response.
  • Mixed precision: Training with a mix of 16-bit and 32-bit numbers to save memory and speed up computation.
  • NF4 (NormalFloat4): A 4-bit number format designed for normally distributed neural network weights, used by QLoRA.
  • Optimizer: The algorithm applying weight updates (e.g., AdamW); it often keeps extra per-parameter statistics.
  • Overfitting: Memorizing training examples instead of learning general patterns; detected when validation loss rises while training loss falls.
  • PEFT (Parameter-Efficient Fine-Tuning): Methods that train a small fraction of parameters (LoRA, adapters, prefix tuning, …).
  • QLoRA: LoRA applied on top of a 4-bit quantized base model, enabling large-model fine-tuning on modest GPUs.
  • Quantization: Storing model weights in fewer bits (e.g., 4-bit instead of 16-bit) to reduce memory.
  • RAG (Retrieval-Augmented Generation): Generating answers grounded in documents retrieved at query time, without changing model weights.
  • Rank (r): In LoRA, the inner dimension of the adapter matrices; controls adapter capacity.
  • Replay data: General examples mixed into task-specific training to prevent catastrophic forgetting.
  • Seed: The random-number seed fixing shuffling and initialization for reproducibility.
  • Test set: Held-out examples evaluated once at the end; the unbiased estimate of performance.
  • Token: The unit of text models process (roughly a word piece); text is tokenized before training.
  • Validation set: Held-out examples used during training to monitor progress and select checkpoints/hyperparameters.
  • Warmup: Gradually increasing the learning rate over the first training steps to avoid early instability.
  • Weight decay: Regularization that penalizes large weights, discouraging overfitting.

Practice Exercises

  1. Decision memo (writing). Pick a real problem from your research area. Write the one-page method-selection memo from Chapter 1: the task, quick prompting and RAG experiments with scores on 50 examples, and a justified choice. Bring it to your advisor.
  2. Baseline bake-off (hands-on). Take 100 examples of a task you care about. Score a base model zero-shot and few-shot (5 examples). Record both scores — these are the baselines your future fine-tune must beat.
  3. Toy fine-tune (hands-on). Run the Chapter 6 walkthrough verbatim on Colab with a tiny model (e.g., a 1–2B model) and 200 toy examples. Confirm the adapter changes the output format. Save the notebook as your debugging baseline.
  4. Data grading (hands-on). Collect or generate 150 candidate instruction examples for your task. Grade 100 against Chapter 3's five quality criteria. Compute the failure rate and fix your pipeline until it's under 10%.
  5. Template consistency check (hands-on). Train two tiny runs identical except for the prompt template, then evaluate each with the other template. Measure the mismatch penalty — this teaches why train/inference templates must match.
  6. Rank ablation (hands-on). On your toy data, train LoRA with r ∈ {4, 16, 64}. Plot validation loss for each. Write two sentences explaining the curve — this is the seed of a paper figure.
  7. Loss-curve diagnosis (analysis). Given these curves — train loss falling smoothly, val loss flat after epoch 2 — write a diagnosis and a three-step remediation plan, citing Chapters 7 and 9.
  8. Forgetting probe (hands-on). Take any fine-tuned adapter you can download (or your toy model). Score base vs. fine-tuned on 30 general-knowledge questions you write. Quantify the regression and write it up as a mini "forgetting report."
  9. Budget plan (writing). Fill in the Chapter 10 budgeting template for your thesis project: data hours, calibration + ablation + final GPU hours, buffer, dollar figure. Identify the single biggest cost and one way to halve it.
  10. Paper autopsy (analysis). Find a published fine-tuning paper in your area. Grade it against Chapter 12's five reviewer checks (baselines, config, data, test discipline, compute). Write a half-page review noting what's missing — then make sure your own paper doesn't repeat it.

End of Book 17 — Fine-Tuning LLMs on Your Own Data. Next in the series: Book 18.