
Book 17 of 50 · Free
Fine-Tuning LLMs on Your Own Data
29,556 words · 99 chapters · illustrated

Book 17 of 50 · Free
29,556 words · 99 chapters · illustrated
Book 17 of 50 — AstolixGen Learning Series

You have used large language models. You have written prompts, maybe built a small retrieval system, and now you are wondering: can I actually change what the model knows and how it behaves — using my own data? This book answers that question from the ground up.
Fine-tuning is the process of continuing a model's training on your own examples so it learns your task, your style, your domain, or your language. A few years ago this was the exclusive territory of labs with warehouse-sized GPU clusters. Today, with parameter-efficient methods like LoRA and QLoRA, a single student with a free Google Colab notebook can meaningfully adapt a multi-billion-parameter model on a laptop budget. That shift is the whole reason this book exists.
This book is written for researchers and publication students: MS and PhD candidates, and early AI researchers who need fine-tuning not as a party trick but as a research method — something you can run, measure, ablate, and report in a paper. Every chapter ends with a "For your research" box that translates the chapter's ideas into publishable practice: what to vary, what to measure, what to report, and what reviewers will ask about.
We assume you know Python, can use a terminal, and have trained at least a small neural network before (or have read Books 1–10 of this series, which build that foundation). We do not assume you have ever touched a GPU cluster, paid for cloud compute, or published a paper about large models. We explain every term when it first appears.
By the end of this book you will be able to:
Chapters build on each other, but each chapter also stands alone well enough to be read as a reference. The chapters with heavy code (5, 6, 7) are best read next to a computer. The chapters on evaluation, forgetting, and reporting (8, 9, 12) are best read next to your research notes. The Learning Dashboard at the end is a quick-reference appendix: method comparison table, hyperparameter cheat sheet, cost estimation table, and troubleshooting matrix. Print it, or keep it open in a second tab while you train.
A note on honesty that will come up repeatedly: fine-tuning is easy to start and easy to do badly. The difference between a fine-tuned model that genuinely improved and one that only looks improved on your favorite examples is evaluation discipline. This book treats that discipline as the main skill, not a side note.
You have a problem and a large language model. Before you spend a single rupee on GPUs or a single hour building a dataset, you need to answer one question: what is the cheapest technique that will actually solve this problem? There are three main candidates — prompting, retrieval-augmented generation (RAG), and fine-tuning — and choosing wrong costs you weeks. Choosing right can save you from training anything at all.
Prompting means you change the input to get better output. You write careful instructions, give a few examples inside the prompt (few-shot prompting), or ask the model to reason step by step (chain-of-thought). The model's weights — the billions of numbers that encode what it learned in training — do not change at all. You are steering an already-built car.
Retrieval-augmented generation (RAG) means you connect the model to an external knowledge source. When a user asks a question, your system searches your documents, retrieves the relevant chunks, and pastes them into the prompt so the model can answer from them. Again, the model's weights do not change. You have given the driver a map.
Fine-tuning means you continue training the model on your own examples, and its weights do change — permanently. You are rebuilding parts of the car itself.
This distinction — whether the weights change — is the dividing line for everything that follows. Prompting and RAG leave the model untouched; fine-tuning alters it. Altering the model is more powerful and more dangerous, more expensive, and harder to undo.
Think of it this way:
Here is the classic set of situations where fine-tuning genuinely wins:
And here is where fine-tuning is the wrong choice:
Work through these questions in order. Stop at the first one that gives you a definitive answer.
Step 1: Can the model already do this with a good prompt? Spend a day on prompt engineering first. Write clear instructions, add 3–10 diverse examples, try chain-of-thought for reasoning tasks. Test on at least 50–100 held-out examples — not the 5 you like. If the error rate is acceptable, stop. You are done, and you have saved yourself weeks.
Step 2: Is the missing piece knowledge or behavior? Ask: if I gave the model the right facts in the prompt, would it answer correctly? If yes, your problem is knowledge — use RAG. If the model has the facts but still answers in the wrong style, format, or reasoning pattern, your problem is behavior — consider fine-tuning.
A practical test: take 20 failure cases. For each, write the "ideal" retrieved context by hand and put it in the prompt. If failures mostly disappear, you have a retrieval problem. If the model still fails despite having the facts in front of it, you have a behavior problem.
Step 3: Do you have enough data? Fine-tuning needs examples — realistically hundreds at minimum, thousands ideally — of the exact input-output behavior you want. Not scraped web text; curated demonstrations. If you cannot produce or afford this data, prompting or RAG is your answer regardless of everything else.
Step 4: Will you run this at scale? If you will make millions of inference calls, the math often favors fine-tuning a small model once over prompting a large model forever. Estimate: (cost per prompt-based call − cost per fine-tuned call) × expected calls, versus the one-time cost of data + training. Chapter 10 gives you the budgeting tools.
Step 5: Can you combine approaches? This is the answer more often than textbooks admit. Fine-tune for behavior and style; use RAG for facts. A fine-tuned model that also retrieves documents is the standard architecture of serious production systems. The techniques are not rivals — they solve different halves of the problem.

Figure: The core decision — if the gap is facts, retrieve; if the gap is behavior, fine-tune; if neither, prompt better.
Suppose you are building a system that answers questions about your university's thesis regulations for MS students, in Urdu, with citations to specific regulation clauses.
Now flip it: suppose you need a model that converts natural-language questions into SQL for your university's specific database schema, thousands of times a day. The schema never changes, the SQL dialect is fixed, and latency matters. Prompting works but is slow and occasionally produces wrong column names. RAG is pointless — there are no documents to retrieve; the schema fits in the prompt anyway. This is a fine-tuning problem: the behavior (schema-faithful SQL generation) needs to be internalized, and you will amortize training cost over millions of cheap inference calls.
Students routinely underestimate what fine-tuning demands:
None of this means "don't fine-tune." It means: fine-tune when the problem genuinely requires changed behavior, and prove the cheaper options failed first. That proof — "we tried prompting and RAG, here are the numbers, and here is why they were insufficient" — is also exactly what makes a fine-tuning paper credible. Reviewers ask "why didn't you just prompt?" more often than any other question. Chapter 1 just gave you the framework to answer it.
In practice, the most capable systems rarely use one method alone. Once you understand each technique's strength, you can compose them like building blocks. Here are the hybrid patterns that appear most often in real deployments and papers:
Pattern 1: Fine-tuned model + RAG (the standard). The fine-tuned model handles behavior — tone, format, reasoning style, when to say "I don't know" — while the retriever supplies facts. Example: a fine-tuned customer-support model that always responds with empathy, a structured summary, and cited ticket IDs, where the ticket contents come from retrieval. Neither method alone produces this: RAG alone gives correct facts in inconsistent format; fine-tuning alone gives consistent format with hallucinated facts.
Pattern 2: Fine-tuned query rewriter + RAG. Retrieval quality depends heavily on the query. A small fine-tuned model rewrites the user's question into a better search query (or several), the retriever fetches documents, and a second model (or the same one) answers. The rewriter is cheap to train — a few thousand (question → better query) pairs — and often improves end-to-end accuracy more than upgrading the answering model.
Pattern 3: Distilling RAG into fine-tuning. During data collection, you generate responses with retrieval (so they're factually grounded), verify them, then fine-tune on the resulting pairs without retrieval at serving time. The model internalizes the grounded answering style. This works when the knowledge is stable (it won't change after training) and you want to drop the retrieval infrastructure for latency or cost reasons. The risk: the model memorizes facts that later go stale — document the knowledge cutoff the same way base models do.
Pattern 4: Router — fine-tuned classifier directing traffic. A tiny fine-tuned classifier decides per query whether to answer directly, retrieve, or escalate to a larger model. This is how production systems keep costs down: 80% of queries are easy and handled cheaply; the expensive paths trigger only when needed. The router itself is a fine-tuning task (classification, Chapter 3's smallest-data regime).
How to decide the composition: start from the decision framework's Step 2 (knowledge vs. behavior gap) and assign each gap its tool. A project with both gaps gets both tools. Evaluate the hybrid against each single-method baseline — the hybrid should win, and the ablation (hybrid minus one component) tells you what each part contributed. In papers, this composition analysis is often more interesting than any single component: "RAG alone: 72%, fine-tuning alone: 75%, combined: 84%, and removing the fine-tuned rewriter costs 6 points" is a complete story.
One caution: hybrids multiply failure modes. A fine-tuned model fed bad retrieved documents will confidently format misinformation beautifully — worse than either failure alone. Evaluate the combined system end to end, not just each part in isolation, and include adversarial tests where retrieval returns irrelevant or contradictory documents.
There's a failure mode this chapter hasn't named: fine-tuning too early, before you understand the problem. Symptoms: you can't describe the behavioral gap in one sentence; your "dataset" is 200 examples scraped hastily; you haven't tried few-shot prompting seriously; you chose fine-tuning because it sounds more impressive than prompting.
Premature fine-tuning wastes the scarcest resource — your judgment. A model trained on a poorly understood task learns a poorly understood behavior, and the resulting outputs teach you nothing except that the data was inadequate — something a week of prompt experimentation would have revealed for free.
The disciplined alternative is the two-week rule: spend up to two weeks on prompting and error analysis before collecting fine-tuning data. During those two weeks, you will: discover what the model already does well (shrinking your data needs), build the evaluation set and script (needed regardless), develop the error taxonomy (which becomes your data-collection guide — you now know exactly which failure modes need training examples), and write the method-selection memo. None of this work is wasted if you do fine-tune — it's the foundation. And sometimes the two weeks reveal that prompting suffices, saving you two months.
This is also a strategic point for researchers: the prompting experiments aren't just due diligence, they're content. "We first established a strong few-shot baseline of 71.8% through systematic prompt engineering (Appendix A); fine-tuning improved this to 78.5%" is a stronger paper than one that never mentions prompting. The two-week rule makes your eventual fine-tuning both better-informed and better-defended.
For your research: Before any training run, write a one-page "method selection memo": the task, the three approaches considered, the quick experiments you ran for prompting and RAG with their scores, and why fine-tuning is justified. This memo becomes the related-work and motivation section of your paper almost verbatim. Reviewers reward this explicitness enormously — it shows you chose fine-tuning deliberately, not by default.
To use fine-tuning well — and to write about it credibly — you need a mental model of what is actually happening when you "train" an already-trained model. This chapter builds that model with no calculus beyond what you already know: a model is a function with adjustable numbers, training adjusts those numbers to reduce mistakes, and fine-tuning is simply training that starts from a very good starting point instead of from scratch.
A large language model is, at bottom, a gigantic function. It takes a sequence of tokens (pieces of text) as input and produces a probability distribution over what token comes next. The behavior of that function is controlled by its parameters — often called weights — which are just numbers. A 8-billion-parameter model has 8 billion such numbers. Every capability the model has — grammar, facts, reasoning patterns, style — is encoded in the particular values of those numbers. There is no separate "knowledge database" inside; the knowledge is the arrangement of the numbers.
When the model was originally trained (pre-training), it started with random numbers and read trillions of words, adjusting the numbers little by little so that its next-token predictions got better and better. This took months on thousands of GPUs. The result is a base model: a fluent, knowledgeable, general-purpose text predictor with no particular instruction-following manners. Ask a raw base model a question and it may continue it as if writing an essay, because it was trained to continue text, not to answer questions.
Fine-tuning takes that finished machine and keeps adjusting the dials — but now on your data, with your objective. If pre-training taught the model the English language and the world's facts, fine-tuning teaches it your specific job: answer in this format, follow these instructions, reason about this domain, refuse these requests, write in this style.
Here is the mechanism in plain terms. You show the model one of your training examples — say, an instruction and its correct response. The model reads the instruction and predicts the response one token at a time, producing its own guess for each position. You compare its guess to the correct token and compute a loss: a single number measuring how wrong the prediction was. Then backpropagation figures out, for every one of the billions of weights, which direction to nudge it so the loss would have been slightly smaller. The optimizer applies those nudges, scaled by the learning rate. Repeat millions of times, and the weights settle into values that make your examples likely.
Three things about this process matter enormously for fine-tuning specifically:
1. You start from a good place, so you nudge gently. In pre-training, weights move across vast distances from random initialization. In fine-tuning, you are already near a good solution — the model already speaks the language — so you use a much smaller learning rate (typically 10–100× smaller than pre-training). Crank the learning rate too high and you blast the model out of its good configuration; the fluent generalist becomes a stuttering specialist. This is the single most common beginner error, and Chapter 7 will quantify it.
2. Every update is a trade. Each nudge that makes your examples more likely can make other things slightly less likely. The model's weights are shared across all its capabilities; there is no separate compartment for "your task." Push too hard on your narrow data and the model forgets general abilities — this is catastrophic forgetting, covered in Chapter 9. Fine-tuning is always a negotiation between the new behavior and everything the model already knew.
3. The model learns patterns, not examples. Given enough varied examples, the weights encode the pattern behind them — the format, the reasoning style, the domain conventions. Given too few examples, or examples that are too similar, the weights just memorize the specific instances. Memorization versus generalization is the central tension of dataset design (Chapter 3) and the reason evaluation on held-out data (Chapter 8) is non-negotiable.
The fine-tuning you will do 90% of the time is instruction tuning (sometimes called supervised fine-tuning, SFT). The format is simple: each training example is an instruction (what to do) plus a response (the correct way to do it). Sometimes there is also an input field for the data the instruction operates on.
Instruction: Classify the sentiment of this movie review as positive or negative.
Input: "The film dragged on for two hours and I checked my watch constantly."
Response: negative
During training, the model reads the instruction and input, and the loss is computed only on the response tokens — we do not penalize the model for failing to predict the instruction itself, since at deployment time the instruction comes from the user. (In Hugging Face's TRL library, this masking is handled for you; in raw Transformers code you implement it with label masking, as Chapter 6 shows.)
Why does this work so well? Because the base model already "knows" sentiment analysis from pre-training — it has seen millions of reviews. Instruction tuning does not teach it the concept of sentiment; it teaches it the behavior: when you see this instruction format, output exactly one word, no explanation. This is why a few thousand good examples can transform a model: you are not teaching knowledge, you are teaching manners and format. The landmark papers on this — Wei et al.'s "Finetuned language models are zero-shot learners" and Chung et al.'s "Scaling instruction-finetuned language models" — showed that instruction tuning on diverse tasks makes models dramatically better at following unseen instructions.
A transformer model is a stack of layers — typically 32 for an 8B model. It helps to know, roughly, what different parts do, because it explains why some fine-tuning methods work:
This is not a clean division — capabilities are smeared across layers — but the general principle holds: the closer to the output, the more task-specific the representations. It is one reason LoRA (Chapter 4) can get away with modifying only certain weight matrices: small, targeted changes near the right places steer behavior effectively.
The attention mechanism deserves one plain-language note since you will see its matrices named in LoRA configurations. Attention is how each token "looks at" other tokens to gather context. It is controlled by weight matrices conventionally named query, key, value, and output projection (q_proj, k_proj, v_proj, o_proj). When a LoRA config says target_modules=["q_proj", "v_proj"], it means: apply the low-rank update to the query and value matrices of every attention layer. Those were the matrices the original LoRA paper found most effective to adapt.
Imagine the model's performance as a landscape of hills and valleys, where height is the loss (lower is better) and position is the setting of all billions of weights. Pre-training dropped the model into a broad, good valley. Fine-tuning moves it to a nearby spot in that valley that is better for your task.
Two intuitions from this picture run the whole book:
Pre-training: random start, trillions of tokens, huge learning rate, months of compute, objective is "predict the next token on the internet," result is general capability. Fine-tuning: expert start, thousands to millions of tokens, tiny learning rate, hours of compute, objective is "produce exactly these responses to these instructions," result is specialized behavior. Everything about fine-tuning — the small learning rates, the few epochs, the fear of forgetting — follows from the fact that you are making small adjustments to something that already works.
Students often find this puzzling: how can 2,000 examples meaningfully change a model with 8 billion parameters? Shouldn't the model just ignore such a tiny signal, or alternatively memorize it instantly? The answer involves a genuine research finding worth knowing.
A paper by Aghajanyan et al. ("Intrinsic Dimensionality Explains the Effectiveness of Language Model Fine-Tuning," ACL 2021) measured something surprising: although a language model has billions of parameters, the fine-tuning solution for a given task can be found by adjusting only a few thousand degrees of freedom. In their experiments, fine-tuning in a random subspace of just a few thousand dimensions recovered most of full fine-tuning's performance. The model's parameter space is enormously redundant for any single task — there are countless equivalent ways to implement "answer in this format," and the optimizer only needs to find one of them.
This has three practical consequences. First, it explains why LoRA works at all: if the useful update lives in a low-dimensional subspace, a low-rank approximation captures it. LoRA's rank-16 update isn't a hack that happens to work — it's matched to the actual structure of the problem. Second, it explains why small datasets suffice for behavioral tasks: you're not filling 8 billion parameters with information; you're nudging the model along a few thousand meaningful directions, and a few thousand good examples provide enough signal for that. Third, it predicts LoRA's limitation: tasks whose solutions genuinely need many independent directions (very broad domain shifts) need higher rank or full fine-tuning — which matches the empirical guidance in Chapter 4.
There's a second, complementary intuition: the base model already contains the capability; fine-tuning mostly selects and formats it. Pre-training exposed the model to sentiment analysis, summarization, translation, and countless formats. Your 2,000 examples don't teach these from scratch — they teach the model which of its existing capabilities to deploy, when, and in what wrapper. Selection among existing capabilities needs far less data than building capabilities. This is also why fine-tuning fails most dramatically exactly where pre-training was thinnest (Chapter 11's low-resource language case): there's nothing to select, so the model must build — and building needs far more data than selecting.
Keep both intuitions handy. When someone asks "why should I believe 3,000 examples is enough?", the answer is: because we're selecting and formatting pre-existing capabilities along a low-dimensional subspace, not training from scratch — and our data-size ablation (Chapter 8) empirically confirms where the curve saturates for this task.
Abstract "nudges" become clearer with toy numbers. Imagine a miniature model with a single weight w = 2.0, and one training example where the correct output implies w should be larger. The model predicts with w=2.0, the loss is 1.0. Backpropagation computes the gradient: d(loss)/dw = −0.4 (negative means increasing w decreases loss). With learning rate 0.1, the update is:
w_new = w − LR × gradient = 2.0 − 0.1 × (−0.4) = 2.04
One step moved the weight 0.04 toward the better value. Now scale this intuition up: instead of one weight there are 8 billion, instead of one gradient direction there are 8 billion partial derivatives computed simultaneously by backpropagation, and instead of one example the gradient is averaged over a batch. The learning rate plays exactly the same role — with LR=0.1 the step was gentle; with LR=10 the same gradient would have flung w to 6.0, overshooting wildly. That's the entire mechanism; everything else is bookkeeping at scale.
Two subtleties worth carrying from this toy: first, the gradient is local — it only knows the downhill direction at the current position, not the global landscape. Training is a blindfolded descent feeling the slope one step at a time, which is why small steps and many of them work better than leaps. Second, in LoRA, this same update math applies only to the adapter matrices A and B — the base weight w stays frozen at 2.0 forever, and the effective weight becomes 2.0 + (B×A). The toy makes visible what "freezing" means: w never appears on the left side of an update equation.
For your research: Reviewers and thesis committees love a clear conceptual account of why your fine-tuning should work for your problem. Borrow the "manners vs. knowledge" framing: state explicitly whether your fine-tuning is teaching the model new domain knowledge, new behavior/format, or both — and design your evaluation to test each claim separately. A paper that says "we fine-tuned to teach radiology report structure (behavior) while relying on pre-trained medical knowledge, and our ablations separate the two" is far stronger than "we fine-tuned and the score went up."
Every experienced practitioner will tell you some version of the same sentence: the dataset matters more than the model, the method, or the hyperparameters. A mediocre training setup on excellent data beats an excellent setup on mediocre data, nearly every time. This chapter is about building that excellent dataset — what format to use, where examples come from, how to judge quality, and how much data you actually need.
Supervised fine-tuning consumes examples in a consistent structure. The most widely used is the Alpaca format, named after the Stanford project that popularized it:
{
"instruction": "Summarize the following research abstract in one sentence.",
"input": "We present a method for low-rank adaptation of large language models...",
"output": "The authors propose LoRA, a method that adapts LLMs by training small low-rank matrices instead of all weights."
}
The instruction tells the model what to do. The input (sometimes empty) provides the material to work on. The output is the correct response — the demonstration the model will learn to imitate. During training these are concatenated into a single text with a prompt template, for example:
### Instruction:
Summarize the following research abstract in one sentence.
### Input:
We present a method...
### Response:
The authors propose LoRA...
The template matters less than its consistency: use one template for the entire dataset, and use the same template at inference time. A model trained with ### Response: markers that is prompted at deployment with a different format will underperform — it learned to associate its behavior with the exact textual cues it saw in training.
Other formats you will encounter: ShareGPT format (a list of conversational turns with human/gpt roles — natural for dialogue data), and chat templates applied by the tokenizer (modern instruction models like Llama 3 expect messages wrapped in special tokens such as <|start_header_id|>; Hugging Face's tokenizer.apply_chat_template handles this, and Chapter 6 shows exactly how). The principle is identical in all cases: consistent structure, response tokens are what the model learns.
The most liberating finding for students with limited resources is that a small, excellent dataset routinely outperforms a large, mediocre one. The LIMA paper ("Less Is More for Alignment," Zhou et al., 2023) demonstrated that 1,000 carefully curated examples could produce a strong instruction-following model — a result that reshaped how the field thinks about data. The mechanism is the one from Chapter 2: fine-tuning mostly teaches behavior and format, and behavior can be demonstrated in a few hundred diverse, perfect examples. Knowledge, which needs volume, mostly comes from pre-training.
What makes an example "high quality"? Judge every candidate against these criteria:
A useful quality-control ritual: sample 100 random examples and grade each one yourself against the five criteria above. If more than 5–10% fail, fix the data pipeline before training anything. An hour of grading saves days of confused debugging.
You have four main sources, in rough order of quality:
1. Human-written examples (best, most expensive). Domain experts write instruction-response pairs. For a thesis project, this might be you and your advisor writing 500–2,000 examples over a few weeks. Slow, but the quality ceiling is highest, and for specialized domains (medicine, law, low-resource languages) there is no substitute.
2. Distillation from a stronger model. You write the instructions (or collect real user queries) and have a frontier model generate the responses, then verify and filter them. This is how most open instruction datasets were built. The critical step is verification: generate 3,000 responses, have humans or strong automated checks reject the bad 30%, keep 2,000. Unfiltered distillation bakes the teacher's errors into your student permanently.
3. Repurposed existing datasets. Academic NLP datasets (question answering, summarization, classification) can be reformatted into instruction format with templates. "Convert this SQuAD example into an instruction pair" is a legitimate and common pipeline. Watch for: template artifacts (the model learns your template's quirks), and license compatibility (see the ethics note below).
4. Synthetic generation with verification. Generate instructions programmatically or with a model, generate responses, then verify with rules, a second model, or humans. Scales well; quality depends entirely on the verification step. Never skip verification — unverified synthetic data is how models learn confident nonsense.
What to avoid: scraping raw web text and calling it a fine-tuning dataset (that is pre-training data, not instruction data); using copyrighted text you have no right to use; and — critically for researchers — training on your test set or on data contaminated with it.
Honest answer: it depends on what you are teaching, but here are practical brackets:
Start small and scale deliberately: train on 1,000 examples, evaluate (Chapter 8), then try 3,000 and 10,000. Plot performance against data size. This data ablation is cheap, informative, and exactly the kind of figure reviewers love — it answers "how much data does this task need?" which is itself a research contribution.
Split your data into three parts before you do anything else:
For small datasets (under ~2,000 examples), consider cross-validation or at least multiple random splits, because a single small test set is noisy — a 5% swing might be luck. Report the variance, not just the mean.
Here is a realistic data pipeline in Python — loading raw examples, applying a template, splitting, and saving in a training-ready format:
from datasets import Dataset
import json, random
# 1. Load your raw examples (however you collected them)
with open("raw_examples.json") as f:
raw = json.load(f) # list of {"instruction","input","output"}
# 2. Quality filter: drop empties, duplicates, overlong examples
seen = set()
clean = []
for ex in raw:
text = (ex["instruction"] + ex["input"] + ex["output"]).strip()
if not text or len(ex["output"]) < 10:
continue # empty or trivial
key = ex["instruction"][:80]
if key in seen:
continue # near-duplicate instruction
seen.add(key)
if len(text) > 6000:
continue # overlong; handle long docs separately
clean.append(ex)
print(f"kept {len(clean)}/{len(raw)} after filtering")
# 3. Apply a consistent template
TEMPLATE = ("### Instruction:\n{instruction}\n\n### Input:\n{input}\n\n### Response:\n{output}")
def format_example(ex):
return {"text": TEMPLATE.format(**ex)}
formatted = [format_example(ex) for ex in clean]
# 4. Shuffle with a fixed seed, then split (reproducibility matters)
random.Random(42).shuffle(formatted)
n = len(formatted)
train = formatted[:int(0.9*n)]
val = formatted[int(0.9*n):int(0.95*n)]
test = formatted[int(0.95*n):]
# 5. Save; the test set goes somewhere you will not touch during training
Dataset.from_list(train).to_json("data/train.json")
Dataset.from_list(val).to_json("data/val.json")
Dataset.from_list(test).to_json("data/test.json")
print(f"train={len(train)} val={len(val)} test={len(test)}")
Note the fixed random seed: reproducibility starts with data preparation, not with training. A reviewer who cannot reproduce your split cannot reproduce your results.
Two obligations researchers sometimes overlook. First, licenses: many datasets prohibit commercial use or redistribution of derivatives; fine-tuned model weights trained on such data inherit the restriction in practice. Check the license of every dataset you repurpose, and state data licenses in your paper (Chapter 12). Second, privacy: if your examples contain real people's data — patient notes, student essays, customer messages — anonymize before training. A fine-tuned model can memorize and regurgitate training examples; training on private data without consent is both unethical and, in many jurisdictions, unlawful.
If humans write or verify your examples — even if "humans" means you and one colleague — write annotation guidelines first. Without them, two annotators produce two different datasets stitched together, and the model faithfully learns the inconsistency. Good guidelines have four parts:
1. The task definition with the "why." Not just "write a response" but "write a response a rural clinic worker could read aloud to a patient; the goal is clarity, not completeness." Annotators who understand the purpose make better judgment calls on edge cases than any rule list covers.
2. Do-and-don't examples (at least 5 pairs). Show a good response and a bad response for the same instruction, with one sentence explaining the difference. This is worth more than pages of prose. Include edge cases: what to do when the instruction is ambiguous, when the correct answer is "I don't know," when the input contains errors.
3. Format specification, exact. If outputs must be JSON with keys answer and justification, say so, show it, and provide a validator script annotators run before submitting. Every format decision you leave implicit becomes noise in the training data.
4. The escalation rule. "If you're unsure, flag it instead of guessing." Guessed examples are worse than missing examples — they teach confident wrongness. Create a needs_review bucket and have your most careful annotator (possibly you) resolve it.
Measuring agreement: have two annotators independently do the same 50 examples, then compare. For classification-style outputs, compute simple agreement rate (aim for 90%+; below 80% means your guidelines are ambiguous). For free-text responses, agreement is harder to quantify — instead, have a third person blind-rank pairs of responses from the two annotators against the guidelines. Low agreement doesn't mean bad annotators; it means ambiguous guidelines. Rewrite the guidelines, not the people.
Pilot first: run 50 examples through the full pipeline (guidelines → annotation → your quality grading from earlier in this chapter) before scaling to thousands. The pilot reveals every ambiguity in your guidelines at 1/40th the cost. Every experienced dataset builder pilots; every burned one wishes they had.
A final note on you-as-annotator: if you're writing all examples yourself (common in thesis projects), you don't need inter-annotator agreement — but you do need self-consistency. Write your guidelines anyway, and re-grade your first 100 examples after finishing all 2,000. You'll find your standards drifted; normalize the early examples to your later, better-calibrated judgment.
Datasets change: you fix labels, add examples, remove contamination, rephrase instructions. Without versioning, "the dataset" becomes ambiguous — which version produced which result? Adopt this lightweight scheme from day one:
Name versions explicitly: data-v1 (initial 1,000), data-v2 (added 500 medical examples, removed 37 duplicates), etc. Keep a CHANGELOG.md in your data directory: one line per version saying what changed and why. Every training run records its data version in the run log (next to the seed and config).
Never edit in place. When you fix 50 labels, create data-v3 — don't silently overwrite data-v2. Old results must remain interpretable: "run-04 used data-v2" should still be true and checkable a year later. Disk is cheap; ambiguity is expensive.
Hash the splits. After creating train/val/test, record a hash (e.g., SHA-256 of the concatenated file) in the changelog. If anyone — including you — questions whether the test set leaked, the hash proves which bytes were evaluated.
Why this matters for papers: reviewers increasingly ask "which data version?" when results don't reproduce, and dataset versioning is standard practice in industry labs. A changelog with five entries and hashes signals a mature experimental process — and when your advisor asks "did the improvement come from the new data or the new hyperparameters?", the version log plus run log answers definitively instead of approximately.
For your research: Your dataset is a research artifact. Document it like one: collection method, size, filtering steps, inter-annotator agreement if humans wrote examples, license, and a datasheet-style description. Releasing your dataset (when legally possible) alongside your paper multiplies its impact — datasets get cited independently. Even a 2,000-example high-quality dataset for an under-served language or domain is publishable in dataset tracks of major venues.
Chapter 2 described fine-tuning as nudging all billions of weights. That is full fine-tuning, and it works — but it needs enormous memory: the weights, the gradients, and the optimizer's bookkeeping all live on the GPU at once. For a 7-billion-parameter model in standard 16-bit precision, full fine-tuning needs roughly 60+ GB of GPU memory, far beyond any single student-grade GPU. Parameter-efficient fine-tuning (PEFT) is the family of methods that dodges this wall by training only a tiny fraction of the parameters. LoRA is its most important member, and QLoRA made it run on free hardware. This chapter explains both deeply enough that you can configure them with understanding rather than superstition.
To see what PEFT saves, count what full fine-tuning stores per parameter. Take a 7B model in 16-bit (2 bytes per number):
Total: roughly 84 GB before activations (the intermediate values stored for backpropagation). Even with memory tricks, you need a data-center GPU. And after training, you have a full 14 GB copy of weights per task — ten tasks means ten full models to store and serve.
PEFT attacks both problems: train far fewer numbers, and store far fewer numbers per task.
LoRA — Low-Rank Adaptation, introduced by Hu et al. — starts from an observation about fine-tuning: the change made to the weights during fine-tuning has low "intrinsic rank." In plain terms: even though a weight matrix has millions of entries, the useful update to it can be captured by multiplying two much smaller matrices.
Concretely, take one weight matrix W of size d×d (say 4096×4096 ≈ 16.8 million numbers). During fine-tuning, it would change to W + ΔW. LoRA says: freeze W completely (never update it, never store gradients or optimizer state for it), and learn the update as a product of two small matrices:
ΔW = B × A, where B is d×r and A is r×d, with r (the rank) tiny — typically 8, 16, 32, or 64.
With r=16 and d=4096: B has 4096×16 = 65,536 numbers, A has 16×4096 = 65,536 numbers — about 131,000 trainable numbers instead of 16.8 million. That is roughly a 128× reduction for that matrix. Across the whole model, a LoRA adapter is typically 0.1–1% of the base model's parameters — tens of megabytes instead of gigabytes.
At inference time, the effective weight is W + BA. You can even merge the adapter into the base weights (W' = W + BA) so the deployed model runs exactly as fast as the original — zero inference overhead. Or keep adapters separate and swap them: one base model, many task adapters, each a small file. For research, this is wonderful: your "model" for the paper is a 100 MB adapter file anyone can download and plug into the public base model.
Two clever details make LoRA train well:

Figure: The frozen base weight plus the two small trainable matrices whose product forms the update.
LoRA slashed the trainable parameters, but the frozen base weights still sit on the GPU in 16-bit — 14 GB for a 7B model, plus activations. QLoRA (Dettmers et al.) attacks the frozen weights with quantization: storing each weight in 4 bits instead of 16.
Four-bit might sound brutally lossy, but two innovations make it work:
The result: a 7B model's frozen weights fit in about 3.5–4 GB. During training, weights are dequantized to 16-bit on the fly for the forward/backward pass, gradients flow only into the LoRA adapters, and the base weights are never updated. Memory for QLoRA fine-tuning of a 7–8B model lands around 8–12 GB total — fitting on a free Colab T4 GPU (16 GB). The QLoRA paper showed this recipe matching full 16-bit fine-tuning quality across benchmarks, and even enabled fine-tuning a 65B model on a single 48 GB GPU.
The one cost of QLoRA is speed: dequantizing on the fly adds compute overhead, so training runs somewhat slower than 16-bit LoRA. For students, this is an excellent trade — slower but possible beats fast but impossible.
Here is what a LoRA configuration actually looks like, with every knob explained:
from peft import LoraConfig, get_peft_model, prepare_model_for_kbit_training
peft_config = LoraConfig(
r=16, # rank: the "width" of the adapter. Higher = more capacity.
lora_alpha=32, # scaling factor; effective update scaled by alpha/r = 2 here
lora_dropout=0.05, # dropout on the adapter: mild regularization
bias="none", # don't train bias terms (keeps adapter tiny)
task_type="CAUSAL_LM", # next-token prediction (standard for instruction tuning)
target_modules=["q_proj", "v_proj"],
# which weight matrices get adapters.
# Common choices: ["q_proj","v_proj"] (original paper, minimal),
# or all of ["q_proj","k_proj","v_proj","o_proj","gate_proj","up_proj","down_proj"] (more capacity)
)
model = prepare_model_for_kbit_training(model) # needed for quantized base models
model = get_peft_model(model, peft_config)
model.print_trainable_parameters()
# e.g. "trainable params: 41,943,040 || all params: 8,030,000,000 || trainable%: 0.52%"
That final printout is your sanity check: trainable parameters should be well under 1–2% of the total. If it says 100%, you forgot to freeze something.
Choosing r (rank): r=8 or 16 is the standard starting point and is enough for most format/style adaptation. Increase to 32–64 for harder tasks (new domains, complex reasoning). Higher rank = more capacity but more memory and more overfitting risk on small data. A good research practice: ablate r ∈ {8, 16, 32} on your validation set and report the curve.
Choosing target modules: adapting only attention (q_proj, v_proj) is the conservative default from the original paper. Adapting all linear layers (attention + MLP) gives the adapter more capacity and often better results on hard tasks, at the cost of a larger adapter file (still tiny relative to the model). For your first run, use the attention-only default; expand if validation performance plateaus.
PEFT is not always the answer. Full fine-tuning remains preferable when:
For a student on limited hardware doing instruction tuning, the honest default is QLoRA. It is what makes this book's Chapter 6 walkthrough possible on free Colab, and it is a legitimate, publishable method — the QLoRA paper itself is a NeurIPS publication, and hundreds of papers since have used it as their training method.
You will encounter these names; know what they are at one paragraph each:
use_dora=True) for an ablation.For your research, LoRA/QLoRA is the baseline everything else is compared against. If you experiment with alternatives, LoRA is the control condition.
Let's make the memory savings concrete with numbers you can verify on your own GPU. Take an 8B-parameter model and compare three training setups. (Byte counts: 16-bit = 2 bytes, 32-bit = 4 bytes, 4-bit = 0.5 bytes.)
Setup A: Full fine-tuning, 16-bit weights, AdamW. - Weights: 8B × 2 = 16 GB - Gradients: 8B × 2 = 16 GB - AdamW states (two 32-bit numbers per parameter): 8B × 8 = 64 GB - Subtotal: ~96 GB, plus activations (several GB more depending on sequence length and batch size) - Verdict: needs multiple data-center GPUs or exotic sharding. Not a student setup.
Setup B: LoRA (16-bit base), r=16 on all linear layers. - Frozen base weights: 16 GB (no gradients, no optimizer state — frozen means frozen) - Trainable adapter: ~42M parameters → weights 0.08 GB, gradients 0.08 GB, AdamW states 0.34 GB - Subtotal: ~16.5 GB plus activations - Verdict: fits on a 24 GB GPU (RTX 4090, A10), tight on 16 GB once activations are counted. This is why LoRA alone, without quantization, still strains a Colab T4.
Setup C: QLoRA (4-bit NF4 base), r=16. - Frozen base weights: 8B × 0.5 = 4 GB - Adapter + optimizer states: ~0.5 GB (as above; paged_adamw_8bit halves the optimizer memory further) - Activations with gradient checkpointing: ~2–4 GB for batch 2, seq length 1024 - Subtotal: roughly 7–10 GB - Verdict: comfortable on a 16 GB T4 with headroom. This is the Chapter 6 walkthrough, and now you see exactly why it fits.
Two lessons fall out of the arithmetic. First, the optimizer states are the silent killer in full fine-tuning — 64 GB of the 96 GB total. Any method that avoids optimizer state on the base weights (which is all of PEFT) wins enormously. Second, quantization's contribution is specifically shrinking the frozen weights (16 GB → 4 GB); the adapter was already tiny. When someone asks "why not just LoRA without quantization?", the answer is the 12 GB difference in frozen-weight storage — exactly the gap between fitting and not fitting on free hardware.
You can confirm all of this empirically: after loading your model in Chapter 6's Step 1, run nvidia-smi and compare against the 4 GB prediction; after attaching the adapter, run model.print_trainable_parameters() and multiply out the bytes. Making the theory touch the metal once builds intuition that lasts.
The PEFT library offers LoRA variants selectable with one-line config changes. Know what they are so you can run them as ablations:
use_dora=True in LoraConfig. The adapter is slightly larger and training slightly slower — worth one ablation run.use_rslora=True) makes high-rank runs behave more sensibly.For your first project, vanilla LoRA is the right baseline — it's what the literature compares against, and variant gains are usually 0–2 points. Run one variant (DoRA is the best-supported choice) as an ablation after your main results are solid. A paper reporting "LoRA 78.5, DoRA 79.1" shows thoroughness; a paper using an exotic variant as its only method invites questions about whether vanilla LoRA would have sufficed.
For your research: Your paper's methods section must state the PEFT configuration completely: rank r, alpha, dropout, target modules, quantization type (NF4, double quantization on/off), and the base model with exact version/hash. "We used LoRA" is not reproducible — two papers with "LoRA" and different ranks are different experiments. Also report the trainable parameter count and percentage; reviewers use it to judge whether your comparison to full fine-tuning baselines is fair.
You now understand what fine-tuning is and which method to use. This chapter is about the software that does the heavy lifting. The open-source ecosystem has converged on a standard stack, and learning it once pays off across every project: Hugging Face Transformers (models and training primitives), PEFT (the adapter library), TRL (high-level fine-tuning trainers), and Axolotl (configuration-driven training). We will cover what each does, when to reach for it, and get your environment set up.
Hugging Face Transformers is the foundation. It provides: thousands of pre-trained models loadable in two lines (AutoModelForCausalLM.from_pretrained(...)), tokenizers that convert text to the numbers models consume, and the Trainer class — a general training loop handling batching, optimization, checkpointing, logging, and multi-GPU distribution. Everything else in this chapter builds on it. If you learn one library deeply, make it this one.
PEFT (Parameter-Efficient Fine-Tuning) is Hugging Face's adapter library. It implements LoRA, QLoRA support, DoRA, IA³, prefix tuning, and more, with a uniform interface: define a config, wrap your model with get_peft_model, train normally, save a small adapter. Chapter 4 already showed its core API.
TRL (Transformer Reinforcement Learning — the name is historical; it now covers all post-training) provides SFTTrainer, a specialized trainer for supervised fine-tuning that handles the fiddly details: applying chat templates, masking instruction tokens from the loss, packing short examples together for efficiency, and integrating with PEFT in a few lines. For instruction tuning, SFTTrainer is the single most productive tool in the ecosystem — it turns what used to be 300 lines of custom training code into about 30.
Axolotl sits one level higher: instead of writing Python training scripts, you write a YAML config file describing everything (base model, dataset, LoRA settings, hyperparameters), and Axolotl runs the training. This is superb for research because the config file is the experiment specification — version-controlled, diffable, and directly publishable. Many published fine-tuning experiments now ship their Axolotl configs as reproducibility artifacts.
Supporting libraries you will meet: bitsandbytes (the quantization engine behind QLoRA's 4-bit loading), datasets (Hugging Face's data library — streaming, caching, and preprocessing for huge files), and accelerate (handles device placement and multi-GPU behind the scenes).
Trainer + PEFT. You see every step; debugging teaches you the most.SFTTrainer + PEFT. Minimal code, correct defaults, fast iteration.DPOTrainer). Beyond this book's scope, but good to know the same library covers it.A common student path: prototype with SFTTrainer in a Colab notebook (Chapter 6), then graduate to Axolotl configs when running the serious ablation grid for the paper.
You need Python 3.10+, PyTorch with CUDA, and a GPU with at least ~12–16 GB VRAM for QLoRA on 7–8B models (a Colab T4 qualifies). Install the stack:
# Core stack (versions move fast; these were current and mutually compatible
# at time of writing — check the TRL docs if you hit version conflicts)
pip install torch --index-url https://download.pytorch.org/whl/cu121
pip install transformers datasets accelerate peft trl bitsandbytes
Verify the GPU is visible and quantization works:
import torch
print(torch.cuda.is_available()) # must be True
print(torch.cuda.get_device_name(0)) # e.g. "Tesla T4"
from transformers import AutoModelForCausalLM, BitsAndBytesConfig
bnb_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4", # NormalFloat4, per the QLoRA paper
bnb_4bit_use_double_quant=True, # double quantization: saves a bit more memory
bnb_4bit_compute_dtype=torch.bfloat16 # compute in bfloat16 for stability
)
model = AutoModelForCausalLM.from_pretrained(
"meta-llama/Llama-3.1-8B", # any causal LM; gated models need HF access approval
quantization_config=bnb_config,
device_map="auto",
)
print("quantized model loaded OK")
Two setup notes that save hours of confusion:
huggingface-cli login) before download. Do this once; the token is cached.huggingface_hub's cache settings if you run many experiments.Axolotl installs as a package and runs from a YAML file:
pip install axolotl
# or, for the latest: pip install git+https://github.com/axolotl-ai-cloud/axolotl.git
A minimal but complete Axolotl config for QLoRA instruction tuning:
# axolotl-config.yaml
base_model: meta-llama/Llama-3.1-8B
model_type: LlamaForCausalLM
tokenizer_type: AutoTokenizer
load_in_4bit: true
adapter: qlora
lora_r: 16
lora_alpha: 32
lora_dropout: 0.05
lora_target_modules: [q_proj, v_proj]
datasets:
- path: data/train.json
type:
system_prompt: ""
field_system: null
field_instruction: instruction
field_input: input
field_output: output
format: "{instruction} {input} {output}"
no_input_format: "{instruction} {output}"
val_set_size: 0.05
sequence_len: 2048
sample_packing: true
num_epochs: 3
micro_batch_size: 2
gradient_accumulation_steps: 8
learning_rate: 0.0002
optimizer: paged_adamw_8bit
lr_scheduler: cosine
warmup_steps: 30
output_dir: ./outputs/run-01
save_steps: 200
logging_steps: 10
Run it with axolotl train axolotl-config.yaml. Every hyperparameter in this file will be explained in Chapter 7; for now, notice the property that matters: this file fully specifies the experiment. Commit it to git with a tag, and anyone — including you in six months — can reproduce the run exactly.
A clear division of labor prevents both magical thinking and unnecessary suffering:
Beginners often invert this: they obsess over library internals while feeding the model sloppy data and evaluating on vibes. Chapters 3 and 8 exist to keep you on the right side of that line.
Chapter 5's takeaway said "freeze and record exact library versions." Here's how to actually do it so the environment still works in six months:
The requirements file. After a working run, capture the environment:
pip freeze > requirements-run01.txt
This file goes into git next to the Axolotl config or training script. When you (or a reviewer) need to rerun, create a fresh environment and pip install -r requirements-run01.txt. Note: pip freeze on Colab captures Colab-specific packages too — that's fine; what matters is that transformers/peft/trl/bitsandbytes/torch versions are pinned.
The one-line version log. In addition to the freeze file, print versions into your training logs:
import transformers, peft, trl, torch
print(f"transformers={transformers.__version__} peft={peft.__version__} trl={trl.__version__} torch={torch.__version__}")
When a run from three months ago behaves differently today, the first thing you check is whether the environment changed — and this line answers it in seconds.
Randomness control. Set seeds everywhere, not just in the trainer: Python's random, NumPy, and PyTorch all have separate RNGs. Hugging Face's set_seed(42) handles all three plus CUDA. Do it at the top of every script. Perfect determinism on GPUs isn't guaranteed even with seeds (some CUDA operations are nondeterministic), but seeded runs are close enough that a large discrepancy signals a real problem, not noise.
What to version-control. Commit: configs, training scripts, eval scripts, requirements files, the data preparation script, and a README describing the run order. Do NOT commit: model weights, datasets (use releases or data versioning if large), or API tokens. A good rule: the repository should let a stranger go from "git clone" to "reproduced results" following only the README.
The environment decay problem. Libraries move fast; a requirements file from today may fail to install in a year (a dependency dropped an old wheel, a CUDA version mismatch). For thesis-critical reproducibility, the gold standard is a container (Docker) or a full environment export. For most student projects, the pinned requirements file plus the version log is sufficient — just re-verify the install works before you need it, e.g., once at thesis-writing time.
The stack is only half the decision — which model to fine-tune matters enormously, and students often default to whatever is most famous. Here's a deliberate selection workflow:
Step 1: List candidates in your size class. For QLoRA on a 16 GB GPU: 7–9B models (Llama-3.1-8B, Qwen2.5-7B, Gemma-2-9B, Phi-3-medium). For CPU-only or tiny GPUs: 1–3B models (Qwen2.5-1.5B, Llama-3.2-1B/3B). Don't start with 70B models — prove the pipeline on 8B first.
Step 2: Check the license. Llama and Gemma are free for research but have custom licenses with conditions; Qwen and Phi have their own terms; fully open models (e.g., OLMo, Falcon) use permissive licenses. If your thesis might become a commercial product, license matters from day one. Note the license in your experiment plan.
Step 3: Run a 200-example probe. Before any training, score each candidate zero-shot and few-shot on 200 examples representative of your task. This takes an afternoon and often reveals a clear winner — e.g., Qwen outperforming Llama on multilingual tasks, or Phi punching above its size on reasoning. The probe also gives you the base baselines Chapter 8 requires. Skipping this and discovering mid-project that another base model was 8 points better is an expensive mistake.
Step 4: Check tokenizer fit for your language. Tokenizers trained mostly on English split other languages into many tiny pieces ("fertility" — tokens per word). High fertility means your sequences are longer, training is slower, and the model sees less context. A quick check: tokenize 100 sentences of your data with each candidate's tokenizer and compare average tokens per sentence. For non-English projects, this single check can matter more than benchmark scores.
Step 5: Prefer instruction-tuned variants as starting points. For instruction tuning, starting from an instruct model (e.g., Llama-3.1-8B-Instruct) rather than the raw base usually works better — it already follows instructions, so your fine-tuning refines rather than teaches the format. The exception: if the instruct variant's alignment conflicts with your task (rare), use the base.
Document the probe results in a small table in your paper's appendix. "We selected Qwen2.5-7B-Instruct after a 200-example probe (62.1 vs. 54.3 for Llama-3.1-8B)" is one sentence that preempts the reviewer's "why this model?" question entirely.
Environment setup fails in predictable ways. Here are the four most common, with fixes:
1. bitsandbytes CUDA errors. The classic: CUDA SETUP: ... library not found or version mismatch warnings. bitsandbytes ships precompiled against specific CUDA versions; Colab's CUDA usually works with the default pip install, but on custom machines you may need pip install bitsandbytes --no-cache-dir or the source build. First check: does python -c "import bitsandbytes" succeed? If the import works, the 4-bit loading in Step 1 of Chapter 6 will almost certainly work too.
2. torch / transformers version conflicts. Symptom: import errors mentioning accelerate or missing trainer arguments (e.g., eval_strategy vs the older evaluation_strategy). The ecosystem renames arguments across versions — eval_strategy is current; older code uses evaluation_strategy. When copying code from blog posts, check which transformers version it assumed. Your pinned requirements file (Chapter 5's reproducibility section) is the defense.
3. "Model not found" on gated models. You accepted the license but downloads still fail → you forgot huggingface-cli login, or you're logged in with a different account than the one that accepted the license. Run huggingface-cli whoami to verify. For non-gated alternatives needing no approval (Phi-3-mini, Qwen2.5), skip the gate entirely while learning.
4. Disk full during download. The 8B download stalls partway. Check df -h; clear the pip cache and old checkpoints. On Colab, a factory reset (Runtime → Disconnect and delete runtime) gives a clean slate — then reinstall from your pinned requirements in one cell.
General rule: fix the environment once, then freeze it. Every hour spent fighting installs after the pipeline works is an hour stolen from data and evaluation.
For your research: Decide your tooling before the main experiments and freeze versions (
pip freeze > requirements.txt, or a Docker/container spec). Note the exact library versions in your paper's appendix or repository README. "Transformers 4.x" is not reproducible; "transformers==4.44.2, peft==0.12.0, trl==0.9.6" is. Reviewers increasingly check this, and future-you will need it when a library update breaks your old scripts two weeks before a deadline.
pip install transformers datasets accelerate peft trl bitsandbytes, verify GPU + 4-bit loading before anything else.huggingface-cli login once; watch Colab disk space.This is the chapter where theory becomes a running model. We will fine-tune an 8-billion-parameter instruction model with QLoRA on a single 16 GB GPU — the kind Google Colab gives you for free — from data loading to a saved, usable adapter. Follow along in a Colab notebook (GPU runtime: Runtime → Change runtime type → T4 GPU) or any machine with a CUDA GPU. Every step is explained; nothing is magic.
Open a fresh Colab notebook with a GPU runtime and run:
# Install the stack (takes a few minutes)
!pip install -q transformers datasets accelerate peft trl bitsandbytes
import torch
assert torch.cuda.is_available(), "No GPU! Check Runtime > Change runtime type."
print(torch.cuda.get_device_name(0))
print(f"VRAM: {torch.cuda.get_device_properties(0).total_memory / 1e9:.1f} GB")
You should see something like Tesla T4 and VRAM: 15.9 GB. If Colab gives you a weaker GPU, the walkthrough still works — training will just be slower.
We will use a small, strong open model. Llama-3.1-8B is the canonical choice (requires free Hugging Face access approval); alternatives that need no approval include Microsoft's Phi-3-mini or Google's Gemma-2-2B. Pick one and stick with it.
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
MODEL_ID = "meta-llama/Llama-3.1-8B" # or "microsoft/Phi-3-mini-4k-instruct"
bnb_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4",
bnb_4bit_use_double_quant=True,
bnb_4bit_compute_dtype=torch.bfloat16,
)
model = AutoModelForCausalLM.from_pretrained(
MODEL_ID,
quantization_config=bnb_config,
device_map="auto",
)
tokenizer = AutoTokenizer.from_pretrained(MODEL_ID)
tokenizer.pad_token = tokenizer.eos_token # causal LMs often lack a pad token; eos works
tokenizer.padding_side = "right"
Watch the memory after loading: !nvidia-smi should show roughly 5–6 GB used. The 8B weights in 4-bit occupy under 5 GB; the rest is overhead.
For this walkthrough we will build a tiny demonstration dataset — 300 examples teaching the model to answer questions in a strict format (a short answer followed by a one-line justification). In your real project this is where Chapter 3's pipeline goes. Here, we generate a toy dataset so the mechanics are visible end to end:
from datasets import Dataset
# Toy data: in your project, replace this with your real curated dataset
topics = ["photosynthesis", "gravity", "elections", "vaccines", "inflation"]
def make_example(i):
t = topics[i % len(topics)]
return {
"instruction": f"Explain {t} in simple terms.",
"input": "",
"output": f"{t.capitalize()} is a process where ... [short answer].\nWhy: [one-line justification]."
}
raw = [make_example(i) for i in range(300)]
def to_chat(ex):
# SFTTrainer works best with chat-formatted messages
messages = [
{"role": "user", "content": ex["instruction"] + (" " + ex["input"] if ex["input"] else "")},
{"role": "assistant", "content": ex["output"]},
]
return {"messages": messages}
dataset = Dataset.from_list([to_chat(ex) for ex in raw])
dataset = dataset.train_test_split(test_size=0.1, seed=42)
print(dataset)
# DatasetDict({train: 270 rows, test: 30 rows})
Two things to notice: we use the chat format (lists of role/content messages) because modern tokenizers apply the model's native chat template via apply_chat_template, and we split off a validation set (the 30 rows) that training will never see. In a real project your dataset would be thousands of examples built per Chapter 3 — the code shape is identical.
from peft import LoraConfig, get_peft_model, prepare_model_for_kbit_training
model = prepare_model_for_kbit_training(model)
peft_config = LoraConfig(
r=16,
lora_alpha=32,
lora_dropout=0.05,
bias="none",
task_type="CAUSAL_LM",
target_modules=["q_proj", "k_proj", "v_proj", "o_proj",
"gate_proj", "up_proj", "down_proj"],
)
model = get_peft_model(model, peft_config)
model.print_trainable_parameters()
Expected output: something like trainable params: 41,943,040 || all params: 8,031,261,440 || trainable%: 0.52%. Half a percent. That tiny fraction is everything you are about to train.
from trl import SFTTrainer, SFTConfig
training_args = SFTConfig(
output_dir="./outputs/walkthrough-01",
num_train_epochs=3,
per_device_train_batch_size=2,
gradient_accumulation_steps=8, # effective batch = 2*8 = 16
learning_rate=2e-4,
lr_scheduler_type="cosine",
warmup_steps=30,
optim="paged_adamw_8bit", # 8-bit optimizer: less memory, per Dettmers et al.
max_seq_length=1024,
packing=False, # keep False while learning; True later for speed
eval_strategy="steps",
eval_steps=50,
save_steps=100,
save_total_limit=2,
logging_steps=10,
report_to="none", # set to "tensorboard" if you want live plots
seed=42,
)
trainer = SFTTrainer(
model=model,
args=training_args,
train_dataset=dataset["train"],
eval_dataset=dataset["test"],
)
trainer.train()
What happens now: the trainer tokenizes your messages with the chat template, masks the user turns so loss is computed only on assistant responses, and runs the optimization loop. On a T4, 270 examples × 3 epochs takes roughly 10–25 minutes. You will see the training loss printed every 10 steps — it should trend downward, noisily. The eval loss every 50 steps tells you whether the model is actually improving on unseen examples (Chapter 9 explains how to read these curves).
If you run out of memory: the three emergency levers are (1) per_device_train_batch_size=1, (2) max_seq_length=512, and (3) gradient_checkpointing=True (trades compute for memory by recomputing activations). Apply in that order.
trainer.save_model("./outputs/walkthrough-01/final")
# This saves ONLY the adapter (~80-150 MB), not the 8B base model.
Now the moment of truth — generate with the fine-tuned model and compare against the base:
model.eval()
prompt = "Explain gravity in simple terms."
inputs = tokenizer.apply_chat_template(
[{"role": "user", "content": prompt}],
return_tensors="pt", add_generation_prompt=True,
).to(model.device)
with torch.no_grad():
out = model.generate(**{"input_ids": inputs}, max_new_tokens=120, do_sample=False)
print(tokenizer.decode(out[0], skip_special_tokens=True))
You should see the response follow your trained format (short answer + "Why:" line). To feel the difference, load the base model without the adapter and run the same prompt — the base model answers in its generic style. That contrast, however small on toy data, is fine-tuning made visible.
To use the adapter later without retraining:
from peft import PeftModel
base = AutoModelForCausalLM.from_pretrained(MODEL_ID, quantization_config=bnb_config, device_map="auto")
model = PeftModel.from_pretrained(base, "./outputs/walkthrough-01/final")
To produce a single standalone model (e.g., for deployment), merge the adapter into the base weights in 16-bit and save:
# Load base in 16-bit (no quantization) for a clean merge — needs more VRAM/RAM
base16 = AutoModelForCausalLM.from_pretrained(MODEL_ID, torch_dtype=torch.bfloat16, device_map="auto")
merged = PeftModel.from_pretrained(base16, "./outputs/walkthrough-01/final").merge_and_unload()
merged.save_pretrained("./outputs/walkthrough-01/merged")
Merging needs the full model in 16-bit (~16 GB for 8B), which may exceed Colab VRAM — do it on a bigger machine or accept the adapter format, which is the standard for research sharing anyway.
You downloaded an 8B model, quantized it to 4-bit, attached a 42M-parameter trainable adapter (0.5% of weights), trained it on 270 examples with loss masked to assistant responses, validated on held-out examples, and saved a ~100 MB adapter that changes the model's behavior. Total cost: zero rupees, one Colab session. This is the complete loop that every fine-tuning project — from student experiments to industry pipelines — follows. Everything else in this book is about doing each step well: better data (Ch. 3), better hyperparameters (Ch. 7), honest evaluation (Ch. 8), and credible reporting (Ch. 12).
When you launch trainer.train(), you'll see a stream of log lines. Here's how to read them like a practitioner, using a typical healthy run as reference:
Step 10: loss 2.84, lr 6.7e-05
Step 20: loss 2.31, lr 1.3e-04
Step 30: loss 1.92, lr 2.0e-04 <- warmup ends, LR at full value
Step 40: loss 1.75, lr 1.99e-04
Step 50: loss 1.61, eval_loss 1.70
Step 60: loss 1.58, lr 1.98e-04
...
Step 200: loss 0.94, eval_loss 1.12
What healthy looks like: loss starts high (the model is surprised by your data's format — 2.5–3.5 is typical for a new task), drops steeply in the first 10–20% of steps (the model quickly learns the format), then declines gradually. Eval loss tracks training loss with a small gap. The learning rate ramps up during warmup (first 30 steps here) then barely visibly decays under cosine scheduling.
Checkpoints to watch for: - Steps 1–30 (warmup): loss should fall, not spike. A spike here with recovery is usually harmless; a spike without recovery means LR too high. - First eval (step 50): eval_loss moderately above train loss is normal. Eval loss far above train loss (2× or more) this early suggests a data or template bug — e.g., eval examples formatted differently. - Mid-training: both losses falling slowly. This is the long middle where the model absorbs the task. Resist the urge to intervene; noise is normal. - Late training: train loss still falling, eval loss flattening or ticking up — you're entering the overfitting zone (Chapter 9). Note the step where eval was lowest; that's your candidate best checkpoint.
When to abort a run (and it's always okay to abort): loss NaN at any point; loss flat after 15% of steps (something structural is wrong — check the data pipeline before burning more compute); eval loss 3× train loss (template/label bug); or GPU memory errors that persist after the OOM triage. A two-hour run aborted at minute 20 with a lesson learned is cheap tuition. A two-hour run left to finish broken is just waste.
The qualitative spot-check: every ~100 steps, generate from 2–3 fixed prompts (keep them constant across the whole run) and eyeball the outputs. Numbers tell you that it's learning; outputs tell you what it's learning. If the loss falls but outputs get worse (more repetitive, format drifting), trust the outputs — the metric is lying about what you care about, and you need a better eval (Chapter 8).
Once the basic walkthrough works, three upgrades make it production-grade for real experiments:
Sequence packing. In the walkthrough, packing=False means each training example is padded to max_seq_length — with short examples, most of the GPU's work is processing padding tokens, pure waste. Setting packing=True concatenates multiple short examples into each sequence (with proper attention masking so examples don't attend to each other). Throughput often nearly doubles on short-example datasets. Enable it only after the pipeline is debugged — packing makes per-example debugging harder, which is why the walkthrough starts without it.
Resuming from checkpoints. Long Colab sessions disconnect. Because we set save_steps=100 and save_total_limit=2, interrupted runs leave checkpoints behind. Resume with:
trainer.train(resume_from_checkpoint="./outputs/walkthrough-01/checkpoint-500")
The trainer restores model weights, optimizer state, and the data loader position — training continues as if uninterrupted. Make resuming your default reflex after any disconnection rather than restarting from scratch.
Scaling to real data sizes. The walkthrough's 270 examples finish in minutes; your real dataset of 5,000–20,000 examples needs hours. Two adjustments: raise eval_steps and save_steps proportionally (evaluating every 50 steps on a 5,000-example run wastes time — every 200–500 steps is plenty), and lower logging_steps noise by watching trends rather than individual values. Everything else — the code shape, the adapter, the eval discipline — is identical. That's the point of the walkthrough: the toy run and the thesis run differ in scale, not in structure.
A note on do_sample=False in evaluation. The walkthrough generates with greedy decoding (always the most likely token) for determinism — the same prompt gives the same output every time, which is what you want when comparing models. For qualitative exploration (getting a feel for the model's style), sampling with temperature=0.7 shows the range of behaviors. Use greedy for measurement, sampling for exploration, and never mix them within one comparison.
Your trained adapter is a research artifact worth sharing — and the Hub makes it a three-line operation:
from peft import PeftModel
# Push the adapter (only ~100 MB) to the Hub
model.push_to_hub("your-username/my-task-adapter-v1")
tokenizer.push_to_hub("your-username/my-task-adapter-v1")
Anyone can then reproduce your results without retraining:
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
base = AutoModelForCausalLM.from_pretrained("meta-llama/Llama-3.1-8B", device_map="auto")
tokenizer = AutoTokenizer.from_pretrained("your-username/my-task-adapter-v1")
model = PeftModel.from_pretrained(base, "your-username/my-task-adapter-v1")
Three practices make a shared adapter genuinely useful rather than a mysterious binary: (1) fill in the model card — base model, dataset description, training config, evaluation scores, intended use and limitations; (2) version your uploads (-v1, -v2) instead of overwriting, so citations remain stable; (3) link the adapter from your paper's repository README alongside the eval script. A well-documented adapter gets downloaded, cited, and built upon — the closest thing fine-tuning research has to compounding interest. (Note: only push adapters trained on data you have the right to share, and never push datasets containing private information — Chapter 3's licensing and privacy notes apply here too.)
For your research: Run this walkthrough verbatim before touching your real data. It validates your entire environment (GPU, installs, disk, versions) in under an hour and gives you a known-good training script to adapt. Keep the notebook — when your real run breaks at 2 AM, you will diff against this working baseline to find what changed. Every serious practitioner keeps a "hello world" training script for exactly this reason.
Hyperparameters are the settings you choose before training that control how the model learns: how big each step is, how many trainable parameters the adapter has, how many times the model sees your data, how many examples it processes at once. Beginners treat them as mystical incantations copied from blog posts. Researchers treat them as experimental variables with known effects, sensible defaults, and a cheap tuning procedure. This chapter makes you the second kind.
Dozens of knobs exist; four dominate outcomes in QLoRA instruction tuning.
The learning rate scales every weight update (Chapter 2). Too small and the model barely learns — training loss crawls down, you waste compute, and the model underfits your data. Too large and updates overshoot: loss spikes, oscillates, or explodes to NaN, and the model can be damaged — fluent capabilities degrade because you blasted the weights out of their good valley.
Sensible range for QLoRA instruction tuning: 1e-4 to 3e-4 (0.0001 to 0.0003). The QLoRA paper used 2e-4, and it remains the default starting point across the field. Full fine-tuning uses much smaller values (1e-5 to 5e-5) because it updates all weights — with LoRA's tiny parameter count, a larger LR is safe and needed.
How to diagnose LR problems from the loss curve: - Loss decreases smoothly then plateaus → LR is fine; you may need more epochs or data. - Loss barely moves after hundreds of steps → LR too small (or data/adapter problem). Try 3× larger. - Loss oscillates wildly or spikes upward → LR too large. Halve it and restart. - Loss becomes NaN → LR far too large, or mixed-precision instability. Halve LR; if it persists, check data for corrupt examples (a single garbage example with extreme tokens can do this).
The cheap LR test: run 200–300 steps at 1e-4, 2e-4, and 4e-4 on a subset of data, and pick the one with the lowest validation loss. Three short runs beat one long guess.
Rank controls how many trainable parameters the adapter has (Chapter 4): roughly, capacity scales with r. Too small and the adapter cannot express your task's patterns — validation performance plateaus below what the data supports (underfitting from capacity). Too large and the adapter memorizes training examples instead of generalizing, especially on small datasets (overfitting from capacity), while also slowing training and inflating the adapter file.
Defaults: r=8 or 16 for format/style tasks; r=32–64 for hard domain adaptation. Keep alpha proportional (alpha = 2× r is the common convention, so the alpha/r scaling stays constant).
How to tune: fix everything else, train with r ∈ {8, 16, 32}, and compare validation loss and a task metric (Chapter 8). If r=32 beats r=8 substantially, your task needs capacity — try 64. If all three are similar, keep 8 or 16 (smaller generalizes better and trains faster). Report this ablation; it is a one-paragraph result reviewers respect.
One epoch = one full pass through the training set. Too few epochs and the model hasn't absorbed the patterns (underfitting). Too many and it starts memorizing specific examples — training loss keeps falling while validation loss rises (overfitting, Chapter 9).
Defaults: 2–5 epochs for instruction tuning on thousands of examples. With very small datasets (a few hundred), even 3 epochs can overfit — watch validation loss. With very large datasets (100k+), 1–2 epochs often suffice because each epoch contains so many updates.
The right procedure: train for more epochs than you need (say 5–8) with checkpointing (save every N steps), then pick the checkpoint with the best validation score — this is early stopping done manually. Never pick the final checkpoint by default; the best one is often in the middle. SFTTrainer/Axolotl can automate this with load_best_model_at_end=True plus a metric.
The batch size is how many examples the model processes before each weight update. Larger batches give smoother, more reliable gradient estimates (less noise per step) and better GPU utilization — but use more memory. Smaller batches are noisier, which can actually help escape sharp, overfit minima, but train less efficiently.
The memory trick — gradient accumulation: you can simulate a large batch on a small GPU. Process micro_batch_size examples at a time, accumulate gradients over gradient_accumulation_steps mini-batches, then update once. Effective batch = micro × accumulation × (number of GPUs). Chapter 6 used 2 × 8 = 16. Effective batch sizes of 16–128 are the normal range for QLoRA instruction tuning.
Practical rule: set micro batch to the largest that fits in memory (try 4, then 2, then 1), then set accumulation steps to reach your target effective batch. If you change the effective batch substantially, consider scaling the learning rate mildly (larger batch → slightly larger LR is the textbook guidance, though in the 16–128 range most practitioners keep LR fixed).
paged_adamw_8bit is the QLoRA-era default — AdamW with 8-bit optimizer states (less memory) and paging (spills to CPU RAM gracefully instead of crashing on spikes). Use it unless you have a reason not to.You cannot grid-search everything — each full run costs hours. Use this staged procedure:
Total: roughly 8–10 training runs, most of them short. This is affordable on Colab (spread over days) or a single cloud GPU (hours), and it is exactly the methodology section reviewers want to read.

Figure: Training loss (falling) versus validation loss (falling, then rising) — the signature of overfitting.
Learn to read these two curves at a glance:
Log both curves from step one (TensorBoard or even the printed logs). The number of student projects that trained blind for six hours and discovered a flat line at the end is too large to count. Don't join them.
The staged tuning procedure in this chapter tunes one thing at a time, which is practical — but hyperparameters interact, and knowing the main interactions prevents confusing results:
Learning rate × batch size. Larger batches produce less noisy gradients, which tolerate (and sometimes benefit from) slightly larger learning rates. If you double your effective batch, nudging LR up by ~1.4× (square-root scaling, the common heuristic) is reasonable — but in the 16–128 batch range, most practitioners keep LR fixed and do fine. The dangerous direction is shrinking the batch (e.g., to fit memory) while keeping LR high: noisier gradients + big steps = instability. If you cut batch size 4× for memory reasons, consider halving LR as compensation.
Rank × data size. Small data + high rank is the classic overfitting recipe (Chapter 9): the adapter has capacity to memorize every example. The interaction runs both ways — with 50,000 diverse examples, r=64 may generalize beautifully where it memorized on 1,000. Rule of thumb: scale rank with data, not with ambition. If you increase data 10×, revisiting rank upward is legitimate; if data stays fixed, higher rank mostly buys memorization.
Epochs × learning rate. More epochs at a high LR drives the model further from base behavior (more forgetting, Chapter 9); fewer epochs at a low LR may underfit. The product that matters is roughly "total distance traveled" — LR × steps. This is why the cosine schedule (which decays LR to near zero) pairs naturally with fixed epoch counts: late epochs take tiny steps, so extra epochs cost less in forgetting than they would at constant LR.
Sequence length × batch size (the memory interaction). GPU memory is shared between them: doubling max sequence length roughly doubles activation memory, forcing you to halve the batch. Since most examples are short (Chapter 7's 95th-percentile advice), cutting max length is usually the cheaper side of this trade — you keep batch size (stable gradients) and only truncate the rare long example.
What this means for your tuning procedure: the staged approach works because these interactions are second-order for typical ranges — tuning LR first, then rank, then epochs gets you to a good neighborhood. But when you change something big (10× data, different model size, 4× batch change), re-check the earlier stages rather than assuming they transfer. And in your paper, report the final configuration as a joint choice ("selected by staged search; LR 2e-4 and rank 16 were co-validated") rather than implying each was optimized in isolation — reviewers know about interactions, and acknowledging them reads as competence.
The cheat sheet gives starting points; this tree handles the cases where they don't work. Start at the top with your symptom:
Training loss won't go below ~2.5 and outputs look untrained. → Is the loss exactly flat, or noisy-flat? Exactly flat: check label masking (is the model training on empty responses? print one tokenized batch with labels). Noisy-flat: LR too small — try 5e-4. Still flat: your data may not be in the format the template expects — decode 3 training examples fully and read them.
Training loss falls beautifully but validation is terrible from the start. → Data distribution mismatch between splits (check: are val examples systematically different — longer, harder, different template?). Or label leakage in reverse: the model learned a spurious training-set pattern. Stratify your split by the important dimensions (topic, length, difficulty) instead of pure random.
Everything worked on the subset but the full run diverged. → The subset wasn't representative (too easy, too clean). Or: you scaled epochs without scaling warmup — warmup steps should scale with total steps; 30 warmup steps for a 3,000-step run is proportionally nothing. Set warmup to ~3–5% of total steps.
Results vary wildly between seeds (±5 points). → Your test set is too small or your task too noisy for the claimed precision. Either enlarge the test set, or report the variance honestly and shrink your claims. Seed variance is a finding — it tells you the task is unstable, which is worth one paragraph.
The model got worse than the base model. → Almost always one of: LR far too high (check for the loss spike signature), training on the wrong objective (e.g., full next-token loss including instructions, teaching the model to generate instructions instead of following them), or catastrophic data corruption (duplicated examples, wrong language). Roll back to the walkthrough baseline and diff every setting.
Work this tree before asking for help or buying more compute — nine times out of ten, the answer is in it.
For your research: Your paper should contain a hyperparameters table (the full config) and ideally one ablation figure — e.g., validation performance vs. LoRA rank, or vs. data size. These are cheap to produce during the tuning procedure above and they transform "we fine-tuned a model" into "we systematically studied fine-tuning for this task." The difference is the difference between a workshop paper and a conference paper.
Training produces a model. Evaluation tells you whether the model is actually better — and "better" is a claim that requires evidence, not enthusiasm. This chapter is about building that evidence: what to measure, what to compare against, and how to avoid the self-deception that fine-tuning makes so easy. If Chapter 3's dataset is the steering wheel, evaluation is the dashboard. Driving blind is not brave; it is just blind.
A fine-tuned model's score, alone, means nothing. "Our model scores 82%" is empty until you know what the base model scored (maybe 81%), what a good prompt scored (maybe 84%), and what chance scored (maybe 50%). Every evaluation needs baselines, and for fine-tuning papers the expected set is:
Report all of them in one table. The pattern reviewers look for — and the pattern that makes a paper convincing — is: base < few-shot < your fine-tune, with ablations showing the gain survives when you vary the details. If few-shot prompting matches your fine-tune, say so honestly; that is itself a publishable finding ("for this task, prompting suffices"), and claiming otherwise will not survive review.
Match the metric to what you actually care about:
One metric is never enough. Report a small panel: your task metric, a format-compliance check, and one broad benchmark for regressions. Three numbers tell a story; one number tells a slogan.
From Chapter 3: your test set is locked until the final evaluation. Here is why this matters so much in fine-tuning specifically. During development you will look at validation results dozens of times, adjusting data, prompts, and hyperparameters. Every adjustment informed by validation leaks a little information about validation into your model. By the end, your validation score is optimistic — you have, in effect, trained on it indirectly. The locked test set is the only unbiased estimate you have. Evaluate it once, on the final model, and report whatever it says — even if it is worse than validation. Reporting the worse number honestly is what separates research from marketing.
Additional hygiene:
Numbers hide as much as they reveal. For every major experiment, read 50–100 outputs by hand — from the base model and the fine-tuned model side by side — and categorize the differences. You will discover things no metric shows: the fine-tuned model fixed formatting but introduced a new hedging tic; it answers correctly but dropped the citation style you wanted; it improved on common cases but got worse on rare ones.
Build a small error taxonomy: the 4–6 most common failure modes, with example counts before and after fine-tuning. "Hallucinated clause numbers: 31/100 → 4/100; wrong output language: 12/100 → 2/100; refused valid requests: 0/100 → 9/100" — this table is worth more than three paragraphs of prose, and it directly suggests your next experiment (the refusal spike means your data over-represents refusals — fix the data mix).
import json, re
# Load locked test set (untouched until now)
test = [json.loads(l) for l in open("data/test.json")]
def format_ok(response):
# Example task-grounded check: response must be valid JSON with required keys
try:
obj = json.loads(response)
return "answer" in obj and "justification" in obj
except Exception:
return False
def evaluate(model, tokenizer):
correct, wellformed, total = 0, 0, 0
for ex in test:
prompt = ex["instruction"] # adapt to your template
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=200, do_sample=False)
resp = tokenizer.decode(out[0], skip_special_tokens=True)
total += 1
if format_ok(resp):
wellformed += 1
# task-specific correctness check goes here
return {"n": total, "wellformed_rate": wellformed/total}
# Run for: base zero-shot, base few-shot, fine-tuned — same script, same test set
for name, model in [("base", base_model), ("finetuned", ft_model)]:
print(name, evaluate(model, tokenizer))
The discipline that matters: the same script, the same test set, the same decoding settings (temperature, max tokens) for every model compared. Changing decoding between baseline and fine-tuned runs invalidates the comparison silently.
Your paper's results section centers on one table. Columns: model/baseline. Rows: metrics. Something like:
| Model | Task accuracy | Format compliance | MMLU (regression) |
|---|---|---|---|
| Base, zero-shot | 54.2 ± 1.1 | 61% | 68.4 |
| Base, few-shot (5 ex) | 71.8 ± 0.9 | 83% | 68.4 |
| Fine-tuned (ours) | 78.5 ± 0.7 | 97% | 67.9 |
This fictional-but-typical table tells the whole story at a glance: fine-tuning beat few-shot prompting on the task (+6.7), massively improved format reliability, and barely dented general knowledge (−0.5 MMLU, the forgetting check from Chapter 9). Write the paragraph around the table, not instead of it.
Using a strong model to evaluate your fine-tuned model's outputs is now standard practice — it's the only scalable option for open-ended generation. But naive LLM judging is biased in documented ways, and reviewers know them. Here's how to do it defensibly:
The known biases (and mitigations): - Position bias: judges favor the first-presented answer. Mitigation: present pairs in both orders and average, or use absolute scoring (rate each output 1–5 alone) instead of pairwise comparison. - Verbosity bias: judges reward longer answers. Mitigation: instruct the judge to penalize unnecessary length explicitly in the rubric, or control for length statistically. - Self-preference: a judge from the same model family as your fine-tuned model favors its sibling's style. Mitigation: use a judge from a different family than your base model, and disclose which judge you used. - Rubric vagueness: "rate quality 1–10" produces noise. Mitigation: write a concrete rubric with anchored descriptions ("5 = correct answer, correct format, cites clause; 3 = correct answer, wrong format; 1 = wrong answer").
The validation step you must not skip: have humans rate 100–200 of the same outputs and measure agreement between human and LLM judge (correlation for scores, Cohen's kappa for categories). Report this number. "Our LLM judge agrees with human raters at r=0.81 on a 200-example validation set" transforms the judge from a black box into a calibrated instrument. Without it, a skeptical reviewer can dismiss all your judged results in one sentence.
Blinding: the judge must not know which output came from which model — strip model identifiers, randomize order, and use identical formatting for both candidates. This sounds obvious; it's skipped surprisingly often.
Cost control: judging 10,000 outputs with a frontier API model gets expensive. Judge a random, pre-registered subset (e.g., 500–1,000) with the strong judge, use programmatic checks (Chapter 8's format verification) on the full set, and human-rate 100–200. Three tiers, each doing what it's best at: programs for scale, LLM judges for nuance at moderate scale, humans for ground truth.
"Mean ± std across seeds" was Chapter 8's advice; here's how to actually compute and interpret it without a statistics degree:
How many seeds? Three is the practical minimum (42, 43, 44 — or any three you fix in advance). Five is better but costs 5× the final-run compute. Never report the best of N seeds as "the result" — that's selection bias, and reviewers check for it by asking whether you reported the max or the mean.
Reading the interval. Suppose your fine-tuned model scores 78.5 ± 0.7 and few-shot scores 71.8 ± 0.9. The intervals don't overlap — the improvement is real beyond seed noise. But 78.5 ± 0.7 vs. 78.1 ± 0.8 do overlap substantially — that "gain" is indistinguishable from luck. Don't claim it; report both numbers and move on. Overlapping intervals don't prove no difference, only that your experiment can't resolve one — a paired test (below) is more sensitive.
A simple paired test for classification. When comparing two models on the same test set, McNemar's test asks: on examples where the models disagree, does one win significantly more often? It's a one-function call in most stats libraries and far more appropriate than eyeballing percentages. For generation metrics, a paired bootstrap (resample the test set with replacement 1,000 times, compute the score difference each time, check whether 95% of differences favor your model) achieves the same thing with no distributional assumptions. You don't need to derive these — just use them and report "paired bootstrap, p < 0.05."
What "significant" doesn't mean. Statistical significance is not practical significance. A 0.4% gain with p < 0.01 on 50,000 test examples is real but probably irrelevant; a 6% gain with p = 0.08 on 200 examples is suggestive but unresolved — collect more test data rather than claiming victory. Report effect size (the actual point difference) alongside the p-value, always.
The pre-registration payoff. Because you wrote the eval script and empty table before training (Chapter 8's discipline), you can't be accused of choosing the metric that happened to look good — the classic "researcher degrees of freedom" problem. Mention this explicitly in the paper: "Metrics and baselines were fixed before training; see our experiment plan." It's a single sentence that substantially raises reviewer trust.
For your research: Design your evaluation before you train — write the eval script and the empty results table first. This "evaluation-driven development" prevents the most common failure mode in student fine-tuning projects: training for two weeks and then discovering you have no clean way to say whether it worked. Your empty table with its baselines and metrics is also the skeleton of your paper's experiments section. Fill it in as results arrive.
Fine-tuning has two characteristic failure modes, and they are mirror images. Overfitting is learning your training data too well — the model memorizes instead of generalizing. Catastrophic forgetting is learning your training data at the expense of everything else — the model gains your task and loses its general abilities. Both are detectable, both are fixable, and both must be checked in any serious fine-tuning project. This chapter gives you the detection toolkit and the standard remedies.
Overfitting means the model performs well on training examples but poorly on new ones. In fine-tuning it is especially sneaky, because your training examples look like success — the model produces beautiful outputs on exactly the cases you show it. The illusion breaks on the first unfamiliar input.
Why fine-tuning overfits easily: your dataset is small relative to the model's capacity. An 8B model with a 42M-parameter adapter can memorize thousands of examples verbatim without breaking a sweat. Memorization is the path of least resistance for the optimizer — reproducing exact training outputs reduces loss faster than learning the general pattern. The model takes the shortcut unless your data and settings force it onto the longer road.
Detection — the loss curves (primary signal): as Chapter 7's figure showed, overfitting's signature is unmistakable: training loss keeps falling while validation loss bottoms out and starts rising. The epoch/checkpoint where validation loss is lowest is approximately where generalization peaked — everything after is memorization. This is why you save checkpoints and pick by validation, not by recency.
Detection — the memorization probe: take 20 training examples and 20 fresh examples of the same type. If the model is dramatically better on the training 20 (near-perfect) than the fresh 20, it memorized. A generalizing model shows a smaller gap. For text, an even simpler probe: feed the model the first half of a training example and see if it reproduces the second half verbatim — verbatim reproduction is memorization, not learning.
Detection — the paraphrase test: paraphrase your test inputs (same meaning, different wording). A model that learned the task handles paraphrases; a model that memorized patterns collapses. Large paraphrase-sensitivity gaps are a red flag reviewers know to ask about.
Fixes, in order of effectiveness:
lora_dropout to 0.1, add weight decay, or use data augmentation (paraphrase your training examples to multiply diversity cheaply).Catastrophic forgetting is the mirror problem: the model learns your task but loses capabilities it had before — general knowledge, other languages, reasoning skills, instruction-following breadth. The name comes from Goodfellow et al.'s 2013 study showing neural networks abruptly lose old tasks when trained on new ones. In LLMs the effect is usually gradual rather than abrupt, but it is real and measurable: fine-tune a model exclusively on medical Q&A for enough epochs and its MMLU score, its poetry, and its low-resource-language ability all degrade.
Why it happens: Chapter 2's trade intuition — every update that helps your examples can hurt others, because capabilities share the same weights. Your narrow data pulls the weights toward a specialist configuration; the generalist configuration erodes. The narrower your data and the longer you train, the worse it gets.
Detection — regression benchmarks: before fine-tuning, score the base model on broad benchmarks (MMLU for knowledge across subjects, plus whatever covers your model's other important abilities — multilingual tests if relevant, code tests if relevant). Score the fine-tuned model on the same benchmarks. A drop of 1–2 points is normal noise and acceptable; a drop of 5+ points is forgetting worth addressing. This is why Chapter 8's results table includes a regression column — it is not decoration, it is the forgetting detector.
Detection — capability probes: hand-write 20–30 prompts testing abilities your fine-tuning shouldn't have touched: general knowledge questions, a different language, simple reasoning, refusal of disallowed content (safety behavior can degrade too — check it). Compare base vs. fine-tuned qualitatively. Automated benchmarks miss subtle degradations; reading outputs catches them.
Fixes, in order of effectiveness:
Notice the uncomfortable symmetry: the cures overlap. More epochs → more overfitting and more forgetting. Higher rank → more capacity to memorize and more disturbance of old weights. Lower LR → less memorization and less forgetting, but slower learning. You are always balancing three goals: learn the new task (fit), generalize to new inputs (not overfit), retain old abilities (not forget).
The practical resolution is the validation panel: don't watch one curve, watch three — task validation score (want: high), held-out generalization probes like paraphrases (want: stable), regression benchmark (want: not much lower than base). Pick the checkpoint and hyperparameters that satisfy all three. A model that scores 80% on your task but lost 8 MMLU points is not "better" than one scoring 78% with no regression — it is differently broken, and the second is usually the one you want.
Imagine your curves and scores look like this after training:
Diagnosis: both failure modes. Rising validation loss + 18-point paraphrase gap = overfitting. 5.3-point MMLU drop = forgetting. Prescription: roll back to the epoch-3 checkpoint (fixes the worst of both), then rerun with: 10% replay data mixed in, r reduced 32→16, dropout 0.05→0.1, and stop at the new validation minimum. Re-measure all three panels. This diagnostic loop — measure three panels, adjust, re-measure — is the core skill of this chapter, and it is worth practicing deliberately on toy runs before your thesis depends on it.
Chapter 9 covered forgetting of capabilities — knowledge, languages, reasoning. There's a second kind of forgetting with sharper consequences: safety degradation. Instruction-tuned base models carry safety behaviors from their alignment training (refusing disallowed requests, hedging on medical/legal advice, avoiding biased outputs). Fine-tuning — even on completely benign data — can erode these.
Why it happens is straightforward: your data never demonstrates refusal behavior, so the gradients never reinforce it, while everything else shifts. The model doesn't "decide" to become unsafe; the safety-relevant weight configurations simply drift as the model optimizes for your task. Research on this is active, but the practical finding is consistent: fine-tuning on narrow benign data measurably increases compliance with harmful requests in several published studies, and the effect grows with epochs and learning rate.
What to check (cheap, non-negotiable if you publish or deploy): 1. Assemble 30–50 prompts spanning the safety behaviors you care about: direct disallowed requests, jailbreak-style framings ("pretend you're…"), and domain-adjacent edge cases (for the medical case study: "should I stop taking my medication?"). 2. Score base vs. fine-tuned responses blindly (the rater doesn't know which model produced which). 3. Report the comparison. A small change is normal; a large one means your training undid alignment — add refusal/hedging examples to your data (the "replay" principle applied to safety) and re-measure.
The data fix: include 2–5% safety-relevant examples in training — refusals of disallowed requests in your domain, proper hedging ("I'm not a doctor; consult one") on advice questions. These don't teach the model new safety behavior; they keep the existing behavior exercised while everything else changes. Think of it as replay data for alignment.
For researchers, this is opportunity, not just obligation. Safety degradation from benign fine-tuning is an active research area with open questions: which training choices (rank? epochs? data mix?) degrade safety most? How little safety data suffices to preserve it? If your thesis involves fine-tuning in a sensitive domain (medicine, law, finance), a careful safety-degradation analysis with ablations is a genuine contribution — and reviewers in those domains increasingly expect at least the basic check.
So your run finished and the three panels (Chapter 9) all look wrong: task score mediocre, paraphrase gap large, MMLU down 6 points. Don't despair and don't start over randomly — diagnose systematically:
Step 1: Check the data first (60% of disasters). Before touching hyperparameters, re-grade 100 training examples (Chapter 3's ritual). In most failed student runs, the root cause is data: mislabeled examples, template bugs (labels including the instruction text), or train/test contamination discovered late. No hyperparameter fixes bad data.
Step 2: Find the best checkpoint, not the last one. Plot validation loss across all saved checkpoints. If the minimum was at 40% of training, your "final model" was overtrained — evaluate the best checkpoint on the test set. You may already have a good model sitting in your outputs directory.
Step 3: Isolate one variable. Change exactly one thing and rerun on a data subset: halve the learning rate, or halve the rank, or add 10% replay data. If the subset run improves, apply to the full run. Changing three things at once teaches you nothing — you'll never know which fix worked, and your paper can't report it.
Step 4: Shrink the problem. If the full task still resists, fine-tune on a narrower slice (one question type, one document type) and check whether that works. Success on the slice proves the pipeline is sound and the problem is task difficulty or data coverage — actionable diagnoses. Failure on the slice too points back to pipeline bugs.
Step 5: Know when to stop. Some tasks genuinely don't yield to the resources available — the base model may simply lack the prerequisite capability (Chapter 11's low-resource warning). Stopping is a decision, not a failure, if it's evidence-based: "three rank settings, two data sizes, and replay ablations all plateau at ~62%; the base model's probe scored 41%, so fine-tuning added real value but the task needs a stronger base model or more data than available." That's a defensible thesis conclusion. Endless tuning without a diagnosis is not persistence — it's the sunk-cost fallacy with GPUs.
Keep a lab notebook (a simple dated markdown file) recording every run: config, data version, result, and your one-sentence interpretation. When you find the fix six runs later, the notebook tells you what changed — and it becomes the methods section's tuning narrative.
For your research: Forgetting analysis is an underused source of paper content. Most fine-tuning papers report only task gains; a paper that additionally reports what was retained and what was lost — with the replay-data ablation showing the trade-off curve — answers a question every practitioner has and few papers address. If your domain is specialized (law, medicine, low-resource languages), the forgetting profile across capabilities is itself a novel empirical contribution.
Fine-tuning used to be priced for institutions. QLoRA and free-tier GPUs repriced it for students — but "cheap" is not "free," and the real budget includes data labor, failed runs, and your time. This chapter gives you honest numbers, a budgeting method, and the tricks that stretch a student budget furthest. No invented price lists: GPU rental prices move constantly, so we give you the estimation method plus the magnitudes to sanity-check against, and tell you exactly where to look up current prices.
1. Data cost (usually the largest). Human-written or human-verified examples cost human hours. Writing 2,000 quality instruction pairs might take one person 3–6 weeks part-time. Even "free" distillation needs verification labor. When students say "fine-tuning is free on Colab," they are ignoring the 100+ hours of data work. Budget it explicitly: hours × your (opportunity) cost. This is also why Chapter 3's quality-over-quantity finding is a budget finding — 1,000 excellent examples cost a fifth of 5,000 mediocre ones and often perform better.
2. Compute cost. The GPU hours for training runs, including failed runs, ablations, and final multi-seed runs. A QLoRA run on 5,000 examples for 3 epochs on a T4 takes roughly 3–8 hours; the full experimental program from Chapter 7 (8–10 runs) might total 30–80 GPU hours. On free Colab, that's wall-clock days with session limits; on paid cloud, it's tens of dollars, not hundreds — for 7–8B models. Costs scale roughly with model size and data size: a 70B run is an order of magnitude more.
3. Storage and misc. Model downloads (tens of GB), checkpoints (each full-precision checkpoint of an 8B model is ~16 GB; adapters are ~100 MB — another reason to save adapters, not merged models, during experimentation), dataset storage. Mostly negligible in cost, occasionally painful in Colab disk limits.
4. Your time. Debugging environment issues, reading loss curves, doing qualitative analysis. Budget 2–3× the pure training time for the human loop around it. This is normal and not a sign you're slow.
Rule of thumb: do all development, debugging, and small ablations on free resources; pay (a little) only for the final multi-seed runs if free tiers can't finish them reliably.
You can estimate training time with a quick calibration run instead of guessing:
Example: 5,000 examples, 3 epochs, effective batch 16 → ~940 steps. At 15 sec/step on a T4 → ~4 hours. Now multiply by your experimental program: coarse LR search (3 short runs on 20% data ≈ 3 × 0.8h), rank ablation (3 runs ≈ 3 × 4h), final seeds (3 runs ≈ 3 × 4h) → roughly 27 GPU hours. On Colab Pro that's a few days of sessions; on a cloud A100 (~$1–2/hr at current market rates — verify on the provider's pricing page) that's roughly $30–60. The method matters more than the numbers: calibrate, multiply by your run count, add 50% buffer for failures.
Cost estimation table (magnitudes, verify current prices):
| Setup | Hardware | ~Cost | Good for |
|---|---|---|---|
| Colab free | T4 16 GB | $0 (time-limited) | Walkthroughs, debugging, small ablations |
| Colab Pro | T4/V100-ish | ~$10/month | Serious student projects, final runs |
| Kaggle | T4/P100 | $0 (weekly quota) | Overflow/parallel experiments |
| Cloud spot/preemptible | A100 40–80 GB | ~$0.50–1.50/hr (verify) | Large runs, 70B models, fast ablations |
| Cloud on-demand | A100 80 GB | ~$1.50–4/hr (verify) | Deadline-driven runs |
| University cluster | Varies | $0 (ask!) | Everything, if available |
Check current prices on the provider's own pricing page before budgeting — they change quarterly. Never trust a blog post's price table, including (eventually) this book's.
save_total_limit to cap disk. A crashed 6-hour run with checkpoints is a 10-minute resume; without them it's a 6-hour loss.packing=True in SFTTrainer, sample_packing: true in Axolotl) once the pipeline works — it can nearly double training throughput on short examples by eliminating padding waste.Fill this in at project start and revisit monthly:
A typical MS-thesis fine-tuning project lands around 30–80 GPU hours and 100–200 human hours, most of it data and evaluation. If your plan says 500 GPU hours, either narrow the task or find cluster access before you start — discovering this in month four is how theses stall.
The estimation recipe in this chapter assumes runs succeed. They won't — at least not all of them. Here's an honest accounting of where GPU hours actually go in a student project, and how to budget for the waste:
Environment debugging (5–15 hours). CUDA version mismatches, bitsandbytes compilation issues, tokenizer quirks, gated-model access approvals. This is front-loaded and mostly one-time — which is exactly why Chapter 6's walkthrough exists: it concentrates the debugging into a single cheap session before the real work starts.
Data pipeline iterations (10–30 hours). You discover mid-project that 15% of examples are malformed, or the template needs changing, which invalidates earlier runs. Mitigation: the Chapter 3 pilot (50 examples end-to-end) and the data-grading ritual catch most of this before any serious training.
Aborted training runs (20–40% overhead). Wrong LR discovered at step 300, OOM at step 50, eval revealing a template bug at epoch 1. This is normal and healthy — aborting early is the skill. Budget it as a multiplier: multiply your "successful runs" estimate by 1.5×, not by wishful thinking.
The "one more ablation" (10–20 hours). Reviewers (or your advisor) will ask for exactly one experiment you didn't run. This is why the budget template includes a buffer — and why keeping your pipeline runnable (Chapter 5's version control) matters: the cost of one more ablation is hours if the pipeline is intact, days if you've let it rot.
Putting it together — a realistic ledger for a thesis project: | Phase | GPU hours | Human hours | |---|---|---| | Walkthrough + env setup | 5 | 15 | | Data collection + verification | 0 | 80–120 | | Pipeline debugging + pilots | 10 | 20 | | Coarse tuning (short runs) | 8 | 10 | | Main ablations | 20 | 15 | | Failed/aborted runs (buffer) | 15 | 10 | | Final multi-seed runs | 15 | 5 | | Evaluation + analysis | 5 | 30 | | Total | ~78 | ~200 |
The GPU column is the smaller number. Internalize that: in fine-tuning projects, compute is rarely the binding constraint for students — human time is. Every hour you spend on data quality and evaluation design pays back more than an hour of extra training. Spend the budget accordingly.
Free Colab is the most common student training environment, and it has sharp edges. Here's how practitioners work within them:
Session timeouts. Colab disconnects idle sessions (~90 minutes idle) and caps total GPU time per day. Never leave a run unattended without checkpointing — save_steps=100 means a disconnection costs you at most 100 steps. For runs longer than a few hours, prefer Colab Pro or split the run: train 2 epochs, save, resume later with resume_from_checkpoint (Chapter 6).
The disk limit. Colab gives roughly 100+ GB of disk, but model downloads plus checkpoints plus datasets add up. Defenses: download the base model once per session (it's cached in ~/.cache/huggingface); keep save_total_limit=2 so old checkpoints delete themselves; save adapters (100 MB) not merged models (16 GB) during experimentation; clean the pip cache (pip cache purge) if space gets tight.
The RAM limit. Colab's CPU RAM (~12 GB free tier) matters when loading datasets or tokenizing large files. If preprocessing OOMs the CPU (not the GPU), process the dataset in streaming mode (datasets.load_dataset(..., streaming=True)) or preprocess once, save to disk, and load the processed version.
Getting throttled. Heavy free-tier use triggers Colab's usage limits ("you've used your GPU quota"). Tactics: do CPU work (data prep, analysis, writing) in non-GPU runtimes; batch your GPU needs into focused sessions; use Kaggle's separate weekly quota as overflow; and treat the $10 Colab Pro as the obvious upgrade the moment throttling costs you more than an hour a week — your time is worth more than $10/month.
Reproducibility across sessions. Colab sessions are ephemeral — the VM resets. Everything durable lives in three places: your git repository (code, configs), your Drive or Hugging Face Hub (adapters, datasets), and your notebook with pinned installs. At the start of each session, the ritual is: mount Drive (or pull from Hub), pip install -r requirements.txt, verify GPU, resume. Write this ritual as the first cells of your notebook so future-you doesn't improvise it each time.
A final budgeting judgment call students face repeatedly: is this slowdown worth paying to fix? A decision heuristic:
Pay (small amounts) when the bottleneck is waiting, not thinking. If your pipeline is debugged, your data is ready, and you're losing days to Colab throttling or queue times before a deadline — $10–50 for Colab Pro or a few cloud GPU hours is obviously worth it. Compute spending should buy calendar time, not substitute for experimental design.
Don't pay to avoid thinking. If runs keep failing for unclear reasons, a bigger GPU just fails faster and more expensively. OOM errors, flat losses, and bad evals are design problems; solve them on the free tier with small runs. Only scale spending after small runs succeed reliably.
The monthly budget rule. Set a fixed monthly compute budget you're comfortable with ($0, $10, $50 — whatever fits) and design the experimental program to fit it, rather than spending reactively. The Chapter 10 template plus the free-tier tactics make a $10/month budget cover a full thesis project comfortably: free Colab/Kaggle for development and ablations, Pro for reliability during final runs. Students who blow past $200 usually did so by launching full-scale runs before the pipeline worked — the most expensive mistake in this book, and entirely preventable.
For your research: Funders, advisors, and thesis committees respond to budgets, not vibes. A one-page compute plan — with the calibration method, the run count, and the dollar figure — attached to your proposal signals that you've done this before (even if this book is your first time). It also protects you: when results need one more ablation, the buffer you budgeted is the reason you can afford it.
Everything so far has been general. This chapter is personal: how to take your research problem — the one your thesis or paper depends on — and turn it into a well-designed fine-tuning study. We work through the design process as a sequence of decisions, illustrated with three realistic case studies drawn from the kinds of problems MS/PhD students actually bring to fine-tuning. Adapt the pattern, not the details.
Step 1: State the behavioral gap in one sentence. Not "we want a better model for X" but "the base model does X, we need it to do Y." Examples: "The base model answers Urdu medical questions in English medical jargon; we need plain-Urdu answers with citations." "The base model writes SQL that references non-existent columns; we need schema-faithful SQL." If you cannot state the gap crisply, you are not ready to fine-tune — go back to Chapter 1's framework and run the prompting/RAG checks first.
Step 2: Define "done" before you start. Write down the metric, the target, and the baseline you must beat. "Task accuracy ≥ 80% on the locked test set, beating few-shot prompting (currently 71%), with MMLU regression under 2 points." This is your contract with yourself. Without it, every result looks publishable and none of them are.
Step 3: Inventory your data honestly. How many examples can you realistically produce or verify? Who writes them? What is the timeline? Map this against Chapter 3's brackets: if your gap needs 10,000 examples and you can produce 800, redesign the task (narrower scope, better base model, RAG assist) rather than training on 800 and hoping.
Step 4: Choose the method and justify it. QLoRA at r=16 is the default; deviate deliberately. Full fine-tuning because your cluster has A100s? Higher rank because the domain shift is large? Write the justification down — it becomes your methods paragraph.
Step 5: Plan the ablations that make it research. A fine-tuned model is an engineering result; a fine-tuning study is research. The difference is the ablations: data-size curves, rank curves, with/without replay, prompting vs. fine-tuning vs. combined. Plan 3–5 ablations that each answer a question someone in your field would ask. Chapter 12 shows how these become figures.
Step 6: Pre-register the evaluation. Write the eval script and empty results table now (Chapter 8). Decide the statistical bar for "improvement" in advance.
The problem. A public-health MS student wants a model that answers common medical questions in plain Urdu, citing which answers need a doctor's visit. The base model (a strong multilingual 8B) answers in English-heavy jargon and never hedges.
Behavioral gap: language register + safety hedging — both behavioral, not factual. RAG over a medical FAQ could supply facts, but the style problem (plain Urdu, consistent "see a doctor when…" framing) needs fine-tuning.
Data plan: 3,000 real questions collected from clinic workers, answers drafted by distillation from a frontier model, then verified and rewritten into plain Urdu by two medical students. Budget: ~6 weeks. Format: instruction (question) → response (plain-Urdu answer + triage line). 10% general Urdu instruction data mixed in as replay.
Method: QLoRA, r=32 (medical domain shift is substantial), base model with the best Urdu among candidates (compare 2–3 base models on a 200-question probe before committing — base model selection is an underrated decision).
Evaluation: locked test of 400 questions; metrics: medical accuracy (rated by the two medical students, blinded, base vs. fine-tuned), plain-language score (rubric), triage-line presence (programmatic check), MMLU + an Urdu benchmark for regression. Baselines: base zero-shot, base few-shot, RAG-only, fine-tuned + RAG.
Ablations: data size {500, 1500, 3000}, rank {16, 32}, with/without replay. The data-size curve alone answers "how much verified data does this task need?" — a citable result for low-resource medical NLP.
The problem. A CS MS student building a natural-language interface to their university's student-records database. The schema is fixed (47 tables), the SQL dialect is fixed, and the system will serve thousands of queries daily. Base model prompting gets 68% execution accuracy; errors are mostly hallucinated column names.
Behavioral gap: internalized schema knowledge — the model must know the 47 tables cold. The schema fits in the prompt, but at thousands of queries/day, stuffing it into every prompt is expensive, and the model still slips. This is Chapter 1's canonical fine-tuning case: narrow task, high volume, behavior to internalize.
Data plan: synthetic generation with verification — programmatically generate thousands of (question, SQL) pairs from the schema (templates × tables × columns), execute every SQL against a test database, keep only pairs that execute and return sensible results. 15,000 examples, near-zero human cost, perfect verification. This is the rare case where synthetic data is high quality — because execution is an objective filter.
Method: QLoRA r=16 (narrow task, small output space), 3 epochs. The interesting research question isn't whether it works — it's the data efficiency: how few synthetic examples suffice?
Evaluation: execution accuracy on a locked set of 1,000 human-written questions (human-written, because the test must not share the synthetic generator's biases — critical design point). Baselines: base + schema in prompt, base few-shot, fine-tuned. Latency and cost-per-query comparison for the deployment argument.
Ablations: synthetic data size {1k, 5k, 15k}, and the key scientific question: does training on synthetic template data generalize to human phrasing? (The train/test distribution gap is the paper.)
The problem. A linguistics PhD student wants an instruction-following model for Sindhi, which the base model barely speaks. Prompting in Sindhi mostly fails; the capability gap is large.
Behavioral gap + knowledge gap: the model lacks both Sindhi fluency (knowledge-ish, from pre-training) and instruction-following in Sindhi (behavioral). Honest assessment required: fine-tuning can teach instruction-following in Sindhi far better than it can teach Sindhi itself — if the base model truly barely saw Sindhi in pre-training, expectations must be modest, and continued pre-training on Sindhi text (Chapter 4's full-fine-tuning case) may be needed first.
Data plan: two-stage. Stage A: continued pre-training on ~1–2 GB of Sindhi web/books text (full fine-tuning, the legitimate exception). Stage B: instruction tuning on 2,000–5,000 Sindhi instruction pairs, created by translating and natively verifying (translation alone produces "translationese" — have native speakers rewrite, not just check).
Method: Stage A full fine-tuning if hardware allows, else high-rank QLoRA (r=64); Stage B QLoRA r=32.
Evaluation: Sindhi instruction-following benchmark (built by the student — itself a contribution), native-speaker ratings, regression on English benchmarks (expect some forgetting; report it). Baselines: base model, Stage A only, Stage B only, both.
Ablations: the stage ablation (A only vs B only vs A+B) is the paper's core figure — it separates "teaching the language" from "teaching instruction-following," exactly the manners-vs-knowledge framing from Chapter 2.
If your fine-tuning work becomes a thesis chapter (rather than a standalone paper), the structure differs slightly — a thesis chapter must also teach and justify, not just report. Here's a structure that has worked for MS theses:
1. Problem and motivation (2–3 pages). The behavioral gap in plain language, the real-world stakeholder (clinic workers, students querying the database), and why existing approaches fall short. Include the Chapter 1 decision memo's numbers — the prompting/RAG baselines that motivated fine-tuning.
2. Background (3–5 pages). Fine-tuning concepts (Chapter 2), LoRA/QLoRA (Chapter 4), and the domain background your examiner needs. Examiners are often not LLM specialists — write this section for a smart non-specialist, defining every term. This section also demonstrates your understanding, which is part of what's being examined.
3. Data (2–4 pages). Collection, guidelines, filtering, statistics, examples (show 3–5 real examples verbatim), license/ethics. Include the quality-grading results — examiners love seeing that you measured your own data quality.
4. Method (2–3 pages). The full config table (Chapter 12), the justification for each choice, the experimental procedure (staged tuning), and compute details. A thesis allows more detail than a paper — include the dead ends briefly ("r=64 overfit; we retained r=16") since they show scientific process.
5. Experiments and results (4–6 pages). Main table, ablations as figures, error taxonomy, forgetting/safety analysis. Every figure needs a caption that states the takeaway, not just describes the axes.
6. Discussion and limitations (1–2 pages). What the results mean, what you'd do with more time/data, honest limitations. Examiners probe here in the defense — writing it yourself first means you've already answered their hardest questions.
7. Conclusion (1 page). The one-paragraph version of the whole chapter: gap, method, result, significance.
A practical note: write sections 1–3 before the main training runs (they don't depend on results), and draft the empty tables/figures for section 5 in advance (Chapter 8's pre-registration). Thesis writing then becomes filling in numbers rather than composing from scratch under deadline pressure — the difference between a calm final month and a panicked one.
Chapter 11's case studies describe full projects. But what if you have four weeks, not four months? Here's the smallest study that still counts as research rather than a demo:
Week 1: Probe and data. Run the base-model probe (200 examples, zero-shot + few-shot). Collect or generate 800–1,500 examples with verification. Write the experiment plan (gap sentence, done-criteria, empty results table).
Week 2: Pipeline. Build the Chapter 3 data pipeline, run the Chapter 6 walkthrough on your data, debug until a 1-epoch run completes cleanly.
Week 3: Tune. Coarse LR check (3 short runs on 30% data), rank ablation {8, 16} on full data, then the full run with checkpoint selection. Two seeds if time permits.
Week 4: Evaluate and write. Locked test evaluation, error taxonomy from 50 hand-read outputs, one ablation figure (data size or rank), and the paper/thesis draft following Chapter 12's structure.
What makes this "viable research" rather than a demo: the baselines (few-shot could win), the locked test set, the ablation (even one), and the honest limitations. A four-week study with these elements is publishable at a workshop; without them, four months of training isn't. Scope is not the enemy of rigor — vagueness is.
A practical note most technical books skip: your advisor's role in a fine-tuning project, and how to use their time well.
The two-page experiment plan (Chapter 11) is your primary collaboration tool. Advisors can't debug your CUDA errors, but they can destroy a flawed experimental design in twenty minutes — which is exactly what you want, before you spend two months executing it. Bring the plan early, when it's cheap to change.
What to ask your advisor: Is the research question well-posed? Are the baselines the right ones? Is the evaluation convincing to someone in our field? Is the scope appropriate for the timeline? These are judgment questions where experience dominates.
What not to ask: hyperparameter values, library debugging, cloud pricing. Those are your job (this book is your reference), and bringing them to meetings wastes the scarcest resource in the project — senior judgment.
The monthly check-in format that works: one page with (a) what you ran since last time, (b) the key numbers in the pre-registered table format, (c) what you concluded, (d) what you'll run next and why. Advisors who see this format regularly can spot design drift ("why did the baseline change?") and keep the project honest. It also means your thesis practically writes itself — twelve such pages are twelve sections of the results chapter.
Finally, if your advisor isn't an LLM specialist, this book's plain-language explanations are your translation layer. "We're teaching the model our format using 3,000 verified examples; the base model scores 62%, few-shot prompting 72%, and we're targeting 80% with no loss on general benchmarks" — that sentence, updated monthly, keeps any advisor oriented regardless of specialty.
For your research: Turn Steps 1–6 above into a two-page "experiment plan" document and review it with your advisor before collecting data. It should contain: the one-sentence gap, the done-criteria, the data inventory with timeline, the method justification, the 3–5 planned ablations, and the empty results table. Advisors catch design flaws in twenty minutes that cost students two months. This document also becomes your paper's introduction-through-methods skeleton.
You ran the experiments, the numbers are good, and now you must write them up so that skeptical experts believe you. Reviewers of fine-tuning papers have seen every shortcut — cherry-picked checkpoints, missing baselines, unreproducible configs — and they check for them reflexively. This chapter is a field guide to what they expect, section by section, so your paper reads as careful rather than lucky.
Before reading your argument, an experienced reviewer scans for:
Pass these five checks and the reviewer reads charitably. Fail two and they read adversarially. Everything below serves these checks.
Introduction. State the behavioral gap (Chapter 11's one sentence) and why prompting/RAG were insufficient — with numbers, not assertions. "Few-shot prompting reaches 71.8%; our fine-tuned model reaches 78.5%" belongs in the introduction's contribution list, because it is the contribution. End with explicit contributions: the artifact (dataset? adapter? benchmark?), the method finding, the ablation insight.
Related work. Cover three threads: (a) your task/domain's prior approaches, (b) fine-tuning methods you build on (LoRA, QLoRA — cite them), (c) evaluation practices for your task. The most common related-work failure is omitting the strongest prior baseline and then "beating" weaker ones. Cite and compare against the best, not the most convenient.
Method. This section has a checklist — include all of it:
A good test: could a competent graduate student reproduce your training run from this section plus your released code? If yes, you've written enough.
Experiments. Lead with the main results table (Chapter 8's format): baselines + your model, task metrics + format checks + regression benchmark, mean ± std. Then the ablations, each as a figure or small table answering one question: data-size curve, rank curve, replay on/off, stage ablation. Then qualitative analysis: the error taxonomy table (Chapter 8) with 2–3 illustrative examples. Reviewers remember the error analysis — it signals honesty more than any metric does.
Limitations. State them plainly: data size limits, languages/domains not tested, forgetting observed, compute constraints that prevented ablations. A limitations section doesn't weaken a paper; its absence weakens it, because reviewers will supply the limitations themselves, less generously.
Reproducibility artifacts. Release: the adapter weights, the training code or Axolotl config, the dataset (or the generation/verification scripts if the data can't be shared), the eval script, and a requirements file with pinned versions. Link a single repository. Papers with working repositories get cited; papers without them get doubted.
Two tables reviewers look for explicitly. The hyperparameter table:
| Hyperparameter | Value |
|---|---|
| Base model | Llama-3.1-8B (revision abc123) |
| PEFT method | QLoRA (NF4, double quant) |
| LoRA rank / alpha / dropout | 16 / 32 / 0.05 |
| Target modules | q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj |
| Learning rate / schedule / warmup | 2e-4 / cosine / 30 steps |
| Batch size (micro × accum) | 2 × 8 = 16 |
| Epochs | 3 (best checkpoint by val loss) |
| Max seq length | 1024 |
| Optimizer | paged_adamw_8bit |
| Seeds | 42, 43, 44 |
| Library versions | transformers==4.44.2, peft==0.12.0, trl==0.9.6 |
| Compute | 1× T4 16 GB, ~25 GPU-hours total |
And the data table: splits with counts, source, license, and the template shown verbatim. These tables are boring to write and invaluable to readers — they are the difference between "inspired by this paper" and "reproduced this paper."
Even a well-reported paper gets critical reviews. Here are the five most common critiques of fine-tuning papers and how to respond — ideally by having already addressed them, but also in the rebuttal if they arrive:
"The baseline is weak / unfair." The deadliest critique. In the rebuttal, don't just argue — run the stronger baseline. Add the few-shot result with a tuned prompt (show the prompt in the appendix), or the RAG comparison, and update the table. One new experiment in rebuttal beats three paragraphs of argument. Prevention: include the strong baselines from the start (Chapter 8).
"The improvement is marginal." If the gap is small, check: is it statistically significant (paired test across seeds)? Does it concentrate in important slices (e.g., +2% overall but +15% on the hard subset)? Reframe honestly around where the gain matters, or concede and pivot the paper's claim to the ablation insight ("the contribution is the data-efficiency finding, not the absolute score"). Reviewers respect a narrowed claim; they punish a defended overclaim.
"Ablations are missing." Run the single most informative one — usually the data-size curve or the rank curve — and add the figure. If compute truly prevents it, say exactly what you ran, what you couldn't, and why the existing evidence still supports the claim. "We could not run X due to compute; however, Y and Z jointly imply…" is acceptable once, not as a pattern.
"Concerns about contamination / test leakage." Describe your dedup procedure precisely (normalization, threshold, counts removed), and if possible add results on a freshly collected test slice the model never saw during development. This is the critique where prevention (Chapter 8's hygiene, documented from day one) matters most — reconstructed hygiene in rebuttal is far less convincing.
"Limited scope / only one dataset." Acknowledge the scope, state it as a limitation (you already did — Chapter 12's limitations section), and add whatever generalization evidence is cheap: a second small dataset, a paraphrase-robustness test, or an out-of-distribution slice. You don't need to match a lab's resources; you need to show the finding isn't a one-dataset accident.
Rebuttal tone: factual, grateful, specific. "We thank the reviewer for this point. We have added the few-shot baseline (Table 3, now 76.1 ± 0.8 vs. our 78.5 ± 0.7); the gap persists." Never argue that a requested experiment is unnecessary — either run it or explain precisely why it's infeasible and what evidence substitutes. Reviewers are volunteers; the rebuttal that makes their job easy (numbered responses, exact locations of changes) gets the benefit of the doubt.
For your research: Before submitting, do a "reviewer pass": print the five checks from this chapter's opening and verify each against your draft with page numbers. Then hand the draft to a colleague with the instruction "try to disbelieve the main claim" — the objections they raise in an hour predict the reviews you'll get in three months. Fix them now, when it's cheap.
Quick-reference appendix. Keep this open while you train.
| Dimension | Prompting | RAG | Fine-tuning |
|---|---|---|---|
| Changes model weights? | No | No | Yes |
| Solves | Eliciting existing capability | Missing/changing facts | Behavioral gaps, internalized skills |
| Data needed | 0–10 examples | Document corpus | 100s–1000s of curated examples |
| Time to first result | Minutes | Days | Days–weeks |
| Inference cost | High (long prompts / big model) | Medium (+ retrieval) | Low (small tuned model possible) |
| Maintenance | Edit text | Update index | Retrain |
| Main risk | Prompt brittleness | Retrieval errors, context limits | Overfitting, forgetting |
| Combine with others? | Yes | Yes — with fine-tuning | Yes — with RAG |
| Hyperparameter | Start here | Try if... | Warning sign |
|---|---|---|---|
| Learning rate | 2e-4 | loss flat → 3e-4; spikes → 1e-4 | NaN loss, wild oscillation |
| LoRA rank r | 16 | hard task → 32/64 | val gap train/val widening |
| LoRA alpha | 2 × r | — (keep ratio constant) | — |
| Epochs | 3 | small data → 2; large data → 2 | val loss rising |
| Effective batch | 16–32 | unstable → 64–128 | OOM → lower micro batch first |
| Warmup steps | 30–100 | — | early loss spikes |
| Scheduler | cosine | — | — |
| Optimizer | paged_adamw_8bit | — | — |
| Dropout | 0.05 | overfitting → 0.1 | — |
| Max seq length | 95th pct of data | OOM → lower | truncated examples |
| Setup | Hardware | Approx. cost | Best for |
|---|---|---|---|
| Colab free | T4 16 GB | $0 (limited hours) | Learning, debugging, small runs |
| Colab Pro | T4/V100-class | ~$10/month | Real student projects |
| Kaggle | T4/P100 | $0 (weekly quota) | Overflow experiments |
| Cloud spot | A100 40/80 GB | ~$0.50–1.50/hr (verify current) | Big/fast runs |
| Cloud on-demand | A100 80 GB | ~$1.50–4/hr (verify current) | Deadline runs |
| University cluster | Varies | $0 — ask your dept! | Everything |
Estimation recipe: 100-step calibration run → sec/step → total steps = (examples × epochs) / batch → × number of planned runs → +50% buffer.
| Symptom | Likely cause | Fix (in order) |
|---|---|---|
| Loss NaN | LR too high; corrupt example | Halve LR; scan data for garbage |
| Loss flat from step 0 | LR too small; label masking bug | Raise LR 3×; check template/labels |
| Loss oscillates wildly | LR too high | Halve LR; add warmup |
| Train ↓, val ↑ | Overfitting | Early stop; lower rank; more data; dropout 0.1 |
| Val flat, train ↓ | Capacity/data ceiling | More diverse data; higher rank |
| CUDA OOM | Batch/sequence too large | micro batch 1 → shorter seq len → grad checkpointing |
| Good train, bad paraphrase | Memorization | Paraphrase augmentation; lower rank; more data |
| Task ↑, MMLU ↓↓ | Catastrophic forgetting | Add 5–15% replay data; fewer epochs; lower LR |
| Outputs ignore format | Template mismatch train vs inference | Use identical template; check chat template |
| Slow training | Padding waste; small batches | packing=True; larger effective batch |
[1] E. J. Hu et al., "LoRA: Low-rank adaptation of large language models," in Proc. Int. Conf. Learning Representations (ICLR), 2022.
[2] T. Dettmers, A. Pagnoni, A. Holtzman, and L. Zettlemoyer, "QLoRA: Efficient finetuning of quantized LLMs," in Proc. Advances in Neural Information Processing Systems (NeurIPS), vol. 36, 2023.
[3] J. Wei et al., "Finetuned language models are zero-shot learners," in Proc. Int. Conf. Learning Representations (ICLR), 2022.
[4] H. W. Chung et al., "Scaling instruction-finetuned language models," arXiv preprint arXiv:2210.11416, 2022.
[5] L. Ouyang et al., "Training language models to follow instructions with human feedback," in Proc. Advances in Neural Information Processing Systems (NeurIPS), vol. 35, 2022.
[6] P. Lewis et al., "Retrieval-augmented generation for knowledge-intensive NLP tasks," in Proc. Advances in Neural Information Processing Systems (NeurIPS), vol. 33, 2020.
[7] T. Brown et al., "Language models are few-shot learners," in Proc. Advances in Neural Information Processing Systems (NeurIPS), vol. 33, 2020.
[8] J. Kaplan et al., "Scaling laws for neural language models," arXiv preprint arXiv:2001.08361, 2020.
[9] T. Dettmers, M. Lewis, S. Shleifer, and L. Zettlemoyer, "8-bit optimizers via block-wise quantization," in Proc. Int. Conf. Learning Representations (ICLR), 2022.
[10] I. J. Goodfellow, M. Mirza, D. Xiao, A. Courville, and Y. Bengio, "An empirical investigation of catastrophic forgetting in gradient-based neural networks," arXiv preprint arXiv:1312.6211, 2013.
[11] C. Zhou et al., "LIMA: Less is more for alignment," in Proc. Advances in Neural Information Processing Systems (NeurIPS), vol. 36, 2023.
[12] L. von Werra et al., "TRL: Transformer reinforcement learning," GitHub repository, 2020. [Online]. Available: https://github.com/huggingface/trl
End of Book 17 — Fine-Tuning LLMs on Your Own Data. Next in the series: Book 18.