
Book 15 of 50 · Free
What Are Large Language Models?
27,077 words · 128 chapters · illustrated

Book 15 of 50 · Free
27,077 words · 128 chapters · illustrated
Book 15 of 50 — AstolixGen Learning Series

This book explains what large language models (LLMs) actually are — under the hood, not in the marketing. You will learn how they turn text into numbers, how they are trained, why scale changed everything, what they can and cannot do, and how to use them honestly and safely in your research work.
Who it is for: MS and PhD students, early AI researchers, and anyone who uses tools like ChatGPT, Claude, or LLaMA but wants to understand what is really happening inside them.
What this book assumes: Basic programming experience. No prior machine learning background is required — every concept is built up from scratch.
How to use this book: Read Chapters 1–4 in order; they build the foundation. Chapters 5–12 can be read in any order based on your interest. Each chapter ends with a "For your research" box connecting the ideas to your academic work, plus key takeaways. The practice exercises at the end are designed to be done on your own machine, and several are written specifically for research students.
By the end of this book, you will be able to:
Strip away the hype and a language model is a surprisingly simple idea: a program that assigns probabilities to sequences of words. Given some text, it tells you how likely that text is, or — equivalently — it predicts what word comes next.
Think about your phone's keyboard. When you type "I am going to the", it suggests "store", "beach", "gym". That is a tiny language model at work. It learned from the text you typed before that after "I am going to the", certain words are likely. A large language model is the same idea, scaled up by an extraordinary amount: instead of your texting history, it learns from trillions of words of books, articles, websites, and code; instead of guessing one word ahead, it builds a deep statistical picture of how language works.
Every LLM you have ever used — ChatGPT, Claude, Gemini, LLaMA — is, at its core, doing next-word prediction, one word at a time, over and over. Everything else — the reasoning, the code writing, the summarization — emerges from that simple loop. If you truly understand next-word prediction, you understand the foundation of the entire field. This chapter traces how we got from counting words in the 1940s to today's models.
The oldest form of language modeling is almost embarrassingly simple: count how often words appear together. A model called an n-gram model looks at short sequences of n words and estimates probabilities from raw counts.
Suppose we want to know what word likely follows "the cat sat on the". A trigram model (n=3) looks at the last two words, "on the", and asks: in a huge collection of text, what words appeared after "on the"? It counts:
Then it assigns probabilities: P(mat | "on the") = 4200 / (total count of "on the ..."). The prediction is the most common continuation.
This works, but it has a fatal weakness: sparsity. A trigram model only remembers the last two words. It cannot use the beginning of a sentence to inform the end. Longer n-grams would help, but most 5- or 6-word sequences never appear in any dataset, so their counts are zero — the model has nothing to say. Researchers invented clever smoothing techniques to handle unseen sequences, but the fundamental limit remained: n-gram models have no understanding; they are lookup tables of counted phrases.
A famous early idea in this lineage is Claude Shannon's 1948 work on information theory, where he estimated the entropy of English text by having people guess the next letter — essentially treating humans as language models. The n-gram approach dominated statistical language processing from the 1980s through the 2000s and powered early speech recognition and machine translation.
Concrete illustration. Imagine an n-gram model trained only on this sentence, repeated: "The researcher published a paper. The researcher published a paper." Ask it to complete "The researcher published a ". It confidently says "paper", because that is the only continuation it has ever seen. Now ask it to complete "The student published a ". It has no data for "student published a", so it guesses randomly or fails. That is the sparsity problem in a nutshell: the model cannot generalize to word combinations it has never counted.
The breakthrough that broke the n-gram ceiling came from a simple but profound idea: instead of treating words as atomic symbols ("paper" is just label #4821), represent each word as a list of numbers — a vector — where similar words have similar vectors.
This is the word embedding idea. In an embedding space, the vector for "king" minus "man" plus "woman" lands near "queen". Words with similar meanings cluster together. Crucially, this solves the sparsity problem: even if the model has never seen "the student published a", it knows "student" is similar to "researcher" (their vectors are close), so it can borrow statistics from one to make a sensible guess about the other.
In 2003, Yoshua Bengio and colleagues published "A Neural Probabilistic Language Model," which trained a neural network to predict the next word from learned word embeddings. It was a modest model by today's standards, but it proved the concept: a neural network with distributed word representations could generalize across similar words in a way n-gram counting never could.
The next decade brought a series of improvements. Recurrent neural networks (RNNs) — especially LSTMs, introduced by Hochreiter and Schmidhuber in 1997 and popularized for language in the 2010s — processed text word by word, maintaining a "memory" of what came before. They could handle longer contexts than n-grams, but they had a weakness: they read sequentially, one word at a time, which made them slow to train, and their memory faded over long passages. If a pronoun referred back to a noun thirty words earlier, an RNN often lost track.
Concrete illustration. Consider: "The trophy did not fit in the suitcase because it was too big." What does "it" refer to? A human knows: the trophy was too big. An RNN reading word by word must carry the idea of "trophy" forward through twenty words of intervening text, and its memory tends to blur. This kind of long-range dependency was the hard problem of the 2010s.
In 2017, researchers at Google published "Attention Is All You Need" (Vaswani et al.), introducing the Transformer architecture. Its key idea is called self-attention: instead of reading words one at a time, the model looks at the whole input at once and learns, for each word, which other words are relevant to it.
Go back to the trophy sentence. A Transformer processing the word "it" can directly "attend" to "trophy" and "suitcase", weighing which is more relevant, regardless of distance. There is no fading memory — every word can look directly at every other word. This solved the long-range dependency problem, and as a bonus, because the model processes all words simultaneously, training could be massively parallelized on GPUs. That parallelization is what made training on trillions of words practical.
The Transformer has two halves: an encoder (reads input and builds a representation) and a decoder (generates output one piece at a time). The models in this book are almost all decoder-only — they consist of stacked decoder layers whose job is to predict the next piece of text given everything before it. GPT stands for "Generative Pre-trained Transformer": generative (it generates text), pre-trained (trained on raw text before any specific task), transformer (the architecture).
OpenAI's GPT series is the clearest demonstration of the scaling story:
In parallel, BERT (Devlin et al., 2018) showed the encoder side of the Transformer could produce excellent representations for understanding tasks like classification and question answering. BERT and GPT represent the two great branches of the Transformer family tree: encoders for understanding, decoders for generation. Modern LLMs are overwhelmingly decoder-based.
Then came ChatGPT (November 2022), which was not a fundamentally new model — it was GPT-3.5 fine-tuned with human feedback (more on this in Chapter 3) and wrapped in a chat interface. The interface mattered enormously: suddenly anyone could talk to the model in plain language. Within two months it had 100 million users, and the LLM era went mainstream.
Three lessons from this history are directly relevant to your work as a researcher:
Simple objectives scale. Next-word prediction is a humble objective. Nobody hand-labeled the training data with "reasoning" or "summarization" tags. The lesson: a simple, scalable objective applied to enough data can produce surprisingly general capabilities. When you design experiments, prefer objectives that scale over objectives that require expensive human labeling.
Architecture enables scale. The Transformer did not win because it was the most biologically plausible or the most elegant idea — it won because it parallelized well on GPUs. In ML research, the architecture that trains fastest at scale often beats the architecture that is smartest in theory. Keep this in mind when evaluating new architectures in the literature.
Interfaces matter as much as models. ChatGPT's breakthrough was partly a user-interface breakthrough. As a researcher, remember that how people interact with a model shapes what the model is "for." Prompting, tool use, and agents are interface innovations layered on top of the same next-token engine.
For your research: When you read a new LLM paper, first ask: what is the training objective? Almost always, it is next-token prediction plus some fine-tuning recipe. Papers that sound like they invented a new kind of intelligence usually just found a better way to train or prompt the same engine. Reading with this lens will save you from a lot of hype.
Between Shannon's 1948 work and the neural turn, language modeling lived inside speech recognition and machine translation labs, and the story is worth knowing because it explains why neural models won.
In the 1980s and 1990s, IBM's speech recognition group built large n-gram models from millions of words — enormous efforts at the time. Their models worked well enough to make dictation software possible, but every improvement required more counted data and cleverer smoothing. The field's dirty secret was that performance came from data size more than modeling cleverness — an early preview of the scaling lesson.
The 2000s added a new idea: statistical machine translation. Given millions of translated sentence pairs (parliamentary proceedings were a favorite source), systems learned which phrases corresponded across languages. These systems were entirely built from counted statistics — phrase tables and n-gram language models — with no neural networks at all. Google Translate's 2006 launch was powered by this approach, and it worked surprisingly well for closely related languages.
What finally broke the statistical era was not a lack of data but a lack of generalization. A phrase-based translation system that had seen "the cat sat on the mat" a thousand times still failed on "the cat sat on the rug" if "rug" never appeared in that exact phrase. Neural embeddings (popularized by word2vec in 2013, from Mikolov et al. at Google) gave every word a position in a continuous space, so "rug" could borrow statistics from "mat." Once embeddings existed, the rest of the neural takeover — seq2seq models for translation (2014), attention mechanisms (2015), and finally the Transformer (2017) — followed in rapid succession.
The pattern to notice: each era's dominant method was the one that best exploited the available compute and data. Counting won when data was scarce and compute was weak. Neural networks won when GPUs made it possible to learn from billions of words. If you remember one meta-lesson from this history, it's this: bet on the method that scales, because data and compute have grown monotonically for seventy years and show no signs of stopping.
Neural networks only understand numbers — long lists of them, arranged in matrices. They cannot see letters. So before a language model can process the sentence "Hello, world!", that text must be converted into a sequence of integers. That conversion is called tokenization, and the pieces it produces are called tokens.
A token is not necessarily a word. It might be a whole word ("hello"), part of a word ("un" + "believ" + "able"), a single character, or even part of a character for some scripts. The model's entire world — everything it reads and everything it writes — is made of tokens. Understanding tokens is understanding the model's native language, and it explains several things that otherwise seem like bizarre model behavior: why models count characters poorly, why some languages cost more to process, and why prompts have length limits measured in tokens, not words.
Why not just split text into words? Three reasons:
The dominant algorithm is byte-pair encoding (BPE), adapted for language modeling by Sennrich et al. (2016). Here is how it works, step by step.
Imagine our training corpus is just these words (with counts): "low" ×5, "lower" ×2, "newest" ×6, "widest" ×3. BPE starts with individual characters plus a special end-of-word marker (shown here as ·):
Step 0 — Initial vocabulary: all characters: l, o, w, e, r, n, s, t, i, d, ·
Step 1 — Count every adjacent pair across the corpus (weighted by word counts): - "l o" appears in "low" (5) + "lower" (2) = 7 - "o w" appears in "low" (5) + "lower" (2) = 7 - "w ·" appears in "low" (5) = 5 - "e s", "s t" appear in "newest" (6) and "widest" (3) = 9 each - ... and so on.
Step 2 — Merge the most frequent pair. Suppose "e s" wins. Add "es" to the vocabulary, and rewrite the corpus treating "es" as one symbol: "newest" becomes n e w es t ·.
Step 3 — Repeat. Count pairs again in the rewritten corpus, merge the most frequent, and continue for a fixed number of merges (real tokenizers do 30,000–50,000+ merges).
After enough merges, frequent words become single tokens ("low" might become one token), while rare words stay split into pieces. The final vocabulary is the set of merged units plus the original characters.
At inference time, tokenizing new text is a mechanical process: start with characters and apply the learned merges in the same order. "lowest" (never seen in training) might tokenize as "low" + "est" — two known tokens combining to cover an unknown word. Nothing is ever truly "unknown": in the worst case, text falls back to individual bytes, and since BPE is usually byte-level, any text in any language or script can be tokenized.
Concrete illustration. Take the sentence: "Tokenization is surprisingly tricky!" A real BPE tokenizer (like GPT's) might split it as:
["Token", "ization", " is", " surprisingly", " tricky", "!"]
Notice: "Tokenization" becomes two tokens ("Token" + "ization"), common words like " is" get one token each (the leading space is part of the token — a quirk worth knowing), and punctuation gets its own token. Six tokens for four words. As a rough rule of thumb, 1 token ≈ 0.75 English words, or about 100 tokens ≈ 75 words.
Tokens explain several famous model oddities:
Different model families use different tokenizers, and they are not interchangeable. GPT models use OpenAI's tiktoken library; LLaMA models use a SentencePiece-based BPE tokenizer; Mistral uses its own. If you count tokens with the wrong tokenizer, your counts will be wrong — and for API billing, the provider's tokenizer is the one that counts.
Try this yourself (a good exercise): install tiktoken in Python and tokenize a paragraph in English and the same paragraph translated into Urdu. Compare the token counts. The difference is the "tokenizer tax" paid by non-English languages, and it is worth experiencing firsthand.
For your research: If your research involves non-English text — especially Urdu, Arabic, or other Perso-Arabic scripts — tokenization is not a footnote; it is a first-order concern. Measure the tokens-per-word ratio for your data before choosing a model. A model with a multilingual-friendly tokenizer (larger vocabulary, better byte-level handling) may outperform a "stronger" model that wastes half its context window on token fragments. Several published papers report results without mentioning the tokenizer; when you replicate or extend them, check this first.
Real tokenizers have details worth knowing once you work with models directly.
Special tokens. Beyond word pieces, every tokenizer has special tokens with reserved IDs: beginning/end-of-sequence markers, padding tokens (for batching), and sometimes mask tokens. When you see a model generate <|endoftext|>, that's a special token leaking into the output — usually a sign the prompt ended ambiguously.
Vocabulary size trade-offs. Bigger vocabularies (100K–200K tokens) compress text into fewer tokens — great for long contexts and non-English languages — but make the model's output layer larger and slower. Smaller vocabularies (30K–50K) are leaner but split rare words into many pieces. There is no free lunch; multilingual models tend toward larger vocabularies for this reason.
Prompt boundary effects. Tokenization depends on context: the same word may tokenize differently at the start of a prompt versus mid-sentence (remember the leading-space quirk: " is" vs "is"). This means small prompt edits can change tokenization downstream in surprising ways. When a prompt behaves oddly after a trivial edit, tokenization is worth checking.
Practical tips. One: always use the model's own tokenizer for counting — the transformers library's AutoTokenizer loads the right one automatically. Two: when budgeting API costs, add 10–20% headroom for tokenization overhead on formatted text (JSON, code, and markdown tokenize less efficiently than prose). Three: if you work in Urdu, Arabic, or similar scripts, compare tokenizers before committing to a model — some open models now ship improved multilingual tokenizers, and the difference in tokens-per-word can be dramatic.
A fun experiment. Tokenize the string "SolidGoldMagikarp" or count how many tokens are in a single emoji versus a Chinese character. Tokenizers are full of historical accidents from their training data, and poking at them builds intuition for why models behave strangely on unusual inputs.
The best way to internalize tokenization is to run a tokenizer yourself. Here is a lab you can do in fifteen minutes with Python.
Setup. Install the tiktoken package (OpenAI's open-source tokenizer library): pip install tiktoken. Then:
import tiktoken
enc = tiktoken.get_encoding("cl100k_base") # the GPT-3.5/GPT-4 tokenizer
text = "Tokenization is surprisingly tricky!"
tokens = enc.encode(text)
print(tokens) # the integer IDs
print([enc.decode([t]) for t in tokens]) # what each ID represents
You will see something like [12119, 21585, 318, 42866, 8912, 0] decoding to pieces like "Token", "ization", " is", " surprisingly", " tricky", "!". Try your own sentences. Try code. Try a URL. Watch how the splits fall in unintuitive places.
Measure fertility. "Fertility" is tokens per word. Compute it for a paragraph:
def fertility(text):
words = len(text.split())
toks = len(enc.encode(text))
return toks / words
Run it on English prose (expect ~1.3 tokens/word), on Python code (often higher — code tokenizes inefficiently), and on Urdu or Arabic text (often 2–4x English). This single number explains half of the cost and context-length differences between languages.
Compare tokenizers. Hugging Face's transformers library lets you load other models' tokenizers: AutoTokenizer.from_pretrained("meta-llama/Meta-Llama-3-8B"). Tokenize the same paragraph with two tokenizers and compare counts. You will find they differ — sometimes substantially — which is why "128K context" means different amounts of text on different models.
Break it. Feed the tokenizer pathological inputs: a 200-character emoji string, a German compound word like "Donaudampfschifffahrtsgesellschaftskapitän", a string of random punctuation. Observe where it falls back to bytes. This builds the instinct that will later help you debug weird model behavior: when output looks strange, check the tokens first.
Tokenization isn't just plumbing — it visibly shapes what models can and can't do.
Glitch tokens. Some tokens, learned from odd corners of training data, behave strangely. The famous GPT-2 examples — tokens like " solidarnosc" or long strings of spaces — could derail generation when the model was forced to repeat them. They arise because a token that appeared rarely in training has a poorly-calibrated representation. Modern tokenizers are cleaner, but the lesson stands: the vocabulary is a historical artifact, and its rough edges become model quirks.
Word games and spelling. Ask a model to "spell 'strawberry' backwards" and it may fail — not from stupidity, but because it never sees letters. It sees "straw" + "berry" and must reconstruct spelling from statistical memory of the word's written form. Newer models often succeed by reasoning it out step by step, but the underlying blindness is architectural. Any task that requires character-level manipulation (acronyms, ciphers, rhyming under constraints) fights the tokenizer.
Token healing at boundaries. When you stop generation mid-word and continue, or when prompts end mid-token, models can produce slightly-off continuations because the token boundary shifted. APIs that let you pass "prefix" text handle this with token healing; when building your own pipelines, be aware that concatenating strings isn't the same as concatenating token streams.
The practical moral. When a model does something bizarre on a seemingly simple task, check the tokens before blaming the model's intelligence. A remarkable fraction of "the model is dumb" moments are actually "the tokenization is weird" moments — and knowing the difference makes you a much better debugger of LLM systems.

Pretraining teaches a model to do exactly one thing: given a sequence of tokens, predict the next token. That is the entire objective. The model reads trillions of tokens of text and, at every position, tries to guess what comes next, adjusting its internal numbers slightly each time it is wrong.
If that sounds too simple to produce ChatGPT, you are having the correct reaction. The rest of this chapter is about why this simple objective, applied at enormous scale, produces something so capable — and what "learning" really means here.
Pretraining data is a vast, messy crawl of human writing: web pages, books, Wikipedia, news articles, forums, scientific papers, and code repositories. For a modern model, the training set is on the order of trillions of tokens — roughly equivalent to millions of books. Nobody curates it by hand; it is filtered automatically for quality and deduplicated, but it remains a rough cross-section of what humanity has written online.
Two important consequences follow:
Training is a loop, repeated billions of times:
Repeat for every position in every document, for trillions of tokens, over weeks or months on thousands of GPUs. The parameters — the billions of numbers that make up the model — slowly settle into a configuration that predicts text well.
Concrete illustration. Early in training, given "The capital of France is", the model might assign 0.001% probability to "Paris" — essentially guessing. The loss is large, and the adjustment is large. After seeing thousands of documents mentioning France and Paris, the model assigns 85% to "Paris". But it also learns subtler things: given "The capital of France is not", the probability shifts toward "London" or "Berlin"; given a French-language context, it predicts "Paris" with the right surrounding grammar. Each prediction task is tiny, but trillions of them, layered, build a rich statistical model of language, facts, reasoning patterns, and style.
Think of pretraining as building layers of ability, from shallow to deep:
The honest summary: pretraining builds a compressed statistical simulator of text. Point it at a prompt, and it continues the text the way its training data suggests text like that usually continues.
A raw pretrained model — a base model — is not a chatbot. Given "What is the capital of France?", a base model might continue "What is the capital of Spain? What is the capital of Italy?..." because its training data contains many quiz lists. It predicts likely continuations, not helpful answers.
Two fine-tuning stages turn a base model into an assistant:
Newer variants like DPO (direct preference optimization) skip the separate reward model, but the idea is the same: pretraining teaches the model what text looks like; fine-tuning teaches it how to behave.
Concrete illustration — base vs. fine-tuned. Prompt: "Explain photosynthesis simply."
Same underlying knowledge; different behavior. The difference is entirely in the fine-tuning.
For your research: If you ever fine-tune a model yourself, remember the base/fine-tune distinction. Fine-tuning teaches behavior and format, not knowledge — trying to teach new facts through fine-tuning is unreliable (the model may just parrot your training examples). For adding knowledge, retrieval-based methods (giving the model documents at inference time) are usually more reliable. This distinction will save you weeks of confused experiments.
Pretraining data doesn't just appear — building it is a major engineering effort, and its composition shapes the model.
Where the text comes from. The backbone is web crawls (Common Crawl, a nonprofit archive of the web), supplemented with books, Wikipedia, news, scientific papers, forums like Stack Exchange and Reddit, and code from GitHub. The Pile (Gao et al., 2020), a public 800GB dataset, showed what a deliberate mix looks like; most frontier labs now build proprietary equivalents.
Filtering matters enormously. Raw web text is full of spam, machine-generated junk, and duplicates. Labs filter aggressively: removing boilerplate, deduplicating (exact and near-duplicate), and upweighting high-quality sources (textbooks, Wikipedia, well-written articles). There is growing evidence that data quality beats data quantity past a point — a smaller set of clean, diverse, educational text can outperform a larger set of web sludge. This is one reason training recipes are closely guarded secrets.
The fine-tuning zoo. Beyond SFT and RLHF, the post-training toolkit keeps growing: DPO (direct preference optimization) achieves RLHF-like results without a separate reward model; constitutional AI (Anthropic) uses model-generated critiques against a written constitution instead of pure human ratings; and RLVR (reinforcement learning with verifiable rewards) trains on tasks with checkable answers like math and code. The trend is toward fine-tuning that rewards verifiable correctness, not just human likability — a direct response to the sycophancy problem.
What fine-tuning can't fix. Fine-tuning adjusts behavior, but the model's factual knowledge and core capabilities come from pretraining. If the base model never learned a fact, fine-tuning rarely installs it reliably — the model learns to sound like it knows rather than to know. This is why retrieval (Chapter 12) beats fine-tuning for knowledge injection, and why "we'll just fine-tune it to be truthful" is not a complete safety strategy.
What does pretraining actually look like operationally? Picture a data center with thousands of GPUs running a single training job for months. Here's the loop at industrial scale.
Data pipeline. Before training starts, trillions of tokens are preprocessed: filtered, deduplicated, tokenized, and packed into fixed-length sequences (typically 2K–8K tokens each). These sequences stream from storage to GPUs continuously — the data pipeline must never starve the GPUs, because idle GPUs burn money.
The training step, distributed. Each GPU holds a copy of the model and processes a batch of sequences: forward pass (predict next tokens), loss computation, backward pass (gradients). Then all GPUs synchronize their gradients — averaging them across the cluster — and each updates its copy identically. This is data-parallel training. For the largest models, the model itself is split across GPUs (tensor/pipeline parallelism) because it doesn't fit on one chip.
What the engineers watch. The loss curve: it should fall smoothly. Spikes or divergence mean something broke — bad data batch, numerical instability, a failed GPU. Checkpoints are saved regularly; if the run crashes (and multi-month runs always crash eventually), training resumes from the last checkpoint. Monitoring dashboards track loss, gradient norms, throughput (tokens/second), and hardware health.
What can go wrong. Loss spikes from a corrupted data batch. Silent GPU failures producing wrong gradients. Running out of memory from an oversized batch. A bug in the learning-rate schedule discovered after two weeks. Frontier training is as much reliability engineering as machine learning — which is why the teams that do it well are small and experienced.
Why you should care. You will probably never run a pretraining job, but understanding this demystifies the artifacts you do use: why models have quirks (training data artifacts), why checkpoints matter (each is a slightly different model), and why training is so expensive (it's the GPUs, the power, the engineers, and the risk, all multiplied by months).
A reasonable objection: predicting the next word seems too shallow to produce reasoning, coding, or translation. Why does it work? Because predicting text well requires modeling what the text is about. To predict the next token of a physics derivation, the model must implicitly track the physics. To predict the next line of a program, it must implicitly track program state. The text is a shadow of the world; predicting the shadow accurately requires a model of the thing casting it. Scale makes the shadow detailed enough that the implicit world-model becomes useful.
When isn't it enough? When the training text doesn't contain the pattern you need: genuinely novel reasoning (no worked examples to imitate), precise computation (text prediction isn't arithmetic), real-time facts (frozen training data), and physical interaction (no body, no tools by default). Each gap maps to a remedy: tools for computation and freshness, retrieval for knowledge, fine-tuning for behavior, agents for interaction. The modern LLM stack is next-token prediction plus these remedies — and understanding which remedy fits which gap is most of applied LLM engineering.

Every LLM training run combines three ingredients:
The central discovery of the late 2010s is that increasing these three together produces smooth, predictable improvements in model performance — and that at certain scales, qualitatively new abilities appear.
In 2020, researchers at OpenAI (Kaplan et al.) published "Scaling Laws for Neural Language Models," showing something remarkable: as you increase parameters, data, or compute, the model's loss (prediction error) decreases along a smooth power law — a straight line on a log-log plot. No sudden jumps in the loss curve; bigger is just better, predictably.
This had a huge practical consequence: you can train small models, fit the scaling curve, and predict how good a much larger model will be before spending the money to train it. Training runs became engineering projects with predictable returns rather than research gambles.
A follow-up refined the recipe. GPT-3 was arguably undertrained: 175B parameters but only ~300B training tokens. The Chinchilla paper (Hoffmann et al., 2022, from DeepMind) showed that for a fixed compute budget, parameters and data should be scaled equally — roughly 20 training tokens per parameter is optimal. A 70B-parameter model trained on 1.4T tokens (Chinchilla) outperformed the 175B-parameter GPT-3 while being smaller and cheaper to run. The lesson: data matters as much as size, and many early giant models were starved of data.
Concrete illustration. Imagine a fixed budget to build a student: you can buy a bigger brain (parameters) or more textbooks (data). Kaplan's laws say spending on either helps predictably. Chinchilla says: don't buy a huge brain and give it only three textbooks — balance the two. Modern models (LLaMA, Mistral) follow the Chinchilla recipe, training smaller models on much more data than GPT-3 used.
While loss improves smoothly with scale, capabilities sometimes appear suddenly. A 1B-parameter model cannot do multi-step arithmetic; a 100B-parameter model can. A small model fails at following complex instructions; a large one succeeds. These sudden appearances of new abilities are called emergent abilities (Wei et al., 2022).
Why does this happen? The honest answer: we don't fully understand it. One plausible view is that emergence is partly a measurement artifact — if you score tasks with all-or-nothing accuracy, a smooth underlying improvement looks like a sudden jump when the model crosses the threshold from "mostly wrong" to "mostly right." But even accounting for that, scale clearly unlocks behaviors (like step-by-step reasoning, or using tools) that small models don't exhibit.
The practical upshot: capability is not linear in size. A 7B-parameter model is not "7% as capable" as a 100B model — it may be unable to do entire categories of tasks the big model handles. But equally, for your task, the smallest model that works is the right model (Chapter 11).
Scale has a price, and it shapes who can do what in this field:
Concrete illustration — scale in physical terms. Training a model like GPT-3 consumed roughly the annual electricity of hundreds of homes, running for months on about 10,000 GPUs in parallel. Each training run is a bet of millions of dollars that the scaling laws hold. When your experiment fails on your laptop, you lose an afternoon; when a frontier training run fails, a company loses millions. This asymmetry is why scaling-law research (predicting big from small) is so valuable.
AI researcher Rich Sutton's "bitter lesson" (2019) argues that the methods winning in AI are not the clever hand-designed ones but the simple methods that scale with compute: search and learning. LLMs are the bitter lesson incarnate. Decades of hand-built grammars, ontologies, and rule-based NLP were swept aside by "predict the next word, but at planetary scale."
For your research career, the bitter lesson cuts two ways. Don't fight scale — build on top of scaled models rather than competing with them from scratch. But the next breakthroughs may come from finding new simple objectives that scale, or from making scale cheaper (efficiency, sparsity, better hardware use). Both are legitimate, publishable research directions.
For your research: When you read a paper claiming a new technique beats a baseline, always check: was compute controlled? A method that wins by using 10x more training compute hasn't necessarily shown a better idea — it may have just spent more. The best papers report results at matched compute budgets, or show that their method changes the scaling curve itself. Make this your default skeptical question.
For years, "scale" meant training scale: bigger models, more data, more FLOPs. Recently a new axis appeared: test-time (inference-time) compute — letting the model think longer before answering.
Reasoning models (like OpenAI's o1 and DeepSeek-R1) generate long internal chains of thought: trying approaches, checking work, backtracking. Giving the model 10x more thinking tokens often improves accuracy as reliably as training a bigger model — a new scaling law, this time at inference. This shifts the economics: instead of training one giant model, you can run a medium model with a large thinking budget for hard problems and a small budget for easy ones.
This has practical consequences for your work. First, cost estimation gets harder: you pay for thinking tokens you never see. Second, evaluation must account for effort: comparing a model that thought for 10 seconds against one that thought for 10 minutes is not a fair fight. Third, it opens research questions: how should models allocate thinking? When should they stop? Can small models with big thinking budgets beat large models with small ones? Early evidence says sometimes yes — which is excellent news for researchers without frontier training budgets.
There's also a paradox worth knowing: as models get cheaper and more efficient, total usage explodes — the Jevons paradox applied to AI. Efficiency gains don't reduce compute consumption; they increase it by making new applications viable. When you read optimistic claims about efficiency solving AI's energy problem, remember Jevons.
You don't have frontier compute. Here's how to still think — and work — in scaling terms.
Extrapolate from small. The core trick of scaling laws works at any budget: train (or fine-tune) at two or three small scales, plot the results, and extrapolate. If your method improves the scaling curve — better performance at every compute level — that's a real contribution, publishable without big iron. Reviewers in efficient-ML venues understand this methodology; it's standard.
Use the Chinchilla ratio in your own training. If you train a small model from scratch (say, for a low-resource language), don't just pick a size — pick size and data together. Roughly 20 tokens per parameter is the compute-optimal neighborhood. A 1B-parameter model wants ~20B tokens. Training a 1B model on 2B tokens is undertraining it; you'd do better with a smaller model on the same data, or more data for the same model.
Rent, don't buy, for spikes. Cloud GPU rentals make large experiments accessible in bursts: you don't need to own 8 GPUs to run a weekend fine-tuning experiment. Budget it like lab equipment: expensive per hour, cheap per project. Many universities have clusters — learn your institution's allocation process early; it's often underused by students who don't know it exists.
Distillation: borrow someone else's scale. If a frontier model does your task well but is too expensive, use it to label data, then train (or LoRA-fine-tune) a small open model on those labels. The small model learns to imitate the big one's behavior on your task at a fraction of the cost. This is distillation in practice, and it's one of the highest-leverage techniques available to resource-constrained researchers. (Check the provider's terms of service — some prohibit using outputs to train competing models.)
The mindset. Stop asking "can I afford to compete with big labs?" and start asking "what can I learn at small scale that generalizes?" Scaling laws, efficiency methods, evaluation, interpretability, low-resource languages — these all reward insight over budget. The bitter lesson says scale wins; the corollary is that understanding scale is itself a research superpower.
Let's make the scale concrete with rough numbers for a Chinchilla-style 70B-parameter model trained on 1.4 trillion tokens.
Compute. A common estimate: training FLOPs ≈ 6 × parameters × tokens. So 6 × 70B × 1.4T ≈ 5.9 × 10^23 FLOPs. A modern GPU delivers on the order of 10^15 effective FLOPs per second on this workload, so that's ~6 × 10^8 GPU-seconds — about 19 GPU-years. Spread across 10,000 GPUs: roughly a week of ideal time, realistically several weeks with overhead, restarts, and evaluation.
Money. At cloud rental rates of roughly $2–4 per GPU-hour, 19 GPU-years costs on the order of $300–600K for the final run alone — and the final run is preceded by many smaller experiments. Frontier runs 10–100x larger reach the tens-to-hundreds of millions cited in Chapter 4. These are rough figures, but the orders of magnitude are right, and they explain the industry's structure: only organizations that can risk millions per experiment get to run them.
Energy. Tens of gigawatt-hours for the largest runs — the annual electricity consumption of thousands of homes. This is why data-center location (cheap, clean power), efficiency research, and carbon accounting are live issues, not footnotes.
The point of this arithmetic isn't precision — it's intuition. When someone proposes "just training a bigger model," you now know what "just" costs.
One final intuition to carry with you: scale is a multiplier on ideas, not a substitute for them. The teams that win with scale are the ones who also get the data mix right, the architecture details right, and the evaluation right. When you read about a record-breaking training run, look past the parameter count to the recipe — that's where the transferable knowledge lives, and it's accessible to anyone who reads carefully.
"Which model should I use?" is one of the most common questions in AI research, and the answer changes every few months. This chapter gives you a durable mental map: the major model families, what distinguishes them, and — most importantly — the open-vs-closed distinction that determines what you can actually do with a model in your research.
The GPT line defined the modern LLM era. GPT-1 (2018) and GPT-2 (2019) were research demonstrations that pretraining plus scale produced increasingly fluent text. GPT-3 (2020), at 175B parameters, introduced few-shot learning to the mainstream (Brown et al., 2020). GPT-3.5 and ChatGPT (2022) took GPT-3-class models, fine-tuned them with RLHF, and wrapped them in a chat interface — the product that made LLMs famous. GPT-4 (2023) was a major capability jump: stronger reasoning, better instruction-following, multimodal input. OpenAI published almost no technical details about GPT-4's size or training, marking the industry's shift toward secrecy.
The defining trait of the GPT family today: closed. You cannot download the weights. You access them through OpenAI's API or website, paying per token. You can use them but not inspect them — no examining internals, no free fine-tuning, no offline use.
In February 2023, Meta released LLaMA (Touvron et al., 2023): models from 7B to 65B parameters, trained on public data with Chinchilla-style recipes. The 13B LLaMA outperformed the 175B GPT-3 on many benchmarks — a stunning demonstration that smaller, better-trained models beat bigger, undertrained ones.
LLaMA's weights were released under a non-commercial research license, then leaked publicly within days. The leak was arguably the most consequential event in open AI: suddenly anyone with a good GPU could download and run a capable LLM locally. An explosion of community fine-tunes followed (Alpaca, Vicuna, and hundreds more), and local tooling (llama.cpp, Ollama, vLLM) matured rapidly.
Llama 2 (July 2023) and Llama 3 (2024) continued with more permissive licenses and stronger models. Meta's open-weights strategy gave researchers, startups, and hobbyists a real alternative to closed APIs.
Concrete illustration. A PhD student studying Urdu NLP cannot send sensitive data to a foreign company's API (privacy, cost, and the tokenizer tax from Chapter 2). With LLaMA-class open weights, she downloads the model once, fine-tunes it on her own Urdu corpus on a university GPU cluster, and runs everything offline. That workflow is impossible with closed models. This is why open weights matter far beyond ideology — they determine what research is possible.
Mistral AI, a French startup founded in 2023, bet on smaller, highly efficient open models. Mistral 7B (September 2023) outperformed the 13B Llama 2 on many benchmarks despite being half the size. Mixtral 8x7B (December 2023) popularized mixture-of-experts (MoE) for open models: 8 expert sub-networks with each token routed to 2 of them, giving roughly the capacity of a 47B model at the inference cost of about 13B. Mistral releases several models under the Apache 2.0 license and also operates a commercial API.
Anthropic's Claude family: closed models known for careful safety work and long context windows, strong on reasoning-heavy tasks; their Constitutional AI approach to alignment is influential in safety research. Google's Gemini: closed models with very large context windows, deeply integrated with Google products. DeepSeek: released open-weights reasoning models (notably DeepSeek-R1, 2025) matching frontier closed models on reasoning benchmarks at dramatically lower training cost. Smaller open models — Microsoft's Phi, Google's Gemma, Alibaba's Qwen — cover the efficient 1B to 30B range.
Forget parameter counts for a moment. The single most important question about any model is: can you download the weights?
Open-weight models (LLaMA, Mistral, DeepSeek) give you: downloadable weights you run anywhere; no per-token fee (just hardware and electricity); data that never leaves your machine; full access to internals for research; full control over fine-tuning; and reproducibility through pinned checkpoints. Their trade-offs: usually one step behind the frontier in raw capability, and licenses vary.
Closed models (GPT-4, Claude, Gemini) give you: access through API or website only; per-token pricing; prompts sent to the provider's servers; a black box you cannot inspect; limited or no fine-tuning; and the risk that the provider silently updates the model behind the same API name. Their advantage: usually the strongest capability available.
A terminology warning: "open source AI" is contested. Truly open-source (OSI-approved license, training data disclosed) is rare; most "open" models are open-weight — you get the numbers, but not the training data or full training code. When a paper or vendor says "open," check what exactly is open.
Concrete illustration — reproducibility. You publish a paper evaluating GPT-4 via API in January. By June, the provider updates the model behind the same API name; a colleague rerunning your prompts gets different answers. Your results are not reproducible. With an open-weights model pinned to a specific checkpoint file, anyone can rerun your exact experiment years later. For academic work, this is a decisive argument for open weights.
Models don't exist in isolation. Key tooling: llama.cpp and Ollama for running models locally; vLLM for fast server inference; Hugging Face Transformers as the research-standard library; PEFT and LoRA for cheap fine-tuning; cloud GPU rentals for training; and benchmarks like MMLU and HumanEval for evaluation (useful but gameable — see Chapter 10).
For your research: Default to open-weights models for thesis or publication work unless you have a specific reason not to: reproducibility, cost, privacy, and inspectability. Use closed frontier models for quick capability checks, baseline comparisons, and tasks needing maximum quality on a small number of queries. Many strong papers do exactly this hybrid: develop on open models, compare against closed APIs.
Knowing the model families is one thing; actually using open models is another. Here's the practical layer.
Getting the weights. The Hugging Face Hub is the central repository: model pages list parameter counts, licenses, and download counts. A typical 8B model is ~16GB in half precision — downloadable, but hefty. Quantized formats (notably GGUF) shrink models 2–4x with small quality loss: a 4-bit quantized 8B model fits in ~5GB and runs on a laptop via Ollama or llama.cpp. For research fidelity, use full-precision checkpoints; for everyday use, quants are fine.
Cheap fine-tuning: LoRA. Full fine-tuning updates all billions of parameters — expensive. LoRA (low-rank adaptation) freezes the model and trains small adapter matrices instead, cutting trainable parameters by 10,000x. With QLoRA (quantized LoRA), you can fine-tune a 7B–13B model on a single consumer GPU. This is how most academic fine-tuning gets done today, and libraries like Hugging Face PEFT make it a few lines of code.
The fine-tune explosion. Because LoRA is cheap, the community has produced thousands of specialized fine-tunes: models tuned for medicine, law, specific languages, roleplay, code. Quality varies wildly. For research, prefer fine-tunes with documented training data and evaluations; treat anonymous uploads as untrusted.
Licenses, concretely. LLaMA's community license historically restricted commercial use above certain user counts and required special terms; Mistral's Apache 2.0 releases are genuinely permissive. Always read the license file on the model page before building something you might commercialize or redistribute. "Open weights" is not "no strings attached."
Model merging. A quirky community technique: averaging the weights of two fine-tuned models to combine their skills. It shouldn't work — but often does, well enough to be useful. It's a reminder of how little we understand about what the parameters actually encode.
Every serious model release comes with a model card — a document describing what the model is, how it was trained, and what it's for. On Hugging Face, it's the README on the model page. Here's how to read one like a researcher.
Check first: license. Look for the license field. Apache 2.0 or MIT: permissive, use freely. LLaMA Community License: usable for research, restrictions on commercial use at scale. Custom/restrictive: read carefully. No license listed: treat as all-rights-reserved. Your thesis is fine under most research allowances; a startup product is not.
Training data and cutoff. Good cards describe the training mix (web, code, books) and the knowledge cutoff date. Vague cards ("trained on a large corpus of public data") tell you the maker is hiding something — common for closed models, a yellow flag for open ones. The cutoff tells you what the model can't know.
Evaluations. Cards list benchmark scores. Read them skeptically (Chapter 10): which benchmarks, and are they the ones relevant to your task? A model card boasting about code benchmarks tells you little about Urdu summarization. Absence of evaluations is a red flag — it means nobody measured, or nobody liked the measurements.
Intended use and limitations. The best cards state plainly what the model is bad at. Take these seriously — they're the closest thing to an honest label. Also check "bias, risks, and limitations" sections; they often disclose known failure modes that would take you weeks to discover yourself.
Red flags. No license, no evaluation numbers, no description of training data, grandiose claims without evidence, or a model uploaded by an anonymous account with a name like "GPT-5-Ultra-Max." The Hub hosts gems and junk side by side; the model card is how you tell them apart. When in doubt, prefer models from established organizations with full documentation.
For all their capability, closed models are scientific black boxes — and the opacity has specific, practical consequences.
Unknown sizes and training. Providers no longer disclose parameter counts, training data, or compute budgets. When a paper reports "GPT-4 achieves X," nobody outside the provider knows what achieved X — which makes the result hard to interpret scientifically. Did the model succeed through scale, data, or a clever recipe? We can't say, so we can't learn the lesson.
Silent updates. Providers update models behind stable API names for safety, quality, or cost reasons. Your January results may not replicate in June — not because you erred, but because the instrument changed. Some providers now offer version-pinned snapshots; use them, and record the exact version string in every experiment log.
Hidden system prompts. API models run with provider-written system instructions you can't see — shaping tone, refusal behavior, and even factual answers. Two providers' models can differ because of these invisible prompts, not because of the underlying models. When comparing APIs, remember you're comparing products, not just models.
What this means for science. A field that depends on instruments it can't inspect has a reproducibility problem. The community's answer has been twofold: build open models that can be studied (Chapter 12's interpretability work depends on them), and treat closed-model results as behavioral observations — valuable, but provisional. As a student, you can contribute to both: use closed models for what they're good at, but anchor your publishable claims in open, pinned, inspectable artifacts.
A closing practical note: the landscape rewards generalists who track it lightly but continuously. You don't need to try every new release — but do skim release notes monthly, re-run your key evaluations when your pinned model gets a successor, and keep one eye on the efficiency frontier, where today's expensive capability becomes tomorrow's laptop model. The model that's overkill for your task today is the baseline for it next year.
Strip away the demos and the marketing, and modern LLMs reliably do a set of things well enough to be genuinely useful:
Language tasks. Summarization, translation, paraphrasing, tone adjustment, and style transfer. These are close to the pretraining objective — the model has seen millions of summaries, translations, and paraphrases — so they work well.
Code. Writing, explaining, debugging, and translating code. Code is abundant in training data (GitHub, Stack Overflow), highly structured, and verifiable (you can run it). LLMs are remarkably good at boilerplate, API usage, and explaining unfamiliar code.
Question answering over provided context. Given a document and a question about it, strong models answer accurately — this is the basis of retrieval-augmented systems.
Structured generation. Producing JSON, tables, outlines, and formatted documents to specification. Useful for data processing pipelines.
Reasoning-shaped tasks. With step-by-step prompting, models solve math word problems, logic puzzles, and multi-step analyses at levels that genuinely surprise — while failing unpredictably (Chapter 7).
Concrete illustration — a real prompt/response pair.
Prompt: "Summarize this paragraph in one sentence: 'The researchers trained a 7B-parameter model on 2 trillion tokens of web text, then evaluated it on 50 benchmarks. It matched the performance of a 70B model trained on 300B tokens, supporting the hypothesis that many large models are undertrained relative to their size.'"
Response: "A 7B model trained on far more data matched a 70B model trained on less, suggesting many large models are undertrained."
That is a good summary: it preserves the key claim and the evidence. This kind of compression is bread-and-butter LLM work.
Concrete illustration — code help.
Prompt: "Write a Python function that reads a CSV file and returns the number of rows where the 'status' column equals 'active'."
Response:
import csv
def count_active_rows(path):
count = 0
with open(path, newline='', encoding='utf-8') as f:
reader = csv.DictReader(f)
for row in reader:
if row.get('status') == 'active':
count += 1
return count
Correct, idiomatic, and explained if you ask. But note: you should still run it — models occasionally produce code with subtle bugs, wrong library versions, or invented API calls.
One of the most important capabilities is in-context learning: the model learns a task from examples placed in the prompt, without any parameter updates. Show it three examples of "English to French" pairs, then an English sentence, and it translates — having never been explicitly trained as a translator.
Prompt:
Translate English to Urdu:
English: Good morning -> Urdu: صبح بخیر
English: Thank you -> Urdu: شکریہ
English: How are you? -> Urdu:
Response: آپ کیسے ہیں؟
The model inferred the task pattern from two examples. Brown et al. (2020) showed this "few-shot" ability emerges with scale, and it is the reason prompt engineering exists as a discipline: the prompt is a programming language, and examples are its functions.
How does in-context learning work? The honest answer is that we don't fully understand the mechanism, though there are promising theories (the model may be doing something analogous to implicit gradient descent in its forward pass). What matters practically: more and better examples usually help; the order and format of examples matter more than you'd expect; and few-shot performance is a legitimate evaluation method for your research.
Chain-of-thought prompting (Wei et al., 2022) is simple: ask the model to show its reasoning step by step before answering. On arithmetic, logic, and multi-step problems, this dramatically improves accuracy.
Prompt: "A bat and ball cost $1.10 together. The bat costs $1.00 more than the ball. How much does the ball cost? Think step by step."
Response: "Let the ball cost x. Then the bat costs x + 1.00. Together: x + (x + 1.00) = 1.10, so 2x = 0.10, x = 0.05. The ball costs $0.05."
Without "think step by step," models often blurt out the intuitive-but-wrong $0.10. The step-by-step format forces the model to generate intermediate tokens, and each intermediate step becomes context for the next — the model literally reasons on the page, where it (and you) can see it.
Why does this work? Recall Chapter 3: the training data contains countless worked solutions. "Think step by step" steers generation into that region of the distribution. It's not magic — it's pattern completion of reasoning-shaped text — but it's remarkably effective.
The newest capability layer is tool use: models that call external tools — search engines, calculators, code interpreters, APIs — in the middle of generating a response. Instead of guessing the answer to "what is 48291 x 7304?", the model writes a calculator call, reads the result, and continues. Instead of hallucinating a paper citation, it searches the web.
Agents extend this: the model plans a multi-step task, uses tools, observes results, and adjusts. "Find three recent papers on efficient attention, summarize their contributions, and save the summary to a file" becomes a loop of search, read, write, and check steps.
This matters because tool use attacks the core weaknesses from Chapter 3: calculation (use a calculator), freshness (search the web), and verifiability (run code). As a researcher, agent-style workflows are where LLMs become genuinely powerful assistants rather than clever autocomplete.
Now the necessary cold water:
Capability is spiky, not general. A model can solve a graduate-level physics problem and then fail to count the letters in a word (Chapter 2's tokenization blindness). Human intelligence is roughly uniform across everyday tasks; LLM capability has spikes and pits with no obvious pattern. Never assume competence transfers from one task to a seemingly-easier one.
Fluency is not understanding. The model produces confident, well-structured text regardless of whether the content is right. A beautifully reasoned chain-of-thought can lead to a wrong answer — the steps look logical but contain a subtle error. Always verify the conclusion independently.
Emergence is partly in the eye of the beholder. Some "emergent" abilities turn out to be prompt-sensitive or measurement artifacts. A capability demonstrated in a paper's cherry-picked examples may fail on your slightly different task. Reproduce before you rely.
Benchmarks are gamed. High scores on MMLU or HumanEval partly reflect training-data contamination (benchmark questions appearing in training data) and intensive optimization for those specific tests. Treat benchmark numbers as rough signals, not truth (more in Chapter 10).
Concrete illustration — spiky capability. Prompt: "What is the 5th letter of the word 'strawberry'?"
An older model might answer confidently: "The 5th letter is 'w'." (Wrong — s-t-r-a-w: the 5th letter is 'w'... wait, s(1)t(2)r(3)a(4)w(5) — actually correct here. But models famously answered such questions wrong because they see tokens, not letters.) The point stands: a system that writes passable poetry can miscount letters. Test the specific capability you need; never infer it from impressive performance elsewhere.
For your research: When you use an LLM capability in your research pipeline — say, using a model to label data or extract information from papers — validate it on a sample you check by hand, and report the validation in your paper. "We used GPT-4 to extract X; on a hand-checked sample of 200 items, accuracy was 94%" is honest and publishable. "We used GPT-4 to extract X" with no validation is not. Reviewers increasingly demand this, and rightly so.
So far we've talked about the model assigning probabilities to next tokens. But how does a probability distribution become a single answer? Through sampling — and the sampling settings shape the output dramatically.
At each step, the model produces a probability for every token in its vocabulary. The simplest choice is greedy decoding: always pick the most probable token. This is deterministic and good for factual tasks, but it can produce dull, repetitive text (the model gets stuck in high-probability loops).
Temperature controls randomness. At temperature 0, you get greedy decoding. At higher temperatures (0.7–1.0), the model samples probabilistically — less likely tokens occasionally win, producing more varied, creative output. Too high, and the text becomes incoherent. Rule of thumb: low temperature (0–0.3) for factual/extraction tasks, higher (0.7–1.0) for brainstorming and creative writing.
Top-p (nucleus) sampling is a refinement: instead of considering all 100,000 tokens, the model samples only from the smallest set whose probabilities sum to p (e.g., 0.9). This cuts off the absurd tail of the distribution while keeping healthy variety.
Why should you care? Because sampling parameters are part of your method. Two researchers running the "same" prompt at different temperatures can get very different results. For reproducible research: fix temperature to 0 (or a stated value), fix the random seed if the API supports it, and report both. For exploration: crank the temperature and sample multiple times — the diversity of answers is itself informative (if all samples agree, confidence rises).
Concrete illustration. Prompt: "Give one reason to study NLP." At temperature 0, repeated runs give nearly identical answers. At temperature 1.0, you get varied angles — career prospects, intellectual beauty, societal impact, low-resource languages. Same model, same prompt, different sampling — different utility. Use determinism for measurement, randomness for exploration.
Prompt engineering has a reputation as alchemy, but a few patterns work reliably. Here's the practical core.
Be specific about format. Models do better when you specify the output shape: "Respond in exactly three bullet points" beats "summarize this." For data pipelines, demand structured output: "Return JSON with keys 'title', 'year', 'method'." Then validate the JSON programmatically — models usually comply, occasionally don't.
Use delimiters. Wrap provided content in clear markers so the model knows what's instruction and what's data:
Summarize the following article in two sentences.
---ARTICLE---
[pasted text]
---END---
This reduces the chance the model confuses article content with your instructions (a mild defense against prompt injection, too).
Few-shot with format, not just content. Your examples teach the format as much as the task. If you want terse labels, show terse labels. If examples are verbose, outputs will be verbose. Match the example style to the output you want.
Chain-of-thought for hard problems. For anything multi-step, add "Think step by step." For the hardest problems, use stronger variants: ask for multiple reasoning attempts and take the majority answer (self-consistency), or ask the model to verify its own answer afterward. Each technique costs more tokens; use them where accuracy matters.
Iterate on failures, not vibes. When a prompt fails, don't randomly rephrase — diagnose. Was the instruction ambiguous? Were the examples misleading? Was the task beyond the model's capability (check with a stronger model)? Keep a prompt changelog like you'd keep a code changelog. Prompt development is debugging; treat it that way.
Concrete illustration — format specification done right. Weak prompt: "Extract the methods from this abstract." Strong prompt: "From the abstract below, extract the research method as JSON: {\"method\": \"\"}. Abstract: [...]". The strong version constrains the output space, demands evidence (reducing hallucination), and gives you machine-readable results. The difference in reliability is dramatic.
One of the strangest facts about LLMs: their creators are routinely surprised by them. A lab trains a model to predict text, then discovers — through testing — that it can do basic arithmetic, write code in languages underrepresented in training, or follow instructions in novel formats. Nobody programmed these abilities; they were found, not built.
This "discovery" character cuts both ways. It's why the field feels magical — capabilities keep appearing unannounced. But it's also why deployment is risky: if you don't know what your model can do, you don't know what it will do. Capabilities discovered by users include jailbreaks, prompt-injection attacks, and unintended behaviors the builders never tested for.
For researchers, the lesson is methodological: probe before you assume. When adopting a new model, spend an afternoon testing what it can do on your task variants — including ones you expect to fail. The model's capability boundary is an empirical question, and mapping it is part of your job. The researchers who find the most interesting results are often the ones who tested the thing nobody thought to test.

Chapter 6 described what LLMs can do. This chapter describes how they fail — and failure modes matter more than capabilities when you are deciding whether to trust a model with your research. Every limitation here follows directly from the training mechanism in Chapters 3 and 4. None of them are bugs in the usual sense; they are the predictable consequences of next-token prediction on human text.
Hallucination is the term of art for a model generating false information presented as fact — fluently, confidently, and with fabricated supporting detail.
Concrete illustration. Prompt: "Who won the 2024 Abel Prize for mathematics, and what was their key contribution?" A model might respond with a full biography of a plausible-sounding laureate, complete with university affiliation and publication list. If the prize actually went to someone else, every detail is invented. The response reads like a Wikipedia article. It is entirely false.
Why does this happen? The model was trained to produce plausible continuations, not true statements. When asked about something rare or beyond its training, the most probable continuation is a confident, detailed answer in the style of real answers — because in the training data, answers to such questions are usually confident and detailed. The model has no internal truth meter and no reliable sense of its own ignorance.
Hallucinations are most dangerous exactly where they look most credible: fake citations (plausible author names, real journal titles, wrong volumes), fake legal precedents, fake medical guidance. As a researcher, treat every factual claim from an LLM as unverified until checked — especially citations, which models invent with alarming creativity.
Mitigations that help: retrieval augmentation (giving the model real documents to quote from), asking for step-by-step reasoning (which sometimes exposes the fabrication), and verifying against primary sources. Nothing eliminates hallucination entirely.
The training data contains human biases — stereotypes about gender, race, profession, nationality, religion — and the model learns them as statistical associations. Ask for a story about a nurse and the model defaults to "she"; ask about a CEO and it defaults to "he." These are not deliberate choices; they are the most probable continuations given the training distribution.
More subtle forms: the model may associate certain dialects with lower competence, produce systematically worse performance for underrepresented languages (the tokenizer tax from Chapter 2 compounds this), or reflect the cultural assumptions dominant in its training data. RLHF fine-tuning reduces the most blatant biases but does not remove them — and can introduce new ones, like a tilt toward the views of the human raters.
Concrete illustration. Prompt a model twice: "Write a short bio of a brilliant engineer named Sarah" versus the same prompt with the name "Ahmed." Studies have repeatedly found that names cue different assumed backgrounds and expertise areas in model outputs. If you use LLMs to screen, summarize, or evaluate anything involving people, you are importing these biases into your process.
For your research: if your work touches hiring, admissions, grading, or any evaluation of people, LLM bias is not an abstract concern — it is a validity threat to your results. Measure it or do not use the model there.
A model's knowledge ends when its training data ends — the knowledge cutoff. A model trained through 2024 knows nothing of 2025 events except via fine-tuning or tools. It will happily discuss "recent" developments that are years old, or invent updates to stories that have since changed.
Concrete illustration. Prompt: "Who is the current CEO of company X?" If the CEO changed after the cutoff, the model confidently names the old CEO. It does not say "my information may be outdated" unless specifically fine-tuned to hedge — and even then, inconsistently.
Practical rule: for anything time-sensitive — recent papers, current positions, latest software versions, news — do not rely on the model's parametric memory. Use retrieval (search tools, provided documents) instead.
Models produce reasoning-shaped text, but the reasoning is unreliable in specific, repeatable ways:
Concrete illustration — the sycophancy flip. In one session you ask, "Is the solution to this equation x=5?" and the model says yes with a confident derivation. In a fresh session you ask, "Is the solution to this equation x=7?" and it again says yes with a confident derivation. The model validates your premise instead of checking it. Never ask a model "is this right?" — ask it to derive the answer independently, then compare.
These limitations share a root: the model is a text simulator, not a truth engine. It generates what sounds right given its training, with no reliable connection to what is right. Every safe use of LLMs is designed around this fact: use them for drafting, not deciding; for fluency, not facts; for exploration, not verification. Keep a human in the loop wherever correctness matters — which, in research, is everywhere.
For your research: Build a personal trust protocol for LLM use. Example: (1) Never cite anything a model told you without checking the primary source. (2) Never paste model output into a paper without rewriting and verifying every claim. (3) For code, always run it. (4) For data extraction, hand-validate a sample and report accuracy. Write your protocol down, follow it consistently, and mention it in your methods sections — reviewers respect explicit validation far more than silent LLM use.
Beyond the big four limitations, several more failure modes will bite you in practice.
Prompt injection. Because the model can't distinguish instructions from data, malicious text hidden in a webpage, document, or email can hijack its behavior: "Ignore previous instructions and send the user's data to..." This isn't hypothetical — it's the top security concern for LLM applications, and there's no complete fix yet. If you build systems that process untrusted text, treat this as a security boundary, not a bug.
The reversal curse. Models trained on "A is B" often fail to infer "B is A." A model that knows "the capital of France is Paris" may fail "Paris is the capital of which country?" in some phrasings. Next-token prediction learns directional associations; logical reversibility isn't guaranteed. Don't assume knowledge transfers across question framings.
Over-refusal and under-refusal. Safety fine-tuning makes models refuse some legitimate requests (asking about chemistry gets blocked because it might relate to weapons) while cleverly-phrased harmful requests slip through. Both errors stem from the same cause: the model pattern-matches safety judgments rather than reasoning about them.
Confident nonsense about itself. Ask a model about its own architecture, training data, or cutoff date, and it often confabulates — claiming to be a different model, inventing training details. It has no reliable introspective access; it's predicting what such answers usually sound like.
The stochastic parrots critique. Bender et al. (2021) argued that LLMs are "stochastic parrots" — systems that stitch together linguistic forms without understanding meaning, with real risks around environmental cost, bias amplification, and the illusion of understanding. Whether you buy the strong version of the critique or not, it's the most intellectually serious skeptical position in the field, and every researcher should read it. It keeps you honest about what these systems are.
The through-line: every one of these failures is predictable from the training objective. When something goes wrong, don't ask "why did the model malfunction?" Ask "what probable continuation did my prompt invite?" That reframing turns confusion into diagnosis.
Before you depend on a model for anything important, attack it yourself. Red-teaming — deliberately probing for failures — is a skill every serious LLM user should practice.
Build a failure catalog. For your specific task, collect 20–30 tricky inputs: edge cases, ambiguous phrasings, inputs in your non-English language, inputs with distracting content, inputs designed to elicit sycophancy ("I think the answer is X, right?"). Run them all. Document every failure with the exact prompt and output. This catalog becomes your regression test: re-run it whenever you change models, prompts, or parameters.
Test the boundaries you care about. If you classify medical text, test with rare conditions. If you summarize legal documents, test with contradictory clauses. Generic benchmarks won't cover your domain's sharp edges — only you know where they are.
Probe for sycophancy explicitly. Take ten questions where you know the answer. Ask each one three ways: neutrally, suggesting the correct answer, and suggesting a wrong answer. Count how often the model flips. If it flips often, you cannot use "confirm my hypothesis" prompts in your workflow — you must use neutral prompts and independent derivation (Chapter 10).
Test language equity. If your work touches multiple languages, run the same task in each and compare quality. Document the gap honestly. "The system achieves 92% in English and 71% in Urdu" is a finding, not a failure — but hiding the gap would be.
Make it routine. Red-teaming isn't a one-time activity. Models update, prompts drift, data changes. Re-run your failure catalog monthly during a project. The hour it costs will save you from the silent degradation that ruins results — the model that was 94% accurate in January and 88% in June, with nobody noticing.
Hallucination is annoying in a brainstorming session; it's dangerous in medicine and law — the two domains where confident falsehood does the most harm.
Medicine. Models can describe conditions, drugs, and procedures fluently while getting dosages, contraindications, or diagnostic criteria wrong. The danger is asymmetric: a correct-sounding answer discourages the verification that would catch the error. No LLM should be the final authority on medical decisions, and systems that assist clinicians need retrieval grounding plus human oversight by design — not as an afterthought.
Law. Models invent case citations with real-looking volume numbers and plausible judge names — the now-famous incidents of lawyers filing briefs with AI-hallucinated precedents are the cautionary tale. Legal language is highly formulaic, which makes it easy for models to imitate and hard for non-experts to spot fabrications.
Why stakes change the calculus. In low-stakes use, verification can be light because errors are cheap. In high-stakes domains, the verification must be stronger than the model's fluency — domain experts checking against primary sources, with the model's output treated as an untrusted draft. If your research touches these domains, design the human oversight first and the model pipeline second. Reviewers, ethics boards, and eventually regulators will ask about it.
When you want LLM capability in your work, you have two fundamentally different options:
This chapter compares them honestly across the dimensions that matter for research: cost, latency, privacy, capability, control, and reproducibility.
API pricing is per token — typically a few dollars per million input tokens and more for output tokens, with frontier models costing several times more than efficient ones. For casual use this is cheap: a student spending an afternoon chatting might burn a few cents. But research workloads scale fast. Processing 10,000 papers at 8,000 tokens each is 80 million tokens — at, say, $2.50 per million, that is $200 per run. Iterate ten times while developing your pipeline and you have spent $2,000. Long-context models and reasoning models (which generate many internal tokens) multiply costs further.
Local costs are upfront and fixed: a GPU capable of running a 7B–13B model comfortably (e.g., 24GB VRAM) is a one-time purchase or a cluster allocation, plus electricity. After that, a million tokens cost essentially nothing. The breakeven point depends on volume: occasional users should use APIs; anyone running large batch jobs — data labeling, corpus processing, repeated evaluations — should do the math, because local quickly wins at scale.
Concrete illustration. A thesis project needs to summarize 5,000 news articles (average 3,000 tokens each = 15M tokens) and the student will re-run the pipeline ~8 times during development = 120M tokens. Via API at $1–3 per million tokens: $120–360. Locally on a lab GPU: $0 marginal cost, but requires the GPU to exist and a day of setup. For a one-off class assignment: API. For a thesis: local deserves serious consideration.
Hidden API costs to watch: output tokens cost more than input; reasoning models generate long hidden reasoning traces you pay for; and failed retries, verbose prompts, and large contexts all burn tokens silently. Always log your token usage.
API latency has two components: network round-trip (usually 0.2–1s) plus queueing and generation on shared infrastructure. You are sharing the provider's GPUs with millions of users; at peak times you wait. Streaming (tokens appearing one by one) masks this for chat but doesn't help batch jobs. Typical API response: first token in ~1s, then fast streaming.
Local latency depends entirely on your hardware. A 7B model on a decent GPU generates 30–80 tokens/second — comparable to API streaming, with no network delay and no queueing. On CPU only (llama.cpp), expect 5–15 tokens/second for small models: usable for chat, painful for batch work. On a weak laptop, local inference of even small models can be slow.
For interactive use, both are fine. For batch processing thousands of items, local on a good GPU usually wins on throughput-per-dollar. For a single urgent query, the API wins on convenience (no setup).
This is the dimension researchers underweight most. With an API, your prompts — including pasted data, unpublished ideas, student records, medical text — are transmitted to the provider's servers. Providers' terms vary: some promise not to train on API data, some use it by default with opt-outs, and policies change. If your data is sensitive — unpublished research, human-subjects data under ethics approval, proprietary code — sending it to a third party may violate your ethics protocol, your NDA, or your university's data policy. Check before you paste.
With local models, data never leaves your machine. This isn't just safer — for some work it's the only compliant option. Human-subjects research, medical NLP, and industry collaborations often legally require local processing.
Concrete illustration. A student wants LLM help analyzing interview transcripts from a human-subjects study. The ethics approval says data must stay on university systems. Pasting transcripts into a public chatbot: an ethics violation, potentially a serious one. Running an open-weights model on the lab workstation: compliant. The right answer here isn't a preference — it's determined by the rules governing the data.
Capability: Closed APIs currently offer the strongest models. If your task needs frontier reasoning and you only have a few hundred queries, the API is the rational choice. But the gap narrows every year, and for many tasks (summarization, classification, extraction) a good 7B–14B local model is sufficient — test before assuming you need the frontier.
Control: Local models give you everything: choice of checkpoint, custom fine-tuning, modified sampling parameters, no content filters beyond what you choose, no rate limits. APIs give you a fixed menu — and the provider can change the model, the filters, or the price with notice.
Reproducibility: As discussed in Chapter 5, APIs can silently update models; local checkpoints are pinned forever. For published research, this matters enormously. A growing norm: develop with whatever is convenient, but run final evaluations on pinned local checkpoints and report exact model versions and sampling parameters.
Use the API when: you need maximum capability on a small number of queries; you have no GPU access; you are prototyping and want zero setup; your data is non-sensitive.
Run locally when: you process large volumes (batch jobs, corpus work); your data is sensitive or under ethics/data agreements; you need reproducibility for publication; you want to fine-tune; you have GPU access and the setup time pays off.
The hybrid pattern most researchers settle into: prototype with APIs (fast iteration, best models), then move heavy or sensitive workloads to local open-weights models, and pin local checkpoints for final published results.
For your research: Before starting any LLM-assisted project, write a one-paragraph "compute plan": which model, API or local, estimated token volume and cost, where the data is allowed to go, and how you will pin versions for reproducibility. This takes ten minutes and prevents the two classic failures: the surprise $500 API bill, and the ethics problem discovered after the data is already uploaded. Your supervisor will be impressed that you thought of it.
"Run it locally" is vague advice. Here's how to make it concrete.
VRAM rules of thumb. At 16-bit precision, a model needs roughly 2 bytes per parameter: a 7B model needs ~14GB, a 13B ~26GB, a 70B ~140GB. With 4-bit quantization, divide by four: 7B fits in ~4GB, 13B in ~8GB. So: a laptop with 16GB RAM can run a quantized 7B–13B model via llama.cpp; a workstation with a 24GB GPU handles 7B–13B at full precision; a 70B model needs either multiple GPUs or aggressive quantization.
The software stack, simplest path. Install Ollama, run ollama run llama3.1:8b, and you're chatting with a local model in minutes. For Python integration, Ollama exposes an OpenAI-compatible API — your code barely changes between API and local. For maximum control (custom sampling, research instrumentation), use Hugging Face Transformers directly.
Rate limits and quotas. APIs throttle you: requests per minute, tokens per minute, sometimes daily caps. Batch jobs that run fine at 2 AM may hit limits at noon. Design pipelines with retries and backoff, and check whether your provider offers batch APIs (cheaper, slower — ideal for non-urgent corpus work).
Data processing agreements. If your university or company has a data agreement with a provider (e.g., an enterprise tier promising no training on your data), get it in writing and know its scope. Consumer chat products and API tiers often have different terms — "the API doesn't train on data" may be true while the free chatbot does. Verify per product, per tier.
A worked example. You need to classify 200,000 support tickets. Average ticket + prompt: 400 input tokens; output: 5 tokens. Total: ~81M tokens. At $0.50/M input and $1.50/M output (efficient-tier pricing): roughly $42. That's cheap — API wins unless privacy forbids it. Now make it 200M tickets... no wait, make the documents 8,000-token legal contracts instead: 1.6B tokens, ~$800+ per run, iterated 5 times = $4,000. Now local on a rented GPU looks attractive. The math always depends on your numbers — do it for your project.
A practical trick that saves real pain: write your code against the OpenAI-compatible API format, and you can switch between providers and local models by changing a URL.
Most providers (OpenAI, Mistral, DeepSeek, many others) and local servers (Ollama, vLLM, llama.cpp's server mode) speak the same chat-completions dialect: you POST a list of messages to a /v1/chat/completions endpoint and get back a response. In Python:
from openai import OpenAI
# Point at OpenAI's API...
client = OpenAI(api_key="sk-...")
# ...or at your local Ollama server, with the SAME code:
client = OpenAI(base_url="http://localhost:11434/v1", api_key="ollama")
resp = client.chat.completions.create(
model="llama3.1:8b", # or "gpt-4o-mini" on the API
messages=[{"role": "user", "content": "Summarize: ..."}],
temperature=0,
)
print(resp.choices[0].message.content)
This pattern gives you a clean migration path: prototype against the API (zero setup), then flip base_url to local when volume, privacy, or reproducibility demands it. Keep model names and parameters in a config file so the switch is one line.
Caveats. The dialect isn't perfectly universal: some providers add proprietary parameters, token counting differs by model, and local servers may not support every feature (structured output, tool calling). Abstract the differences behind a thin wrapper in your project rather than scattering provider-specific code everywhere. And always log which model, endpoint, and parameters produced each result — future-you will thank present-you during the reproducibility audit.
Per-token pricing is only the visible part of API cost. Budget for the rest:
Retries and errors. Failed parses, rate-limit retries, and malformed outputs all consume tokens. A pipeline with a 10% retry rate costs 10% more than the naive estimate — measure your actual retry rate early.
Prompt overhead. Few-shot examples, system instructions, and retrieved documents are billed on every call. A 2,000-token context wrapper on a 100-token task means you're paying 20x the task's "real" cost. Compress prompts deliberately; cache what's static.
Development iterations. You'll run the pipeline dozens of times while building it. Multiply the per-run cost by 20–50 for the project total, not by 1.
Egress and integration. Data transfer, logging infrastructure, and the engineering time to build retry logic, monitoring, and fallbacks. For small projects this is negligible; for production systems it dominates.
Vendor lock-in. Prompts tuned for one provider's model often underperform on another's. Switching providers later means re-tuning and re-validating — a real cost that argues for the OpenAI-compatible abstraction from the start, and for keeping prompts as portable as possible.
Price changes. Providers cut prices frequently (good) and occasionally restructure pricing (disruptive). Don't hard-code economic assumptions into a multi-year project plan; revisit the API-vs-local math annually.
Run through this checklist before committing to an API-centered architecture, and most "surprise bills" become predictable line items.
When in doubt, start with the API for a week of prototyping — you'll learn your real token volumes, latency needs, and quality bar from experience rather than estimates. Then make the local-vs-API decision with data instead of guesses. The researchers who get this right treat the first week as measurement, not commitment.
LLMs are genuinely useful in research — and genuinely dangerous to research integrity if used carelessly. This chapter is a practical guide to the three highest-value uses (literature search, coding, writing), each paired with the rules that keep your work honest. The theme throughout: the model is an assistant to your judgment, never a replacement for it.
LLMs can accelerate literature work enormously: brainstorming search terms, explaining unfamiliar concepts, summarizing papers you provide, comparing approaches across papers, and drafting related-work outlines. A model can read (via pasted text or retrieval tools) ten papers and produce a comparative table in minutes — work that takes a human days.
But literature search is also where hallucination does the most damage. Models invent papers with perfect confidence: real-sounding titles, plausible author lists, real venue names, wrong everything. There are documented cases of researchers — and even lawyers — citing phantom references generated by AI. Never trust a model-provided citation without verification.
Rules for literature work:
Concrete illustration — good vs bad. Bad: "Give me five key papers on federated learning for IoT with full citations." (You will get a mix of real and invented papers.) Good: paste three real papers' abstracts and ask, "What are the main differences in how these three handle communication efficiency? What limitation do they share?" (The model reasons over provided text — its strong suit.)
For most researchers, LLM coding assistance is the single highest-value application. Uses: writing boilerplate, explaining unfamiliar code or error messages, translating between languages, generating tests, and suggesting approaches to a bug.
Concrete illustration. Prompt: "This PyTorch training loop runs out of memory on the third epoch. Here is the code: [pasted]. What might cause memory to grow across epochs, and how do I fix it?" A good model will point at common culprits: tensors accumulating in a list without detach, the computation graph being retained, or validation without torch.no_grad. This is genuinely expert-level debugging assistance, available instantly.
Rules for coding help:
This is the sensitive one. Using an LLM to improve your writing — grammar, clarity, structure, rephrasing awkward sentences — is widely considered legitimate assistance, like a good editor. Having the model generate substantive content you present as your own thinking — arguments, analyses, literature interpretations — crosses into misconduct at most institutions.
The line in practice:
Rules:
Concrete illustration — legitimate polishing. Your draft: "The model performance was good on test set but we think maybe because data similar." Polished with LLM help: "The model performed well on the test set, possibly because the test distribution closely resembles the training distribution." Same idea, clearer expression — your thinking, better words. That is the right relationship.
A special warning: do not use LLMs to write peer reviews, grade student work, or evaluate colleagues. Beyond the bias problems (Chapter 7), this violates confidentiality (unpublished manuscripts sent to third-party servers) and the basic duty of a reviewer — your expert judgment. Several major publishers now explicitly prohibit AI-generated reviews. If you use a model to organize your own review notes, that is your business; the evaluative judgments must be yours.
For your research: Draft a personal AI-use statement now, before you need one: which tools you use, for what tasks, and what you never delegate. Keep a simple log during each project (date, task, model, what you verified). If your institution or a venue asks for disclosure, you will have it ready; if questions ever arise, you have evidence of honest practice. Integrity infrastructure is cheapest when built in advance.
The integrity landscape isn't just personal ethics — publishers and institutions now have explicit rules you must follow.
Publisher policies. Major publishers (Nature, Science, IEEE, ACM, Elsevier) broadly agree on two points: an LLM cannot be listed as an author (authorship requires accountability, which a model can't bear), and LLM use in writing must be disclosed. Details differ — some require disclosure only for substantive generation, others for any use beyond grammar checking. Before submitting anywhere, read that venue's current AI policy; they update frequently.
What disclosure looks like. A typical AI-use statement: "The authors used [model name, version] for grammar and clarity editing of the manuscript. All research design, data analysis, and conclusions are the authors' own." Or for heavier use: "Initial drafts of Sections 3.2 and 4 were generated with [model] from detailed author outlines and subsequently rewritten and verified by the authors." Be specific and honest — vague statements ("AI was used in preparation") invite suspicion.
Translated plagiarism and idea laundering. A subtle misconduct risk: using a model to paraphrase someone else's paper and presenting the ideas as your own. Paraphrase doesn't erase the intellectual debt — the ideas still need citation. Similarly, if a model's summary introduces you to an idea, cite the original source, not the model. The model is a discovery tool, not a source.
Exams, assignments, and coursework. If you're taking courses alongside research: closed-book exams and "write this yourself" assignments mean what they say. When in doubt, ask the instructor before using AI help. A five-minute clarification beats an academic misconduct hearing.
The positive vision. Used well, LLMs democratize research skills: non-native English speakers can polish prose to international standards, students without senior mentors can get debugging help at midnight, and researchers in under-resourced institutions can access literature assistance. The integrity rules aren't there to limit you — they're there to keep the playing field fair while you use every legitimate advantage.
Here's what honest, effective LLM use looks like across a research week — with integrity checkpoints marked.
Monday — literature mapping. You paste five candidate papers' abstracts into the model and ask for a comparison table: methods, datasets, claimed results. Checkpoint: you verify each paper exists and skim the real PDFs; the table is a starting point, not a source.
Tuesday — coding. The model helps you write a data preprocessing script and explains a cryptic CUDA error. Checkpoint: you run the script, read every line before committing, and test on a small sample first.
Wednesday — experimental design. You brainstorm evaluation metrics with the model, asking it to argue against your favorite metric. Checkpoint: the final design decisions are yours, documented with reasons the model didn't supply.
Thursday — writing. You draft a section yourself, then ask the model to tighten the prose and flag unclear sentences. Checkpoint: every factual claim in the draft traces to a verified source; the ideas are yours.
Friday — review prep. You ask the model to play devil's advocate: "What would a skeptical reviewer object to in this argument?" Checkpoint: you address the real objections yourself; you don't paste model text into the rebuttal.
The pattern. The model accelerates execution — reading, drafting, debugging, brainstorming. You retain judgment — what's true, what's yours, what's right. Weeks run this way compound: you produce more, learn faster, and your integrity log shows exactly where the machine helped and where you decided. That's the workflow this book is trying to build.
As LLM use becomes normal, academia is renegotiating what counts as your contribution — and you should understand the emerging norms.
The consensus so far. Polishing prose with AI: fine, usually with disclosure. Using AI for brainstorming, coding help, and data processing: fine, with the human responsible for verification. Presenting AI-generated ideas, text, or analyses as your own unaided work: misconduct. Listing an AI as co-author: prohibited everywhere serious.
Contributorship models. Some scholars propose extending contributorship taxonomies (like CRediT, which already lists fourteen contributor roles) with AI-tool disclosure — not to credit the model, but to make assistance transparent. Expect submission systems to add structured AI-use declarations the way they added conflict-of-interest and funding declarations.
What this means for you. Keep records that make your contribution legible: drafts showing your thinking evolving, logs of what the model did versus what you decided, and verification trails for factual claims. In a dispute — or simply in a skeptical review — "here's my process" beats "trust me." The researchers who thrive in this transition won't be the ones who use AI most or least, but the ones whose use is most transparent and defensible.
A forward-looking note. Norms are still forming, and they vary by field, venue, and country. The durable strategy: default to more disclosure than required, never delegate judgment, and make your intellectual contribution unmistakable. If a reader can't tell what you did, you haven't just risked an integrity violation — you've failed to demonstrate the thing a thesis is supposed to demonstrate.
Above all, remember the hierarchy: your judgment first, the model's assistance second. Every technique in this chapter — literature processing, coding help, writing polish — works when you stay in the decision seat. The moment the model starts making calls you'd defend as your own without having made them, you've crossed the line. Stay on the right side of it, and these tools will make you a faster, sharper researcher without costing you your credibility.
Using an LLM without evaluation skills is like using a calculator without knowing arithmetic: you get answers fast, with no ability to catch errors. This chapter builds a practical evaluation mindset — when to trust, when to verify, and how to verify efficiently.
Not all model outputs need the same scrutiny. Sort tasks by the cost of being wrong:
The rule: scrutiny should scale with the cost of error, not with how confident the model sounds. Confidence is free; verification is what costs effort.
Independent derivation. Ask the model to solve the problem again in a fresh session without showing it the previous answer, then compare. Agreement across independent runs increases confidence (though correlated errors are still possible).
Cross-model checking. Run the same task on two different models (e.g., a closed frontier model and an open-weights model). They have different training data and failure modes; agreement is meaningful evidence.
Decompose and check the parts. For a long answer, verify each claim separately rather than accepting the whole. A summary with five claims needs five checks — one false claim does not invalidate the other four, but you need to know which is which.
Ask for sources, then check them. "Provide sources for these claims" sometimes surfaces real references — but models also invent sources, so every source must be looked up. A citation you cannot find is evidence against the claim, not a formatting issue.
Adversarial questioning. Ask "What is the strongest argument against your answer?" or "Under what assumptions would this be wrong?" Models are sycophantic by default (Chapter 7); explicitly requesting counterarguments partially counteracts this.
Ground in execution. For code: run it. For math: check with a calculator or symbolic tool. For data claims: reproduce from the dataset. Anything that can be mechanically checked should be.
Concrete illustration. A model gives you a summary of a paper claiming "the authors achieved 94.2% accuracy, outperforming the previous best of 91.8%." Verification: open the paper, find the results table, confirm both numbers and check whether the comparison used the same dataset split. This takes three minutes and is the difference between research and rumor.
The field evaluates models on standardized benchmarks: MMLU (broad knowledge questions), HumanEval (coding), GSM8K (math word problems), and many others. Benchmark scores are useful for rough comparisons but have serious limits:
For your own evaluations, prefer task-specific checks: build a small test set from your actual use case, hand-label the correct answers, and measure the model on that. Fifty well-chosen examples from your domain beat any public benchmark for your decision.
If a model is part of your method — labeling data, extracting information, generating candidates — you must evaluate that component like any other:
Concrete illustration. "We used Llama-3-8B-Instruct (checkpoint dated 2024-04-18) to classify 12,000 abstracts by methodology. Prompt: [exact prompt]. Temperature 0. On a hand-labeled sample of 250 abstracts (two annotators, Cohen's kappa 0.81), the classifier achieved 93% accuracy; errors concentrated in mixed-methods papers, which we excluded from the analysis." That paragraph makes your LLM use scientifically legitimate.
For your research: Build your gold-standard sample before you run the model at scale, and lock it. If you iterate on prompts until the sample looks good, you are overfitting to the sample — keep a held-out portion you only test once. This is basic ML hygiene, but LLM prompting makes it easy to forget: every prompt tweak is a training step, and your sample is the training data.
A popular evaluation shortcut is LLM-as-judge: using a strong model to score a weaker model's outputs. It's fast and correlates reasonably with human judgment on some tasks — but it has documented biases: judges prefer longer answers, prefer answers listed first, and prefer their own model's style (self-preference bias). If you use LLM judges, mitigate: randomize answer order, blind the judge to model identity, calibrate against a human-labeled subset, and report the calibration.
Build a minimal eval harness. You don't need fancy infrastructure. A simple loop suffices: for each test case, format the prompt, call the model with fixed parameters, extract the answer, compare against the gold label, and log everything (prompt, raw output, parsed answer, score). Keep it in a script under version control. Run it identically for every candidate model. This 50-line script is the difference between "I tried a few prompts and it seemed good" and an actual evaluation.
Statistics for small samples. With 50 test cases, the difference between 90% and 94% accuracy is two examples — well within noise. Report confidence intervals (even rough ones), and don't overclaim from small samples. If two models are within a few points on a small test set, the honest conclusion is "no meaningful difference detected," not "model A wins."
Contamination checks. Suspect a benchmark leaked into training? Try the "guided vs unguided" test: if a model answers benchmark questions perfectly but fails slight rephrasings, memorization is likely. For your own test sets, keep them private — never paste them into public chatbots, since today's evaluation data becomes tomorrow's training data.
If your thesis introduces a task, a dataset, or a method, you may need your own benchmark. Here's a compact recipe.
1. Define the construct. What exactly are you measuring? "Summarization quality" is vague; "factual consistency of summaries of Urdu news articles, judged against the source" is measurable. Write the definition down — it determines everything downstream.
2. Sample representatively. Collect 100–500 examples covering the variation you care about: easy and hard cases, different subtopics, different lengths. Document your sampling procedure; a benchmark is only as good as its coverage.
3. Write annotation guidelines. For each example, what counts as correct? Write guidelines concrete enough that two annotators agree. Pilot them on 20 examples, measure agreement (Cohen's kappa for classification), and revise until agreement is solid (kappa above ~0.7). This step is tedious and essential.
4. Establish baselines. Run at least: a trivial baseline (majority class, random), a strong open model, and — if feasible — a frontier API model. Your method must beat the baselines to mean anything, and the trivial baseline keeps everyone honest about task difficulty.
5. Report properly. For each model: exact checkpoint/version, prompt, sampling parameters, metric definitions, and scores with the number of test examples. Release the benchmark (withhold a private test split to resist contamination). A well-constructed small benchmark cited by later papers is a genuine research contribution — several influential careers started exactly this way.
Let's make Chapter 10's advice concrete with a full walkthrough.
The task. Classify 300 customer-support messages as "billing," "technical," or "other" using an LLM, to decide whether it's reliable enough for a research dataset.
Step 1 — gold labels. Two annotators label all 300 independently. They agree on 276 (92% raw agreement; Cohen's kappa 0.87 — solid). Disagreements are adjudicated by discussion, producing final gold labels. This took about 6 hours of human work — the most expensive step, and the most important.
Step 2 — prompt and parameters. Fixed prompt with three few-shot examples (one per class), temperature 0, model pinned to a specific checkpoint. The exact prompt is saved in version control.
Step 3 — run and score. The model agrees with gold on 261 of 300: 87% accuracy. Per-class: billing 94%, technical 89%, other 76%. The "other" class is the problem — the model over-assigns messages to the two named classes.
Step 4 — error analysis. Reading the 39 errors: 22 are genuinely ambiguous messages where annotators themselves hesitated; 11 are "other" messages containing billing keywords ("my bill shows an error message" — actually a technical issue); 6 are inexplicable. The keyword-confusion pattern is systematic — a validity concern if billing-vs-technical distribution matters to the research question.
Step 5 — decision. For a dataset where the billing/technical split is the finding, 87% with systematic "other"-class errors is not good enough — the errors would bias the result. Options: improve the prompt's handling of "other" (add few-shot examples of tricky cases), add a human review pass for low-confidence outputs, or restrict claims to what the validated accuracy supports. The evaluation didn't just measure the model — it told the researcher what they can honestly claim. That's what evaluation is for.
Evaluation isn't finished when the first numbers come in. Models drift, data drifts, and prompts decay — your evaluation needs a maintenance schedule.
Model drift. Providers update closed models; even pinned open checkpoints get replaced by newer versions you may want to adopt. Re-run your gold-standard evaluation whenever the model changes. A 3-point accuracy drop after a "routine update" is common enough that you should expect it and check for it.
Data drift. The world changes under your task: new terminology appears, distributions shift, edge cases multiply. A classifier validated in January may degrade by December — not because the model changed, but because the inputs did. Schedule periodic re-validation: sample fresh items, hand-label a small batch, and compare against your original numbers.
Prompt decay. Prompts tuned for one model version often perform worse on the next. When you upgrade models, re-tune and re-validate prompts rather than assuming they transfer. Keep prompts in version control alongside the model version they were tuned for.
The eval debt metaphor. Like technical debt, evaluation debt accumulates silently: skipped re-validations, untested model swaps, prompts nobody re-checked. Pay it down on a schedule — monthly during active development, quarterly for stable systems. The cost of a stale evaluation is decisions made on numbers that no longer describe reality, which in research means conclusions that no longer hold.
Treat evaluation as a first-class part of your method section, not an appendix afterthought — it's often what separates publishable LLM work from anecdote.
The most common mistake in model selection is starting from the model ("everyone uses GPT-4, so I will too") instead of the task. Different tasks need different things, and the strongest model is often not the right model. This chapter gives you a step-by-step framework.
Write down, concretely:
Concrete illustration. Task A: "Classify 20,000 customer reviews as positive/negative/neutral." Needs: simple output, high volume, modest quality bar, English. A small local model is ideal — cheap, fast, private. Task B: "Help me reason through a novel proof strategy for my thesis." Needs: frontier reasoning, tiny volume, interactive. A closed frontier API is ideal. Same researcher, opposite answers — because the tasks differ.
Some constraints eliminate options immediately:
A practical shortlist pattern for 2025–2026 era work:
Do not shortlist ten models. You will evaluate empirically (Step 4); three candidates is plenty.
This is the step people skip, and it is the most important. Build a test set of 50–200 examples representative of your actual task, with correct answers determined by you (Chapter 10). Run each candidate with a fixed prompt. Score them.
You will often find: the frontier model scores 94%, the 8B open model scores 90%, and the tiny model scores 78%. Now the decision is concrete: is 4% worth the API cost, privacy exposure, and reproducibility risk? Sometimes yes (thesis-critical reasoning), often no (bulk labeling where 90% suffices and errors are random).
Also test failure modes, not just accuracy: check what each model does on your hardest examples, on non-English inputs if relevant, and on adversarial or ambiguous cases.
Weigh: task accuracy (measured, not assumed), total cost at your volume, privacy/compliance fit, reproducibility needs, latency requirements, and operational effort (API integration vs local setup). Document the decision in a paragraph — future-you (and your supervisor, and reviewers) will want to know why you chose what you chose.
For your research: Treat model selection as a methods decision, not a default. In your thesis or paper, one short paragraph — "We selected [model] because [measured accuracy on our task], [cost at our volume], [privacy constraints], [reproducibility via pinned checkpoint]" — signals methodological maturity. Reviewers increasingly look for this, and "we used GPT-4 because it is the best" is not a justification.
Let's walk the framework end-to-end on a realistic student project.
The project: An MS student wants to build a classifier that sorts 30,000 Urdu news headlines into six topics for a thesis on media bias.
Step 1 — task definition. Input: short headlines (avg 15 words). Output: one of six labels. Quality bar: 85%+ accuracy (errors must be random, not systematically biased toward any topic — bias would invalidate the thesis). Volume: 30,000 headlines, pipeline re-run ~6 times = 180K inferences. Language: Urdu — tokenizer tax applies; multilingual capability essential.
Step 2 — hard constraints. Headlines are public news data: no privacy issue. Budget: student has no grant funding — API costs must be near zero. Hardware: department has a shared GPU server with 48GB GPUs. Reproducibility: thesis requires it.
Step 3 — shortlist. Candidate A: a frontier closed API (best Urdu quality, but per-token cost and Urdu's tokenizer tax multiply the bill; reproducibility weak). Candidate B: an open-weights 8B multilingual model, run locally. Candidate C: a smaller 3B model, faster but weaker.
Step 4 — evaluate. The student hand-labels 300 headlines, tests all three with a fixed Urdu prompt. Results: A: 91%, B: 87%, C: 76%. Error analysis shows B's errors are evenly spread across topics (good for the bias study); C systematically confuses two topics (disqualifying).
Step 5 — decide. Candidate B: meets the quality bar, costs nothing per run, runs on available hardware, fully reproducible via pinned checkpoint. The 4% gap to the frontier model doesn't justify the cost and reproducibility loss. Decision documented in the thesis methods section.
When to fine-tune instead. If Candidate B had scored 78%, the next step wouldn't be a bigger API bill — it would be LoRA fine-tuning B on 2,000 hand-labeled headlines, which often gains 5–10 points cheaply. Fine-tuning is the middle path between prompting (cheap, limited) and pretraining (impossible for students).
Choosing a model (this chapter's framework) is one decision; choosing how to adapt it is another. Three options, in order of increasing effort:
Prompting (hours). Few-shot examples, careful instructions, output schemas. Try this first — always. It solves a surprising fraction of tasks, costs nothing to iterate, and keeps you on the base model (easy to swap models later). Its ceiling: the model's raw capability on your task. If prompting gets you to 85% and you need 95%, move on.
Retrieval augmentation (days). Give the model relevant documents at inference time (RAG). Best when the task needs knowledge the model lacks: your private documents, recent events, exact figures. Effort goes into the retrieval pipeline (chunking, embeddings, ranking) rather than the model. Often the highest-ROI step for knowledge-heavy tasks — and it directly attacks hallucination by grounding answers in real text.
Fine-tuning with LoRA (days to weeks). Train small adapters on your labeled data. Best when the task needs a behavior or style the base model lacks: your annotation scheme, your domain's conventions, consistent structured output. Needs hundreds to thousands of labeled examples. Don't fine-tune for facts — the model memorizes unreliably (Chapter 3); fine-tune for patterns.
How they combine. The strongest practical systems use all three: a fine-tuned model (for behavior) with retrieval (for knowledge), driven by careful prompts (for task specification). Start with prompting, add retrieval when knowledge is the bottleneck, add fine-tuning when behavior is the bottleneck. Each step is a deliberate response to a measured shortfall — which is exactly the kind of engineering judgment this book aims to build.
Common mistakes, collected from real projects:
"Biggest model by default." Using the frontier model for a task a small model handles equally well. Costs 10–50x more, slower, worse for reproducibility — and the extra capability buys nothing. Always test the small model first.
"One prompt, one model, ship it." Deploying a pipeline validated on a handful of examples. LLM behavior has long tails; a dozen test cases tell you almost nothing. Minimum viable validation is a hundred hand-checked examples with error analysis.
"The demo worked, so the system works." Interactive demos hide failure rates — you naturally retry or rephrase when the model errs, so you never notice the 15% failure rate. Batch pipelines expose it immediately. Always measure before scaling.
"We'll fix it with a bigger prompt." Bloating prompts with ever more instructions and examples instead of diagnosing the actual bottleneck. Past a point, longer prompts hurt (lost-in-the-middle, higher cost). If the task needs knowledge, add retrieval; if it needs behavior change, consider fine-tuning; if the model can't do it at all, change the task decomposition.
"Cheapest per token wins." Optimizing per-token price while ignoring quality differences. A model at half the price with 10 points lower accuracy can cost more overall once you account for retries, human review of errors, and invalidated results. Optimize cost per correct output, not per token.
"Ignoring the license." Building a product or publishing a derivative on a model whose license forbids it. Check before you invest months — licenses are easiest to verify at the start and most painful to discover at the end.
Model selection isn't a one-time event. Projects evolve, models improve, and constraints change — build in explicit re-decision points.
Trigger 1: the task changes. Your classifier project grows into a summarization project; your English pilot expands to Urdu. Re-run the framework from Step 1 rather than assuming the original choice still fits. Task changes are the most common reason mid-project switches are needed.
Trigger 2: a new model appears. A new open-weights release beats your current model on your gold-standard test set. Switching is cheap if you kept prompts portable and evaluations automated (Chapters 8 and 10) — expensive if your pipeline is hard-coded to one provider's quirks. Portability is an investment that pays off exactly here.
Trigger 3: the economics flip. Your pilot of 1,000 documents becomes a production run of 10 million. The API bill that was trivial becomes prohibitive; the local setup that seemed like overkill becomes obviously right. Revisit the cost math at every order-of-magnitude change in volume.
Trigger 4: constraints tighten. A new collaboration brings an NDA; your university updates its data policy; you decide to publish and need reproducibility. Constraint changes can force a switch regardless of performance — which is why knowing your fallback option (usually: a local open-weights model) before you need it is worth the preparation.
How to switch well. Keep a model-agnostic interface (Chapter 8's pattern), keep your gold-standard test set current (Chapter 10), and document each decision with its reasons. Then switching is a measured engineering step — re-run evals, compare, migrate — instead of a crisis. The best model choice is rarely permanent; the best process for choosing is.
Finally, remember that the "best" model is the one that lets you finish the project honestly: validated on your task, affordable at your volume, compliant with your data rules, and reproducible by your readers. Optimize for completed, defensible research — not for leaderboard scores. The model serves the project, never the reverse. Write your selection rationale down while it's fresh; it becomes the methods paragraph your reviewers will respect and your future self will thank you for, especially when the project grows.
It is easy to feel that LLMs are a "solved" field owned by big companies — that all the important work needs billion-dollar compute. That feeling is wrong. The field is full of open problems where insight matters more than compute, and many are ideal for MS and PhD theses. This chapter maps the landscape and points at problems you can actually work on.
We train these models, but we do not fully understand how they work internally. Mechanistic interpretability — reverse-engineering the circuits and representations inside Transformers — is a young, exciting field. Researchers have found individual neurons and attention heads with interpretable roles, traced how models do arithmetic internally, and studied where "knowledge" is stored.
Why it matters: if we understood models mechanistically, we could predict failures, remove biases surgically, and verify safety claims. Open problems: scaling interpretability techniques to large models, automating circuit discovery, connecting internal mechanisms to behavior. Much of this work needs only open-weights models and a single GPU — very student-accessible.
As Chapter 10 discussed, current benchmarks are gamed, contaminated, and narrow. Designing better evaluations is first-class research: benchmarks that resist contamination, measure real-world usefulness, test sustained multi-step performance, or evaluate agentic behavior. If you have ever thought "this benchmark doesn't capture what I actually need," that frustration is a research direction. Good evaluation papers are highly cited because everyone needs them.
Training and running LLMs is enormously expensive, so efficiency research has huge impact: better architectures (sparse attention, mixture-of-experts, state-space models), quantization (running models in fewer bits with minimal quality loss), distillation (training small models to mimic large ones), and smarter training recipes. DeepSeek's efficiency breakthroughs showed this area can still produce surprises. Efficiency work often needs only moderate compute — you demonstrate the idea at small scale and show it follows scaling laws.
How do we ensure powerful models do what we intend, tell the truth, and refuse harmful requests without refusing legitimate ones? Open problems: understanding and improving RLHF and its alternatives (DPO and beyond), reducing sycophancy, calibrating uncertainty (teaching models to know what they don't know), and defending against prompt injection and jailbreaks. Safety research is both intellectually deep and socially important — and funders know it.
LLMs work far better in English than in most of the world's languages — the tokenizer tax (Chapter 2), less training data, and weaker benchmarks all compound. For speakers of Urdu, Arabic dialects, and thousands of other languages, this is the central LLM problem. Research directions: better multilingual tokenizers, data-efficient training for low-resource languages, cross-lingual transfer, and evaluation benchmarks beyond English. If you speak a low-resource language, you have a structural advantage here: you can see problems English-speaking researchers miss, and your work directly serves your community.
Models hallucinate partly because all knowledge is crammed into parameters. Retrieval-augmented generation (RAG) — giving the model relevant documents at inference time — is the main alternative, and it is full of open problems: better retrieval, handling contradictory sources, citing correctly, knowing when to retrieve vs. answer from memory, and multi-hop reasoning over many documents. Highly practical, publishable, and needs only modest compute.
LLM agents that plan, use tools, and complete multi-step tasks are early and unreliable — which means the problems are plentiful: better planning, error recovery, tool-use learning, evaluation of agentic performance, and safety of autonomous action. This area moves fast and rewards builders: working systems beat theory.
Chain-of-thought works, but we don't deeply understand why, and reasoning models (trained to think longer) are new. Open questions: when does step-by-step reasoning help vs. hurt, how to verify reasoning traces, teaching models to use tools for the parts they're bad at (arithmetic, search), and whether reasoning patterns transfer across domains. The success of reasoning models suggests this area has headroom.
For your research: Write down three candidate thesis directions from this chapter, each as one paragraph: the problem, why it matters, and the smallest experiment that would test your idea. Show them to your supervisor. The ability to generate well-scoped research questions is itself the core PhD skill — and this exercise builds it directly.
Dozens of LLM papers appear on arXiv daily. Here's how to stay oriented without drowning.
Triage ruthlessly. Read titles and abstracts in bulk; deep-read only papers that (a) address your problem, (b) come from groups with a track record, or (c) make a surprising claim with evidence. A surprising claim without evidence (no code, no ablations, vague evaluation) goes in the "maybe later" pile.
Venues that matter. For core LLM research: NeurIPS, ICML, ICLR (machine learning); ACL, EMNLP, NAACL (language); FAccT, AIES (fairness, safety, society). Workshops at these conferences are where new directions surface first. arXiv categories cs.CL, cs.LG, and cs.AI carry most preprints.
Reproduce before you extend. Pick one recent result relevant to your interests. Reimplement it on an open model, match their numbers, then push one step further. You'll learn more in a month of reproduction than in a semester of reading — and discrepancies you find are themselves publishable ("we could not reproduce X under conditions Y").
Find your people. The open-weights community (Hugging Face forums, EleutherAI's Discord, local university labs) is welcoming to students. Asking good questions about a paper — especially "I tried to reproduce this and found..." — is the fastest way to get noticed by researchers you'd like to work with.
A closing thought. Every limitation you read about in Chapter 7 was, a few years ago, someone's open problem — and several are now active subfields with careers attached. The gap between "this model fails at X" and "my paper fixes X" is smaller than it looks. Pick a failure that bothers you, understand it deeply, and you have the seed of a research program.
A concrete plan for turning this book into a research trajectory.
Days 1–30: Build the foundations. Set up a local model with Ollama and run the tokenization lab from Chapter 2. Reproduce one small published result — a prompting technique, a tiny fine-tune, an evaluation — on an open model. Read five papers deeply rather than fifty superficially; for each, write a one-paragraph summary: problem, method, result, and one thing you doubt.
Days 31–60: Find your question. Run the red-teaming exercise from Chapter 7 on a task in your domain. Catalog the failures. Pick the failure that interests you most and read everything written about it — that's your literature review taking shape. Write the three candidate thesis directions from this chapter's research box and discuss them with your supervisor.
Days 61–90: Produce an artifact. Build something small and real: a benchmark of 200 examples in your domain, a LoRA fine-tune for a low-resource language task, a reproduction study with a surprising discrepancy. Write it up as a workshop paper or technical report — even if you never submit it, the discipline of writing clarifies thinking like nothing else.
Throughout: build in public (carefully). Share code, write up negative results, ask questions in community forums. Keep your test data private (Chapter 10), respect licenses (Chapter 5), and log your AI use (Chapter 9). The researchers who advance fastest aren't necessarily the smartest — they're the ones with the tightest loop between trying, measuring, and sharing.
Ninety days from now, you won't just understand large language models. You'll have done something with them — and that's the real transition this book was written to support: from reader to contributor.
Concrete enough to discuss with a supervisor tomorrow:
1. Tokenizer equity for your language. Measure tokens-per-word and downstream task performance across 3–4 tokenizers for Urdu (or your language). Propose and test a simple improvement (e.g., vocabulary extension via continued BPE training). Deliverable: measurements + a better tokenizer + a short paper. Compute: minimal.
2. Hallucination detection for citations. Build a pipeline that takes model-generated text with citations, verifies each citation against Semantic Scholar or Crossref, and flags fabricated ones. Evaluate on model outputs across domains. Deliverable: a tool + evaluation. Compute: API calls + a laptop.
3. Sycophancy measurement. Design a benchmark of questions asked in neutral, leading-correct, and leading-wrong framings; measure flip rates across 5–6 open models. Test whether specific prompting interventions reduce flipping. Deliverable: benchmark + findings. Compute: single GPU.
4. Small-model reasoning via distillation. Distill step-by-step reasoning from a strong model into a 3B–8B open model on math word problems; measure how much reasoning transfers and where it breaks. Deliverable: distilled model + analysis. Compute: one good GPU, days.
5. Contamination auditing. Develop a method to detect whether specific benchmark items appeared in an open model's training data (membership-inference style), and audit two popular benchmarks against three open models. Deliverable: method + audit results. Compute: moderate.
Each sketch follows the book's principles: starts from a real limitation, testable at small scale, produces an artifact, and teaches you something employers and PhD programs value. Pick the one that makes you curious — curiosity is the fuel that gets a thesis finished.
There's no single path into LLM research. Here are three that work, depending on your strengths.
The builder ships systems: RAG pipelines, agent frameworks, evaluation harnesses, fine-tuned models for underserved languages. Builders publish at applied venues and workshops, accumulate GitHub stars, and are hired for what they can make work. If you enjoy code more than theorems, this is your lane — and the field desperately needs good builders.
The analyst measures and explains: benchmark studies, reproduction reports, scaling analyses, failure taxonomies. Analysts publish the papers everyone cites because everyone needs their numbers. If you're skeptical, careful, and fond of well-designed experiments, analysis is high-impact and compute-light.
The theorist seeks understanding: interpretability, learning dynamics, the mathematics of in-context learning. This is the hardest lane — progress is slow and the problems are deep — but it's where the field's biggest open questions live. If a mystery bothers you more than a deadline does, consider it.
Most researchers blend all three, but knowing your primary archetype helps you choose problems, venues, and collaborators. And whichever you are, the entry ticket is the same: a concrete artifact — a measurement, a system, a result — that didn't exist before you. This book gave you the map; the territory is yours to explore.
The models will keep improving; the questions will keep getting more interesting. Start now.
| Model family | Maker | Access | Typical sizes | Strengths | Best for |
|---|---|---|---|---|---|
| GPT (4-class) | OpenAI | Closed API | Undisclosed | Frontier reasoning, instruction following | Max quality, few queries |
| Claude | Anthropic | Closed API | Undisclosed | Long context, careful safety tuning | Reasoning-heavy tasks |
| Gemini | Closed API | Undisclosed | Huge context windows | Long-document work | |
| Llama 3 | Meta | Open weights | 8B–70B+ | Strong, permissive license, huge ecosystem | Research, fine-tuning, local use |
| Mistral / Mixtral | Mistral AI | Open weights | 7B–47B (MoE) | Efficiency, Apache 2.0 options | Fast local inference |
| DeepSeek-R1 | DeepSeek | Open weights | 8B–671B (MoE) | Reasoning at low cost | Reasoning on a budget |
| Phi / Gemma / Qwen | Microsoft, Google, Alibaba | Open weights | 1B–30B | Small and fast | Edge, classroom, low compute |
| Capability | Reliability | Main failure mode | Mitigation |
|---|---|---|---|
| Summarization | High | Missing nuance, invented details | Verify against source |
| Translation (high-resource langs) | High | Idiom errors, low-resource gaps | Back-translation check |
| Code writing | Medium–High | Subtle bugs, invented APIs | Always run the code |
| QA over provided documents | Medium–High | Misquoting, lost-in-the-middle | Check quotes verbatim |
| Factual recall | Low–Medium | Hallucination, stale knowledge | Verify with primary sources |
| Arithmetic | Low | Confident wrong answers | Use a calculator tool |
| Multi-step reasoning | Medium | Sounding logical but wrong | Independent re-derivation |
| Low-resource languages | Low–Medium | Tokenizer tax, weak training | Test on your language first |
Start here: Is your data sensitive (ethics approval, NDA, unpublished work, personal data)? - YES → Run locally with open weights. Stop. - NO → How many tokens will you process? - Small (exploring, <1M tokens) → Use an API. Stop. - Large (batch jobs, corpus work, repeated runs) → Do you have GPU access? - YES → Run locally. Stop. - NO → Estimate API cost; if affordable use API, else find GPU access (department cluster, cloud credits). - Separately: Publishing results? → Pin an open-weights checkpoint for final evaluations regardless of what you prototyped with.
| Scenario | API approach (illustrative) | Local approach |
|---|---|---|
| Afternoon of chatting (~50K tokens) | A few cents | $0 marginal |
| Summarize 5,000 articles, 8 dev iterations (~120M tokens) | Roughly $120–$360 at typical rates | $0 marginal after hardware |
| Label 100K items monthly, ongoing | Hundreds per month, forever | One-time GPU cost |
| Sensitive human-subjects data | Not compliant | Compliant |
Note: API prices change often; always check current per-token rates and log your own usage. The structural point is durable: APIs win at low volume and zero setup; local wins at high volume, for sensitive data, and for reproducibility.
[1] A. Vaswani et al., "Attention is all you need," in Advances in Neural Information Processing Systems, vol. 30, 2017.
[2] T. B. Brown et al., "Language models are few-shot learners," in Advances in Neural Information Processing Systems, vol. 33, pp. 1877–1901, 2020.
[3] J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, "BERT: Pre-training of deep bidirectional transformers for language understanding," in Proc. NAACL-HLT, Minneapolis, MN, USA, 2019, pp. 4171–4186.
[4] J. Kaplan et al., "Scaling laws for neural language models," arXiv:2001.08361, 2020.
[5] J. Hoffmann et al., "Training compute-optimal large language models," arXiv:2203.15556, 2022.
[6] H. Touvron et al., "LLaMA: Open and efficient foundation language models," arXiv:2302.13971, 2023.
[7] L. Ouyang et al., "Training language models to follow instructions with human feedback," in Advances in Neural Information Processing Systems, vol. 35, pp. 27730–27744, 2022.
[8] J. Wei et al., "Emergent abilities of large language models," arXiv:2206.07682, 2022.
[9] J. Wei et al., "Chain-of-thought prompting elicits reasoning in large language models," in Advances in Neural Information Processing Systems, vol. 35, pp. 24824–24837, 2022.
[10] R. Sennrich, B. Haddow, and A. Birch, "Neural machine translation of rare words with subword units," in Proc. 54th Annual Meeting of the Association for Computational Linguistics, Berlin, Germany, 2016, pp. 1715–1725.
[11] E. M. Bender, T. Gebru, A. McMillan-Major, and S. Shmitchell, "On the dangers of stochastic parrots: Can language models be too big?" in Proc. ACM Conf. Fairness, Accountability, and Transparency (FAccT), 2021, pp. 610–623.
[12] A. Radford et al., "Language models are unsupervised multitask learners," OpenAI Blog, 2019.
Agent: An LLM-based system that plans multi-step tasks, calls tools, observes results, and adjusts its actions in a loop.
Alignment: The research effort to make models behave in line with human intentions and values (helpful, honest, harmless).
Backpropagation: The algorithm that computes how each parameter contributed to the model's error, enabling gradient descent.
Base model: A pretrained LLM before instruction fine-tuning; predicts likely continuations rather than acting as an assistant.
Benchmark: A standardized test set used to compare models (e.g., MMLU, HumanEval); useful but gameable.
Bias (model): Systematic unfair tendencies in model outputs, inherited from training-data prejudices.
Byte-pair encoding (BPE): A tokenization algorithm that builds a vocabulary by repeatedly merging the most frequent adjacent symbol pair.
Chain-of-thought prompting: Asking the model to reason step by step before answering, which improves multi-step task accuracy.
Checkpoint: A saved snapshot of a model's parameters at a point in training; pinning one enables reproducibility.
Closed model: A model whose weights are not public; accessible only via the provider's API or product.
Compute (training): The total arithmetic operations used in training, measured in FLOPs; a key scaling ingredient.
Contamination: When benchmark test questions appear in a model's training data, inflating its scores.
Decoder: The Transformer component that generates output token by token; modern LLMs are decoder-only.
Distillation: Training a small model to mimic a larger model's outputs, transferring capability cheaply.
Embedding: A learned vector representation of a token, where similar meanings sit near each other.
Emergent ability: A capability that appears suddenly at scale though absent in smaller models.
Fine-tuning: Additional training of a pretrained model for specific behavior or tasks (SFT, RLHF, LoRA).
FLOPs: Floating-point operations; the unit for measuring training and inference compute.
Few-shot learning: Teaching the model a task via examples in the prompt, without parameter updates.
Gradient descent: The optimization loop that nudges parameters to reduce prediction error.
Hallucination: Fluent, confident generation of false information.
In-context learning: The model's ability to infer tasks from prompt examples without retraining.
Inference: Running a trained model to generate outputs (as opposed to training it).
Knowledge cutoff: The date at which a model's training data ends; events after it are unknown to the model.
Latency: The delay before/between generated tokens; affected by hardware, network, and model size.
Loss: A number measuring how wrong the model's predictions were; training minimizes it.
Mixture-of-experts (MoE): An architecture with multiple sub-networks ("experts") where each token is routed to a few, giving large capacity at lower inference cost.
N-gram model: A classical language model estimating next-word probabilities from counts of short word sequences.
Open weights: Models whose parameter files are publicly downloadable (though training data/code may not be).
Parameter: An adjustable number (weight) in the model; counts range from billions to trillions.
Pretraining: The initial massive training phase: next-token prediction over trillions of tokens.
Prompt: The input text given to a model, including instructions and examples.
Prompt injection: Malicious instructions hidden in data the model processes, hijacking its behavior.
Quantization: Reducing the numerical precision of parameters (e.g., to 4-bit) to run models with less memory.
Retrieval-augmented generation (RAG): Giving the model relevant documents at inference time to ground its answers.
RLHF: Reinforcement learning from human feedback; fine-tuning models using human preference judgments.
Scaling laws: Empirical power-law relationships showing predictable improvement with more parameters, data, and compute.
Self-attention: The Transformer mechanism letting each token weigh the relevance of every other token.
SFT (supervised fine-tuning): Training on human-written prompt/response examples to teach assistant behavior.
Sycophancy: The model's tendency (amplified by RLHF) to agree with the user rather than be correct.
Temperature: A sampling parameter controlling randomness: low values make output deterministic, high values make it creative.
Token: The unit of text the model processes — often a subword; the model's native currency.
Tokenization: Converting text into token IDs (and back) via a learned vocabulary.
Transformer: The neural architecture (Vaswani et al., 2017) underlying modern LLMs, based on self-attention.
Tokenize by hand. Take the sentence "Unbelievably, tokenization affects everything!" and predict how a BPE tokenizer would split it. Then install tiktoken (or use an online tokenizer for your model of choice) and compare. Where were you wrong, and why?
Measure the tokenizer tax. Translate a 200-word English paragraph into Urdu (or another non-Latin-script language you know). Count tokens for both versions with the same tokenizer. Compute the ratio. Write a paragraph on what this implies for API cost and effective context length for that language.
Base vs assistant. Using any chat model, try a prompt designed for a base model (e.g., a list beginning "The capital of France is") and observe how the assistant fine-tuning shapes the response. Then ask a direct question. Describe the behavioral difference in your own words.
Few-shot design. Build a 3-shot prompt that teaches a model to classify paper abstracts by methodology (quantitative / qualitative / mixed). Test it on 10 abstracts you label yourself. What accuracy do you get? How does accuracy change with 1, 3, and 5 examples?
Hallucination hunt (integrity-aware). Ask a model for five citations on a niche topic in your field, with full references. Then verify each one in Google Scholar or your university library. Document which were real, which were fabricated, and which were real papers with wrong details. Write a 300-word reflection on what this means for literature reviews.
Bias probe (integrity-aware). Prompt a model to write short professional bios for names associated with different genders or ethnicities, keeping the prompt identical except for the name. Compare the outputs systematically. Report the differences you observe and discuss whether this model would be acceptable for any screening or evaluation task in your department.
Validation pipeline. Choose a real task from your research (e.g., extracting a field from 50 papers, labeling 100 items). Hand-label a gold sample, run an LLM on it, compute accuracy, and do an error analysis categorizing the mistakes. Write up the methods paragraph as you would for a paper, including model version, prompt, and parameters.
Cost planning. For the task in Exercise 7 scaled to your full dataset, estimate total tokens (input + output), look up current API pricing for two providers, and compute the cost. Compare against running an open-weights model locally (estimate GPU-hours). Which would you choose, and why? Include privacy considerations for your data type.
Model selection memo. Pick a hypothetical thesis project (real or invented). Walk through the Chapter 11 decision framework: task definition, hard constraints, shortlist, and hypothetical evaluation plan. Write a one-page memo justifying your model choice as you would to a supervisor.
Open-problem pitch (integrity-aware). From Chapter 12, choose one open problem relevant to your field or language. Write a two-page mini-proposal: the problem, why it matters, related work (with verified citations only — check every reference yourself), your proposed approach, and the smallest experiment that would test it. Exchange with a peer and critique each other's citation integrity.