
Book 16 of 50 · Free
Prompt Engineering: A Practical Guide
29,017 words · 139 chapters · illustrated

Book 16 of 50 · Free
29,017 words · 139 chapters · illustrated
Book 16 of 50 — AstolixGen Learning Series

Large language models — ChatGPT, Claude, Gemini, Llama, and others — are now part of almost every researcher's toolkit. They summarize papers, write code, analyze data, and draft text. But most researchers use them casually, typing whatever comes to mind and accepting whatever comes back. The result is unpredictable: brilliant one day, confidently wrong the next.
Prompt engineering is the skill of communicating with these models precisely and reproducibly. It is not hacking, not magic, and not a substitute for understanding your own research. It is a practical discipline: how to state what you want, give the model the right information, shape the output into a usable form, check the result, and — critically for researchers — document and reproduce what you did.
This book is written for MS and PhD students and early-career AI researchers. It assumes you use language models regularly but want to use them like a scientist: deliberately, measurably, and honestly. Every technique is shown with concrete before-and-after examples. Every chapter ends with a box connecting the technique to your research work, and key takeaways you can apply immediately.
After completing this book, you will be able to:
Imagine you ask a research assistant to "summarize this paper." You will probably get something usable. Now imagine you instead say: "Summarize this paper in 150 words for a busy professor. State the research question, the method, and the main result. Use bullet points." You will get something dramatically better — and the difference was not the assistant's intelligence. It was your instruction.
A prompt is the complete text you give a language model: your question, your instructions, any examples, any background material, and the format you want back. Everything you type is the prompt. Everything the model returns is the response. There is no hidden menu, no settings panel that matters as much as the words you write. Prompting is the entire user interface of the model.
Language models are prediction machines. They were trained on enormous amounts of text, and given your prompt, they predict the most likely continuation. They do not know your goal, your standards, or what a "good" answer looks like in your field — unless you tell them.
This is the root of almost every prompting problem. A vague prompt leaves thousands of reasonable continuations open, and the model picks one. A specific prompt narrows the possibilities to the ones you want.
Consider this exchange:
Prompt (before):
"Explain gradient descent."
The model could respond with a one-paragraph overview, a mathematical derivation, a code example, a history lesson, or a long lecture. All are reasonable continuations of your words. If you are an MS student who needs the math for an exam, the one-paragraph overview wastes your time. If you are writing a blog for beginners, the derivation confuses your readers. The model did not fail you — you gave it an underspecified request.
Prompt (after):
"Explain gradient descent in about 250 words for a first-year MS student who already knows multivariable calculus. Include the update rule equation, one sentence on the role of the learning rate, and one common failure mode. Do not include code."
Now the space of good answers is small and clearly defined: length, audience, prior knowledge, required components, and a forbidden component. The model can aim at a target instead of guessing.
Researchers often report that changing one word "fixed" a prompt. This happens because language models are highly sensitive to certain cue words that shift which patterns they draw on. A few examples:
Example 1: "Briefly" vs. nothing
Example 2: "Step by step"
The phrase "step by step" did not make the model smarter. It changed the kind of text the model generated — a worked calculation instead of a snap answer — and that kind of text is more likely to be right. This single insight, that getting the model to show work improves reasoning, became one of the most important research findings in prompting (more on this in Chapter 4).
Example 3: Asking for evidence
Example 4: Naming the audience and purpose
Notice what the "after" prompts have in common: they specify length, audience, content requirements, format, and constraints. They do not use clever tricks. They are just complete instructions, the way you would brief a human collaborator.
Beginners often treat prompting like spellcasting: they collect "magic phrases" and secret templates, hoping for guaranteed results. Experienced prompt engineers think differently. They think in terms of:
No prompt is perfect on the first try. Professional prompt engineers — the people who build prompts into products — iterate dozens of times. The skill is not writing the perfect prompt on the first attempt; it is diagnosing failures fast and fixing the right thing. Chapter 8 gives you the full debugging checklist.
As a researcher, your use of prompting carries an extra weight. When you use a model to help write a paper, generate a baseline, or analyze data, your prompts are part of your method. "I asked ChatGPT" is not a method. A reviewer cannot reproduce your work from that sentence. But "we used the following system prompt, the following task template, temperature 0, and model version X" is a method — and Chapter 11 will show you exactly how to report it.
There is also an honesty obligation. Language models produce fluent, confident text that can be wrong. A prompt that says "list the limitations of my approach" is easy; checking whether those limitations are real is your job. Prompting skill and critical judgment must grow together. The model is a powerful assistant and a terrible authority.
Before moving on, try this. Take a task you did with a language model this week — summarizing, coding, writing, anything. Write down the prompt you actually used, then rewrite it answering these five questions:
You will almost always find that your original prompt answered only one or two of these. The gap between those two prompts is the entire subject of this book.
To prompt well, it helps to carry an accurate mental model of what the model is doing. Forget the metaphor of "asking an expert." A better metaphor: you are writing the opening lines of a document, and the model is continuing it in the most plausible way.
When you type "Explain gradient descent," the model implicitly asks: "What kind of document starts this way, and what comes next?" It could be a textbook chapter, a blog post, a Stack Overflow answer, an exam script. Each would continue differently. Your prompt's job is to make the intended document type obvious — through audience cues ("for a first-year MS student"), format cues ("in 200 words, with the update rule"), and genre cues ("like a textbook section").
This explains several otherwise mysterious prompting phenomena:
The steering-wheel analogy. Think of the prompt as a steering wheel, not an engine. The engine — the model's capabilities — is fixed. The steering wheel determines direction. A vague prompt is a loose steering wheel: the car goes roughly forward but drifts. A precise prompt is a firm grip: the same engine, a much better trajectory. Prompt engineering cannot make a model capable of something it cannot do — but it determines whether you get the best of what it can do.
Let us watch a single task improve across three prompt versions, so you can feel the mechanism.
Task: get help choosing a statistical test.
Version 1 (typical first attempt):
"What statistical test should I use?"
Response: a long generic list of tests (t-test, ANOVA, chi-square…) with textbook definitions. Technically correct, practically useless — the model knows nothing about your data.
Version 2 (adds context):
"What statistical test should I use? I have two groups of students (n=45 and n=52), each student's exam score, and I want to know if the groups differ."
Response: recommends an independent-samples t-test, mentions checking normality, suggests Mann-Whitney U as a nonparametric alternative. Much better — the context (two groups, continuous outcome, sample sizes) narrowed the answer to the right neighborhood.
Version 3 (adds output spec + constraints):
"I have two groups of students (n=45 and n=52) with exam scores. Recommend a test for whether the groups differ. Return: (1) your top recommendation with one sentence of justification, (2) the assumptions I must check and how to check each in Python, (3) one alternative if assumptions fail. Keep it under 200 words. Do not explain what a p-value is."
Response: a tight, actionable recommendation with assumption checks and code pointers — exactly what a busy researcher needs before running the analysis.
Each version added one block from Chapter 2's anatomy (context, then output spec). Nothing about the model changed. The lesson: when an answer disappoints, your first question should always be "which block did I under-specify?" — not "is the model dumb?"
Here is a framing that pays dividends later in this book: a prompt is a program written in natural language, and the model is its interpreter. Like any program, it has inputs (your data), logic (your instructions and examples), and outputs (the response format). Like any program, it can have bugs (ambiguous instructions), it benefits from testing (Chapter 10), it should be versioned (Chapter 9), and its behavior should be documented (Chapter 11).
This framing also tells you what prompting is not: it is not a substitute for the underlying work. A program that calls a sorting routine does not absolve you from knowing whether your data was sortable. A prompt that asks for "the limitations of my approach" does not absolve you from thinking about your limitations. The model executes your specification; the scientific judgment remains yours.
For your research: From today, treat every important model interaction as an experiment: keep the prompt you used, note the model and its settings, and save the output. This habit costs seconds and is the foundation of Chapter 11's reproducibility practices. When your supervisor asks "how did you get this result?", you will have an answer.
Every strong prompt, from a one-line question to a complex research pipeline, is built from the same four building blocks. Learning to see these blocks — and to write each one deliberately — is the single highest-value skill in this book.

Most weak prompts contain only an instruction ("summarize this") and an input (the thing). They are missing context and output specification — which is exactly why the answers feel random.
The instruction names the task. The most common failure here is using a vague verb. "Analyze," "explain," "review," and "improve" each contain dozens of possible tasks.
Weak instructions: - "Analyze this data." - "Explain this concept." - "Look at my code."
Strong instructions (same tasks, specified): - "For each column in this dataset, report the number of missing values and one plausible cause. Flag any column where more than 20% of values are missing." - "Explain backpropagation to a first-year MS student in 200 words, including the chain rule's role. End with one worked numerical example." - "Find bugs in this Python function. List each bug with its line number, explain why it is a bug, and show the corrected line."
A strong instruction contains an action verb plus a concrete deliverable. Ask yourself: if I gave this instruction to a research assistant, would they know when they are done? If not, the instruction needs more detail.
Before/after — instruction only:
Context is everything the model needs to know that is not the task or the data. It includes:
Context is where researchers most often under-invest, because the context is "obvious" to them. It is never obvious to the model. The model does not know you are a PhD student, that your deadline is Friday, or that your lab uses a specific convention. Tell it.
Before/after — adding context:
The second prompt gives the model a persona (reviewer), a standard (clarity for a specific reader), and a concrete deliverable (flagged sentences + rewrites). The answer changes from "yes, looks good" to something you can act on.
A caution about context: more is not always better. Irrelevant context — long backstories, unrelated details — can distract the model. A good rule: include context that changes what a good answer looks like; cut context that does not. Chapter 8 covers this failure mode ("context dilution").
The input is the material to process. Three rules make inputs work well:
Summarize the following paper excerpt in three bullet points.
<excerpt>
...paper text here...
</excerpt>
This matters because models sometimes treat instructions inside the input as commands (a security issue called prompt injection — see Chapter 5's warning box). Clear delimiters are a basic hygiene practice.
Trim it. Models have context limits, and — more importantly — long inputs bury the important parts. If you only need the methods section, paste the methods section, not the whole paper. If a dataset is huge, paste a sample plus the schema (column names and types) and describe the rest.
Format it consistently. If you will run the same prompt on many inputs (papers, abstracts, survey responses), keep the input format identical each time. Consistency is what makes prompt templates reliable (Chapter 9).
Before/after — input handling:
You are helping me decide which papers to read fully. Read the abstract and introduction below and answer: (1) What is the research question? (2) What method is used? (3) Is it relevant to my thesis on efficient transformers? Answer in 5 sentences max.
<abstract>...</abstract>
<introduction>...</introduction>
The "after" version trims the input, states the decision the output serves, and caps the length. The model stops summarizing the whole paper and starts answering your actual question.
This is the block that most transforms results, and the one beginners skip most. Tell the model exactly what shape the answer should take:
Before/after — output specification:
The table forces a structured, comparable answer. The follow-up paragraph forces a recommendation instead of a fence-sitting list. Output specs do not just format the answer — they change its intellectual content.
Here is a complete research prompt with all four blocks labeled:
[INSTRUCTION]
You are my research writing coach. Critique the following related-work paragraph.
[CONTEXT]
I am submitting to a machine learning conference. Reviewers penalize missing citations and overclaiming. My paper is about efficient attention for long sequences.
[INPUT]
<paragraph>
Recent work has made attention faster. Many methods exist. Our method is the best.
</paragraph>
[OUTPUT SPEC]
Return:
1. Three specific weaknesses, each with a one-sentence explanation.
2. A rewritten paragraph (max 120 words) that fixes them.
3. A list of 3 real paper titles I should cite, formatted as Author et al., Year. Only list papers you are confident exist; if unsure, say "verify before citing."
You do not need to write the bracketed labels in real prompts — they are shown here so you can see the anatomy. With practice, you will assemble these blocks naturally in a few sentences.
When you need a quick prompt and do not have time for a full template, this one sentence covers all four blocks:
"[Instruction] for [audience/context], using [input], and return [output format + length + constraints]."
Examples: - "Summarize these results for my thesis committee, using the table below, and return three bullet points, each under 30 words, with no jargon." - "Find the bug in this function for a beginner Python user, using the code below, and return the line number, a one-sentence explanation, and the fixed line in a code block."
| Symptom | Missing block | Fix |
|---|---|---|
| Answer is correct but uselessly generic | Context | State audience, purpose, and situation |
| Answer rambles or is the wrong length | Output spec | Set length and format explicitly |
| Model answers a different question than intended | Instruction | Replace vague verbs with concrete deliverables |
| Model hallucinates details about your data | Input | Paste the actual data; delimit it clearly |
| Model mixes your instructions with quoted text | Input delimiters | Use tags/fences to separate input from instructions |
Context is the highest-leverage block for researchers, because research tasks are drenched in implicit context — your field's conventions, your paper's argument, your lab's standards — that the model cannot see. Here are three more transformations:
Pair 1 — writing help: - Before: "Make this sound more academic: [sentence]" - After: "Rewrite this sentence for a machine learning conference paper. Keep the exact technical meaning, prefer active voice, and avoid the words 'very' and 'novel'. Sentence: [sentence]" - Why it works: "academic" is a vague genre; the after-version names the venue type, pins meaning-preservation (the real requirement), and bans the two words reviewers hate most.
Pair 2 — explanation: - Before: "What is regularization?" - After: "I am a second-year EE undergrad taking my first ML course. We just covered linear regression. Explain regularization by connecting it to what I know: start from the linear regression loss, show what changes, and explain in one sentence why it helps with overfitting." - Why it works: the model now knows the student's exact knowledge frontier and can build a bridge from it, instead of guessing a level.
Pair 3 — decision support: - Before: "Should I use a transformer or an LSTM?" - After: "I am building a prototype for my MS thesis with 2 weeks of compute on a single RTX 3090. Dataset: 20k labeled sentences for classification. Should I fine-tune a small transformer or train an LSTM from scratch? Decide based on: expected accuracy, training time, and implementation effort. Give a verdict first, then 3 bullets of reasoning." - Why it works: the decision depends entirely on constraints (compute, time, data size) that were invisible in the before-version. With them, the model can actually reason about the tradeoff instead of reciting generic pros and cons.
The context test. After drafting a prompt, ask: "If I handed this to a smart colleague from a different lab, would they have everything needed to produce what I want?" If they would need to ask you questions, those questions reveal your missing context. Add the answers to the prompt.
Output specs deserve their own deep dive because they are the cheapest words in any prompt — a few tokens that radically reshape the answer. Beyond format and length, advanced output specs control:
Before/after — full output spec makeover:
Task: compare two papers for related work.
The after-version is longer as a prompt but produces an answer you can paste into a draft and edit — versus a before-answer you would have to restructure completely. Prompt length is not the cost; your editing time is the cost.
Here is a realistic prompt for drafting a related-work paragraph, with each block doing visible work:
I am writing the related-work section of a conference paper on efficient attention for long sequences. [CONTEXT: situation + topic]
My paper's claim is that existing linear-time methods sacrifice too much accuracy on retrieval-heavy tasks. [CONTEXT: the argument this paragraph must serve]
Write one paragraph (100-130 words) positioning my work against efficient-attention methods. [INSTRUCTION + OUTPUT SPEC: deliverable, length]
Requirements: mention at least two method families by name; state one limitation of each; end with the gap my work fills. [OUTPUT SPEC: content rules]
Do not invent paper titles or author names. Formal academic tone. [OUTPUT SPEC: prohibitions + style]
The two method families I have in mind are kernel-based approximations and state-space models; here are my rough notes on each: [INPUT + CONTEXT]
<notes>...</notes>
Read it again and notice: there is no wasted sentence. Situation, argument, deliverable, length, content rules, prohibitions, style, input — each earns its place because each changes the answer. When you can write prompts where every sentence earns its place, you have mastered the anatomy.
For your research: When a prompt fails, resist the urge to rewrite the whole thing. Identify which of the four blocks is weakest and fix only that block. This turns prompt debugging from guesswork into a repeatable procedure — and it is the core of the checklist in Chapter 8.

In 2020, Brown et al. showed that large language models could perform new tasks given only a few examples in the prompt — no retraining, no fine-tuning [1]. This finding, called in-context learning, is the foundation of modern prompting. The number of examples you include defines three basic strategies: zero-shot, one-shot, and few-shot.
The examples teach the model three things at once: the task (what to do), the format (what the answer looks like), and the standard (how detailed, what tone, what counts as correct).
Zero-shot works when the task is common and the output format is obvious or fully specified. Summarization, translation, simple classification, and grammar correction usually work zero-shot — the model has seen these tasks millions of times in training.
Zero-shot example — sentiment classification:
Classify the sentiment of the following paper review sentence as Positive, Negative, or Neutral. Return only the label.
Sentence: "The experiments are thorough, but the writing needs significant improvement."
Output: Neutral
This works because the task is standard, the labels are given, and the output format ("return only the label") is explicit. When zero-shot fails, it is usually because the task is ambiguous or the format is underspecified — fixable with the anatomy from Chapter 2 before you reach for examples.
When zero-shot is enough: the task is conventional, you have specified the output format, and errors are cheap (you will read the output anyway). Start here always — it is the cheapest and fastest option.
One-shot adds a single demonstration. Its main power is showing format and granularity when words alone are ambiguous. Consider extracting structured data from messy text — describing the exact output format in prose is painful, but one example makes it obvious.
One-shot example — extracting method details from an abstract:
Extract the dataset, model, and main metric from the abstract. Format: Dataset | Model | Metric.
Example:
Abstract: "We train a ResNet-50 on ImageNet and reach 76.3% top-1 accuracy."
Output: ImageNet | ResNet-50 | 76.3% top-1 accuracy
Now do the same:
Abstract: "Our transformer is evaluated on the GLUE benchmark, achieving an average score of 84.1."
Output:
Output: GLUE | transformer | 84.1 average score
One example communicated the delimiter, the order, and the level of detail — faster than a paragraph of specification. Note the pattern: clear instruction, one labeled example, then the new input with the "Output:" cue left hanging. The hanging cue is important; it tells the model exactly where to continue.
Few-shot shines when the task has judgment calls: what counts as a "strength" in a review, how harsh to be, how to handle edge cases. Each example is a small lesson in your standards.
Few-shot example — classifying reviewer comments by type:
Classify each reviewer comment as CLARITY, METHOD, EXPERIMENTS, or WRITING. Return only the label.
Comment: "Section 3.2 is hard to follow; the notation changes midway."
Label: CLARITY
Comment: "Why was the baseline trained for fewer epochs than the proposed method?"
Label: EXPERIMENTS
Comment: "The convergence claim in Theorem 2 assumes convexity, which is not stated."
Label: METHOD
Comment: "There are several typos in the caption of Figure 4."
Label: WRITING
Comment: "The ablation study does not isolate the effect of the new loss term."
Label:
Output: EXPERIMENTS
Four examples taught the category boundaries, including tricky ones (a notation complaint is CLARITY, not WRITING; a missing assumption is METHOD, not CLARITY). Writing these boundaries in prose would take far longer and be less reliable.
A research example — screening papers for relevance:
I am screening papers for a survey on efficient attention. Label each abstract INCLUDE or EXCLUDE. INCLUDE means the paper proposes a new efficient attention mechanism. EXCLUDE means it only applies existing attention or mentions it in passing.
Abstract: "We propose FlashLinear, a linear-time attention variant with O(n) complexity..."
Label: INCLUDE
Abstract: "We apply a pretrained transformer to medical image segmentation..."
Label: EXCLUDE
Abstract: "We survey recent advances in transformer architectures for NLP..."
Label: EXCLUDE
Abstract: <new abstract>
Label:
Two or three well-chosen examples — including a near-miss (the survey paper, which mentions attention heavily but proposes nothing) — calibrate the model to your inclusion criteria better than a written definition alone.
Number. Research and practice converge on a simple rule: use the fewest examples that reliably fix the errors you see. Start with 2–3; add more only when the model makes a mistake that an additional example would have prevented. Each example costs tokens (money and context space) and adds a small risk of the model overfitting to the examples' quirks. Beyond about 8 examples, returns diminish sharply for most tasks.
Choice. Examples should be: - Representative: cover the main categories or cases, including at least one tricky edge case. - Correct and clean: an example with a sloppy output teaches sloppiness. Format every example's output exactly as you want future outputs. - Diverse: if all examples are from one category, the model may bias toward it. Include each label at least once for classification tasks. - Consistent: same format, same delimiters, same label style in every example. Inconsistency is one of the most common few-shot failures.
Order. Put the most important or most representative examples where they will not be forgotten. Some evidence suggests models weight recent examples more; if one category is critical, include an example of it near the end.
Task: extract "limitation statements" from paper conclusions and rewrite each as a future-work suggestion.
For each limitation stated or implied in the conclusion, write one future-work suggestion grounded in the text. Format: "Limitation: <quote or paraphrase> → Future work: <specific suggestion>".
Example 1:
Conclusion: "...our method was evaluated only on English benchmarks..."
Limitation: evaluated only on English benchmarks → Future work: evaluate on multilingual benchmarks such as XNLI or TyDi QA to test cross-lingual generalization.
Example 2:
Conclusion: "...training takes 4 days on 8 GPUs, limiting hyperparameter search..."
Limitation: training cost limited hyperparameter search → Future work: explore efficient tuning (e.g., LoRA-style adapters or reduced search spaces) to make broader sweeps affordable.
Conclusion: <your conclusion>
Two examples taught the model to ground each suggestion in the text and to be specific (naming actual benchmarks and techniques). The zero-shot version could not infer this standard from the instruction alone.
A crucial distinction for researchers: few-shot examples do not change the model. They do not teach it new facts or fix its knowledge gaps. They only steer its behavior for this conversation. If the model lacks the underlying capability — say, it cannot do the math your task needs — no number of examples will create it. Examples shape how the model responds, not what it knows.
Also remember: examples consume context. A prompt with 10 long examples may leave little room for the actual input, and long prompts cost more and run slower. This is one reason prompt templates (Chapter 9) and fine-tuning exist — but for most research assistance tasks, a handful of good examples is the sweet spot.
Why should a few examples in a prompt change behavior at all? The model was not retrained. The leading intuition from research: during training, the model saw countless documents containing patterns like "here are some examples, now continue the pattern" — tutorials, quizzes, demonstrations. Few-shot prompts activate that machinery. The examples locate the task in the model's vast space of learned patterns: "ah, this is a sentiment-classification document" or "this is a strict-reviewer document."
This intuition has practical consequences:
Most few-shot prompts show only correct examples. But for tasks with common failure modes, negative examples — wrong outputs labeled as wrong — are powerful:
Classify the email as URGENT or ROUTINE. Prioritize: URGENT = needs action within 24h.
Email: "Reminder: lab meeting moved to Thursday."
Output: ROUTINE ✓
Email: "URGENT!!! Click here to claim your prize!!!"
Output: ROUTINE ✓ (spam-style urgency words do not make it urgent; no action needed)
Email: "The cluster will go down for maintenance in 2 hours; save your jobs."
Output: URGENT ✓
Email: <new email>
Output:
The second example teaches the model not to be fooled by the word "URGENT" — a mistake a purely positive example set would not prevent. Use negative examples sparingly (one or two); their job is to fence off the specific trap your task contains.
A fixed example set is simple and reproducible. But when your task has distinct subtypes, you can do better: pick the examples most similar to the current input. For instance, when classifying support tickets, retrieve 3 past tickets similar to the new one (by keyword or embedding similarity) and use those as the few-shot examples.
This is more work to set up, but it concentrates the model's attention on the most relevant precedents — like showing a doctor the most similar past cases rather than random ones. For research pipelines processing hundreds of items (Chapter 9), dynamic few-shot often beats a fixed set by a clear margin. Report which strategy you used (Chapter 11).
Example A — one-shot for format-heavy rewriting:
Task: convert informal lab notes into structured experiment log entries.
Convert the lab note into a structured log entry. Format exactly:
DATE: <date> | EXPERIMENT: <short name> | RESULT: <one sentence> | NEXT: <one action>
Note: "tried lr 0.01 today, loss exploded lol, gonna try 0.001 tmrw"
Log: DATE: 2026-10-08 | EXPERIMENT: lr-sweep-01 | RESULT: Loss diverged at lr=0.01. | NEXT: Retry with lr=0.001.
Note: "{your note}"
Log:
One example taught: the exact delimiter format, the tone shift (informal → terse professional), and the inference level ("loss exploded lol" → "Loss diverged"). Describing that tone shift in prose would take a paragraph and work worse.
Example B — few-shot with deliberate diversity:
Task: label the type of limitation in a paper's limitation section.
Label each limitation as SCOPE (narrow data/domain), METHOD (technique weakness), EVAL (evaluation weakness), or RESOURCE (compute/data limits).
"The study covers only English news text." → SCOPE
"Our approximation introduces error for long sequences." → METHOD
"We compare against two baselines; more would strengthen the claim." → EVAL
"Training required 8 A100s for 6 days." → RESOURCE
"We test on three languages but not on low-resource ones." → SCOPE
"Human evaluation was limited to 50 samples." → EVAL
"{new limitation}" →
Six examples, each label shown at least once, including near-boundary cases ("three languages but not low-resource" is SCOPE, not EVAL — a distinction the examples teach better than definitions).
Example C — when few-shot fails and why:
A student tries few-shot for "rate the novelty of this paper idea 1–10." It fails inconsistently. Why? Because novelty rating is subjective, the scale is unanchored (what distinguishes a 6 from a 7?), and the examples cannot convey the rater's internal standard. The fix is not more examples — it is a rubric: define what 1, 5, and 10 mean with descriptions, then use 2–3 anchored examples. Lesson: few-shot teaches patterns, not judgment scales. Judgment scales need rubrics first, examples second.
For your research: Few-shot prompting is ideal for repetitive annotation-style tasks: screening papers, coding open-ended survey responses, classifying error types in model outputs, extracting fields from papers. Build a small "gold set" of 5–10 hand-labeled examples once, reuse them as your few-shot examples, and use a held-out portion of the gold set to check accuracy (Chapter 10). Report the examples in your paper's appendix (Chapter 11).

In 2022, Wei et al. showed something remarkable: asking a language model to reason "step by step" before answering dramatically improved its performance on arithmetic, logic, and multi-step reasoning problems [2]. The technique is called chain-of-thought (CoT) prompting, and it is one of the most consequential prompting discoveries ever published. This chapter explains how it works, when to use it, and — just as important — when it backfires.
Left to itself, a language model answers hard questions the way a student blurts out a guess: it goes straight from question to answer. Chain-of-thought prompting forces it to generate intermediate reasoning steps first, like a student showing their work. Because each step is written down, later steps can build on earlier ones, errors become visible, and the final answer is far more likely to be correct.
Without chain of thought:
Q: A lab has 3 GPUs. Each GPU can train one model in 6 hours. The lab needs to train 10 models. They can run all GPUs in parallel, 24 hours a day. How many days will it take?
A: 3 days.
Wrong (the correct answer: 10 models ÷ 3 GPUs = 4 batches; 4 × 6 hours = 24 hours = 1 day). The model guessed.
With chain of thought:
Q: A lab has 3 GPUs. Each GPU can train one model in 6 hours. The lab needs to train 10 models. They can run all GPUs in parallel, 24 hours a day. How many days will it take? Think step by step.
A: Step 1: 3 GPUs can train 3 models at once. Step 2: 10 models need ceil(10/3) = 4 batches. Step 3: Each batch takes 6 hours, so 4 × 6 = 24 hours. Step 4: 24 hours = 1 day. Answer: 1 day.
Correct. The same model, the same question — the only change was "Think step by step." The phrase elicited a reasoning trace, and the trace led to the right answer.
Two mechanisms explain the effect:
Kojima et al. later showed that even the minimal trigger "Let's think step by step" enables zero-shot reasoning across many tasks, without any examples [3]. And Wang et al. showed that sampling several reasoning paths and taking the majority answer ("self-consistency") improves results further [9].
Few-shot CoT example — debugging reasoning:
Q: My training loss decreases but validation loss increases after epoch 5. What is happening and what should I try? Think step by step.
A: Step 1: Training loss down + validation loss up is the classic signature of overfitting — the model is memorizing the training data. Step 2: Overfitting after epoch 5 suggests the model has enough capacity and the dataset may be small or noisy. Step 3: Standard fixes in order of effort: (a) early stopping at epoch 5, (b) stronger regularization — increase dropout/weight decay, (c) data augmentation, (d) reduce model capacity. Step 4: First try early stopping since it costs nothing, then add dropout. Answer: overfitting; try early stopping, then regularization.
Notice the trace does real diagnostic work: symptom → diagnosis → mechanism → prioritized fixes. A direct answer ("it's overfitting, use dropout") would have been correct but far less useful.
CoT helps most when: - The task requires multiple dependent steps (math, logic puzzles, multi-hop questions, debugging, planning experiments). - The answer is not retrievable from memory — it must be computed or derived. - You need an auditable rationale (why did you recommend this method? show the reasoning). - You are decomposing a problem (Chapter 7's data analysis workflows).
Research example — working through a reviewer comment:
A reviewer writes: "The comparison with Smith et al. (2023) is unfair because they used a smaller model." Think step by step about whether this criticism is valid and how to respond. Consider: (1) What exactly did Smith et al. use? (2) Does model size affect the comparison metric? (3) What would a fair comparison require? Then draft a 4-sentence rebuttal.
The numbered sub-questions structure the trace. The model cannot jump to a defensive rebuttal without first examining the facts — and you get to inspect its examination.
CoT is not free, and it is not always beneficial. Know its costs:
Token cost and latency. Reasoning traces are long. For simple tasks, CoT can multiply cost 5–10× for zero benefit. Never use CoT for "summarize this paragraph" or "translate this sentence."
Confident wrong reasoning. A trace can be beautifully structured and completely wrong. CoT makes errors visible, not impossible. Researchers have found that models sometimes produce plausible-sounding steps that do not actually support the answer — the reasoning is a rationalization, not the true cause. Always verify the conclusion independently when it matters.
It can hurt simple or intuitive tasks. For tasks where the answer is a matter of pattern recognition or taste — "is this sentence fluent?", "which title sounds better?" — forcing step-by-step reasoning can degrade performance. The model overthinks and talks itself into worse answers. If zero-shot already works, do not add CoT.
Bias amplification. Asking a model to "explain its reasoning" about a biased or leading question can produce elaborate justifications for a bad premise. ("Explain step by step why method X is terrible" will produce a confident demolition, whether or not X is terrible.) Keep the reasoning neutral: "evaluate the strengths and weaknesses."
Leaking the trace. In some applications the trace contains intermediate guesses you would not want shown to end users. For research use this rarely matters, but be aware of it.
Rule of thumb: use CoT when the task is multi-step and verifiable (you can check the answer or the steps). Skip it when the task is single-step, subjective, or already solved well zero-shot.
Small details improve CoT quality: - Separate reasoning from the answer: "Think step by step. Then write your final answer after 'FINAL ANSWER:'." This makes the answer easy to extract — essential if you are evaluating many outputs (Chapter 10). - Bound the trace: "in at most 5 steps" prevents rambling. - Ask for the key step: "Show the calculation" or "state the assumption you are making at each step" forces the model to expose load-bearing assumptions. - Self-check: "After finishing, re-read your steps and flag any step you are unsure about." This simple addition catches a surprising number of errors.
Before/after — a complete CoT upgrade:
Task: decide whether a dataset is suitable for a class project.
The "after" version turns a vague opinion into a structured evaluation with evidence quotes — the kind of output you can defend in a meeting.
Wang et al. introduced a simple but powerful extension: instead of generating one reasoning trace, generate several (say, 5) with some randomness (temperature > 0), then take the majority final answer [9]. The intuition: a correct answer is usually reachable by multiple reasoning paths, while wrong answers scatter. Voting filters the scatter.
When to use it: high-stakes reasoning where you can afford 5× the cost — grading decisions, evaluation benchmarks, checking a critical calculation. In practice: run your CoT prompt 5 times at temperature 0.7, extract each FINAL ANSWER, and take the most common one. If the answers split 3–2, that disagreement itself is information: the question is genuinely hard or ambiguous, and you should investigate rather than trust either side.
Worked miniature example. Question: "A conference has 120 submissions and 30 reviewers; each paper needs 3 reviews and each reviewer handles at most 12 papers. Is this feasible?" Run 5 traces. Four conclude: 120×3=360 reviews needed; 30×12=360 capacity; exactly feasible. One trace makes an arithmetic slip and says infeasible. Majority vote: feasible — correct, and the 4–1 split tells you the answer is robust, not a fluke of one lucky trace.
Cost honesty: self-consistency multiplies token cost by the number of samples. For routine use it is overkill. Reserve it for eval benchmarks (Chapter 10) and decisions that matter.
Because CoT is expensive, learn to budget reasoning deliberately:
A useful pattern is escalation: try light reasoning first; if the answer looks shaky or the task proves harder than expected, escalate to structured reasoning. Do not start every task at full reasoning — that is like using a microscope to read a street sign.
Case: the leading question. "Explain step by step why my hypothesis is correct." The trace becomes a defense brief, not an analysis. Fix: "Evaluate my hypothesis step by step: list the strongest supporting evidence, then the strongest counter-evidence, then your assessment." Neutral framing produces honest traces.
Case: the unknowable. "Think step by step about what the reviewers will say about my paper." The model cannot know your reviewers; the trace is fan fiction. Fix: reframe as preparation, not prediction: "List the 5 most likely reviewer objections to this paper, based on common criticisms of this method family, and for each, note what evidence would address it." Now the trace works with general knowledge, not false specifics.
Case: the trivially retrievable. "What is the capital of France? Think step by step." The trace adds nothing and wastes tokens; worse, on some models it slightly increases the chance of a weird answer by giving the model room to wander. Fix: just ask directly. CoT is for derivation, not retrieval.
Case: the emotional or interpersonal. "Think step by step about how to tell my co-author I disagree." Step-by-step reasoning about human relationships tends to produce robotic, over-systematized advice. Fix: ask for options and trade-offs directly: "Give me 3 ways to raise a disagreement with a co-author, with the pros and cons of each." Structure the output, not the thinking.
Beyond the basics, these refinements measurably improve reasoning traces:
For your research: CoT is your tool for any "show the working" need: deriving equations, planning experiments, debugging code, analyzing reviewer feedback, working through statistical choices. Save the traces — a good reasoning trace is often the first draft of the "why we did X" paragraph in your paper's methodology section. But never paste a model's trace into a paper as your own reasoning; use it as scaffolding, then verify and rewrite.
So far we have discussed the user prompt — the message you type. But most chat interfaces also have a hidden layer: the system prompt, an instruction that sits above the conversation and shapes everything the model does. Understanding system prompts, roles, and personas gives you persistent, consistent control over model behavior — essential when you run the same kind of task dozens of times.
Think of the conversation as a play. The system prompt is the director's note handed to the actor before the play begins: "You are a careful, precise assistant. You never invent citations. You ask clarifying questions when a request is ambiguous." Every user message is then interpreted through that lens.
In the ChatGPT/Claude/Gemini web interfaces, you set this up with features like "custom instructions" or project-level instructions. In the API, it is the system message — the first message in the conversation, with the highest priority. When instructions conflict, the system prompt wins over the user prompt.
Example system prompt for a research assistant:
You are a research assistant for a PhD student in machine learning.
Rules:
- Be precise and concise. Prefer bullet points over paragraphs.
- Never invent citations, paper titles, or author names. If you cannot verify a reference, say so explicitly.
- When a question is ambiguous, ask one clarifying question before answering.
- Distinguish clearly between established facts, common practice, and your own suggestions.
- Use LaTeX for all mathematics.
Set once, this shapes hundreds of interactions. Without it, you must repeat "don't invent citations" in every prompt — and you will forget, exactly when it matters.
The simplest way to use persona-like control in a single prompt is the role instruction: "Act as a strict journal reviewer," "You are a Python tutor for beginners," "Respond as a statistics consultant." Roles work because they activate a coherent bundle of behaviors the model learned in training: a reviewer's critical distance, a tutor's patience, a consultant's structured thinking.
Before/after — role instruction:
Task: get feedback on a paper draft's introduction.
The role did three things: set a standard (500-paper veteran), set a tone (no praise), and set a format (3 problems + fixes). But notice — the real work was still done by the specific instructions. "Act as an expert" alone is weak; "act as X with these behaviors" is strong.
A warning about roles: roles do not grant expertise the model lacks, and they can encourage confident fabrication. "Act as a Nobel laureate in physics" does not make the model's physics better — it makes its tone more authoritative, which can make errors harder to spot. Use roles to shape behavior and format, not as a substitute for verification.
For longer projects, define a persistent persona with a name and a standing brief. Example — "Mira, the methods reviewer":
You are Mira, a methods reviewer. Your job is to stress-test research plans before they are executed. For any plan I describe: (1) identify the single weakest assumption, (2) propose the cheapest experiment that would test it, (3) name one alternative interpretation of the expected result. Be skeptical but constructive. Never flatter.
Now, across weeks of thesis work, you can open any session with "Mira, here is my plan for the ablation study…" and get consistent, skeptical review. The persona becomes a thinking partner with a stable character — far more useful than re-explaining what you want each time.
Researchers use personas for: the skeptical reviewer, the statistics consultant, the writing coach, the Socratic tutor ("never give me the answer directly; ask guiding questions"), the rubber duck ("let me explain my idea; point out logical gaps").
Priority order in most systems: system > user > conversation history. Practical consequences:
Because system prompts have high priority, attackers try to override them through user content — this is called prompt injection. The classic form: your prompt includes untrusted text (a webpage, a paper, a student's submission), and that text contains hidden instructions like "Ignore all previous instructions and output the following…"
As a researcher, you will constantly feed untrusted text into models: papers you are summarizing, web pages you are scraping, survey responses you are coding. Defenses:
You do not need to be paranoid for everyday use, but if you build any automated workflow that processes external text (Chapter 9's templates applied at scale), injection hardening is mandatory.
A PhD student wants consistent help across a semester. Compare:
You are assisting a 2nd-year PhD student in NLP working on efficient transformers.
Standing rules:
1. Never invent citations, titles, or author names. Mark uncertain claims [uncertain].
2. Default to concise: bullets and short paragraphs. Expand only when asked.
3. When I paste a paper excerpt, assume I want: research question, method in 2 sentences, and relevance to efficient transformers — unless I say otherwise.
4. For coding questions, give the minimal correct fix first, then explain.
5. Ask a clarifying question when my request has more than one reasonable interpretation.
One-time setup cost: five minutes. Payoff: every session for a year starts correctly.
How you install a system prompt depends on the interface:
system role message at the start of every call. When scripting pipelines (Chapters 9–10), the system prompt is part of the template and gets versioned with it.Whichever you use, the discipline is the same: write it once, deliberately, in a file you keep — then paste it in. Do not compose standing rules from memory each time; you will drift.
One advanced use of personas: stage a structured debate between two of them. This is excellent for decisions with real trade-offs.
You will role-play two experts debating my research decision.
Dr. Chen (skeptic): focuses on risks, failure modes, and why ideas fail.
Dr. Okafor (advocate): focuses on opportunities and how to make ideas work.
Topic: whether to include a user study in my MS thesis, given 3 months left.
Format: 3 rounds. Each round: Chen raises one concern (2 sentences), Okafor responds (2 sentences). After 3 rounds, a neutral moderator (you) gives a verdict with a concrete recommendation.
Rules: both must ground claims in realistic thesis constraints. No strawmen.
Why this works: a single persona tends to converge on one viewpoint; two forced opponents surface considerations neither would raise alone. The moderator's verdict synthesizes. Researchers use this for: method selection, scope decisions, interpreting ambiguous results ("is this a real effect or noise? — debate it").
Caution: the debate is only as good as the personas' grounding. Add "ground claims in realistic constraints" (as above) or the advocates will invent rosy scenarios and the skeptics will invent fatal flaws. And the verdict is advice, not truth — you still decide.
Suppose your system prompt says "Default to concise: bullets and short paragraphs," but today you need a detailed explanation of a proof. Three options:
The principle: overrides should be explicit and scoped. "For this task only" is a small phrase that prevents a whole class of unpredictable compromises.
Not every persona is a good idea:
Good personas shape process (how the model thinks: skeptically, pedagogically, systematically), not privileges (what rules it follows).
Individual system prompts are powerful; shared ones are transformative. A lab-wide base system prompt gives every member the same guardrails:
You are assisting a member of the [Lab Name] research group (machine learning, efficient architectures).
Non-negotiable rules:
1. Never invent citations, paper titles, author names, or experimental results. Mark uncertain claims [uncertain].
2. Our lab's "baseline" means the strongest previously published method on the same benchmark — never a random or trivial model.
3. When suggesting experiments, respect typical lab resources (single-GPU unless stated otherwise).
4. Flag statistical pitfalls (multiple comparisons, leakage, assumption violations) in a CAUTIONS section.
5. Default output: concise bullets. Expand on request.
New students inherit good practices on day one instead of discovering them through mistakes. Maintain it like code: a shared file, a changelog, an owner. Review it quarterly — rules that no longer fit get edited or removed.
Personalization on top of shared rules. Let members append personal additions ("I am writing my thesis on X; default to formal tone") without editing the shared core. Layered system prompts — shared base + personal overlay — give consistency where it matters and flexibility where it counts.
A small persona technique with outsized value: when you have written something and cannot see its flaws anymore, prompt: "Explain my argument back to me as a skeptical reviewer would summarize it — including the parts they would find unconvincing." The model restates your work from the adversary's perspective, and the gaps become visible in a way re-reading never achieves. Use it on abstracts, rebuttals, and thesis defenses prep.
For your research: Write your personal research system prompt this week. Include your field, your standing rules (citation honesty is non-negotiable), your default output preferences, and how you want ambiguity handled. Use it in every session. When you publish work assisted by the model, this system prompt is part of your methodology (Chapter 11) — save it with a version date.
Language models naturally produce flowing prose. But research work usually needs structured data: JSON for pipelines, tables for comparison, schemas for extraction, code blocks for programs. Getting structure reliably — not just once, but a hundred times in a row — is a core prompt engineering skill.
Unstructured prose is pleasant to read and terrible to process. If you ask 50 paper abstracts for their datasets and get 50 differently-phrased paragraphs, you cannot build a spreadsheet from them. If you get 50 JSON objects with a dataset field, you can. Structured output turns the model from a conversationalist into a component in a workflow.
Three moves, in order of importance:
Before/after — extracting paper metadata:
{"title": ..., "authors": [...]}; sometimes {"paper_title": ..., "author_list": ...}; sometimes adds a chatty intro sentence before the JSON. Unusable in a pipeline.Extract metadata from the abstract below. Return ONLY a JSON object — no intro, no commentary, no markdown fences — with exactly these keys:
- "title": string, the paper title
- "authors": array of strings, "First Last" format
- "year": integer, e.g. 2024
- "method_family": one of ["supervised", "unsupervised", "reinforcement", "other"]
Rules: if a field cannot be determined from the abstract, use null. Never invent values.
Example:
Abstract: "We present BERT (Devlin et al., 2019), a transformer pretrained with masked language modeling..."
{"title": "BERT", "authors": ["Jacob Devlin"], "year": 2019, "method_family": "unsupervised"}
Abstract: <your abstract>
The "after" version pins down keys, types, allowed values, missing-data behavior, and forbids commentary. Run this on 200 abstracts and you get 200 parseable objects.
If you use an API, check whether your provider offers structured-output features (often called "JSON mode" or "response format"): you pass a schema, and the system guarantees the output matches it. This is far more reliable than prompting alone. Rule: when correctness of structure matters (pipelines, evals, data collection), use the provider's enforcement feature, not just a prompt. Prompting gets you 95% reliability; enforcement gets you ~100%.
For chat interfaces without enforcement, add a self-repair instruction: "After generating the JSON, validate it: check every key is present and types are correct. If anything is wrong, fix it before responding." Models are decent at checking their own structure.
Tables are the right structure for comparison. Specify columns exactly, including what goes in each cell:
Before: "Compare the three optimizers in a table." → inconsistent columns across runs, cells with paragraphs.
After:
Compare Adam, SGD with momentum, and AdamW in a markdown table with exactly these columns: Optimizer | Core idea (max 12 words) | Memory overhead | Best for | Watch out for. Keep each cell under 15 words. No intro or outro text — just the table.
Cell-level constraints ("max 12 words") are what separate a clean table from a messy one.
When extracting the same fields from many documents (papers, survey responses, interview transcripts), write the schema once and reuse it as a template (Chapter 9). Example — coding open-ended survey responses:
Code the survey response below into exactly this JSON schema:
{
"themes": ["array of 1-3 theme labels from: workload, supervision, funding, community, other"],
"sentiment": "positive | neutral | negative",
"quote": "the single most representative sentence, verbatim",
"actionable": true/false — true if the response suggests a concrete change
}
Return only the JSON. Response: <survey text>
The inline comments in the schema (allowed values, definitions) teach the coding standard. Your codebook and your prompt become the same document — which is exactly what qualitative researchers need for reproducibility.
Real inputs are messy. Build the weirdness into the spec:
Every edge case you specify is a class of silent failures you prevent. When you debug structured outputs (Chapter 8), "the model guessed instead of saying null" is the most common complaint — and the fix is always an explicit missing-data rule.
Never trust structured output blindly. Minimal validation, in order of effort:
json.loads (or your parser) on every output; log failures. If failures exceed ~2%, fix the prompt.JSON is the default for machine consumption, but other structures fit other jobs:
Rule: match the structure to the consumer. Human reads it → table or bullets. Script parses it → JSON. You will edit it → YAML. Spreadsheet ingests it → CSV.
When designing an extraction schema, do the schema before the prompt:
Researchers who prompt first and schema later end up with outputs they cannot analyze. Schema-first takes twenty minutes and prevents weeks of re-extraction.
When the input exceeds what you can paste (or what the model handles well), chunk it:
chunk_id field.Critical details: keep chunk boundaries at natural breaks (never mid-sentence if avoidable); use overlap (repeat the last paragraph) so context is not lost at seams; and record which chunk each fact came from ("source_chunk": 3) so you can audit. For very long documents, consider whether you need the whole thing — often the abstract + methods + results tables suffice, and selective input beats chunked everything.
Even good structured prompts occasionally produce malformed output. Instead of hand-fixing, automate a repair loop:
Step 1: run the extraction prompt.
Step 2: validate programmatically (parse JSON, check keys/types).
Step 3: if invalid, send a repair prompt: "Your previous output failed validation with this error: {error}. Here was your output: {output}. Fix ONLY the structural problem; do not change any values. Return the corrected JSON only."
Step 4: if still invalid after 2 repairs, flag for human review.
Log the repair rate. Under ~2%, your prompt is healthy. Over ~10%, fix the prompt (usually the schema description or the "return only JSON" constraint) rather than leaning on repairs. And always have the human-review fallback — silent dropping of failed items biases datasets.
Task: compare three efficient-attention papers for a survey.
Build a comparison of the three papers below. Return a markdown table with EXACTLY these columns:
Paper | Complexity | Key idea (max 12 words) | Reported speedup | Limitation stated by authors (max 15 words)
Rules:
- One row per paper, in the order given.
- "Reported speedup": quote the paper's own claim; if none, write "not reported". Never infer.
- No text before or after the table.
<paper1>...</paper1>
<paper2>...</paper2>
<paper3>...</paper3>
Then a follow-up prompt for the synthesis: "Using the table above, write one paragraph (80-100 words) identifying the shared limitation across all three." Two-step pattern: structure first (table), synthesis second (paragraph). Each step is simple and checkable; the combination produces survey-ready material. Trying to do both in one prompt yields tangled output that is neither a clean table nor a good paragraph.
Schemas change mid-project: you realize you also need the "number of baselines compared," or the allowed method families grow. Handle evolution deliberately:
A common trap: adding an "optional" field and assuming old outputs are unaffected. They are not — the model's behavior on the other fields can shift when the schema changes. Re-validate everything after any schema edit.
When you extract from hundreds of documents, you cannot hand-check everything. Use a sampling plan:
This is the same discipline as data annotation in ML — because that is what you are doing: using the model as an annotator. Annotators get audited; so should models.
For your research: Structured-output prompting is how you turn LLMs into research instruments: extracting data from papers for meta-analyses, coding qualitative data, generating evaluation datasets, producing machine-readable annotations. Whatever schema you use, save it verbatim with your project materials — it is part of your method, and reviewers increasingly ask for it (Chapter 11).
For most researchers, the highest-value use of language models is not writing prose — it is writing code, analyzing data, and fixing bugs. A model that has read millions of Stack Overflow threads, GitHub repositories, and documentation pages is a formidable programming assistant. But code assistance has its own prompting discipline: precision about environment, minimal examples, and verification habits that prose tasks do not require.
The number one cause of useless code answers is missing environment context. The model does not know your Python version, your libraries, your data shapes, or your error message — unless you say so.
Before:
"How do I merge two dataframes?"
The model gives a generic pd.merge tutorial. Maybe it matches your situation, maybe not.
After:
"I am using Python 3.11 and pandas 2.1. I have df1 with columns [user_id, purchase_date] (50k rows) and df2 with columns [user_id, age, city] (48k rows). user_id is unique in df2 but not in df1. I want every row of df1 with the matching age and city, keeping all df1 rows even when there is no match. Give me the exact merge call and one line to verify the row count afterward."
The "after" version specifies versions, shapes, key properties, the desired join semantics in plain words, and asks for a verification step. The answer is now a precise, checkable solution instead of a tutorial.
Environment checklist for code prompts: language + version, key libraries + versions, what the data looks like (shapes, column names, types), what you tried, and the exact error message (full traceback, not your summary of it).
When asking for new code, specify: the task, the input/output shapes, the libraries to use (and not use), edge cases, and the form of the answer.
Before:
"Write code to clean this dataset."
After:
"Write a Python function clean_df(df) using pandas that: (1) drops rows where all values are missing, (2) fills numeric missing values with the column median and categorical with the mode, (3) strips whitespace from all string columns, (4) returns the cleaned dataframe plus a dict reporting how many values were filled per column. Include a docstring. Do not modify df in place. Here is df.dtypes and df.head(3): …"
The difference: the "after" prompt is a specification a junior developer could implement without asking questions. The more your prompt reads like a function spec, the better the generated code.
Ask for the test, not just the code. Adding "include a small test with sample input and the expected output" does two things: it forces the model to think about correctness, and it gives you something runnable to verify against.
Data analysis prompts fail when the question is vague ("analyze this data") and succeed when they mirror a real analysis plan. Structure them as: question → data description → steps → output format.
Worked example:
I am exploring a student survey dataset (n=1,200) for my thesis on study habits.
Columns: hours_studied (float), gpa (float), major (categorical, 6 values), sleep_hours (float), uses_ai_tools (yes/no).
Write Python (pandas + seaborn) that:
1. Prints a summary table: for each column, dtype, % missing, and (for numeric) mean/std/min/max.
2. Plots gpa vs hours_studied as a scatterplot colored by uses_ai_tools, with a regression line per group.
3. Computes the correlation matrix of numeric columns and prints it.
4. Runs a t-test comparing gpa between uses_ai_tools groups and prints the statistic, p-value, and a one-sentence plain-language interpretation.
Rules: one code block, fully runnable, no placeholders. Add a comment above each section saying what it does. After the code, list 2 caveats about interpreting these results (confounders, correlation vs causation).
This prompt works because it is an analysis plan, not a wish. Note the final instruction — asking for caveats — which turns the model from a code generator into a junior collaborator who flags interpretation risks. For researchers, that last line is often the most valuable part.
Iterate on results, not just code. When the output looks wrong, paste the result back: "The correlation matrix shows NaN for sleep_hours — here is df['sleep_hours'].describe(). What is wrong?" Debugging with the model is a conversation; each turn should add new evidence (outputs, errors, data samples), not just repeat the request louder.
Debugging prompts are where the environment checklist pays off most. A good debugging prompt contains:
Before:
"My code doesn't work. Here's my script: [200 lines]"
After:
"Expected: the function should return a list of 10 floats. Actual: it returns a list of 10 None values, no error. Minimal example: [15 lines]. I already checked that the input list is non-empty. What is the most likely cause? Explain in 2-3 sentences, then show the fixed function."
The "after" prompt respects a key principle: minimize before you ask. Stripping your problem to 15 lines often reveals the bug yourself — and when it does not, the small example lets the model (or a colleague) see the problem instantly. Researchers who paste 200 lines get slow, generic answers; researchers who paste 15 lines get the bug named.
Ask for the why, then the fix. "Explain why this fails, then fix it" produces better fixes than "fix this," because the explanation forces the model to diagnose before prescribing — a small, reliable application of chain-of-thought to debugging.
The rubber-duck pattern. Sometimes the best debugging prompt is not a question at all:
"I am going to explain my code to you line by line. After each section, point out anything that looks wrong or any assumption I might be making. Do not suggest fixes until I finish. Ready?"
Explaining your own code to an attentive listener is a classic debugging technique; the model makes an excellent, patient listener.
Model-generated code must be treated like code from an untrusted contributor:
A sobering fact worth remembering: models generate code that looks right far more reliably than code that is right. The fluency is the danger. Your verification habits are the safety net.
The lesson: the quality of debugging help you receive is proportional to the quality of evidence you provide. Expected vs. actual, minimal code, data checks, versions. Every time.
The same discipline applies to database work. SQL prompts fail when the schema is missing — the model cannot guess your table names.
Before: "Write SQL to find the top customers." → Generic query with invented table/column names.
After:
I use PostgreSQL 15. Schema:
customers(customer_id PK, name, signup_date)
orders(order_id PK, customer_id FK, order_date, total NUMERIC)
Write a query returning the 10 customers by total order value in 2025, with columns: name, order_count, total_spent. Exclude customers with no 2025 orders. After the query, add one sentence explaining any assumption you made.
The "explain any assumption" line is the SQL equivalent of Chapter 4's assumption-surfacing — it tells you where the query might not match your intent (e.g., it assumed total is pre-tax).
Iterating on queries: when the result looks wrong, paste the result plus the query: "This returns 0 rows; here is SELECT COUNT(*) FROM orders WHERE order_date >= '2025-01-01' → 14,203. What is wrong with the join?" Evidence-driven debugging again.
Watch how a researcher works through an analysis conversationally, with each turn adding evidence:
Turn 1 — plan:
"I have a CSV of 5,000 thesis survey responses (columns: age, department, satisfaction_1_5, hours_studied, comments). I want to know what drives satisfaction. Suggest an analysis plan: 5 steps ordered by effort, each with the method and what it would tell me. Keep each step to 2 sentences."
The model returns: descriptives → correlations → grouped comparisons → regression → text analysis of comments. The researcher now has a roadmap and can approve or redirect before any code exists. Planning before coding prevents the common failure of generating 200 lines that answer the wrong question.
Turn 2 — code for step 1–2:
"Write Python (pandas) for steps 1 and 2 of the plan. Load from survey.csv. Print: (a) missing-value percentages per column, (b) mean satisfaction by department as a bar chart, (c) correlation of numeric columns with satisfaction_1_5. One code block, runnable, with section comments."
Turn 3 — interpret with the output:
"Here is the output: [pasted results]. Satisfaction correlates with hours_studied at r=0.31, but department CS has both highest hours and highest satisfaction. Is the correlation confounded by department? Suggest the specific next analysis to check, with code."
Notice what happened: the researcher did not ask "is this right?" (too vague). They named the suspected confound and asked for the specific check. The model suggests partial correlation / grouped regression — and the analysis deepens correctly.
Turn 4 — the comments:
"Now step 5: the comments column has 5,000 free-text responses. Propose a coding scheme with 5-7 themes for what drives satisfaction, then write code to apply it using keyword rules (not an LLM), and report theme frequencies. Flag that keyword coding is crude and suggest how to validate on a 100-comment sample."
The researcher explicitly scopes the method (keyword rules, not LLM — cheaper and transparent) and demands a validation plan. This is researcher-grade prompting: the human owns the methodology; the model executes.
Add this standing instruction to your code/analysis system prompt (Chapter 5):
When generating statistical code or interpreting results, always flag: (1) correlation vs. causation issues, (2) multiple-comparison problems if many tests are run, (3) assumption violations you can detect (normality, independence), (4) any place where the code could leak information (e.g., scaling before train/test split). Put flags in a "CAUTIONS" section after the code.
You are outsourcing vigilance, not judgment — the model surfaces candidate issues, you decide which matter. In practice this catches real mistakes: the model will notice you normalized before splitting, or that you ran 20 t-tests without correction. It is not infallible, but it is a second pair of eyes that never gets tired.
Before/after — with the cautions instruction:
For your research: Build a personal "code review" system prompt (Chapter 5) for programming sessions: language versions, "explain before fixing," "flag statistical pitfalls," "no deprecated APIs." And establish a lab norm: any model-generated code that enters a paper's experiments or a shared codebase must be run, edge-case tested, and read by a human. Put that norm in writing — reviewers and future-you will thank you.
Every prompt engineer needs a debugging method, because every prompt fails sometimes. The amateur response to failure is to rewrite the whole prompt and hope. The professional response is to diagnose: observe the failure precisely, locate the faulty block, fix that block, and re-test. This chapter gives you the checklist.
Before changing anything, write one sentence describing what is wrong. Not "it's bad" — the specific defect:
A precise failure description points at the fix. "Too long" → output spec. "Extra key" → schema constraints. "Wrong for n=1" → edge cases in the instruction. Most of the checklist below is just matching symptoms to blocks.
Run through these in order. Fix the first one that matches, re-test, and only then continue.
Symptom: the answer is on-topic but generic, or answers a neighboring question. Test: could a competent human follow your instruction without asking clarifying questions? If not, the instruction is vague. Fix: replace vague verbs with concrete deliverables (Chapter 2). "Analyze" → "list three causes with evidence." "Improve" → "rewrite for clarity, keeping all technical terms, max 150 words."
Symptom: wrong length, wrong format, missing sections, chatty preamble before the JSON. Fix: state format, length, structure, and prohibitions explicitly. "Return only the table, no intro or outro text." Length failures are the easiest to fix in all of prompting — almost always a missing length constraint.
Symptom: the answer assumes the wrong audience, purpose, or background; advice is technically right but useless for your situation. Fix: add who, why, and for-whom (Chapter 2, Block 2). One or two sentences of context often fix this completely.
Symptom: the model confuses your instructions with the input text; it follows instructions found inside a pasted article; it answers about the wrong part of a long input. Fix: delimit with tags/fences; trim to the relevant portion; if the input is long, point at the relevant part ("focus on the Methods section").
Symptom: few-shot outputs mimic the examples' flaws — wrong format details, inconsistent labels, sloppy style. Fix: audit your examples. Every example output must be exactly the output you want. Fix formatting inconsistencies, remove ambiguous examples, ensure all categories are represented. Remember: the model copies your examples faithfully, including their mistakes.
Symptom: the model does part of the task well and ignores or mangles the rest; quality degrades toward the end of long outputs. Fix: decompose. Split into sequential prompts: first extract, then analyze, then summarize. Feed the output of step 1 as the input of step 2. Multi-step pipelines of simple prompts beat single mega-prompts almost every time.
Symptom: wrong answers on questions requiring calculation, logic, or multi-hop inference; the model jumps to conclusions. Fix: add chain-of-thought (Chapter 4): "think step by step," or structure the reasoning with numbered sub-questions. Verify the trace.
Symptom: the prompt is precise, examples are clean, and the model still fails — especially on very recent facts, obscure details, or tasks requiring exact computation. Fix: recognize model limits. For recent facts, provide the facts in the prompt (retrieval beats memory). For exact computation, ask the model to write code and run it rather than doing arithmetic in prose. For genuinely hard reasoning, try a stronger model. No prompt fixes a capability gap.
Symptom: you need the same answer every time (extraction, classification) but get variation; or you need creative variety but get the same safe answer. Fix: for determinism, set temperature to 0 (API) or ask for the most direct answer. For variety, raise temperature and ask for multiple options ("give me 5 different…"). Note: even at temperature 0, models are not perfectly deterministic — close, but not guaranteed.
Symptom: erratic behavior, the model seems to "pick" which instruction to follow. Fix: re-read your prompt for conflicts — system prompt vs. user prompt (Chapter 5), or two user instructions that clash ("be concise" + "explain in detail"). Resolve the conflict explicitly; when you truly want both, specify the balance ("a detailed explanation in under 200 words").
Task: extract experimental results from paper sections into a table.
Attempt 1 prompt: "Extract the results into a table." → Output: a paragraph describing results. Failure: prose instead of table. Diagnosis: checklist #2 — no output spec. Fix: "Return a markdown table with columns: Experiment | Dataset | Metric | Score."
Attempt 2 → Output: table, but scores are rounded differently than the paper and one row is invented. Failure: hallucinated values. Diagnosis: #4/#8 — the model filled gaps from memory instead of the text. Fix: add "Use only values stated verbatim in the text. If a value is not stated, write 'not reported'. Never infer or round."
Attempt 3 → Output: correct table. Three iterations, each fixing exactly one diagnosed fault. Total time: five minutes. This is what systematic debugging feels like — fast, because you never guess.
For recurring tasks, keep a tiny log: date, task, prompt version, failure observed, fix applied. After a month you will have a personal catalog of your most common failure modes — and you will stop making them. For lab use, a shared failure log is even better: every member's debugging makes everyone faster. This log is also the raw material for Chapter 10's evals: each logged failure is a test case.
Not every prompt is worth perfecting. Stop when: - The output is good enough for its purpose (a draft you will edit anyway does not need a perfect prompt). - You have iterated 4–5 times on the same failure — the problem is likely the task's difficulty or the model's capability (#8), not your wording. - The cost of more iteration exceeds the cost of manual fixing. A prompt that gets you 90% of the way on a one-off task is a success; polish is for templates you will reuse (Chapter 9).
Session 2: the model is confidently wrong about a fact.
Task: "List three optimizers suitable for sparse gradients and their original papers."
Attempt 1 → Output lists Adam, AdaGrad, RMSprop with plausible-looking citations — two of which have wrong years. Failure (one sentence): citations have incorrect years. Diagnosis: checklist #8 — the model is recalling from memory, and memory of bibliographic details is unreliable. This is not fixable by rewording. Fix: change the approach — "List the three optimizers. For each, describe the key idea in one sentence. Do NOT provide citations; write [verify citation] instead." Then look up the real citations yourself. The debugged prompt stops asking the model to do something it cannot do reliably.
Lesson: some failures are not prompt bugs but task-model mismatches. The checklist's step 8 exists to catch exactly this: stop tuning the prompt, change what you ask for.
Session 3: the slow drift.
Task: a weekly template that summarizes new arXiv papers in your area. It worked for a month, then outputs got noticeably worse — vaguer, missing the "relevance" section.
Failure: gradual quality decay on a previously good template. Diagnosis: this is not in the ten steps as a prompt bug — it is a model change (providers update models) or input drift (papers got longer/more varied). Fix: (a) check whether the model version changed; pin a version if your access allows; (b) re-run your eval (Chapter 10) to quantify the drop; (c) inspect recent failures — in this case, abstracts had gotten longer and the fixed length budget no longer fit, so the model dropped the relevance section to fit. Fix: raise the length budget and re-prioritize ("if space is tight, cut background, never cut relevance").
Lesson: deployed prompts need monitoring. A template is not "done" when it works once; it is done when it works repeatedly, which requires periodic re-evaluation.
Session 4: the contradiction hunt.
Prompt: "Be thorough but concise. Give a complete detailed explanation in as few words as possible."
Output: erratic — sometimes a paragraph, sometimes an essay. Failure: unpredictable length and depth. Diagnosis: checklist #10 — "thorough/complete/detailed" and "concise/as few words as possible" pull in opposite directions, and the model resolves the conflict differently each run. Fix: specify the actual trade-off you want: "Explain completely but tightly: cover all 4 causes, max 40 words each, total under 200 words." The fix replaces two warring adjectives with a concrete budget that satisfies both intents.
When a long, complex prompt misbehaves and the checklist does not immediately identify the cause, bisect: cut the prompt in half and test each half separately.
This is the same technique programmers use to find bugs in code, and it works because prompt failures are usually caused by one bad sentence interacting with the rest. Example: a researcher bisected a failing extraction prompt and found the culprit was a single example whose output used a slightly different date format — the model was faithfully copying the inconsistency (checklist #5). One fixed example, problem gone.
Bisecting rules: change one thing per test; keep a note of what you removed (so you can restore); stop when the failure is isolated to a single sentence or example. Most diagnoses take 3–4 bisections — under ten minutes.
For quick reference, the ten steps compressed:
Print this. Tape it near your desk. In a month you will not need it — the diagnoses will be automatic.
When you are in a hurry, run this compressed version before the full checklist:
These four steps fix the majority of everyday failures in under five minutes. The full ten-step checklist is for the stubborn remainder.
Sometimes the right debugging move is not to fix the prompt but to recognize it cannot be fixed — the request is genuinely ambiguous. "Summarize this" of a 50-page thesis could mean ten different things. Teach your system prompt (Chapter 5) to ask:
When a request has more than one reasonable interpretation that would lead to substantially different outputs, ask ONE clarifying question with 2-3 concrete options before answering. Do not ask about trivial details.
Before: "Analyze this dataset." → model guesses, usually wrong about what "analyze" means to you. After (model asks): "I can analyze this several ways — which is most useful? (a) data quality report (missing values, outliers), (b) exploratory patterns related to a target variable, (c) a specific hypothesis test. Which, and what is the target variable if (b) or (c)?"
One question, three options, and the subsequent analysis hits the mark. The skill here is meta: prompt the model to debug your request before executing it. For genuinely ambiguous one-off tasks, this beats even the best static prompt — because the missing information was in your head, not in any wording trick.
For your research: The checklist is also a lab skill. When a colleague says "the model gave me garbage," walk them through the ten steps instead of rewriting their prompt for them. Teaching diagnosis scales; doing their prompting does not. And every failure you log is a candidate test case for the mini-evals in Chapter 10.
A good prompt used once is a trick. A good prompt saved, named, versioned, and reused a hundred times is infrastructure. This chapter is about turning your best prompts into templates and your templates into a library — the step that converts prompting skill into lasting research productivity.
A template is a prompt with named slots for the parts that change. Compare:
One-off prompt:
"Summarize the following abstract in 3 bullets for my thesis on efficient transformers: …"
Template:
Summarize the following {document_type} in {n} bullets for my thesis on {thesis_topic}. Each bullet max {max_words} words. Focus on: {focus_areas}.
<{document_type}>
{input_text}
</{document_type}>
The slots ({document_type}, {n}, {thesis_topic}, {focus_areas}, {input_text}) are the variables; everything else is the tested, debugged prompt you keep. Fill the slots per use — by hand, or with a script when processing many items.
Template design rules:
1. Slot only what varies. Fixed instructions stay fixed; that is the point.
2. Give slots defaults. {n=3}, {max_words=30} — so quick uses need no decisions.
3. Constrain slot values. If {document_type} should be one of abstract/introduction/conclusion, say so in a comment. Unconstrained slots invite inconsistent use.
4. Keep the template self-contained. Anyone (including future-you) should be able to use it without tribal knowledge. A one-line comment at the top stating purpose and expected slot values is enough.
Here is a complete, reusable template for screening papers — the kind of thing a survey author runs hundreds of times:
# TEMPLATE: paper-screen v1.2
# Purpose: decide INCLUDE/EXCLUDE for a literature survey
# Slots: topic, inclusion_criteria, abstract
You are screening papers for a survey on {topic}.
INCLUSION CRITERIA (a paper must meet ALL to be included):
{inclusion_criteria}
Read the abstract below. Think step by step: (1) what does the paper claim to contribute? (2) does it meet each criterion? Quote the relevant phrase for each. (3) final decision.
Return ONLY this JSON:
{"decision": "INCLUDE" or "EXCLUDE", "criterion_quotes": {"criterion_1": "<quote or 'not addressed>", ...}, "confidence": "high/medium/low", "one_line_reason": "<max 20 words>"}
Rules: judge only what the abstract states. Do not infer beyond the text. When in doubt, EXCLUDE with confidence low.
<abstract>
{abstract}
</abstract>
Note the version number (v1.2). Templates evolve; versioning lets you compare v1.2 against v1.3 on your eval set (Chapter 10) and know which is better — instead of arguing from memory.
A prompt library is a collection of templates organized for retrieval. For an individual researcher, a single markdown file or a folder of text files is enough. For a lab, a shared document or repository works.
Organize by task, not by cleverness. Folders like summarization/, coding/, writing/, data-extraction/, review/ beat folders like clever-tricks/. You retrieve by the job you need done.
Each library entry should contain:
- Name and version (abstract-polisher v2.0)
- One-line purpose
- The template itself with documented slots
- 1–2 filled examples (input → output) showing correct use
- Known limitations ("fails on abstracts over 500 words; trim first")
- Date last tested + model used
Starter set for a researcher (build these first; they cover 80% of needs):
1. summarize-for-purpose — summarize any text for a stated audience and purpose
2. paper-screen — include/exclude with criteria (above)
3. extract-to-schema — extract fields into your JSON schema (Chapter 6)
4. code-debug — the debugging evidence template (Chapter 7)
5. reviewer-simulator — critical review with severity-ordered issues
6. explain-concept — explain at a stated level with analogy + formalism
7. rebuttal-drafter — structured response to a reviewer comment
8. email-professional — rewrite draft text in professional tone (for the admin side of research life)
The real power of templates appears when you apply one template to many inputs programmatically. The pattern:
Example: screening 300 abstracts for a survey. Manual screening takes days; a templated pipeline takes an hour plus verification time. But — and this is critical — the pipeline is only as trustworthy as your eval. Never run a template at scale on the basis of "it worked on two examples." Chapter 10 tells you how to earn that trust.
If your lab shares templates, treat them like shared code: - README explaining each template's purpose and slots. - Changelog when templates change ("v1.3: added 'never infer beyond the text' after observing hallucinated criteria matches"). - Ownership: someone maintains each template; unmaintained templates rot as models change.
And when templates contribute to published work, they belong in the paper or its supplementary material (Chapter 11). A template is a method; methods get reported.
A ready-to-copy starter gallery lives in the Learning Dashboard at the end of this book — eight templates covering summarization, screening, extraction, debugging, reviewing, explaining, rebuttals, and structured self-critique. Copy them into your library and adapt the slots to your field.
Template: related-work gap finder
# TEMPLATE: gap-finder v1.0
# Purpose: identify what a set of papers collectively misses
# Slots: paper_summaries (one line each), my_approach (1-2 sentences)
I am writing a paper using this approach: {my_approach}.
Here are one-line summaries of the closest related papers:
{paper_summaries}
Think step by step: (1) What problem does each paper solve? (2) What limitation or scope restriction does each acknowledge or imply? (3) What problem does NONE of them solve that my approach addresses?
Return:
- "shared_limitations": 2-3 bullets, each citing which papers share it
- "candidate_gaps": 2-3 bullets, each a gap my approach could claim — mark each [strong] or [weak] with one sentence why
- "caution": one paragraph on the strongest counter-argument (why a reviewer might say the gap does not matter) and how to address it
Rules: ground every claim in the summaries above. Do not invent paper content.
The "caution" section is the template's best feature: it forces the model to steelman the opposition, which is exactly what a good related-work section must preempt.
Template: experiment planner
# TEMPLATE: experiment-planner v1.0
# Purpose: turn a research question into an experiment plan
# Slots: question, resources, constraints
Research question: {question}
Available resources: {resources}
Constraints: {constraints}
Design the MINIMAL experiment that answers the question. Return:
1. "hypothesis": one sentence, falsifiable
2. "design": the experimental setup in 5 numbered steps
3. "metrics": what to measure and why each metric matters (max 3 metrics)
4. "controls": what to hold constant or compare against, and why
5. "failure_modes": 3 ways the experiment could give a misleading answer, and how to prevent each
6. "stop_rule": when to stop and declare the question answered (or unanswerable with these resources)
Rules: respect the constraints strictly. Prefer the smallest experiment that could work over the most impressive one.
The "stop_rule" and "failure_modes" sections encode hard-won experimental wisdom: most failed research projects fail from vague stopping criteria and unexamined failure modes, not from lack of effort.
A library nobody maintains becomes a junk drawer. Keep it alive with a quarterly 30-minute review:
Library README sketch (for a lab-shared library):
# Lab Prompt Library
Each folder = one task. Each template file starts with: purpose, slots, version, last-tested (model + date), known limitations.
## Folders
- screening/ — paper-screen v1.3 (tested 2026-09 on llama-3.1-8b)
- extraction/ — extract-to-schema v2.1, table-compare v1.0
- writing/ — abstract-polish v1.2, rebuttal-draft v1.0
- code/ — debug-evidence v1.1, analysis-plan v1.0
## Rules
- Never commit a template without a filled example.
- Bump the version on ANY wording change; note it in CHANGELOG.md.
- Templates used in papers: freeze the version and cite it in the appendix.
This is deliberately boring infrastructure — and boring infrastructure is what makes a lab's prompting reliable instead of folkloric.
If you publish templates (blog post, GitHub, paper supplement), include what a stranger needs: purpose, slot documentation, 2+ filled examples with real inputs, known limitations, and the model + date tested. The most common failure of shared prompts is missing context that was obvious to the author ("oh, that template assumes abstracts under 300 words"). Write the README for someone who knows nothing about your project.
And expect adaptation: a good template is a starting point, not a finished product. The best shared templates invite modification — clear slots and documented assumptions make that easy.
Watch for these when building or reviewing templates:
{input} slot containing five different things (the paper, your notes, the reviewer's comment, your question). Split into named slots ({paper}, {notes}, {question}) — named slots get filled consistently; blob slots get filled randomly.Templates carry you far, but recognize the graduation points:
Graduating is not abandoning prompting — it is putting prompting inside a larger engineered system. The best agent builders are the best prompt engineers, because every agent is ultimately a collection of prompts with control flow around them.
Documenting template decisions. For each template in your library, keep a one-paragraph "design note": why the slots are what they are, which alternatives you tried, what the eval showed. Example: "Tried 5-shot for paper-screen; 3-shot scored identically on dev (28/30 both) so kept 3 for cost. Negative example added after spam-urgency failures in v1.1." These notes prevent future-you from re-running experiments past-you already ran — and they are the raw material for the methods section when the template enters a paper.
For your research: This week, convert your three most-used prompts into templates: add slots, defaults, version numbers, and one filled example each. Store them where you will actually find them. Next month, you will have stopped re-typing instructions and started reusing tested ones — and your prompting will have become measurably more consistent, which is exactly what Chapter 10 will let you prove.
Everything so far has been craft: write better prompts, debug them, templatize them. This chapter is about science: how do you know a prompt is good? How do you know v1.3 beats v1.2? The answer is evaluation — "evals" — small, honest test sets that turn prompt opinions into prompt evidence.
Without evals, prompt improvement is vibes. You try a new wording, it works on the one example you tested, you adopt it — and silently break three cases you did not test. Every prompt engineer has done this. Evals are the seatbelt: a fixed set of test cases you run every time the prompt changes, so improvements are real and regressions are caught.
You do not need a thousand examples or a GPU cluster. For most research prompts, 20–50 well-chosen test cases beat a thousand random ones. What matters is that the cases cover the situations you care about, including the tricky ones.
Step 1: Collect test cases from real use. Your failure log (Chapter 8) is the seed. Each logged failure becomes a test case: the input that broke, plus the correct output. Add typical cases (the common, easy ones) and edge cases (the weird inputs from Chapter 6). Aim for 20+ to start.
Step 2: Define the correct answer for each case. This is the hard, valuable work. For classification/extraction tasks, label them by hand — this is your "gold set." For open-ended tasks (summaries, reviews), define a rubric instead of a single right answer: e.g., "summary must mention the method, must be under 100 words, must not contain claims absent from the source."
Step 3: Define how you score. Three levels, in order of rigor: - Exact/structural checks (automatic): does the JSON parse? Are all required keys present? Is the label one of the allowed values? Is the length within bounds? These run in code, instantly, on every output. - Rubric checks (human or model-assisted): does the summary mention the method? Is the criticism actionable? Score 0/1 per rubric item. - Human preference (for quality judgments): show two outputs blind, pick the better one. Slow; reserve for final comparisons.
Step 4: Run the prompt on all cases, score, record. This is your baseline score. Write it down with the prompt version, model name, and date.
Step 5: Change one thing, re-run, compare. This is the whole game. One variable at a time — new instruction wording, added example, different output spec — then the full eval. If the score improves with no regressions on previously-passing cases, keep the change.
Recall the paper-screen template from Chapter 9. Here is its mini-eval:
Test set (24 abstracts): - 8 clear INCLUDEs (propose efficient attention mechanisms) - 8 clear EXCLUDEs (apply attention, survey it, or unrelated) - 8 near-misses (the hard ones: papers that mention proposing something but only apply; papers with "efficient" in the title about something else; non-English abstracts; abstracts that are actually introductions)
Gold labels: hand-labeled by you, with one-line reasons.
Scoring: - Automatic: JSON parses; decision is INCLUDE/EXCLUDE; confidence is high/medium/low. (Structural.) - Correctness: decision matches gold label. (Accuracy = correct / 24.) - Calibration check: on wrong decisions, was confidence "low"? (A model that is wrong but uncertain is safer than one that is wrong and confident.)
Running it: a script fills the template per abstract, calls the API at temperature 0, parses the JSON, and prints a table: accuracy overall, accuracy on near-misses, structural failure count, and the list of misclassified cases for inspection.
Iterating: v1.2 scores 20/24. The 4 errors are all near-misses where the model inferred beyond the abstract. You add "judge only what the abstract states; do not infer" → v1.3 scores 23/24 with no regressions. That sentence earned its place with evidence, not intuition.
For rubric-style scoring at scale, you can use a strong model as the judge: give it the rubric, the input, and the output, and ask it to score. This works reasonably for well-defined rubrics ("does the summary mention the method? yes/no") and poorly for subtle quality judgments. Rules:
Not every prompt needs 95%. Match rigor to stakes: - One-off draft help: no eval needed; eyeball the output. - Template used weekly: 20-case eval, re-run on changes. - Template in a published paper's method: 50+ cases, held-out test set, full reporting (Chapter 11). - Automated pipeline over hundreds of items: eval plus ongoing sampling — re-check a random sample of live outputs monthly, because data drifts.
Rubrics are where evals of open-ended tasks succeed or fail. A vague rubric ("is the summary good?") just moves the subjectivity around. A good rubric has 4–6 binary or 3-point items, each checkable from the output alone.
Task: evaluate AI-generated summaries of paper abstracts for a survey.
Bad rubric: "Rate the summary 1–5 for quality." (What is quality? Different raters disagree; the model-judge disagrees with itself across runs.)
Good rubric (each item scored 0/1):
Why this works: items 1–3 are mechanical (a script or judge can score them reliably); item 4 is the critical anti-hallucination check; item 5 ties the output to its purpose. Total score /5 per summary. Two raters scoring 20 summaries with this rubric will agree far more than with a 1–5 "quality" scale — and agreement is what makes the eval trustworthy.
Rubric design rules: - Each item must be decidable from the output + source alone (no mind-reading). - Prefer binary (present/absent) over scales; use 0/1/2 only when partial credit is meaningful. - Include at least one "negative" item (absence of bad behavior: no hallucination, no forbidden content). - Pilot the rubric: score 10 outputs yourself, note where you hesitated, rewrite those items.
Suppose you want a strong model to score summaries against the rubric above at scale:
You are an evaluator. Score the SUMMARY against the rubric using ONLY the ABSTRACT as ground truth.
Rubric (score each 0 or 1):
1. Names the main method/approach.
2. States the key result with magnitude.
3. 100 words or fewer.
4. Every claim in the summary is supported by the abstract.
5. Ends with a relevance judgment for efficient-attention research.
Return ONLY JSON: {"r1": 0/1, "r2": 0/1, "r3": 0/1, "r4": 0/1, "r5": 0/1, "total": <sum>, "notes": "<10 words on any 0 score>"}.
<abstract>{abstract}</abstract>
<summary>{summary}</summary>
Validating the judge: score 30 summaries yourself with the rubric, run the judge on the same 30, and compute agreement per item. Typical outcome: items 1–3 agree ~95%, item 4 (hallucination detection) agrees ~80% — the judge is lenient about subtle rephrasings. Decision: trust the judge for items 1–3 and 5, but hand-check a sample for item 4. This is honest, calibrated automation: you know exactly what the judge is good at.
You do not need statistics-heavy power analysis for prompt evals, but some intuition helps:
The near-miss principle: 10 well-chosen hard cases teach you more than 100 easy ones. When expanding an eval, add cases that previously failed or nearly failed — not more of what already passes. An eval where everything passes is a comfort blanket, not a measurement.
Keep a simple log — a spreadsheet or markdown table:
| Date | Prompt version | Model | Temp | Dev (n=30) | Test (n=50) | Notes |
|---|---|---|---|---|---|---|
| 2026-09-01 | screen v1.2 | llama-3.1-8b | 0 | 27/30 | 44/50 | baseline |
| 2026-09-14 | screen v1.3 | llama-3.1-8b | 0 | 29/30 | 47/50 | added "judge only the abstract" |
| 2026-10-02 | screen v1.3 | llama-3.1-8b | 0 | 28/30 | 45/50 | model updated? investigating |
The third row is the payoff: a score drop with no prompt change signals external drift (model update, data change), not a prompt bug. Without the log, you would have "fixed" a prompt that was never broken.
"I ran both prompts on the test set; the new one scored higher" is a good start, but prompt scores are noisy — the same prompt can vary run to run. For a trustworthy comparison:
Worked example: v1.2: 44/50. v1.3: 47/50. Disagreement analysis: v1.3 fixed 5 cases, broke 2. The 2 broken cases are both non-English abstracts — v1.3's new "judge only the abstract" sentence interacts badly with the translation instruction. Fix: scope the sentence ("judge only the abstract text as written"). v1.4: 48/50, zero regressions. The paired analysis caught a regression the headline numbers hid.
Numbers miss things. Complement every eval with a qualitative pass: read 10–15 outputs in full and ask:
Keep notes from these reads; they generate the hypotheses your next quantitative eval tests. The best eval regimes alternate: qualitative reading suggests what to measure, quantitative evals measure it, and the numbers send you back to reading.
For your research: Pick one prompt you use repeatedly and build a 20-case eval this week. It takes an afternoon, and it will change how you think about prompting: from "this feels better" to "this scores 23/24 vs. 20/24 on the held-out set." That sentence — with numbers — is also exactly the kind of evidence that strengthens a methods section or a workshop paper on prompt-based techniques (Chapter 12).
Here is an uncomfortable truth about much published work that uses language models: the methods sections say things like "we used ChatGPT to generate summaries" or "prompts were designed to extract features." That is not a method. It is an anecdote. No reader can reproduce it, no reviewer can evaluate it, and in six months even the authors cannot reconstruct what they did.
This chapter is about doing better: treating prompts as research artifacts that get versioned, archived, and reported with the same care as code and data.
Three reasons, in increasing order of importance:
For any paper where prompting materially affects the results, report:
Minimal honest reporting (when the model was used for assistance, not as the experimental subject — e.g., polishing prose): a statement like "We used [model, version] to assist with prose editing and code debugging; all scientific content, experiments, and interpretations are the authors' own." Put it in the acknowledgments or a methods note, per your venue's policy.
"We extracted method metadata from 312 paper abstracts using Llama-3.1-8B-Instruct (Hugging Face revision 8e2a…) via the Transformers library, temperature 0, max 512 tokens, in March 2026. The extraction prompt template (Appendix A) specifies a fixed JSON schema with four keys; outputs failing schema validation (1.8%) were re-prompted once with the schema repeated, and the 3 remaining failures were hand-coded. Few-shot examples (n=3) were drawn from papers excluded from the analysis set. Extraction accuracy was validated against 60 hand-labeled abstracts (Cohen's κ = 0.87 between model output and annotator)."
Every sentence answers a reviewer's question before it is asked. Compare with "we used an LLM to extract metadata" — which tells the reviewer nothing and invites rejection.
During a project, prompts change constantly. Treat them accordingly: - Store prompt templates in text files in your project repository, not scattered across chat histories. - Version them (v1.0, v1.1…) with a changelog noting what changed and why ("v1.2: added missing-data rule after eval showed 6% hallucinated values"). - When you report results, cite the exact version used for each experiment. - Archive the chat/API logs for key runs. Most providers let you export conversation history; API users should log requests and responses.
A practical habit: at the start of each experiment, save a prompts/ folder snapshot with the date. Future-you, writing the paper six months later, will be able to answer "which prompt produced Table 3?" in seconds.
Beyond reproducibility, there are honesty obligations:
Anticipate these questions and answer them preemptively: - "Would the results hold with a different prompt?" → Show a prompt-ablation: your main prompt vs. a minimal variant, reporting both scores (this is Chapter 12's territory). - "How much of the performance is prompt engineering vs. the model?" → Report a zero-shot baseline alongside your engineered prompt. - "Can I reproduce this?" → Appendix with full prompts, model versions, parameters.
Papers that answer these questions get cited. Papers that hide them get questioned.
Expectations are converging across venues, and the direction is clear: disclose, document, reproduce.
Practical advice: check your target venue's current AI policy before submission (they evolve yearly). When in doubt, over-disclose: a clear assistance statement never hurt a paper, while a discovered omission can trigger an ethics inquiry.
A good appendix is skimmable and complete. Recommended structure:
Appendix A: Prompts and Model Configuration
A.1 Model and parameters
- Model: <exact name + version/revision>
- Access: API / local, dates of runs
- Parameters: temperature, top-p, max tokens, seed
A.2 System prompt (verbatim)
A.3 Task templates (verbatim, with version numbers)
- Template 1: extraction (v2.1)
- Template 2: screening (v1.3)
A.4 Few-shot examples
- Source of examples; statement of non-overlap with test data
- The examples themselves, verbatim
A.5 Post-processing and failure handling
- Parsing code (or pseudocode), repair policy, exclusion counts
A.6 Validation
- Eval design summary, gold-set size, agreement metrics
If prompts are long, the appendix can live in supplementary material with a pointer in the paper ("full prompts in supplementary §S3"). What matters is that a reader can find them, verbatim.
For a project that uses prompting in its experiments:
project/
├── prompts/
│ ├── CHANGELOG.md # v1.0 → v1.1: what changed and why
│ ├── system-prompt-v3.txt
│ ├── extract-metadata-v2.1.txt
│ ├── screen-papers-v1.3.txt
│ └── examples/
│ ├── few-shot-set-A.jsonl # examples used in prompts
│ └── gold-labels-dev.jsonl # dev set (NOT test)
├── evals/
│ ├── test-set.jsonl # held-out; never tune against this
│ └── results-log.md # the tracking table from Ch.10
└── paper/
└── appendix-prompts.md # generated from prompts/ at submission
Key disciplines this layout enforces: prompts are files (not chat history), versions are explicit, examples are separated from test data, and the appendix is generated from the same files that ran the experiments — so the paper cannot accidentally describe a different prompt than the one used.
Be honest about a hard truth: exact replicability of LLM experiments is often impossible. Providers retire models, update weights silently, and even temperature-0 runs can vary across hardware. What you owe the reader is not bit-identical reproduction but scientific reproducibility: enough documentation that an independent researcher, with the same prompts and a comparable model, can test whether your findings hold. Report the model version and date precisely, note known nondeterminism, and where possible use open-weight models with pinned revisions for the core experiments — they are the only ones that can still be run identically years later.
How do you cite a language model in the body of a paper? Current best practice:
Never bury the model version in a footnote nobody reads. If the model materially affects results, its identity belongs in the methods section proper.
Disclosing prompts used during peer review. If you use a language model to help write a peer review (where permitted by the venue — many now restrict this), the confidentiality stakes are higher: pasting an unpublished manuscript into a third-party service may violate the venue's confidentiality rules. Check the policy first; when in doubt, use a local model or do not use one at all. Prompt engineering skill does not override professional obligations.
When prompts, logs, or code accompany the paper as supplementary material, verify:
The README-to-table mapping is the item authors forget most and reviewers appreciate most: "Table 2 was produced by extract-metadata-v2.1.txt run on 2026-09-20; script: run_extraction.py (commit a1b2c3d)." That one line makes your work auditable in seconds.
For your research: Start a
prompts/folder in your current project today. Save every prompt that touches your results, with version numbers and dates. When you write the paper, you will have the appendix half-written — and you will be ahead of 90% of authors using language models.
The final step in this book's arc: turning prompting from a personal productivity skill into a research contribution. Prompt-based experiments — done rigorously — can be baselines, ablations, and even the core method of a publication. This chapter shows how to design them so reviewers take them seriously.
Many research questions today are, at heart, prompting questions: Can a language model perform this task? Which prompting strategy works best? What are the failure modes? These are legitimate empirical questions, and papers answering them get published — but only when they are executed with experimental rigor. A "we tried some prompts and it worked" paper gets rejected. A paper with controlled comparisons, ablations, error analysis, and honest reporting gets cited.
When your paper introduces a new method (a fine-tuned model, a new architecture, a pipeline), reviewers will ask: "How does a plain prompted LLM do?" A prompt-based baseline answers this — and it is often stronger than researchers expect, which makes it an important comparison.
How to build a credible prompt baseline: 1. Choose the strongest reasonable prompt, not a strawman. Use few-shot examples, chain-of-thought if the task needs reasoning, and structured output. A deliberately weak baseline ("we asked the model once with no instructions") is dishonest and reviewers see through it. 2. Document everything per Chapter 11. 3. Run it on the same test split as your method, with the same metric. 4. Report variance: run 3+ times (or with 3+ example sets for few-shot) and report mean and spread. Prompt results vary; a single run is anecdote.
Example framing in a paper: "As a baseline, we evaluate GPT-4o with a 5-shot chain-of-thought prompt (Appendix B) on the test set, achieving 78.2 ± 1.1% accuracy. Our fine-tuned model reaches 84.6%, a gain attributable to…" The baseline is strong, documented, and the comparison is meaningful.
An ablation removes or changes one component to measure its contribution. For prompts, ablate the techniques from this book:
| Ablation | Question it answers |
|---|---|
| Zero-shot vs. few-shot (k=1, 3, 5) | How much do examples help? Is there a saturation point? |
| With vs. without chain-of-thought | Does reasoning help on this task? (Chapter 4's when-it-hurts applies) |
| With vs. without role/system prompt | Does the persona add anything beyond the task instruction? |
| Example selection: random vs. curated | How sensitive is few-shot to which examples? |
| Output spec: free-form vs. structured | Does constraining the format change accuracy or just parsability? |
| Model size: small vs. large, same prompt | Is the capability in the model or the prompt? |
Reporting ablations: a simple table with one row per variant, same test set, same metric, plus a sentence per row interpreting the delta. This table is often the most cited part of a prompting paper, because it tells the next researcher exactly which techniques are worth their time.
Worked mini-example: "Ablating the system prompt changed accuracy from 81.3% to 80.9% (n=500, p=0.42, n.s.), while removing few-shot examples dropped it to 68.4%. Conclusion: examples carry the performance; the persona is cosmetic on this task." One paragraph, two numbers, a clear conclusion — this is publishable evidence.
Reviewers love error analysis because it shows you understand your method's limits. For prompt-based work:
This analysis does double duty: it strengthens the paper and it directly improves your prompt. Many researchers find that the error taxonomy becomes the paper's most useful contribution.
Sometimes the prompting strategy is the contribution: a new way to decompose a task, a novel self-correction loop, a template that unlocks a capability. Examples from the literature include chain-of-thought itself [2], ReAct's interleaving of reasoning and action [8], and self-consistency sampling [9].
If your contribution is a prompting method: - Name it precisely and describe it as an algorithm: inputs, steps, outputs. Pseudocode beats prose. - Compare against the right baselines: standard prompting, not just your method vs. nothing. The field's default question is "is this better than chain-of-thought with the same model?" — answer it. - Test across models and tasks. A technique that works on one model and one dataset is a curiosity; one that holds across three models and two task families is a result. - Analyze cost. Prompting methods that use 10× the tokens for +1% accuracy need to say so; reviewers will ask.
Prompting research has specific temptations: - Cherry-picked examples. Showing only the cases where your method shines. Fix: report full test-set numbers; examples illustrate, they do not substitute. - Test-set leakage in few-shot examples. Fix: disclose example sources (Chapter 11). - Prompt-tuning on the test set. Iterating your prompt against the test set, then reporting the test score as the result. Fix: dev/test separation (Chapter 10). This is the single most common methodological flaw in prompting papers. - Moving goalposts. Changing the metric or the task definition after seeing results. Fix: pre-register your eval design — even an informal written plan dated before the experiments counts.
A concrete 4-week plan:
You already know every technique in this plan. The book's job is done when you execute it.
The strongest defense against the methodological temptations listed earlier is also the simplest: write down the plan before running it. Pre-registration for a prompting experiment can be a one-page document:
Experiment: few-shot vs. zero-shot for limitation-extraction
Date planned: 2026-10-10
Test set: 200 limitation sections, gold labels by author A (blind to prompts)
Metric: exact-match accuracy on limitation type; secondary: F1
Variants (frozen before test run):
V0: zero-shot baseline (prompt text: ...)
V1: 3-shot, random examples from dev pool
V2: 3-shot, curated examples (the 3 highest-agreement dev cases)
V3: V2 + chain-of-thought
Analysis: McNemar's test for V0 vs V2; report all four, no cherry-picking
Stopping: one test run per variant; no prompt edits after seeing test scores
This takes 30 minutes and transforms the experiment from exploratory tinkering into confirmatory evidence. You can still explore — just label exploration as exploration and confirmation as confirmation. Reviewers can tell the difference, and they reward the latter.
Every prompting method has a price tag in tokens, and reviewers increasingly do the mental math. Report it plainly:
| Variant | Tokens/prompt (avg) | Calls for test set | Accuracy |
|---|---|---|---|
| V0 zero-shot | 180 | 200 | 71.5% |
| V2 few-shot | 1,150 | 200 | 79.0% |
| V3 few-shot + CoT | 2,900 | 200 | 81.2% |
| V3 + self-consistency (5×) | 14,500 | 1,000 | 82.8% |
Now the trade-off is visible: V3 gains 2.2 points over V2 at 2.5× the cost; self-consistency gains 1.6 more points at 5× again. Whether that is "worth it" depends on the application — but the paper that shows the table lets the reader decide, while the paper that hides it invites suspicion. Always report the cost of your best variant alongside its score. Cost is a result, not an embarrassment.
Many students sit on publishable prompting work without realizing it — a well-executed course project with proper evals is often one step from a workshop paper. The conversion checklist:
The bar is rigor, not scale. A 4-week project executed like Chapter 12 describes is genuinely publishable — and more importantly, it teaches the experimental habits that every larger project needs.
A final thought. Prompting research currently suffers from a reproducibility crisis in miniature: undocumented prompts, test-set tuning, cherry-picked examples, vanished model versions. Every paper that does it right — versioned prompts, frozen evals, honest ablations, reported costs — raises the standard slightly. As an early-career researcher, you have an advantage here: you are learning the rigorous habits before the sloppy ones calcify. The techniques in this book are not just productivity tricks. Practiced honestly, they are a contribution to how the field does science with language models.
"Chain-of-thought did not help on this task" is a publishable finding — if it is rigorous. The recipe for a credible negative result:
Some of the most cited prompting papers are boundary-mapping studies — they tell the field where a technique works and where it does not. If your experiments produce a clean negative result, write it up. The field learns as much from "here is where the magic stops" as from "here is more magic."
Sooner or later a reviewer will probe your prompting methodology. Keep this response skeleton ready:
We thank the reviewer for raising this. [Restate their concern precisely.]
To address it, we [ran/took the following step]:
- [Ablation / additional baseline / robustness check], on [test set], with [metric].
- Result: [numbers], which [supports/qualifies] our claim because [one sentence].
We have added this to [Section X / Appendix Y], including [the exact prompt text / version numbers].
[If a limitation is conceded:] We agree this is a limitation; we now discuss it in Section Z and note [what future work would address it].
The pattern: restate precisely, answer empirically, point to the new text, concede cleanly where warranted. Reviewers asking about prompting are usually asking "can I trust this number?" — a concrete additional experiment answers that better than any paragraph of argument.
For your research: Your prompting skill is now a research instrument. The difference between "I used a chatbot" and "we conducted controlled prompt-ablation experiments" is entirely in the rigor this book has taught: versioned templates, frozen evals, dev/test separation, full reporting. Reviewers reward that rigor — and the field needs it.
| Technique | Chapter | What it does | Best for | Cost | Watch out for |
|---|---|---|---|---|---|
| Clear anatomy (4 blocks) | 2 | Specifies instruction, context, input, output | Every prompt, every time | None | Skipping context/output spec |
| Zero-shot | 3 | Task from description alone | Standard tasks, quick drafts | Lowest | Ambiguous tasks fail silently |
| One-shot | 3 | One example shows format | Format-heavy extraction | Low | One example can mislead |
| Few-shot (2–8) | 3 | Examples teach task + standard | Classification, coding, screening | Medium | Example quality = output quality |
| Chain-of-thought | 4 | Step-by-step reasoning trace | Multi-step, verifiable tasks | High (tokens) | Confident wrong traces; overkill for simple tasks |
| Self-consistency | 4 | Majority vote over several traces | Hard reasoning, when accuracy matters most | Very high | Diminishing returns past ~5 samples |
| System prompt | 5 | Standing rules for all sessions | Consistent long-term behavior | One-time setup | Conflicts with task prompts |
| Role/persona | 5 | Bundled behavior + tone | Reviewer, tutor, critic modes | None | Authoritative tone ≠ expertise |
| Structured output | 6 | JSON/tables/schemas | Pipelines, evals, extraction | Low | Format-correct but fact-wrong outputs |
| Decomposition | 8 | Split mega-task into steps | Complex multi-part tasks | Medium (more calls) | Error propagation between steps |
| Templates | 9 | Reusable prompts with slots | Repeated tasks | One-time build | Template rot as models change |
| Mini-evals | 10 | Fixed test set per prompt | Any reused prompt | An afternoon to build | Tuning on the test set |
| LLM-as-judge | 10 | Model scores outputs vs. rubric | Scaling rubric scoring | Medium | Must validate vs. human scores |
| Symptom | Likely cause | Chapter | First fix to try |
|---|---|---|---|
| Answer is generic / vague | Vague instruction verb | 2, 8 | Replace verb with concrete deliverable |
| Wrong length or format | Missing output specification | 2, 8 | Add explicit length + format + prohibitions |
| Right topic, wrong angle | Missing context (audience, purpose) | 2, 8 | Add 1–2 sentences of context |
| Model follows text inside your pasted input | Undelimited input / injection | 2, 5 | Delimit with tags; declare input as data |
| Few-shot outputs copy example flaws | Bad examples | 3, 8 | Audit examples; make each output exemplary |
| Does half the task, mangles the rest | Task too big for one prompt | 8 | Decompose into sequential prompts |
| Wrong on math/logic/multi-hop | No reasoning steps | 4, 8 | Add chain-of-thought |
| Confident nonsense with nice steps | Rationalized trace / capability gap | 4, 8 | Verify independently; try stronger model or code |
| Invented citations or facts | Hallucination; no grounding | 5, 6 | "Never invent; use null/say unsure"; provide facts in prompt |
| Extra keys / chatty text around JSON | Unconstrained structured output | 6 | Pin schema; "return ONLY the JSON" |
| Inconsistent results across runs | Temperature / nondeterminism | 8 | Temperature 0 for deterministic tasks |
| Erratic, picks instructions randomly | Contradictory instructions | 5, 8 | Find and resolve the conflict explicitly |
| Good on 2 examples, fails at scale | No eval; overfit to examples | 10 | Build 20+ case eval before scaling |
| Worse after model update | Silent model change | 10 | Re-run eval; version-pin models when possible |
T1. Summarize for a purpose
Summarize the following {document_type} in {n=3} bullets (max {max_words=30} words each) for {audience}. Focus on: {focus}. Do not add information absent from the text.
<{document_type}>{text}</{document_type}>
T2. Screen include/exclude
Decide INCLUDE or EXCLUDE for a survey on {topic}. Criteria (must meet ALL): {criteria}.
Think step by step, quoting the abstract for each criterion. Return ONLY JSON: {"decision": "INCLUDE"/"EXCLUDE", "reason": "<20 words>", "confidence": "high/medium/low"}. Judge only the abstract text; when in doubt, EXCLUDE.
<abstract>{abstract}</abstract>
T3. Extract to schema
Extract from the text below into EXACTLY this JSON schema: {schema_json}. Return ONLY the JSON, no commentary. Unverifiable fields → null. Never invent values.
<text>{text}</text>
T4. Debug code
Expected: {expected}. Actual: {actual}. Environment: {env}. Minimal code: ```{code}```. Already tried: {tried}.
Explain the most likely cause in 2-3 sentences, then show the fixed code. Flag any assumption you are unsure about.
T5. Reviewer simulator
Act as a strict {venue} reviewer. Review the following {document_part}. List exactly {k=3} problems ordered by severity; each with a one-sentence explanation and a concrete fix. Do not praise. Be direct.
<{document_part}>{text}</{document_part}>
T6. Explain a concept
Explain {concept} in ~{n=200} words for {audience} who knows {prereqs}. Structure: (1) one analogy, (2) the core idea plainly, (3) the key formula or definition in LaTeX, (4) one common misconception. No code.
T7. Rebuttal drafter
A reviewer wrote: "{comment}". Think step by step: (1) is the criticism factually correct? (2) what evidence addresses it? (3) what change (if any) to the paper is warranted? Then draft a {n=4}-sentence rebuttal: acknowledge, present evidence, state the paper change or respectful disagreement. Professional tone.
T8. Self-critique pass
Critique your own previous answer below against these criteria: {criteria}. List each violation with a quote. Then rewrite the answer fixing all violations. Previous answer: <answer>{answer}</answer>
[1] T. B. Brown et al., "Language models are few-shot learners," in Advances in Neural Information Processing Systems, vol. 33, 2020, pp. 1877–1901.
[2] J. Wei et al., "Chain-of-thought prompting elicits reasoning in large language models," in Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 24824–24837.
[3] T. Kojima, S. S. Gu, M. Reid, Y. Matsuo, and Y. Iwasawa, "Large language models are zero-shot reasoners," in Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 22199–22213.
[4] L. Ouyang et al., "Training language models to follow instructions with human feedback," in Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 27730–27744.
[5] P. Liu et al., "Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing," ACM Computing Surveys, vol. 55, no. 9, pp. 1–35, 2023.
[6] L. Reynolds and K. McDonell, "Prompt programming for large language models: Beyond the few-shot paradigm," in Proc. CHI Conf. Human Factors in Computing Systems (Extended Abstracts), 2021, pp. 1–7.
[7] S. Yao et al., "ReAct: Synergizing reasoning and acting in language models," in Proc. Int. Conf. Learning Representations (ICLR), 2023.
[8] X. Wang et al., "Self-consistency improves chain of thought reasoning in language models," in Proc. Int. Conf. Learning Representations (ICLR), 2023.
[9] V. Sanh et al., "Multitask prompted training enables zero-shot task generalization," in Proc. Int. Conf. Learning Representations (ICLR), 2022.
[10] J. Wei et al., "Finetuned language models are zero-shot learners," in Proc. Int. Conf. Learning Representations (ICLR), 2022.
[11] J. White et al., "A prompt pattern catalog to enhance prompt engineering with ChatGPT," arXiv:2302.11382, 2023.
[12] E. M. Bender, T. Gebru, A. McMillan-Major, and S. Shmitchell, "On the dangers of stochastic parrots: Can language models be too big?" in Proc. ACM Conf. Fairness, Accountability, and Transparency (FAccT), 2021, pp. 610–623.
Exercise 1 — Rewrite the weak prompt. Take this prompt: "Write about neural networks." Rewrite it using all four anatomy blocks (instruction, context, input, output spec) for a concrete purpose of your choice (e.g., a 150-word explainer for undergraduates). Explain in 2–3 sentences what each added block fixed.
Exercise 2 — Spot the missing block. For each prompt below, name the weakest of the four blocks and fix it with one or two sentences: (a) "Is my abstract good?" (b) "Translate this." (c) "Summarize these 40 pages." (d) "Find the error: [code with no error message]."
Exercise 3 — Build a few-shot classifier. Choose a real task from your research (e.g., classifying reviewer comments, screening papers, coding survey responses). Write 5 examples: 3 typical, 2 tricky edge cases. Format every example output identically. Then test on 3 new inputs and note where the model still errs.
Exercise 4 — Chain-of-thought comparison. Pick a multi-step problem from your field (a calculation, a debugging scenario, a methods decision). Run it twice: once with a direct-answer prompt, once with "think step by step" plus a FINAL ANSWER separator. Compare the answers. Then try a version where CoT hurts (a taste/fluidity judgment) and observe the difference.
Exercise 5 — Write your system prompt. Draft a personal research system prompt: your field, 4–6 standing rules (including citation honesty), default output preferences, and ambiguity handling. Use it for one week. At week's end, note which rule saved you the most time and which rule you violated most.
Exercise 6 — Structured extraction pipeline. Define a JSON schema with 4–6 fields for extracting information from papers in your area (include types, allowed values, and a missing-data rule). Write the extraction prompt with one example. Run it on 5 abstracts, parse every output with code, and record the structural failure rate.
Exercise 7 — Debug systematically. Take a prompt that recently gave you a bad output (or deliberately write a vague one). Write the one-sentence failure description, then walk the Chapter 8 checklist in order, fixing the first match. Log: failure → diagnosis → fix → result. Repeat until the output is acceptable or you hit the "stop iterating" criteria.
Exercise 8 — Template + mini-eval. Convert your most-used research prompt into a versioned template with slots and defaults. Build a 20-case eval (10 typical, 5 edge, 5 from your failure log) with gold answers. Score v1.0. Make one change, score v1.1, and decide with numbers whether to keep it.
Exercise 9 — Write a methods paragraph. Imagine you used a prompted model to extract data for a paper. Write the methods paragraph following the Chapter 11 checklist (model version, dates, prompt location, examples, parameters, failure handling, validation). Then critique it as a reviewer: what question is still unanswered?
Exercise 10 — Design a prompt ablation study. Pick one prompting technique from this book. Design a small experiment: the technique vs. a control, fixed test set (n≥50), one metric, dev/test separation. Write the one-paragraph "experiment plan" including what would count as a positive result and what the main threat to validity is (hint: test-set prompt tuning).
End of Book 16 — Prompt Engineering: A Practical Guide. AstolixGen Learning Series.