
Book 49 of 50 · Free
From Student to Researcher: Publishing Basics
25,803 words · 17 chapters · illustrated

Book 49 of 50 · Free
25,803 words · 17 chapters · illustrated
Book 49 of 50 — AstolixGen Learning Series (Detailed Edition) For researcher and publication students

Every published researcher was once a student staring at a blank document, wondering how ordinary people turn coursework into peer-reviewed papers. This book closes that gap. It walks you through the entire publishing journey in plain language: how academic publishing works and who the players are, how to choose a research problem that is worth your months of effort, how to read and organize the literature, how to design honest experiments, how to write a paper section by section, how to format it to IEEE standards, how to respond to reviewers without taking criticism personally, where to submit your work, and how to do all of it ethically. You do not need prior publications to benefit — only curiosity and a willingness to work through the steps. By the end, you will have a complete mental map of the research process and a set of templates and checklists you can reuse for every paper you ever write.
By the end of this book, you will be able to:
| Chapter | Guiding Question | Key Takeaway |
|---|---|---|
| 1. How Academic Publishing Works | Who decides what gets published, and how? | Publishing is a structured, human process of submission, peer review, and revision — not a mystery. |
| 2. Choosing a Research Problem Worth Solving | How do I pick a problem that matters and is doable? | A good problem is original, significant, and feasible — and you verify all three before committing. |
| 3. Literature Review | How do I find, organize, and synthesize prior work? | A literature review maps what is known, what is contested, and where the gaps are. |
| 4. Reading Papers Critically | How do I evaluate a paper instead of just reading it? | Critical reading asks what was claimed, what was shown, and what was left out. |
| 5. Research Methodology | How do I design a study that produces trustworthy answers? | A sound methodology links your question to data, controls, and evaluation in a defensible chain. |
| 6. Running Experiments and Recording Results Honestly | How do I run experiments I can stand behind? | Honest records, fixed protocols, and reported failures are what make results believable. |
| 7. Academic Writing: The Anatomy of a Paper | What goes in each section of a paper? | Every section has one job; write each one to do that job and nothing else. |
| 8. IEEE Formatting, Citations, and References | How do I format and cite correctly? | Consistent IEEE formatting and complete citations signal professionalism and respect for prior work. |
| 9. Figures, Tables, and Visuals That Clarify | How do I show my work visually? | Good visuals make one point each and are understandable without reading the text. |
| 10. Responding to Reviewer Feedback Without Fear | How do I handle criticism of my paper? | Reviewer comments are free expert consulting — respond point by point, calmly and completely. |
| 11. Choosing Where to Submit | Journal or conference — and which one? | Match your paper's maturity, scope, and timeline to the venue's expectations. |
| 12. Ethics: Plagiarism, Authorship, and Integrity | What are the rules I must never break? | Integrity is non-negotiable: credit others, report honestly, and never fabricate. |
| Section | Purpose | Common Mistakes |
|---|---|---|
| Title | Tell the reader exactly what the paper is about | Too vague, too long, or promising more than the paper delivers |
| Abstract | A complete miniature of the paper: problem, method, results, conclusion | Background-heavy with no results; exceeding the word limit |
| Keywords | Help indexing and discovery | Using terms nobody searches for; repeating the title verbatim |
| Introduction | Motivate the problem, state contributions, outline the paper | Literature dump with no narrative; contributions buried or missing |
| Related Work | Position your work among existing approaches | Listing papers without comparison; omitting the closest competitors |
| Methodology | Describe what you did in enough detail to reproduce | Missing hyperparameters, data splits, or preprocessing steps |
| Experiments/Results | Present evidence fairly, with baselines and analysis | Cherry-picked results; no statistical grounding; missing ablations |
| Discussion | Interpret what the results mean and admit limitations | Repeating the results section; hiding weaknesses |
| Conclusion | Summarize contributions and point to future work | Introducing brand-new claims; vague "future work" with no direction |
| References | Credit every source you built on | Incomplete entries; citing papers you never read; formatting errors |
| Your Journey Stage | Book Parts That Help | What You Gain |
|---|---|---|
| "I have no topic yet" | Chapters 1–3 | Understanding of the system, a method for finding a worthwhile problem, and a map of existing work |
| "I have a topic but no plan" | Chapters 4–5 | Critical reading skills and a defensible study design before you spend months on experiments |
| "I have results but no paper" | Chapters 6–9 | Honest experiment records, a section-by-section writing guide, IEEE formatting, and clear visuals |
| "My paper is written" | Chapters 10–11 | A venue strategy and a calm, systematic way to survive peer review |
| "I want to do this right" | Chapter 12 (and every chapter's ethics notes) | A working knowledge of plagiarism, authorship, and integrity rules that protect your career |
When Sharmeen finished her first year of MS in Artificial Intelligence, her supervisor said something that changed how she saw her degree: "Your thesis is not the goal. A published paper is." Sharmeen had assumed publishing was something professors did — distant, political, and mysterious. It took her one full paper cycle to learn that academic publishing is actually a very human, very procedural system. Once you understand the machinery, it stops being intimidating and starts being navigable. This chapter explains that machinery.
Researchers publish for reasons that go beyond career advancement, though career advancement is certainly one of them. At its core, publishing is how science accumulates. An experiment run in a lab in Karachi is useless to a researcher in Berlin unless its methods and findings are written down, reviewed, and made accessible. Publication is the mechanism by which private effort becomes public knowledge. When you publish, you are doing four things at once: claiming a contribution ("here is something new"), inviting scrutiny ("check my work"), giving credit ("here is whose shoulders I stand on"), and enabling reuse ("here is how you can build on this"). Every rule in the publishing system exists to protect one of these four functions.
Understanding this helps you write better papers. A paper is not a report of everything you did; it is an argument that a specific contribution is real, new, and useful. Reviewers are not grading your effort; they are testing that argument. Keep this in mind and much of what follows will make sense.
Research appears in several kinds of venues, and they differ in speed, selectivity, and purpose.
Journals are periodicals that publish on a rolling basis. You submit whenever you are ready, and the review process typically takes several months. Journals allow longer papers, multiple rounds of revision, and in many fields they are the most prestigious venue. In computer science and AI, well-known journals include IEEE Transactions on Pattern Analysis and Machine Intelligence, the Journal of Machine Learning Research, and Nature Machine Intelligence. Journal papers are usually archival: they are the final, polished record of a piece of work.
Conferences are events with fixed deadlines, held annually or biannually. You submit by the deadline, reviewers evaluate within a few weeks or months, and accepted papers are presented at the event and published in proceedings. Conferences are fast — from submission to publication can be under six months — and in computing fields, top conferences like NeurIPS, ICML, ICLR, CVPR, and ACL are as prestigious as the best journals. For a student, a conference paper is often the ideal first publication: the deadline forces you to finish, and presenting gives you feedback and visibility.
Workshops are smaller, focused events usually attached to larger conferences. They accept shorter papers, preliminary results, and position papers. Workshops are excellent for a first submission because the bar is calibrated for work in progress, and you get expert feedback early.
Preprint servers, most famously arXiv (pronounced "archive"), let you post your manuscript publicly before or during peer review. Preprints establish priority — proof that you had the idea first — and get your work read quickly. In AI, posting to arXiv is standard practice. But note: some venues have rules about preprints, so check the policy of your target venue before posting.
Several roles interact in the publishing process:
In double-blind review, authors and reviewers are anonymous to each other, which reduces bias. Many AI conferences use double-blind review; many journals use single-blind (reviewers know who the authors are). You must anonymize your manuscript for double-blind venues — no names, no "in our previous work," no revealing acknowledgments.
Here is the typical journey, using a conference as the example:
For journals, step 5 often yields "major revision" or "minor revision" rather than a clean accept, and the cycle of review and revision can take a year or more. Rejection is normal — most submissions to selective venues are rejected, and experienced researchers collect rejections routinely. A rejection is not a verdict on your worth; it is one set of reviewers' judgment on one version of one paper.
Reviewers generally evaluate five things, whether or not the review form names them:
No paper is perfect on all five. Reviewers weigh the balance. A paper with a modest but solid and clearly explained contribution often beats a flashy paper with shaky experiments.
Traditionally, readers (or their libraries) paid for journal subscriptions. Open access flips this: the paper is free for anyone to read, and the cost is covered by an article processing charge (APC) paid by the authors' institution or funder, or absorbed by the publisher. Many venues now offer open-access options. When you sign a copyright form, you typically transfer copyright to the publisher but retain rights to share preprints and use the work in your thesis. Always read the copyright agreement. As a student, check whether your university covers APCs before choosing an open-access route that charges fees.
Be wary of venues that charge large fees while promising suspiciously fast acceptance with little review — these are predatory journals, covered in Chapter 12. Legitimate venues never guarantee acceptance.
Sharmeen's first paper took eleven months from idea to camera-ready: two months of reading, three months of experiments, one month of writing, one month waiting for reviews, one month of rebuttal and revision, and the rest in formatting and waiting. That is normal. Plan your thesis timeline around it: if your degree requires a publication, start the paper at least a year before you need it accepted.
It helps to understand the economics, because it explains many venue behaviors. Traditional subscription journals charge libraries for access; your university library pays so you can read. Open-access journals and conferences charge authors instead, through article processing charges (APCs) or conference registration fees. Neither model is free — someone always pays for editing, hosting, and organizing review.
As a student, three practical consequences follow. First, always check costs before falling in love with a venue: an open-access journal with a $2,000 APC is not an option unless your university or grant covers it — ask your supervisor early, not after acceptance. Second, conference travel costs real money; factor registration, flights, and visas into your plan, and ask about student travel grants (many conferences offer them, but deadlines are early). Third, be deeply suspicious of any venue whose business model seems to be "collect fees from authors as fast as possible" — that is the economic signature of predatory publishing (Chapter 12).
Your supervisor is not just an approver — they are your co-pilot through a process they have flown many times. A good supervisor helps you scope the problem, points you to the right literature, reads drafts with a reviewer's eye, suggests venues, and interprets reviews. In return, they are typically a co-author (often last author), reflecting their intellectual contribution and responsibility.
Make the collaboration work: bring drafts early and often (a rough draft beats a perfect draft that arrives the night before the deadline), ask specific questions ("is this baseline set fair?" beats "is this good?"), and keep them informed of timeline risks. If your supervisor is unresponsive, be politely persistent and put requests in writing — and cultivate a second reader, like a senior PhD student, as backup. Your first paper teaches you the process; your supervisor's experience is the textbook.
Publication is not the finish line — it is the starting line for impact. After your paper appears: post the preprint to arXiv if the venue allows (many researchers discover papers there, not in proceedings); share the code and data repository link; present the work at your department's seminar; and add it to your Google Scholar profile and ORCID record so citations find you. When others cite your work, read those citing papers — they tell you how the field understood your contribution and often spark your next problem. Sharmeen's first paper earned modest citations, but one citing paper's "future work" section gave her the seed of her second. The publication cycle is a loop, and each loop makes the next one easier.
Not every submission reaches reviewers. Desk rejection — rejection by the editor or program chair without review — happens for predictable reasons: the paper is out of scope, wildly over the page limit, badly formatted, clearly unfinished, or submitted to the wrong track. Desk rejection is the most avoidable failure in publishing, and it stings because you wait weeks for what a careful hour would have prevented. Before submitting, check scope fit against recent proceedings, respect every formatting rule, and have someone outside your project read the abstract and tell you what the paper is about. If they cannot, reviewers will not get the chance to try.
For your research: Before you write a single word of your paper, pick one real venue you might target — a conference whose proceedings you have actually read. Download its call for papers, note the deadline, page limit, and formatting template. Working toward a concrete target changes how you plan experiments and how you write. Vague plans produce vague papers; a real deadline produces a finished one.
Key takeaways: - Publishing turns private effort into public, scrutinized, reusable knowledge — the paper is an argument for a contribution, not a diary of your work. - Journals, conferences, workshops, and preprints serve different purposes; in AI, top conferences rival journals in prestige and are often the best first target. - The review pipeline (submit → review → rebut/revise → decide → publish) is procedural and human; rejection is routine, not personal. - Reviewers judge originality, significance, soundness, clarity, and reproducibility — aim for a solid balance, not perfection. - Understand open access, copyright, and predatory journals before you submit anywhere. - Start early: a first paper commonly takes close to a year from idea to publication.
The most expensive mistake in research is solving the wrong problem — spending eight months on work nobody needs, or that was done five years ago, or that cannot be finished with the resources you have. Choosing well is a skill, and like all skills it can be learned. This chapter gives you a practical method.
Three tests, all of which must be passed:
Sharmeen learned this the hard way. Her first idea was "build a better large language model than the state of the art" — original in ambition, significant in theory, but completely infeasible on a single university GPU. Her supervisor helped her narrow it to something feasible: evaluating how well existing multilingual models handle Urdu news headline classification, and improving them with a targeted fine-tuning strategy. Same curiosity, feasible scope.
Good problems rarely arrive as lightning bolts. They come from systematic exposure:
Use this funnel to go from vague interest to a committed problem:
Stage 1 — Brainstorm (1–2 weeks). Write down 10–15 candidate problems as single sentences. Each should name the task, the approach or angle, and the expected contribution. Example: "Fine-tune multilingual transformer models for Urdu news headline classification and compare data-efficient strategies for low-resource settings."
Stage 2 — Literature check (2–3 weeks). For each candidate, search the literature (Chapter 3). Kill any candidate where the exact work already exists. This is the stage where most candidates die, and that is good — each death saves you months.
Stage 3 — Feasibility audit. For survivors, answer honestly: - Do I have (or can I get) the data? - Do I have the compute? (Estimate GPU hours; ask seniors what similar work cost.) - Do I have the skills, or can I learn them in the time available? - Is my supervisor able to guide this? - Can a first result appear within 3–4 months? (If not, the scope is too big.)
Stage 4 — The one-paragraph proposal. Write one paragraph: the problem, why it matters, what you will do, how you will evaluate, and what success looks like. Show it to your supervisor and one senior student. If they cannot poke a hole in it, you have a problem.
Students typically overscope. "I will build a complete Urdu NLP platform" is a product, not a paper. A paper answers one question well. Sharmeen's final scope: three fine-tuning strategies, one dataset she built from public news archives, one clear evaluation. That is a paper. The platform can come later.
A useful test: can you state your contribution in one sentence without the word "and" more than once? "We show that strategy X outperforms standard fine-tuning on Urdu headline classification with limited data" — one claim, testable, bounded.
Not every chosen problem survives contact with reality. Pivot (change direction) when: - The literature check reveals the work is done — pivot early, pivot proudly. - Experiments consistently show the approach cannot beat a simple baseline — the problem may be wrong, not just the method. - The data you need does not exist and cannot be created in time.
Do not pivot when: results are merely disappointing but the question is still open (that is called research); a reviewer dislikes the topic (that is called taste); or you are bored (that is called Tuesday — push through).
Keep a decision log: a dated document recording why you chose the problem, what you ruled out, and pivot decisions. It prevents circular rethinking and becomes gold when you write the introduction.
Most students start with an interest too broad to be a problem: "I want to work on NLP for Urdu." Narrow it by repeatedly asking "for whom?" and "for what task?" until you reach something concrete:
Each answer cuts the scope and sharpens the contribution. Five rounds of this drill typically take a sprawling interest down to a paper-sized problem. If you cannot answer "for whom," the problem may lack significance; if you cannot answer "compared against what," it lacks an evaluation plan — both are fixable now, painful later.
Here is how Sharmeen's ten candidate sentences fared. Candidates like "build a better LLM for Urdu" died at the feasibility audit (no compute). "Apply BERT to Urdu sentiment analysis" died at the literature check — three papers had done nearly exactly that. "Compare fine-tuning strategies for Urdu headline classification under limited data" survived all three tests: the literature covered high-resource settings and other languages, newsroom editors confirmed the need, and the compute fit one GPU.
Her one-paragraph proposal read: "Newsrooms handling Urdu wire copy need automatic categorization, but only ~4,000 labeled headlines are available. We compare three fine-tuning strategies for multilingual transformers under this constraint, evaluating with macro-F1 across five seeds. Success means identifying a strategy that beats standard fine-tuning by a clear margin while remaining computationally practical." Her supervisor's only change: add a simple baseline. Two weeks of funnel work replaced what could have been six months of drifting.
Learn to recognize these early:
Walking away from a bad problem in week three is wisdom. Walking away in month eight is tragedy. The funnel exists to make quitting cheap and early.
Borrow a concept from startups: what is the smallest version of this work that still makes a contribution? For Sharmeen, the minimum viable paper was: one dataset, three strategies, one clear comparison. Everything else — the larger dataset, the extra language, the demo system — was cut from the paper (not from her dreams; some became her second paper). Define your minimum viable paper explicitly with your supervisor. When time runs short, as it always does, you protect the core and cut the extras — instead of delivering a sprawling, unfinished mess.
After the funnel, stress-test significance with the "so what" drill. State your expected result, then ask "so what?" — and answer concretely. "Adapter tuning beats full fine-tuning on Urdu headlines with limited data." So what? "Newsrooms and low-resource practitioners can deploy accurate classifiers without massive annotation budgets." So what? "That lowers the cost barrier for Urdu NLP applications in media and civic tech." Three rounds of "so what" should land on a real beneficiary doing something they could not do before. If you cannot get past round one, the problem's significance is asserted, not real — go back and find the beneficiary (Chapter 2's newsroom visit is how Sharmeen found hers).
Bring your three surviving candidates as one-page briefs: the problem sentence, the gap in one paragraph, the feasibility audit results, and your honest ranking. Then ask three questions: "Which of these is most publishable in a year?", "Which best fits your research program?", and "What would you change?" Supervisors think in portfolios — they know which problems have momentum, which venues are receptive, and which ideas they can actually guide. Sharmeen arrived convinced her top-ranked candidate was best; her supervisor pointed out that candidate #2 connected to a funded project with an available dataset, making it twice as feasible. She switched, and never regretted it. The meeting takes thirty minutes and can redirect a year of work — have it before you commit, not after.
Whatever timeline you estimate, multiply by three. This is not pessimism — it is calibration from thousands of student projects. Data collection hits access problems, experiments reveal bugs, writing takes longer than the "just write it up" fantasy suggests, and review cycles add months. Sharmeen estimated four months for experiments; they took seven. The students who finish on time are not faster — they planned with the 3× rule and protected their buffers fiercely. When your supervisor suggests a venue deadline, work backwards with tripled estimates: if the math does not fit, choose a later venue rather than compressing the science. A rushed paper submitted on time is worse than a solid paper submitted next cycle.
Keep your decision log simple — a dated document with three columns: Date, Decision, Reason. Example entries: "Oct 12 — Dropped candidate 'Urdu LLM pretraining' — Reason: needs 8× our GPU budget; literature check shows two similar efforts. Oct 28 — Chose adapter-tuning comparison — Reason: passes all three tests; supervisor has relevant expertise; dataset obtainable." Review it monthly. The log's real power shows at month six, when you are tempted to revisit a discarded idea — the log reminds you why you quit, with evidence, instead of letting nostalgia restart a dead end. It also becomes the basis of your thesis's "research design" narrative: examiners love seeing reasoned choices.
For your research: This week, write 10 candidate problem sentences using the template: "I will [approach] for [task] in [domain/context], and evaluate by [metric/benchmark], because [who benefits]." Run each through the three tests (originality, significance, feasibility) and keep the best three. Bring those three to your supervisor. This single exercise will save you more time than any other in this book.
Key takeaways: - A worthwhile problem passes three tests: original, significant, and feasible for you specifically. - Problems come from gap-hunting in papers, reproducing work, talking to real users, datasets, and your supervisor's agenda. - Use the funnel: brainstorm → literature check → feasibility audit → one-paragraph proposal. - Scope to the smallest publishable unit: one clear, testable claim per paper. - Pivot on evidence (problem is done, infeasible, or beaten by baselines), not on boredom; keep a decision log.
The literature review is where most students either drown (reading 200 papers with no system) or cheat (citing 15 papers they never read). Neither works. A literature review is a structured investigation with a deliverable: a clear map of what is known, what is contested, and where your work fits. This chapter gives you the system.
Start with the right sources. For AI and computing: Google Scholar (broadest coverage, citation tracking), arXiv (latest preprints), IEEE Xplore and the ACM Digital Library (published versions), and Semantic Scholar (good filters and TL;DR summaries). Your university library likely provides access to paywalled papers — learn to use it in week one.
Search strategy matters more than search tools:
Two rounds of snowballing from good seeds usually surfaces 80% of what matters. Set up Scholar alerts for your key terms so new work comes to you.
There is no magic number, but there are stopping rules. Stop expanding when: new papers mostly cite papers you already have; you can predict a new paper's approach before reading it; and you can explain the field's main lines of work from memory. For an MS thesis in AI, a working set of 40–80 papers, read at varying depths, is typical. Depth beats breadth: 30 papers truly understood beat 150 skimmed.
Pick a reference manager on day one — Zotero (free, excellent) or Mendeley — and put every paper you touch into it immediately, with the PDF attached. Future you will thank present you when it is time to format references.
For each paper that matters, write a structured note. Use this template:
Paper: [full citation]
One-sentence summary: [what did they do?]
Problem: [what question did they answer?]
Method: [how, in 3-5 lines]
Data/evaluation: [datasets, metrics, baselines]
Key result: [the headline number or finding]
Strengths: [what is genuinely good]
Weaknesses/limits: [what is missing, shaky, or assumed]
Relevance to me: [how does this connect to my problem?]
Sharmeen kept these notes in a simple spreadsheet with one row per paper and columns for method family, dataset, metric, and result. When she later needed "all papers that used data augmentation for low-resource text classification," she filtered the sheet instead of re-reading everything.
Collecting papers is not a literature review. Synthesizing means organizing them into a story. Group papers by approach, not by author: "three families of solutions exist — feature-based methods, fine-tuning approaches, and data augmentation strategies." Within each family, note the trajectory: what improved over time, what stalled, what assumptions changed.
Then identify: - Consensus: what everyone agrees on (e.g., "pretrained multilingual models outperform training from scratch on small Urdu datasets"). - Contestation: where papers disagree (e.g., whether adapter-based tuning beats full fine-tuning at small data sizes). - Gaps: what nobody has tried (e.g., "no published work evaluates these strategies specifically on Urdu news headlines with fewer than 5,000 examples").
Your gap — stated precisely — is your problem's justification. Sharmeen's introduction practically wrote itself once she could say: "Prior work covers X and Y, but not Z; we address Z."
A related-work section is not a list; it is an argument with citations as evidence. Structure it by theme, and end each theme with a sentence connecting it to your work: "Unlike [A], which assumes large labeled datasets, our setting provides only 4,000 labeled headlines, motivating data-efficient strategies." Every cited paper should earn its place — if you cannot say why it is there, remove it.
Cite the closest competitors prominently and fairly. Reviewers often wrote those papers. Misrepresenting or omitting the most relevant prior work is the fastest route to rejection.
A few techniques multiply your search power. On Google Scholar, use quotes for exact phrases ("low-resource text classification"), the author: operator to find a researcher's other work, and the "since year" filter to focus on recent advances. Sort by date when surveying the last two years, by citations when finding foundational work. On arXiv, browse the relevant categories (cs.CL for language, cs.CV for vision, cs.LG for learning) weekly — fifteen minutes keeps you current.
Learn the field's vocabulary as you go. You might start searching "Urdu AI" and discover the literature says "low-resource language" and "cross-lingual transfer." Each newly learned term unlocks a new search. Keep a running list of terms; it becomes your search arsenal and, later, your keyword list.
Set up Scholar alerts for your two or three core queries. New papers will arrive by email. Also follow the key researchers' lab pages and Google Scholar profiles — prolific authors in your niche are human RSS feeds for what matters.
Collecting PDFs is easy; processing them is the skill. Use this workflow:
Sharmeen's rule: no PDF sits in the inbox longer than two weeks. The pile never grew beyond twenty unprocessed papers, and her spreadsheet became the lab's shared resource — her juniors still use it.
You will find papers that contradict each other: one says adapters beat full fine-tuning, another says the opposite. Do not panic and do not pick a side by citation count. Instead, ask what differs: the datasets, the data sizes, the base models, the tuning budgets? Contradictions usually dissolve into "it depends on X" — and X is often your research opportunity. Sharmeen found exactly this split in the literature, and her paper's framing became: "we resolve this disagreement for the low-resource news setting." A contradiction, investigated, is a contribution waiting to happen.
If a full experimental paper feels far away, consider writing a focused survey or mini-review of your sub-area as a first publication. Surveys are publishable (workshops and some journals welcome them), they force you to master the literature, and they are genuinely useful to the field. A good mini-survey does not just list papers — it proposes a taxonomy, compares approaches on common dimensions, and identifies open problems. Your "state of the field" memo from this chapter is already half a survey. Discuss with your supervisor whether your area needs one; if the last survey is three years old, it probably does.
When your pile grows past thirty papers, reading notes alone stop being enough — you need to compare. Build a synthesis matrix: rows are papers, columns are the dimensions that matter for your problem (method family, data size, language/domain, key technique, reported metric, limitation). Fill it from your notes. Patterns jump out: you will see at a glance that every paper uses datasets above 50k examples (your low-data angle is the gap), or that no one reports variance (your multi-seed evaluation is a differentiator). Sharmeen's matrix had a column for "tuning budget reported" — nearly every cell was empty, which became a methodological point in her paper: she reported hers. The matrix also writes your related-work section: each column with interesting variation becomes a paragraph comparing approaches on that dimension.
Beyond manual snowballing, use tools that map citation networks. Semantic Scholar and Connected Papers visualize which papers cite which, revealing clusters — a dense cluster is a sub-community you must understand, and a paper bridging two clusters is often influential. Scite shows how papers cite: supporting, contrasting, or merely mentioning. A paper with many contrasting citations is contested territory — read it carefully before building on it. These tools do not replace reading, but they tell you where reading will pay off most. Spend one afternoon mapping your area's citation graph; it often reveals the two or three papers everyone orbits, which become your Pass-3 priorities.
A literature review is a snapshot; the field keeps moving. Build a staying-current routine that costs thirty minutes a week: skim new arXiv listings in your categories every Monday, read your Scholar alert emails instead of archiving them, and browse the proceedings of the top two conferences in your area when they appear. Once a quarter, update your synthesis matrix with the important new papers and ask: does anything change my gap or my framing? Sharmeen caught a newly published paper that overlapped with her work three months before submission — because she spotted it early, she added a comparison and cited it as concurrent work instead of being blindsided by a reviewer. Staying current is also how you find your next problem: the future-work sections of this year's papers are next year's opportunities.
Not everything worth reading is a published paper. PhD theses often contain the fullest account of a research program — literature reviews, failed approaches, and details cut from papers for space. Technical reports from industry labs document systems and datasets papers only summarize. Research blogs (from labs and companies) explain ideas intuitively and announce new resources early. Treat grey literature as scaffolding: excellent for learning and for finding leads, but cite the peer-reviewed version when one exists, and verify claims before building on them — blogs are not peer-reviewed. Sharmeen found her dataset-cleaning approach in a PhD thesis's appendix, a detail that never made it into the author's papers. The lesson: search beyond proceedings when you are stuck; the answer is sometimes in the document nobody cites.
Make forward snowballing a reflex, not a project. Whenever a paper matters to you, click "Cited by" once and scan the first page of results for anything that extends, challenges, or applies it. Five minutes, once per important paper, keeps your map alive between major review efforts. It is also the fastest way to find the researchers currently active in your niche — the names that recur in "cited by" lists are the people whose new work you should watch, whose labs you might one day join, and who might one day review your paper.
For your research: Build your reading-note spreadsheet this week with the template above, and fill it for your 10 most important papers. Then write a two-page "state of the field" memo: three approach families, the consensus, the contestation, and your gap in one paragraph. This memo becomes the skeleton of your related-work section and the core of your introduction's motivation.
Key takeaways: - Search systematically: seed papers → backward and forward snowballing → keyword expansion → Scholar alerts. - Stop when new papers stop surprising you; depth of understanding beats raw count. - Use a reference manager from day one and structured reading notes for every important paper. - Synthesize into consensus, contestation, and gaps — your gap justifies your problem. - Write related work as a themed argument, cite closest competitors fairly, and give every citation a reason to exist.
Most students read papers the way they read textbooks: assuming everything is correct and trying to absorb it. Research papers are not textbooks. They are arguments made by people with limited time, limited data, and an interest in presenting their work well. Critical reading means engaging with the argument: what is claimed, what is actually shown, and what is missing. This chapter teaches you how.
Read every paper three times, with different goals:
Pass 1 — The survey (5–10 minutes). Read the title, abstract, introduction, section headings, conclusion, and glance at figures. Ask: what problem, what approach, what claim? Decide: is this relevant enough for Pass 2? Most papers stop here, and that is fine.
Pass 2 — The understanding (30–60 minutes). Read the full paper but skip proofs and implementation minutiae. Focus on the method's logic and the experiments: what are the baselines, the datasets, the metrics? Try to summarize the paper in your own words afterwards. If you cannot, you did not understand it — re-read the method section.
Pass 3 — The interrogation (1–2 hours, only for key papers). Virtually reimplement it in your head or on paper. Question every choice: why this baseline and not a stronger one? Why this dataset split? Would the result hold if the random seed changed? Check the references for the claims you doubt. This is the pass where you find gaps worth exploiting.
Sharmeen applied Pass 3 to the two papers closest to her problem and discovered that both evaluated on English-centric benchmarks with only a small Urdu subset — a weakness she could address directly in her own evaluation design.
Keep this checklist beside you:
A table full of bold numbers is rhetoric, not proof, until you check: How many runs? (One lucky run is not a result.) What is the variance? Are improvements of 0.3% meaningful or noise? Do the gains hold across datasets or just one? Does the method win because it is better or because it used more compute, more data, or more tuning than baselines? Train yourself to read the experimental setup section before the results table — the setup tells you how much to trust the table.
Recognizing these is not cynicism — it is how you learn what "good" looks like, and it directly improves your own experimental design (Chapters 5–6).
Critical reading pays off twice. First, your related work becomes honest and precise because you actually understand what each paper showed. Second, you internalize the genre: how introductions motivate, how methods are described, how results are presented. Keep a file of "well-written passages" — an introduction paragraph you admire, a clear method description — and study them when you write (Chapter 7).
Not all papers argue the same way. Empirical papers (most of AI) claim "method X works" and prove it with experiments — interrogate their baselines, data, and variance as this chapter teaches. Theoretical papers claim "statement Y is true" and prove it with math — interrogate their assumptions: what do the theorems assume, and do those assumptions hold in practice? A theorem proved under unrealistic assumptions is elegant but may not guide your experiments. Survey papers claim "the field looks like this" — interrogate their coverage: what did they leave out, and is their taxonomy fair to all camps? Dataset and benchmark papers claim "this resource enables research" — interrogate collection methods, labeling quality, and licensing. Knowing which kind you are reading tells you where its weak points hide.
Many students freeze at equations. Do not read math linearly like prose — read it strategically. First, identify what each symbol means (usually defined nearby; keep a symbol glossary for dense papers). Second, understand the equation's role: is it a definition, an objective to optimize, or a result? You can often grasp the role without following every derivation. Third, check dimensions and edge cases: does the formula behave sensibly at extremes? Finally, connect math to code: ask "how would I implement this?" — implementation thinking exposes what the notation hides. You do not need to verify every proof on first reading; you need to understand what is being claimed and why it should be true. Proofs can wait for Pass 3, and even then, focus on the proof's key insight rather than every algebraic step.
Reading alone is slow; reading together is fast and fun. Start or join a journal club: a weekly meeting where one person presents a recent paper in 20 minutes and the group interrogates it. Presenting forces Pass-2-level understanding; the group's questions routinely surface weaknesses the presenter missed. Sharmeen's lab runs a Friday club with a simple format: 5 minutes on the problem, 10 on the method, 5 on results, then open critique. Her rule for presenters: end with "one thing I would do differently" — it trains the gap-hunting reflex from Chapter 2. Within a semester, club members read more deeply and write sharper related-work sections, because they have practiced criticism out loud.
Keep a short list — ten to fifteen papers — that you re-read once a year: the foundational works of your area and the best-written papers you have found. Foundational papers (like the deep learning review by LeCun, Bengio, and Hinton [2] or the Transformer paper [1]) reward re-reading because your growing experience reveals layers you missed. Best-written papers are your style teachers. A canon gives you roots: when the literature feels overwhelming, you return to the papers that define what "good" means in your field.
Once a month, pick a paper in your area and spend exactly ten minutes writing a mini-review: one paragraph summary, two strengths, two weaknesses, and an accept/reject verdict with one sentence of justification. This trains the reviewer's mindset faster than any other exercise — you start noticing the same weaknesses (weak baselines, missing ablations, overclaimed abstracts) that real reviewers notice, and you stop committing them yourself. Sharmeen's lab keeps a shared document of these mini-reviews; before submitting, authors check whether their draft would survive their own lab's ten-minute test. It is humbling how often the answer is "not yet" — and how much stronger the paper becomes after fixing what the test found.
Some of the best ideas come from adjacent fields. A technique from computer vision (data augmentation) transformed NLP; causal inference from statistics is reshaping ML evaluation. Once a month, read one paper from a neighboring area with no goal beyond curiosity. You will not understand everything — that is fine. You are collecting lenses: new ways to frame problems, new evaluation habits, new mathematical tools. Keep a separate "outside reading" list; when you are stuck on your own problem, browse it. Sharmeen's augmentation idea came from a vision paper a labmate presented at journal club — she would never have found it searching only NLP venues. Breadth feeds originality.
Keep an idea journal — a running document where reading turns into thinking. After each Pass-2 or Pass-3 read, spend ten minutes writing freely: what surprised you, what you disagree with, what experiment the paper suggests, what connection you see to your work. Do not edit; do not structure. Over months, this journal becomes a goldmine: half your research ideas will be traceable to entries you barely remember writing. Sharmeen's journal entry from October — "why does no one test adapters below 5k examples? the disagreement in the literature might be a data-size effect" — became, nearly verbatim, her paper's central framing. Reading without writing is entertainment; reading with a journal is research. Review the journal monthly with fresh eyes and tag entries that have grown into real ideas.
Passive highlighting creates the illusion of understanding. Replace it with margin questions: every time you highlight a claim, write a question beside it. "Why this dataset and not X?" "Would this hold with half the data?" "Is this assumption realistic for low-resource settings?" Questions do three things highlights cannot: they record your confusion precisely (so you can resolve it later instead of re-reading), they generate research ideas (Chapter 2's gaps often start as margin questions), and they make your reading notes actionable. When you finish a paper, copy your unanswered questions into your idea journal — they are a personalized research agenda written by your own curiosity. Sharmeen's margin question "do they tune baselines equally??" — with two question marks — eventually became a methodological pillar of her paper. Read with a pen, and your confusion becomes your compass.
A paper's reference list is a curated map of its intellectual neighborhood, drawn by someone who knows the area. When a paper matters to you, read its references as a list: which works are cited repeatedly in the text (foundational), which appear once in passing (context), which are recent (the live frontier)? The authors' choice of what to cite — and what to omit — tells you how they position themselves. Then read the two or three references you have never heard of; they are often the hidden roots of the idea. This habit compounds: each paper's references lead you deeper into the field's structure, and after a few months you will navigate the literature the way locals navigate a city — by landmarks, not by map.
A final reading discipline: for every ten papers you survey, deep-read only one or two — the ones closest to your problem or most cited by the others. The rest get Pass 1 or Pass 2 and an honest tag in your manager. Students feel guilty about papers they "haven't really read," but strategic depth beats guilty breadth every time. Your job is not to read everything; it is to understand deeply the papers that shape your work and to know the rest well enough to place them. Give yourself permission to stop at Pass 1 without guilt — that is what the passes are for.
For your research: Take the single paper closest to your planned work and do a full Pass 3 interrogation. Write a one-page critical summary: the narrow claim vs. the broad implication, the strongest and weakest experimental choices, and three specific things your work will do differently. This page is the seed of both your methodology and your related-work argument.
Key takeaways: - Read in three passes — survey, understand, interrogate — and reserve deep reading for papers that earn it. - Interrogate claims, baselines, data, metrics, ablations, limitations, and reproducibility with a fixed checklist. - Read the experimental setup before the results table; variance, run counts, and tuning fairness determine trust. - Learn to spot weak baselines, dataset monoculture, and buried limitations — then avoid them in your own work. - Critical reading directly feeds your related work, your method design, and your writing style.
Methodology is the bridge between your research question and your results. It is the set of reasoned choices — about data, methods, baselines, and evaluation — that makes your conclusions trustworthy. Students often treat methodology as paperwork to write after the experiments. Reverse that: design the methodology first, then run the experiments it calls for. A week of design saves months of rework.
A common student trap is tool-first thinking: "I want to use transformers" instead of "I want to answer whether data-efficient fine-tuning helps Urdu headline classification." Tools serve questions, not the reverse. Write your research question as a single interrogative sentence, then derive everything else from it:
If you cannot name your variables, you do not have a methodology yet — you have an activity.
Your method section must answer "why this approach?" not just "what did I do?" Justify choices against alternatives: "We chose adapter-based tuning because it updates fewer than 5% of parameters, which the literature suggests reduces overfitting on small datasets — a central concern in our setting." Every significant choice (model, optimizer, data split, metric) should carry a one-sentence justification. "Because it is popular" is not a justification; "because prior work shows it is the strongest baseline for this task" is.
Baselines are the existing methods you compare against. Weak baselines make your method look good and your paper look dishonest. Your baseline set should include:
Tune baselines with the same care you tune your method. Nothing destroys credibility faster than a reviewer noticing the baseline was trained for 3 epochs while yours got 30.
Describe your data as if a stranger must rebuild it: source, size, collection procedure, cleaning steps, and labeling process. Then split it properly: train (for learning), validation (for tuning decisions), and test (for final reporting, touched once). The test set is sacrosanct — every time you peek at it to make a decision, your reported numbers become optimistic. This is called test-set leakage, and it invalidates results.
Other leakage traps: normalizing using statistics from the full dataset before splitting; duplicate or near-duplicate examples across splits (common when scraping news); and temporal leakage, where the model trains on future data to predict the past. Sharmeen caught duplicates in her Urdu headlines by hashing normalized text and found 6% overlap between her initial splits — fixing it changed her results and saved her from an embarrassing review.
Choose metrics that match the task and the data. Accuracy misleads on imbalanced data (a classifier that always predicts the majority class can score 90% accuracy and be useless); use precision, recall, and F1 per class, and say which you optimize. For ranking or generation tasks, use the field's standard metrics so your work is comparable.
Then ground your numbers: run each experiment multiple times with different random seeds (at least 3–5) and report mean and standard deviation. If you claim method A beats method B, check that the difference survives a basic significance test rather than eyeballing two means. You do not need advanced statistics — consistent multi-seed results with reported variance already put you ahead of many published papers.
An ablation removes one component of your method at a time and measures the effect. If removing component X changes nothing, X is decoration, not contribution — either justify it differently or drop the claim. Plan ablations during design, not after, because they determine what you log. Sharmeen's design included ablating her data-augmentation step; when augmentation turned out to contribute only 0.4% on average, she reported it honestly and reframed her contribution around the adapter strategy. The paper was stronger for it.
Before running anything, write a one-to-two-page experiment plan: the question, variables, data splits, baselines, metrics, ablations, compute budget, and the exact conditions under which you would declare success or failure. Share it with your supervisor. This is not bureaucracy — it is protection against fooling yourself. When results disappoint (they will), the plan tells you whether to adjust the method or question the question, instead of drifting into endless tweaking.
Most AI work is quantitative: numbers from controlled experiments answer the question. But some research questions need qualitative methods — interviews, user studies, thematic analysis of texts — especially when studying how people use systems. Sharmeen's newsroom visit was informal, but a structured interview study with ten editors about categorization pain points could itself be a publishable contribution in a human-centered venue. Mixed methods combine both: quantitative performance plus qualitative user evaluation. Match the method to the question: "which strategy is most accurate?" is quantitative; "why do editors distrust automatic categorization?" is qualitative. Reviewers judge you against the standards of the method you chose — a user study with three participants will be criticized, just as an experiment with one seed would be. If you use qualitative methods, learn their standards (sampling, coding, saturation) rather than improvising.
Before finalizing your design, list everything that could make your conclusions wrong — then address each. Common threats:
Write these into your paper's limitations discussion. A limitations section that mirrors your actual threat analysis reads as mature; one that says "future work will explore more datasets" reads as boilerplate. Sharmeen's paper explicitly noted that results might not transfer to morphologically richer languages — a reviewer praised the honesty and suggested it as follow-up work.
Before the full experimental campaign, run a pilot: one baseline and your method, one seed, a data subset, the full pipeline end to end. Pilots catch the unglamorous failures — broken preprocessing, label mismatches, out-of-memory errors, metric bugs — in hours instead of weeks. Sharmeen's pilot revealed her tokenizer was silently truncating Urdu headlines (longer token sequences than English), which would have corrupted every result. Budget one week for the pilot and treat its findings as design input, not results.
Your experiment plan becomes your methodology section almost directly — which is another reason to write the plan well. Translate each plan element into prose: the question becomes the section's opening, variables become the experimental design description, baselines become the comparison subsection, and the protocol details become the reproducibility paragraph. Because the plan was written before results existed, the section naturally describes what you intended, keeping you honest about deviations (which you disclose: "we originally planned X; pilot results led us to Y because..."). Reviewers can feel the difference between a methodology designed upfront and one reverse-engineered from results — the former reads as science, the latter as storytelling.
If your research involves people — user studies, interviews, surveys, annotators — you likely need ethical approval from your university's review board before collecting data, not after. The process takes weeks, so start early: describe what participants will do, the risks (usually minimal for interviews, but state them), how you will obtain informed consent, and how you will anonymize data. Even when formal approval is not required, follow its principles: voluntary participation, informed consent, anonymity, and the right to withdraw. For annotation work, pay fair rates and document the guidelines annotators received — annotation quality is a methodological detail reviewers increasingly ask about. Plan ethics in the methodology design phase (Chapter 5's checklist), because retroactive approval does not exist.
For extra credibility, consider pre-registering your study: publishing your experiment plan (question, hypotheses, methods, analysis plan) in a timestamped public registry before running experiments. Pre-registration proves your hypotheses preceded your results — the strongest possible defense against p-hacking and HARKing (hypothesizing after results are known). Some journals offer registered reports, where the plan itself is peer-reviewed before data collection, and acceptance is guaranteed regardless of outcome. While still uncommon in AI conferences, pre-registration is growing in applied ML and is standard in neighboring fields. Even informal pre-registration — emailing your plan to your supervisor with a timestamp, or committing it to a git repository — sharpens your thinking and protects you from self-deception. The plan you wrote in this chapter is already 90% of a pre-registration.
Open the methodology section with an overview paragraph that gives the whole design in five sentences: the question, the approach, the data, the evaluation, and the baselines. Readers decide in this paragraph whether to trust the details that follow — and many reviewers skim the rest. Write it after the full section is drafted, distilling rather than previewing. A strong overview paragraph also serves double duty: it is the paragraph you will reuse (adapted) in talks, posters, and the thesis. If you cannot summarize your methodology in five sentences, the design is probably too complicated — simplify the study, not the paragraph.
A surprisingly effective design test: explain your study to a peer using only the words "we vary X, we measure Y, we hold Z constant." If you stumble — if X, Y, and Z will not stay still — the design has a confound hiding in it. Sharmeen's first attempt went: "we vary the fine-tuning strategy and the data size, and we measure accuracy, holding... the model constant, except the adapter one has fewer parameters, which is the point..." — and there it was: parameter count varied with strategy, a confound she then addressed explicitly by reporting it and discussing it. Saying the design aloud forces the variables into the open where you can see them clearly.
For your research: Write your two-page experiment plan this week using the checklist above. Be specific: name the datasets, the exact splits, the three baselines, the metrics, and the number of seeds. Then ask your supervisor one question: "If I run exactly this plan, will the results — whatever they are — be publishable?" If the answer is no, fix the plan before spending a GPU-hour.
Key takeaways: - Design methodology before experiments: question first, then variables, method, baselines, data, and evaluation. - Justify every major choice; "popular" is not a justification. - Baselines must be strong and fairly tuned — include simple, state-of-the-art, and ablated versions. - Protect your test set; check for duplication, normalization, and temporal leakage. - Report multi-seed means with variance, choose metrics suited to your data, and plan ablations up front. - A written, supervisor-approved experiment plan protects you from self-deception and wasted compute.
Experiments are where research meets reality, and reality is messy: jobs crash at 3 a.m., results contradict your hypothesis, and the "obvious" improvement does nothing. How you handle this mess determines whether your paper is science or storytelling. This chapter is about running experiments you can defend and recording them so completely that your future self — or a skeptical reviewer — can verify everything.
Your experiment plan from Chapter 5 is a contract with yourself. Before the first run, freeze the decisions that affect validity: data splits, preprocessing, hyperparameter search spaces, the number of seeds, and the stopping criteria. Write them in a PROTOCOL.md file in your project directory. If you change the protocol mid-stream (sometimes necessary), record the change, the date, and the reason. An undocumented change is indistinguishable from cherry-picking.
Keep a running experiment log — a simple dated document or spreadsheet — with one row per run:
Date | Run ID | Config (model, hyperparams, seed) | Data version | Result (metric ± std) | Notes
Include failed runs. "Run 14 crashed: out of memory with batch size 64; reduced to 32" is valuable information, not embarrassment. Sharmeen's log showed that her best result came from seed 7 of 10 — without the log, she might have reported only that seed and claimed a stronger result than the average supported. The log kept her honest, and the honest average still beat her baselines.
Log these for every run: the exact code version (git commit hash), the data version, the full configuration, the random seeds, the hardware, the wall-clock time, and the raw outputs. Disk space is cheap; irreproducible results are expensive.
Use git for code from day one, and commit before every experiment batch. Tag the commit that produced each reported number. Keep datasets versioned too — at minimum, record the download date, source URL, and a checksum, and never silently modify a dataset file. If you clean the data, script the cleaning and version the script. "I tweaked the data by hand in Excel" is how results become unrepeatable.
Your hypothesis will be wrong sometimes. That is not failure; it is information. Report negative results plainly: "Contrary to our expectation, data augmentation did not improve performance (mean F1 0.71 vs. 0.71, n=5 seeds)." Reviewers respect this. What they do not respect is the suspicion — always present when only positive results appear — that negative runs were buried. A paper that reports what did not work is more believable about what did.
There is a practical reason too: your thesis needs a story, and "we tried X, it failed because Y, which led us to Z" is a stronger, more honest story than pretending Z was the plan all along.
The most common honest mistake in student experiments is test-set overfitting by iteration: trying dozens of configurations, checking the test score each time, and reporting the best. Each peek leaks information, and the reported number becomes an optimistic fiction. The discipline is simple: make all decisions on the validation set; evaluate on the test set once, at the end, for the final configurations. If you must revisit, say so in the paper ("test scores guided early stopping" is a limitation to disclose, not hide).
Similarly, do not tune hyperparameters on the test set, do not select the best seed's result as "the" result, and do not compare your tuned method against untuned baselines. These are the quiet corruptions that separate real science from theater.
Report what things cost: training time, hardware, and approximate energy where relevant. "Training each model took ~6 hours on a single NVIDIA RTX 3090" lets readers judge practicality and reproduce your budget. If your method needs 10× the compute of baselines for a 1% gain, say so — that tradeoff is part of the scientific record, and hiding it misleads the field.
When results look wrong, resist random tweaking. Debug like an engineer:
Keep a bug log alongside your experiment log: symptom, hypothesis, test, outcome. Sharmeen's bug log shows she once spent two days tuning a model that was training on shuffled labels — the log entry now saves every junior in her lab from repeating it.
Tune systematically: define the search space in your protocol, use the validation set only, and give baselines the same budget as your method. Two honest approaches: grid/random search with a fixed budget (e.g., 20 configurations), or manual tuning guided by learning curves — both fine if the budget is fixed in advance and equal across methods. What is not fine: tuning your method for a week, the baseline for an afternoon, then comparing. Report the search spaces and the selected values in the paper or appendix; "we tuned learning rates on validation" without details is not reproducible.
Plan from the start to release your code and, where possible, your data. A clean repository with a README (setup, data preparation, how to reproduce the main table, expected runtime) dramatically increases your paper's credibility and citation chances — researchers cite work they can build on. Clean as you go: delete dead code weekly, or the pre-deadline cleanup becomes a rewrite. If data cannot be shared (licensing, privacy), share the collection and preprocessing scripts so others can recreate it, and say so explicitly in the paper. "Code and data available at [link]" is now expected at top venues; "available on request" increasingly is not.
Sooner or later, your numbers will disagree with a published paper. Do not assume you are wrong — and do not assume they are. Investigate: different data versions? Different preprocessing? Different tuning budgets? Undisclosed details? Document your exact setup and try to reconcile the difference; if you cannot, report your results honestly with full details and note the discrepancy. Science advances through such contradictions. Sharmeen's adapter results initially contradicted a well-known paper — the difference turned out to be data size (theirs was 50× larger), which became a key insight of her paper rather than an embarrassment. A contradiction you can explain is a contribution; one you hide is a liability.
Treat GPU hours as a budget, not an unlimited resource. Before the campaign, estimate: (time per run) × (number of configurations) × (number of seeds) × (methods including baselines). Sharmeen's estimate was sobering — 4 methods × 15 configs × 5 seeds × 2 hours = 600 GPU hours, about 25 days on one GPU. That forced prioritization: fewer seeds for exploratory configs (2), full seeds only for finalists (5), and baselines tuned with the same budget her method got. Track actual vs. estimated weekly; when the budget slips, cut exploratory breadth before cutting evaluation rigor — a paper with fewer methods but solid statistics beats one with many methods and shaky numbers. And record the final totals for the paper: reviewers and future researchers need to know what the work cost.
Two weeks before the writing deadline, freeze the numbers: no more experiments change the reported results. This discipline prevents the death spiral of endless "one more run" tweaking that has killed many submissions. After the freeze, runs are only for verification (reproducing a table entry) or for reviewer-requested additions during revision. Announce the freeze to yourself in the log: "NUMBERS FROZEN 2026-11-01 — results in results/final/." Sharmeen's lab enforces freezes because they learned the alternative: a paper that is perpetually "almost done" while the deadline passes. A frozen, honest set of results beats a perpetually improving set that never ships.
One month before submission, do the bravest thing in this book: reproduce your own main result from scratch. On a clean machine (or clean environment), follow only your written documentation — the README, the scripts, the logged configs — and regenerate the paper's headline table. If it works, your reproducibility claims are real and your documentation is complete. If it fails, you have found the gaps while there is still time to fix them: the undocumented preprocessing step, the hardcoded path, the missing dependency. Sharmeen's replication test caught a data-loading script that silently depended on a file only present on her laptop — the exact kind of failure that makes "code available" an empty promise. Schedule this test in your calendar when you freeze the numbers; treat failures as gifts, not disasters.
The best lab notebook is the one you actually maintain. Paper notebooks work for quick thoughts and sketches but fail at searchability — you will never find "that learning rate from March." Digital logs (a spreadsheet, a markdown file in git, or tools like Weights & Biases / MLflow for ML experiments) are searchable, timestamped, and shareable with your supervisor. For ML work, experiment-tracking tools that automatically log configs, metrics, and artifacts are worth the setup hour — they eliminate the "I forgot to write down the seed" failure mode entirely. Whatever you choose, follow the non-negotiable rule: the log entry is written when the run starts, not when it finishes. A run without a log entry officially never happened.
A week before submission, audit your experiments against this book's checklists with fresh eyes: open the log and verify every reported number traces to a run; confirm the test set was never used for tuning decisions; check that baselines received equal tuning budgets; confirm variance is reported everywhere it should be. Do this with your supervisor or a labmate — a second pair of eyes catches what familiarity hides. Fix what you find, even if it means re-running something; a correction made now is a strength, while the same correction demanded by a reviewer is an embarrassment. Sharmeen's audit caught one table where she had reported validation scores as test scores — a twenty-minute fix that likely saved her paper.
For your research: Set up your experiment log and git repository today, before your next run. Create the PROTOCOL.md with your frozen decisions, and make your next five runs fully logged with commit hashes. This feels slow for a week; then it becomes automatic, and when a reviewer asks "how many seeds?" or "was the test set used for tuning?" you will answer from records, not memory.
Key takeaways: - Freeze your protocol before running; document any changes with date and reason. - Log every run — including failures — with code version, data version, config, seeds, and raw results. - Version-control code and data transformations; never hand-edit datasets silently. - Report negative results plainly; they increase your credibility, not decrease it. - Never iterate on the test set; tune on validation, evaluate on test once. - Report compute costs honestly; reproducibility is part of the contribution.

You have results. Now you must turn them into a paper — a persuasive, precise document written for busy experts. Academic writing is a genre with conventions, and conventions exist to help readers extract meaning fast. This chapter walks through each section: what it must do, how to write it, and what goes wrong.
Do not write the paper front to back. Use this order:
A good title names the task, the approach or key idea, and sometimes the setting: "Data-Efficient Fine-Tuning of Multilingual Transformers for Urdu News Headline Classification." Avoid cute titles, questions as titles, and acronyms nobody knows. Check: would a researcher searching for this work use these words? The title is a retrieval device first, a marketing device second.
The abstract is the most-read part of your paper — many readers never go further. In 150–250 words, cover four moves: context/problem (1–2 sentences), method (2–3 sentences), results (2–3 sentences with key numbers), conclusion/implication (1 sentence). Write it last and check every claim against the paper body. Never include citations, figures, or undefined acronyms in the abstract.
Example skeleton: "Text classification for low-resource languages remains challenging due to scarce labeled data. We study fine-tuning strategies for multilingual transformers on Urdu news headline classification with only 4,000 labeled examples. We compare full fine-tuning, adapter-based tuning, and augmentation-enriched training under a fixed compute budget. Adapter-based tuning achieves the highest mean F1 of 0.78 (±0.02) across five seeds, outperforming full fine-tuning by 4.1 points while updating fewer than 5% of parameters. Our results suggest parameter-efficient methods are particularly suited to low-resource settings, and we release our dataset and code to support further work."
Structure the introduction as a funnel: broad context → specific problem → why it is hard/unsolved → your approach → your contributions → paper outline. The contributions are the contract of the paper — a bulleted list of 3–4 items, each a verifiable claim: "We release the first publicly available dataset of 12,000 labeled Urdu news headlines across 8 categories." Every contribution bullet must be delivered somewhere in the paper; reviewers check.
End with a roadmap paragraph ("Section 2 reviews related work..."), which orients the reader in one paragraph.
Organize by theme (Chapter 3), compare approaches on the dimensions that matter for your problem, and close the section by stating precisely how your work differs. A useful sentence pattern: "While [A] studies X in high-resource settings and [B] explores Y for a related task, neither addresses Z, which is our focus." Be generous and accurate — the authors you cite may review your paper.
Write so that a competent peer could reimplement your work. Include: data description and splits, preprocessing steps, model architecture and configuration, training procedure (optimizer, learning rate schedule, epochs, batch size), hyperparameter search spaces and selection method, evaluation metrics, and compute environment. Put standard details in briefly and unusual choices in detail — readers need to know what is standard (so they can replicate) and what is novel (so they can evaluate). If space is tight, move exhaustive details to an appendix or supplementary material, but never omit them entirely.
Present results in a logical order: main comparison first (your method vs. baselines on the primary metric), then ablations, then additional analyses. For each table or figure, the text should state the takeaway, not recite the numbers: "Adapter tuning outperforms full fine-tuning across all data sizes (Table 2), with the gap widening as data decreases — consistent with our hypothesis that fewer trainable parameters reduce overfitting." Let the visual carry the numbers; let the prose carry the interpretation.
Report negative and unexpected findings (Chapter 6). A results section that only contains victories reads as advertising.
The discussion (sometimes merged with results) answers "what does it mean?" Interpret patterns, connect findings to your hypotheses, and — crucially — state limitations plainly: data limitations, compute constraints, settings where the method may not generalize. Authors fear limitations sections; reviewers reward them, because acknowledged limitations show you understand your own work's boundaries. Never introduce new results in the discussion.
Restate the problem, summarize what you showed (mirroring the contribution bullets), and point to concrete future work — specific next questions your work opens, not vague wishes ("we will explore deep learning further"). Do not introduce new claims or citations-heavy arguments here.
Academic writing rewards clarity, not ornament. Prefer short sentences, active voice where natural ("we train" rather than "training was performed"), and precise terms defined on first use. Cut throat-clearing ("It is well known that..."), hedge only where uncertainty is real, and keep paragraphs to one idea each. Read each paragraph and ask: what is its one point? If you cannot answer, rewrite it.
Get feedback before submission: your supervisor, a senior student, and ideally someone outside your subfield. Fresh eyes catch the gaps your familiarity hides. Sharmeen's senior labmate flagged that her introduction never defined "data-efficient" — a term she had used 40 times without noticing.
Each paragraph makes a contract with the reader in its first sentence: "this paragraph is about X." Everything after must serve X. When you catch yourself writing "another important aspect is..." mid-paragraph, you are breaking the contract — start a new paragraph. Between paragraphs, use transition sentences that carry the argument forward: "Having established the method, we now evaluate whether it works under limited data." Readers skim first sentences to map your argument; make those sentences carry the skeleton of the paper. A useful test: read only the first sentence of every paragraph in a section. If the story holds, your structure is sound.
Most researchers write in English as a second (or third) language — you are in good company, and reviewers judge ideas, not native fluency. Still, a few habits help enormously. Prefer short sentences; long sentences are where grammar breaks. Use the same term for the same concept every time (not "method," "approach," "technique," "framework" rotating randomly). Avoid idioms and cultural references. Learn the genre's standard phrases by imitation: "We evaluate on...", "As shown in Table 2...", "In contrast to prior work...". Use a grammar checker, then have a fluent reader review — not to rewrite your ideas, but to catch the small errors that distract reviewers. And read your draft aloud: awkward phrasing you miss on screen becomes obvious to the ear.
Revise in passes, each with one job. Pass 1 — structure: does the argument flow? Are sections in the right order? Is anything missing or repeated? Move blocks freely; this is the time for surgery. Pass 2 — paragraphs: does each paragraph keep its contract? Are transitions smooth? Cut ruthlessly — every paper is 15% too long on first draft. Pass 3 — sentences: grammar, word choice, citation placement, number formatting. Never do Pass 3 work during Pass 1 — polishing sentences you later delete wastes hours. Sharmeen does Pass 1 on printouts with a red pen (distance from the screen helps), Pass 2 on screen, and Pass 3 the day before submission, when she is too tired to restructure anything anyway.
Headings are signposts — write them to carry meaning, not just label. "4.1 Adapter-Based Tuning" tells the reader where they are; "4.1 Why Adapters Suit Low-Resource Settings" tells them where they are and what to learn. Prefer the second style for key sections. Within the paper, use headings to make the argument skimmable: a reviewer reading only your headings and figures should grasp the paper's arc. And keep heading levels consistent — a subsection that appears once under a section suggests the structure is lopsided; either add its sibling or fold it into the parent.
Watch for these patterns in your drafts: the mystery introduction (the problem appears on page 3 — state it in paragraph one); the methodology maze (implementation details before the big picture — give the overview diagram and intuition first, details after); the results dump (tables with no narrative — every visual needs its takeaway sentence); the related-work apology (citing competitors defensively instead of positioning confidently — you belong in this conversation); and the conclusion that concludes nothing (vague future work — name specific next questions). Sharmeen's first draft had three of the five; her supervisor's margin notes were mostly structural, not technical. Structure is the highest-leverage revision: it changes how every sentence is received.
Multi-author papers need a workflow, or they become version-control nightmares. Agree upfront on: the writing order (usually lead author drafts, co-authors comment, supervisor does a final pass); the commenting convention (suggesting mode in Word, or issues/pull requests in Overleaf with git); and response discipline (comments addressed within a week, or flagged as "for later"). One person owns the master document at a time — parallel edits to the same file create merge conflicts in prose that are painful to resolve. Sharmeen's lab uses a simple rule: the lead author integrates all comments and is the only one who edits the master; co-authors comment but do not directly rewrite. It sounds rigid, but it eliminated the lost-paragraph incidents of her first paper. And keep a changelog for major drafts (v0.1 outline, v0.5 full draft, v1.0 submission) so feedback always targets the right version.
Give your introduction to someone intelligent who knows nothing about your area — a friend in another department works perfectly. After reading, ask them: what problem does this solve, and why should anyone care? If they cannot answer, your introduction is written for insiders and will lose everyone else, including reviewers outside your exact niche (and at least one reviewer always is). The fix is usually in the first two paragraphs: add the concrete stakes before the technical framing. Sharmeen's introduction originally opened with transformer fine-tuning details; the stranger test revealed a non-NLP reader was lost by sentence three. She added two opening sentences about newsrooms drowning in uncategorized wire copy — and suddenly the technical story had a reason to exist. Write the opening for the stranger; the experts will keep up.
Before submission, read the entire paper aloud — every word, including captions. Your ear catches what your eye forgives: run-on sentences, repeated words, abrupt transitions, and the dreaded paragraph that made sense at midnight but not in daylight. It takes about forty minutes for an eight-page paper, and it is the highest-value revision pass per minute at the end of the process. Mark awkward spots with a pen as you go, then fix them in one sitting.
For your research: Draft your method and experiments sections this week, while the work is fresh — aim for completeness over polish. Then write your contribution bullets (3–4 verifiable claims) and check: is each bullet delivered somewhere in the draft? If a bullet has no corresponding evidence, either run the experiment or cut the bullet. The paper's contract must balance.
Key takeaways: - Write method and experiments first, introduction and abstract last; build figures before describing them. - The abstract is a four-move miniature (problem, method, results, conclusion) with real numbers. - Introduction funnels from context to contributions; contribution bullets are a contract the paper must fulfill. - Related work positions your work thematically and states the precise difference. - Methodology must enable reproduction; results sections show then tell; discussions interpret and admit limitations. - Prefer clarity over complexity: one idea per paragraph, precise terms, feedback from fresh eyes.
Formatting is not vanity — it is professionalism. Reviewers form impressions within minutes, and a sloppy manuscript suggests sloppy research. IEEE style is the standard across much of engineering and computing. This chapter covers the mechanics: document setup, citation practice, and reference lists.
IEEE provides official LaTeX and Word templates (IEEEtran). Use them — do not approximate the layout by hand. Key rules:
Learn basic LaTeX if you plan a research career — most venues expect it, and it handles numbering, cross-references, and bibliographies automatically. Overleaf (online LaTeX editor) removes the installation barrier and enables supervisor collaboration.
IEEE uses numbered citations in order of appearance: the first cited work is [1], the next new one [2], and so on. Cite with brackets in the text: "Transformers [1] have been applied to..." When citing multiple works, group them: [2]–[5]. Every citation number must correspond to exactly one reference list entry, and every entry must be cited at least once.
Citation placement matters: put the citation immediately after the claim it supports. "Urdu is a low-resource language [7]" — the reader should never wonder which part of the sentence the citation covers.
Cite: every claim that is not common knowledge or your own contribution; all methods, datasets, and tools you used; the closest related work (fairly); and any figure, table, or text adapted from elsewhere. Do not cite: papers you have not read (read at least the relevant parts); Wikipedia or blogs as primary evidence (use them to find primary sources); and your own unrelated work just to inflate citation counts (reviewers notice).
A good rule: for every paragraph in related work and every methodological choice, ask "says who?" If the answer is a paper, cite it.
Each entry must be complete and consistent. IEEE formats:
Use "et al." for four or more authors in the text citation context (in the reference list, IEEE style lists all authors up to a limit, then et al. — check the current template). Include DOIs where available. Use a bibliography manager (BibTeX with LaTeX, or Zotero's Word plugin) rather than hand-typing — hand-typed reference lists are where formatting errors breed.
Verify every reference: open each one and confirm the title, venue, year, and pages. Sharmeen once caught a wrong year in an imported BibTeX entry that would have cited a 2019 paper as 2016 — a small error that signals carelessness to reviewers.
You can be productive in LaTeX within a day by learning a small core. On Overleaf, start from the IEEEtran template. The essentials: \section{}, \subsection{} for structure; \cite{key} for citations; \label{} and \ref{} for cross-references ("as shown in Fig.~\ref{fig:results}"); \begin{figure}...\end{figure} with \includegraphics and \caption; \begin{table}...\end{table} with tabular; and math mode $...$ for inline symbols. Compile often — errors compound. The golden rule: never type a figure number, citation number, or section number by hand; always use labels and references, so reordering never breaks numbering. When LaTeX throws a cryptic error, comment out the last thing you added — binary search beats staring.
Your .bib file is a database — keep it clean. Use meaningful keys (vaswani2017attention, not paper1). Fill every field the style needs: author, title, booktitle/journal, year, pages. Watch for common import corruptions: HTML entities in titles, wrong capitalization (protect capitals with braces: {BERT}), and conference names mangled by Scholar. One corrupted entry can break compilation the night before the deadline — audit the .bib file when you add entries, not at submission time. Keep a single master .bib across projects; it compounds in value.
After acceptance, you will sign a copyright transfer or license. Typically you grant the publisher exclusive publication rights while retaining rights to: share the preprint, include the work in your thesis, and reuse figures in future work with credit. Read the form — it is short. If your funder or university requires open access, choose the appropriate option before signing; some agreements are hard to reverse. For conference papers, the form is usually electronic and takes five minutes; for journals, it may arrive with the proofs. When in doubt, ask your supervisor — they have signed dozens.
Errors reviewers see repeatedly: mixing citation styles mid-paper; references in the list never cited in text (or cited but missing); figure captions below figures but table captions also below tables (IEEE: table captions go above); equations unnumbered; "Fig." vs "Figure" inconsistency; author names in the anonymized version's PDF metadata (check Document Properties before submitting); and exceeding page limits by tweaking margins — program chairs notice, and some venues desk-reject for it. Run through the checklist in this chapter 72 hours before the deadline, when there is still time to fix what you find.
Modern papers build on more than papers — cite all of it. Datasets get full citations like papers (many have published data papers; cite those, e.g., the ImageNet database paper [10]). Software and libraries should be cited or at least named with versions ("implemented in PyTorch 2.0") — some packages have preferred citations; check their docs. Preprints are citable: include the arXiv identifier. A common student error is using a dataset or tool extensively while citing nothing — it looks like the work sprang from nowhere and denies credit to the creators. When no formal citation exists, cite the project page or documentation URL in a footnote. Generous citation of infrastructure is a hallmark of careful researchers.
Zotero, Mendeley, and JabRef all work; what matters is using one consistently. Zotero is free, open-source, and excellent at grabbing metadata from publisher pages with one click. Whichever you choose, develop these habits: import at discovery time (never "later"), verify metadata on import (authors, year, venue — Scholar imports are often wrong), attach the PDF immediately, and use collections/tags for projects. Connect it to your writing: Zotero's Word plugin or Better BibTeX for LaTeX keeps citations and the reference list synchronized automatically. The students who struggle with references at deadline time are invariably the ones who managed citations by hand. Do not be one of them.
This chapter teaches IEEE style, but you will encounter others. ACM uses author–date or numbered citations depending on the venue, with its own acmart template. NeurIPS/ICML/ICLR provide their own LaTeX styles with specific rules (e.g., NeurIPS requires line numbers in submissions). Some journals want author–date citations (Smith, 2020). The principle never changes: find the venue's official template and author guidelines, and follow them exactly — do not adapt an IEEE-formatted paper by hand for an ACM venue. Reference managers handle style switching gracefully (one click in Zotero), which is another reason to use one. When a venue's guidelines contradict your habits, the guidelines win, every time. And always check the current year's guidelines — templates and rules change, and last year's paper is not a reliable template.
Every reference should help a reader find the work. Include the DOI whenever one exists — it is the permanent link that survives publisher website redesigns. For arXiv preprints, the arXiv identifier serves the same role. Use URLs sparingly and only for works with no DOI (project pages, datasets, software documentation); when you do, add an access date, since web pages change. Never cite a bare URL without author, title, and year — "http://example.com/paper.pdf" in a reference list is not a citation, it is a shrug. And verify every link before submission: one dead link in your references suggests the rest were not checked either. Your reference manager can store DOIs and URLs in dedicated fields — fill them at import time, not at 2 a.m. before the deadline.
Learn these by recognition, not by suffering: unescaped special characters (&, %, _ in titles break compilation — write \&); lost capitalization (BibTeX lowercases titles, so protect proper nouns: {Urdu}, {BERT}, {ImageNet}); mangled author names ("LeCun, Yann" vs "Yann LeCun" — pick the Last, First format consistently); wrong entry types (conference papers as @article — use @inproceedings); and duplicate keys (two entries sharing a key silently shadow each other). When compilation fails with a cryptic bibliography error, the culprit is almost always the most recently added entry — check it first. Keep a "known good" example entry at the top of your .bib file as a template to copy.
Do one last dedicated formatting pass after the content is frozen: search for double spaces, inconsistent hyphenation (fine-tune vs fine tune — pick one), straight quotes that should be curly in LaTeX (use `` and ''), and units with missing non-breaking spaces (6~hours in LaTeX). Check that every in-text citation has its bracket closed and every reference-list entry is actually cited — automated tools catch most of this, but a human sweep catches the rest. Print the PDF and flip through it physically: formatting glitches invisible on screen (a widowed heading, a figure overlapping text) jump out on paper. This unglamorous hour is what separates manuscripts that look professional from ones that merely read well.
For your research: Create your BibTeX/Zotero library now and add every paper from your reading-note spreadsheet with complete, verified entries. Then take one page of your draft and audit every claim for its citation — each factual or borrowed statement should have one. This audit habit, done early, prevents the last-minute citation scramble that introduces errors.
Key takeaways: - Use the official IEEE template (LaTeX/IEEEtran preferred); page limits and layout rules are strict. - IEEE citations are numbered in order of appearance; place each citation directly after the claim it supports. - Cite everything borrowed or factual; never cite papers you have not read. - Keep reference entries complete, consistent, and verified — use a bibliography manager, not hand-typing. - Number and caption all figures, tables, and equations; reference each in the text. - Anonymize thoroughly for double-blind review, including PDF metadata.
Readers skim papers visually before they read them: the figures, the tables, the section structure. Strong visuals can carry your argument; weak ones bury it. This chapter is about making every visual earn its place — each one should make a single point clearly enough to be understood almost on its own.
Before creating any figure or table, write down its one message in a single sentence: "Adapter tuning beats full fine-tuning at all data sizes, and the gap grows as data shrinks." If you cannot state the message, you do not need the visual. If a visual tries to make three points, split it. Sharmeen's first results figure crammed six curves into one plot; splitting it into two focused figures made both the main result and the ablation instantly readable.
Avoid: 3D charts (distort perception), pie charts for precise comparison (humans judge angles poorly), and screenshots of code or terminals (retype as formatted text).
Tables present precise comparisons. Rules: a clear title/caption above the table; column headers that name the metric and dataset; one consistent number of decimal places (two or three — more is false precision); bold best results; and variance reported for every number that comes from multiple runs. Include the baseline rows a reviewer expects — a table without the obvious competitor invites the question "why did they hide it?"
Example structure for Sharmeen's main table: rows = methods (TF-IDF baseline, full fine-tuning, adapter tuning, augmentation-enriched), columns = F1 at 1k/2k/4k examples plus parameter count. One table, one message: adapter tuning wins, especially with little data, while training far fewer parameters.
Draw your architecture or pipeline early — it forces you to clarify the method before writing about it. Show inputs on the left, outputs on the right, and label every transformation. A reader who understands your diagram will forgive dense prose; a reader confused by your diagram will not trust the prose. Keep it to one figure; if the method has two stages, two panels in one figure beat two figures.
Every figure and table must be referenced in the text, in order: "As shown in Fig. 2..." Do not write "the following figure" — with two-column layouts, floats move. And never leave a visual undiscussed: if the text never mentions Table 3, delete Table 3.
Use whatever produces vector graphics (PDF/SVG/EPS): matplotlib with PDF output, or drawing tools that export vector formats. Vector figures stay sharp at any size; PNG screenshots turn to mush. Set your plotting library's font to match the paper body where possible — visual consistency signals care.
About 1 in 12 men has some color-vision deficiency — including reviewers. Never encode meaning in color alone: combine color with line styles, markers, or direct labels. Choose colorblind-safe palettes (blue/orange instead of red/green) and test your figure in grayscale: print it or use a simulator. Keep sufficient contrast between text and background, and avoid tiny fonts that strain every reader. Accessibility is not extra credit; it is part of clarity, and unclear figures get papers rejected.
Do not hand-tweak figures in a GUI the night before the deadline — you will never reproduce them. Instead, script everything: a Python script that reads your logged results and writes the PDF figure. The script lives in git next to the experiment code. When a reviewer asks for a new baseline in the revision, you re-run the script instead of rebuilding the figure by hand. Sharmeen's plotting script takes her results CSV and regenerates all six paper figures in seconds — her revision round, which added two methods, cost her an afternoon instead of a week. Version your figures like code, because they are code output.
Sharmeen's first main-results figure had six methods, a legend covering data points, 8pt axis labels, and a y-axis from 0.72 to 0.80 that made a 1% gap look enormous. The rescue took thirty minutes: split into two figures (main comparison; ablation), cut to four methods, moved the legend outside, bumped fonts to 10pt, set an honest axis, and wrote a caption stating the takeaway. Same data, completely different persuasiveness. When your figure feels "off," run this rescue checklist: fewer elements, bigger text, honest axes, outside legend, takeaway caption. Nine times out of ten, the figure was fine — the presentation was not.
Students often default to tables for everything. Use this guide: choose a figure when the message is a pattern — trends, gaps widening, clusters, distributions. Choose a table when the message is precise values — exact scores a reader might cite or compare against. When both matter, do both: a figure for the pattern in the main paper, the full table in the appendix. Never present the same numbers as both a figure and a table in the main text — it wastes space and suggests you could not decide what matters. And remember the hierarchy: the main paper carries the 2–4 visuals that prove your claims; everything else (extra ablations, per-class breakdowns, hyperparameter tables) belongs in supplementary material, referenced once.
The best figures in top papers often contain a small annotation — an arrow, a highlighted region, a text label — pointing at the key observation: "gap widens below 2k examples." Annotations respect the reader's time: instead of hunting for the pattern, they see it, then verify it in the data. Keep annotations minimal (one per figure) and factual (describe what is shown, not what it proves — the text argues the proof). Sharmeen annotated her main figure with a bracket around the low-data region where adapters pulled ahead; reviewers mentioned the figure positively by name in two reviews. A figure that teaches is a figure that persuades.
You do not need design training — you need defaults that work. For categorical data (comparing methods), use a colorblind-safe palette like blue, orange, teal, and purple — never red/green. For sequential data (heatmaps, intensity), use a single-hue gradient from light to dark. In matplotlib, set a global style at the top of your plotting script: larger default fonts (11–12pt), thicker lines (2pt+), and no top/right spines for a clean look. Save directly to PDF. These five settings take ten lines of code and instantly make student figures look professional. Keep a shared paper_style.py in your lab so every member's figures match — visual consistency across a group's papers builds a recognizable, credible identity.
Treat figure files like code: name them meaningfully (fig2_adapter_vs_full.pdf, not figure_final_v3_REAL.pdf), keep the generating script alongside, and never hand-edit a figure file — regenerate it. Store figures in a figures/ directory in git, with the script that produced each one. When a reviewer asks for a change, you edit the script and regenerate; the history shows exactly what changed. Sharmeen once spent an evening hunting for which of seven plot_final.png files was actually in the paper — after that, her lab adopted the rule: if the script cannot regenerate it, it does not go in the paper. This discipline pays off most during revisions, when figures change under time pressure and mistakes are easiest to make.
Your presentation needs different visuals than your paper. Slides are seen from meters away for seconds at a time: enlarge fonts dramatically (nothing below 24pt), show one message per slide, and simplify ruthlessly — a paper figure with six curves becomes a slide with two. Keep the same colors per method so the audience connects talk to paper. And never screenshot your paper's PDF into slides; regenerate from the script at slide size. Sharmeen rebuilds her three key figures as large-format versions for every talk — thirty minutes of work that makes the difference between an audience that follows and one that squints.
Before finalizing any visual, show it to someone for ten seconds, hide it, and ask what they remember. If they recall your message, the visual works; if they recall confusion ("there were a lot of lines"), it does not. This test is brutally honest and takes almost no time — run it on every figure that carries a core claim. Sharmeen's labmate failed her first main figure in four seconds flat ("something about green being higher?"), which is what triggered the rescue described above. A figure that passes the ten-second test will survive a reviewer's thirty-second skim, which is the real examination every visual faces.
For your research: Take your single most important result and sketch three different visual forms for it (e.g., line plot, bar chart, table). Show all three to a labmate and ask which message they get in ten seconds. Use the winner. This ten-minute test consistently produces better figures than hours of solitary polishing.
Key takeaways: - One visual, one message; split visuals that try to do more. - Match the form to the message: line plots for trends, bars for few-way comparisons, tables for precise numbers, diagrams for methods. - Design for column-width legibility: large fonts, self-contained captions, consistent colors, honest axes, minimal clutter. - Tables need consistent decimals, bold best results, variance, and the expected baselines. - Reference every visual in order in the text; delete any visual the text never discusses. - Export vector graphics; test readability by shrinking to print size.

The reviews arrive. Your stomach drops. Reviewer 2 "does not see the novelty." Reviewer 1 wants three new experiments. This moment — the first round of real criticism — is where many students stall for weeks. This chapter reframes peer review as what it actually is: free, expert consulting on your work, delivered bluntly. Your job is to extract every useful bit and respond with calm professionalism.
Do not respond for 48 hours. Read the reviews once, then close them. Your first reading will be emotional — you will notice the harshest sentences and miss the helpful ones. After two days, re-read with a pen: mark every actionable point (a specific request, question, or criticism tied to the paper) and separate it from tone (blunt phrasing, which carries no information). "The baseline comparison is unfair because X used less tuning" is actionable; "this is incremental" is a judgment you address with evidence and framing, not feelings.
Sharmeen's first reviews included "the Urdu dataset is too small to support the claims." Her first reaction was despair; her second reading recognized it as actionable — she could add a data-size analysis showing at what size the effect appears, and soften the claim's scope. That response turned a weakness into one of the paper's better sections.
For journals (and conferences with rebuttals), you write a response letter addressing every reviewer point. Structure it:
Dear Editor / Reviewers,
We thank the reviewers for their careful reading and constructive feedback.
Below we address each comment. Reviewer comments are in italics; our
responses follow. Changes to the manuscript are highlighted in blue.
Reviewer 1:
1. Comment: "..."
Response: ... [what we changed and where: "see Section 4.2, paragraph 3"]
...
Tone rules: thank sincerely (they did unpaid work), never argue about tone, concede valid points plainly ("The reviewer is correct; we have..."), and push back on invalid ones with evidence, not attitude ("We respectfully disagree because...; see new Figure 4"). Never insult, never grovel — professional peers discussing work.
Number every reviewer point and answer every one — including the ones you disagree with and the minor typos. A reviewer who sees their point 7 unanswered assumes you ignored it. For each point, do three things: (1) state what you understood, (2) describe what you changed (with location), (3) explain briefly why the change addresses the concern. If you did not change something, explain why with evidence: "We did not add baseline X because it addresses a different task (citation); however, we now discuss this explicitly in Section 2."
Make the changes visible: many venues ask for highlighted diffs. More importantly, make each change actually address the point — superficial edits ("we added a sentence") on substantive criticisms ("the evaluation is flawed") will be caught in the next round. After revising, re-read the full paper: piecemeal edits often break flow, and a paper that reads as patched-together invites new criticism.
If the decision is rejection, do the 48-hour rule, extract the actionable feedback, revise, and submit to the next venue on your list (Chapter 11). Do not resubmit the identical manuscript elsewhere — reviewers overlap between venues, and an unchanged resubmission wastes everyone's time including yours. Every published researcher has a rejection collection; the difference between them and others is that they kept submitting.
Many AI conferences include a rebuttal phase: reviewers ask questions, and you get a short response (often one page) plus a brief window. This is not the time for new experiments — it is the time for clarity. Prioritize: correct factual misunderstandings first (a reviewer who misread your method may flip with one clear paragraph), then answer questions with pointers to existing evidence ("see Table 3, row 4, which already compares..."). Be concise, be specific, and never promise experiments you cannot deliver — but you may briefly describe planned additions for the final version if accepted. Sharmeen's rebuttal corrected a reviewer's belief that she had tuned on the test set (she had not — Section 3.2 said so, now clarified), and that single correction moved her paper from borderline to accept. Rebuttals reward papers that were clear but misread, so write the original paper to survive misreading.
Sometimes a reviewer demands what you cannot do: a dataset that does not exist, compute you do not have, a comparison against a proprietary system. Do not panic and do not fake it. Respond with a three-part pattern: acknowledge the value of the suggestion, explain the concrete constraint honestly, and offer what you can do instead. "We agree multilingual evaluation would strengthen the paper; the annotated corpus for language X is not publicly available, and collecting one is beyond this paper's scope. Instead, we have added an analysis of performance across headline categories (new Section 4.4), which tests generalization within our setting." Reviewers are researchers too — they recognize genuine constraints when stated plainly, and they punish evasiveness.
Journal papers often go through two or three rounds. Each round should converge: round one addresses the big issues (method, evaluation), round two polishes (clarity, additional analysis), round three is minor. If round three raises brand-new fundamental objections, something went wrong earlier — usually a misunderstanding that should have been caught in round one. Keep a revision map: a table of every reviewer point across rounds, your response, and where the change lives. It prevents you from contradicting an earlier response and shows the editor the paper's trajectory. And keep perspective: a "major revision" decision means the editor believes the paper can be accepted. That is good news wearing a scary costume.
Rejection hurts, especially the first time. Use a protocol so emotion does not make decisions:
Sharmeen's first submission was rejected with a painful but accurate critique of her baseline tuning. Three weeks of re-tuning later, the paper was accepted at her target venue — and the reviews there were notably warmer, because the work was genuinely better. Rejection, processed through this protocol, is just peer review with extra steps.
To make the point-by-point method concrete, here is how Sharmeen answered her toughest review comment:
Reviewer 2, point 3: "The baseline comparison seems unfair — the full fine-tuning baseline was trained for fewer epochs than the proposed method."
Response: "We thank the reviewer for catching this. The reviewer is correct: the baseline used 10 epochs while our method used 15. We have re-run the full fine-tuning baseline for 15 epochs with the same learning-rate schedule (see updated Table 2). The baseline's mean F1 improved from 0.71 to 0.73, and our method's advantage remains at 5 points (0.78 vs. 0.73). We have also added the epoch counts for all methods to Section 3.3 for transparency."
Note the anatomy: thanks, concession, concrete fix, new evidence, and a transparency improvement that prevents recurrence. The reviewer did not just get an answer — they got a better paper. This is the standard to aim for on every substantive point.
If you believe a rejection resulted from a clear procedural error — a reviewer with an undisclosed conflict, a factual mistake the rebuttal could not address, a review of the wrong paper — most venues allow an appeal to the program chair or editor. Appeals succeed only when you can document the error concretely and briefly; "the reviewers didn't appreciate my work" is not grounds. In practice, appeals are rare and rarely succeed, and they cost weeks. The usually better path is the rejection protocol: revise and route to the next venue. Sharmeen's supervisor's rule: appeal only if you can state the procedural error in one sentence with evidence; otherwise, the energy goes into the next submission. Your paper's goal is publication and impact, not winning an argument with one venue.
When you resubmit a revised manuscript, include a brief cover letter to the editor: thank the reviewers, summarize the major changes in three to five bullet points, and note anything unusual (e.g., a requested experiment you could not do, with reasons). Keep it to one page — the detailed responses live in the response letter. The cover letter's real audience is the editor, who decides whether the revision goes back to reviewers or is accepted directly. A clear, confident summary ("we re-ran all baselines with equal tuning budgets; the main conclusions hold and are now stronger") helps the editor see the trajectory at a glance. Attach the response letter and the highlighted manuscript as the venue instructs, and double-check that the uploaded files are the final versions — submitting a draft with tracked changes visible is a surprisingly common mishap.
Not all criticisms weigh the same. Learn to read sentiment: a reviewer who writes "the paper would be stronger with..." is broadly positive — address the suggestion and you likely gain an ally. One who writes "I am not convinced that..." is undecided — your rebuttal or revision must supply the missing evidence, not just words. One who writes "the contribution is unclear" is signaling a framing problem, often fixable in the introduction rather than the experiments. And a reviewer whose every point is minor (typos, clarifications) has essentially accepted your work — thank them warmly and fix everything promptly. Calibrate your effort to the sentiment: major energy for the unconvinced, careful completeness for everyone.
For your research: Before you ever receive reviews, practice: ask your supervisor or a senior student to review your draft as harshly as a real reviewer would. Then write a practice response letter addressing their points. This rehearsal teaches the point-by-point discipline while the stakes are zero — and the paper you submit will already be stronger for having survived one hostile reading.
Key takeaways: - Wait 48 hours before responding; separate actionable points from tone on the second reading. - Answer every point, numbered, with what you changed and where — including points you disagree with. - Tone is professional peer-to-peer: concede valid points, rebut invalid ones with evidence, never engage with rudeness. - Do requested experiments that test core claims; set honest scope boundaries for the rest. - Make changes visible and re-read for flow; revise-then-resubmit elsewhere after rejection, never resubmit unchanged.
Where you submit shapes how you write, how long you wait, and who reads your work. Students often default to wherever their supervisor suggests without understanding the tradeoffs. This chapter gives you the framework to choose deliberately — and a shortlist strategy so rejection never leaves you stranded.
| Dimension | Journals | Conferences |
|---|---|---|
| Timeline | Rolling submission; 6–18 months to publication | Fixed deadlines; ~4–6 months submit-to-publish |
| Paper length | Longer (often 10–15+ pages); room for depth | Shorter (typically 6–9 pages); forces focus |
| Revision | Multiple rounds with the same reviewers | Usually accept/reject; rebuttal clarifies but rarely adds experiments |
| Prestige (AI/CS) | Top journals highly prestigious | Top conferences equal or higher prestige; faster visibility |
| Presentation | None required | You present; networking and feedback |
| Best for | Mature, archival, comprehensive work | Timely results; first publications; community feedback |
In AI and machine learning, conferences (NeurIPS, ICML, ICLR, AAAI, CVPR, ACL/EMNLP for language work) are the primary venues — fast, prestigious, and presentation-driven. Journals (IEEE Transactions, Journal of Machine Learning Research, Nature Machine Intelligence) suit longer, consolidated contributions. Neither is "better" universally; match the venue to the work.
Ask five questions:
Sharmeen's paper — a focused empirical study with one dataset — fit a conference better than a journal: the contribution was timely, the page limit forced clarity, and she wanted presentation feedback before extending the work into her thesis.
Never target a single venue. Before submitting, build a ranked shortlist of three:
If the reach rejects, the paper goes to the target with revisions informed by the reviews. This pipeline thinking removes the emotional catastrophe of rejection — it becomes a routing decision, not a verdict.
The call for papers (CFP) is your specification document. Extract: submission deadline (and timezone — deadlines are unforgiving), page limits and whether references count, formatting template, anonymization rules, preprint policy, supplementary material rules, and review timeline. Put the deadline in your calendar with reminders at 4 weeks, 2 weeks, and 3 days. Sharmeen's lab rule: the paper is "done" 72 hours before the deadline, leaving a buffer for formatting disasters, which always happen.
Warning signs: unsolicited emails inviting you to submit ("Dear esteemed researcher"); guaranteed fast publication; vague or absent peer review description; fees that appear only after acceptance; names mimicking reputable venues; no indexed proceedings. Check: is it indexed in Scopus, Web of Science, or DBLP? Do researchers you respect publish there? When in doubt, ask your supervisor — one question now saves a line on your CV you will later want to remove. (More in Chapter 12.)
Acceptance brings a short, intense phase: format the camera-ready version to the publisher's template, sign the copyright form, check proofs carefully (errors in proofs are your responsibility), and register for the conference. Then prepare the presentation: a 10–15 minute talk is not the paper read aloud — it is the story of the paper (problem → idea → evidence → takeaway), with the single best figure on screen as much as possible. Practice with your lab. Presenting well multiplies the paper's impact.
Venues publish acceptance rates — 20% for a top conference, 40% for a mid-tier one — and students treat them as odds. They are not odds; they are outcomes of a self-selected pool. A 20% acceptance rate at a top venue reflects submissions from the world's best labs; your personal chance depends on your paper's strength relative to that pool, not on the raw number. More useful signals: does the venue publish work like yours (scope fit beats prestige), do people you respect attend, and will the feedback improve your work? A thoughtful paper at a well-matched 35%-acceptance venue beats a rushed paper desk-rejected from a famous one. Also note: acceptance rates vary wildly by track — workshops attached to top conferences often accept 50–60% and give you the same audience for feedback.
Two underused routes deserve attention. Journal special issues — themed collections on a topic — often have higher acceptance rates than regular issues and put your work beside directly related papers, increasing its visibility to exactly the right readers. Workshops (Chapter 1) are the ideal first submission: shorter papers, faster cycles, expert audiences, and reviewers who expect preliminary work. Sharmeen's lab has a standing rule: every new project first targets a workshop paper within six months. The workshop deadline forces an early result, the feedback shapes the full paper, and the student gets a publication on the board fast — which does wonders for morale during the long conference cycle.
Turn venue selection into a calendar, not a one-time decision. Once a year, sit with your supervisor and map: which conferences have deadlines in the next 12 months, which fit your project's trajectory, what each requires. Work backwards from each candidate deadline: experiments frozen 6 weeks before, full draft 3 weeks before, internal review 2 weeks before, submission buffer 72 hours. Put every date in a shared calendar with reminders. Research runs on deadlines — self-imposed ones slip, venue ones do not. Sharmeen's second paper existed because a workshop deadline nine months out went into her calendar in January; without it, the experiments would have drifted indefinitely. The calendar is the difference between "working on a paper" and "submitting a paper."
Beyond traditional journals, open-access mega-journals (such as IEEE Access, Scientific Reports, or PLOS ONE) publish technically sound work across wide scopes with faster review. They are legitimate, indexed venues — useful when your work is solid but not a fit for a selective topical venue, or when you need an archival publication on a timeline. The tradeoff is prestige signaling: a mega-journal paper says "sound work," not "top-tier selective." For a student's early publications, that is often exactly right — a published, citable, openly accessible paper beats a perpetually "under review" one. Just ensure the mega-journal is genuinely peer-reviewed (check indexing and editorial boards), since the open-access model is also where predatory journals hide.
Major conferences host workshops, tutorials, and competitions alongside the main event. A single trip can thus serve multiple goals: present a workshop paper early in the week, attend tutorials to learn, and network at the main conference. Some students also enter shared-task competitions (with leaderboards and overview papers) — a strong competition entry with a system-description paper is a legitimate publication and excellent training in rigorous evaluation. Plan conference travel around these co-located opportunities; the marginal cost of adding a workshop submission to a trip you are already taking is small, and the feedback is disproportionately valuable.
Here is what Sharmeen's three-venue shortlist actually looked like for her fine-tuning study:
She submitted to the reach first, was rejected with useful reviews, revised for six weeks, and was accepted at the target. The safety net was never needed — but knowing it existed made the reach submission psychologically possible. Build your shortlist with this level of concreteness: named venues, real deadlines, honest reasoning. Vague shortlists ("some IEEE conference") produce vague plans.
The best venue information is not on websites — it is in people's heads. Ask senior students and recent graduates: "Where did you submit this kind of work? How was the reviewing? Would you go back?" They know which venues give constructive reviews, which are slow, which have scope quirks the CFP does not mention, and which program chairs run a tight process. Ask your supervisor which venues they trust as a reviewer — people review for venues they respect. One caution: treat secondhand prestige opinions skeptically ("everyone knows venue X is bad") unless backed by specifics; the academic rumor mill is strong but not always accurate. Combine network intelligence with your own reading of recent proceedings. Between the two, you will rarely choose badly.
Here is what a realistic first-paper year looks like, working backwards from a conference deadline in month 12:
Pin this to your wall. When month 5 arrives and experiments are not done, the timeline tells you exactly how much trouble you are in — and what to cut (scope, not rigor) to recover.
After each submission cycle, write a short diary entry: the venue, the outcome, what the reviews taught you, how the timeline felt, and whether you would submit there again. Over a PhD, this diary becomes a personal venue guide more valuable than any ranking — it records your experience with reviewing quality, turnaround times, and audience fit. Share anonymized notes with your lab; a shared lab diary prevents every student from learning the same venue lessons the hard way. Sharmeen's lab diary now covers fourteen venues, and new students consult it before building their first shortlist. Your future self, choosing venues for paper three and four, will thank you.
For your research: Build your three-venue shortlist this month: one reach, one target, one safety — each with its next deadline, page limit, and preprint policy written down. Then plan backwards from the target venue's deadline: writing complete 3 weeks before, experiments frozen 6 weeks before. A venue-driven schedule is the most reliable cure for the never-finishing paper.
Key takeaways: - In AI/CS, top conferences rival journals in prestige and offer speed plus presentation; journals suit longer archival work. - Match on five fits: scope, maturity, timeline, selectivity, and practical costs. - Keep a ranked shortlist (reach, target, safety) so rejection becomes routing, not catastrophe. - Treat the call for papers as a specification: deadlines, page limits, anonymization, and preprint policy. - Vet venues for predatory signs; check indexing and ask your supervisor when unsure. - Plan backwards from the deadline; finish 72 hours early; prepare the talk as a story, not a recitation.
Everything in this book is useless — worse than useless — without integrity. Research is a trust system: readers trust that your data is real, reviewers trust that your words are yours, and the public trusts that science is honest. One ethical breach can end a career before it starts. The rules are clear, and this chapter makes them plain.
Plagiarism is presenting someone else's words, ideas, figures, or data as your own without attribution. It includes:
The defense is simple and absolute: when in doubt, cite. Paraphrase genuinely (read, close the source, write from understanding), cite the source, and keep notes on which ideas came from where (Chapter 3's reading notes do exactly this). Universities run similarity checkers (Turnitin, iThenticate) — assume your paper will be screened, because it will be.
A note on AI writing tools: using them for language polishing is increasingly accepted, but passing off generated text as your own analysis, or generating fake references, is misconduct. Check your venue's AI-use policy and disclose as required. Never let a tool invent citations — fabricated references are a fast-growing cause of retractions.
Authorship credit follows contribution, not hierarchy. Standard criteria (used by IEEE and most publishers): substantial contribution to the conception, design, execution, or interpretation of the work; drafting or critically revising the manuscript; and approving the final version. Everyone listed must meet these; everyone who meets these must be listed.
Practical rules: - Discuss authorship early — at the project's start, not at submission time. Agree on author order then; first author is typically the main contributor, last is often the supervisor. - Gift authorship (adding someone who did not contribute) and ghost authorship (omitting someone who did) are both misconduct. - Acknowledge non-author contributions (data providers, helpful colleagues, funders) in the acknowledgments section. - The corresponding author handles submission and correspondence — usually the supervisor or the lead student.
Sharmeen's lab holds a 15-minute authorship conversation at project kickoff and revisits it if contributions shift. Awkward for fifteen minutes; prevents months of conflict later.
Fabricating data, falsifying results, or selectively reporting to mislead are the gravest offenses in research. The lines are bright:
The honest practices from Chapter 6 — logging everything, reporting negative results, protecting the test set — are not just good science; they are your ethical armor. A complete lab notebook is the best defense against any question about your integrity.
As an author: do not attempt to discover your reviewers' identities in double-blind review; do not submit the same manuscript to two venues simultaneously; disclose conflicts of interest. As a reviewer (which you may become sooner than you think): keep manuscripts confidential, review promptly or decline, judge the work not the authors, and disclose conflicts. The system runs on these norms.
Predatory venues exploit the pressure to publish: they charge fees, promise rapid publication, and provide little or no real peer review. Publishing in them can permanently taint your record — the paper is effectively unpublishable elsewhere afterwards, and the venue on your CV signals poor judgment. Red flags from Chapter 11 apply. When uncertain, consult your supervisor and check indexing in Scopus, Web of Science, or DBLP. Legitimate venues never solicit you by flattery email.
If you suspect misconduct in work you encounter — plagiarized text, duplicated figures, impossible results — raise it privately with your supervisor first, not publicly. Institutions have formal processes for a reason: accusations are serious, evidence standards matter, and mistaken public accusations harm innocents. Your responsibility is to not participate, to keep your own work clean, and to escalate concerns through proper channels.
For your research: This week, do three things: (1) run your current draft through your university's similarity checker and inspect every flagged passage; (2) hold the 15-minute authorship conversation for your project and write down what was agreed; (3) verify that every figure and dataset in your draft is either yours or properly credited with permission. These three actions close the most common ethics gaps students fall into.
Key takeaways: - Plagiarism includes copy-paste, paraphrase, idea theft, self-plagiarism, and figure/data copying — when in doubt, cite. - Paraphrase from understanding, keep source notes, and assume similarity screening; never let AI tools invent references. - Authorship follows contribution: discuss and agree on authors early, avoid gift and ghost authorship, acknowledge the rest. - Never fabricate, falsify, or selectively report; pre-decide analyses and keep complete records. - Honor peer-review norms: no simultaneous submissions, no unblinding attempts, confidentiality as a reviewer. - Avoid predatory venues; escalate suspected misconduct privately through proper channels.
[1] A. Vaswani et al., "Attention is all you need," in Proc. 31st Int. Conf. Neural Information Processing Systems, Long Beach, CA, USA, 2017, pp. 5998–6008.
[2] Y. LeCun, Y. Bengio, and G. Hinton, "Deep learning," Nature, vol. 521, no. 7553, pp. 436–444, 2015.
[3] A. Krizhevsky, I. Sutskever, and G. E. Hinton, "ImageNet classification with deep convolutional neural networks," in Proc. 26th Int. Conf. Neural Information Processing Systems, Lake Tahoe, NV, USA, 2012, pp. 1097–1105.
[4] I. Goodfellow, Y. Bengio, and A. Courville, Deep Learning. Cambridge, MA, USA: MIT Press, 2016.
[5] D. P. Kingma and J. Ba, "Adam: A method for stochastic optimization," in Proc. 3rd Int. Conf. Learning Representations, San Diego, CA, USA, 2015.
[6] K. He, X. Zhang, S. Ren, and J. Sun, "Deep residual learning for image recognition," in Proc. IEEE Conf. Computer Vision and Pattern Recognition, Las Vegas, NV, USA, 2016, pp. 770–778.
[7] S. Hochreiter and J. Schmidhuber, "Long short-term memory," Neural Computation, vol. 9, no. 8, pp. 1735–1780, 1997.
[8] D. E. Rumelhart, G. E. Hinton, and R. J. Williams, "Learning representations by back-propagating errors," Nature, vol. 323, no. 6088, pp. 533–536, 1986.
[9] R. Kohavi, "A study of cross-validation and bootstrap for accuracy estimation and model selection," in Proc. 14th Int. Joint Conf. Artificial Intelligence, Montreal, QC, Canada, 1995, pp. 1137–1143.
[10] J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, "ImageNet: A large-scale hierarchical image database," in Proc. IEEE Conf. Computer Vision and Pattern Recognition, Miami, FL, USA, 2009, pp. 248–255.
End of Book 49.