# Research, Rewritten > Using AI to Produce Knowledge You Can Trust. Deep research gives you plausible. This book gives you reliable. It wires general-purpose AI into every step of truth-seeking work and says where you can hand it to AI, where you must gate it yourself, and where nobody knows yet. --- # Preface Β· This Book Was Put on Trial !!! info "Chapter companion" πŸ’» [`code/smol-army`](https://github.com/hallieren/research-rewritten/tree/main/code/smol-army/) Β· πŸ’» [`code/persona-panel`](https://github.com/hallieren/research-rewritten/tree/main/code/persona-panel/) You have probably just read a report written by AI, thought it looked better than the last one you wrote by hand, and stopped with your hand over the forward button. That one second of stopping is the whole problem this book deals with. Before saying how it helps you, first what it did to itself. For a book that teaches you not to trust AI output too readily, the most embarrassing way to die is for its own claims to fail its own process. Before writing I set one house rule. **Every step of the process this book teaches gets run on the book itself first.** This preface puts the bill in front of you. The bill is also the first piece of evidence for the book's claim. The claim first. AI drove the cost of producing plausible to the floor. A report with a clean structure, full citations, and confident wording went from three weeks to three hours. The cost of reliable barely moved. The crack that opened between them is the whole content of this book. The craft of research has a set of steps whose only job is to turn plausible into reliable. Sharpen the question, lock in the criteria, run it so it can be reproduced, interrogate the results, deliver honestly, red-team your own work. A criterion is the pass-or-fail line written in advance, and it cannot be changed afterward. Red-teaming means picking holes in your own claim before anyone else does, the same move as red-teaming a model in AI safety. As long as a wrong answer means something real gets hurt, these steps are your protection, and academic etiquette is the least important of their identities. This book writes out the full process for one kind of person, the software or AI engineer who has to do research at work. You are asked to sign off on a technical judgment (can a small model replace the large one, should the retrieval layer be swapped, can the numbers in this report be trusted), and the evidence is a computational experiment anyone can rerun. The signature is yours. If you are not that person, you can still read it. Researchers, analysts, and people who make decisions use the same process, with the shape of the case converted using the conversion table in Chapter 3, section 3.6, and the paragraphs below make that clear. Now the bill. This book took several cuts of its own, each one on record. A claim in the first draft of Chapter 4 was shot dead on the spot by the book's own process. I wrote, with some excitement, "cost-matched comparisons are a gap in the literature." The wording was pretty, and it happened to raise the value of my own experiment. An independent channel check, which is to say opening a brand-new session and handing the question to a different model that did not know which answer I was hoping for, traced the forward citations, the papers that later cited it, and came back. The comparisons exist, and they cluster on the skeptics' side. The claim died. Its corpse is pinned to the map in Chapter 13 that labels the status of every claim. **Two experimental lines were run with real money, and one of them did not survive to the finish.** The book's spine case (a team of small open-source models vs a single frontier model, where frontier means the strongest commercial large models of the moment) and its subplot case (can AI personas replace interviews with real people) are both preregistered real experiments. The criteria went into the repo first, and after that only a change log could be appended. Which line died, how it died, and how a pretty number I almost celebrated early got overturned in the interrogation before it was announced, I leave to the main text. Only one thing can be said now. I did not know any of these endings before the runs started, and the timestamps on the criteria will testify to that. That sentence itself needs a discount. When this book started its runs, there was no third-party preregistration. Chapter 12 will teach you to check the timestamp on the criteria before accepting any output, but that check has a hole of its own. A git timestamp can be forged. You only need to change the local clock. So "the timestamps will testify," in front of a reader who really presses, counts as self-attestation and falls short of evidence. To become evidence it needs a third party I do not control to stamp the time, and the common practice on the academic side is a preregistration hosting platform. When writing this book I did not do that from the start. This is a real omission. The remedy and the current status are recorded in the notes of the online case library, the part of this book that keeps updating on the web, with the address in Chapter 16. I did not delete that sentence, and this paragraph stays in the preface, because it demonstrates the book's basic move, **take apart your own chain of evidence first, and write down the holes you find**. **Every number points to its crime scene.** Every accuracy, cost, and confidence interval in the book ends up pointing at specific files and commits in two public repositories (smol-army and persona-panel, under https://github.com/hallieren/research-rewritten/tree/main/code). You do not need to believe me. You can rerun it. One more entry has to go in the preface, because by this book's own discipline, hiding it would be cheating. **AI was used deeply in producing this book.** Chapter drafts, literature checks, code scaffolding, a large share of the steps were executed by AI. That is exactly why I use the verification process in this book every day to protect myself first, and only then recommend it to you. The independent channel check has overturned my claims. The interlock check, which is to say making the same number in different chapters answer to each other, caught number drift in the drafts. In the subplot case's interviews there was also a survey question I misremembered, and only the preregistration discipline caught it. That story is in Chapter 7. A book with AI this deep in it is the first test bed for whether "move at its speed, hold the line of science" holds up. Whether it does, you judge for yourself when you finish. What this book does not cover is also stated here. It is not an encyclopedia of cases by discipline, not a tool manual, and not a job-hunting guide. One more point, and it matters more. It is for people who are not the core reader. Every first-hand case in the book is a computational experiment that can be rerun, which is what my craft covers. If your evidence is wet lab, fieldwork, one-off interviews, or clinical data, you will find the process in this book and not a case of your shape, and forcing one would only produce cliches. The process sharpens how judgment is done right, and it does not care about the shape of the case, so the process still works for you, but **the conversion is your job**. Chapter 3, section 3.6 gives a conversion table with five premises. People for whom all five hold are the core reader described above. No need to memorize the five now. Check them one by one when you get there, and wherever you break one, patch yourself at that spot. There is one corner that is an exception. The output you sign off on came from someone else's hands and you never ran it yourself, which is Premise 3 in Chapter 3, section 3.6. Spot checks and full verification you can still do, and Chapter 12 is written for exactly that situation. How much to sample before it is enough, how to escalate when a check fails. This book has no answer. I labeled them "still exploring" there, and did not talk my way past them. **How to read**, four paths, depending on who you are. - **The impatient** enter through Start Here, the piece right after this preface. Two hours later you will have a number that makes your heart race, and a list that shuts it up. - **People with a real project to use it on right now**, the route is Start Here to pick your question β†’ Chapter 3 (especially the five premises in 3.6, which decide whether each later chapter needs converting) β†’ Chapter 6 (lock in the criteria first) β†’ Chapter 12 (acceptance) β†’ Chapter 11 (failure modes). Those four chapters together are the most compressed part of this book. - **People who mainly accept other people's output**, whose inbox holds more AI reports than experiments they ran themselves, the route is Start Here for the seven-item list only β†’ Chapter 3, section 3.6 β†’ Chapter 11 (failure modes) β†’ Chapter 12 (the two-hour acceptance workflow) β†’ Chapter 10 (the four attack surfaces as an acceptance checklist). - **People deciding whether this book is worth finishing** go straight to Chapter 13. That is the book's master ledger, and the one place where I put my own claims and other people's claims on the same table and label their status, including two rows downgraded by my own rules and two tombstones of my own. If that table does not convince you, the twelve chapters before it probably will not either. The map of the book's chapters is in Chapter 1, section 1.5. One last ugly sentence, the one the whole book keeps repeating. Every claim in this book carries an evidence-status label, verified, still exploring, or falsified, and each one records "what evidence would change its status." This preface included. --- # Start Here Β· A Two-Hour Win !!! info "Chapter companion" πŸ’» [`code/smol-army`](https://github.com/hallieren/research-rewritten/tree/main/code/smol-army/) > This chapter skips the argument and gets results first. Two hours from now you will hold an experiment number that makes your heart race, and a list that calms you back down. The two together are the whole book. This chapter hands you three things. An experiment number you can reproduce in two hours, a seven-item "don't trust it yet" list that shuts it up, and a one-sentence real problem you pick yourself. --- ## Run it first This book chases one question with you. **Can a team of small open-source models tie a single frontier model?** Frontier here means the strongest batch of commercial closed-source models of the moment. The question is worth real money. Companies that believe "yes" are saving money, companies that believe "no" are paying the bill, and the research literature has not settled the fight. We go through the literature in Chapter 4. Today we read no papers at all and run the most naive version, no team, one small model alone against the frontier. The materials list has three items. - A computer that can run Python. The models all go through APIs, no GPU needed. - Two API keys, one for OpenRouter (calls the open-source small model) and one for OpenAI (calls the frontier model). - A budget of a few cents. If you would rather not run code, you do not have to leave. The repository keeps the raw results of my run in `results/demo_naive.jsonl`. Open it and follow the scoring logic along the four steps below. Honestly, reading along does not get you the thrill of a run that works. That thrill comes from a number you produced yourself, and no prose makes up for it. The good news is that the heavy part of this chapter's deliverable is the list that follows, and the list can be read. Readers who follow along can skip straight to the section "19 to 20". From installing the environment from zero to seeing the score, budget two hours, of which the machine's actual working time is a small fraction. This "two hours" is my two hours, and I write this kind of code every day. If the last time you set up a Python environment was under someone else's guidance, double the budget, and accept one thing in advance. The first time you get any unfamiliar toolchain running, getting stuck is normal, and it has nothing to do with how smart you are. Before you run, spend 15 minutes on the check below. > **The 15-minute pre-run check** > > - [ ] OpenRouter account registered, API key in hand > - [ ] OpenAI account registered, API key in hand > - [ ] $1 loaded on each (the demo spends a few cents, but a zero balance stalls on the first problem) > - [ ] Python 3.11+ and uv installed (uv is the package manager this book uses throughout, `uv --version` prints a version number) The item on the check most easily underestimated is the third, and what stalls you is often not the money. Both platforms require a foreign credit card, and in some countries and regions that gate is harder to pass than every technical step after it combined. No card, no entry. That is the real admission condition of this demo, and effort will not get around it. I write it here so you do not run into it yourself at minute 40. If you cannot get past it, you still need not leave. Follow along as the previous paragraph describes, and the list this chapter delivers is yours in full. Two more places in the toolchain where I made an assumption you may not share. First, package management uses uv, not conda or pip. Mixing them produces strange problems, and the repository README guarantees only the uv path. Second, the `export` line that sets environment variables comes from macOS/Linux. On Windows use `set` instead, or `$env:` in PowerShell. Then, four steps. **Step 1, install the environment.** Pull down the book's companion repository smol-army (at ), install uv, and set the two keys as environment variables. The README has a one-line export instruction, copy it as is. **Step 2, get the problems.** The exam is HumanEval+, a code generation benchmark. The plus sign means it added more test cases on top of the original problems, and the problems are the same set. Each problem gives a function description, the model writes the implementation, passing the test cases counts as right, failing counts as wrong, and scoring needs no human opinion. We take the first 20 problems of the dataset. **Step 3, each contestant answers once.** On the small-model side is gpt-oss-20b, a 20B open-source model whose weights anyone can download. The frontier side's model has most likely been replaced by the time you read this. Swap in the frontier of your day. The number will change, and the seven-item list coming up will not. When I ran it I used GPT-5.6-terra. Each problem gets one chance, single-shot, one call only, no retries, no rerun on failure, no ensembling, no merging answers from several models. The core loop is these few lines. ```python for it in items: text, usage = llm.chat([{"role": "user", "content": it["prompt"]}], seed=0) score = tasks.score("code", tasks.extract_code(text), it) ``` **Step 4, count the score.** ```text $ uv run python scripts/demo_naive.py ... gptoss: 19/20 frontier: 20/20 ``` Done. This is the number I really got, and the raw results sit in the repository at `results/demo_naive.jsonl`. If you ran along and your score does not match 19/20, do not rush to doubt yourself. Remember that feeling. Item 2 on the list is there to serve it. ## 19 to 20 ![Naive demo result, gpt-oss-20b 19/20, frontier 20/20, one point short of a tie, but don't trust this number yet](../assets/images/demo-19-20.svg) Pause a second and look at what this number says. GPT-5.6-terra is one of the strongest models money can buy right now, and priced per token it costs tens of times what the small model does. gpt-oss-20b is a 20B model with open weights that anyone can download. Twenty problems, one point apart. If you are the person your manager asked "can we bring inference cost down," this number looks like the answer. **The small model looks like it can tie.** Make one slide, put this 19 to 20 on it, and Wednesday's review will go very well. The first time I saw those two lines of output, my reaction was excitement. That excitement is right. It means you have hit a question worth being serious about. The word "looks," though, is the book's number one enemy. The pair of words is **plausible** and **reliable**, and this book calls them by those names from the first page to the last. Everything in this demo stops at the former. Before sending this number to anyone, I forced myself to write a list. **This conclusion, where I don't trust it yet.** ## Slam the brakes, seven "don't trust it yet" items The seven items below are all real defects of this demo. Not one is a straw man. 1. **Too few problems.** With 20 problems, one problem is 5 percentage points. The gap between 19/20 and 20/20 looks exactly like one swing of luck, and the eye cannot tell them apart. How many problems are enough, and how to compute a gap so it counts, is the statistical debt of Chapter 8. 2. **Each problem answered only once.** A single sample, no repeats. Rerun at another time and does the number change? I do not know, I did not rerun. If you got 18/20 or 20/20, neither is a surprise. The repeat count and the seed should be locked in before the run, Chapter 6. 3. **Cost not accounted for.** "Tie" is a price word. Frontier tokens cost tens of times more, but how much more, and how many cents each side spent, I did not compute. The `usage` field sits right there in the results file, and I did not read it. A "tie" with no price tag means nothing. The cost basis is set in Chapters 5 and 6. 4. **Only one subject tested, and possibly a leaked one.** No task outside code generation was touched. HumanEval is also one of the best-known code benchmarks, and most likely sits in both models' training corpora. Testing memorized problems tests memory, not ability. How to pick tasks and how to check for contamination, Chapter 6. 5. **No control arm.** When a team really gets formed later and the score really rises, does the credit go to "teaming up" or to "spending a few more samples"? A control arm is the group in an experiment set up to rule out one alternative explanation. The group missing here is "same budget, let one model answer several times and take the majority." Without it, the two explanations cannot be told apart in the numbers. How to design the control arm, Chapter 6. 6. **Nobody reviewed the scorer.** I wrote the scoring code in passing. It says 19, so it is 19? One problem misscored and the whole story gets told differently. The scorer is code too, and code has bugs, Chapter 7. 7. **The first 20 problems are laziness, not sampling.** Who says the first 20 problems of a dataset represent the whole dataset? The hand of whoever arranged them hides in the problem order. The sampling rule should be locked in before the run, Chapter 6 again. Count them and you will see that most of these seven debts are booked to the same account. **The rules were not locked in before the experiment ran.** That is no coincidence. Every chapter that follows repays a few of them, and this list is the roadmap of the whole book. ## This is not a death sentence for the demo To be clear. After listing the seven items I did not delete this demo, and I do not regret running it. **An experiment that takes one coffee's worth of time should be run.** It cost 4 cents in total. The arithmetic takes the token counts from the `usage` field in the results file and multiplies by the list prices in the repository's config/prices.toml, checked on 2026-07-25. The frontier side 4.0 cents, the small-model side 0.14 cents. Roughly 30 times apart. This is a simulated account at list price, not a bill, and Chapter 10 will find its deviation on the bill basis. Yes, that is the debt item 3 on the list owes. Those 4 cents bought three things. A question worth being serious about. First-hand feel, what the two models' APIs look like and what HumanEval+ problems look like. And a code starting point that every serious experiment later reuses. I did only one extra thing. In the repository I labeled it "not preregistered." Preregistration means locking in, before the run, the claim to be tested and the standard for right and wrong. This demo locked in none of that, so this number may only inspire questions, never support conclusions. There is only one wrong way to treat the demo, which is to run it and go make slides. Between "looks like a tie" and "dare to sign off at the review" lies a whole process. The question has to be sharpened into a falsifiable hypothesis, a judgment the results can overturn. The criteria have to be locked in first, meaning what counts as right and wrong is decided before the run. The experiment has to run reproducibly. The ghosts in the results have to be dragged out one by one. The conclusion has to be written as a deliverable that survives being taken apart. And at the end someone has to beat it up hard. That process is this book. This demo is the first act of the spine case. Every chapter from here on, I push it one gate further, and every item on the seven-item list gets settled along the way, one by one. How it ends, right now I know no more than you. ## One thing while you are here, pick your real problem Watching someone else run experiments will not teach you to run experiments. While you are here, pick a problem of your own. There are only three criteria. If the answer is wrong, something real gets hurt. All you have right now is a "plausible" answer. Within the next two or three months you really have to take a position on it. What it looks like in a few different trades. - **Technology selection**, "Should the retrieval layer be swapped from the current design to a vector database?" - **Market research**, "Will target users really pay $20 a month for this product?" - **Policy argument**, "Does the new 'three days a week back in the office' rule have evidence behind it?" - **Academic topic**, "That line in my thesis proposal, 'A regulates B through some pathway,' how far does the existing evidence carry it?" Write your problem as one sentence and stick it to the edge of your screen. From Chapter 4 on, every chapter ends with a "Swap in your project" block that translates the step the spine case just took into a concrete action on your project, and this sentence is what it acts on. When you finish this book, what you hand over is a real project that has been run through once in full, not a reading reflection. **Want an agent to run it with you?** Paste the block below into Claude Code, Codex, or any coding agent: ```text Clone https://github.com/hallieren/research-rewritten, go into code/smol-army, run uv sync --extra dev and uv run pytest, then run uv run python -m smol_army.run --mock to walk the whole chain offline, and show me the output verbatim. Then open results/demo_naive.jsonl and explain to me, problem by problem, how the first 3 problems were scored, following the four-step scoring logic in Start Here. Do not total the score for me, I want to count the 19 to 20 number myself. If I have OPENROUTER_API_KEY and OPENAI_API_KEY configured, wait until I say go before running uv run python scripts/demo_naive.py, it will cost a few cents. The seven-item "don't trust it yet" list is mine to write, do not write a single item for me. If any command errors, stop and show me the output. ``` ## The unfair advantage you now hold An experiment number reproduced in two hours, a seven-item list that shuts the number up, and the real problem you just wrote down, are all in your hands. This combination is the first place the book's whole attitude touches the ground. **Move at its speed, hold the line of science.** --- # Chapter 1 Β· After Coding, Research !!! info "Chapter companion" πŸ“‹ [Chapter 1 templates](../appendices/ch01-templates.md) Β· πŸ—‚ [Template index](../appendices/template-index.md) > An AI research report can dazzle you within ten minutes and embarrass you after you forward it. Those are two faces of the same thing. This chapter states its true shape, and states what this book plans to do about it. (The AI autonomy ladder, the ruler that measures how deep AI can go for you inside one step of the work, is not built until Chapter 3, and the ladder lines at the head of each chapter begin in Chapter 4.) This chapter hands you three things. A pair of concepts that run through the whole book, plausible and reliable. An operational definition, research is a standard, not a profession. And a map of the whole book, with a self-test that helps you decide whether to keep reading. --- ## 1.1 The last question before you forward it Friday, four in the afternoon. Your manager messages you. Next Wednesday's review has to settle one thing, whether to self-host the inference service or buy a managed one. He knows you have no time for two weeks of research, so he adds, "A directional call is enough for now." You open a deep research tool, type in the question, and get up for a glass of water. When you come back the report is on the screen. Nine pages, four sections, two comparison tables, more than twenty cited links, a clear conclusion, a restrained tone. Honestly, it looks better than the survey you spent three days writing by hand last quarter. You move the mouse to the forward button and stop, maybe because a "too smooth" instinct kicks in. You take a second look at one number. The managed option's total operating cost is 40% lower than self-hosting. Once that number enters the meeting room, the review's conclusion follows it. You ask yourself one question. Where did this number come from? The report gives two citations. The first is a technical blog post. It exists, but it compares a completely different set of workloads, and the 40% comes from dividing two unrelated numbers. The second is a "2024 industry benchmark report." You search for ten minutes. It does not exist. Most of the report is right. The framework is useful, and several citations are real and relevant. But a report that is 80% correct, where you cannot locate the 20% that is wrong, may have negative decision value. Yet the confidence it hands you is 100%. This has happened for real. In 2023 a New York lawyer filed court papers drafted with help from ChatGPT. Several of the cases it cited were invented, and the judge checked and fined him (Mata v. Avianca, Southern District of New York, 2023; a five-thousand-dollar fine). Where he fell was in taking "looks like a precedent" for "is a precedent," and not asking that one question before filing. Using AI itself was not what the court punished. You may want to say, fine, I will check every citation in the report. You can, and the direction is right. Notice what you just agreed to. Verification takes work, and that work never appeared in the "report in thirty minutes" pitch. The cost of production collapsed. The cost of verification did not. This asymmetry is a structure the book keeps coming back to. Coding hit the same wall two or three years ago, and Chapter 2 covers it in detail. Two common reactions to this experience, both wrong. The first is to ban it. AI is unreliable, go back to writing by hand. That throws away a machine that really can compress three days of research into thirty minutes, and it does not stop your colleagues or your competitors from using it. The second is to forward it as is, since it is mostly right. That puts your name behind a machine that answers for none of its conclusions. This book offers a third way. **Move at its speed, hold the line of science.** Treat plausible as raw material, and work it into reliable with a method you can run. ## 1.2 Two worlds, the one in the headlines and the one on your desk First, take apart a common confusion. Right now the phrase "AI does research" covers two worlds that barely touch. **The first world is in the headlines.** AlphaFold predicted the structures of nearly every known protein, and half of the 2024 Nobel Prize in Chemistry went to Demis Hassabis and John Jumper behind it. The same family holds specialized models for materials discovery, weather forecasting, and gene sequences. They are real scientific achievements, and what they share is that each was built for one specific scientific question, eats specialized data, and produces a specialized answer. You cannot use AlphaFold to answer Friday afternoon's self-host-or-managed question. It has no interface to your work. **The second world is on your desk.** General-purpose models and agents, model programs that call tools on their own and work through many steps, are entering every step of truth-seeking work. Scanning literature, reading a codebase, proposing hypotheses, writing analysis scripts, running computations, interpreting results, drafting reports, finding fault with their own work. Each step looks too ordinary to matter. Added up, what gets rewritten is **the workflow that produces knowledge itself**. This book is only about the second world. It has an interface to you, it is far larger in volume than Nobel-grade discovery, and nobody is in charge of it. The only inspector of the report on your desk is you. From here on the first world appears only as background. The distinction earns its own section because confusing the two produces two costly errors at once. One is **overrating by borrowed halo**. If AI can win a Nobel, how could its report be wrong? AlphaFold's reliability was checked point by point against decades of accumulated experimental structure data. The report on your desk went through no corresponding check. No credibility passes between the two. The other is **dismissing by borrowed crash**. If AI even fabricates citations, this whole wave is a bubble. A fabricated citation is one specific, preventable failure mode in the second world (Chapter 11 files it), and using it to dismiss the whole shift is like concluding from early cars breaking down that the combustion engine had no future. One more thing needs saying plainly. In this book, "AI does research" means AI enters the steps of research. It cannot take the researcher's seat. Asking, deciding, signing, answering for it, those seats are still yours. Chapter 3 gives a four-level ladder for how high AI can stand in each step, and spells it out level by level. ## 1.3 Plausible is not reliable Now the book's core concept, head on. The pair of words is **plausible** and **reliable**. This book calls them by those names from the first page to the last. Credit first, honestly. Frontier labs, the handful of labs that train the strongest general-purpose models, built deep research modes that let a model search, read, and synthesize on its own over many rounds, then hand back a cited report. That is a real jump in capability. The breadth is real. In tens of minutes it scans sources you could not cover in days. The quality of the first-pass synthesis holds up too. The structure, the prose, the arrangement of different sources, often beat what a tired human produces before a deadline. The flaw is in the optimization target. These systems are trained to produce answers people are satisfied with. Fluent, confident, well structured, comprehensive. Those properties get rewarded in training because human raters like them. And to the model, "telling the truth" and "sounding like the truth" come from the same generative skill. The consequence is a structural decoupling. **A report's persuasiveness and a report's correctness are no longer correlated.** A fabricated citation and a real one are identical on the surface of the text. Errors carry no marker, and they do not cluster where you would think to be suspicious. What you saw in section 1.1 is exactly this. The tone of the 40% sentence matched the tone of every true sentence in the report. Credit and flaw are both on the table. Whether plausible is enough depends on one variable. **What happens if the answer is wrong.** In plenty of situations it is entirely enough. Planning a trip, getting the gist of a new topic, feeding a brainstorm, writing a stretch of code that a compiler and tests will catch. In these situations the cost of error is low, or the error exposes itself quickly, or a cheap automatic correction catches it. Code that fails to run and errors on the spot is one example. Paying for verification there is waste. This book carries no hostility toward deep research mode. I use it every week, and it is a competent starting point in this book's workflow. The trouble starts only when the starting point is taken for the end. In another class of situations, plausible is fatal. A review board makes an architecture decision on your numbers. A funding committee decides from your literature review whether to fund the direction. A client decides whether to pay based on your due diligence report. These situations share one thing. **If the answer is wrong, something real gets hurt.** That is this book's operational definition of research. **Research is a standard, not a profession.** If a wrong answer from you would really hurt something, you are doing research and you need reliable, whatever your title. A PhD student, an engineer doing technology selection, an analyst doing due diligence, a parent checking treatment options for their child, are the same kind of person under this definition. Of that kind, this book writes out the full process for only one, the engineer doing technology selection, whose evidence can be rerun and who signs off in person. The others read the same process, but the shape of the case has to be converted, and Chapter 3, section 3.6 gives the conversion table. Conversely, a casual searcher who is satisfied once "the answer sounds fine" is not doing research, however advanced the tool, and this book cannot help them. Which side you are on, the self-test in section 1.5 will help you judge honestly. The definition also settles a familiar quarrel in passing. Someone will say, I do not publish papers, why talk about research method. Wrong. Method is discipline forced out by the cost of error, and it never belonged to academia alone. Preregistration, peer review, reproducibility, which is to say writing the hypothesis down and filing it before the run, handing it to peers to find fault, letting others redo it from your description. These academic mechanisms all do one thing. Where a wrong answer costs a lot, humans were forced to invent process to fight their own credulity. Your review meeting, your due diligence, your selection report cost no less when wrong than a paper does, and the same discipline holds for you. It used to cost too much to run, so only academia could afford it. AI brought the cost of running it down, and for the first time this discipline is affordable for everyone. So what is reliable? It is more practical than "absolutely correct," which science never promises. Its operational definition is **traceable, checkable, and any error can be located**. Every key claim points to a source, and the source exists and really says that. Every key number can state its basis and its method, which is to say what scope and what calculation produced it. Where something is wrong, you can find which step introduced it. In one sentence, this is a report you dare to sign, and dare to let someone who knows the field take apart in front of you. **Reliable is a property of the process, not of the text.** The same report, checked line by line or not checked at all, reads identically and differs enormously in credibility. That is why the only upgrade path runs through method. A smarter model will not buy it. Will the gap close on its own once models get strong enough? My judgment is no, on two levels. The first is the verified present. As of this writing, the strongest deep research products still produce fabricated citations and numbers whose basis has drifted, meaning a number whose statistical scope or method changed without the report saying so. Public evaluations back this up, not just my impression from use. A full-trajectory hallucination evaluation of six mainstream products (OpenAI, Gemini, Perplexity, Qwen, Grok, Salesforce) found that no deep research agent (DRA) could stay free of fabrication across the whole research trajectory, and two were described as "confident fabricators." The original conclusion reads "no single DRA achieves robust performance across the full trajectory." See "Why Your Deep Research Agent Fails? On Hallucination Evaluation in Full Research Trajectory," 2026, arXiv:2601.22984. The second is a prediction anchored in history. Every time a tool made "production" cheap, spreadsheets made modeling cheap, statistical software made computing a statistic cheap, "verification," checking whether the output is right, never got cheap along with it. It became the new bottleneck instead. Chapter 2 lays out this pattern with its evidence. The prediction is falsifiable, meaning future facts can overturn it. If a system someday keeps key claims at zero fabricated citations and every number's basis traceable, over a long run and with no human spot checks, this cornerstone of the book should be overturned. I am willing to go that far because the whole book speaks to one standard, **verified, still exploring, and falsified get separate labels**. ## 1.4 Why now If this tension, production sped up and verification did not keep pace, looks familiar, that is because it just played out in full next door. AI coding ran two or three years ahead of AI research. The trajectory compresses into one sentence. From code completion, to conversational coding, to coding agents that execute multi-step tasks on their own. Chapter 2, section 2.2 lists the products and years of the three stages. More valuable is the road programmers as a group traveled in those two or three years. First they laughed, a toy. Then they were dazzled, it writes faster than I do. Then it crashed, it confidently called a function that does not exist. Last came calibration. Learning which work to hand off, which to watch yourself, and which mechanisms (tests, review, tiered authorization) turn a collaborator you cannot fully trust into productivity. Today most developers' daily workflow already has a place for AI. In the Stack Overflow 2025 Developer Survey, 84% use or plan to use AI tools, and 51% of professional developers use them daily. In the same survey, only a third trust the output. Using it but not trusting it, that group has reached the calibration step. That expensive curve accumulated a batch of lessons that transfer straight into research. Chapter 2 sorts them, along with earlier precedents, into one map. Here only one claim gets planted. **The shift in front of you has a precedent, so you need not guess from zero.** Beyond a mature precedent, two conditions belong to "now" alone. First, agent capability crossed the threshold of usefulness. A model can take your question and work for tens of minutes to hours without you feeding it sentence by sentence. Steps of the research workflow that were counted in days can, for the first time, be squeezed into hours. Second, cost crossed the threshold of wide access. Capability at this level once belonged only to institutions that could pay large bills. Now, by Epoch AI's March 2025 tracking of six benchmarks, the inference price to reach a given score dropped by a factor of 9 to 900 each year over the past three years. The spread is that wide because different benchmarks lose price at very different speeds, and Epoch itself warns that the steepest drop came in the most recent year and may not persist. The capability line and the cost line crossed recently. That is why this book was written in 2026, not 2023. Expectations need managing too. This book will not tell you which model or which product to use. Version-level comparisons expire in three months and pull your attention to the wrong layer. Workflows and criteria outlive tool lists by a wide margin. The state of the tools is left to the book's companion online case library to track. There is one harder reason, and it has nothing to do with whether you use AI. **You can choose not to do research with AI. You cannot choose not to consume research that others did with AI.** Your inbox already holds reports, reviews, and due diligence with deep AI involvement, and most do not say so. They all look right. Telling whether a document "searched fast" or "did research," whether it is plausible or reliable, is moving from a bonus skill to basic self-defense. The earlier you build it, the more it is worth. It compounds. ## 1.5 How to use this book Cards on the table first. I am an engineer. My daily work is building agents and doing evaluation, the tests that set questions for models and score them, and I live inside the question "do I dare believe this conclusion." The methods in this book are the ones I earn my living with, including the times they fail. The reader it writes the full process for is that same person, a software or AI engineer asked at work to sign off on a technical judgment, where the evidence is an experiment anyone can rerun. Such a person has two kinds of days. The producer's day, when your manager says "look into this" and you have to run it into a conclusion you would sign. The seven chapters of Part II are arranged around that day. And the acceptor's day, when a director forwards a report with deep AI involvement and wants to know by three in the afternoon whether it can be trusted. Chapters 11 and 12 are written for that day. Most people alternate between the two. The book has four parts, each with one job. **Part I (Chapters 1 to 3) is the frame.** The chapter you are reading sets the stance. Chapter 2 answers "has this kind of change happened before," distilling AI coding and several earlier tool revolutions into a transfer map, which lays the road other tool revolutions traveled alongside the new road of AI research. Every forward-looking judgment in the book hangs on it. Chapter 3 breaks "truth-seeking" into a seven-step workflow panorama and raises the four-level AI autonomy ladder, **tool β†’ assistant β†’ collaborator β†’ autonomous researcher**, so you can place yourself and your project. **Part II (Chapters 4 to 10) is the main line, the operational meat.** One chapter per step of the seven-step workflow. Master the field, questions and hypotheses, test design, execution, read and catch errors, delivery, red team. Red team means having someone attack your conclusion on purpose, the same move as red-teaming a model in AI safety. Every chapter gives templates and checklists, specific enough that you can close the book and do one concrete thing on your own project. **Part III (Chapters 11 to 13) is the cutting edge.** Reliable is cashed in here. File the failure modes unique to research, review AI research output with a verification workflow that works like code review, and hand over an honest map of the present, labeled "verified / still exploring / falsified." **Part IV (Chapters 14 to 16) looks at roles.** Once the steps are rewritten, where the center of the researcher's craft moves, how teams collaborate, and how this book itself stays current in a field that changes every month. Two real cases run through the book. **The spine case** advances chapter by chapter from Part II on. It asks whether a team of small open-source models, voting, debating, dividing the work, can tie a single frontier model on a cost-matched basis, meaning both sides spend about the same before scores are compared. Consensus is unsettled, and papers from both camps fight head on in the literature. I run it myself from doubt to conclusion, literature, hypotheses, experiments, delivery, taking criticism, with every pit written up as it happened. **The subplot case** appears in key chapters as a counterpoint and asks, can synthetic data stand in for real data? Use persona-bearing models ("a 25-year-old mom with one child who works at X") in place of real people for user interviews. The answers read like a real person, and it saves time and money, but can it be trusted? What counts as ground truth, and where do the real answers used for checking come from? Engineers get asked this class of question every week. Can synthetic users replace real user testing, can synthetic samples stand in for an eval set. This line asks the same thing as the spine case. Swap the expensive one for a cheap substitute, and do you still dare trust the step you saved? It also holds a suspense that is not resolved until Part III. ![The narrative arc of the spine case, doubt β†’ verify β†’ contribute β†’ deliver and take the hits](../assets/images/spine-arc.svg) There is also a Start Here at the very front of the book. It skips the argument and walks you through building something that runs in two hours, listing on the spot "where I don't trust it yet." That is the spine case's first act. At the same time you pick one real problem of your own. Every chapter in Part II ends with a "Swap in your project" block that translates the step the spine case just took into an action on your own project. When you finish the book, what you deliver is a real project you ran through the whole workflow with your own hands. A reading reflection does not count. Finally, four questions to help you judge honestly whether to keep reading. The first three decide whether this book is for you. The fourth decides how you read it. 1. If your answer is wrong, does someone, some money, or some decision really get hurt? Yes, this book is written for you. No, what you need is a good search tool, and you can put this book back. 2. Can you accept the premise that "verification takes work"? This book does not sell "reliable in one click." It promises to make the work of verification affordable and reusable. It cannot bring it to zero. 3. Are you willing to work in an unsettled field under a "still exploring" label? This book would rather say "nobody knows here" than fake certainty. If you want a book that is categorical on every page, there are plenty. Do not pick this one. 4. Is your evidence something that can be rerun, code, data, a ledger that can be recomputed? Yes, every step in the book runs on your project as is. No, wet lab, interviews, clinical data, the steps still apply, but the shape of the case has to be converted per Chapter 3, section 3.6, and that section decides how you read Part II. The full version of this self-test (with scoring) is in this chapter's templates. It takes a minute. **One thing you can do right now.** Find the most recent AI report you received or generated this week, pick the one number in it that most affects a decision, and spend ten minutes finding its source. There are only three outcomes. Found, and the basis matches. Found, but the basis differs. Not found. Write the result down. The fate of this one number is the trust you should place in the whole report right now. **Want an agent to run it with you?** Paste this to your AI assistant or coding agent: ```text Help me do the task at the end of Chapter 1. I will paste you a recent AI report. You do one thing only, list every number in it that would affect a decision, and after each number write the source the report claims for it, or "none" if it gives no source. I pick which number matters most, and I check the source myself. You may not search for me, and you may not tell me "this number is probably right." When I come back with the result, you record it as one of three outcomes, found with matching basis, found with a different basis, not found. At the end I answer the seven-question self-test in the templates file ch01-templates.md myself. You only total the first six questions by the scoring rules. Question seven is not scored. If any command errors, stop and show me the output. ``` ## 1.6 The unfair advantage you now hold When the next AI report lands in front of you, whether you generated it or someone forwarded it, you can first tell which world it belongs to, searched fast or did research. Then test it on the spot with one question, where did this number come from? The test has only two answers, plausible or reliable. Most people's mouse, at that moment, is still hovering over the forward button. --- # Chapter 2 Β· A Map Stolen from Paradigm Shifts !!! info "Chapter companion" πŸ“‹ [Chapter 2 templates](../appendices/ch02-templates.md) Β· πŸ—‚ [Template index](../appendices/template-index.md) > The change you are watching has run the same script at least five times, and the audience's notes are still around. This chapter's job is to steal those notes back and bind them into a map that every later chapter cites. The ladder does not get raised until Chapter 3. This chapter is about seeing which wall to lean it against. This chapter hands you three things. Five lessons AI coding already paid tuition for, an eleven-row transfer map, one row per pattern the precedents walked, set against research, and five transfer failure conditions. --- ## 2.1 The panic of 1979 A headline this morning. Some AI system "wrote a paper on its own," or "did a PhD student's week of literature review in ten minutes." You cannot tell whether to be excited or to sneer, and you certainly cannot tell whether it touches the report you owe next month. This chapter hands you a map. Put that kind of news on it and you see which old pattern it is, which rerun of that pattern, and what half-sentence it left out. The map has to be drawn starting from an earlier panic. In October 1979 a program called VisiCalc went on sale for the Apple II at around a hundred dollars. Its inventor, Dan Bricklin, was still a student at Harvard Business School. By his own account, half the idea came from his frustration. One wrong number in a case assignment voided every number computed after it. The other half came from a scene in class. A professor worked a financial model on the blackboard, and every change meant rewriting a string of cells by hand. VisiCalc made "recompute it" cost nothing. Bricklin later described a reaction that kept recurring. Programmers watched the demo and said "nice, so what." People who had actually built the reports trembled, because what took seconds on screen was a full week of their work, and when they were done trembling they pulled out a card on the spot to buy it. The panic followed. Recomputing tables was the daily core of the accounting craft. The core got automated, so what is left of the profession? Forty-some years later the accounts can be settled. On the US Bureau of Labor Statistics job classification, bookkeeping and accounting clerk jobs are down by about four hundred thousand since 1980, while accountant and auditor jobs are up by about six hundred thousand over the same period (Planet Money, citing US Bureau of Labor Statistics (BLS) data). What got wiped out was the mechanical step of "recompute it." What did not get wiped out was judgment. How this sheet should be built, whether the numbers are right, what they mean. Demand for judgment grew, because modeling got cheap and everyone started modeling. The story has another half that the inspirational version usually skips. In the same era that democratized modeling, a whole class of unprecedented errors was born in bulk. A formula pointing one cell off, a copy that skips a row, a sum where an average belonged. They hide behind a polished interface without a sound. In 2010 Reinhart and Rogoff published "Growth in a Time of Debt," and policymakers in the US, the UK, and the EU cited it to justify austerity. In 2013 the graduate student Herndon, working with Ash and Pollin, reproduced it and found that the range of the Excel sum formula left out the rows holding five countries. After the correction the negative growth turned positive, and the average growth rate of high-debt countries went from -0.1% to +2.2%. That is the full shape of a paradigm shift, the kind of change where the old tools and the old standards get replaced wholesale. **The old bottleneck disappears, judgment appreciates, a new class of error is born, the panic misses, and it misses in a way nobody expected.** You will see this shape five more times. One of them just happened to coding. The other four are earlier precedents, and counted together with the spreadsheet they come to five below. This chapter distills those patterns into a **transfer map** the whole book keeps citing. ## 2.2 The main precedent, the tuition AI coding paid first The precedent closest to research, and the most isomorphic in mechanism, is AI coding. It started two or three years earlier, it ran messy enough, and its tuition bill is itemized enough. You have probably walked this curve yourself. Completion (the 2021 Copilot preview, AI finishing half a line behind your cursor), conversation (pasting whole blocks of code into a chat box after ChatGPT at the end of 2022), agent (from 2024 on, AI editing files, running tests, and opening PRs while you fall back to acceptance). Every new stage renegotiates the division of labor between you and the AI. So the five lessons below do not retell the coding story. Each keeps one sentence you already know and spends its length on the research counterpart, and on where that counterpart takes a discount. **Lesson one, the cost of production collapsed, the cost of review did not, and the bottleneck moved to review.** Your team is probably stuck in the review queue already. This is a permanent change in the cost structure. When production is nearly free, your throughput ceiling equals your review bandwidth. The research counterpart is already around you. First drafts, literature summaries, and analysis scripts all got cheap. "Can this conclusion be trusted" got no cheaper at all. Part III of this book is built entirely on this lesson. The discount is in section 2.6, failure condition one. Research review has no compiler underneath it, which makes it more expensive than code review. **Lesson two, the demo dazzles, the merge is a disaster.** In 2024 a demo video of a "fully autonomous AI software engineer" caused a sensation. A month later the engineer Carl Brown (YouTube channel Internet of Bugs, April 2024) went through that demo frame by frame, found several exaggerations, and showed that the Upwork task in it was never finished to the client's requirements. The programmer community ended up counting only production-grade numbers, merge rate, rework rate, incident rate. The research counterpart is the same. How a deep research report performs in a demo and how it performs once it enters a decision chain are two different things. You can judge it only on the second, whether the citations hold up under checking, whether the numbers can be traced. On the research side this one bites harder. Section 2.6, failure condition four, says a research question worth doing is by definition not in the corpus, and demos are almost always recorded on problems that are already solved. **Lesson three, self-perception is not trustworthy, calibration has to come from measurement.** You may never have measured this one yourself, so the numbers stay in full. In 2025 METR, a research group that evaluates AI capability, ran a randomized controlled trial. It had 16 open-source developers with an average of five years on their projects complete 246 tasks on real projects they knew well, with each task randomly allowed or not allowed to use AI. Before starting, they predicted AI would make them 24% faster. After finishing, they still rated themselves 20% faster. The measurement came out about 19% **slower** on average (the confidence interval is wide, +2% to +39%, so the real slowdown probably falls in that range and is not nailed down). Prediction beforehand, feeling afterward, measurement itself, three levels of contrast with the first two wrong in the same direction. Fluent interaction systematically manufactures the illusion of speed and of mastery, and veterans fall for it too. For research this is a doubled warning, because objective acceptance is harder for research output than for code. Between "I feel like AI helped me master this field" and actually mastering it sits a whole process (Chapter 4 expands). **Lesson four, trust calibration moves from two poles to graded delegation.** The trust-it-all camp paid in incidents, the trust-nothing camp paid in speed, and what survived was delegation graded by the risk and the verifiability of the task. Let boilerplate go, watch the core logic, keep it away from architecture decisions. The version you can act on is a question. For this step, for this task, what level do I delegate to? The autonomy ladder that Chapter 3 raises is this lesson institutionalized for research. The discount is in what the grading rests on. In coding the grading rests on verifiability, and in research most steps have low verifiability to begin with, so the whole ladder sits a notch lower. **Lesson five, verification infrastructure sets the radius of letting go.** A team with solid test coverage dares to let an agent change core modules on its own. A team with no tests has to watch even an edit to a comment, and no amount of model intelligence moves that ceiling. Transferred to research, this points to an uncomfortable inference. If you want AI to climb high on autonomy inside your research, you first have to write "what counts as right and what counts as wrong" into executable criteria, rules you can follow to rule on right and wrong, and that is one of the hardest parts of research itself. Part II runs into it again and again. ## 2.3 Five earlier rehearsals AI coding is only the most recent precedent. Turn back further and there are at least five full performances of "a tool rewriting a craft." One main lesson flashed per case, with the weight spread unevenly on purpose. **Spreadsheets (1979), the same door that democratized capability also let in a whole class of new errors.** Both sides were covered in section 2.1, so here is one addition, a sense of the scale. A coarse audit method applied to 367 spreadsheets in corporate use found 24% contained errors. A stricter audit method put the share containing errors above 86% (Panko, 2000). The gap between the two numbers comes from how tight the audit was, not from a trend over time. Errors did not stop spreadsheets from taking over the world, and "the spreadsheet is right" never became the default. "The AI output is right" works the same way. Statistical software in the 1970s-80s turned a significance test from hours of hand computation into one click. The bar to operate dropped, and the bar to abuse dropped with it. Serious people tested their hypotheses faster, and opportunists found their way to p<0.05 faster too. In 2011 a famous paper demonstrated the power of those degrees of freedom on the spot, degrees of freedom meaning the room to switch analysis methods however you like. Using only flexibility that stayed inside the rules, it "proved" that listening to the Beatles' "When I'm Sixty-Four" made people almost a year and a half younger, p=.040, statistically significant (Simmons et al.). A few years later a large replication effort took 100 studies from three leading psychology journals and reproduced statistical significance in only 36%, against 97% in the original papers (Open Science Collaboration, 2015). What got industrialized was "looking rigorous," not fraud. Chapter 11 takes this warning head on. AI pushes the unit price of "looking rigorous" lower still. From 1998 on, search engines devalued "remembering a fact" and raised the value of "knowing what to look up and judging whether what you found can be trusted." "Found" is not "understood." That boundary applies unchanged to the new version. Swap "found" for "AI can answer it," and swap "understood" for "you have mastered it." **Open-source collaboration (from the 1990s), output from strangers gets trusted through mechanism, not through goodwill.** Code written by tens of thousands of people who never met each other runs inside your bank's systems, and it holds up because of a whole trust machine. Review, continuous integration, version control, a maintainer hierarchy, a traceable origin for every changed line. The famous line "given enough eyeballs, all bugs are shallow" holds only while the mechanism is running. When Heartbleed broke in 2014, everyone saw that the crypto library the whole internet depended on had a core team of about four volunteers for years, only one of them full time. The 2024 xz backdoor was nastier. It went around the code and attacked the chain of trust itself, with the attacker spending more than two years cultivating a trusted maintainer identity. Trust can be engineered, and the trust mechanism itself becomes an attack surface. Chapter 12 builds the answer to "how does AI research output get trusted" out of the same thinking. After AutoCAD arrived in 1982, the drafter's occupation shrank and the designer's did not. The US Bureau of Labor Statistics Occupational Outlook Handbook lists CAD as the main cause of the decline in drafter jobs, while design jobs grew over the same period. What disappeared was the transcription work of copying design intent onto the drawing. The people who decide what to draw got a faster iteration loop instead. This is the most informative historical answer to "will researchers be replaced." Do not argue yet, split the account first. In your daily work, which parts are transcription and which are decisions. Chapter 14 expands. ## 2.4 Three books worth stealing from Some of the people who lived through a paradigm shift wrote the lessons into books. Three are worth stealing from most, one thing from each. The most famous law in The Mythical Man-Month (Brooks, 1975) says that adding people to a late project only makes it later, because communication cost grows on the order of the square of the head count. Agents still have to agree on a basis, merge conflicts, and check each other, and coordination cost does not vanish because they draw no salary. "Does adding agents actually add speed" has to be marked "still exploring" on the transfer map, and the small-model army, the team of small open-source models in the spine case, tests exactly one special case of it head on. "No Silver Bullet" adds another cut. Tools only kill "accidental complexity," the part the tools and the process bring, and cannot touch "essential complexity," the part the problem itself brings. The essential complexity of research is roughly "deciding which question is worth asking and what counts as evidence," and whether AI can touch it is the undercurrent running through the whole book. The Structure of Scientific Revolutions (Kuhn, 1962) gives no operating advice. It gives you a pair of glasses. The chaos of a paradigm shift, two generations who cannot read each other, old and new standards fighting, all of that is normal, not a malfunction. Put the glasses on and look at the mess of "AI research" in 2026, demos everywhere, standards missing, opinion at two poles, and what you read is the ordinary turbulence of a shift. It also warns you that a regularity seen during turbulence cannot be taken as how things will stay. The Pragmatic Programmer (Hunt & Thomas, 1999) has two engineering instincts worth stealing. One is the **tracer bullet**, get a minimal end-to-end path working first and thicken it step by step, which is the road the rough two-hour demo in Start Here took. The other is **orthogonality**, parts that do not tangle with each other, so an error in one place does not infect another. Transferred to AI research, that means setting "generation" and "verification" as two independent stages, and never letting the same AI both produce the conclusion and judge it. Part II uses this over and over. ## 2.5 The transfer map Now distill the precedents into a table. The status column uses the book's three labels, and what it labels is the research side. Verified means enough evidence is already observable in research settings. Still exploring means it makes sense mechanically and the research-side evidence is undecided. Falsified means the research side already has counterevidence. No row currently dares to be marked falsified, and that by itself says the field is young. Read the first three rows first, they fit the work in your hands most closely, and come back to the rest when a concrete situation calls for them. | # | Pattern | How it happened in the precedent | The counterpart in research | Status | |---|---|---|---|---| | 1 | Production cost collapses, review cost does not, the bottleneck moves to review | AI writes code an order of magnitude cheaper, team throughput jams in the review queue | First drafts, summaries, analysis scripts nearly free; "can it be trusted" got no cheaper, verification becomes the throughput ceiling | Verified | | 2 | The demo dazzles, production is a disaster | An autonomous coding agent demo causes a sensation, production-grade numbers (merge rate, rework rate) expose it | A deep research report dazzles in the demo, then gets exposed on citation checking and number tracing once it enters a decision chain | Verified | | 3 | Self-perception is not trustworthy, calibration comes from measurement | Developers rated themselves about 20% faster and measured about 19% slower (METR 2025 controlled trial) | Fluent question and answer manufactures a sense of having "already mastered it"; that sense of mastery decouples from mastery itself and needs an outside check | Still exploring (verified on the coding side) | | 4 | Trust moves from two poles to graded delegation | Both the trust-it-all and the trust-nothing camps lost money, delegation graded by risk and verifiability won | The autonomy ladder, delegation level set on "step Γ— task" (Chapter 3) | Still exploring | | 5 | Verification infrastructure sets the radius of letting go | Only teams with good test coverage dare let an agent into core modules | Criteria locked in first, evals built first, only then is letting go on the table (Part II, III) | Still exploring | | 6 | Lowering the bar to operate = lowering the bar to abuse | Statistical software turned significance into one click, p-hacking (switching methods until it comes out significant) got industrialized | AI turns "looking rigorous" into one generation; The Lancet audited 2.5 million biomedical papers in May 2026. The share of papers carrying fabricated citations rose about 12 times in three years, to about 1 in 277 (see the table note) | Verified | | 7 | New capability manufactures new error types in bulk | Spreadsheet formula errors, invisible, in bulk, hidden behind a polished interface | Hallucination, fabricated citations, spurious significance, equally invisible, in bulk, fluent (Chapter 11) | Verified | | 8 | "Findable" reshapes "known," judgment does not die | Search engines devalued memory and raised the value of "what to look up and what to trust" | "You can ask it" is not "you have mastered it"; AI being able to answer is not you having judgment | Verified | | 9 | Output from strangers gets trusted through mechanism, not goodwill | Review, CI, and traceable origin let strangers' code into production; when the mechanism fails you get Heartbleed | A traceable, checkable workflow for AI output (Chapter 12); the mechanism has not settled yet | Still exploring | | 10 | Adding people does not add speed, communication cost grows as the square | Brooks's law, adding people to a late project makes it later | Adding agents does not necessarily add speed; the literature has started to run the numbers, but nobody has done a full decision-grade comparison with real cost accounting, and the small-model army tests one special case of it head on | Still exploring | | 11 | What gets eaten is the mechanical part of the craft | Drafters disappeared, designers did not, and iteration got faster instead | Transcription work (formatting, boilerplate literature reviews, copying out) gets eaten; "what to ask and what to trust" appreciates (Chapter 14) | Still exploring | Table note, row 6 cites Topaz et al., "Fabricated citations: an audit across 2Β·5 million biomedical papers," *The Lancet*, 2026, DOI: 10.1016/S0140-6736(26)00603-3. Three notes on use. First, later chapters cite this table in a fixed format, "row N of the transfer map," so turn back here when you see it. Second, the status column will expire. This table belongs to the book's living book mechanism, and the online edition keeps updating. You should maintain one for your own field too, and the fillable blank template and the guiding question list are in this chapter's templates. Third and most important, this table is a set of hypotheses backed by historical collateral. The historical evidence is a guarantee, not a payout. What that collateral is worth depends on the next section. Here is one demonstration of how to use it. The example is academic, and for a headline about benchmark scores the checking runs the same way. Suppose tomorrow morning you see a headline. Some tool claims to finish a literature review fully automatically, a PhD student's week of work in ten minutes. Run it against the map. Was the scenario it demonstrated cherry-picked? Row 2, treat it as a demo and wait for production-grade numbers. Whose time do those ten minutes save? The production cost of a review draft collapsed, and whether each citation is real and whether each claim drifted in the restating is still your review. Row 1, the bottleneck did not disappear, it just moved onto your desk. Whose bar did it lower along the way? The serious people and the people padding their output each saved a week, row 6. Three rows in, the headline goes from "unprecedented" to "three old patterns bound into one rerun," and you know where to look. ## 2.6 Transfer failure conditions The most dangerous moment for a map is when you forget it is a map. An analogy is a loan, and the interest is that you must inspect the collateral, meaning check whether the mechanisms in the two settings really are the same. Between research and coding there are at least five places where the mechanisms are not isomorphic. Each one discounts several rows of the table above. **Failure condition one, research has no compiler.** Code has a string of cheap ground truth sources, ground truth meaning a reliable basis that tells you right from wrong. The compiler tells you in seconds that the syntax is wrong, the tests tell you in minutes that the behavior is wrong. The whole trust mechanism of AI coding, the graded delegation of lesson four and the radius of letting go of lesson five, is all built on the foundation of "verification is cheap." Research takes its ground truth from reality itself. One experiment takes weeks, one round of peer review, meaning having researchers in the same field review the paper, takes months, an independent reproduction runs to years, and for a large share of claims the ground truth never arrives in principle. Every lesson that depends on a cheap verification loop takes a discount in transfer. At the same model capability, the radius of letting go in research is structurally smaller than in coding. This inference has a falsifiable shape, meaning it can say what observation would overturn it. AI autonomy should climb first on the research steps where cheap ground truth exists, such as rerunnable analysis code and tasks with public data to compare against, and lag on steps like interpretation and judgment where there is no cheap ground truth. If the opposite order is ever observed, this judgment is void and the map gets redrawn. A prediction once made has to be reconciled. In 2026-08 it was rechecked. Every publicly visible climb in autonomy happened on closed tasks where ground truth is cheap, benchmark scores, training time, test pass rates, and no counterexample was observed. The reconciliation record updates with the online edition. **Failure condition two, the feedback loop is one to four orders of magnitude slower.** A programmer runs the "write, run, fail, fix" loop dozens of times a day, the AI coding community paid its tuition and worked out its norms within three years, and a high-frequency loop with millions of people in it is an enormous learning machine. The research loop is measured in weeks and years, so that machine turns far more slowly. Two inferences follow. One, bad habits survive longer before they get punished, and hype keeps a longer shelf life. Two, "best practice for AI research" will converge visibly more slowly than in coding. The mess of 2026 (hold up Kuhn's glasses) will last a good while yet. **Failure condition three, the error contaminates knowledge itself, and it is almost impossible to recall.** Bad code usually has a bounded blast radius. Roll back, patch, at worst tear it down and rewrite, and the loss is counted in dollars. A bad conclusion has no rollback button. It enters the literature, gets cited, gets written into reviews, gets fed into the training data of the next generation of models, and gets used to support policy. One fraudulent medical paper published in 1998 took twelve years to be retracted. Twenty years after publication its citation count was still rising rather than falling, and among the papers citing it between 2011 and 2018, 28% still did not mention that it had been retracted (JAMA Network Open, 2019). With AI in the loop, compounding is added. An AI-generated error gets published, then gets retrieved and learned by the next generation of AI. Virtues from coding culture like "fail fast, ship boldly" have to be cut off at the publication step when they transfer to research. Keep the trial and error inside your private loop. What enters the public knowledge base is held to another standard. That is also why this book puts the red team (Chapter 10), the step whose whole job is to get people to pick holes in it, after delivery and before submission. **Failure condition four, there is no corpus at the frontier.** AI writes code well partly because it has seen a huge volume of similar code. Your CRUD endpoint has ten thousand relatives in the training data. If a research question is worth doing, that almost by definition means the answer is not in the corpus. The inference is that capability is spread unevenly. AI performs best on the already-solved parts of research, standard methods, boilerplate procedures, and field common sense included, and it is weakest exactly at the novel core where the value is. That gives a rule of suspicion. Any "AI did research" demo recorded on an already-solved problem gets treated as row 2 of the transfer map first, not as a counterexample to row 2. **Failure condition five, the precedents are themselves survivors.** You have heard the VisiCalc and CAD stories because those shifts ended well. Nobody writes biographies of the tools that drove an industry into a ditch. This map has survivorship bias built in, and the antidote is to pin the unit of transfer to mechanism. How the cost structure changed, where the bottleneck moved, what the new errors look like, all of that is on the record. History rhymes on mechanism, not on endings. In practice, run three questions before you rule with the map. How far do the premises this pattern depends on (cheap ground truth, fast feedback, recallable errors) hold in my setting? Am I transferring the mechanism, or transferring the ending? How would this judgment be falsified? The full self-check version of the three questions is in this chapter's templates too. **Want an agent to run it with you?** Paste this to your AI assistant or coding agent: ```text Help me run a transfer analysis for my own field. Walk me through the five steps of Template 2 in docs/appendices/ch02-templates.md, one question at a time, waiting for my answer before you ask the next. In step 1 I pick the precedent. In step 2 you may list candidate mechanisms for me, but mark each one "your guess" or "mine." In step 4, the failure condition audit, I score how far each of the five holds, and you do not score. In step 5 I write the falsifiable judgment, and you only check whether it states "what observation would overturn it." Once the whole map is filled in, ask me the three quick-card questions again. If any command errors, stop and show me the output. ``` ## 2.7 The unfair advantage you now hold The next time a blockbuster "AI does research" headline lands in your feed, do not do what the accountants at the 1979 trade show did and pick between panic and awe. Turn to the transfer map, find which row's pattern it is and which rerun of that pattern, judge whether the status column should change, then run the three failure condition questions and see which one it did not tell you. Settle the accounts before you decide whether to forward it. --- # Chapter 3 Β· The Truth-Seeking Workflow and the Autonomy Ladder !!! info "Chapter companion" πŸ“‹ [Chapter 3 templates](../appendices/ch03-templates.md) Β· πŸ—‚ [Template index](../appendices/template-index.md) > **This chapter's ladder.** No ladder line. The ladder itself is born in this chapter. From Chapter 4 on, every chapter opens with a ladder line that tells you how high this step can safely climb right now. Everything you need to read that line, this chapter hands over at once. > > Two maps. One shows you what serious truth-seeking looks like. The other shows you where AI can stand inside it. By the end of this chapter you can pin your own project onto them. This chapter hands you four things. A panorama of the seven-step truth-seeking workflow, a four-level autonomy ladder (with four behavioral criteria), a 2026 snapshot, and a conversion table for the five premises that bound where all of this applies, so you can place your own project. --- ## 3.1 One subscription, two outcomes What follows is a composite scene. The people are synthetic, every step of the workflow is real. Two engineers, same company, same Monday morning, same assignment. A startup claims its vector search engine "cuts latency by 90%", sales is pushing hard, and management wants a call on whether it is worth replacing the current system, answer due Friday. Both engineers use the same AI tool on the same subscription tier. The first one opens the chat box, pastes the task in as is, and adds "please do deep research." Twenty minutes later he has a fourteen-page report, market landscape, technical architecture, competitor comparison, risk analysis, all there, tidy subheadings, a respectable list of citations. He polishes the wording, adds a cover page, and turns it in on Wednesday. The second one starts from the same chat box, but he does not hand over the task in one piece. He first has the AI draw a map, which technical routes exist for this kind of engine and what their backers are arguing about. Then he grinds "is it worth replacing" into a question evidence could overturn, "under our real load distribution, can p99 latency, the time taken by the slowest one percent of requests, drop below 50 ms, with the migration cost recovered within a year." Before running any test, he locks in what counts as a win and what counts as a loss. Only then does he have the AI build a small benchmark and run one round on the company's real queries, anonymized. When the numbers come in he does not celebrate, he first checks a few common sources of illusion. Finally he has the AI play the harshest reviewer and attack his own conclusion, patches two holes, and turns in a four-page memo on Friday. At the review meeting, someone asks both the same question: "That 90%, compared to what, and measured under what load?" The first one flips through the report and finds only a paraphrase of the startup's own blog. The second one opens his own test basis and scripts: "Their 90% is an idle cold-start comparison. Under our load it is 34%, still considerable, but the migration cost takes 20 months to recover, which exceeds the criterion. Recommendation, wait." The gap between them is in how they walked. How well the prompt was written explains only a small part. The first one let AI take one step, the second let AI take seven, and stood in a different position on each. One deliverable stopped at plausible. The other reached reliable. That difference in position is the two maps this chapter delivers. The first map answers which steps truth-seeking is made of. The second answers how high AI can stand on each one. ## 3.2 The truth-seeking workflow, seven steps in a loop The first map first. In this book, "research" goes by the anchor in Chapter 1, not by profession. So this map has to be domain-neutral. A doctoral thesis, a due-diligence report, a technology selection, a quant strategy, strip off the domain clothing and the skeleton underneath is the same one. That skeleton has seven steps. ![The seven-step truth-seeking workflow in one view: master the field β†’ questions and hypotheses β†’ test plan β†’ execution β†’ read and catch errors β†’ deliver β†’ red team, matching Chapters 4–10](../assets/images/workflow-panorama.svg) **Step 1, master the field.** Turn "can't read it all" into "can ask it something," and work out who claims what in this field, what they are arguing about, and which argument your question lands in. The engineering version is digesting an unfamiliar framework's ecosystem in two weeks, docs, issue tracker, competitor threads, more volume than bandwidth. The academic version is the literature review before a PhD student's proposal, that is, reading the field's papers before starting the research to learn the terrain, finding three camps and five load-bearing papers among four thousand papers. Same problem. **Step 2, questions and hypotheses.** Grind a blur of curiosity into a question worth answering, one whose answer might embarrass you. The engineering version grinds "should we replace the retrieval layer" into "under our load, can p99 be pushed below 50 ms, with the migration cost recovered within a year." The first can be talked about for a day with no conclusion, the second can be settled in three days of checking. The academic version grinds "this phenomenon is kind of interesting" into a testable hypothesis, and states which observation would kill it. **Step 3, test plan.** Before running anything, lock in the plan, the method, and the criteria, above all what counts as losing. The engineering version of a performance evaluation sets the test load, the comparison baseline, and the pass threshold before turning on the machines. The academic version is called preregistration. Sample size, test method, exclusion rules signed off first, no door left open for picking data afterward. In an evaluation whose criteria were added later, the conclusion always "happens to" support the plan that was finished first. **Step 4, execution.** Write the code, run the analysis, run the computational experiment, turn the design into data. The engineering version is building the eval pipeline, running backtests, pulling and cleaning data and fitting models. The academic version has simulations, statistical analysis, experimental pipelines. Of the seven steps, this is the one that overlaps most with AI coding. **Step 5, read and catch errors.** Sort the results into findings, noise, and bugs. Engineering version, a metric jumped 40% overnight. Real signal, or did an upstream tracking field change yesterday? The academic version asks whether this significance survives a multiple-comparison correction, that is, after testing many groups at once, has the luck of guessing one right been subtracted? Did the data leak? In both settings, the most expensive error looks like the most exciting finding. **Step 6, deliver.** Turn "I know" into "others can trust," pick the right vehicle, and write the evidence, the limits, and the chain of reasoning into something that stands up to scrutiny. The engineering version is a one-page decision memo for the CTO, an evaluation report for the technical committee. The academic version is the paper, the figures, the point-by-point response letter to reviewers. The vehicles differ. The standard, every number traceable, is the same. **Step 7, red team.** Before release, let the harshest criticism happen at home first. Review yourself, look for counterevidence, attack your own conclusion. The engineering version is the pre-launch premortem, assume the project has already failed and reason backward to the cause of death, or the colleague whose job in the review meeting is to disagree. The academic version rehearses the reviewers before submission, asking "if I were Reviewer 2, famously the pickiest reviewer on the panel, where would I strike first." The only difference is that now you can hire a tireless opponent at any hour. Two things must be said clearly now. **First, this is a loop, not an assembly line.** The seven steps are numbered in teaching order, not in marching order. The most common thing the red team does is send you back to step 2. Interpretation often sends you back to step 3. A paragraph you cannot write clearly at delivery usually exposes a step 1 you never mastered. The second engineer's Friday memo in the opening scene is what came out after a lap and a half around the map. ```mermaid flowchart LR S1[1 Master the field] --> S2[2 Questions and hypotheses] --> S3[3 Test plan] --> S4[4 Execution] --> S5[5 Read and catch errors] --> S6[6 Deliver] --> S7[7 Red team] S7 -. sent back, the question was asked wrong .-> S2 S5 -. sent back, the criteria did not block an illusion .-> S3 S6 -. sent back, step 1 was never mastered .-> S1 ``` **Second, this map is the table of contents for Part II.** Chapters 4 to 10, one chapter per step, in the order above. How to do each step, its templates and its pitfalls, all live there. This chapter's only job is to let you see the whole map first. ## 3.3 The autonomy ladder, four levels defined by behavior The second map answers a different question. In the sentence "AI helps me do research," what does "helps" actually mean? The same "helps" can mean reformatting a citation for you, or deciding for you how the experiment should be designed, and the risk between those two differs by orders of magnitude. So the four levels of the ladder are not defined by capability adjectives. "Smarter" and "more powerful" are the marketing department's language. The definitions use **behavioral criteria**. There are only four questions. **Who drafts? Who reviews? Who decides? Who answers when it's wrong?** Change the answer to any of the four and the level changes. These four criteria only set the level. They are not the win-or-lose criteria you lock in at step 3, they just share a word. ![The four levels of the autonomy ladder: tool β†’ assistant β†’ collaborator β†’ autonomous researcher](../assets/images/ladder.svg) **Tool level.** You draft, you review, you decide, you answer for it. AI is a single-point executor, translating a methods section, converting thirty citations to another format, turning a table into a chart. One explicit instruction at a time, every output glanced over by you, errors visible on the spot and discarded on the spot. Take one day as an example. You are rushing a review to final draft, you have AI convert the bibliography from one format to another, and along the way it translates two passages from a German paper. The engineer's version of that day has it reformat an API doc and translate two chunks of error output in passing. Through the whole process it never made a single "decision." **Assistant level.** AI drafts, you review every part in full, you decide, you answer for it. AI takes on bounded subtasks, summarizing a paper, writing a plotting function, drafting the related-work section. The key criterion is the word "review." Every part of the output passes your eyes before it enters your project. Review is the only line of defense, so it has to cover everything. Take one day as an example. You hand ten papers to AI one at a time to distill in a fixed format, and check each card's key passages against the original. In the afternoon you have it draft a section, then edit sentence by sentence until it is unrecognizable, and it is still three times faster than starting from a blank page. **Collaborator level.** AI drafts and self-checks, and proposes alternatives you had not thought of. You no longer review sentence by sentence. You review the plan, review the key checkpoints, and spot-check by ratio. You decide, you answer for it, but answering for it no longer rests on your eyes alone. It also rests on a process locked in beforehand, tests, criteria, cross-validation, where cross-validation means different sources checked against each other, not the dataset-splitting kind from machine learning. The ticket into this level is **an executable standard of verification**, and no model, however strong, can buy it for you. Without locked-in criteria you have no standing to talk about spot checks. Take one day as an example. In the morning you and AI each propose a version of the experiment plan, pick each other's apart, and merge into a final one. In the afternoon it writes the whole data pipeline end to end, tests all green, and you look closely only at the interface design and three pieces of key logic. Before you log off, the day's numbers are read straight against the thresholds locked in the day before. **Autonomous researcher level.** AI drafts the whole way, makes the intermediate decisions itself, and reviews itself. The human does two things only, set the goal and the acceptance criteria beforehand, and accept the final product afterward. **Who answers when it's wrong? Right now, nobody.** AI bears no consequences, and the human never reviewed the process step by step, so when something goes wrong nobody catches it in time and nobody can be held responsible. That is exactly why in 2026 this level is common in demo videos and rare in settings where, if the answer is wrong, something real gets hurt. Take one day as an example (still mostly imagined). You leave a research goal before bed and receive a complete analysis report in the morning. The problem is, if one conclusion in it is confidently worded, fluently reasoned, and happens to be wrong, who finds it? The four levels side by side. | Level | Who drafts | Who reviews | Who decides | Who answers when it's wrong | |---|---|---|---|---| | Tool | You | You (a glance in passing) | You | You; errors visible on the spot | | Assistant | AI | You, every part in full | You | You; human review is the only line of defense | | Collaborator | AI (with self-checks and alternatives) | You, checkpoints + spot checks + a locked-in process | You | You + the process; errors hit the criteria first | | Autonomous researcher | AI (including intermediate decisions) | Mostly AI self-review, the human accepts only the end product | Goals and acceptance criteria stay with the human, the process goes to AI | **Nobody**, the root reason this level is rare | From assistant to collaborator, the biggest change is that your reviewing switches trades, from reading the output line by line to designing and maintaining a verification process. AI doing more work is the secondary change. For every notch a programmer hands to an agent, the tests and CI to catch it came first, the letting go second (Chapter 2's transfer map, trust calibration moving from two poles to graded delegation). Research works the same way. How high you can safely climb depends on how hard your criteria are written, not on how new the model is. ## 3.4 Levels live on step Γ— task, not on people Now the most important argument of this chapter. The whole book's ladder-line mechanism rests on it. **The autonomy level lives on "one kind of task at one step." It is not a global property.** The same person on the same project can work at different levels within a single afternoon, and should. The question "how much autonomy should I give AI" is wrongly posed, like asking "what grade of sterilization does this hospital run." The operating room and the outpatient lobby should never share one standard. The correct picture is a row of knobs, one per step, with further subdivisions by task inside each step. There is no master switch. Why must the levels be uneven? The three variables that decide whether letting go is safe are themselves spread very unevenly across the seven steps. - **Visibility of errors.** The execution step has a built-in alarm. Code that does not run throws an error, a broken pipeline leaves logs. The interpretation step has none. A wrong interpretation sits quietly, fluently worded, until it poisons a downstream decision three months later. - **Cost of correction.** Miss a paper while mastering the field and you add it later. Send out a deliverable with one wrong number and recalling it costs tenfold at least. - **Whether cheap ground truth exists.** The execution step sits closest to "compiler-style ground truth." Whether the tests pass is a cheap, instant, unambiguous signal. The questions step sits farthest from it. No machine can rule on "is this question worth answering" (Chapter 2 said research as a whole has no compiler, but the seven steps sit at different distances from cheap ground truth). Where the alarm is sharp, correction is cheap, and ground truth is cheap, the knob can turn up. Where errors are silent, pollute downstream, and have no ground truth to lean on, the knob must stay low, however strong your model. The execution step can climb to collaborator level because errors there are the hardest to hide. The interpretation step must fall back to assistant level, which has nothing to do with how dumb the AI is. Errors at that step are the best at disguise. Within one step, the level also varies with task granularity. Both inside the execution step, "write the plotting code" can go to collaborator level, a wrong chart is obvious at a glance. "Choose the statistical test" has to stay at assistant level. Choose wrong and numbers still come out, only their meaning has quietly changed. Both inside the delivery step, "make the paragraph read smoothly" is assistant-level work, while "dare we state this conclusion unconditionally" should never fall within AI's remit at all. Applied to the whole book, this argument becomes the ladder line at the top of every chapter from Chapter 4 on. It marks "the current safe ceiling for this step," how far the field has verified it, which door is half open, which is still welded shut. Where each concrete task in your own project stops is for you to rule on the spot with the four criterion questions. The chapter-head line is a ceiling, not an order. "Climbing the ladder" does not mean "turning every knob to maximum." The snapshot in section 3.5 will show that for some steps the right target is to stop at assistant level. Forcing "fully automatic" on those steps is running an operating room whose sterilization fails the standard. Nothing advanced about it. ## 3.5 The 2026 snapshot, which step has climbed to which level Below is the ladder water line for each of the seven steps as of this writing (the middle of 2026). It is a snapshot, not a law, and it will expire. The book's living book mechanism (the online edition keeps updating, Chapter 16) and the honest map (the table in Chapter 13 that marks the status of every claim) are responsible for updating it. The judgment leans conservative. To call something "stable," a large number of everyday users must be able to reproduce it. A few dazzling demos do not count. In 2026-08 it was rechecked against public accounts of the automated research systems of the time. The seven water lines did not move. The notes for the execution and questions rows were updated. Read the first two columns to set the level. The notes column is evidence for looking back, read it when you need it. | Step | Stable water line | Half-open door | Conservative note | |---|---|---|---| | Master the field | Assistant | Collaborator, having AI find "what the literature is arguing about" is already feasible | "Which argument is worth entering" remains a human call; the output of fully automatic review tools works as a first draft, not yet as a map (Chapter 4 expands) | | Questions and hypotheses | Assistant (batch-generating candidate questions and hypotheses) | None | "Which question is worth answering" has no reliable automated path, still exploring. The closest current systems come to "autonomous questioning" is picking up ideas already in the public literature, abandoned by people, and combining and landing them, which is an extension of execution (Chapter 5 expands); the "autonomous" claims here are mostly marketing | | Test plan | Assistant | Collaborator, still exploring | AI drafts a plan fast, but the most expensive flaws in a plan (unfair baseline, basis drift, a criteria backdoor) are exactly the kind it does not flag itself. This is an inference from adjacent evidence (models are systematically insensitive to flaws in their own output, the self-correction blind spot, arXiv:2507.02778; LLM-judge silent failure, that is, AI used as a scoring judge gets it wrong without a sound, arXiv:2509.20293), with no direct measurement yet of the "drafting a test plan" setting | | Execution | Collaborator | Local autonomy, feasible on closed subtasks with tests and criteria guardrails in place | Highest of the seven, the most direct dividend from coding transfer; the word "local" is load-bearing, leave the guardrails and it downgrades, expanded below the table | | Read and catch errors | Assistant (nominally) | None | Actually demands the strongest human presence. Models tend to say what you expect. Five frontier assistants consistently showed sycophancy across many task types, and human preference data itself rewards the behavior (Sharma et al., 2023, arXiv:2310.13548); how to make AI disagree reliably is still exploring | | Deliver | Assistant | None | Drafting, rewriting, figure captions can all be handed off; responsibility for claim strength and wording cannot be delegated | | Red team | Assistant, and widely underrated | Collaborator, systematic search for counterevidence, still exploring | Having AI attack your draft costs almost nothing and pays off at once; but it finds holes in reasoning, not the "everyone in your field knows this, only it doesn't" kind of problem | The execution row deserves a few more words. On closed tasks with full guardrails, public cases of unattended runs lasting days to weeks have appeared in 2026. The area of "local" is growing. The definition has not changed. One more ugly truth, "feasible" does not mean "necessarily faster." The METR measurement in row 3 of the transfer map is the counterexample. The most useful way to read this table is as an X-ray. Whenever a tool or a news story claims "fully autonomous research," check it against the seven steps. Do not count the steps it demonstrates, that count will cheat you. The area a demo covers keeps growing. Early on it was only execution plus delivery, now exploration, building the eval, running the experiment, and writing the retrospective can all be strung into one smooth recording. What you look for is the steps it does not cover, and that absentee list is quite stable. Who chose the question, who decided "this task is safe to hand off," who signed the criteria, who set the claim strength. Those tools are not useless. Their correct name is "collaborator on these steps," not "autonomous researcher." Of the four criterion questions, the one that pierces the packaging best is the last, who answers when it's wrong? Where the manual has no answer, treat it as nobody. ## 3.6 Where this workflow applies Both maps are standing. Before you use them to look at any real project, one thing must be said that this chapter has assumed all along without writing down. Readers in a hurry to place their project can skip ahead to section 3.7 and return to this section afterward, but you must finish it before entering Part II. Whether each chapter from Chapter 4 on needs a conversion depends entirely on this section. Section 3.4 said levels live on "step Γ— task." That sentence is incomplete. Levels also live on a third thing, your situation. Same step, same kind of task, a different person doing it, and the safe ceiling can differ by a whole level. What decides whether letting go is safe is not only the nature of the process but also what shape your evidence takes and whether you get a second chance. This workflow has five premises, and together they are the definition of the core reader in the Preface. All five hold for this book's spine case, and all five hold for most engineers' technical investigations. Your evidence is code and data, a bad run can be rerun, criteria can be written first, and there is a next round. If all five hold, you are the kind of person this book writes the full process for, and you follow Part II as written. That is the merit of the spine case as a teaching vehicle, every step of the process can be demonstrated in full, and it is also its bias, it demonstrates the best case. For each premise that fails on your project, you patch yourself in the matching place. This section tells you where the patches go. **Premise 1, errors can be found cheaply.** The data is still on disk and recomputing is free. In Chapter 8, an audit reversed a +18 into a βˆ’2.3, the army ahead at first glance, behind after the recompute, both numbers percentage points of accuracy gap. The whole thing cost a few lines of code and one rerun. It fails for wet-lab experiments (the sample is used up, the antibody is spent, the cells have been passaged until their morphology changed), one-off interviews, decisions already in production. What fails first is the whole interrogation process at step 5, which assumes you can still recompute when you find a problem. The compensation is to move the interrogation forward. The step where criteria are locked in (step 3) is promoted from "good habit" to your only chance at interrogation. This is the second-best answer under a hard constraint, not an equal trade, and the cost is recorded as is. **Premise 2, the evidence is machine-readable.** The 3,580 lines of jsonl results the spine case produced can be handed straight to AI to pull examples, cluster, and recompute. It fails for handwritten lab notebooks, csv files with improvised column names, instrument exports with merged cells, terminal screenshots from behind a paywall, data-room documents that may not leave the room. What fails first is the Chapter 12 lesson "hand mechanical checks to the machine," because the machine cannot get in. The compensation is to count the cleanup cost into the verification budget, knowing that this bill is usually larger than the check itself. Chapter 8 will say "interrogation is cheap enough that the excuses run out." Here that sentence gets discounted. The excuses did not run out, they moved from "reviewing is too expensive" to "preparing to review is too expensive." **Premise 3, the verifier is the producer.** Every verification step in the book up to this point assumes the person reviewing is the person who ran it. You can rerun your own pipeline, you know where every number comes from. It fails when a manager signs for a subordinate, an advisor signs for a student, an investment committee signs for an analyst, any setting where "the person whose name is on it is not the person who did it." Which part fails first must be stated precisely. Acceptance in Chapter 12 has three layers, and this is a separate grading used for acceptance, not the four levels of the ladder. L0 is the spot check, L1 is full verification, L2 is the adversarial recompute. The first two check whether citations exist, whether numbers trace to their source, whether the criteria's timestamp precedes the results, and the signer can do them with the output in hand, no rerun capability required. What fails is only L2's "independent re-derivation," which assumes you have the raw materials and the ability to rerun, and the signer usually has neither. So this cell is not blank, it is missing a corner. L0 and L1 are the signer's ready answer, and section 12.5 in Chapter 12 is written precisely for accepting someone else's report. What is **still exploring** is that corner, the acceptance economics of second-person sign-off, how to set the spot-check rate, what to sample, how to escalate when a check fails. This book has no deliverable answer. I mark it here rather than pretend it has been answered. All I can offer is half a rule. Use the four attack surfaces of Chapter 10 (the scorer, the data composition, the independence assumption, the cost basis) as an acceptance checklist, which beats reading the output end to end by a wide margin. **Premise 4, the criteria can be written before the run starts.** On HumanEval+ right and wrong are ruled by test cases and cost is ruled by the ledger, so locking in criteria is feasible. It fails when the criterion is itself the research question, "should this retrieval be trusted," "is this interviewee telling the truth," "does this piece of user feedback count as a real need." In these settings, the standard you want to lock in is exactly the thing you do not yet know. What fails first is step 3, and with it the whole chain. If the criteria are not hard, the collaborator-level ticket from section 3.3 is void. The compensation is to lock in an operable proxy criterion first, and at the same time write "the gap between the proxy and the real target" into the deliverable as a formal limitation. The subplot case is a living specimen of this shape. The answer distribution of real people is only a proxy, and the distance between it and the real question, "does the persona behave like a real person," was written as a caveat running through the whole case, that is, a limiting clause attached to the claim (Chapter 12). **Premise 5, there is a next round.** A conclusion that dies under review can be rerun. Criteria written badly get fixed in the next experiment. It fails when there is no budget, no sample, a deadline nine hours away, or this is your last batch of data before graduation. When this one breaks, what falls is the conclusion the whole process produces, and naming any single step is not enough. Followed strictly under these constraints, the process derives "then publish nothing." That answer is wrong. It only shows the process has hit its own boundary. The compensation is the downgrade ladder in section 12.6 of Chapter 12, which cut to make first when the budget is down to a tenth, what each cut costs, and where the honest floor is. The five premises side by side in one table. | Premise | This book's spine case | What fails first when it does not hold | Where the compensation is | |---|---|---|---| | 1 Errors can be found cheaply | Holds | The step 5 interrogation process | Move the interrogation forward to step 3 | | 2 Evidence is machine-readable | Holds | Dispatching mechanical checks | Cleanup cost goes into the verification budget | | 3 Verifier = producer | Holds | L2's independent re-derivation | L0/L1 as written (Chapter 12), the economics of the spot-check rate **still exploring** | | 4 Criteria can be locked in beforehand | Holds | Step 3, and with it the whole chain | A proxy criterion + the gap written into the limitations | | 5 There is a next round | Holds | The conclusion of the whole process | The downgrade ladder in section 12.6 of Chapter 12 | The subplot case is the only record in the book of a premise actually breaking. The real-interview arm was cut because no interviewees could be recruited, so Premise 1 (errors can be found cheaply) did not hold on that arm. No re-collection was possible, and a step that had been a bonus was promoted to mandatory (section 12.7 in Chapter 12 keeps the account). When a premise breaks, admit it first, then compensate, and write the cost next to the conclusion. The alternative, delivering in the posture of all five holding, produces a deliverable that looks identical to the real thing. One last sentence for readers who do not run computational experiments. Every case in this book is a rerunnable computational experiment, because that is inside the author's craft. The book will have no first-hand cases from wet labs, fieldwork, or clinical trials. Forcing them would only produce cliches. The process still works for you, and the five premises are the exchange-rate table for your conversion. But the conversion is your job, not something I did for you, and I write that here honestly rather than let you discover it on your own at Chapter 8. ## 3.7 Pin the spine case to the maps Now a demonstration of using the two maps on a real project. The one pinned up is the book's spine case. **Can a team of small open-source models (voting, debating, dividing the work) tie a single frontier model on a cost-matched basis?** First, verify it with the Chapter 1 anchor. Is this "research"? If the answer is wrong, does something real get hurt? Yes. Whether a company chooses "a local team of small models" or "a paid frontier API" is an architecture decision with real money on it, and it drags in data privacy and vendor lock-in. There is no consensus on the question yet either. The literature has papers on both sides (the scene of that brawl is where I take you in Chapter 4). Answer unknown, wrong answer costly. It qualifies. Then run it through the seven steps. At each step I mark in advance the level this case plans to climb to. This is a preview, not a battle report. **Master the field (Chapter 4).** Whose evidence is harder, the two camps', and where the real disagreement lies. Draw the controversy map first. Expected mostly assistant level. Scanning and distilling are handed off, the crux I judge myself. **Questions and hypotheses (Chapter 5).** Right now "tie" is a marketing word. Which task family? How is cost matched? How is a tie defined? Without a falsifiable statement, everything after is wasted runs. At this step AI generates candidate definitions for me, choosing one is my job. Assistant level. **Test plan (Chapter 6).** Pick the benchmark, match the baseline, lock in the criteria first. Assistant level drafts, I check every line. **Execution (Chapter 7).** Build the eval harness, the homemade scaffolding the experiments run in, and actually get the small-model army and the large-model baseline running. This is the step the whole case most hopes to climb to collaborator level on, guardrails in place. **Read and catch errors (Chapter 8).** The day the first batch of numbers comes out is the most dangerous day of the whole case. At this step I downgrade deliberately. AI sits as a juror, never the judge. **Deliver (Chapter 9).** The same evidence written into two vehicles, a technical report plus a one-page decision memo for the CTO. **Red team (Chapter 10).** First let AI be the harshest critic, patch, then really send it out and take the hits from humans. Note that this preview table is itself a living specimen of the argument in section 3.4. One case, and the planned levels across the seven steps range from collaborator to "deliberate downgrade." When someone asks me "what level of AI does your project use," the only honest answer is, which step? ## 3.8 Now it's your turn Before closing the two maps, spend ten minutes pinning yourself onto them. While reading this book you should have a real question of your own in hand. If you do not, pick one now, the kind where you have to give a judgment and a wrong judgment has consequences. Then do the following four things. 1. **Mark on the seven-step map the step you are stuck at now.** Be honest before you mark. Most people place themselves at execution or mastering the field, and one question, "have you locked in your criteria," exposes them. They have never reached step 3. 2. **Fill in your actual current level for every step.** Rule with the four criterion questions, who drafts, who reviews, who decides, who answers when it's wrong. A typical first fill looks like this. Two or three steps at assistant level, one step mistakenly climbed higher than it should, and the rest of the cells blank. Blank means you are not doing that step at all right now, and it is usually the red team. 3. **Circle two cells, the step you most want to climb, and the step you should least let go of.** The first is your private focus for Part II. The second is your operating room. This book teaches you to gate it there, not to save effort there. 4. **Go through the five premises in section 3.6 and circle the ones that do not hold.** This takes three minutes and decides the posture you read Part II in. All hold, follow it as written. Two or three broken, you do one extra conversion per chapter, and the last column of section 3.6 tells you which direction to convert in. Readers who broke the third premise (you sign off on someone else's output), take note. Spot checks and full verification you still run. The missing corner has no answer in this book. Do not read "the book didn't write it" as "it doesn't matter." Complete fillable versions of this chapter's three sheets are in the appendix (Ladder Self-Rating Sheet + Seven-Step Workflow Check Card + Five-Premise Comparison Table). The filled-in sheet is your personal navigation for Part II. Each time you open a chapter, read the ladder line at the top first, then check it against your own cell. **Want an agent to run it with you?** Paste this to your AI assistant or coding agent: ```text Help me pin myself onto the two maps in Chapter 3. Build an empty table from the self-rating matrix in Template 1 of docs/appendices/ch03-templates.md, then walk the seven steps one at a time and ask me the four criterion questions, who drafts, who reviews, who decides, who answers when it's wrong. Fill the level from my answers, and do not upgrade me. Which step I am stuck at, the step I most want to climb, and the step I should least let go of are mine to circle. You only bold the two cells I circle. Finally go through the five premises in section 3.6. Whether each holds is my call. For the ones that do not, tell me the conversion direction from the last column of Sheet 3. If any command errors, stop and show me the output. ``` ## 3.9 The unfair advantage you now hold Given any new "AI research" tool or news item, you can say within ten seconds which of the seven steps it lands on and which rung of the ladder it is trying to climb, and also who does the six steps it does not mention and who answers when it's wrong. The gap between the two outcomes in section 3.1 hides in the other six steps it never brought up. It was never opened by the one step it demonstrates. --- # Chapter 4 Β· Master a Field !!! info "Chapter companion" πŸ“‹ [Chapter 4 templates](../appendices/ch04-templates.md) Β· πŸ—‚ [Template index](../appendices/template-index.md) > **This chapter's ladder.** At this step AI sits stably at assistant level. Scanning, distilling, and tracking can all be handed off. The collaborator door is half open, and having AI find what the literature is arguing about is already feasible. Which argument is worth entering is still yours to judge. > > **Spine update.** I am about to start looking into whether a team of small models can tie the large model. By the end of this chapter you will see that I expected to find a consensus and found a brawl. > > **This chapter delivers.** A controversy map, a stack of paper cards, a coverage check built into the process. --- ## 4.1 Forty-seven browser tabs Half past five on a Thursday afternoon, your manager stops you in the hallway. "Give me a call by next Wednesday. Can we get our inference cost down? I hear you can stitch a pile of small models together to replace the large one. Go find out whether that holds up." You go back to your desk, open arXiv, and search "multi-agent LLM". Over four thousand results. Try another term, "model ensemble inference", another two thousand. You open one that looks like a survey, and its abstract cites thirty papers you have not read. You open five of them, and each cites thirty of its own. Ten at night, your browser holds forty-seven tabs. You have finished three and a half papers, and your notes hold a pile of conclusions that contradict each other. Your grip on "does it hold up" is, honestly, weaker than it was at half past five. Back then you at least did not know what you did not know. Change the skin and this scene is everyone's scene. The literature review before a PhD student's proposal (defending a research plan before the work formally starts), an analyst handed due diligence on an unfamiliar industry, an engineer who has to master a new framework's ecosystem in two weeks. A field's total knowledge passed any single person's reading bandwidth long ago, and your task happens to require you to "master" it. The standard answer to this used to be "grind it out." Grind for a few years, read until the returns diminish, and you are an expert. Now there is a new answer. Many people hear the new answer as "let AI read for you." That illusion is the first thing this chapter takes apart. ## 4.2 From "can't read it all" to "can ask it something" AI rewrites this step, but "read a hundred papers" did not become "read zero." What actually changed is scanning, distilling, and how you talk to the literature. The cost of scanning collapsed. Knowing who is in a field and what they are arguing about used to cost years of soaking in it. Now one afternoon of AI-assisted scanning gives you a first-draft map good enough to use. The map will have errors, but from day one it tells you the shape of the continent. Distilling can be outsourced. What is this paper's core claim, what is the evidence, where are the limits. AI does structured distilling of that kind fast and steadily. Two conditions. You force it to output in a fixed format, and you spot-check (section 4.4 gives the template). The literature became something you can talk to. The old move was read first, then think. Now you enter carrying a question. "Do these two papers contradict each other?" "Has anyone made this comparison on a cost-matched basis?" "Who proposed this method first, and who overturned it later?" Cost alignment means making several options spend roughly the same money or compute before comparing, otherwise the winner may just be the one that burned more. Every question comes back on the spot with an answer you can chase down. The working definition of "master a field" changed with it, from "read enough" to "can ask it something." You can put an insider's question to the field, and you know how to check the answer. That is what "can ask it something" means. Deciding what to believe did not change, and it is worth more than before. Which evidence is solid, which is the authors gilding their own record, which argument touches the crux of your question. Nobody can outsource that judgment. ## 4.3 Lessons stolen from an unfamiliar codebase You have taken over an unfamiliar codebase, so you know the four rules. Map before detail. Let AI be the tour guide, not the driver. It invents APIs that do not exist. Cutting across with a question beats reading file by file. "Entering an unfamiliar literature" is the same problem in a different shape. The volume of knowledge far exceeds bandwidth, the structure is implicit and written nowhere, and old errors are buried inside. The four rules move over as is, and below I look only at where they break after the move. **Lesson one, map before detail.** The literature version asks first which camps this field has, what each camp's representative works are, and what their core disagreements are, then reads specific papers. The reading order changes from "sorted by search results" to "navigated by map," and most of the efficiency gap comes from there. Where it breaks, a codebase's structure has a directory tree and a call graph holding it up, while the literature's "camps" cannot be ruled on by any machine. The first version of the map is certain to be wrong, and the coverage check in section 4.4 is the patch for that break. **Lesson two, let AI be the tour guide, not the driver.** The literature counterpart is direct. Having AI summarize a paper is fine, but the moment a claim will enter your decision, you must go back to the original and check that passage. In this workflow, "a human laid eyes on it" is the quality gate. Where it breaks, when the driver crashes there is a diff to roll back, and when a claim that drifted in paraphrase enters your judgment there is no diff to look at. **Lesson three, it invents APIs that do not exist, and it invents papers that do not exist.** The same defect in the literature setting is called a **fabricated citation**. The title is plausible, the authors are common names, the journal is real, the paper does not exist. This is exactly where it breaks. Code is lucky, a compiler error catches it, while the literature has no compiler and you have to be one. Before a citation enters your notes, confirm in a paper search engine such as Semantic Scholar or Google Scholar that it exists, then confirm the paper really contains the sentence AI paraphrased. Section 4.4 hardens these two checks into a fixed step in the workflow. **Lesson four, the unit of "can ask it something" is the question, and asking document by document is a losing trade.** You already know how to enter a codebase with "why does this request time out". The literature is the same. Paper by paper, "summarize this for me," is inefficient. Taking your own question and cutting it across the whole field is what pays. "Who has made a cost-matched comparison?" Where it breaks, a codebase is closed, and when a thread runs out it has really run out, while the literature is open, and running out usually only means your search terms ran out. Section 4.5 will show how I missed an entire continent. ## 4.4 The three-layer intake workflow What follows you can copy straight out. Split "master a field" into three layers. Each layer has a definite deliverable, and the line between handing off and gating is drawn in the open. ### Map layer, draw a controversy map in one afternoon This layer asks only for **orientation**, and leaves understanding for later. Orientation means working out which camps this field has, what they are arguing about, and which argument my question lands in. The opening prompt template for AI is below, rewrite it as needed. ```text I need to master a field, starting with a map. Field: [your field]. My specific question is: [the question you need to answer]. Give me: 1. The 3-6 main positions/camps in this field, and each one's core claim; 2. 2-3 representative works per camp (title, authors, year, venue); 3. The real points of disagreement between camps, not differences in wording, substantive conflicts of the form "if A is right, B is wrong"; 4. The 1-2 disagreements most relevant to my question. Requirement: list only papers you can give a real source for; mark anything you are unsure exists as "unsure". ``` The deliverable is a **controversy map**. A table whose rows are claims and whose columns are the papers that support it, the papers that oppose it, and the substance of the disagreement. This table is the skeleton of your whole investigation, and every paper you read afterward gets filled into it. The map layer has one hard step you cannot skip. Every paper AI lists, confirm one by one in an academic search engine that it exists. The step is mechanical and boring and takes about twenty minutes. It is the only cordon between you and "a castle in the air built out of fabricated citations." ### Skeleton layer, read five to fifteen papers closely, one card each The map tells you which papers are load-bearing. Each camp's representative works, the repeatedly cited sources, and the empirical studies tied directly to your question are worth reading closely. Close reading is not bare reading. First have AI generate a **paper card** for each one in a fixed format. ```text Read this paper and distill it in the format below, no embellishment: - Core claim (one sentence, in the paper's own wording) - Evidence (what experiment/data/task, at what scale) - Where the claim applies (limits the authors admit, in the limitations section and hidden in footnotes) - Which prior work this paper refutes or depends on - [Blank] How much I believe it: ``` Limitations, in the template, is the section of a paper that states its own limits. The last field is always left blank, and you fill it by hand after reading the key passages of the original. That field is the soul of the card. It forces you to take a position, and taking a position forces you to check. A card where you filled in "how much I believe it" and an AI-generated summary are two completely different assets. Feed the paper to AI and ask what you actually care about. "Is its baseline fair?" "How much of this gain is left after cost alignment?" "What did the authors dodge in the limitations?" ### Frontier layer, turn tracking into a subscription A field map expires. What the frontier layer solves is staying present. - Use the citation alerts in Google Scholar / Semantic Scholar to watch the three to five load-bearing papers in your skeleton layer. Whoever cites them may be shaking or reinforcing your map. - A fixed half hour every week, have AI scan the week's new literature and output in a fixed three-column format, new evidence relevant to your controversy map, signals that the map needs changing, noise you can ignore. - Lock in the criterion. A new paper is worth entering the skeleton layer if and only if it could change who wins a row of the controversy map. Let the rest flow past. Most of the anxiety of chasing the frontier comes from having no rule that permits letting things flow past. ### Coverage check, how do you know you missed nothing The blind spots of a single search path are systematic. Before you close out, run a four-way cross-check. 1. **Keyword multipath**. Have AI generate 5-8 sets of search terms from different terminology systems. The same thing goes by different names in different communities. Run a round on each. 2. **Citation graph**. Start from the load-bearing papers, look forward at who cited them and backward at who they cited, one layer each way. 3. **Reverse test**. Ask AI directly, if one paper could overturn my current map, what would it most likely look like and which community would it sit in, then go search whether it exists. 4. **Human anchor**. Find someone who really knows the field, show them your controversy map, and ask "what did I miss". You come carrying a map, and an expert can point out a structural omission in ten minutes. The complete fillable version of this chapter's templates is in the appendix (ch04-templates). ## 4.5 Spine update Β· I expected a consensus and found a brawl Now I run the workflow above in front of you. My question was planted back in Start Here. Can a team of small open-source models, on a cost-matched basis, tie a single frontier model? Now the debt comes due, and step 1 is to work out what this field actually knows. The map layer's first prompt went out, and the map that came back was blunter than I expected. This field has two contradictory answers, each with papers behind it. **The enthusiasts** have three load-bearing papers, and this time I really did open every one. "More Agents Is All You Need" (Li et al., TMLR 2024; inside the parentheses, authors, the journal or conference it appeared in, year) uses the plainest sampling and voting, and performance rises with the number of agents. The paper itself writes that the gain rises then falls with task difficulty. The condition for "catching up" also needs a clear look. It takes 15 Llama2-13B stacked together to match the score of a single Llama2-70B on one query, and the other side did not stack 15 of its own. "Mixture-of-Agents" (Wang et al., ICLR 2025) does layered aggregation with pure open-source models and reaches a 65.1% length-controlled win rate on AlpacaEval 2.0, above GPT-4o's 57.5%. AlpacaEval has two models answer the same batch of instructions, with an LLM judge comparing whose answer is preferred. The length-controlled win rate is the win rate computed after subtracting the preference that answer length buys. The judge prefers long answers to begin with. Later papers set out to correct that bias, and the correction itself was criticized for not going far enough. Add the source paper on debate (Du et al., ICML 2024), where multiple models debate each other and factuality, math, and strategic reasoning all improve. String the three together and it is still the seductive story. Intelligence can be assembled sideways, and the poor can eat the large model's meal too. Except each paper carries more qualifiers than the secondhand retelling does. **The skeptics** are more systematic than I expected. "Are More LLM Calls All You Need?" (Chen et al., NeurIPS 2024) finds that voting systems rise then fall in performance as calls increase, non-monotonically, and the mechanism is that easy and hard queries are mixed inside a task. They even give an analytical model that infers the optimal number of calls from a small sample. "Should We Be Going MAD?" (Smit et al., ICML 2024) reports that under default settings, multi-agent debate cannot reliably beat old single-model tricks like self-consistency, which has one model answer the same problem over and over and takes the majority answer. The authors also write the other side honestly, after tuning some debate systems can pull ahead. In the coverage check stage, the "title needs verifying" line flagged red in my notes came back too, and it is "Rethinking the Bounds of LLM Reasoning: Are Multi-Agent Discussions the Key?" (Wang et al., ACL 2024). A single agent with a strong prompt nearly matches the best multi-agent discussion, and discussion holds a clear edge only when the prompt carries no examples. Walk one more layer along the forward citations and it gets harsher. A systematic evaluation (arXiv:2502.08788; arXiv is a preprint repository and that string is the number; later retitled "Position: Stop Overvaluing Multi-Agent Debate", NeurIPS 2025 Position track) pulls 5 debate methods across 9 benchmarks and 4 model families. Debate among homogeneous models burns visibly more compute and still often loses to chain of thought plus self-consistency. The same paper leaves an opening. Swap the debate pool for heterogeneous models (Heter-MAD) and the average gain over single-model CoT reaches 5.8%. The 2026 equal-token-budget study (Tran and Kiela, arXiv:2604.02460), which is to say matching the token count each option may spend before comparing, gives an information-theoretic argument on multi-hop question answering. At equal budget and with the context fully used, a single agent usually matches or beats the best multi-agent system. The authors draw the boundary clearly themselves. When the context is heavily polluted, a multi-agent pipeline with filtering and validation pulls ahead instead, and the conclusion has so far been checked only on text-only multi-hop tasks. When the first round of the map was done, I wrote a "structural finding" into my notes that excited me. The two camps were not exchanging fire on the same battlefield. The enthusiasts chase accuracy and keep no cost account, the skeptics keep the account but on a narrow band of task families. Read that way, "a head-to-head comparison on a cost-matched basis" was an unclaimed patch of open ground, and it happened to catch my question. The coverage check took out half of that finding. Once the forward citations were walked, cost-matched comparisons exist, and they cluster on the skeptics' side. Smit's abstract states a three-way cost, time, accuracy tradeoff outright, and the equal-budget study put "matched" in its title. The charge of "a narrow band of task families" does not hold against these papers either. Three observations are left standing, and the first row of the controversy map gets rewritten to match. The first three columns of the last two rows are left blank, meaning they continue the substance of the disagreement from the row above. | Claim | Supports | Opposes | Substance of the disagreement | |---|---|---|---| | A team of small models can reach large-model level | Li 2024; Wang 2025; Du 2024 | Chen 2024; Smit 2024; Wang (ACL) 2024; the two 2025-26 systematic evaluations | The enthusiasts rarely disclose strict cost alignment | | | | | After alignment the conclusion is highly task-dependent. On reasoning tasks a single agent often ties or pulls ahead, and on some structured tasks debate still holds the edge | | | | | A "decision-grade" comparison at a real enterprise procurement basis (dollars per query, amortized local deployment included), with an equal-budget self-consistency control arm, across task families, is still what nobody has done in full | My question narrowed from "fill the gap" into drawing the boundary of task dependence clearly, then using a real accounting basis to supply the comparison nobody has done in full. The gap is still there, much narrower than I thought on Thursday night. One note in the book's honesty layering, this row counts as **still exploring**. My private odds at this moment, for the small-model army narrow tasks have a shot, and a tie in general settings is marketing talk. That is intuition, not a conclusion. Whether it qualifies as a falsifiable hypothesis is the next chapter's business. Last, the three pitfalls of that afternoon, each worth more than the one before. First, a paper I wrote into my first-draft notes from memory, whose title I misremembered by one word. Wang et al. (ACL 2024) asks "…Are Multi-Agent Discussions the Key?", and I remembered it as "the Answer". The verification channel doing the forward check, which is to say opening another model session that did not know which conclusion I was hoping for, caught it. One word off looks harmless, but by the time you search, cite, or get checked by someone else, a wrong title is a paper that does not exist. Lesson three guards against AI inventing papers. This time the inventor was my own memory. Second, the skeptics' cost-matched literature was hiding under a term I had not thought of, "compound AI systems", meaning systems that stitch several models or tools together to do a job. My first round of search terms all circled "multi-agent", and only when the verification channel widened the search (2026-07-18) did this batch of literature get filled in. Without that step, my map would confidently be missing a continent. The third pitfall is the most expensive. That "literature gap" that excited me survived a whole round of scanning and died on the forward-citation check. AI's hallucinations have a process guarding against them. The more dangerous one is your own hallucination when you want a gap to exist, and AI only wraps it to look more real. The eight papers read that afternoon, conclusions and qualifiers laid side by side below. When you look at the numbers, the qualifiers column is the part that decides whether any of it applies to you. | Representative work | Conclusion | Qualifiers | |---|---|---| | More Agents Is All You Need (Li et al., TMLR 2024) | Performance rises with the number of agents under sampling and voting | The gain rises then falls with task difficulty, and 15 stacked Llama2-13B match a single Llama2-70B's score on one query | | Mixture-of-Agents (Wang et al., ICLR 2025) | 65.1% length-controlled win rate on AlpacaEval 2.0, above GPT-4o's 57.5% | The judge is an LLM, and its taste for long answers was corrected only in part | | The source paper on debate (Du et al., ICML 2024) | Multiple models debating each other improve factuality, math, and strategic reasoning | Like the other enthusiasts, it rarely discloses strict cost alignment | | Are More LLM Calls All You Need? (Chen et al., NeurIPS 2024) | Voting systems rise then fall in performance as calls increase, non-monotonically | The mechanism is easy and hard queries mixed inside a task | | Should We Be Going MAD? (Smit et al., ICML 2024) | Under default settings multi-agent debate cannot reliably beat self-consistency | After tuning some debate systems pull ahead | | Rethinking the Bounds of LLM Reasoning (Wang et al., ACL 2024) | A single agent with a strong prompt nearly matches the best multi-agent discussion | Discussion holds a clear edge only when the prompt carries no examples | | Position: Stop Overvaluing Multi-Agent Debate (arXiv:2502.08788, NeurIPS 2025 Position track) | Homogeneous-model debate costs more compute and still often loses to chain of thought plus self-consistency | Swapped for heterogeneous models (Heter-MAD), the average gain over single-model CoT reaches 5.8% | | The equal-token-budget study (Tran and Kiela, arXiv:2604.02460) | At equal budget and with the context fully used, a single agent usually matches or beats the best multi-agent system | When the context is heavily polluted a multi-agent pipeline pulls ahead, and it has been checked only on text-only multi-hop tasks | ## 4.6 Swap in your project You picked your own problem back in Start Here. Now run the map layer on it, budget one afternoon. 1. **Write down your question**, one sentence, stuck to the edge of your screen. The most common way the map layer crashes is forgetting what you wanted while you search. 2. **Fire the map prompt** (the template in 4.4), and get camps, representative works, and points of disagreement. 3. **Verify paper by paper that each one exists**, twenty minutes, no skipping. 4. **Build your controversy map**, even if it holds only two rows. The point is finding the line of "if A is right, B is wrong". 5. **Answer one test question.** Which disagreement does your question land in? If it lands in none, be wary. Either your question already has an accepted answer and you can just go look it up, or you have not found the real battlefield yet, so go back to step 2 and change your search terms. Finish those five steps and your grip on the field beats "two weeks of reading papers with no map." The skeleton layer and the frontier layer are not urgent. They grow on their own as your project moves through the later chapters. **Want an agent to run it with you?** Paste this to your AI assistant or coding agent: ```text Help me run the Swap in your project of Chapter 4. I will give you my question in one sentence first, and you paste it at the top of every reply so I do not forget it while searching. Then use the controversy map prompt in Template 1 of docs/appendices/ch04-templates.md to produce a first-draft map, camps, representative works, points of disagreement, listing only papers you can give a real source for, and marking anything you are unsure exists as "unsure". Whether a paper really exists I check one by one at Semantic Scholar or Google Scholar. You may not check on my behalf, and you may not call this map "trustworthy" before I have finished checking. The line of "if A is right, B is wrong" I find myself, and you hint only if I cannot. If any command errors, stop and show me the output. ``` ## 4.7 Sober reminders - **The coverage illusion** is the biggest hidden pit at this step. AI's answers are always fluent and complete, and it never volunteers "I missed a community". My missing the entire "compound AI systems" battlefield in section 4.5 is the live example. The only antidote is a coverage check built into the process, not a cleverer prompt. - **Secondhand retelling drifts.** Between the claim AI summarized and the paper's own text, each hand it passes through drops a little of the qualifier. "Improves on arithmetic tasks" becomes "improves", and "when cost is unconstrained" disappears outright. Load-bearing papers must go back to the original. - **Do not mistake "can ask it something" for "have mastered it."** Fluent question and answer manufactures a sense of mastery. The test is the hard line in section 4.4. Can you predict what new evidence would change your map? If you cannot answer, you only toured the place, you have not moved in. - **Still exploring**. Tools that let an AI agent run a whole literature review end to end are iterating fast, autonomous search, filtering, and map-building in one pass, and the book's online case library tracks them. As of this writing, their output works as a first draft, not yet as a map. The distribution of their errors and omissions lands exactly on the most valuable judgment of all, which argument matters. ## 4.8 The unfair advantage you now hold Give you any unfamiliar field and one afternoon, and you can produce a controversy map that marks the camps, the load-bearing papers, and the substantive disagreements. Most people, in the same four hours, produce forty-seven browser tabs. --- # Chapter 5 Β· Questions and Hypotheses !!! info "Chapter companion" πŸ“‹ [Chapter 5 templates](../appendices/ch05-templates.md) Β· πŸ—‚ [Template index](../appendices/template-index.md) Β· πŸ’» [`code/persona-panel`](https://github.com/hallieren/research-rewritten/tree/main/code/persona-panel/) > **This chapter's ladder.** At this step AI is stable at **assistant level**. Batch-generating candidate questions, candidate hypotheses, and counterexample angles can all be handed off. But this step has no half-open door. "Which question is worth answering" has no reliable automated path, still exploring. Any "autonomous" claim you see at this step is mostly marketing. > > **Spine update.** Last chapter I found a brawl, and watched the "gap" I had found for myself get narrowed a notch by a citation check. This chapter I grind the surviving question into a sentence I dare sign, and the word "tie" gets forced into a number. > > **This chapter delivers.** Three things. The five steps for sharpening a question, one hypothesis with a falsification condition, and a question-sharpening card. --- ## 5.1 Two weeks of reading, still not started Your boss drops a line in the weekly meeting, "look into whether we can rebuild the support workflow with agents." The question is as big as the weather. You cannot take it and you cannot refuse it. You open a new document titled "Agent Research." A week goes by and the body is still empty. You have not been idle. You just do not know which first stroke would not be the wrong one. The academic version is the same play. A PhD student says at group meeting, "I want to study interpretability in LLMs." The advisor nods, good direction, go pin down the question. Two weeks later his progress is a reading list thirty papers longer. The advisor asks what the question is, and he says he needs to read a bit more. Both scenes have the same root illness. What you hold is a direction, not a question. Direction and question differ by whether there is a finished state, not by size. A direction has no finished state. Under the word "interpretability" you can always "read a bit more." A question has a finished state. It gets answered, or it gets falsified, and the thing is over. People stuck in a direction are not necessarily lazy. Squeezing out a question is a hard step with no feedback. Think in silence for three days and you cannot tell whether what you squeezed out is gold or waste paper, so the safest move is always to go back to reading. Chapter 3 defined this step, grind a blur of curiosity into a question worth answering, one whose answer might embarrass you. In Chapter 4 you got the controversy map. The map tells you where the battlefield is. It will not write the declaration of war for you. This chapter teaches the step from map to declaration. ## 5.2 From squeezing out questions to filtering them The way AI rewrites this step is direct. **The cost of producing candidate questions has collapsed.** Growing a question out of a direction used to take talent plus soaking, staying in a field until one day a question surfaced on its own. Now you hand AI the direction, the controversy map, and your constraints, and in a few minutes you get twenty candidates back. Much of it is silt, but for the first time the candidate space is fully on the table. Your work shifts from making something out of nothing to trimming and choosing. That is a change of trade, not a demotion. Filtering needs a sieve. A good question passes three tests, and none can be skipped. - **Testable**, you can say what evidence would kill it. A question that cannot be killed is not a question, only a position. - **Worth answering**, the answer has to change action. If you do the same thing whether the answer is yes or no, the question is decoration. - **Affordable**, the road from question to evidence is walkable with your time, money, data, and skill. A good question out of reach belongs to someone else. The three tests frame the human/machine line for this step. Generating candidates can be handed off. AI enumerates faster than you and brings more angles. Judging which candidate deserves the next few weeks of your life cannot be delegated. The reason is that the three variables in Chapter 3, section 3.4 all press onto the human side at this step. "Is it worth answering" has no cheap ground truth, picking the wrong question raises no error, and picking again is priced in months. That is where this chapter's ladder line comes from. The half you hand off has a hidden pit too, named **the mirror risk**. AI's candidate questions come from the corpus, and questions in the corpus are by definition questions already asked. Chapter 2's transfer failure condition four said the same thing. So the candidate space AI lays out has a systematic shape. "Already asked" is covered densely, "nobody has asked" thinly. The practical corollary follows. Use AI's candidate list as a floor, not a ceiling. The genuinely novel one or two, you will mostly have to add yourself. ## 5.3 Lessons stolen from writing issues You have dispatched work to a coding agent many times. An issue, a spec, that is the question you feed the agent. The tuition you have paid on this step comes to three lessons. Below, only where each carries over to research and where it breaks. **Lesson one, the clearer the question, the less rework.** Throw "this page feels slow, optimize it" at an agent and it optimizes the wrong place. You know this. The research side is identical. Feed "I want to study X" to AI and back comes fluent, hollow skimming. Feed it a question with criteria and back comes evidence. Chapter 4's lesson four said the unit of "can ask it something" is the question, not the document, which is the front of this lesson. This chapter is its back. The quality of the question itself is the first load-bearing wall of the whole workflow. Where it breaks. Write a bad issue and the agent's code fails the tests, and you know the same day. Sharpen a question badly and it takes weeks before you notice you are precisely answering a question nobody cares about. **Lesson two, fire a tracer bullet first, sharpen the smallest testable version first.** Do not expect to grind a "perfect question" at your desk in one pass. Grind the smallest version you can start testing today, and let the evidence flow back and correct the question itself. This also explains again what Chapter 3 said, "the seven steps are a loop." Your first version of the question is almost certain to be sent back for regrinding by a later step. That is part of the process, not a failure. Where it breaks. A tracer bullet in code flows back every few hours. In research one round takes days to weeks, so the first round has to be smaller. **Lesson three, where candidates get cheap, judgment gets expensive.** Row 8 of the transfer map wrote it. Search engines devalued memory and raised the value of "what to look up and what to trust." Judgment did not die. Judgment went up in price. The same pattern replays at this step. Candidate questions go from scarce to surplus, and the scarce resource moves, from "being able to think of a question" to "being able to see which question is worth answering." Read that row in full, only generation was devalued. Hand judgment over to AI along with it and you have opted out of this round of appreciation. ## 5.4 The five steps for sharpening a question Below is a workflow you can copy as is. The input is a direction plus the controversy map you built in Chapter 4. The output is one hypothesis with a falsification condition, or an honest "it will not sharpen." Budget one hour. **Step 1, diverge candidates.** Have AI put the candidate space on the table. ```text My direction: [one sentence]. Background: here is my controversy map: [paste or summarize the table from Chapter 4]. My constraints: [time budget / available data and hardware / the edge of my skills]. Generate 20 candidate research questions. Requirements: 1. Cover different grain sizes, from "worth a paper" to "worth an afternoon"; 2. Cover different positions, at least 5 of them questions the opposition or a skeptic would ask first; 3. Attach to each question one line, "what evidence could refute it." If you cannot write that line, do not list it. At the end, mark separately which questions have most likely already been asked in the literature, and what the clue is. ``` That last requirement is a detector fitted for the mirror risk. It makes AI confess which items on the list are recitation. Candidates marked "already asked" still have a use, they amount to a free novelty check. Your novelty space sits near the lines that were not marked. **Step 2, score the three tests.** You do it, AI sits as a juror. Run each candidate through the three tests and score 0, 1, or 2 on each. The value of scoring is that it forces you to state a reason. The score itself is secondary. Any candidate you dare give a 0 on "testable" is out on the spot, and the top three by total go to the next step. You can have AI take the other side ("attack the testability of this question"), but the pen stays in your hand. **Step 3, rewrite it falsifiable.** Rewrite the survivors into a falsifiable shape. Practice this sentence pattern. ```text Under [conditions/basis], the [measurable metric] of [subject], compared with [control], is [direction and threshold]. If [specific observed outcome] is observed, the hypothesis is falsified. ``` If you cannot write both lines, it is still a direction. Go back to step 2. The most common disease here is the wish sentence, "explore the possibility of X," "validate the effectiveness of Y," and what they share is that no observation could make them wrong. The cure is to pour nouns into the sentence. Which metric, compared with whom, how big a gap counts. **Step 4, rehearse "does the answer change action."** Assume the answer is yes, then assume it is no, and write down what you, or your reader, or your boss, would do differently in each case. Same action either way? The question is decoration. Go back to step 2 and take the next one. **Step 5, reality-check the resources.** Take the falsifiable sentence and ask one last time. Is the road to that "specific observed outcome" walkable for you? Can you get the data? Are the budget and the compute enough? Are the skills you need inside the range of you plus AI? If it is not walkable, shrink it. Shrink the task scope, shrink the metric, and a control arm can be cut too, until it is walkable, or until you honestly admit that this question does not belong to you right now. The output of the five steps lands on a **question-sharpening card**, with fields for direction, candidates, three-test scores, the falsifiable sentence, the action rehearsal, the resource check, and the sign-off date. The full fillable version of this chapter's template is in the appendix. ## 5.5 Spine update Β· forcing "tie" into a number The sharpening below is a synthetic narrative reconstructed after the fact, with the sequence rearranged for teaching. The hypothesis H at the end, the sentence this chapter grinds out, together with its criteria and the reasons for the values, is a recorded fact. Now watch the five steps run. My starting point is the rewritten controversy map at the end of Chapter 4. The empty ground it left was the decision-grade comparison nobody had done in full, and my private odds were that narrow tasks have a shot and a general tie is doubtful. But odds are intuition, and intuition cannot be preregistered. Preregistration means locking in the criteria and filing them before the run, with no changes after. To go further, "can small models team up" had to become a sentence the data could condemn to death. At the diverge step AI taught me a lesson first. I fed in the controversy map and the constraints and asked for twenty candidates, and more than half of what came back was recitation of the literature brawl. "The effect of debate rounds on accuracy," "the scaling curve for the number of agents," all battlefields Chen 2024 and Smit 2024 had fought over long ago, the two papers from the skeptics' side on last chapter's controversy map. The mirror risk needs no argument. It is the default shape of a candidate list. What was worth money in this round was the "already asked" column. It pushed me back to the empty ground left after Chapter 4's narrowing. The direction did not change. Three words had to be sharpened, **task family, cost alignment, tie**. Task family knocked out a tempting option first. The candidate I liked most in the first version was "can the army, meaning the small-model team system, tie GPT-4o on the AlpacaEval benchmark." Chapter 4 checked it. Mixture-of-Agents scored above GPT-4o exactly there, a ready-made comparison target. It died on "testable" in the three-test scoring. AlpacaEval uses an LLM as judge. The judge's taste for long answers has a paper devoted to correcting it, and the correction itself was criticized as incomplete (arXiv:2404.04475), a limitation that hangs on the Mixture-of-Agents paper card from Chapter 4. Test the army on a task judged by an LLM and what am I measuring, "the army is stronger" or "the army is better at pleasing the judge"? There is no telling them apart. Judge preference would contaminate the measurement into a second research question. One question with two unknowns in it equals zero answerable questions. So I set an exclusion rule. No open-ended generation judged by an LLM. Scoring has to be programmatic, cheap, and uncontested. Out of the space that remained, I chose three families. One family alone will not do, since my question is exactly "which tasks have a shot." All of them will not do either, affordability goes bankrupt on the spot. I took three that sit far apart in type, β‘  math and logic reasoning (a contamination-resistant variant set) β‘‘ knowledge Q&A (an MMLU-Pro subset, a knowledge quiz in multiple-choice form) β‘’ code (a HumanEval+ or LiveCodeBench slice, two programming test sets that score themselves). The three families stand for reasoning, knowledge, and automatically verifiable generation, and every score goes through a program. The first family says "variant set" on purpose, not the original problems. The originals may long since have entered the training corpus, and Chapter 8 opens that contamination account. Team topology is the independent variable, taking the three plays the two camps in the literature have fought over, voting (sample and take the majority), debate (multi-round mutual review), division of labor (role decomposition). Then the hardest word, "tie." Right now it is a marketing word anyone can claim. Forced into a number, it means a gap ≀ Ξ΅ counts as a tie. How big Ξ΅ should be is the first thing I signed in this case. Set it to 0 and you demand scores match to the decimal, when the jitter across repeated runs alone is bigger than that. Set it to 5 points and "clearly worse" counts as a tie, and the skeptics laugh first. I signed Ξ΅ = 2 percentage points, with the reason written down for the record. Below that gap, a team really choosing a technology mostly will not change its decision. Above it, the word "tie" does not deserve its name. This is a judgment, not a theorem. Together with all the criteria it gets its final sign-off at the Chapter 6 preregistration. The action rehearsal in step 4 gave me an unexpected bonus. Rehearsing "yes" and "no" both went smoothly. H holds, local small-model teams become a serious candidate in enterprise technology selection, and my memo would recommend a pilot. H is falsified, everyone saves the trouble and the paid API continues. But halfway through the rehearsal I ran into a middle outcome. The army neither ties nor loses badly, it catches up at three times the cost. My first version of the hypothesis was completely silent on that outcome. It wrote a gap threshold and no cost boundary, so "tie" could be bought with unlimited money. A tie bought at more than double the money is not a tie, it is burning cash. So the falsification condition gained a cost clause. Lesson noted in passing, the action rehearsal rehearses not only the answer but the falsification condition itself. A job budgeted at one hour actually took close to two, recorded honestly. The final version follows, where pp is short for percentage points. **Hypothesis H, under cost alignment, the accuracy gap between an open-source small-model team system and a single frontier model on the selected task families is ≀ Ξ΅ (Ξ΅ = 2 percentage points), which counts as a "tie."** **Falsification condition, if the army trails by > 5pp on all three task families, or catches up only at > 2Γ— the cost, H is falsified.** Note the shape of the falsification condition. Losing one task family does not kill H, since "which tasks have a shot" was part of the question to begin with. Total defeat, or catching up only by burning cash, is what earns a death sentence. Reading is done family by family with no aggregation across families, because a total averaged over three families represents nobody's decision. There is also a gray band. A single family trailing by 2 to 5 percentage points is neither a tie nor a falsification, and it sits between the tie line (≀2pp) and the falsification line (>5pp). Its verdict gets locked in when Chapter 6 writes the criteria. One word has still not been cashed in. How exactly is **cost alignment** accounted, in dollars or in compute, and how is amortization defined? That is the first pillar of the test design, and Chapter 6 opens it. The three readings together make the table below. Read it family by family, and it goes into the file as is when Chapter 6 writes the criteria. | Result on a single task family | Verdict | |---|---| | Trails by ≀ 2 percentage points | Tie | | Trails by 2 to 5 percentage points | Undecided, neither a tie nor a falsification, report it as is, no picking a side | | Trails by > 5 percentage points | This family lost. All three families lost, or catching up only at cost > 2Γ—, and H is falsified | ## 5.6 The same rasp on another field, persona research The book's second case line enters here, and it comes back in Chapters 6, 8, 11, and 12. It is one concrete shape of the question, can synthetic data stand in for real data? Product teams use persona-bearing models for user interviews and survey rehearsals, the "a 25-year-old mom with one child who works in manufacturing" kind, and save a round of recruiting real people, in time and in money. The answers read a lot like a real person. So the question arrives, can AI personas replace interviews with real people? That proposal on your desk to "run it with synthetic users before launch" is the same question. Hold it up to the three tests and it is a direction, not yet a question. "Replace" cannot be killed. Supporters demo ten interviews indistinguishable from the real thing, opponents point out ten distortions, and both sides can talk past each other forever. Run the same five steps on it. The questionnaire itself does not matter in this case. Watch the process. Rewriting it falsifiable, first swap "replace" for something measurable, the distribution of persona answers against the distribution of answers from real people. That needs a ready-made control. Some large public survey question bank, where how real people answered is on the record, let the personas answer the same set, and compare the distributions. The first falsifiable sentence takes shape. Does the agreement between a persona model's answer distribution on a public survey question bank and the answer distribution of the corresponding real-population subgroup reach a preset threshold? The "questions the opposition would ask" from the diverge step earned its keep here. The opposition's hardest question, the averages match, is that enough? If every "25-year-old mom" gives a textbook-consistent answer, matching averages become the most dangerous illusion of all. The decision maker holds a stereotype recitation machine and takes it for a population. So the question gained a second clause, and no systematic stereotype drift appears (answers for a subgroup more extreme and more homogeneous than real people's). The action rehearsal passed cleanly. Both clauses met, persona rehearsal can enter the formal research process as a coarse screen. Either one missed, it is only a toy for brainstorming, and its output is banned from decision documents. The resource check walks through too, public survey data is free to get. Which question bank to use as ground truth, the real answers used for checking, and where to set the threshold, that is test design work, and Chapter 6 picks up the hook. What this demonstration is really about lives in the contrast between the two lines. The spine compares models, the subplot compares data, one substitute is a small model and the other a synthetic person, and the same five steps carried both through. What this process sharpens is **the shape of the question**. The metric is measurable, the control is ready-made, the falsification condition can be written, and domain knowledge counts for little of it. ## 5.7 Swap in your project Last chapter you built a controversy map. Now give it a declaration of war. Budget one hour, and run the clock. 1. **Write down your direction**, however vague. This is raw material, not output; 2. **Run the diverge prompt** (the template in 5.4), get 20 candidates, look first at the "already asked" marks, your novelty space sits near the lines that were not marked; 3. **Score the three tests**, pick the top three, and remember "testable" is a veto; 4. **Write a falsifiable sentence for each**, and whichever you cannot write is out on the spot; 5. **Run the action rehearsal and the resource check on the survivors**, put the last one standing into the question-sharpening card, and sign the date. The five steps can also end in total defeat, with no candidate clearing all three gates. The hour was still not wasted. You just confirmed that this direction will not yield a question of your own right now, and what you saved is the weeks you would have spun in place. Go back to Chapter 4 for a different battlefield, or loosen the resource constraint, and run another round. **Want an agent to run it with you?** Paste this to your AI assistant or coding agent: ```text Help me run the Swap in your project of Chapter 5, clock running, one hour. I give you a vague direction, and you use the diverge prompt of step 1 in docs/appendices/ch05-templates.md to produce 20 candidate questions, marking each "already asked" or "not seen," with a source for the mark, and "unsure" when you cannot give one. The three-test scoring is mine, you only write my scores into the table, testable is a veto, and if I forget to veto you remind me. The falsifiable sentence is mine to write, and you judge one thing only, whether the sentence says "what counts as losing." Run the action rehearsal and the resource check with me using the prompts of steps four and five, and the last one standing I put into the question-sharpening card and sign the date. If any command errors, stop and show me the output. ``` ## 5.8 Sober reminders - **Mistaking a direction for a question** is the most common crash at this step, and the person rarely feels it, since a direction can keep you very busy. There is only one test. Does it have a finished state? If you cannot say "under what conditions this thing is over," what you hold is still a direction. - **Wish-list hypotheses** come second. "Validate the effectiveness of X," "explore the potential of Y," grammatically research, logically a wish, with no observation that could make them fall through. The falsifiable sentence pattern (5.4, step 3) is the targeted antidote. A statement that will not fit that pattern, do not call it a hypothesis. - **Letting AI pick your topic** turns the knob past the safety line. This step has no cheap ground truth. When AI says "I recommend number 3," hold it up to the four criterion questions from Chapter 3. Who answers when it's wrong? Nobody. Stack the mirror risk on top and what it recommends is usually the mode of the corpus, while what you are looking for sits by definition at the corpus edge. Generation handed off, the ruling taken back, a line this chapter has now drawn three times. - **Over-sharpening** is the pit in the other direction. Sharpening a question can become an advanced form of procrastination, the question always one round of polish short, which conveniently means no work has to start. The five steps are budgeted at one hour. Run well over and what you lack is the first tracer bullet (lesson two), not a better question. Walk into the next step with a "good enough" question and let the evidence keep sharpening it for you. - **Still exploring** deserves its own entry. Research systems that generate hypotheses automatically, some claiming end-to-end "propose and verify a finding," are iterating fast. Look closely at the loudest results and every one carries qualifiers. Sakana's AI Scientist-v2 passed a workshop at ICLR 2025, meaning a session hanging under the main conference with a lower bar, and the review was only semi-informed, reviewers knew AI papers were mixed into the batch but not which ones. Zochi claims acceptance at the ACL 2025 main conference, the highest venue, meaning place of publication, among this batch of systems, but the result is vendor self-reported with no independent audit. The overall signal is still weak, and the case-by-case interrogation is in Chapter 13. You also have to see clearly how this class of system actually wins. For the places they take in public competitions, the ideas come almost entirely from human published papers and community discussion, and their most striking skill is picking up ideas others abandoned as too hard to implement, combining them, landing them. That is an extension of execution, strong and valuable, but it is not asking questions. Packaging combinatorial search as "autonomously proposing research directions" is the most common line on this front. As of this writing they are good at mass-producing candidates in the neighborhood of solved problems, which is assistant level doing its job. On the judgment of "which one is worth answering," no system's performance makes me willing to write a name here. The book's online case library tracks this front. ## 5.9 The unfair advantage you now hold Give you any vague direction and one hour, and you can grind out a question with a falsification condition, a rehearsed action, and a walkable resource path, or honestly find that it will not grind. Both outcomes are worth more than "I need to read a bit more." --- # Chapter 6 Β· Turn an Idea into a Falsifiable Test Plan !!! info "Chapter companion" πŸ“‹ [Chapter 6 templates](../appendices/ch06-templates.md) Β· πŸ—‚ [Template index](../appendices/template-index.md) Β· πŸ’» [`code/persona-panel`](https://github.com/hallieren/research-rewritten/tree/main/code/persona-panel/) > **This chapter's ladder.** At this step AI is stable at **assistant level**. Drafting the plan and enumerating confounders can both be handed off. **Collaborator level** is still exploring. AI drafts a plan fast, but the most expensive flaws in a plan, unfair baseline, basis drift, a criteria backdoor, are exactly the kind it does not flag itself. Reviewing and signing off the criteria is the hard step that stays in your hands. > > **Spine update.** Hypothesis H is already sharpened, and this chapter turns it into a plan signed off before the run. By the end you will see that the soul of the whole plan is one control arm I nearly left out. > > **This chapter delivers.** The seven items of a test plan, the confounder checklist, plus a plan red-team prompt. --- ## 6.1 The eleventh minute of the review meeting Friday afternoon review meeting, and you are on the third page. For the past two weeks you evaluated a new retrieval design, it beats the production baseline by a clear margin, pretty curves, a clean conclusion. At the eleventh minute the engineer sitting in the corner looks up. "The recall parameters on the baseline, those are the defaults from when it shipped, right? How long did you tune your new design?" Two weeks. Your new design ate two weeks of your careful tuning, the baseline got not one minute. The room goes quiet. That is the verdict. How much of the gain comes from the design itself and how much from those two weeks of tuning, your data cannot answer. Two weeks of experiment, void. Change the skin and the same scene plays out after submission in one line from a reviewer, "the baseline appears undertuned," and in an investment committee in one question, "how were the comparison companies picked?" The structure does not change. **When the fairness of a test is not designed before the run, the hole always gets found after the results are in, by someone else, for you.** Rerun two weeks and that bill can still be counted. The bigger loss is that you have already seen the numbers. On the rerun you know which configuration produces the good result and which slice favors you, and from then on every "reasonable technical choice" you make carries a direction of preference. The first run, you are the experimenter. The rerun, you are an interested party. Chapter 3 named this step already. In an evaluation whose criteria were added later, the conclusion always "happens to" support the plan that was finished first. This chapter takes that sentence apart at the mechanism and hands you a workflow that keeps it from happening to you. ## 6.2 Degrees of freedom colluding with motive First be clear who the enemy is. Set fraud aside. It is rare, shameful, and relatively easy to catch. The enemy is a collusion between two things that are each innocent on their own. The first is **degrees of freedom**. Not the parameter you compute with in a statistics class. Here it means how many decisions can be argued either way. Between an idea and a number, any test has dozens of decisions. Which slice of the task, which metric, how much tuning budget the baseline gets, whether outliers are dropped, when to stop running, which runs count. Each decision on its own has a defensible case both ways. The second is **motive**. You want a certain result to hold. That is no disgrace. Someone with no preference would never take on this problem at all. But preference means one thing. Once dozens of "either way is reasonable" decisions are deferred until after you see the data, what pulls them needs no saying. Statisticians call this the "garden of forking paths" (proposed by Gelman and Loken in a 2013 working paper, formal version in *American Scientist* 2014, titled "The Statistical Crisis in Science"). No single act of wrongdoing is needed. Follow the result at every fork and the destination is a beautiful false conclusion. Feynman said the same thing in his 1974 Caltech commencement address. "The first principle is that you must not fool yourself, and you are the easiest person to fool." ("Cargo Cult Science," in *Engineering and Science*, 1974) This collusion is not a new disease of the AI era. Statistical software turned significance into one click and p-hacking got industrialized, the old plot of row 6 of the transfer map. After psychology got burned by the replication crisis (it erupted around 2011, when a large share of published results failed to reproduce once a different group reran them), the prescription written was **preregistration**. Sample size, test method, exclusion rules, decision thresholds, all locked in, filed, and signed before the data is seen. In 2013 OSF launched, a platform that hosts preregistrations, and Registered Reports arrived, a submission format that reviews the plan before it reviews the results. Those two are where the prescription landed (registration for clinical trials came earlier, and its spread through the social sciences really did follow the crisis). The pharmacology is simple. **Lock the degrees of freedom before the data is seen and motive can find no fork.** What this chapter delivers is the field-neutral version of the preregistration idea. You do not have to be an academic researcher, and you do not have to use any platform. There is one core action. Before the run, write the plan, the basis, and the criteria, above all "what counts as losing," into a file with a timestamp on it. After the run, that file is your conclusion's alibi. ## 6.3 Drafting collapsed in price, sign-off did not The way AI rewrites this step is isomorphic to the way it rewrote the literature review. Three things really changed. Drafting a plan collapsed in price. Writing a plan is dull, running an experiment is tempting, and most people skip the first and go straight to the second. That is how the person in section 6.1 started. Now, from your hypothesis to a structurally complete first draft of a plan, minutes. The draft will have errors, but "a draft you can attack" and "nothing at all" are two different projects. Enumerating confounders became AI's strong suit. A confounder is another cause that could explain the same result. Enumerating them is a classic breadth problem. It tests how many ways you have seen an experiment crash, not how deep you think. You dry up at five confounders, it lists twenty without breathing. Fifteen of them do not apply, and among the remaining five there are usually one or two you genuinely had not thought of. Counter-plan generation went from a luxury to a commodity. It used to take a senior collaborator to think for you, "if I wanted to overturn this conclusion, how would I design the experiment?" Now that sentence is a prompt and the cost is near zero. Section 6.5 freezes it into a template. One thing did not change, and it is the vital point of this step. **Reviewing and signing off the criteria.** The note in this chapter's ladder comes from the Chapter 3 snapshot. Unfair baseline, basis drift, a criteria backdoor, these three kinds of flaw AI does not flag itself. The reason is no mystery. They raise no error, they do not stand out, each one on its own looks like a reasonable technical choice, and they tend to grow along your preference. A model inclined to go along with the user will not pick a fight with your criteria on its own. Drafting, enumerating, playing the contrarian, hand them off. The final wording of every criterion you read word by word, then sign. What you sign is "if it's wrong, I answer for it," and no model today can carry that sentence. ## 6.4 Lessons stolen from test-first You know test-first (TDD), write the test before the code, watch it go red before you make it green. The three lessons below look only at where it carries over to research and where it breaks. **Lesson one, locking the criteria first is the research version of test-first.** A test written afterward grows into the shape of the code, and this one has bitten you. It verifies "what the code actually does," not "what the code was supposed to do." Criteria work the same way. Criteria set after the run grow into the shape of the result, and wherever the result lands is exactly where the "reasonable threshold" gets drawn. Writing the criteria first is the only way to make criteria independent of the result, and the virtue of rigor is secondary. Where it breaks. A test written afterward at least still catches regressions. Criteria written afterward are worth nothing at all. They only endorse the result. **Lesson two, watch it go red first.** A test that has never been red does not count when it is green. The research counterpart is the **falsification rehearsal**. Once the criteria are written, invent a concrete set of numbers and check whether they would really rule against you. If you cannot think of any set of results that would trigger "H is falsified," what you wrote is decoration, not criteria. Hypothesis H from Chapter 5 came with its falsification condition (trailing by more than 5 percentage points on all three task families, or catching up only at more than 2Γ— the cost) exactly for this moment. It can go red. Where it breaks. Watching a test go red takes one run. Watching criteria go red takes imagination, and nobody runs it for you, which is why this step is the easiest one to skip. **Lesson three, verification infrastructure sets the radius of letting go.** Row 5 of the transfer map, a team with solid test coverage dares to let an agent into core modules. The research counterpart is in the plan in front of you. Chapter 3 said the ticket into collaborator level is an executable standard of verification, and no model, however strong, can buy it for you. Without locked-in criteria you have no standing to talk about spot checks. That ticket gets printed in this chapter. In the next step, execution (Chapter 7), what lets you turn AI loose on a whole eval pipeline without watching every line is that every number ends up hitting the criteria this plan locked in. Trust is no help here. ## 6.5 The seven items of a plan Below is a practice you can copy as is. A test plan that holds up under examination has seven items, and each one missing leaves a class of accident waiting for you. ```text # Test plan (signed before the run; after sign-off, only appended change-log entries, no edits) 1 Question: what this test has to answer. One sentence. 2 Hypothesis H: falsifiable statement with its scope. (The Chapter 5 output, copied as is.) 3 Arm design: - Main arm (your design): - Baseline arm (control): lock in the tuning budget the baseline gets, equal to the main arm. - Steelman arm: write down the sentence the strongest opponent would use to call your result an artifact, then name the arm built to block it. 4 Basis: one primary basis (how it is computed, down to the formula); sensitivity basis listed separately. Locked-in commitment: if the two bases reach opposite conclusions, report it honestly, no picking. 5 Criteria and falsification condition: what counts as a win, as a loss, as undecided, written down to "what number triggers it." Statistical test, repeat count, random seed, stopping rule, all set beforehand. 6 Contamination and confounder checklist: run every line (see the appendix); for lines you cannot clear, write a mitigation, or write it honestly into the limitations. 7 Filing: stamp a timestamp (git commit / an email to yourself or the team / OSF, AsPredicted and other preregistration platforms, both free to use as of 2026). ``` One sentence each on the fourth and fifth items first. The sensitivity basis is a second way of computing, used to check whether the conclusion changes when the computation changes. The stopping rule says how many runs count as done and under what condition you may stop early. Of the seven, item 3 needs the most unpacking, above all the steelman arm inside it, the control arm that makes a result interpretable. The core of arm design is growing the control arm on the right enemy. Merely "having a control" is not enough. The most common crippled design compares only against "doing nothing" or a straw-man baseline, which proves your design beats inaction and proves nothing about it beating the cheapest alternative explanation. When you design the arms, ask yourself one question. If the result comes out as I want, what would the person who least wants this conclusion to hold say? That sentence of his has to be pinned by an arm built for it. In section 6.6 you will watch me nearly crash on exactly this. The confounder checklist in item 6, the in-text version carries only the five most painful lines, with the full version in the appendix. One, is the baseline fair. Two, are the budgets on both sides measured on the same basis. Three, is there a cheaper explanation that could eat your effect, and does it have an arm of its own. Four, could the evaluation data have been "seen" by the model or by your process. Five, are the metric, the slice, and the stopping rule unique and set beforehand. Last, the plan red-team prompt. Once the plan is written, before you sign, hand it to AI for one round of attack. ```text This is my test plan: [paste the full plan] Your job is to overturn it, not to improve it. Assume you are the reviewer who least wants this conclusion to hold. 1 List every degree of freedom in the plan that "can still be moved after the results are in"; 2 For each one, say which conclusion it would favor if adjusted after the fact; 3 Give one cheapest alternative explanation that, with the arm design unchanged, would produce the same result; 4 Point out which criterion's wording leaves a backdoor (words like "as appropriate," "a reasonable range"); 5 If you were to add one arm built to make trouble for my conclusion, which arm would you add and why. Raise only problems specific enough to act on, no general methodology advice. ``` There is one iron rule of use. What an AI red team produces is a draft list. The ruling is not in its hands. Go line by line, what you adopt turns into changes to the plan, what you reject gets one line of reason, and that "rejection record" becomes your ammunition at the defense later. Remember the note in this chapter's ladder. It can find wording backdoors, but the most expensive class of flaw it may not flag, and what answers for that is still the checklist and you. Red-teaming your **conclusions** is Chapter 10's business. What gets red-teamed here is a plan that has not run yet. The complete fillable versions of this chapter's three templates are in the appendix. ## 6.6 Spine update Β· I nearly left out the steelman arm Now watch the seven items run. At the end of Chapter 5 I held hypothesis H. Under cost alignment, the gap between an open-source small-model army and a single frontier model across three task families is ≀ 2 percentage points, which counts as a tie. The three families are math and logic reasoning, on a contamination-resistant variant set, knowledge Q&A, on an MMLU-Pro subset, and code, on a HumanEval+ or LiveCodeBench slice. Team topology takes three forms, voting, debate, division of labor. My first draft of the arm design had two arms, a single call to the frontier model as the baseline, the small-model army as the main arm, topology as the independent variable. It looked complete, and for a while I thought item 3 could be ticked. What stopped me was the third line of the checklist, **is there a cheaper explanation that could eat your effect?** I sat with that line for a while, came up with nothing, and nearly wrote "none." Then, following the process, I went back to the Chapter 4 controversy map, to the skeptics' column. On the Smit 2024 paper card was a distillation I had written by hand two weeks earlier. Under default settings, debate cannot reliably beat old single-model tricks like strong prompting plus self-consistency. The Chen 2024 card said performance is non-monotonic in the number of calls. Rereading the two cards side by side, I finally saw where the skeptics' core objection lands. The sentence they hold onto is **"the gain from teaming is only the gain from spending more sampling budget,"** far more precise than "teaming is useless." Rehearse my two-arm design against that sentence. Say the army beats frontier and I announce "teaming works." A reviewer on the skeptics' side takes it apart in one line. "You gave the army a budget of ten calls and gave the single model one. Give the same budget to one small model sampling ten times and taking the majority, is the gap still there?" My plan had no defense against that sentence at all. The army winning might only prove "spending works," and I would misreport it as "teaming works." So a third arm enters, **a single small model doing self-consistency at the same budget**, the self-consistency those paper cards were talking about. One small model, sampling repeatedly at exactly the army's budget and taking the majority. If the army beats the third arm and the third arm beats the baseline, the team structure itself contributes. If the army only matches the third arm, the "team" in the army is decoration and the gain comes entirely from sample count. **Without this arm, however pretty the result, "teaming works" cannot be told apart from "spending works."** That is the steelman arm of this plan. Looking back, the way this arm was nearly lost is representative. I knew Smit 2024 perfectly well. I filled that paper card in myself. Whoever designs a plan stands by default on the side of his own hypothesis, while the steelman arm grows on the enemy's argument, and you do not spontaneously design for the enemy. What pulled me back was one line on a checklist, nothing to do with inspiration. A process does not depend on the state you are in that day, which is exactly why you want a process. This chapter's ladder says the most expensive flaws in a plan "are exactly the kind AI does not flag itself." In fairness, I did not flag it either. Once the third arm stood, the words "same budget" turned around and forced me to lock in the cost basis, or budget alignment itself becomes one more degree of freedom to fiddle with afterward. I set the primary basis as **dollars per query**, with two sets of books, one at API list price, one at amortized local deployment. The two sets of books match two real situations enterprise readers are in, and the conclusions may differ. The sensitivity basis is listed separately, accounted in compute. The parameters actually active in one model run, times the tokens it produces, is the unit of this ledger. This ledger has nothing to do with list price or discounts. It is cost at the physical level. One commitment is locked into the plan. **If the two bases reach opposite conclusions, report it honestly, no picking.** The reason for the ranking goes into the plan too. Chapter 9 writes a one-page memo for the CTO, and that kind of reader thinks in dollars, so dollars are the primary basis. The compute basis answers "does the conclusion still hold when list prices change?" The criteria item, locked in line by line, preregistration style. The concrete task sets and slices for the three families, the random seed, the repeat count per configuration, significance by paired bootstrap (paired resampling, used to put an interval around the gap), the data contamination check as a step of the process (whether the evaluation set might appear in the model's training data, run family by family), and H's falsification condition filed as is. Trailing by more than 5 percentage points on all three task families, or catching up only at more than 2Γ— the cost, and H dies. The verdict for the gray band is locked in first too. A single family trailing by 2 to 5 percentage points is neither a tie nor a falsification, ruled "undecided," reported honestly, no picking a side. For a few parameters still hanging I give a provisional basis, and the three sign-off items are listed here. - The gap threshold Ξ΅, the tie line signed last chapter, stays at 2 percentage points. - The small-model size tier is set at "the smallest practical tier among open-source models in service today," total parameters ≀21B. - The frontier control is provisionally one of the mainstream flagships at the time of writing. There are two reasons for setting the size tier there. One, quantized it runs on a single consumer GPU, which keeps the promise that readers can reproduce it. Two, and this matters more, it keeps the word "small model" honest. Team up 100B-class "open-source large models" to tie frontier and a win does not answer the original question. The army's concrete members, together with the frontier model, are deferred and locked just before the runs start in Chapter 7, since list prices and availability change month to month and locking too early is false precision. The locking itself goes into the plan's change log. It does not happen quietly. Once the plan is final, stamp a timestamp and file it. It will be the first file in the Chapter 7 eval harness repo, and the harness may only implement it, never revise it. Honesty layering as usual, label whatever needs labeling. At this moment I hold a design and no data, so H remains **still exploring**. The only thing this plan guarantees is that whatever Chapters 7 and 8 produce, you can reconcile it against today's file. ## 6.7 Subplot Β· what does the persona study use as ground truth Move the same step onto the subplot and the difficulty jumps a level. Chapter 5 ground the persona study into a testable question. Does the answer distribution of a persona model on a public survey question bank agree with the distribution of the corresponding real population subgroup to a preset threshold, with no systematic stereotype drift. Now it is the test design's turn, and the first question stops you cold. **That "distribution of the real population" in the criteria, where does it come from?** The spine case has it easy at this step. Math problems have gold answers, code has test cases, ground truth is ready-made, and all I have to design is how to use it. The persona study has no such luxury. "How a real person would answer" is itself something you have to pay to find out. There are two candidate answers, using a large public survey dataset as ready-made comparison, plus small-sample real-person calibration. Public data means long-running cross-national social attitude surveys like the World Values Survey, with the public question bank of Pew, a polling organization, as backup. Each road has pits of its own. Public survey data, for instance, is very likely already in the model's training corpus, which is the textbook case of the "has the data been seen" line on the confounder checklist. !!! note "WVS licensing (skip on first read)" The seventh wave of WVS is free for non-commercial research. The price is registration, a citation obligation, and no redistribution of raw data. **When ground truth is not ready-made, half the work of verification design is designing the ground truth itself.** The full version of this plan, its pits and its criteria, is revealed in Chapter 12. ## 6.8 Swap in your project In Chapter 5 you ground out your own hypothesis. Now give it a plan, budget one evening. 1. **Write down the strongest opponent's sentence.** "If the result comes out as you want, how would he call it an artifact?" Then check your arm design. Does that sentence have an arm built to pin it? If not, add the arm before you go on; 2. **Write the criteria down to "what number counts as losing."** Then rehearse the falsification, invent a concrete set of numbers, and confirm it really triggers "I lost." If it cannot trigger, rewrite until it can go red; 3. **Rank the bases.** One primary basis, down to the formula; the sensitivity basis listed separately; write down "opposite conclusions, report it honestly"; 4. **Run the confounder checklist** (full version in the appendix). Do not force through lines you cannot clear, write a mitigation, or write it honestly into the limitations; 5. **Run one round of the plan red-team prompt**, rule on every line, change what you adopt, write the reason for what you reject; 6. **Stamp a timestamp and file it.** git commit, an email to yourself, the team wiki, any form will do as long as it cannot be altered. Finish with one test. Hand the plan to a colleague or an advisor, let them read only the plan, and have them guess "what result would make you admit you lost." **If they cannot guess, the criteria are not locked in yet**, go back to step 2. **Want an agent to run it with you?** Paste this to your AI assistant or coding agent: ```text Help me run the Swap in your project of Chapter 6. Build the test plan of Template 1 in docs/appendices/ch06-templates.md as an empty form, and ask me through the seven sections one by one, with hypothesis H copied as is from the Chapter 5 output. The strongest opponent's sentence is mine to write, and you check one thing only, whether the arm design has an arm built to pin it. After I write the criteria down to "what number counts as losing," you invent a concrete set of numbers and run the falsification rehearsal, confirming it really triggers "I lost," and if it cannot trigger, send it back to me to rewrite. The confounder checklist is mine to run line by line, and you record the lines I cannot clear into the limitations. Then open a separate session, paste only the plan and not my expectations, run the plan red-team prompt of Template 3, and the charges that come back are mine to rule on one by one. At the end, remind me to git commit for the timestamp. If any command errors, stop and show me the output. ``` ## 6.9 Sober reminders - **Adjusting criteria after the fact** is the number one failure mode at this step, and it always happens under a respectable name, such as "only once it was running did I find the original criteria unreasonable." Set the rule in advance. Criteria can change, but only by appending a change-log entry that states the change and the reason. Any conclusion under criteria changed after the results were seen is downgraded to exploratory, a lead and not a conclusion, and getting confirmatory status back means rerunning. - **A missing control arm** typically shows up as controlling against the wrong enemy. There usually is a control in the plan, it just compares against "doing nothing" or a straw man instead of against the cheapest alternative explanation. There is one test. Can every one of your main opponents' arguments point at some arm in the plan? - **Reporting sensitivity analysis selectively** is the basis version of "picking data afterward." Both bases went into the plan, so both go into the report. Presenting only the flattering one is the same act as deleting the ugly data points, just dressed in the clothes of sensitivity analysis. - **Preregistration does not forbid exploration.** Exploratory analysis outside the plan is free to do and is often the source of the next hypothesis, as long as it comes on stage labeled "exploratory" and does not wear the clothes of a confirmatory conclusion. - **Still exploring.** Tools that let AI draft a whole test plan end to end are iterating fast, and the output is already usable as a first draft. As of this writing there is no evidence that models reliably catch unfair baselines and criteria backdoors in their own drafts, so before criteria get signed as a contract, the step where a human reads them word by word cannot be skipped. The book's online case library tracks progress. ## 6.10 The unfair advantage you now hold Before you run anything, you hold a plan with "what counts as losing" locked in. Three weeks later, however ugly the result, your conclusion is natively immune to the charge of "picking data afterward." Most people do not discover they need this file until the eleventh minute of the review meeting. --- # Chapter 7 Β· Execution !!! info "Chapter companion" πŸ“‹ [Chapter 7 templates](../appendices/ch07-templates.md) Β· πŸ—‚ [Template index](../appendices/template-index.md) > **This chapter's ladder.** Execution is the step AI climbs highest on among the seven. Collaborator level is solid, and the door to autonomous researcher level opens a crack here, which counts as local autonomy. The reason is not mysterious. This step has cheap ground truth. Tests go red, the ledger adds up, a number can be recomputed, errors do not stay hidden long. The ladder carries one footnote. Building the pipeline and writing the tests can both be handed off. The one exception, "the harness itself is wrong," is the kind AI does not flag, and all five bugs in this chapter are that kind. > > **Spine update.** The plan was signed off, and it was time to spend real money. I thought the pilot would be out in twenty minutes. Three hours later I had caught five bugs, and every one of them pinned the strongest arm to the floor at zero first. > > **This chapter delivers.** The four pillars of a harness, a minimal checklist, and a pilot process that spends 3-5% of the total budget. --- ## 7.1 A twenty-minute plan, a three-hour reality The day of the run my confidence had grounds. The Chapter 6 plan went into the repo ahead of any result. The harness, the homemade scaffolding the experiments run in, holds the data loading, the model calls, the scoring and the accounting, about four hundred lines, a model client, a ledger, tasks, topologies, a batch runner, plus two small pieces for reports and statistics, with 19 unit tests, all green. Under mock mode the whole chain had run end to end more than once. What was left, swap the fake model for the real API and spend a dollar or so on a 20-problem pilot, a small-scale trial run. I glanced at the clock and put twenty minutes on it. The first real call, a 400 error. The plan says the frontier arm runs at temperature 0, that is, no random sampling, always take the most likely token. But GPT-5.6-terra, locked in just before the run, is a reasoning model, one that generates a stretch of internal thinking before it answers, and its API refuses a custom temperature and wants the token ceiling passed under a different parameter name. The code itself was not wrong. What was wrong was the code's assumptions about the world, and mock mode had faithfully mocked the whole world away. Fixed, run again. This time it started, and halfway through the process died. qwen3.5-9b in the army is also a reasoning model, a 1024 token ceiling was not enough for it to finish thinking, the reply cut off midway came back with content null, my code used null as a string, and it crashed on the spot. The way it crashed was more embarrassing. My first batch runner was a naive serial loop. One problem throws, the whole run lies down, and the problems already finished have to be fished back by the resume mechanism. The mechanism was there, thankfully, but every crash cost me one more restart and one more stretch of watching it. The third time was OpenRouter, the middle layer that forwards requests to the various model providers, hiccupping. Status code 200, response body not JSON. Parsing blew up and the whole run lay down again. By now I had to admit the crashes kept landing in the same place. The design was brittle. Why should one problem's failure put the whole experiment to death? The fix went into the skeleton of the batch runner. Tasks are flattened into small independent pieces, a single problem's failure goes onto a failed list only, and the run carries on. While I was there I opened concurrency to 32 lanes, since the tasks were flat anyway. That one change is the direct reason the full run later finished in a single pass. Serial, the full run would take about fifteen hours, and I would have to pray that not one bug showed up in those fifteen hours. The pilot finally finished and the ledger stopped at $1.23. Then I looked at the numbers, and my blood pressure went higher than at any of the crashes. Frontier on math, 0. Code problems across every arm, all 0. Two new bugs. GSM-Symbolic, the contamination-resistant variant set the plan uses for the math family, has answers that often carry a % suffix. My scorer ran float parsing on them, a parse failure scored 0, and frontier happened to answer in the tidiest format of all, so it was wiped out. Scoring the code problems runs evalplus, an off-the-shelf test suite for code problems, whose tests import numpy, and this project's dedicated Python environment (venv) had no numpy, so the grader quietly recorded 0 for every problem in every arm. Three hours. Five bugs. The part that stings most in review, not one bug came from "I can't write code," and not one of the 19 tests could stop them. My twenty minutes estimated the time for "the code is not wrong." It did not estimate the time for "the world does not follow my assumptions," and the real work of the execution step lives in the second kind of time. Readers who do not write code should not be shut out by this section. Swap the pipeline for your own data-gathering and accounting process, and all five bugs have counterparts there. ## 7.2 The step where it is easiest to look like you are working Good news first. In the seven-step workflow, execution carries the thickest transfer dividend. Ten years of accumulated coding practice moves over almost as is, and AI happens to be most skilled at this step too. Most of my harness was written by AI, and going from an empty directory to the whole chain running under mock mode took hours. Three years ago this was one person's week or two. The bad news lives next door. Execution is also the step where "looking like work" most easily passes for work. Logs scroll, the progress bar moves, the bill climbs, all of it the appearance of work, none of it proof that trustworthy numbers are coming out. The numpy bug is the perfect counterexample. With the dependency missing, the pipeline still ran diligently, still threw zero errors, and the code still looked elegant. Every number it produced was garbage. The core proposition of this step follows. **The credibility of execution comes from the structure of the harness, and the structure has to leave errors nowhere to hide.** "Looks right" is an aesthetic judgment, and AI's code almost always looks right. "Cannot hide" is a structural judgment. Does a crash lose data, is there an account for the spending, can the criteria still be edited, can the scoring be replayed. Section 7.4 breaks this structure into four pillars. The division line gets drawn here too. Writing the pipeline, writing the tests, writing the scaffolding, hand them off. But the kind it does not flag has to be seen clearly, and that kind is the harness itself being wrong. It does not know what terra's API contract looks like, does not know your venv is missing numpy, does not know this batch of answers carries a %. These bugs live on the seam between the code and the world, and the seam sits outside the view of whoever writes the code. Chapter 6 said the most expensive flaws in a plan are the kind AI does not flag. The execution step's counterpart, a green test suite only proves the code matches the world you defined, and the pilot is the only probe that reaches into the world's side. ## 7.3 Lessons stolen from pair programming Pairing, tests, tracer bullets, all three lessons of this step are your daily routine, so I will only cover where they carry over to a research experiment and where they break. **Lesson one, pair on execution.** The keyboard changes hands, AI writes, you review, and the bottleneck of Row 1 of the transfer map lands right here in the execution step. "Review" has to be layered, and you do not read every line along with it. Boilerplate (retries, writing to disk, command-line arguments) gets a glance and a pass. Lines that touch the criteria, the scoring function, the ledger's money formula, the resume key covered in the next section, get read word by word. The reason is cold. An error in these few places is directly an error in the number you sign off on two weeks later, and the program is the small matter. Seven tenths of the human hours I spent on the harness went into these three places, and in hindsight not one of them was wasted. Two of the five bugs, the % and the null, were caught right around this review line through scoring and parsing. Where it breaks. Pairing on production code reviews logic. Here you review the definition of right and wrong itself, and when the scoring function is wrong the tests stay green. **Lesson two, verification infrastructure sets the radius of letting go.** Chapter 6's lesson three already established this, so here I only add the execution step's increment. Ground truth at this step is unusually cheap. Tests go red or green in seconds, every ledger line adds up, and scoring can be replayed offline. That is why collaborator level is real here and local autonomy is worth discussing. The same model downgrades on purpose at the interpretation step of Chapter 8, where there is no cheap ground truth. Section 7.1 already demonstrated the other side. The infrastructure covers only the world you defined, and all five bugs lived on seams the 19 tests could not reach. Whether the infrastructure itself can be trusted waits for the pilot to test. **Lesson three, fire the tracer bullet first.** Walk the whole chain with mock first, then spend real money. This is the practice Chapter 2 stole from The Pragmatic Programmer, and at the execution step it lands as one switch, `--mock`. A fake model that always answers 42 with a fixed token count walks data loading, team topology, scoring, writing to disk and reporting through the whole chain offline. It cleared out a whole class of bugs for free, data formats, voting logic, the report pipeline, so that the first time I spent real money only the seams were left to blow up. Its boundary deserves an honest label too. Mock by definition cannot clear the bugs on the world's side, and all five are proof of what slipped through. So the tracer bullet is fired twice. Mock is the first shot, free, clearing structural bugs. The pilot is the second, a few percent of the total budget, clearing seam bugs. Skip the second shot and go straight to the full run, and the five bugs will ride along through your entire budget. ## 7.4 The four pillars of a harness Four pillars, each one matching a class of "places errors hide," a few dozen lines of code all together, and the highest-value few dozen lines of the execution step. **Pillar one, a resume key.** Every result written to disk carries a key that uniquely rebuilds it. Mine is a four-tuple (task family, problem, arm, seed). The run starts by scanning the result file for finished keys and striking them out of the task pool. The effect is that a crash mid-run drops from an accident to a lossless event. Problems already run are not rerun, money already spent is not spent again. Two things come with it. Tasks must be flattened into small independent pieces, because the mid-state of a long serial chain cannot be expressed as a key, and that was the root of my crashes in section 7.1. A single piece's failure goes onto the failed list and may not take down the run, and the resume retries it naturally. Two lines are enough to show it. ```python done = {(r["family"], r["item_id"], r["arm"], r["seed"]) for r in results} jobs = [j for j in all_jobs if key(j) not in done] ``` **Pillar two, an append-only ledger with a budget hard cap.** Every API call writes one JSONL line, a text format with one record per line, recording the timestamp, the model, tokens in and out, and dollars. Append only, never edited. The moment cumulative spending passes the $35 hard cap, throw, and the whole run stops. This pillar later paid two dividends. First, during the pilot a batch of self-consistency result lines had been run with the wrong model, before the probe selection rule covered in the next section was executed. Cleaning up, the lines in the result file were deleted and rerun, and not one ledger line moved. The principle is one sentence. Results record the current understanding, and when the understanding is corrected they should be rewritten. The ledger records history, and history does not accept edits. Second, "how much did this experiment actually cost" went from a number recalled by impression to the sum of a column. Chapter 6's sign-off promise on cost basis is redeemed right in these JSONL lines. **Pillar three, mock mode.** Lesson three in concrete form. Besides serving as the tracer bullet, it is the free regression test for every later change. Changed the topology logic? Mock reverifies the whole chain in three minutes. One calibration mark from experience, building mock cost a few dozen minutes and paid back an order of magnitude on the first afternoon. **Pillar four, the criteria enter the repo ahead of the results.** prereg.md goes into the repo ahead of any result. The harness may only implement it, never revise it. Every change after the run starts may only be appended to the change log, with the date, what changed and why. The bug where terra refused temperature 0 is the textbook disposition under this discipline. The plan's assumption about the world was wrong, the plan itself does not change, the difference goes into the change log, the harness runs by reality, and when Chapter 8 reconciles, every entry is out in the open. Why not just fix the plan? "Just this once" has no stopping condition, and the change log does. The more changes there are, the uglier that file looks on its own. The most beautiful execution of this discipline happened on the subplot, and it is worth a digression. Readers who care only about the spine case can skip to the fifth pillar. In the preregistration of the persona survey (its execution story opens up in Chapter 12), I had labeled Q120 of the WVS question bank a "competition" question from memory. While writing the data loader, following the step "a first-hand document you cite must be checked," I opened the official WVS-7 questionnaire. Q120 is the "risk of being held to account for taking a bribe" question, and the competition question is Q109. Same disposition. The locked preregistration does not change, the question number stays, the official wording replaces mine, and the difference goes into the change log. The host of this bug is worth recording. The code was not wrong, AI was not wrong, my own memory was wrong. It turned "checking" from a virtue into a step, and a step catches its own designer along with everyone else. There is a fifth pillar. It was too cheap to list on its own, and it earned its name during the bug-catching day in the next section. Write the original text of the answers to disk, and let the score be a derived column. The result file stores the extracted original answer text, not only the score. So when the scorer's bug was fixed, one replay script recomputed every score at zero cost. The grade of this pillar gets tested once more in Chapter 10. ### How readers who do not write code use these pillars The pillars are about structure, not code. Doing interviews, running questionnaires, digging through archives, checking reports, all of them need the same pillars with different materials. - A resume key. Every time you finish one minimal unit (one interview, one document, one report), write it down and number it, and after an interruption pick up from the number, redoing nothing and missing nothing. Same precondition, break the work into small independent pieces. - An append-only ledger with a budget hard cap. Record the money and the hours spent on a sheet that only grows and never changes, lock in a ceiling before you start, and stop at the line, with no "just a little more." - Mock mode. Walk the whole process on one set of fake material, from gathering to scoring to aggregating, confirm every step connects, and only then touch the real material. After that, every time you change the process, walk it again on the fake material first. - Criteria leave a trace ahead of results. Once the criteria are written, send yourself or a colleague a timestamped email, or save them into a document with version history. Later changes append a note, they do not overwrite. - Raw records go to disk. Keep the recordings, the source excerpts, the screenshots, and derive the scoring and coding from them. When the scoring rule changes, recompute from them instead of collecting again. ## 7.5 Spine update Β· the five bugs go for the strongest arm Now an autopsy report for those three hours. Five bugs, each one by where it landed. | # | Bug | The seam it hid in | Symptom | Who zeroes out first | |---|---|---|---|---| | 1 | The reasoning model's API contract | plan ↔ API | temperature 0 refused, a 400 error | The frontier arm, not one problem run | | 2 | Truncated thinking returns null | token budget ↔ reasoning model | content is null, the harness crashes | The strongest reasoning member of the army | | 3 | A 200 status code with a non-JSON response | provider ↔ client | the status says success, the body is garbage | Whoever it hits (the only one whose landing point is random) | | 4 | The % suffix scoring miscarriage | scorer ↔ data | correct answers carrying % scored 0 | Frontier's math all zero | | 5 | The missing numpy dependency | grader ↔ runtime environment | code scoring silently all 0 | The code set zeroed out across every arm | Read the landing column straight down and the pattern shows itself. **Bugs zero out the strongest arm first.** The mechanism is not mysterious. The strongest arm uses a reasoning model, whose contract is the most special and whose thinking burns the most tokens. Its answers are also the tidiest, carrying units and percent signs, feeding exactly into the parser's most fragile path. The grader it depends on is the heaviest too, since it has to actually run tests, so its dependency chain is the longest. The more capable it is, the more seams it has, and the higher the density of bugs. The weak arm is safer. A model that was going to answer wrong anyway cannot be wronged by much. The corollary is ugly. **If your harness has bugs and you did not catch them, the most likely direction of bias in your readings is systematically wronging the strongest contestant.** Then you get a surprising "the weak beat the strong" result, exactly the kind you wanted most, and you write it into the report full of excitement. What saved me this time was absurdity. Frontier scoring 0 on math cannot be true. But bugs make no promise of absurdity. A bug that chews a whole arm down to zero screams on its own. A bug that only chews off 3 points says nothing, and 3 points is already enough to flip a verdict at Ξ΅=2pp, that is, two percentage points. So the alarm has to come from process. Go through the pilot numbers arm by arm, plus one dedicated question. Does the strongest arm's performance make sense? The repair process verified the fifth pillar along the way. The fix for the two scoring bugs (the % and numpy) was to change the scorer and run the replay script, rescoring every historical answer for free. The 38 null answers from the truncation era were deleted and rerun. The pilot has one more piece of preregistered business. The change log of the Chapter 6 plan locked in a rule, the self-consistency arm uses "the strongest army member on a 20-problem probe," a harder control against the objection that teaming is just multi-sampling. The probe readings came in. Of the three candidates gpt-oss-20b was highest at 0.867, ministral-14b 0.850, qwen3.5-9b 0.533. The rule executed, the SC arm locked to gpt-oss-20b, and the readings and the choice both went into the change log. The probe brought back something else. qwen's 0.533 incidentally exposed its chronic illness. With too low a token ceiling it thinks past the timeout, and under a 4096 ceiling it gave no answer on 28 of the 60 probe problems. The ceiling went to 8192, and one rule was locked in on the spot, any further timeout gets reported honestly as a property of that member, with no more patching. Even "when do you stop tuning" has to be agreed in advance. Then the full run. Five bugs fixed, the mock regression run once more, 32 lanes of concurrency started. Three numbers at the end, 3,580 result lines, $5.58 in the ledger (hard cap $35), 0 failures. Three hours of crashing bought a full run that passed in one go, and that trade comes out ahead however you count it. The full run's numbers also held a surprise that made my heart race, and this chapter's five bugs are the whole reason I later dared to face it in the right posture. Chapter 8 opens it up. ## 7.6 Swap in your project In Chapter 6 you gave your hypothesis a plan. Before you start, give it a harness, and accept it with this minimal checklist, where every item matches a class of real accident from this chapter. Budget half a day. 1. **Fire the mock tracer bullet first**. Build a fake data source or fake model and walk the whole chain of "input β†’ processing β†’ scoring β†’ writing to disk β†’ aggregation" offline. Every break you hit is a break you would otherwise have paid real money to find; 2. **Run a crash drill**. Halfway through a real run (or a mock run), kill it by hand and restart. Check that results are neither duplicated nor missing, count the lines, reconcile the ledger. If you cannot, add the resume key before you start. The version for people who do not write code, close the laptop halfway through and see whether the records alone let you resume the next day; 3. **Blow the fuse once**. Set the budget hard cap to nearly zero, run, and watch it abort with your own eyes. A fuse that has never blown is not a fuse. Chapter 6's discipline of "watch it go red first" applies to infrastructure too. The version for people who do not write code, set the ceiling at a number you will hit within the hour and see whether you actually stop; 4. **Check the timestamps**. The commit of the criteria file must be earlier than the first line of results. If it is later, you already know which chapter to go back to; 5. **Plant a poison pill**. Deliberately drop in a task that must fail, and confirm it only goes onto the failed list and does not take down the run. The version for people who do not write code, mix in a piece of material you know is bad, say a fabricated citation, and see whether the process flags it on its own; 6. **Run the pilot on 3-5% of the total budget**. Look at the numbers arm by arm and configuration by configuration, spot-check the original text of the answers behind every 0 and every perfect score (the faithful extract, meaning the answer pulled out of the model's reply as is, without passing through the scorer), and finish with that sentence. Does the strongest configuration's performance make sense? This chapter's checklist, design patterns and audit prompt are in the appendix in full fillable form. **Want an agent to run it with you?** Paste the block below into Claude Code, Codex, or any coding agent: ```text In code/smol-army of the research-rewritten repo, help me run Chapter 7's Swap in your project. First run uv run python -m smol_army.run --mock to fire a mock tracer bullet, and show me the whole chain's output verbatim. Then run the crash drill, kill the mock run halfway and restart it, count the lines, reconcile the ledger, and tell me whether results are duplicated or missing. Set the budget hard cap in config/run.toml to nearly zero and run once so I can watch the fuse blow with my own eyes, then change it back. Check the timestamps again, is the commit of docs/prereg.md earlier than the first line of results in results/. All of that is practice on the book's harness. My own project's harness I accept myself, against the checklist of Template 1 in docs/appendices/ch07-templates.md, with Template 3's "have AI audit the harness" prompt run in a separate session. I sign the conclusion. If any command errors, stop and show me the output. ``` ## 7.7 Sober reminders - **Taking "it is running" for "it is producing"** is this step's number one illusion. Logs, progress bars and the bill are all moving, and that does not make the numbers trustworthy. There is only one test, can every number point to the criteria, the ledger and the original answer text (the faithful extract). - **The false safety of an all-green test suite** comes second. Green only proves the code matches the world I defined, and all five bugs lived on seams outside that definition. The only test for a seam is a pilot that spends real money. Budget for it, do not save on it. - **The temptation to edit while running** is fiercest during execution. Every friction between the plan and reality argues for fixing the plan while you are in there. The rule again, the harness obeys reality and the document obeys discipline. Any change made after you have seen the results is downgraded to exploratory. - **Running bare with no ledger**. In an experiment with no ledger the cost numbers rest on later recall, and recall always leans toward making your method look cheap. Put that into a cost-matched comparison and it already counts as measurement failure. - **Dating this**. Every number in this chapter will expire, model versions, list prices and probe readings alike. The four pillars and the five classes of bug will not. - **Still exploring**. Frameworks that let an agent run experiments autonomously for long stretches and find and fix harness defects on its own are iterating fast. As of this writing, my observation is that they catch syntax bugs and logic bugs faster than people, and there is no evidence they catch seam bugs reliably. Contracts, dependencies, data quirks, exactly where all five bugs of this chapter are registered. The book's online case library tracks this front. ## 7.8 The unfair advantage you now hold Before any experiment that costs money and days, you can walk the whole chain for free under mock mode, turn "crashing mid-run" into a lossless event and "blowing the budget" into an impossible one. Most people only find out on their own three-hour afternoon that these are everybody's floor, having treated them until then as a luxury for the cautious. --- # Chapter 8 Β· Read the Results, Catch the Errors !!! info "Chapter companion" πŸ“‹ [Chapter 8 templates](../appendices/ch08-templates.md) Β· πŸ—‚ [Template index](../appendices/template-index.md) Β· πŸ’» [`code/smol-army`](https://github.com/hallieren/research-rewritten/tree/main/code/smol-army/) Β· πŸ’» [`code/persona-panel`](https://github.com/hallieren/research-rewritten/tree/main/code/persona-panel/) > **This chapter's ladder.** At this step AI downgrades on purpose. Batch splitting, recomputing, pattern scanning can be handed off. Those jobs have cheap ground truth, a recomputed number comes out right or wrong. The ruling on "does this count as evidence" is yours. The reason was written in the Chapter 3 snapshot, interpretation is sycophancy's home ground, you ask with excitement and it answers along with the excitement. Let it sit as a juror. The judge's seat is not for it. > > **Spine update.** The Chapter 7 harness finished the full run. In this chapter you will see the most dopamine-rich line of numbers in the whole project, and how within twenty-four hours it went from +18 to βˆ’2.3. > > **This chapter delivers.** The result interrogation checklist, the preregistered/post-hoc side-by-side report template, the surprise-result red flags. --- ## 8.1 The verdict column reads army_ahead The evening of July 25, 2026, the full run closed out. 3,580 result lines, $5.58 in the ledger, 0 failures. How the harness was built and the five bugs the pilot caught, Chapter 7 covered. I opened results/report.md and went straight to the paired bootstrap table. Paired bootstrap, paired resampling, used to put an interval on the gap between two arms over the same set of problems. The code row, the army 93.7%, frontier 96.0%, CI crossing zero. CI, confidence interval, the range the gap could fall in, and crossing zero means that range contains 0. Respectable and expected. The mmlu_pro row, the army 15 percentage points behind, CI [βˆ’21.3, βˆ’8.7]. The skeptics' script, also within the odds. Then math. Army vote 0.856, frontier 0.673. A gap of +18.2 percentage points. The raw difference is 18.3, this book uses the bootstrap point estimate throughout. CI [+12.7, +24.2], and the verdict column, the column in the results table where the ruling is written, reads **army_ahead**. What went through my head in those minutes, I can report faithfully. In Chapter 4 the private odds I gave myself were "narrow tasks have a shot," and now a narrow task had come in person to collect. This one line of numbers was enough for the cover story of the whole book, enough for the opening paragraph of the Chapter 9 memo. I wanted to send it to every colleague who had poured cold water on it, and I even started thinking about what to call this chapter. I wanted to announce. That urge, by itself, is the first signal this chapter teaches. Hitting the brakes was no virtue. This line of numbers was good beyond bounds. The same table held two more things that made no sense. First, the cost-matched self-consistency arm also beat frontier on math, by six points. If a 20B small model can overtake frontier on multi-sampling alone, the story should not be "team magic." Second, frontier scored only 0.673 on grade-school probability problems. A frontier model cannot do a third of grade-school math problems? Too good to be true and too bad to be true, on the same table. The evening the results arrive is the most dangerous moment in the whole workflow. ## 8.2 The brake rule and the three interrogation questions First see why this step is dangerous. Motivation is at its peak, weeks of work are waiting for this line of numbers to pay out, the degrees of freedom are still wide, how to split, how to tell it, which slice to stress are all undecided, and the error is silent, a wrong conclusion and a right one look identical in a report. The criteria locked in Chapter 6 govern "what counts as a win." They do not govern how fast you run out the door holding the word win. The way AI rewrites this step has the same shape as the earlier steps, only the labor collapses in price. Splitting the results by problem, clustering the wrong examples, recomputing the statistics under a different assumption, these interrogation moves used to cost so much that you did them only when a reviewer, or the checker before delivery, forced you to. Dispatch them to AI now and the full set comes back in an hour. The guilt of not interrogating went up, not down. Not interrogating used to have an excuse, interrogation was too expensive. Now interrogation is cheap enough that the excuses run out. The ruling half is another matter. The three variables of Chapter 3 all press toward the human here. No machine can rule on whether a number is qualified to carry a conclusion, a wrong ruling raises no error, and retracting an announced conclusion is priced in reputation. So the ladder at this step falls rather than rises. Splitting and recomputing are handed off. The judge's seat is taken back. The brake rule is one sentence, written into the process. **Any result that makes you want to announce it at once goes through the interrogation before you announce it.** You do not need to feel something is wrong to interrogate. Noticing that you want to announce is enough. The interrogation is three questions. **Question one, can the scorer be trusted?** Do not stop at the metric. Pull out the original text of the answers scored wrong and scored right and look. What do the wrong answers look like? Messy wrongness looks like a normal capability boundary. Tidy wrongness looks like the scorer or the gold answer itself being sick. All five bugs of Chapter 7 were this question's prey. You will see shortly that this question can dig deeper than the harness. **Question two, what does the data look like?** Open the problem text itself. How many kinds of problem are there? Independent of each other, or batch variants from one mold? Your CI was computed under the assumption "samples are independent," and what that assumption is worth you only know after seeing the raw data. **Question three, where is the win concentrated?** Spread the gap out by problem and by slice. Is the win spread evenly, or concentrated in a small handful of problems? Concentration is not a crime by itself, but it means your conclusion hangs on the quality of that handful, and they deserve a separate interrogation. All three are jobs AI can be handed. How you dispatch decides the quality. Do not ask "is this +18 credible," that hands the judge's seat to sycophancy, the tendency to flatter in the direction of the question, which Chapter 11 will dissect. Dispatch mechanical work. "List the original text of every frontier wrong answer, the matching gold answer, and the numerical relation between the two, ordered by problem id." Let it sit as a juror and lay out the facts. "Does this count as evidence," you rule. Before the interrogation begins, one more thing has to be set up, a **reporting discipline**. Report the preregistered numbers as is. Report the post-hoc breakdowns the interrogation produces side by side, labeled "post-hoc." At no time may a post-hoc number replace a preregistered one. This discipline must stand before the work starts, or there is no line between "interrogate" and "interrogate until I like it." ## 8.3 Spine update Β· the twenty-four hours from +18 to βˆ’2.3 Now I run the three questions in front of you. Every number below can be recomputed in the smol-army repo (scripts/audit_math.py, matching the 2026-07-25 post-hoc audit entry in docs/CHANGES.md). The audit itself is outside the preregistered scope and is labeled post-hoc by the discipline. **Question one, executed.** I had AI pull out every frontier wrong answer on math, 49 in all. The first pattern surfaced on the spot. All 49 wrong answers fell in math-100 to math-149, one continuous stretch of the 150 problems. The second pattern sent a chill down my back. **Every wrong answer was exactly 4 times the gold answer, with a percent sign attached.** 49 times, no exception. A model that "cannot do the problem" does not look like this. It was doing a different problem, with extreme consistency. Question two, executed. Open the problem text. math-100 to 149 are 50 variants of one GSM-Symbolic probability template. That benchmark generates different versions of the same problem from symbolic templates, built to detect fragile mathematical reasoning and data contamination (Mirzadeh et al., Apple, 2024, arXiv:2410.05229). The question reads "how much more likely...(as a percentage)". That English has two readings. The gold answer reads it as the **absolute percentage-point difference** (p₁ βˆ’ pβ‚‚). Frontier answered the **relative increase** every time ((p₁ βˆ’ pβ‚‚)/pβ‚‚). The base probability in this template is 1/4 in every variant, so the relative reading is always exactly 4 times the absolute one. All 49 "wrong answers" are 4 times with a percent sign, and that is the whole solution, **the question itself is ambiguous, the math is not wrong**. On these 50 problems frontier matched the gold answer's reading on only 1. It had firmly chosen the other, perfectly defensible, English reading. Question three, executed. The win is entirely concentrated on the ambiguous template. On these 50 problems, army vote scored 0.61, self-consistency 0.23, frontier 0.02. The distribution of readings of one ambiguous English sentence across three model lineages happened to tip the majority toward the gold answer's side. Multi-lineage voting really does have an advantage, but the advantage is in guessing the reading. Reasoning gets no credit. Remove the ambiguous template and recompute. Frontier 1.000, the army 0.977, self-consistency 0.983. The paired bootstrap gives the army **βˆ’2.3 percentage points, CI [βˆ’4.0, βˆ’0.7], direction reversed**. There is one more quiet but important comparison. On the clean subset the army's 0.977 against self-consistency's 0.983, teaming is about equal to multi-sampling. The steelman arm that Chapter 6 nearly left out earned its whole wage right here. The theme sentence of Chapter 7's five bugs was "bugs zero out the strongest arm first." The sixth hid deeper, in the data's gold answers. The harness was innocent, the code all correct. This bug zeroed out nobody either. It handed me 18 points. **A bug that zeroes an arm, you will run into sooner or later in an error message. A bug that hands you points, only the interrogation catches.** Let me say the ugly part here first. This round of post-hoc breakdown is itself a new batch of numbers, and the same discipline applies. The Chapter 10 red team will come back for it, and will find something. ## 8.4 150 problems that are not 150 problems Question two had one more aggravating finding. Clicking through the problem text family by family, the 150 math problems are really only about **3 template families**, each with 50 variants in sequence. That family count was clicked out problem by problem, not estimated backward from variance, and "about" is there only because the family boundaries were drawn by eye. The sampling script took the first 150 rows of the dataset without shuffling. The damage to the statistics is structural. The 50 variants of one template are highly correlated. Template unambiguous, the whole family is right together. Template ambiguous, the whole family is wrong together. The preregistered bootstrap resampled them as 150 independent samples, so the CI it computed is **overconfident**. That respectable narrow interval [+12.7, +24.2] bought its narrowness with the assumption "problems are independent." The effective sample size does not reach 150. It is on the order of those roughly 3 template families that were clicked out. This pit goes on my own account. Chapter 5 chose a "contamination-resistant variant set" to guard against memorized problems, but a variant set is by nature copies of a few templates. It blocked one contamination and introduced one correlation. Not shuffling at sampling time pushed that correlation to its maximum. The Chapter 6 preregistration locks in motive. It cannot lock in ignorance. Locking the criteria first guarantees I cannot pick data afterward. It does not guarantee the data I picked beforehand is healthy. For this class of error, the only detection mechanism is interrogation question two, seeing with your own eyes what the raw data looks like. Now produce the deliverable by the reporting discipline. In the table, pp means percentage points. | Task family | Preregistered result (reported as is) | Post-hoc breakdown (labeled post-hoc) | |---|---|---| | math | Army vote +18.2pp [+12.7, +24.2], army_ahead | After removing the ambiguous template, the army βˆ’2.3pp [βˆ’4.0, βˆ’0.7], direction reversed; the army 0.977 β‰ˆ self-consistency (SC) 0.983; this family has only ~3 template families, the preregistered CI is overconfident | | mmlu_pro | The army βˆ’15pp [βˆ’21.3, βˆ’8.7], army_behind | No overturn. The SC arm βˆ’18pp, the fault lies in the small models' knowledge base, not in teaming | | code | The army 94% vs frontier 96%, CI crossing zero; the army costs 1/5 of frontier | No overturn. The SC arm likewise level | Last, H goes through the interrogation. The preregistered falsification conditions were two. (a) The army trails by more than 5 percentage points on all three families. Not triggered, code's CI crosses zero. (b) The army catches up only at more than 2Γ— the cost. Not triggered either, the army's bill is a fifth of frontier's on code, under half on mmlu, about equal to frontier on math, and no family bought a tie with money. **So H was not falsified, but neither did it hold across the board.** The landing point is the question Chapter 5 planted, **which tasks have a shot**. Knowledge Q&A has none, teaming cannot rescue a weak knowledge base. On math, trust neither side, this exam has to be reissued first, with a template-shuffled problem set that has enough families. Code has a shot, and that is the sturdiest sentence in the whole case. Read strictly by the preregistration, a gap has to be within Β±2pp to count as a tie, and code has only 100 problems, so the interval cannot be squeezed that narrow. It can only count as "undecided." As a decision-grade conclusion, directionally no frontier advantage is visible, and the bill is hard. How that sentence is written into a paper's limitations, and into the first line of the CTO memo, is Chapter 9's business. The two readings side by side are the table below. They do not contradict each other. They answer two different questions. | Reading | code | mmlu_pro | math | |---|---|---|---| | Preregistered reading, Ξ΅=2pp, per family | Undecided, CI crosses zero | 15 percentage points behind, a clean loss | The "overtake" died in the interrogation, the problem set has to be reissued | | Decision-grade reading | Directionally no frontier advantage visible, at a fifth of the price | No shot | Trust neither side | ## 8.5 Subplot mirror Β· FAIL gets interrogated too The persona study's verdict is in. All three criteria failed outright, the preregistered falsification shape was reached, and "persona can replace real interviews" is dead. The full table of numbers has to wait for Chapter 12. Here it first fills in one lesson for this chapter, **when "like a real person" counts as evidence, and why a falsified conclusion gets interrogated all the same**. Readers who care only about the spine case can skip to the next section and take one sentence with them, the interrogation does not depend on direction. The first thing takes one sentence. A single answer that "sounds like" a person never counts as evidence. Fluency is the model's default property, with zero correlation to distributional fidelity (the first item on the list when Chapter 11 does the subplot's case review, the fluency illusion). The only thing that counts as evidence is a match at the distribution level against criteria locked in beforehand. If the distribution distances for all six subgroups had cleared the thresholds, variance had not collapsed, and the cross-slices had all been green, at that moment "like a real person" would rise to evidence, and the scope of that rise would reach only as far as the question bank's domain. The persona case is a long way from that moment. The second thing is the point of the mirror. For this book, FAIL is a **desirable result**. It confirms Chapter 11's average face prediction beautifully, that is, persona answers are the averaged-out result of a crowd of people smoothed together, more homogeneous than real people, and narratively you could not ask for better. By the symmetry discipline, the more desirable, the more it gets interrogated. Walk the three questions. The scorer, all three criteria are computable statistics, no LLM judge, program rerunnable. The data, 10,800 interviews, zero invalid answers. Where is the loss concentrated, and the answer is that it is not. All six subgroups fail, variance collapsed on eight in ten clean questions, and no slice can carry the blame alone. **The more spread out the loss, the sturdier the falsification.** Only after this pass does FAIL qualify for the Chapter 9 deliverables. Only one ambiguity remains. In the contamination test, 10 of 15 questions were flagged, the answer distributions on the original question and the reworded one clearly disagree. Two readings. Persona memorized the original questions, or it is highly sensitive to wording (a close relative of sycophancy drift, and sycophancy drift is the sycophancy described earlier). The single-source design cannot tell these apart. Single-source means the real-interview arm was cut before the run started, leaving only the public question bank as a comparison, and the cost of cutting the real arm is booked in Chapter 12. The shape of the interpretation discipline here is **do not force a single reading**. Report carrying both readings, and check whether the conclusion stands under each. If it is memorization, what was memorized is the real people's question bank, and the effect can only push persona's distribution toward the real people's side, which makes FAIL only more conservative and cannot rescue persona. If it is wording sensitivity, that is itself another kind of distortion, and it supports "cannot replace" just the same. The falsification survives under both readings. Forcing one interpretation would be the real defect. ## 8.6 Lessons stolen from debugging When a result surprises you, suspect your own code first. That is your muscle memory. Three lessons move over as is, and below I look only at where they break when moved onto interpretation. **Lesson one, an unexpected victory, suspect the scorer first, the world second.** "'select' Isn't Broken" (The Pragmatic Programmer, first edition Tip 26, twentieth anniversary edition Tip #33). When you think the compiler is broken, it is almost always your own code that is broken. The interpretation version is identical. The data tells you "the frontier model cannot do grade-school probability," and there are two explanations on the table. The world turned over, or your gold answer is wrong. The second is boring, but it is cheap and common. The size of the surprise should be proportional to the strength of your suspicion. Most people make it proportional to the strength of their excitement. Where it breaks, in code, suspecting yourself has a stack trace to help you locate it. Here, suspecting yourself means pulling out the original text of the answers and looking with your eyes, so interrogation question one must be written into the process. **Lesson two, a green light proves the tests passed, not that the tests are right.** Row 5 of the transfer map says verification infrastructure sets the radius of letting go. This chapter adds its dark side. **When the infrastructure itself is wrong it raises no alarm, and the whole pipeline runs green all the way to the wrong answer.** Chapter 7's five bugs were all the harness's fault, and tests and the ledger could catch them. The sixth bug was in the gold answers. Every line of the harness was right, every point was scored correctly, and what was wrong was the definition of "right." So the interrogation has to go all the way down to the problem text itself. Locking the criteria first locks in your degrees of freedom. It cannot lock out the error buried in the data. **Lesson three, excitement is the least trustworthy gauge.** Row 3 of the transfer map, developers rated themselves 20% faster and measured 19% slower, self-perception decoupled from fact, calibration comes from measurement (Chapter 2 covered it, not repeated). The interpretation step's counterpart is the Chapter 11 law, the more desirable the result, the stronger the sense of mastery, the less checking. The brake rule's trigger is written as "wanting to announce" because the self-perception of that moment is the least trustworthy gauge reading in the whole workflow. ## 8.7 Swap in your project Dig out your most recent "want to announce" result. Last week's experiment, that good-looking curve, the p-value that finally came out significant, any of them. Give it the interrogation it missed, budget two hours. 1. **Write down the announcement sentence.** The sentence you originally wanted to say out loud, written down without changing a word, and pinned beside you. When the interrogation ends it is either alive or dead. First give the interrogation an autopsy subject; 2. **Question one.** Pull out the original text of the key wrong and right examples that support the sentence (dispatch it to AI, with none of your expectations in the brief, template in the appendix). Look for a pattern. Messy wrongness or tidy wrongness? 3. **Question two.** Count your effective sample size, and do one thing based on the count. The counting is short. First write down the n in your report, then ask "how many groups among these n share one source," such as one template, one batch of cells, one crawl, one surveyed institution. Whatever shares a source counts as one independent unit. Counted, act on it. **If the number of effective units is below the reported n, write the number of effective units into the body text** (not a footnote), in the format "n=150, 3 independent units". Keep the original interval as is, and add one sentence, "this interval was computed under the assumption that problems are independent, and at this number of independent units that assumption does not hold." Do not quietly swap in a wider interval. Recomputing the interval is next round's job. This round's move is **to turn the assumption from implicit into explicit**. This step goes only this far. What comes after (a mixed-effects model? average within groups first?) depends on your field's conventions and your reviewers' standards, and this book cannot give a general prescription. "Write n and the number of independent units side by side" holds in every field, and it is enough to let whoever reads your conclusion reprice it; 4. **Question three.** Spread the effect out by slice and look at concentration. The small handful of samples the effect concentrates in gets its own quality pass; 5. **Side-by-side report.** Report the numbers under the original criteria as is, report the breakdowns the interrogation produced side by side, labeled "post-hoc"; 6. **Closing self-test.** How is the announcement sentence now? Alive as is, alive after narrowing, dead. All three are qualified outputs, and the second is the most common. Only one is unqualified, sent out without the interrogation. This chapter's three tools, the result interrogation checklist, the preregistered/post-hoc side-by-side report template, the surprise-result red flags, are in the appendix in full fillable form. **Want an agent to run it with you?** Paste this to your AI assistant or coding agent: ```text Help me put a result through the interrogation, the Swap in your project of Chapter 8. I will first write you the announcement sentence without changing a word, and you paste it as is at the top of every reply. Question one, you follow the instructions in Template 1 of docs/appendices/ch08-templates.md to pull out the original text of the key wrong and right examples, with none of my expectations in the brief, and whether the wrongness is messy or tidy is for me to say after I have read the originals. Question two, the n in my report and the groups sharing a source are mine to count, you only write it into the body text in the format "n=how many, how many independent units", and you may not quietly swap the interval. Question three, you spread the effect out by slice, and where it concentrates is my ruling. The ruling field is signed by a human, you may not sign on my behalf. Finally the side-by-side report, the numbers under the original criteria as is, the interrogation breakdowns labeled "post-hoc". If any command errors, stop and show me the output. ``` ## 8.8 Sober reminders - **Not interrogating before you announce** is the number one failure mode of this step, and it dies quietly. The +18 flows on into the abstract, the weekly report, the next round's budget request, gaining value at every stop, until some stranger runs the audit for you. You have only two choices, be the first auditor yourself, or wait for someone else to be. - **Replacing preregistered numbers with post-hoc ones is worse than not interrogating**, because it wears the clothes of rigor. βˆ’2.3 cannot replace +18 as "the result of this experiment." It can only stand beside it, labeled post-hoc, and only the next round's experiment is entitled to retest it confirmatorily. The moment you replace, the interrogation degrades into another round of picking data, Chapter 6's degrees of freedom and motive colluding again, with a different moment to strike. - **Interrogating only the undesirable results** is the most hidden kind. The interrogation is triggered by surprise and by stakes, and direction is not in the trigger conditions. The honest account for this case, mmlu's βˆ’15 followed the literature's expectations, and I really did ask it fewer questions than I asked the +18. That asymmetry goes into the limitations. When you run your own interrogation, check once, did the favorable results and the unfavorable ones go through the same checklist? - **Still exploring**. Tools that let AI audit results end to end on its own are iterating. As of this writing, dispatched mechanical interrogation, pulling examples, clustering, recomputing, can be handed off. The nose for "what to be suspicious of" is not there yet, and which item on the interrogation checklist to run first is still the human's job. The book's online case library tracks it. Generalizing this interrogation into a systematic attack procedure against any conclusion, attacking in turn the scorer, the data composition, the independence assumption, the cost basis, is Chapter 10's business. ## 8.9 The unfair advantage you now hold For any result that excites you enough to want to announce it at once, you hold an interrogation checklist that runs before the announcement, three questions, two hours. In this chapter it interrogated a +18pp "victory" until the direction flipped, and later it did not spare its own output either (Chapter 10). The people without this checklist are, right now, ordering champagne on their own +18pp. --- --- # Chapter 9 Β· Delivery !!! info "Chapter companion" πŸ“‹ [Chapter 9 templates](../appendices/ch09-templates.md) Β· πŸ—‚ [Template index](../appendices/template-index.md) Β· πŸ’» [`code/smol-army`](https://github.com/hallieren/research-rewritten/tree/main/code/smol-army/) > **This chapter's ladder.** At the delivery step AI sits steadily at **assistant level**, and this is one of the steps whose price collapsed hardest in the whole book. First drafts, audience rewrites, figure captions, all can be handed off. There is no half-open door here. **Responsibility for claim strength and wording cannot be delegated.** One sentence gets pulled out and quoted on its own. Do you stand behind it? No model can answer that for you. > > **Spine update.** The numbers have been interrogated. In this chapter I load the same batch of evidence into two completely different heads, a technical-report skeleton and a one-page CTO memo, and every number in the two documents has to interlock and point back to the repo. > > **This chapter delivers.** A claims list, which is the single source of truth. Templates for the technical-report skeleton and the one-page memo. Plus a prompt set for audience rewriting and the number interlock check. --- ## 9.1 A twelve-page document, dead in thirty seconds Friday evening. You turn three weeks of experiments into a twelve-page document. Method, every results table, sensitivity analysis, plus four appendix links. You send it to the CTO and copy the team. This feels like the easiest step of the three weeks. The work is done, and all that is left is writing it up. Monday morning, the CTO replies with one line. "So, should we switch or not?" Your first reaction is that this is unfair. The answer is on page 7, table 3, spelled out. The second reaction is the right one. He never reached page 7, and probably did not finish page 1. That is not his fault. His job is to make a decision in fifteen minutes, and reading your twelve pages was never in his job description. Your document did not answer his question. It dumped the raw material for answering it on his desk. The academic mirror takes one sentence. A reviewer asks "why did you not control for variable X," the answer is in appendix C, and reviewers never read appendix C. Whichever page the objection rises on, the answer has to be buried on that same page. One page late counts as unwritten. Two scenes, one structure. **The quality of the evidence and the quality of the delivery are two independent variables.** Three weeks of evidence can die inside thirty seconds of reading, and die silently. Nobody will tell you "your conclusion was right, I just never got to it." The delivery step loads evidence into the audience's decision loop. Writing down what you did does not finish it. Different audiences, different loop shapes. One batch of evidence should almost never have only one vehicle. ## 9.2 Evidence does not speak for itself Start with what got cheap, and this time it got cheap all the way down. First drafts got cheap. The blank page used to be delivery's first wall. Going from a claims list to a structurally complete first draft is now a matter of minutes. Audience rewriting got cheap. This is the chapter's real new dividend. "Write another version for management" used to mean half a day. Now the same batch of evidence gets rewritten into a paper, a memo, a slide script, an email summary, in batch. "Reorder the detail by audience" went from luxury to default move. Then the part that did not get cheap, and it decides why this step's ladder stops at assistant level. **Which sentence you dare sign has not dropped a cent in price.** The signature test has a plain definition. A sentence leaves your document, gets quoted on its own, and travels with your name on it. Do you stand behind it? "Proved," "looks like," "undecided," every choice among the three strength tiers is a signature. Row 1 of the transfer map reads like this at this step. The production cost of text collapsed, the review cost of claims did not. AI also brings this step a new risk of its own. **Claim strength drifts quietly upward during drafting and polishing.** The shortest path to fluency is confident wording. "The results show" reads far smoother than "under condition X it looks like." Qualifiers shed a layer with every round of AI rewriting. Chapter 4 said secondhand retelling drifts. Here is the home-brewed version of it. The model did not lie. This is a legitimate side effect of optimizing wording, and it happens to land on the step where errors are silent. So this chapter's workflow has exactly one core move. **Lock the facts first, then let go of the wording.** While you are here, give the step an operational definition so it can be accepted. In a deliverable that passes, the audience arrives with their own question and hits the answer on the first screen. Any number gets challenged, and you point to its source within fifteen seconds. ## 9.3 Lessons stolen from PR descriptions You write three kinds of text for one diff every day, the commit message, the PR description, the release note. Three lessons carry straight over. The only thing to watch is where they break on research delivery. **Lesson one, a vehicle can wear many faces, a fact gets only one body.** The three texts differ in detail, terminology, and length, and they all point at the same diff. The diff is the single source of truth. Research's counterpart sits right there. The report, the memo, and the slides can look completely different, but every number has to point at the same results file in the same repo. The counterexample is the norm. One number in the report, another in the slides, a third in the memo, all three "roughly right." Three roughly right numbers add up to a system with zero credibility. Here is where it breaks. The diff is machine-generated, so the three texts cannot lie against it. A results file will not jump up and reconcile itself, which is why the interlock check in step 4 of section 9.4 has to be dispatched on purpose. **Lesson two, the most valuable field in a PR description is "what this change does not do."** Seasoned engineers state boundaries and known defects up front, for a cold reason. A reviewer who finds a pit you did not declare costs you ten times as much, and from that moment he rereads all your code with "what else is hidden" in his eyes. Research's counterpart is limitations, the section of a paper that states its own limits. Do not treat it as a humility ritual. It is a fight over who speaks first. A defect you declare is called rigor. A defect someone else finds is called an incident. **Lesson three, once documents generate in one click, "written well" stops being evidence of "done right."** Row 6 of the transfer map, lowering the bar to operate is lowering the bar to abuse. AI turned "looking rigorous" into a single generation. A fully formatted limitations section, a respectable confidence interval, modest restrained wording, all of it can decouple completely from the actual evidence. The coding side's answer is to grow documents out of the code. Interface docs come from types, changing the code changes the docs, and decoupling becomes mechanically hard. Research's counterpart, **no number may be typed by hand**. Every number in both vehicles is quoted from the results file, then interlocked against the other. What readers take as a signal is changing generations too. A number that can point to a repo is replacing "properly written" as the new mark of trust. ## 9.4 One source of truth, two vehicles Here is the workflow you can copy straight out. Five steps, half a day. **Step 1, write the claims list, not the document.** One line per claim, four fields. The claim in one sentence, the evidence pointer (file/table/commit), the honesty tier (verified / still exploring / falsified), and the signature field (dare / do not dare). There is a trick to writing claims. Write **the strongest form you dare sign**. Turning "seems somewhat useful" into a conclusion is cowardice. Write it up to the point where one more notch of strength would stop you from signing. On the same batch of evidence, "tie" has at least four ways to be written. - "Proved a tie" - "Looks like a tie" - "Undecided, but no gap shows in direction" - "A tie under these three premises" Which tier you pick is the signature itself. This list is the single source of truth, and both vehicles get generated from it. **Step 2, vehicle one, the technical-report skeleton.** The audience is peers and the technical committee, and their question is "how do you know." This skeleton is shaped like a workshop paper. Question, method, results, limitations, reproducibility. Anyone submitting can expand it directly. Detail tilts toward method and statistics, with criteria, arm design, and test methods given in full. Honesty layering, which is to say the tier column in the previous step's list, lands in this vehicle as the limitations section, the section that states your own limits, and it has to be written for real. Every item the Chapter 8 interrogation turned up, template concentration, the post-hoc breakdown, the overconfident confidence interval, goes in, written with numbers, not in camouflage wording like "certain limitations may exist." **Step 3, vehicle two, the one-page memo.** The audience is decision makers, and their question is "what should I do." The structure is fixed. The one-sentence conclusion on top, three lines of risk right behind it, then a number table of at most five rows, one next-step recommendation, and one line of evidence pointers. Decision makers do not read the limitations section, so honesty layering lands in this vehicle as those three risk lines. The same fact changes shape across the two vehicles. The template concentration problem in the math problem set is a paragraph of statistical discussion with numbers in the report, and one line in the memo, "the math evidence is void, do not cite it." **Tune the depth of the layering to the audience. Never tune its honesty.** **Step 4, the interlock check.** Dispatch AI to pull every number out of both documents and check them one by one against the results file. Any disagreement among the three means at least one of them is lying. This is pure mechanical work, so hand it off. The prompt is in the appendix. **Step 5, close with the signature test.** Go through both documents sentence by sentence and ask of each one, if this sentence gets pulled out on its own with my name on it, do I stand behind it? A sentence you dare not sign has two roads only, lower the strength or delete it. This is the one step in the whole flow that may not be outsourced. Charts get their own item, and both vehicles use it. Before drawing any chart, answer one question. What judgment do you want the reader to make? Generating charts and captions can be handed to AI. Two things stay with a person, the error bars and the axes. A "large lead" drawn with the y axis starting at 0.9 is the graphical cousin of wording drift. The numbers are right, the strength lies. The AI division of labor fits in one line. First drafts, rewrites, number extraction, checking, hand them off. Step 1 and step 5, fixing the facts and signing, stay with you. The full fillable version of this chapter's templates is in the appendix. ## 9.5 If you are submitting, three corners of the same discipline Most readers deliver memos and technical reports. This section is for people submitting to a venue or writing a proposal. Everyone else can skip to section 9.6. Choosing where to submit is also an audience judgment. The shape of this batch of evidence, a mixed ending, one post-hoc overturn, decision-grade accounting, fits a workshop, the venue for trading half-finished work and negative results. It cannot carry the packaging of a main-conference claim. Upgrade the vehicle and readers automatically upgrade how they read your claim strength. Inflating a workshop skeleton into a conference paper means you completed the wording drift yourself. In an engineering setting this is the decision of who sees the study first, an internal team review or a direct report to management. The point-by-point response letter to reviewers works the same way. Let AI draft the point-by-point replies. How far each point concedes, and which comment is worth resisting, is a negotiation over claim strength and stays with a person. Review comments are also a free round of red team, which Chapter 10 opens up. The "preliminary results" section of grants and proposals is the disaster zone for strength drift, where pilot data grows into "an established method" under AI polishing. The discipline does not change. Numbers point to the repo, strength matches the evidence, and you submit only what you dare sign. That "preliminary results" paragraph in a budget request is its engineering cousin. ## 9.6 Spine update Β· writing βˆ’2.3 and 1/5 into two documents Now I run the workflow in front of you. Chapter 8 closed owing two debts. How the sturdiest conclusion in the whole case, the code one, gets into the first line of the CTO memo, and how the math overturn gets into the technical report's limitations. First the lazy route, having AI draft the report abstract straight from the results file. Feed it only the preregistered results table, one call, raw output stored as is in the repo at `results/abstract_draft_raw.txt` (that run was in Chinese, so the specimen stays Chinese and the quotes here are translated). The first few sentences of the draft are actually quite restrained. The code differences are "all not significant," cost 15%~40%, all correct. The closing sentence gives it away. "Overall, an army of open-source small models can match or surpass frontier models on structured reasoning tasks at lower cost, though its gains are clearly task-dependent…" Every word of that sentence has a source, and not one word can be signed. "Surpass" comes from the math preregistered table, where the verdict column really does read army_ahead, and the middle sentence honestly reported "army vote leads significantly by 18.2 percentage points." But that +18 already died in the Chapter 8 audit. The model did not know. What it was handed was the preregistered table. "Match" is stamped on code's CI crossing zero. Read strictly by the preregistration, crossing zero is called "undecided," not "match." As for that handsome class name "structured reasoning tasks," it is an honorary title invented on the spot for one ambiguous-template artifact, an artifact being an illusion the measurement process manufactures itself, not a real capability gap. AI fabricated not one number, and it even carried the "task-dependent" caveat for me. Inside the legitimate wording space, it parked the summary sentence in the most flattering corner. Strength drift raises no alarm. This is what it looks like. Not one number is wrong. What crossed the line is the nerve of the half sentence after "overall." So the process backs up to step 1, write the claims list first. The tier column gained one tier beyond the three, outstanding, meaning the parts the preregistration promised and this round did not do. | Claim (the strongest form I dare sign) | Evidence pointer | Tier | Signature | |---|---|---|---| | On the code task the gap between the army and frontier cannot be told apart within measurement precision (93.7% vs 96.0%, CI crossing zero), at 1/5 the cost | the code row of report.md | Verified (decision-grade) | Dare | | On knowledge QA the army trails by 15pp [βˆ’21.3, βˆ’8.7]; the SC (self-consistency) arm trails as well, teaming cannot rescue a weak knowledge base | the mmlu row of report.md | Verified | Dare | | Trust neither side of the math evidence. Preregistered +18.2pp, post-hoc removal of the ambiguous template βˆ’2.3pp, direction reversed; the problem set holds only ~3 template families, the CI is overconfident | audit_math.py, the post-hoc audit entry in CHANGES.md | Verified (for "the problem set is sick") | All I dare sign is "this exam is void" | | This experiment cannot detect any contribution of "teaming" over "multi-sampling" (clean subset 0.977 vs 0.983; on code SC is cheaper still) | report.md + the audit | Verified | Dare | | The self-hosted amortization basis, promised in the preregistration, never run in the repo | the cost section of prereg.md | Outstanding | I dare sign "not done" | The last row is worth a stop. The Chapter 6 plan signed off on two cost accounts, API list price as the main one, self-hosted amortization as the secondary. The repo delivered only the first. The amortization account has to be estimated from GPU rental prices, and its slot died on budget and time. The temptation at a moment like this is silence, since nobody remembers a Chapter 6 promise. The discipline goes the other way. **What you promised and did not do goes into limitations, recorded as outstanding**, and not one number gets invented. Then the two vehicles take shape. The technical-report skeleton puts the task-dependent framing "when does it pay off" straight into the title, so the abstract does not get read halfway. The abstract reports numbers under the Chapter 8 side-by-side discipline. Seven limitations, each carrying numbers. The three ugliest, the math slice's ~3 template families, the direction reversal from the post-hoc audit, and the outstanding amortization basis, happen to be the three most informative. Writing that section I noticed one thing. It is the only section in the whole document that felt steadier the longer I wrote. The memo's first line is the claims list's first line verbatim. Three risk lines follow in order. This is a public benchmark, not your workload. The gain from "teaming" cannot be told apart from "multi-sampling on one model," so the deployment recommendation gets simpler instead, one model plus majority vote. Cost is counted at API list price only, and the self-hosted amortization account was never measured. The "no shot" on knowledge QA trailing by 15 points goes into the judgment column of the number table. "No shot" and "evidence void" are the two most money-saving words in a memo. The second risk line's origin is worth noting. The most valuable line of advice in the memo came from the finding most damaging to the army narrative. Once the SC arm audited "team magic" away, what was left was a recommendation with half the ops, half the story, and twice the credibility. Honesty often improves the recommendation itself. It is far more than a moral pose. Close with the interlock check. Dispatch AI to extract every number in both documents and match them against report.md. All of them matched. Catching nothing is still worth it. The value of a step is that it runs every time. The subplot got delivered along the way. The variance-collapse red line in the text below is one of the three criteria. It means the spread of answers within one persona group is far smaller than among real people, like a room of recitation machines. The text reads, "The persona panel cannot replace real interviews for a decision-grade rehearsal. All six subgroup distributions fail the threshold, and the variance-collapse red line is tripped. 10,800 interviews, $0.47, bought one clean negative." The most common delivery failure for a negative result is **burial**, spreading a FAIL across ten pages of "worth further study." A negative put on top in one sentence is how a team saves itself one more $0.47. All the numbers and caveats get on the table only in Chapter 12. Here only the conclusion sentence is delivered. Finally, the honest identity of these two finished pieces. They sit in full in this chapter's appendix, every number pointing at a specific file and commit in the smol-army repo. Nothing was submitted, no CTO ever signed off on them, and the book does not act it out. Right now they have passed only their own signature test. Chapter 10 will let them take a real beating. ## 9.7 Swap in your project Dig out your most recent batch of evidence that is "done but not delivered." Budget half a day. 1. **Write the claims list**. One sentence per claim (the strongest form you dare sign) + evidence pointer + honesty tier + signature field. Write what you did and what you did not, and "not done" is a line too; 2. **Pick two real audiences**. One wants "how do you know" (peers, reviewers, the technical committee), one wants "what should I do" (your boss, a client, a funder); 3. **Have AI produce two drafts from the list**. The technical-report skeleton + the one-page memo, with "do not change the strength of any claim" locked into the prompt (templates in the appendix); 4. **Reconcile the strength**. Compare sentence by sentence against the claims list and mark every upward drift. You will usually find more than three. That is normal, not an incident; 5. **Interlock check**. Dispatch AI to extract every number in both documents and match it to the source. Release only when all three agree; 6. **Signature test**. Ask of each sentence "pulled out on its own, do I stand behind it," and for anything you dare not sign, lower the strength or delete it. The acceptance criterion in one sentence. Anyone points at any number in either document and asks "where did this come from," and you point to the source file within fifteen seconds. **Want an agent to run it with you?** Paste this to your AI assistant or coding agent: ```text Help me run the Swap in your project of Chapter 9. First I write the claims list with Template 1 of docs/appendices/ch09-templates.md, one sentence per claim plus an evidence pointer and an honesty tier, "not done" is a line too, and you only build the table, you never write my conclusions. Once the list is final, use the prompt set to produce two drafts from it, a technical-report skeleton and a one-page memo, with "do not change the strength of any claim" locked into the prompt. Then I reconcile the strength, and you mark every wording stronger than the list. Then open a separate session, give it only the two documents and the list, run the number interlock check, and list every disagreement among the three for me to rule on. I do the signature test sentence by sentence, and lower the strength or delete anything I dare not sign. If any command errors, stop and show me the output. ``` ## 9.8 Sober reminders - **One draft sent to everyone is the number one failure mode**. Sending the same full text to every audience outsources the reordering of detail to the reader. Readers do not reorder. Readers just do not read. - **Polishing is not a free operation.** Every "smooth out the tone for me" reshuffles claim strength again. Fix the rule. After polishing you must rerun the strength reconciliation. - **Limitations is not body armor.** Once you have written template concentration and an overconfident CI, the abstract may not cite that CI as a victory. Honesty layering is a downgrade consistent across the whole document, and a disclaimer at the end cannot stand in for it. Using limitations to insure an overclaim in the main text is worse than writing no limitations at all. - **Do not invent a recipient.** A deliverable's credibility comes from being recomputable, and the social proof of "already adopted" does not help with that. If it was not submitted, say it was not submitted. A pilot is called a pilot. - **Still exploring**. Fully automatic tools from results to paper are iterating fast. As of this writing, my observation is that the draft structure is usable and the claim-strength calibration is not. Their default wording setting is selling, not reporting. The book's online case library tracks it. One step remains before delivery. Kill your own conclusion before someone else does. Chapter 10, red team. ## 9.9 The unfair advantage you now hold With a batch of evidence in hand you can produce a technical-report skeleton and a one-page memo in half a day, numbers interlocked, every number traceable to the repo in fifteen seconds. What most people deliver is a twelve-page document that answers nobody's question and earns a one-line "so what" on Monday morning. --- --- # Chapter 10 Β· Red Team !!! info "Chapter companion" πŸ“‹ [Chapter 10 templates](../appendices/ch10-templates.md) Β· πŸ—‚ [Template index](../appendices/template-index.md) Β· πŸ’» [`code/persona-panel`](https://github.com/hallieren/research-rewritten/tree/main/code/persona-panel/) > **This chapter's ladder.** On the red team step AI sits steadily at **assistant level**, and this is the step of the seven that gets underrated hardest. Having it attack your conclusion costs almost nothing and pays off at once. **Collaborator level** (systematically searching out counterevidence for you) is still exploring. Keep the ceiling in mind too. It finds holes in reasoning and holes in statistics, and it does not find the "everyone in your field knows this, only it doesn't" kind of problem. So it writes the charges, and the ruling is yours. > > **Spine update.** The last chapter packed the evidence into two vehicles, a technical report and a memo. The deliverables sit in the repo, one step short of going out. This chapter is the last step before that. I take the Chapter 8 interrogation apart into parts, assemble them into an attack process that can be run on any conclusion, and then lay the whole case file out where anyone can rerun it. > > **This chapter delivers.** Three tools. A red-team dispatch brief template, a four-attack-surface prompt set, and a disposition table. --- ## 10.1 The email three weeks later Three weeks ago you sent that memo to the whole team. The new design was thirty percent cheaper than the old one, the data was complete, the charts were clean. This afternoon an engineer from the next team over, known for being exacting, replied with your boss copied in. The wording was very polite. Attached was a notebook that reran all of your data. Your scoring script counted one class of timed-out requests as successes, and the timeouts happened to cluster in the old design's peak hours. Drop them and thirty percent becomes seven percent, with a confidence interval crossing zero, meaning the interval contains zero, so you cannot rule out that the new design saves nothing at all. You stared at that attachment for a long time. The conclusion flipping is the lesser part. The sharpest sting is remembering that you had seen that timeout field, hesitated two seconds over whether to handle it separately, and then the deadline won. Anyone who has submitted a paper knows the same scene. The review comment that hurts most always aims at the paragraph you were shaky about yourself and hoped nobody would notice. Shaky paragraphs are all written the same way. The qualifiers suddenly multiply, the evidence suddenly thins, and exacting readers have a hound's nose for it. The two scenes share one law, and I call it **the conservation of criticism**. A conclusion with weight will meet its first serious attack sooner or later. You do not decide whether it happens. You decide two things. Whether it happens before release or after. And whether the attacker is employed by you, or by their own curiosity and hostility. That choice has a name, the red team, **kill your own conclusion before someone else does**. Chapter 2 put this step after delivery and before submission, and the reason was given then. Keep the trial and error inside your private loop. What enters the public knowledge base is held to another standard. Now it is time to execute. ## 10.2 From one interrogation to an attack process In Chapter 8 you already saw one attack from start to finish, a "victory" of +18.2 percentage points interrogated down to βˆ’2.3 by three questions. But that was the interrogation. The interrogation and the red team differ in three places. **The trigger differs.** The interrogation is triggered by excitement, and its object is a result you want to announce. The red team has no trigger condition. Every claim in the deliverable takes a beating, dazzling or dull. The interrogation is firefighting. The red team is acceptance. **The timing differs.** The interrogation happens during interpretation, and it hits numbers. The red team happens after the deliverable is final and before it goes out, and it hits "the sentence you intend other people to read," wording included. "The army ties frontier on code" and "on code the gap between the army and frontier cannot be measured in either direction, at one fifth the cost" come from the same numbers and differ by a factor of two in how much surface they expose. **AI's role differs.** The Chapter 8 discipline was to seat it as a juror, laying out facts and passing no judgment. In the red team it is promoted to a hired prosecutor, and the job is to construct charges, as hard as it can. This does not cross the division of labor of the interpretation step. What the prosecutor hands in is only a list of charges, each with an executable check attached. The verdict is not in it, and not one millimeter of ruling power has moved. The real risk runs the other way. It will hit too softly. Put weeks of work in front of it, add "take a look and see if there is anything wrong," and it will dutifully list three harmless items. Sycophancy, the habit of agreeing with you to please you, turns in a red-team setting into running a death-penalty review as an awards ceremony. How to dispatch so you get a real attack is what section 10.5 is for. This kind of systematic hostility used to be scarce and left to luck. Now you can hire a tireless opponent at any hour, as Chapter 3 promised. The excuse "I could not find an opponent" no longer exists. ## 10.3 Lessons stolen from security engineering You may not do security engineering daily, but you know the names threat modeling and fuzzing, and its daily business is paying people to attack its own systems. Three lessons transfer directly. The only question is where they break when they land on research conclusions. **Lesson one, attacks come from enumerating a list, not from inspiration.** A security review does not ask "does this system have a hole." It walks a fixed threat taxonomy surface by surface, and Microsoft's STRIDE goes through six surfaces one class at a time (Kohnfelder and Garg, 1999 internal document, "The Threats to Our Products"). "Is there a problem" is a prayer. "Here is what each of these surfaces turned up" is a process. The counterpart for research conclusions is the four attack surfaces of the next section. Where it breaks. STRIDE's six surfaces are cut by the attacker's motive, and a research conclusion has no attacker, so the four surfaces are cut by which parts the conclusion stands on. Section 10.4 shows how those four got drawn. **Lesson two, machines do volume, people do triage.** Google's OSS-Fuzz has reported more than fifty thousand bugs and more than thirteen thousand security vulnerabilities across about a thousand open-source projects (official figures, as of May 2025), and most of that output is duplicates and false positives. The valuable step is triage. Row 1 of the transfer map reads in one line at the red team step, **the production cost of charges collapsed, the cost of ruling on them did not**. AI can hand you dozens of charges in an hour. Which one is worth a check and which is noise is still your work, and it is the one part of this step you cannot cut. Where it breaks. A fuzzer's false positive is ruled out by one reproduction, while the check for a charge has to be designed by you, which is why the fourth discipline in section 10.5 requires a charge to carry its own check. **Lesson three, trust comes from mechanism, not from self-assessment.** Row 3 of the transfer map says self-perception has to be calibrated by outside measurement, and row 9 says output from strangers gets trusted through mechanism, not goodwill. Put together, the corollary at the red team step is cold. When you say "I checked it myself," what carries to another person is about zero information, because everyone who never checked says the same thing. Transferable trust has only one shape. Put the criteria, the ledger, and the scripts on the table, and bring the cost of "recheck it" down to one command. That is the whole reason for the public case file in section 10.9. ## 10.4 The four attack surfaces First, why these four. Take any empirical conclusion apart and it stands on four parts. Who judged right from wrong, what problems it was tested on, the sample size you think you have, and whether all of it was worth the price. Across the seven chapters of this case, every one of those parts took a real hit. **The four attack surfaces are the shape you get by drawing lines between every bullet hole of this case on the wall**, and I claim no credit for inventing a taxonomy. | Attack surface | The prosecutor's question | The bullet hole in this case | |---|---|---| | The scorer | Who defined "right"? Is there a second defensible definition? Has every link of the scoring chain been replayed? | The % suffix miscarriage of justice; gold (the standard answer) read ambiguous English as having one correct reading (49 "wrong answers" identically equal to 4Γ—gold) | | Data composition | Where did the problems come from and how were they drawn? How many molds are there? Are they in the training corpus? | The draw took the first 150 rows unshuffled; the 150 problems are really ~3 template families | | Independence assumptions | Is n the number of rows or the number of independent units? What does the CI assume? Are the errors of the two arms correlated? | Within a template family they are highly correlated, and the preregistered CI is overconfident | | Cost basis | Who chose the basis? Does the conclusion flip on another set of books? Was the price the control arm got fair? | frontier used the family's balanced tier; the local amortization account is an analytic estimate, not measured | Now the shape of the questions surface by surface. The full prompt set is in the appendix, and the chapter body demonstrates only one. **Attack the scorer.** The question is the definition of "right" itself. Is there a second defensible reading of the gold answer? On the chain of parsing, matching, and scoring, which link silently swallows points or hands them out? Do the wrong answers carry a tidy pattern, always k times, always in one format, clustered in a continuous stretch? Those 49 "wrong answers" of Chapter 8 that were identically 4 times gold are the standard quarry on this surface. The surface holds for research that uses no benchmark too, only the scorer goes by another name. A human annotation rulebook, an LLM judge's preferences, the operational definition of "valid," all of them are scorers. The work order to the prosecutor looks like this, and the core is three things, charge, consequence, check. The rest is formatting detail. ```text Role: you are hired to attack the claims list below, on the attack surface "the scorer / the criteria," that is, how right and wrong, success and failure, get judged (gold answers, scoring scripts, annotation rules, operational definitions). Input: the claims list and the paths to the materials (see the dispatch brief). Output: a list of charges, ordered by lethality. Every charge must carry three things: 1 Charge: which point of the judging chain could score a right answer wrong, score a wrong answer right, or where the definition of "right" itself has a second defensible reading; 2 Consequence: if the charge holds, which claim dies, and what direction and magnitude of effect gets manufactured; 3 Check: one check executable today (a script sketch / a sampling plan / a replay experiment) whose result can confirm or rule out this charge. Rules: attack only, no balanced coverage; do not evaluate whether the conclusion as a whole is credible; do not list a charge you cannot give an executable check for; if you find nothing here, write "nothing found on this surface," do not pad. ``` The work orders for the other three surfaces differ only in the attack surface line. Everything else is the same. **Attack the data composition.** This asks about the material itself. Does the coverage hold up the wording of the claim? The two bullet holes of this case are in the table above. The nastier one is GSM-Symbolic. It was chosen for contamination resistance, and the perturbed variants did block memorized problems, but they turned the problem set into copies of a few templates. The move that seals one attack surface can open the door on another. **Attack the independence assumptions.** This asks about the quality of n. What do the samples share, and has the data structure broken the independence assumption behind the CI and the bootstrap? The bullet hole of this case is the one in section 8.4. That narrow interval [+12.7, +24.2] bought its narrowness with the assumption that problems are independent, and once the assumption goes bankrupt the effective sample size drops to 3 template families. **Attack the cost basis.** Every claim carrying the words "cheaper," "better value," or "cost-matched" owes this surface a round. Two soft spots are on the record in this case. The frontier arm used the balanced tier of the GPT-5.6 family, the mid-priced tier of that family, not the flagship, and the change log carries a line added for it, "the book must state the exact model," a preventive narrowing of "ties frontier." The "local deployment amortization" account was an analytic estimate at public rental prices, never run on real hardware, and it travels with the deliverable as a caveat. When I drafted this section I parked an attack here that I thought unbeatable, as an example, asking whether the army's one-fifth cost advantage might be propped up by frontier's hidden reasoning tokens. My answer at the time was yes, but the basis signed was billed dollars, hidden tokens are money you really pay, so the charge holds and the conclusion does not move an inch. It sounded airtight. Then the real trial of section 10.7 knocked down the "unbeatable" itself. The prosecutor went and checked the usage field. frontier's reasoning tokens measured zero. That "money you really pay" I held up as a shield does not exist. Keeping this corpse here is useful. **The finest defense you rehearse for yourself usually rests on a fact you never verified.** ## 10.5 How to hire a prosecutor who does not applaud With the four work orders written, the way you dispatch decides whether you get attacks or applause. Five disciplines, shaped after this book's own verification process. The check that overturned my "literature gap" illusion in Chapter 4 was run exactly this way. Another model, a fresh session, a brief holding only the claims list, not one word of expectation. **One, an independent session, zero history.** The red-team session may not carry the narrative, the excitement, or the wording habits of your last few weeks. All of them leak expectations. Switching models is better still. Letting the model that paired with you to write the conclusion also red-team it violates the orthogonality principle of Chapter 2, section 2.4. **Two, the brief gives only the claims list and the raw materials.** Compress the deliverable into a column of neutral claims ("on the code family the CI of the gap between the army and frontier crosses zero, at one fifth the cost"), and attach the paths to the data, the scripts, and the ledger. Which one you hope survives, not one word. "I am worried" and "confirm this for me" are barred too. **Three, one surface, one order.** Open a separate order for each of the four attack surfaces, and do not merge them. Dispatch them merged and the model fills the page with the easiest surface, while the one that hurts most gets two perfunctory lines. **Four, a charge must carry an executable check.** A charge without a check is a comment. A charge with one is a work order. Writing this into the rules section of the prompt lifts the output a whole grade. It forces the model to translate empty phrases like "the data may have a problem" into "cluster the problem texts by template and count how many kinds there are." **Five, file the output and rule line by line.** Every charge either gets its check run or gets rejected in writing with one line of reason. The rejection record becomes your ammunition at the defense later, and Chapter 6 set the same rule. These independent channel disciplines do not serve the red team alone. Generalizing them into a verification process that cuts across the whole workflow is what Chapter 12 opens up. ## 10.6 After you are hit, overturn, narrow, caveat Once you rule that a charge holds, there are only three legal dispositions. The fourth is called "I am aware of it," and it does not exist. A charge that hit you and got no written disposition is worse than no red team at all, because now you know. **Overturn, the charge hits load-bearing structure and the announcement sentence dies.** The specimen is math's +18.2. gold read ambiguous English as having one correct reading, that charge held, and after the ambiguous template was removed the direction reversed to βˆ’2.3, so "the army pulls ahead on math" came off the deliverable. Note that overturning is not hiding. The preregistered numbers get reported as is, by the Chapter 8 discipline. **An overturn targets what you intended to say. Not one number that happened gets touched.** **Narrow, the charge holds, but what it cuts away is the range or the strength, and the conclusion itself still stands.** The specimen is math's CI. The charge that "the 150 problems are really 3 template families" held, and the narrowness of the preregistered interval is false confidence. The disposition is to downgrade the claim to "this problem set cannot detect a conclusion, reissue it and test again." Announcing the opposite conclusion would cross the line. There is one criterion for a qualified narrowing. The new statement is strictly weaker than the old one and still checkable. The cost-surface line "the book must state the exact model" is the same move, narrowing "ties frontier" into "ties GPT-5.6-terra (the balanced tier of that family)." Narrowing is nothing to be ashamed of. From "can small models team up" to "code has a shot, knowledge QA has none, math needs a retest," the main conclusion of this case got narrowed into shape the whole way. **Caveat, the charge holds or cannot be ruled out, but it is unfixable, untestable, and not fatal.** The disposition is to keep the conclusion and let the caveat travel at the same address as it, not buried in the third paragraph of limitations. Qualification for a caveat is strict. What can be fixed gets fixed, what can be tested gets tested, what is truly fatal gets overturned or narrowed, and only when all three fail does the caveat get its turn. The specimen for the caveat tier is in the subplot, and it is the most complete "red-team your own plan" in the book. The ground truth of the persona study was designed with two sources, public survey data plus a small sample of real interviews. The real-interview arm was cut before the run started. I could not recruit interviewees. A real constraint of the verification budget, undignified but true. After cutting the arm I ran a round of soft-spot self-check on the single-source plan that was left, on exactly the last two of the four surfaces above. On data composition, WVS, the World Values Survey, is a famous public question bank that has been out for years and anyone can look up, very likely inside the training corpus. If persona answers like a real person, that could be likeness, or it could be memorized questions. On independence, once the real-interview arm was cut, no control left in the plan is naturally immune to corpus contamination. The arm that could tell "memorized questions" from "sensitivity to wording" is exactly the one that got cut. The 10/15 double reading of the Chapter 8 subplot, the contamination test flagging two thirds of the questions, has its root here. Soft spot confirmed, and unfixable (there is no second ground truth to swap in) and untestable (the budget is the reason the arm was cut). The disposition has two steps. The reworded contamination test got promoted from a bonus to mandatory and written into the change log. Every subplot conclusion carries the same caveat, **"the ground truth itself may be inside the training corpus."** Even the FAIL carries it. Analyzed in the direction of the Chapter 8 trial, even if the contamination is real, falsification only gets more conservative. The caveat goes on anyway, because a caveat states the boundary of the evidence and has nothing to do with whether the conclusion feels shaky. When you red-team your own plan, the most honest output is sometimes a caveat written on its face. ## 10.7 Trial record, one round on each surface, all ten charges hold With the process written, it is time to eat my own dog food. After the two deliverables of Chapter 9 were final, I really did dispatch a round under the five disciplines of section 10.5. Four attack surfaces, four independent sessions, one order per surface, a brief holding only a neutral claims list and file paths, every charge carrying an executable check, every output reviewed and ruled on by me line by line. Ten charges. All of them hold. Here are the six that hurt most. **On the scorer surface, the first shot hit the audit itself.** That βˆ’2.3pp reversal of Chapter 8 was overturned. Correct answers in the clean subset wrapped in format residue like `**` and `}` got scored zero, and only small reasoning models produce that residue, so it kills small models in one direction only. Rescored under lenient parsing the three arms come in at 1.000 / 0.9967 / 1.0000, all saturated, meaning all near full marks with nobody able to beat anybody, so clean math is uninformative in both directions and can detect a gap in no direction at all. The pilot had fixed the same bug for the % suffix, and `**` and `}` slipped past under its nose. The audit itself has to be audited too. **The cost surface turned up a breach.** The preregistration stated that the SC arm, the control arm where a single small model samples itself many times at the same budget and takes the majority, would have its k, the number of self-consistency samples of the single small model, tuned after the pilot to within Β±10% of the army vote's dollars per query. It was never executed. k stayed at 5, and SC actually paid only 0.21 to 0.77 times the army. The phrase "same-budget control" is withdrawn across the book and honestly restated as SC spent less money and still tied or did better. The criticism that "teaming β‰ˆ sampling" got harder instead. **The independence surface reported a mechanical death sentence.** Voting on code is exact string tallying with ties resolved to the first sample. Extracted code is almost never identical character for character, so the code army's score is identically the single sample of qwen, the small model rotated into first place, and four of the five calls are wasted. **The data composition surface seized the twin of the math sampling bug.** "mmlu_pro stratified by category" is really the three subjects business / law / psychology. The draw pulled the first 2000 rows of a test set blocked by subject, with zero STEM coverage. The math sampling bullet hole was on the record at the time, and this one slipped through the mesh of the audit. The math family itself did not escape either. Rerun the bootstrap clustered by template and the CI is [βˆ’4.0, +59.3], crossing zero. Add that gold clusters too, 24 of the 50 problems of one template family sharing the standard answer "25," and voting harvests the dividend of the modal gold, meaning the answer voting lands on may just happen to match the standard answer that appears most often in this batch of problems, which in this design cannot be separated from real error correction. The math family retires in both directions. **The ledger confessed last.** That $5.58 is a simulated account of list price Γ— tokens, and it does not match what the bill charged. The army side actually paid about 19% more (from the author's bill reconciliation, not recomputable from repo data). Another 12.6% of real spending (cleaning, reruns) sits in the ledger and appears in no reported number. Every disposition got filed under the three tiers of section 10.6, and all of it went into the case file. The yield of this round is not mainly the bugs it caught. The point is the relation between the red team and the audit. **The red team attacks the audit itself.** The audit hunted math down. The red team then seized three more cases exactly where the audit had declared things clean, the mmlu sampling, the code mechanism, and the ledger basis. The full table of ten charges follows, and each line matches the case file at docs/redteam-2026-07-25.md. Read the last column first. Two overturns (β‘  and β‘’). Narrowings are the majority. The caveats are all open items and travel with the deliverable. The middle column, "how it was checked," is for people who want to rerun it, and you can skip the terms you do not know. Only ⑨'s "zero-variance degenerate-interval artifact" needs one line. In that sub-slice all three sides score full marks, every resample comes out the same, and the interval width is zero, an illusion produced by identical scores rather than a real tie. | # | Charge | How it was checked | Verdict | Disposition tier | |---|---|---|---|---| | β‘  | The βˆ’2.3pp reversal on the clean subset is a scoring-residue artifact, correct values wrapped in `**`/`}` scored zero, killing small models in one direction | Rescore under lenient parsing, three arms 1.000/0.9967/1.0000, clean math saturated, uninformative in both directions | Holds | Overturn | | β‘‘ | The preregistered SC Β±10% cost-alignment clause was never executed (k stayed 5) | Recheck the ledger against the prereg clause, SC actually paid only 0.21 to 0.77 times the army | Holds | Breach log + narrow (the "same-budget" label withdrawn across the board) | | β‘’ | The code voting mechanism degenerated, exact string tallying with ties resolved to the first sample | Replay the tally, the code army's score is identically qwen's first sample, 4 of the 5 calls wasted | Holds | Overturn (mechanism) | | β‘£ | The mmlu slice is really the three subjects business/law/psychology, zero STEM coverage | Look the categories up by problem id (50/50/50); split by subject, law βˆ’25.3 / business βˆ’9.3, clustered CI [βˆ’25.3, βˆ’9.3] | Holds | Narrow | | β‘€ | Independence bankrupt across all of math + gold clustering | Bootstrap clustered by 3 template families [βˆ’4.0, +59.3] crossing zero; 24 of 50 problems share the gold "25" | Holds | Narrow (the math family retired in both directions) | | β‘₯ | The ledger is a simulated list-price account, not what was paid | Bill reconciliation, the army paid about 19% more; 12.6% of real spending enters no reported number, and on the full basis code SC/vote cost the same | Holds | Narrow (wording) + caveat | | ⑦ | frontier ran on an undisclosed default reasoning tier, and the budget remedy was asymmetric | Check the usage field, reasoning tokens measured zero; qwen raised to 8192 while frontier stayed at 4096 | Holds | Narrow + caveat | | β‘§ | runs.jsonl stores extracted answers, not raw output | Inspect the written fields and the extraction regex (first match taken, the loss not quantifiable) | Holds | Caveat + correction to the ch7 statement | | ⑨ | The math "tie" of debate/division is a zero-variance degenerate-interval artifact | Replay the bootstrap, all three sides full marks on the sub-slice, every resample zero, CI [0,0] | Holds | Narrow | | β‘© | The cost comparison is sensitive to the procurement basis | Recompute on another basis, at Batch half price the code cost ratio goes 1/5β†’2/5; the frontier unit price was only checked against an aggregator | Holds | Caveat (open item, the author to check first-hand billed prices) | ### After the beating, what that one page says now With ten charges disposed of, the one-page memo of Chapter 9 cannot go out as is. Part One of the appendix keeps the pre-red-team version. Both versions stand, and the difference is the lesson. The final version changes three places only, the first line, the second risk line, and the next step. The first line now reads. On code generation tasks the accuracy gap between one call to a single open-source small model and GPT-5.6-terra cannot be told apart within measurement precision, qwen single-shot 93.7% against 96.0%, the interval crossing zero, read as uninformative, not as a tie, at about 1/25 the latter's list price per call. The three risk lines now read. First, this is a saturated slice of a public benchmark, the problems may have leaked, and it is not our workload. Second, "teaming" showed no gain over single-model sampling this round, and the voting mechanism on code even degenerated into a single shot with four of the five calls wasted, so any proposal sold on the army gets evaluated at zero gain. Third, cost is a simulated account of list price times tokens, actual payment runs about 19% higher, self-hosted amortization was not measured, and procurement discounts can push the cost ratio from 1/5 to 2/5. The next step now reads. Run a two-week shadow-traffic pilot on code workloads, starting from single shot, the cheapest, adding sampling only if that is not enough, and re-evaluating k=5 after the tallying mechanism is fixed. Knowledge QA stays on frontier, since this round tested only the three subjects business, law, and psychology. For math, wait for the problem set to be reissued, and do not use this round as a basis. All three changes come from charges β‘‘β‘’β‘₯β‘© of the table above, and each one can point at the case file. The final version is shorter than that Chapter 9 page and cheaper, and it recommends one thing less, the army. That is what the red team does to a deliverable. What it deletes is never the numbers. It is what you had intended to say. ## 10.8 Both sides of peer review Readers who submit and review should linger on this section, and everyone else can skip to the next one. Before you submit, dispatch the manuscript under the disciplines of section 10.5, one round per surface, and the list of charges that comes back is a pre-review report. Then write the response now. What you can write stays as ammunition, and what you cannot write is a narrowing signal. Revising today beats forcing an argument three months later in the rebuttal, the reply that argues back against review comments. When you review someone else's manuscript, the four attack surfaces are your review checklist. Is the scorer trustworthy, what does the data look like, is n real, is the cost basis fair, one paragraph per surface, ten times more useful than "the contribution feels insufficient." One red line. Feeding someone's unpublished manuscript to an outside AI service is a direct violation in many settings. NIH has barred reviewers from using generative AI in grant review since June 2023 (NOT-OD-23-149, the ban covers reviewers, not applicants), Elsevier and Springer Nature bar reviewers from uploading manuscripts, and ACL's reviewing guidelines forbid generative tools for drafting a first review. The checklist itself works without AI, and the duty of confidentiality outranks the convenience of a tool. ## 10.9 No submission, but a case file anyone can rerun Now an old debt gets settled. When Chapter 3 previewed this step, what I wrote was "first let AI be the harshest critic, patch, then really send it out and take the hits from humans." To report it honestly, there was no submission. The deliverables of this case are a manuscript and a repo, and there is no paper in my hands waiting on a journal's verdict. Inventing a submission and three fictional reviewers to complete the narrative would violate exactly every discipline of this book. The substitute is harder than an apology, **the case file is public**. The preregistration entered the repo before any result. The ledger is append-only, and 5.58 dollars adds up line by line. The audit script reruns on one command, and `uv run python scripts/audit_math.py` prints every number of the Chapter 8 overturn. The change log keeps even the erratum for a mistyped concurrency setting exactly as it was (the subplot's accident of misrecording a problem id has its erratum in the persona-panel change log, also on file as it was). The red-team case file itself is in the repo too, docs/redteam-2026-07-25.md, all of it on record. Any reader can be my Reviewer 2, and gets more than a real reviewer would. A reviewer only gets the text in a PDF. You get an executable chain of evidence. Security people call this play a bounty, turning hole-hunting into a legitimate business for the whole world, which beats vouching with your own word that there are no holes. This is also the last cell of the spine case in Part II. That two-hour demo of Start Here, 19 to 20, "it looks like a tie." It went through the literature brawl of Chapter 4, the hypothesis with Ξ΅ equal to two percentage points of Chapter 5, the three-arm sign-off of Chapter 6, the 5.58 dollar execution of Chapter 7, the interrogation of Chapter 8, and the two vehicles of Chapter 9, and arrived at this chapter. Here it took the beating from its own side, the math family retired, and the case file went public. The endgame is H not falsified and not holding across the board. A single small model (self-consistency arm 94.7%, army arm 93.7%) cannot be told apart from frontier (96.0%) in direction on the saturated code slice, the reading is uninformative, do not read it as a tie, at about 1/25 per call (list-price basis), and the one fifth quoted earlier is the full bill for the army's five calls on one problem, while this figure counts a single call. Knowledge QA was measured on only the three subjects business, law, and psychology, and has no shot. The math family retires in both directions. "Real teaming beats sampling a single small model" never got measured once in the whole case. What is interesting is that the naive demo guessed the direction right on the code family. Now that direction comes with criteria, a ledger, control arms, an audit script, and a red-team record. **What lies between a lucky guess and something you can trust is these seven chapters.** The seven steps are done. But a process governs the workflow, not the person executing the process. Why a fabricated citation looks so much like a real one, and why you believed it exactly when you wanted it most, these failure modes of judgment itself are what Chapter 11 opens up. ## 10.10 Swap in your project Dig out the deliverable you are about to send, a paper, a memo, a report, whichever is closest to the send button. Budget half a day. 1. **Compress out the claims list.** List the attackable factual claims of the deliverable as 5 to 10 neutral statements, with the conclusion and its qualifying conditions split into separate lines; 2. **Write the dispatch brief** (appendix template), holding only the claims list and the paths to the raw materials, then run a leak check on it, and delete every expectation, worry, and adjective; 3. **Open one independent session per surface and run a round on each of the four**, the one you feel shakiest about first; 4. **Triage.** Run the check each charge carries, hold it or reject it, and write one line of reason for a rejection; 5. **Disposition.** Assign a tier to every charge that holds, overturn, narrow, or caveat, one of the three, and write it into the disposition record; 6. **Revise the deliverable, then run one quick round again**, since a patch can introduce a new handle; 7. **Publish what can be published.** Data, scripts, criteria, bring the cost of "recheck it" down to one command. If you cannot publish everything, at least attach the disposition record to the deliverable. It buys more trust than the conclusion itself does. The full fillable versions of this chapter's three tools are all in the appendix. **Want an agent to run it with you?** Paste this to your AI assistant or coding agent: ```text Help me run the Swap in your project of Chapter 10. I give you the deliverable, and you do one thing, use Template 1 of docs/appendices/ch10-templates.md to compress the attackable factual claims into 5 to 10 neutral statements, with the conclusion and its qualifying conditions split apart. Then help me write the dispatch brief, holding only the claims list and the paths to the raw materials, and when it is written run the leak check, delete every expectation, worry, and adjective, and I will delete again what you miss. You do not do the red team itself, you are the generation channel. I open four independent sessions of my own for the four attack surfaces and run one round on each, pasting only the brief into each session. When the charges come back you only run the check each one carries and record it. I rule on each, holds or rejected, and I decide the disposition, one of three (overturn, narrow, caveat). If any command errors, stop and show me the output. ``` ## 10.11 Sober reminders - **An empty report earns no medal.** When the red team comes back saying "no major problems found," suspect the dispatch first. Expectations leaked, or the material you handed over was your summary and the originals never went in. Switch models and dispatch a second round, and only when both come back empty are you allowed to celebrate, quietly. - **The red team can knock you back to any earlier step.** Chapter 3 said the seven steps are a loop. A hit on the scorer sends you back to Chapters 6 and 7, and a hit on data composition can send you all the way back to the task-family choice of Chapter 5. Budget for "fix it and run another round," and do not treat the red team as a last-second glance in the mirror on your way out the door. - **A caveat is not a trash can.** Of the three tiers the caveat is the most comfortable and the easiest to abuse. Hanging a caveat where an overturn belongs is the best-dressed form of self-deception. The qualification check is in section 10.6. - **It cannot find your field's blind spots.** The four attack surfaces catch diseases of method, not "this measurement has been notorious in your field for years." The antidote is still the fourth route of Chapter 4. Find someone who really knows the field, and show them your conclusion and your red-team record. Those ten minutes still have no process that can replace them. - **Still exploring**. End-to-end automated red teams, generating the attacks, running the checks, assigning the tiers with no human hand, are iterating. As of this writing, charge generation can be let go of, and no tool yet dares sign for triage and disposition. The book's online case library tracks it. ## 10.12 The unfair advantage you now hold Before delivering any conclusion you can have AI run a round on each of the four attack surfaces, dispose of every hit as an overturn, a narrowing, or a caveat, and attach a case file anyone can rerun. For people without this step, the first red team is that email three weeks later with the boss copied in. --- --- # Chapter 11 Β· Failure Modes Unique to Research !!! info "Chapter companion" πŸ“‹ [Chapter 11 templates](../appendices/ch11-templates.md) Β· πŸ—‚ [Template index](../appendices/template-index.md) Β· πŸ’» [`code/persona-panel`](https://github.com/hallieren/research-rewritten/tree/main/code/persona-panel/) > **This chapter's ladder.** This chapter cuts across all seven steps. Every failure mode grows on a step of its own, so the chapter cannot be filed under any single step. Recognizing failure modes is itself work AI currently does at **assistant level**. Formal errors, fabricated citations, numbers that do not add up, you can let it scan for those. You will also meet a class of error it structurally cannot catch, and in that class your accomplice is you. The conservative note on the "read and catch errors" row of the snapshot in Chapter 3, section 3.5 is about exactly this. > > **Spine update.** Part III opens no new level. It settles a debt. The pit I fell into in Chapter 4 goes on the dissection table here, as the most expensive specimen in the chapter. > > **This chapter delivers.** Three things. Three error amplifiers, an "AI attribute Γ— research step" failure-mode table, and a thirty-minute failure-mode census, with an AI self-check prompt set attached. --- ## 11.1 The ninth hour before the meeting Eleven on a Tuesday night, the architecture review nine hours away. You are on the last pass. On the desk is a technology selection report AI was deeply involved in, twenty-six pages, clean structure, plenty of citations, your group's main output of the past two weeks, concluding that the retrieval layer should be swapped for a new engine. The core input to the migration benefit model is one performance judgment, that the new engine cuts "p99 latency by 31%" under comparable load. The report gives a source, the 2025 Vector Retrieval Benchmark Review, with a page number. You want to see the test workloads in the original, so you search. Nothing. Three different keyword sets, still nothing. You go back to the chat window and ask the AI. It apologizes, then "corrects" the source to the name of another evaluation report. That one does not turn up either. It is half past midnight. The report is not garbage. Most of the other citations you spot-checked are real, and the framework is genuinely useful. The trouble is that the number holding up the entire migration benefit now hangs on a citation nobody can find, in exactly the same tone as every true sentence in the report. The next thought is colder. What if you had not run that one extra search tonight? The 31% goes into the benefit model, the benefit model goes into the architecture decision record, the decision record goes into quarterly planning. Six months later another team runs a selection and cites your decision record. The lawyer in the Avianca case from Chapter 1, who filed AI-invented precedents with the court and was fined after the judge checked, is the man who failed to ask one more question at exactly this spot. This chapter does one thing. It builds a case file for moments like this. First the format of the file. It is not organized as a bug list for AI. You will see later that reading these errors as a bug list is precisely why they do not get stopped. The questions it answers are structural. Why are these errors more lethal in research than elsewhere? Where do they grow from? Handed a polished output, how do you know where to poke? ## 11.2 Old errors, a new production line Start by taking apart a convenient illusion, "these are all new AI defects, and one more model generation will fix them." Turn back to Chapter 2 and not one of these errors is new. Spurious significance means a conclusion that looks statistically sound but is a false positive squeezed out by adjusting the data and the analysis method over and over, p-hacking in the jargon. The Simmons paper demonstrated it live in 2011. That was eleven years before ChatGPT. Bad data riding into an authoritative conclusion, the Reinhart-Rogoff Excel formula that skipped five rows, moved fiscal policy debates in several countries. Citation rot is an older disease still, one that bibliography has always had. Dead links, drifting paraphrase, qualifiers shedding layer by layer through secondhand citation, all of it dates to the print-journal era. A meta-analysis of 28 studies of medical journals, meaning the results of many studies computed together (Jergas and Baethge, 2015, PeerJ), measured a total citation error rate of 25.4%, and about one in ten of those citations simply did not support the claim they were cited for. The comparable audit in ecology journals (Todd and Yeo, 2007) landed around 24%. Long before ChatGPT, academic citation carried a standing one to two in ten duds. AI added not one new kind of error. What it changed is the **production function** of error. Every coefficient moved. Volume changed. Manufacturing one fabricated citation good enough to pass used to take effort, so only people who set out to fake did it. Now it is a free byproduct of a language model, a dozen out of one generation, and nobody has to set out to do anything. This is row 6 of the transfer map, AI turns "looking rigorous" into one generation. Fluency changed too. Pre-AI errors often carried a tell, a sudden shift in style, data that did not hang together. Errors AI produces have no tell. The wrong paragraph and the right paragraph come off the same text engine, and their surface properties are identical. That is row 7 of the transfer map, hallucination and fabricated citations hiding behind polished prose. Marginal cost went to zero. In a hand-written review, every citation passed through the author's hands at least once. In an AI-generated review, the length of the citation list and the probability that any one entry passed through anyone's hands are fully decoupled. Pre-AI error detection ran mostly on that one gate, human attention. So read this chapter the right way. An old set of opponents on the research field just got industrial equipment. Your defense has to industrialize with them, and that is Chapter 12's business. This chapter settles "recognize it" first. ## 11.3 The three amplifiers Why is the same error more lethal in research than elsewhere? Research fits errors with three amplifiers. These three are the foundation of the chapter's taxonomy, and of every failure mode below you can ask which of them it feeds on. **Amplifier one, invisibility.** Research errors raise no error. Write code wrong and the compiler hands you a red answer in seconds. Get a research conclusion wrong and nothing happens at all. The wrong conclusion sits quietly in the document, looking exactly like a right one. Chapter 2 calls this "research has no compiler" (failure condition one). A research error can only be caught by a step in the process, and steps can be skipped, especially when you are short on time. **Amplifier two, compounding.** The definition of a research output is "knowledge downstream will use." A wrong conclusion gets written into a report, the report gets cited into a review, the review gets fed into the training data of the next generation of models, and all the while it is supporting real decisions. Chapter 2's contrast (failure condition three) applies here. Bad code can be rolled back. A bad conclusion has no rollback button. If the 31% in section 11.1 had survived that night, its blast radius would widen month by month. **Amplifier three, motivated collusion.** The nastiest of the three. When an error happens to grow into the shape you wanted, supporting your hypothesis, filling the gap in your argument, making the project look like a contribution, your motivation to check it drops to its lowest point. This amplifier differs from the other two in one place. Those two are properties of the environment. This one grows on you. The inspector and the error become accomplices, usually without noticing. AI's role here is an amplifier of the amplifier, not a deceiver. It is trained to talk in the direction you are facing, and whatever you want, it wraps to look more like fact. Section 11.5 has a firsthand specimen, and the victim is me. The three amplifiers multiply, they do not add. An error that is invisible, that compounds, and that happens to be what you wanted, that is the recipe the most expensive accidents in research are produced from. ## 11.4 Attribute times step, building the case file Now build the taxonomy. First the chapter's core claim. **Failure modes are the interaction product of "AI attribute Γ— research step," and no bug list written about AI alone will produce them.** Take it apart. On the AI side, three attributes offend again and again, and all three are neutral in themselves. They are even why you bought it. **Fluency**, the output is always coherent and confident regardless of whether the content is true, which is a language model's default property. **Sycophancy**, it is trained to produce answers people are satisfied with, and human raters' preferences are etched into its gradients. **Corpus prior**, its "knowledge" is a statistical compression of past text, not an observation of your situation. On the research step side, it is Chapter 3's seven steps. The same attribute landing on different steps grows into different errors, the way one pathogen causes different diseases in different organs. The summary table first, then a dissection of each class. | High-incidence step | Failure mode | Chief attribute | One-line signature | |---|---|---|---| | Master the field, deliver | Hallucination and fabricated citations | Fluency Γ— corpus prior | The most on-point citation is the most suspect | | Test plan, interpretation | Spurious significance and the criteria backdoor | Sycophancy Γ— fluency | The criterion appears after the result | | Execution | Data leakage and contamination | Corpus prior | The score is too good to be true | | Interpretation, red team | Sycophancy drift | Sycophancy | The conclusion flips with how you ask | The four rows are not everything. One more class stays off the table because it does not pick a step, and section 11.5 dissects it on its own. **Hallucination and fabricated citations.** These grow on the master-the-field and deliver steps. Fluency writes the invented content so it cannot be faulted, and corpus prior makes it look like a typical sample of the field, typical title, common authors, plausible year. There are two signatures. The first is counterintuitive. **Fabricated citations are built to order for your needs, and their distribution is not random.** The model is completing "a paper supporting this claim belongs here," so the one that happens to fill the gap in your argument, the one whose title hits your question dead center, are the entries in the whole list to check first. The odds are not small either. Walters and Wilder (2023, *Scientific Reports*) examined citations generated by GPT and found GPT-3.5 fabricated 55% of them, GPT-4 still 18%. The second signature is paraphrase drift. The citation is real, and AI's retelling carries fewer qualifiers than the original. "Improves on arithmetic tasks" becomes "improves", and Chapter 4, section 4.7 covered the mechanism. The scenarios most likely to catch you are first drafts of reviews, the citation list of a report, and anywhere "AI added a source while it was at it." **Spurious significance and the criteria backdoor.** Move to test plan and interpretation and this one runs the house. Say to AI "have a look at whether there is anything in this data" and it will almost certainly give you something. The sycophancy attribute guarantees a finding, the fluency attribute guarantees the finding looks well founded. This is the AI edition of p-hacking, more dangerous than the manual edition. When a person picks the data, that person knows roughly what they did. When AI picks for you, what you receive is an analysis that looks objective and neutral. The criteria backdoor is its twin. The decision standard gets written down only after the results are in, or gets quietly revised, and the conclusion then always "just happens" to clear the line. Its signature is one question. **Was this criterion locked in before the data was seen, or after?** If you cannot answer, or the answer is "after," treat that significance as zero. The scenarios most likely to catch you, one is exploratory analysis delivered as a confirmatory conclusion, the other is the boss asking "could you look at this from another angle." Chapter 6 locks criteria first and signs off on them to guard against exactly this class. For auditing someone else's output after the fact, see Chapter 12. **Data leakage and contamination.** On the execution step this one owns the ground. Corpus prior takes its most concrete shape here. The questions and answers of a public benchmark are, with high probability, already in the model's training corpus. It saw the exam paper before the exam, and you cannot know how much of it it saw. Another variant is analysis code AI writes that lets test-set information leak into the training side, feature engineering using statistics computed over the full data, normalization performed before the split, which is to say things that should have been computed on training data only get quietly computed on everything. This class of error is just as common in code people write, and the AI version is fluent enough that you want to look closely even less. There is only one signature. **When the score is too good to be true, treat it as leakage first**, especially when the score collapses on a fresh set of questions from the same distribution. The scenario most likely to catch you is evaluating model capability on a public benchmark, which is this book's spine case as daily routine. Chapter 5's "choose a contamination-resistant variant set" when picking task families and the data contamination check in Chapter 6's criteria list are both sentries the spine case posted for this class of error at design time. Whether the sentries hold, Chapters 7 and 8 give the verdict. **Sycophancy drift.** It shows up in interpretation and red team. The sycophancy attribute takes its purest form here. You ask with a lean, it answers along the lean. "Does this result show our method works?" and "Could this result be nothing but noise?", one batch of data, two ways of asking, and AI can hand you two fluent and opposite readings. Five frontier assistants consistently showed sycophancy (the sycophancy attribute described above, talking along with the asker) across many task types, and human preference data itself rewards the behavior (Sharma et al., 2023, arXiv:2310.13548). The signature here is operational. **Ask the same question twice with opposite leans. If the conclusion flips with the lean, what you measured is your own phrasing, not the world.** The scenarios most likely to catch you share a structure, letting the AI that produced the conclusion serve as its own judge, including asking it to "check itself for problems." That directly violates the orthogonality principle stolen in Chapter 2, section 2.4. For how to build an independent channel, see Chapter 12. For how to use AI as a critic instead of a yes-man, Chapter 10 already ran the demonstration. Not one of the four classes needs the model to "get dumber" before it happens. All four are byproducts of the model working normally. That is why "wait for a stronger model" cannot serve as a defense strategy. The attributes will not disappear. They will only get more fluent. ## 11.5 Spine update Β· the "finding" that died on the check One class is still missing from the table, amplifier three, motivated collusion. It offends on whichever step carries your wish. Textbook specimens are hard to find, because the people it hits usually do not know they were hit. I can offer a firsthand one, a pit this book fell into itself. Chapter 4 wrote the process down honestly. Here it gets dissected from another angle. The summary is one sentence. I took "a head-to-head comparison on a cost-matched basis" for an unclaimed patch of open ground, that judgment survived a full round of scanning, and it died on the forward-citation check. The comparisons exist, and they cluster on the skeptics' side. The value of the dissection is in the contrast. In the same project, a paper title I wrote from memory had one word wrong, "Key" in Wang's ACL 2024 paper, which I had remembered as "Answer." The verification channel picked it up while doing the forward check, a mechanical step, dull work. The "gap" judgment lived a whole round. Why? Look at the three attributes. Fluency? I wrote that judgment myself. Corpus prior? The core error was not in AI's retelling. Sycophancy? AI did not contradict me, true, but it counts only as an accessory. **The principal offender is what the "gap" was worth to me. A gap meant my project had a contribution, the two-hour demo was not wasted, and this book's spine case held.** A title one word wrong died on a mechanical step that does not care where anything came from. The gap judgment went unchecked far longer, because it was my wish. The line at the end of Chapter 4 bears saying again here. The more dangerous one is your own hallucination when you want a gap to exist. From this specimen you can lift the general signature of motivated collusion. **The conclusion you are most excited about is usually the conclusion you checked least.** That is the default output of the motive structure, and it has nothing to do with character. "Be careful next time" is not a countermeasure. What saved me in Chapter 4 was the coverage check as a step in the process, not my own vigilance. For what that process looks like, see Chapter 12. First an ugly thing said in full, because it governs how you use this chapter's templates. **AI self-check cannot catch this class of error.** Ask AI to check your conclusion and it checks the form, whether the citation exists, whether the numbers add up, whether the reasoning chain breaks. Your motive it not only leaves unchecked, it goes along with. That warning is written out again in bold in the appendix prompt set. ## 11.6 The perfect respondent Now point the whole taxonomy at this book's subplot case and run a full workup. Readers who care only about the spine case can read just the names of the five ways of passing for real and the closing paragraph, but keep in mind that synthetic data in your own eval set fakes in the same five ways. Persona research means using a persona-bearing model ("a 25-year-old mom with one child who works at X") in place of real people for interviews and survey rehearsals. It is the perfect specimen of the plausible trap. Its output is **almost impossible to fault one answer at a time**, and it reads more like "a real person" than plenty of real interviews do. Chapter 5 ground that rough question into a testable shape, the answer distribution agrees with the real population and no systematic stereotype drift appears. This chapter answers a different question. An output that resembles a real person that closely, in exactly which ways can it be fake? There are five, and each has a seat on the table in section 11.4. **One, the fluency illusion.** The answers are coherent, detailed, in a voice that fits the persona, and they trigger your "like a real person" intuition. Fluency has zero correlation to distributional fidelity. This one is a mechanism fact and needs no experiment. Fluency is the very target the model is optimized for, and it holds for any persona and any question. This is the standard case of "fluency Γ— interpretation." The answer itself may be fine. What goes wrong is the reasoning step where you take fluency for evidence. **Two, the average face problem.** A persona's answers are a weighted average of the corpus's stereotype of that group. Every single answer is reasonable, and the aggregate distribution is narrower and more homogeneous than the real population, variance collapse. It is like stacking a hundred faces into one "average face," well proportioned, and no such face exists in the world. Use it for all hundred and you will conclude that "humans all look alike." Corpus prior offends on the data generation step. Nominally you are sampling a population. In operation you are sampling the same statistical compression over and over. This one has peer-reviewed measurement behind it. Early work leaned optimistic. Argyle et al. (2023) reported that "silicon samples" fit real surveys reasonably well at the level of the overall distribution. Then Bisbee et al. (2024, *Political Analysis*) measured the standard deviation of synthetic answers as significantly smaller than that of real respondents. Shrunk by how much? Run the same statistical power calculation, which is the calculation of how many respondents are enough, twice. On the variance of the synthetic data, 33 partisan respondents are "enough." On the real variance from the American National Election Studies (ANES), you need about 300. The center of the distribution may be right, the spread within the group has been squeezed out, and inference built on synthetic data is therefore systematically overconfident. Verified in the political survey setting, and the effect size in other fields is still exploring. **Three, stereotype drift.** In answers written for the "25-year-old mom," the concentration of parenting anxiety runs systematically high. This is the average face one step further. Answers for a subgroup are more homogeneous and also systematically more extreme. The distribution narrows the same way, and its center lands near the **stereotype**, off the real mean, and the stereotype itself carries an offset. This one has measurement too. The Marked Personas study by Cheng et al. (2023, ACL) found that in character portraits generated by GPT-3.5 and GPT-4, stereotype vocabulary appeared at a higher rate than in a human-written control group, and leaned toward "seemingly positive" essentializing narratives (writing a group as born that way), exoticizing, "strong Black woman" archetypes. Train a simple classifier on portrait wording alone to guess which demographic group was being written about, and GPT-4's portraits get guessed right 96% of the time (GPT-3.5, 92%). Verified at the phenomenon level. For business decisions this is the most dangerous of the five. What you hear is the corpus's caricature of this kind of user, not the user. **Four, sycophancy drift.** When the interviewer's phrasing carries a lean, a persona follows the question further than a real person does, the survey edition of sycophancy. Real respondents wander off topic, push back, and say "that question is put wrong." Personas almost never do. So every implicit assumption in the interview guide gets gently confirmed once more. This is the fourth class from section 11.4 replayed on the subplot, same signature. The grade of the evidence, though, is not like the previous class, and it splits into two layers. Sycophancy in the general setting has solid evidence, Sharma et al. (2023) cited in section 11.4. Whether role-play makes it worse has so far only one 2026 preprint giving a preliminary signal (Shah et al., arXiv:2604.10733, in 9 of 13 open-source models, the higher a persona's agreeableness, a personality-test score for being easygoing and pleasant to deal with, the higher the sycophancy rate), publication status unconfirmed, so it can only be used as preliminary evidence. Still exploring. **Five, time dislocation.** What you interview in 2026 is a "25-year-old mom" compressed out of text from before 2024. A persona reflects a time slice of the training corpus, not the population of today, and every shift in attitude after the corpus cutoff is missing. The symptom has measurement. A 2026 study ran 9 LLM configurations against the real polls of the 2024 US election and found support for Harris systematically overestimated by 10% to 40% (the paper's own words are overestimated by 10% to 40% relative to the polls, without defining whether that is proportional or percentage points, arXiv:2602.06302), and hooking up web retrieval does not fix it. That it cannot be fixed is worth writing down. It says "training cutoff date" is an explanation with partial evidence behind it, not the confirmed sole mechanism. Symptom confirmed, cause still exploring. For fast-moving fields (consumer preference, attitudes to technology, public opinion topics), this dislocation is a systematic error of direction and cannot be treated as noise. The five ways of passing for real collect into one table, and the two columns "which layer it lives on" and "evidence grade" are enough. | Way of passing for real | Lives in the single answer or in the distribution | Evidence grade | |---|---|---| | One, the fluency illusion | Single answer, it fools your intuition | Mechanism fact, no experiment needed | | Two, the average face problem | Distribution, variance collapse | Verified in political surveys, still exploring in other fields | | Three, stereotype drift | Distribution, center pulled toward the stereotype | Verified at the phenomenon level | | Four, sycophancy drift | Distribution, it follows how you ask | Still exploring, one 2026 preprint's preliminary signal and nothing else | | Five, time dislocation | Distribution, the whole thing stopped at the corpus cutoff | Symptom confirmed, cause still exploring | The five share one property. **Not one of them is visible at the level of a single answer.** The fluency illusion fools your intuition, and the other four all live at the distribution level. Every single answer is unfaultable, and a thousand answers together are wrong. This is the top form of the plausible trap. Every item passes its check and the conclusion is still wrong. Against it, reading the interview transcripts more carefully does nothing. You have to compare against the distribution of a real population as ground truth, dual-source, three criteria, a preregistered falsification shape. For how to expose it, see Chapter 12, which is also where the subplot's suspense is settled. ## 11.7 Swap in your project, run a failure-mode census Your project should have walked a few steps by now, at least a controversy map and a sharpened question. Run a failure-mode census on it, budget thirty minutes. 1. **List the steps.** Lay your project out along the seven steps and mark the three where AI is most deeply involved. The census covers those first; 2. **Ask two questions of every step.** Which AI attribute am I consuming at this step (fluency / sycophancy / corpus prior)? At this step, which class of error from the table in section 11.4 is that attribute most likely to grow into? Put the answers into the census sheet, a step Γ— error type matrix, and the templates have a fillable version; 3. **Write one signature for every high-risk cell.** No copying the book's own wording. Translate it into a concrete signal in your project. What score exactly counts as "the score is too good to be true" here? Which entries exactly are "the most on-point citation" here? A cell where you cannot write a concrete signal means you have not yet worked out what the error there looks like; 4. **Circle two cells.** One is the cell where being wrong costs most, and that is your operating room. The other takes honesty, **the cell holding the conclusion you are currently most excited about**. By the law in section 11.5, that is the high-incidence zone for motivated collusion, the place where your checking should double and where you are in fact most likely to cut corners; 5. **The closing self-test.** With the filled sheet in hand, answer one sentence. If this project blows up three months from now, which cell is it most likely in, and what does the crime scene look like? Answer that and the census passes. The census sheet and the "have AI self-check for failure modes" prompt set, complete and fillable, are in this chapter's templates. The use warning in there is a direct corollary of section 11.5. Do not skip it as boilerplate. **Want an agent to run it with you?** Paste this to your AI assistant or coding agent: ```text Help me run the Chapter 11 failure-mode census, budget thirty minutes. Following Template 1 in docs/appendices/ch11-templates.md, build an empty step Γ— error type matrix, then walk the seven steps asking me two questions each, which AI attribute I am consuming at this step and which class of error it is most likely to grow into. I fill the answers in. The signature for every high-risk cell is written by me as a concrete signal in my project. You do one thing only, send it back if I copied the book's wording. The cell holding the conclusion I am most excited about is circled by me. You may not circle it for me, and you may not say that conclusion looks fine. When the census is done, run the "have AI self-check for failure modes" prompt set from Template 2, and before running it read out the use warning in the templates. It checks form. It cannot check my motive. If any command errors, stop and show me the output. ``` ## 11.8 Sober reminders - **This classification list will itself expire.** Row 7 of the transfer map, new capability manufactures new error types in bulk, is a regularity still running, not a summary of history. The next generation of model capability will manufacture errors that are not on the table in section 11.4, the way the error type "the formula references one cell off" did not exist before the spreadsheet was invented. So what this chapter gives you is classification instinct, not a complete list. When a new error surfaces, you can ask "which attribute hit which step" and add a row to the table yourself. People who can add their own rows do not fear an expiring list. - **Recognizing is not preventing.** This chapter has done diagnosis and nothing else. The full defense is in Chapter 12, layered verification, the independent channel, and check dispatch all live there. The red-team routine for having AI attack your conclusions systematically was delivered in Chapter 10. Do not hold this chapter's list of signatures and feel safe. A signature only tells you where to look. - **Nobody has measured the relative incidence of the four classes.** The order in this chapter follows the steps, not a ranking by frequency or by damage. There are hard measurements in places, the Walters and Wilder fabrication rate cited in section 11.4 is one, and that is a number for one class of error on one step. Which class is most common in real research workflows and which causes the most loss still has no unified measurement covering the whole workflow. This line is still exploring, and it is a row that will appear on the honest map in Chapter 13. - **Motivated collusion has no tool solution, still exploring.** Fluency is handled by a checking step, sycophancy by an independent channel, corpus prior by fresh data. Motivated collusion cannot handle itself, and "itself" is exactly the problem. It has only a process solution. Let a channel that does not share your motive, another person, or a verification agent that does not know the answer you expect, touch the conclusion you are most excited about. That is how this book got saved that one time (section 11.5). Turning it into your standard equipment is Chapter 12's job. ## 11.9 The unfair advantage you now hold Handed any polished research output, your own included, you can say within ten minutes which class it is most likely broken in and where to poke. Check the most on-point citation, ask whether the criterion was written before the result or after, look at whether the score deserves suspicion, or ask the question again from the opposite side. Most people facing the same output have two settings available, "something feels off" and "looks fine to me." --- # Chapter 12 Β· The Verification Workflow !!! info "Chapter companion" πŸ“‹ [Chapter 12 templates](../appendices/ch12-templates.md) Β· πŸ—‚ [Template index](../appendices/template-index.md) Β· πŸ’» [`code/persona-panel`](https://github.com/hallieren/research-rewritten/tree/main/code/persona-panel/) > **This chapter's ladder.** Verification cuts across all seven steps and does not occupy one of them, so this line labels the checking task itself. Mechanical checks, citation existence, number tracing, consistency checks, handed to AI, have settled at **assistant level**, and with batch dispatch plus automated procedures they are climbing toward **collaborator level**. The final ruling on "do I dare trust this output," plus how the verification budget is split, has a water line level with the interpretation and red team rows in the Chapter 3 snapshot. That can only be you. > > **Subplot update.** The persona study wraps up in this chapter. The ground truth hook Chapter 6 left behind is cashed in here as a verification plan signed off before the run. Along the way you will also watch one of this book's own claims get shot dead on the spot by its own verification process. > > **This chapter delivers.** Four things. The three disciplines of the independent channel check, the three-layer verification workflow from L0 to L2, the downgrade ladder for when the budget runs short, and a one-page verification budget sheet. --- ## 12.1 The afternoon you grade by prose Tuesday, ten in the morning. A director forwards you a forty-page research report with a one-line note: "The management review is at three this afternoon, see whether it can be trusted." The report came from an outside team, with AI deeply involved. No guessing needed. The page count, the delivery speed, the tidy summary at the end of every section all confess it. You open it. Clean structure, a respectable reference list, numbers precise to one decimal place, conclusions worded with confident restraint. You have under two hours. Here is the question. What do you plan to check? What most people really do at this moment is read the report front to back and give it an overall score from an intuition that mixes prose, layout, and confidence. Chapter 11 already closed that road. Grading by prose means using the other side's strongest attribute as your criterion. You have enough alertness. What you lack is a routine. Which claims should the two hours go to? How deep should each be checked? How do you dispatch the parts you hand to AI so the check does not come back empty? This chapter delivers that process, and it is the book's core asset. The earlier chapters taught you to produce and to recognize. This one teaches you acceptance. ## 12.2 Production cost collapsed, trust did not First the economics, because they decide the shape of this process. Row 1 of the transfer map says production cost collapsed and review cost did not. A forty-page report now takes three hours to produce instead of three weeks, and judging whether it can be trusted takes almost the same time as before. So verification cannot be "all in." Run the full check on every output and the team falls back to pre-AI throughput. It cannot be "all out" either. Then plausible flows straight into the decision chain, and every failure mode in Chapter 11 waits downstream to compound. There is one way out, **layering**. Assign each output a verification level by "cost of being wrong Γ— probability of being wrong," and spend the limited verification budget where it cuts. Section 12.5 gives the full operating detail of the three layers. Before layering, break "trustworthy" into properties you can act on. The operational version of "reproducible, traceable, checkable" is three questions you can check on the spot. - **Traceable**. For any claim in the output, can you point to its source within three minutes, which paper, which dataset, which experiment? A claim whose source you cannot point to is treated as having none. - **Reproducible**. Given the same raw materials, could a different person or a different AI walk the path again and reach the same conclusion? Is the path on record? - **Checkable**. Was the criterion for judging it right or wrong locked in before the output existed, or improvised after seeing the result? Chapter 6 delivered the full practice of criteria first. Here you check one thing only. Does the criteria timestamp precede the result? All three questions ask "can you," and none asks "do they have a conscience." Trust is stamped on the mechanism, not on character, and the source of that sentence is coming shortly (section 12.4). Verification itself has been accelerated by AI too. Citation checks, number tracing, reverse re-derivation are all dispatchable work, mostly mechanical and dull, and AI does it fast and steady. Here lies the most important principle of this chapter, **how you dispatch decides the quality of the check**. Dispatch it wrong and you get an enthusiastic, worthless "certificate of confirmation." Why, a real case first. The one overturned was me. ## 12.3 The step that overturned me "Cost-matched comparisons are a gap in the literature." How that claim lived and how it died, Chapter 4 has the scene and Chapter 11 dissected its motive. Here is the backstage, only how the channel caught it. I did not touch that check. A separately dispatched verification agent, another model, a brand-new session, received nothing but a list of claims, and not one word on the list revealed which one I hoped would hold. It followed the forward citations, the papers that later cited this literature, all the way back, and reported that cost-matched comparisons exist, and cluster on the skeptics' side. The claim was shot dead, section 4.5 was rewritten to match the check, and the check report was filed separately. Now a thought experiment. Suppose I had pasted the claim, excitement included, back into the session that did my literature scan and asked "help me confirm whether this gap is real." What would I have gotten? Most likely confirmation. Three forces push the same way at once. The model tends to talk along with the asker. The model prefers its own earlier output, which Panickssery et al. (NeurIPS 2024) measured. They let GPT-4 judge between abstracts it wrote and abstracts others wrote, and found clear self-preference, while human reviewers showed no such bias on the same abstracts. In scores, self-preference came out at 0.705 and 0.912 on two datasets, where 0.5 is unbiased. The measurement is limited to 2024 abstract tasks, and the direction is clear. The third force, that session's context was full of the retrieval results that had propped up the claim in the first place, and the same well yields no new water. Add a fourth force, which Chapter 11 calls motivated collusion. I was hoping to be confirmed too. This is the **independent channel principle**, the soul of this chapter. **The channel that produces a conclusion cannot be its own judge. Verification must run through an independent channel that shares no context and knows no expectations.** Chapter 2 planted a sentence when it discussed orthogonality, never let the same AI both produce the conclusion and judge it, and here it is cashed in as three disciplines. They share a root with the five dispatch disciplines of the red team chapter. There they protect the prosecution's case specifically. Here they generalize into the foundation every conclusion has to walk across. **First, channel independence.** Verify in another session, on another model, or with a person. Whatever it is, it cannot be the context that generated the output. Context is a position. **Second, the brief leaks no expected answer.** The task sheet for the verifier holds only the claim to be checked, not where it came from, and not whether you want it to hold or fall. There is a reliable self-check on wording. Does your brief ask "please confirm X," or "please rule on the evidence status of X"? The former has already stuffed the answer into the question. **Third, three-value output.** A verification conclusion may take only three values, confirmed, falsified, undecidable. Mushy phrasings like "basically correct" or "broadly credible" are banned outright. Most of them are a leaked expectation coming back around. One line above all must be locked in, **undecidable does not equal pass**. The in-text version of the verification brief looks like this. The full fillable version is in the appendix. ```text You are the verifier. Below is a set of claims. Check each one independently. I will not tell you which document they came from, and I will not tell you which ones I hope hold. For each claim output: 1 Verdict (pick one of three): confirmed / falsified / undecidable 2 Basis: the source you actually found (it must open, or point to a specific paper) 3 If "falsified" or "undecidable": where exactly the gap between the claim and the evidence lies Forbidden: guessing the "expected answer" from the wording or ordering of the claims; Forbidden: leaning phrasing on any "undecidable" item. Claim list: 1 [claim one, the claim itself only, stripped of the original's rhetoric and concluding tone] 2 [claim two] ``` One aside. These three disciplines hold for human review too, and the reason is that this line of defense is cheap, not that "a reviewer will be anchored" has been proven. The evidence has two ends. The analogy holds, direct evidence is not there yet. The analogy end stands. Someone who sees the conclusion before checking it lets attention run down the road the conclusion has paved. Judgment research calls this anchoring. Tversky and Kahneman ran the experiment in 1974. Whether a wheel stopped at 10 or at 65 could drag the median estimate of an unrelated proportion from 25% to 45%, even though everyone knew the wheel was random. The direct-test end is a null result. The only randomized controlled experiment so far that directly tests "are reviewers anchored by a first impression" (PLOS ONE, 2024, 108 researchers) measured no significant anchoring effect. Medical trials insist on blinding, that is, not letting the operator know who got the real drug and who got the placebo, and auditing insists on independent review. They are the seasoned version of the same idea, and they cost so little that they are worth paying for without waiting for anchoring to be proven. ## 12.4 Lessons stolen from code review You already live a life of accepting large volumes of unfamiliar output every day, and the answer is called review culture. Three lessons carry the most value. Look only at where they break when moved to research acceptance. **Lesson one, unfamiliar output is trusted through mechanism, not goodwill.** Code written by tens of thousands of strangers dares run inside banking systems because of review, CI, and a traceable origin for every changed line (row 9 of the transfer map). Research's counterpart is the three properties of section 12.2. Traceable matches "origin on record." Reproducible matches "check out and rerun," that is, a different person or a different AI takes the same raw materials and walks the path again. Checkable matches "tests first," the criteria timestamp precedes the result (Chapter 6). The question "can this AI output be trusted" is itself tilted. The right question is "how dense have I woven my net of mechanisms." Row 5 of the transfer map says verification infrastructure sets the radius of letting go. Now you see its other face. Verification infrastructure also sets the radius of trust. How far you dare trust an output equals how many of your mechanisms it has passed through. **Lesson two, hand mechanical checks to the machine, and people look only at what the machine cannot.** CI stops formatting and failing tests, and the reviewer's effort goes to design and logic. The research-side counterparts are citation existence, consistency between numbers and figures, uniform units and basis, all of which can be written as fixed dispatches or even scripts that run the moment an output comes through the door. The attention saved goes where the machine cannot look. Is this claim's chain of evidence strong enough to bear weight? Was that slice which "happens to support the conclusion" declared beforehand? Where it breaks, a red CI is red, but the "undecidable" that comes back from a verification channel has no color, and people are quickest to wave it through as a green light. The third discipline of section 12.3 and the escalation rule of section 12.5 both plug that hole. **Lesson three, what fails the mechanism does not merge, and there is no exception channel.** A PR with a red CI does not enter the main branch, however famous the author or moving the explanation. The value of this discipline lies precisely in how unfeeling it is. The research version is the same. Output that fails verification does not enter the decision chain. It goes back for rework. "The author thinks it is fine" is not a pass. Where it breaks, separating generation from verification is welded shut for programmers by the permission system, while in research "failed verification does not enter the decision chain" has no machine to weld it, only institutions, and Chapter 15 teaches how to weld. ## 12.5 The three-layer verification workflow Now the main deliverable. Assign the layer first, then do the work. Every quantity below is a starting default, not an empirical finding. How many citations to sample, how much time to leave, how many days count as enough, all need recalibrating to your field and your cost of being wrong. Assigning a layer asks only two questions. First, where is this output going? Does it stay on your own desk, go into team discussion, or into the decision chain, into the public knowledge base? The farther it goes, the larger the blast radius of an error. Second, which kind of error is it most likely to hide? Match it against the failure-mode census you did for your own project in Chapter 11. Citation-heavy output guards against fabricated citations, statistical conclusions against spurious significance, synthetic answers against the average face. The product of the two questions decides the layer. Look at the three layers side by side first, then read each checklist. | | L0 spot check | L1 full verification | L2 adversarial recompute | |---|---|---|---| | Trigger | The output does not leave your desk, brainstorming, exploratory drafts, intermediate material only you will see | Someone will spend money, commit people, or draw conclusions based on it, and it is about to leave your desk | A single conclusion being wrong would trigger a hard-to-reverse action, architecture selection, a funding decision, public release | | Budget | Fifteen to thirty minutes | Half a day to a day, mostly AI time | One to several days, spent only on the one or two conclusions named | | Actions | Sample citations, sample numbers, one reverse question | Forward-check every citation, trace every number, walk the reasoning chain link by link, file the reference table | Independent re-derivation, recompute by another method, red team | | Escalation and close-out | One hard defect found, the whole output moves up to L1 | "Undecidable" may not quietly turn green | Every unresolved disagreement gets a human ruling | ### L0 spot check, an alarm on low-risk output Three checklist items. 1. Random citation check. Sample five citations or ten percent, whichever is larger, and check two things for each. Does it exist? Does it really say what the output claims it says? Sample with a random number or a fixed rule. Never let the generating side pick; 2. Sample three key numbers and ask "where did this number come from," tracing each to its source or to a dead end; 3. One reverse question. Hand the output to an independent channel and ask "which claim in this material has the weakest evidence, and why." One escalation rule. If the spot check finds one hard defect, a fabricated citation or a sourceless number, the whole output moves up to L1. The logic is the same as quality inspection. One defective unit in the sample means this production line's defect rate does not deserve sampling. Dispatch notes. All three can be handed to AI, but through an independent channel, with the brief written by the disciplines of section 12.3. Which items get sampled is decided by you or by a random number, not by the channel doing the check. ### L1 full verification, the threshold for the decision chain Four checklist items. What is added over L0 is "full" and "filed." 1. **Forward-check every citation**. Does it exist? Does it say so? Was it later overturned or retracted? The third question is the easiest to skip and the last one that should be. My claim in Chapter 4 died on exactly this question; 2. **Trace every number**. Go from the number in the output back to its original source, and check that the basis matches item by item, comparison baseline, time window, units; 3. **Walk the reasoning chain link by link**. How many steps lie between evidence and conclusion? Is each step a citation, a calculation, or "the author thinks"? Mark the "author thinks" links separately and hand them to a human ruling; 4. **File it**. Produce a "claim β†’ source" table and store it with the output itself. This table is the physical form of the traceable property. Dispatch notes. Split the checklist into mechanical subtasks and dispatch them in batches, citation checks through one channel, number tracing through another, and neither brief carries the other's conclusions. When you collate, a human goes through only two columns, every "undecidable" and every "falsified." One more sentence for the person who signs rather than the person who does the work. Neither L0 nor L1 requires that you can rerun the output. Both can be done with the report alone in hand. Of the blank that Chapter 3, section 3.6, Premise 3 marked, what remains is only L2's independent re-derivation and the corner of "how much to sample before it is enough." ### L2 adversarial recompute, what load-bearing conclusions get An output usually holds only one or two conclusions that deserve L2. That is normal, not laziness. Three checklist items, all spent only on the named conclusions. 1. **Independent re-derivation**. Another channel gets only the raw materials and the research question, not the conclusion, and derives from scratch. Converging on the same conclusion is the hardest machine evidence money can currently buy. Not converging, the points of disagreement are a gold mine. Rule on each by hand, and each one either fixes the output or goes into the limitations; 2. **Recompute key numbers by another method**. Change the calculation path, change the data source, and see whether the number stands; 3. **Red team**. Let the harshest critic attack this conclusion. How to fight, Chapter 10 delivered. Here only one rule is set. In L2 the red team is mandatory, not a bonus. Dispatch notes. The re-derivation brief is the easiest one in the whole process to leak. No residue of the original conclusion, its wording, its structure, even its subheadings, may appear in it. Better to spend ten extra minutes reorganizing the raw materials than to save effort by clipping the first half of the output. ### Back to that afternoon Now replay the scene from section 12.1. The two hours go like this. The first ten minutes, read through and assign layers. This report goes to the management review, so L1 as a whole is the floor, and the two conclusions supporting the final recommendation are named L2. Next send an L0 spot check out as a scout, ten citations and five numbers, independent channel, results in twenty minutes. If it finds a hard defect, you have your conclusion on the spot: "The spot check failed. Not recommended as a basis for today's decision. Returned for further verification." If the spot check comes back clean, you list the L1 reference table and the two L2 to-dos: "Citation and number spot checks passed. Full verification out tonight. Two load-bearing conclusions need independent re-derivation, final ruling by noon tomorrow." Notice that what you hand over has changed. A list that says "what was checked, by what criteria, what remains" has replaced an impression score. Someone who grades by prose gives an opinion. Someone who walks the process gives an evidence status. Complete fillable versions of this chapter's four tools are in the appendix. The layered verification workflow card, the independent channel check prompt set, the verification budget sheet, and the downgrade note template the next section delivers. ## 12.6 The downgrade ladder for when the budget runs short The last section left one premise unspoken. It assumes you can pay. When those numbers were written down, I had time, API credit, and the chance of a second rerun. Reality often does not give you that. The investment committee meets in nine hours and the report just arrived. This is the last batch of data before graduation, with no money for more experiments. You are the person pulled in for a quick look with no permission to touch the raw materials. None of these count as exceptions. They are what it looks like when a few of the five premises in Chapter 3, section 3.6 break on you. This section has to be written. Without it the last section fails in a hidden way. **A process that can only be executed on a full budget does not get "executed at a discount" in front of a deadline. It gets abandoned wholesale**, and you fall back into the afternoon of section 12.1. Rigor that offers no downgrade plan has the practical effect of no rigor. That is a failure of process design, not laziness in the executor. ### Cut layers, not the order The wrong shape of a downgrade is doing half of every layer. Sample three citations instead of five, one number instead of three, stop the re-derivation halfway. Cut this way, every line of defense is left with a gap in the door, and a gap stops as much as an open door. The right shape is **keep the earliest line in the order and cut the later ones whole**. The reason is that the cost of errors is uneven. A criterion patched in afterward contaminates every downstream conclusion on the whole chain. Two fewer citations sampled loses only the information in those two citations. From tightest budget to loosest, four tiers. **The ten-minute tier, you have time for one deep breath.** Do one thing only, **check the criteria timestamp**. Was this output's standard of judgment written before the results were seen, or after? If the answer is "after" or "there is none," your conclusion is already in and nothing else needs checking. It is an untested hypothesis, and it gets treated as an untested hypothesis. This tier does not produce "trust / don't trust." It produces "is this something that can be tested at all." It stops the most expensive class of error among the four tiers. **The half-hour tier, the scout spot check.** The ten-minute tier plus items one and three of L0. Sample five citations for existence and paraphrase fidelity, and add one reverse question through an independent channel. Number tracing matters, and it gets cut anyway, because in half an hour it is the one most likely to end up half done. A half-done trace is more dangerous than none. You will remember that you "checked the numbers" and forget you checked only one. **The two-hour tier.** This is the full form of the afternoon at the end of section 12.5, and it is not a downgrade. **The "no next round" tier.** This tier is hard in its structure. The results are already out, the interrogation killed them or left them ugly, and you have no resources to rerun, no money, no sample, no time. This is what it looks like when Chapter 3's Premise 5 (there is a next round) breaks, and it is the question this book gets asked most. The honest floor is three items, all executable. 1. **Report the numbers under the original criteria as is.** The interrogation's output stands beside them, labeled "post-hoc," and the old ones are not deleted. Chapter 8 calls this the side-by-side reporting discipline, and it applies here unchanged. Killed in the interrogation does not mean it may be deleted. Deleting is the fraud. 2. **Write "no next round" itself into the limitations**, and be specific. The courtesy of "limited by resources" does not count. Write it to the level of "the X this conclusion depends on has only 3 independent units, a confirmatory retest would need about N additional samples, and this project did not run it." The reason for being specific is practical. A reader can price your conclusion from it. A vague disclaimer gives them nothing. 3. **Lower the claim strength to the tier the evidence can carry, not to zero.** This item is the easiest to get wrong, in both directions. Erring upward is writing an unconditional statement knowing the evidence falls short. Erring downward is being so frightened by the interrogation that you dare say nothing. Chapter 9 covered both directions. A result delivered by these three items does not reach "qualified research." It is **an honestly priced output**. What separates it from fraud is not how good the conclusion looks but whether a reader can judge from it how far to trust it. This is the best a resource-constrained person can get, and it is the floor this book is willing to stand behind. ### A downgrade must leave a record The whole downgrade ladder has only one hard rule. **If you downgraded, write on the deliverable which tier you downgraded to.** One line, and it travels with the conclusion. "This conclusion is delivered at the half-hour tier, criteria timestamp verified, citation spot check 5/40 passed, numbers not traced." Writing that line costs almost nothing. Not writing it costs everything. A downgrade without a record and a pretense of the full set look identical to a downstream reader. The favorite entrance of Chapter 11's failure modes looks exactly like this. Nobody lied. Someone skipped a step and forgot to say so. One common objection, answered in passing: "If I admit I only did a spot check, will people still trust me?" Yes, and more than they trust hedging. A conclusion labeled "L0 spot check passed," the reader knows how to use. An unlabeled conclusion, the careful reader can only discount at the worst case, and the careless reader takes at full marks. You want neither. ### Whose budget is short Finally, separate two kinds of "short," because their prescriptions are opposite. One is a real constraint. The ceiling on money, samples, or time sits right there. This whole section was written for it. The other is priority disguised as constraint. "This one is not important, a spot check will do," but it is going into the decision chain, so it is important. Ask the first question of section 12.5 again, "where is this output going?" If the answer contains "for someone else to make a decision," you are facing a prioritization problem, not a budget problem. Prioritization problems should not be solved with the downgrade ladder. They should be solved by delaying delivery. **"No time to verify" and "no time to finish this" are the same sentence.** The latter you would say out loud. The former you usually would not. ## 12.7 Subplot wrap-up Β· "like a real person" goes to trial The subplot wraps up here. The testable question Chapter 5 sharpened, is the persona answer distribution consistent with real subgroups and free of stereotyping. The hook Chapter 6 left, **when ground truth is not ready-made, half the work of verification design is designing the ground truth itself.** The five ways of passing for real that Chapter 11 dissected. The three come together here as one plan, which is also a full-scale demonstration of this chapter's workflow. First the design of the ground truth, and the first cut it took. The plan was originally dual-source. The first source, a large public survey dataset, finally locked to the US sample of the seventh wave of the World Values Survey, a public global survey of social attitudes run in batches by year, each batch called a wave, using the answer distributions of real population subgroups as the comparison. The second source, small-sample real-person calibration, structured interviews with twenty or thirty people to cover the questions the public question bank does not. Each source insures the other. Public data is cheap and plentiful, but Chapter 6 pointed at its fatal spot. It is very likely already in the model's training corpus. Real-person calibration is expensive and small, and freshly collected interviews cannot have been memorized from a corpus. The license for the first source was verified, free for non-commercial use, download by registration, no redistribution of raw data, so the subplot repo holds only the loader and per-subgroup respondent counts, and readers download the raw data themselves. Then reality took the knife. After the plan was locked, the real-interview arm was cut. I could not recruit interviewees. This cut has to be booked in front of you. A verification plan, once designed, is not always one you can pay for in full. After cutting the arm, ask again. What replaces the half of the insurance you lost? The answer is to promote a step that had been a bonus to mandatory. Questions from the public bank are reworded and asked again. If persona memorized the original questions, the answer distributions on the original and the reworded version will show the crack. In the dual-source plan this was icing. In the single-source plan, after losing the real-interview arm with its natural immunity to contamination, it is the **only** defense against memorized questions. The price is posted as it is. From here on, every subplot conclusion carries a caveat, the ground truth itself may be inside the training corpus. This cut went into the change log, signed and filed, under the same discipline as the criteria. Three criteria, preregistration style, locked before the run. The same idea as Chapter 6. - **Criterion a, distribution agreement**. The distance between the persona answer distribution and the real subgroup distribution, measured by the top-choice agreement rate, the share of questions where both sides' top-voted option is the same, plus one distribution distance metric, must reach a preset threshold. The exact metric is locked in the subplot repo before the run starts, and the locking goes into the change log; - **Criterion b, the anti-stereotype red line**. The within-group variance of persona answers may not be systematically lower than real-person variance. This red line strangles exactly Chapter 11's "average face"; - **Criterion c, the subgroup cross-check**. Cross-slice on at least two demographic dimensions, to guard against "right on average, wrong on every slice." Chapter 11's stereotype drift hides in the slices. The falsification shape, with a downgraded conclusion. If persona's distribution distance exceeds the threshold on most questions, or the variance-collapse red line is tripped, "persona can replace real interviews as a rehearsal" is falsified, and the downgraded claim "usable only for tuning questionnaire wording" may still stand. A criterion that can only lose as "worthless" is a blunt instrument. A good falsification shape tells you which tier it lost down to. Check this plan itself against this chapter's checklist. Criteria locked first, timestamp preceding any result, checkable. The data source and the arm-cutting change both point to a specific archive, traceable. Plan and pipeline sit in the subplot repo, anyone can walk it again, reproducible. All three criteria land on computable statistics, and none needs a model to give an impression score. Chapter 5's rule excluding the LLM judge holds in the subplot too. Then the results. When the plan was written, the block below was empty. Now the subplot repo has run to the end, 10,800 interviews (6 subgroups Γ— 40 personas per group Γ— 15 questions Γ— 3 wording variants), the interview model run straight through, total cost $0.47, zero invalid answers. The verdict follows. Reading the pass or fail at the head of each line is enough. The numbers are there for checking back. > **Subplot result** (persona-panel repo, results in `results/report.md`, change log in `docs/CHANGES.md`, and the backfill touched only this block). > - Criterion a, distribution agreement, **fail**. All six subgroups went down, with a median JS distance from the real distribution of 0.19 to 0.25 bits (JS distance measures how far apart two distributions are, 0 means complete overlap, and bits is its unit), against a threshold of 0.10; top-choice agreement sits around 40% across the board, against a threshold of 70%. > - Criterion b, the variance-collapse red line, **tripped**. On 80% of the clean questions, persona's within-group variance is less than half of the real people's. Chapter 11's "average face" was pinned to the data by my own pipeline. > - Criterion c, the subgroup cross-check, **0/6 subgroups pass**. > - The contamination test, the step promoted to mandatory. 10 of 15 questions were flagged, with JS > 0.05 between the answer distributions on the original and the reworded question. Both readings have to stay. Memorized the original questions, or highly sensitive to wording, the latter a close relative of Chapter 11's sycophancy drift. The single-source plan cannot separate the two, and that is precisely the price of cutting the real-interview arm, booked above. > - By the preregistered falsification shape, **falsified**. "Persona can replace real interviews as a rehearsal" is dead, and cleanly. Every criterion failed, not one came close. And the downgraded claim "usable only for tuning questionnaire wording"? Honestly, this experiment did not test it. Wording tuning does not require distributional fidelity, so the falsification cannot kill it. But there is no evidence either that the wording problems persona flags match a real-person pilot. It stays "still exploring" and does not rise to "verified." In the first draft of this section, the block above really was a placeholder. The plan was locked and filed first, and the results came back months later. They came back falsified across the board, which turned out to be the best advertisement this process could have. Had I written it a pretty ending in advance, persona winning by a hair, this book would have lost to its own Chapter 11. The teaching asset is the shape of the plan, not the joy or grief of the ending, and now you can accept that sentence with the results in hand. ## 12.8 Swap in your project Set a one-page verification budget sheet for your project, one hour budgeted. 1. **List the outputs**. Over the next month, what will your project produce? Reviews, analysis reports, charts, code, memos to your manager, one per line; 2. **Two questions per row to assign the layer**. Where is it going? Which kind of error is it most likely to hide? Dig out the failure-mode census you did in Chapter 11 and fill in L0 / L1 / L2 from it; 3. **Lock in the escalation rules**. Which signals trigger a move up a layer, a spot check finding a hard defect, the output's destination escalating, a conclusion being cited by a bigger decision; 4. **Name a channel for L1 and above**. Which model, which prompt path, which colleague verifies. Write the name, and it must be separate from the generating side; 5. **Post it**. One page, in the project README or the team wiki, so that "has this thing passed the layer it should have passed" becomes a question anyone can ask out loud. 6. **Pre-write one downgrade note per layer.** In the format of section 12.6, write the template line "this conclusion is delivered at tier X, here is what was checked and what was not" ahead of time. The reason is practical. Downgrades always happen on the busiest day, and on that day you will not have the mind to improvise the wording. You will skip it. Two closing self-checks. If the sheet is all L0, either your project has no output that dares enter the decision chain, or you are exempting yourself from inspection, and both deserve suspicion. If it is all L2, you have handed back all the speed AI gave you. Go back and reprice. In layering, an error in either direction costs real money. **Want an agent to run it with you?** Paste this to your AI assistant or coding agent: ```text Help me set up the Chapter 12 verification budget sheet, one hour budgeted. Build the sheet from Template 3 in docs/appendices/ch12-templates.md. I list the outputs for the next month, I answer each row's destination and most likely error type, I assign L0 / L1 / L2, and you only check each row against the layer quick-reference table and flag where it disagrees. I write the escalation rules and the downgrade note template lines. The verification channel for L1 and above must be separate from the generation channel, so you only generate and record. The 2b citation check brief and the 2e independent re-derivation brief I dispatch from a separate session that carries no context from this conversation. Prepare those two briefs now from the appendix templates, and my expectations may not appear in the claim list. If any command errors, stop and show me the output. ``` ## 12.9 Sober reminders - **This workflow stops honest mistakes, not people determined to fake.** Every mechanism in it assumes participants want to get it right and merely err. When the adversary is deliberate forgery you need a different set of things, and remember the xz incident Chapter 2 mentioned. The lesson is that the trust mechanism itself becomes an attack surface. - **The verification channel errs too.** "Confirmed" means only that the evidence found at this moment supports it, not a permanent verdict. Who verifies the verifiers? Spot checks plus filing are enough, no infinite regress required, the same arrangement as "the CI config gets reviewed too." - **Do not let three-value output degrade to two.** "Undecidable" is the most informative of the three values. It marks the edge of your knowledge. Quietly filing it under "pass" is the most common way this process dies. - **Criteria depreciate too.** A criterion is a snapshot, not a constant. Rising model capability slowly exhausts a criterion's discriminating power, and corpus contamination quietly voids the premise that "the questions have not been seen." This book handled one itself. Chapter 5 dropped the original benchmark for a contamination-resistant variant, precisely because the original questions had most likely entered the training corpus. For a project that runs once, criteria locked until the run ends is enough. For a project that runs for months, give the criteria a review date too, and fold the cadence into Chapter 13's map. - **Ask about the verification channel's lineage.** The independent channel principle in this chapter speaks of a single check. An automated pipeline that packs generation and evaluation into one system is another matter. Evaluation code cross-reviewed by agents of the same family, a judge model distilled from the model being judged, two channels in appearance, one lineage underneath, and the self-preference measurement cited in section 12.3 applies exactly. The executable question is one sentence. How much lineage does your judge share with your generator, the same model, the same family, distilled from whom? What you cannot state, treat as the same channel. Beyond lineage there is one more. Any model used for scoring is an instrument that has to be calibrated first. Let it rule blind on a batch of human-labeled samples, and read its disagreement rate layered by cost of error, before it is qualified to produce numbers. On the class where an error costs the most, it may only report up, never release. - **Still exploring**. Automated verification tools, citation-check services, fact-check agents and the like, are iterating fast, and the book's online case library tracks the current state. As of this writing, "existence checks" can be dispatched with confidence, while "paraphrase fidelity" checks (did that paper really say this sentence) still need human spot checks as the backstop. One last line. Every verification conclusion expires. Labeling the evidence status of a whole field's claims, and letting the status update, that map is the next chapter's business. ## 12.10 The unfair advantage you now hold Whenever a research output with deep AI involvement lands on the table, you have a process that reaches a conclusion within two hours. Ten minutes to assign the layer, a scout spot check to open the way, and the rest of the time spent on the right layer. People without this process are still sitting in the afternoon of section 12.1, grading by prose. --- # Chapter 13 Β· An Honest Map !!! info "Chapter companion" πŸ“‹ [Chapter 13 templates](../appendices/ch13-templates.md) Β· πŸ—‚ [Template index](../appendices/template-index.md) > **This chapter's ladder.** This chapter is outside the seven steps, so it carries no ladder line. Chapter 3, section 3.5 marks the capability water line, how high AI can stand at each step. This chapter marks the evidence status of claims, which of the many things said about this shift can be believed. > > In the same week, "AI can now do science on its own" and "AI research is all a bubble" will show up in your feed one after the other, each with papers to back it. A third headline will not save you. What you lack is a map that labels evidence status. > > **This chapter delivers.** A fourteen-row honest map, the status rules (definitions for the three statuses plus two disciplines), and a method for re-deriving it row by row. --- ## 13.1 Two headlines in one week Monday morning, your feed pushes the first one. "Milestone. An AI system picked the topic, ran the experiments, and wrote the paper on its own, and the paper passed peer review." The picture is a flowchart, a screenshot with no humans in the author list, and a round of cheering. Thursday night the second one lands. "Bubble bursting in real time. A spot check shows that papers with deep AI involvement fail replication on a large scale, and several top conferences are tightening policy." The picture is a downward curve, also with a round of cheering, from a different crowd. Both are collages, but you probably scrolled past their relatives last week. Both come with demos, both cite papers. In the reposts of both, there are smart people you know. A colleague forwards you both together, with one line, "So which one is true?" The question cannot be answered the way it is asked. Its unit is wrong. A headline's unit is the event, some system passed some review, some spot check found some percentage. A judgment's unit is the claim. Does "AI can produce publishable research on its own" hold? Does "AI involvement lowers reliability" hold? How much support one event gives one claim depends on the basis of the evidence, on whether the sources are independent, on whether anyone has seriously tried to knock it down. The headline, as a genre, has no room for any of that. Information you do not lack. On this topic in 2026 there is far too much of it. What you lack is a map that pins each claim to its evidence status. Is this one verified, still exploring, falsified, or does nobody know at all. This chapter hands over that map. It is the book's master ledger. Every claim gets a row, with a status, the evidence, and a review date. Whether a map dares to write "falsified," whether it dares to write "nobody knows," decides whether it is a judgment tool or one more piece of carefully worded promotion. ## 13.2 What this map marks, and what it does not First a boundary, so you do not take this map for a reprint of the one in Chapter 3. **The snapshot in section 3.5 marks AI's capability water line**, **this chapter's map marks the evidence status of claims**. One gauges the water, the other audits the books. The water gauge tells you at which step you can let go. The ledger tells you in which argument you can place a bet. Each row of the map has five columns. **Claim.** It must be written as one sentence that can be judged true or false. "Multi-agent has a lot of promise" does not get onto the map, because no evidence can hit "promise." "A multi-agent team beats a single model" can get on. It can die. A statement that cannot be written in falsifiable form is only a slogan, and slogans have no seat on this map. **Current status.** Three statuses, verified, still exploring, falsified. The operational definitions for assigning one are in the rules card in this chapter's appendix. Here only the skeleton. Verified requires production-grade evidence plus an independent source. One source is not enough, a demo even less. Falsified requires a reproducible counterexample, or a criterion written down in advance that was triggered. Whatever reaches neither end is still exploring. Inside that status there are two grades. **Evidence accumulating**, the direction shows, the volume does not. **Nobody knows**, not even a direction yet. Daring to write "nobody knows" is the first thing that separates this map from a headline. Counting the two grades inside still exploring, the figure below draws four boxes. ![The honest map's four evidence statuses, verified / evidence accumulating / nobody knows / falsified. When new evidence arrives, the status has to move](../assets/images/honest-map.svg) **Key evidence.** The hardest few items that hold up the current status. Each must be identifiable. Papers, controlled measurements, and practice that many users can reproduce all count. A controlled measurement puts the AI's result next to a baseline, for example the result of doing it by hand, instead of taking AI's word for how well it did. "Industry consensus" is not accepted. "Everyone says so" is not accepted. **What evidence would change its status.** The most valuable column on the whole map. It forces you to admit, at the moment you write a status down, that the status can die, and to write down in advance how it dies. With this column the map is alive. With it empty, the map decays into one more list of positions. This is the muscle the "how much I believe it" column on the Chapter 4 paper card trained, doing the same work here. **Last review date.** This edition's baseline is 2026-07. Five of the rows were reviewed once more in 2026-08 against public accounts of the automated research systems of the time, and are marked 2026-08 separately. What this column is for waits for section 13.6. For now, one sentence. A judgment without a review date will not notify you when it expires. Assigning a status has two more disciplines, both stolen from the transfer map, and they happen to be what kills most headline claims. Row 2 of the transfer map says, **a dazzling demo does not mean production-ready**. So demo videos, case write-ups, and vendor material only ever qualify to trigger an investigation on this map, never to decide a status. Status recognizes production-grade evidence only. Row 3 of the transfer map says, **self-perception is not trustworthy, calibration comes from measurement**. So "everyone who used it says it's great" does not go into the evidence column. Controlled measurements do. ## 13.3 The map itself, fourteen rows Where do the rows come from? Half from this book itself. Earlier chapters made judgments, and now the bill is due. The other half from the field, the big claims that get reposted the most. If you do not label their status yourself, they move into your judgment with "true by default" as their posture. Read the status column first, then the fifth column. The other columns can be skipped on a first pass. Come back to them when a specific decision needs them. | # | Claim | Current status | Key evidence | What evidence would change its status | Last reviewed | |---|---|---|---|---|---| | 1 | Literature scanning and structured distillation can be handed to AI (assistant level) | Still exploring (evidence accumulating) | Many everyday users can reproduce it, but by the second discipline in this section that is self-perception, not controlled measurement, and it cannot hold up "verified"; the only hard evidence is the operating form itself (the three-layer intake workflow in Chapter 4), which proves the procedure can be executed, not that the output is reliable | A controlled measurement. AI distillation's qualifier loss rate vs a human baseline β†’ meets the bar, upgrade to "verified"; exceeds it, downgrade to full verification | 2026-07 | | 2 | The output of a fully automatic literature review can serve as a map of the field | Still exploring | Usable as a first draft; errors and omissions cluster in the judgment of "which argument matters" (Chapter 4, 4.7) | Its controversy structure stays consistent with domain experts' maps in blind tests β†’ upgrade | 2026-07 | | 3 | With criteria guardrails in place, the execution step can be raised to collaborator level | Still exploring (evidence accumulating) | Verified on the coding side; on the research side it is **a transfer argument, not a direct measurement**, meaning the coding-side conclusion is carried over and reasoned from, not measured on the research side itself. Transfer is this book's strongest reasoning tool, but by the same standard as row 4, an argument does not get into the "verified" column; the execution step sitting closest to cheap ground truth (Chapter 3, 3.4) is mechanism support, not production-grade evidence | Controlled measurement in a research setting, collaborator-level execution with full guardrails vs full human review, error rate and output volume β†’ meets the bar, upgrade to "verified"; or a batch of research accidents of the "all tests green, conclusion wrong" kind β†’ downgrade to "falsified" | 2026-08 | | 4 | A multi-agent team beats a single model (unconditional form) | **Falsified** | Chen et al., NeurIPS 2024, performance is non-monotonic in the number of calls; Smit et al., ICML 2024, under default settings debate does not reliably beat self-consistency; Wang et al., ACL 2024, a strongly prompted single agent nearly catches up | None. The unconditional form is dead; the live argument is in the next row | 2026-07 | | 5 | After cost alignment, teaming still has a net gain on specific task families | Still exploring (evidence accumulating) | The two camps face off in the Chapter 4 controversy map; after alignment the conclusion is highly task-dependent; two systematic evaluations in 2025–26 each carry counterexamples, homogeneous debate usually loses, heterogeneous model pools and multi-hop tasks come out differently, full names and qualifiers in Chapter 4, 4.5 | Someone does the full version, dollar basis + same-budget self-consistency as the third arm + decision-grade comparison across task families. The spine case delivered the first data point, but its same-budget arm was ruled a breach by its own red team (see row 13). By the middle of 2026 companies began publishing automated-research technical reports (such as AlphaLab from the Morgan Stanley team), with no cost-aligned basis in sight, so they count only as signals that trigger investigation | 2026-08 | | 6 | AI can produce publishable research on its own | Still exploring (weak signal) | Four vendor claims, each with qualifiers, workshop and semi-informed review, no independent audit, undisclosed submission, effect size revised down by hand, the case-by-case interrogation at the second stop in 13.4; by the middle of 2026 two more channels were added, public competitions and internal company deployment, both self-reported | Stable acceptance at mainstream venues under informed review + independent replication + someone assigned to answer when it's wrong | 2026-08 | | 7 | AI-generated fabricated citations and junk papers have entered knowledge production in bulk | Verified (as a phenomenon) | Row 6 of the transfer map, the unit price of "looks rigorous" has been pushed down; the Lancet 2026 audit (2.5 million papers, fabricated citations up 12-fold in three years, 1/277), scale evidence in Chapter 2 | Verification infrastructure spreads and the inflow rate drops significantly β†’ relabel "contained" | 2026-07 | | 8 | Personas / synthetic data can replace surveying real people | Still exploring (counter-mechanisms accumulating) | Of the five ways of passing for real, the fluency illusion is a mechanism fact; variance collapse (Bisbee et al. 2024) and stereotype drift (Cheng et al. 2023) have peer-reviewed measurements (the Chapter 11 dissection) | The results of the Chapter 12 three-criteria verification, see row 14 | 2026-07 | | 9 | Users' self-perception of AI speedup can be trusted | Falsified (coding side); still exploring on the research side | METR controlled study (numbers checked, detailed in Chapter 2), predicted beforehand +24% / self-rated afterward +20% / measured βˆ’19%, wide confidence interval (row 3 of the transfer map) | Controlled measurement in a research setting; if self-estimate and measurement agree β†’ rewrite the research side | 2026-07 | | 10 | Model upgrades will solve reliability problems on their own | Still exploring (every historical precedent says no) | Five paradigm shifts without exception, new capability manufactures new error types in bulk (row 7 of the transfer map, verified) | A model generation appears whose error rate, with no external verification process, is no higher than that of teams that have one | 2026-08 | | 11 | What AI eats is the transcription part of the research craft, not the judgment | Still exploring | The CAD precedent (row 11 of the transfer map); research-side evidence unsettled (Chapter 14 expands) | Stable evidence of automated competence on judgment work (picking the topic, setting criteria, setting claim strength) appears β†’ redraw | 2026-08 | | 12 | "Cost-matched comparisons are a gap in the literature" (this book's first-draft judgment in Chapter 4) | **Falsified** (2026-07-18) | Forward-citation check, cost-matched comparisons exist, and they cluster on the skeptics' side (Chapter 4, 4.5) | None. Kept as a tombstone, a specimen of motivated collusion | 2026-07 | | 13 | Spine hypothesis H, cost-matched, the small-model army trails the frontier model by ≀ 2 percentage points on the chosen task families | Mixed, not falsified and not holding across the board (task-dependent) | The code slice cannot detect a direction difference (reading unsettled, about 1/25 the cost per call, list-price basis, meaning billed at third-party inference vendors' posted prices); knowledge QA (only three subjects measured, business, law, psychology) trails by 9–25 percentage points; the math family retired in both directions; "real teaming beats sampling a single small model" got no evidence anywhere in the case. The falsification condition was not triggered. Numbers and how to read them in 13.5 | A retest on a larger task sample with template concentration removed; or rerun the three arms with an army from a larger size tier | 2026-07 | | 14 | Persona surveys can be used to rehearse interviews / questionnaires | **Falsified** (2026-07-25); the downgraded claim "wording tuning only" still exploring | All three criteria failed, six subgroup distribution distances close to or above twice the threshold (lowest subgroup 1.94Γ—), top-choice agreement rate (the share of questions where both sides' most-voted option is the same) 40% against a threshold of 70%, variance collapse on 80% of clean questions (the Chapter 12 verdict block, traceable in the repo) | Change the persona construction method (such as real-person seed calibration) and retest all three criteria β†’ reopen this row. The three limits on this row's falsification must be read with it, see 13.5 | 2026-07 | The short literature citations in the table have their full names, baselines, and qualifiers in the battlefield record of Chapter 4, section 4.5. This chapter does not retell them. It stops at four places only, the four that best show this map's temperament. ## 13.4 Four places worth stopping at **First stop, row 4, daring to mark "falsified."** The unconditional form of "multi-agent beats single model," that is, the line in marketing copy and repost captions that "a team always beats a solo," has been punched through, and not on a single piece of evidence. Chen et al. ("Are More LLM Calls All You Need? Towards the Scaling Properties of Compound AI Systems", NeurIPS 2024, arXiv:2403.02419) measured the performance of Vote / Filter-Vote systems as **non-monotonic** in the number of calls, and the mechanism is that easy and hard queries are mixed within a task. Smit et al. ("Should We Be Going MAD?", ICML 2024, arXiv:2311.17371) found that **under default settings** debate does not reliably beat self-consistency, that is, sampling the same model several times and taking the majority answer. After tuning it can partly pull ahead, but that is already a conditional proposition. Wang et al. ("Rethinking the Bounds of LLM Reasoning: Are Multi-Agent Discussions the Key?", ACL 2024, arXiv:2402.18272) had a strongly prompted single agent nearly catch up with multi-agent discussion. All three papers hit the same word, "unconditional." Note that the enthusiasts' load-bearing papers did not collapse. The gain in Li et al. ("More Agents Is All You Need", TMLR 2024, arXiv:2402.05120) is real, only the gain rises and then falls with task difficulty, and what Llama2-13BΓ—15 caught up with was a **single query** of Llama2-70B. Wang et al. ("Mixture-of-Agents", ICLR 2025, arXiv:2406.04692) took 65.1% on AlpacaEval 2.0 against GPT-4o's 57.5%, with a "length-controlled" qualifier. Add the source paper of debate, Du et al. (ICML 2024, arXiv:2305.14325). Both camps' evidence is real. What died is only the unconditional proposition, and the surviving argument moved into row 5. The conclusion is highly task-dependent, and decision-grade evidence has only begun to accumulate. "Only begun" is to be read literally. The comparison the fourth column of row 5 describes has not been done in full by anyone to date, this book included. **Second stop, row 6, "AI can produce publishable research on its own."** This is the loudest claim of the moment, and its status only rates "still exploring, weak signal." Spread the vendors' claims out, and none of the four survives being repeated in full, and none passes the three questions below. Sakana's AI Scientist-v2 got a paper generated end to end by AI through a workshop at ICLR 2025. The organizers knew and approved in advance, the reviewers were told AI papers were mixed into the submissions but not which ones, and after acceptance Sakana withdrew the paper as promised beforehand, so it never entered the formal publication record. Intology's Zochi claims a main-conference acceptance at ACL 2025 (acceptance rate about 20%), the highest-venue case in this batch of systems, but the rebuttal, the written reply to reviewers' comments, was written by humans, the result is vendor-reported, and there is no independent third-party audit. Autoscience's Carl submitted without disclosing to the organizers, and the paper was withdrawn once it was found out. FutureHouse's Robin made it into Nature, the wet lab work was done by humans, and the 7.5Γ— effect size AI reported was revised down to 1.75Γ— on human reanalysis. Reports of this kind have to pass three questions. What level is the venue? Was the review informed? How far did humans intervene in the process? A harder question comes from the ladder in Chapter 3. The fourth criterion of the autonomous researcher level, who answers when it's wrong, still lands on nobody. One paper being accepted only says it passed one spot check. It does not say the process that produced it deserves the name "autonomous researcher." By the middle of 2026 two more kinds of channel had grown outside this list, worth recording separately. One is public competition. Weco's Aiden ran for three straight weeks in Parameter Golf, hosted by OpenAI, and set more leaderboard records than any single human entrant (March to April 2026, results vendor-reported, the leaderboard publicly checkable). The "venue" here is a leaderboard plus community reuse, not informed peer review. On the three questions it changed exam halls, it did not pass. The other is inside companies, and what to watch on this road is that "who answers when it's wrong" lands on the company's own risk-control regime. The AlphaLab technical report published by the Morgan Stanley team describes a fully automatic quantitative research pipeline whose models must pass internal risk control before going live (details self-reported, report and code public). So "who answers when it's wrong, still nobody" needs to be said more finely. In the academic publishing setting it is still nobody. In the in-company setting an institutional answer is beginning to appear, at the price that the output never enters the public body of knowledge and trust circulates only inside the walls. The fourth column of this row is written plainly. Three things together, then it upgrades. Until they come together, every "milestone" on this row is handled by row 2 of the transfer map. **Third stop, row 1 and row 3, downgraded one notch by this map's own rules.** The first drafts of both rows were marked "verified" and got sent back in a round of reader feedback. The round was run by exactly the kind of AI persona panel the book's subplot had just ruled unable to replace real people. This has to be labeled clearly on the spot, or it becomes a slap in the book's own face. What the personas raised here was **a logical inconsistency that can be verified independently**. Checking it against the written disciplines in section 13.2 settles it true or false on the spot, with no need to trust the judgment of whoever raised it. What Chapter 12 falsified was **distributional fidelity**, whether persona answers can represent the distribution of opinion in a real population. Something that cannot replace real people as respondents can still handle "checking written rules line by line," which is mechanical cross-checking. That is exactly the boundary row 2 and row 3 of this chapter mark out together. The execution step can be let go (row 3), judging which argument matters cannot (row 2). AI can audit your books. It cannot place your bets. The evidence column of row 1 originally read "many everyday users can reproduce it." Section 13.2 states in black and white, "everyone who used it says it's great" does not go into the evidence column, controlled measurements do, and row 9 had just used METR's three-level contrast to falsify "self-perception" wholesale. **Using a type of evidence this map has just shot dead to issue a pass certificate to another row of the same map**, this is the hardest kind of inconsistency to catch yourself. The two rows sit eight rows apart, and each reads fine on its own. Row 3 likewise. Its evidence is a transfer argument carried over from the coding side. Measured by the same ruler as row 4, "three empirical papers against one proposition," reasoning is not measurement, and one transfer argument does not get into the verified column. Both rows were downgraded to "still exploring (evidence accumulating)," and the fourth column was rewritten as a concrete upgrade measurement. It is recorded here because it exposed a reusable self-check move. **Ask of each of your own rows, "how did I treat this column's type of evidence elsewhere on this map?"** The same type of evidence enjoying different treatment in different rows means one of the rows has your preference mixed in. Where the preference lands also follows a pattern. Row 1 and row 3 are both "AI works well in the places I know," which happen to be the two rows I most wanted to be true. The line in section 13.8, "the most dangerous row is the one you want to be true," I had thought was written for the reader. **Fourth stop, row 12, the book's own tombstone.** How the claim "cost-matched comparisons are a gap in the literature" survived a whole round of scanning and then died in a forward-citation check, that is, looking up who later cited a paper, Chapter 4 has the scene and Chapter 11 has the full dissection from the angle of motivated collusion. It is not retold here. A map's credit depends on whether it dares to pin its own corpse to the board. How many rows it got right comes second. ## 13.5 The last two rows, left blank when the plan was written, backfilled with status only Rows 13 and 14 were empty in the first draft of this chapter. The criteria were signed off first, and the filing slot was built to wait for the results. Now the results are back, and the act of filling them in is itself worth a look. The task families, three-arm comparison, cost basis, significance test, and falsification conditions of spine hypothesis H were all signed off in Chapter 6, before the runs started. So when the result landed as the mixed ending "on some task families no gap can be detected," row 13 got one status and a link to the results file, with no reopening of the debate. In Chapter 4 I placed private odds, narrow tasks have a shot, a general tie is doubtful. The second half came true, a general tie really did not happen. The first half can only be called not lost. The sturdiest finding in the whole case is that a single small model cannot be told apart from frontier on the saturated code slice. The reading is uninformative, not a tie. The numbers are 94.7% for the self-consistency arm, 93.7% for the army arm, 96.0% for frontier, at about 1/25 the cost per call (list-price basis). The "army" itself never beat sampling. The knowledge QA cell is marked "read as the joint CI interval of the two preregistered arms." The reading goes like this. Each of the two preregistered arms yields a confidence interval for its gap to frontier, and the lag is taken as the union of the two intervals, the lower bound the lower of the two, the upper bound the higher. Plainly, the two intervals are joined into the widest one. A confidence interval itself is the range the gap most likely falls in. The phrase "list-price basis" has to pin down the extrapolation radius, because it decides whether this result can be used in your decision. The whole case was billed at API list prices, the posted prices third-party inference vendors charge for open-source models. So what it answers is **renting small models vs renting frontier**. "Building your own cluster vs buying the API" is a different question. The cost drivers change to GPU depreciation, utilization, batch throughput, and ops staff, which can differ from list price by multiples, with no guarantee even the direction agrees. The company decision, "self-host or pay for the API," gets no answer from this case, only a method and one data point on a rental basis. Read this limit as main text. It is part of the "not in full" in row 5's "has not been done in full by anyone to date," this book included. Betting in public means one thing only, that "I knew it all along" has nowhere to hide. The subplot likewise. The three criteria for the persona survey were locked in Chapter 12, and all three failed. Row 14 was filled with "falsified" in the preregistered falsification shape, and the downgraded claim was marked "still exploring" by what the experiment actually covered. No new trial was opened, no defense was filed. The "falsified" in row 14 carries three limits, and reading that row means reading them too. One, the preregistered criteria are a single experiment, with thresholds set by this book. Two, the ground truth (WVS-7, the questionnaire of the seventh wave of the World Values Survey) is most likely in the training corpus, and with the real-person arm cut this cannot be ruled out, so every conclusion carries that proviso. Three, 10 of the 15 questions were flagged by the contamination test, and from a single source "memorized the original item" cannot be separated from "sensitive to wording." The directional evidence (Bisbee 2024, Cheng 2023) stands independently. This experiment is the third data point, not the verdict. Both rows can point to specific commits in the repo. What you just saw is what "criteria locked first, results fill in status only" looks like in a book. Most headlines cannot do this. Their criteria get written only after they see the result. ## 13.6 How to re-derive this map yourself This map starts expiring the moment it is printed. So its value is in the method. The five-column structure plus the two ruling disciplines are enough for you to re-derive it any time. A review does not require rereading every paper. **Review = run the fifth column, row by row.** The "what evidence would change its status" column is itself a ready-made search instruction. All you do is go and check whether that kind of evidence has shown up. After checking, only three moves are allowed. Change the status (the evidence arrived), change the evidence (a harder source replaced it), change the date (checked, nothing moved). The third is the easiest to underrate. Even when only the date changes, it must change. A map whose review dates do not move is a dead map, and a dead map is more dangerous than no map, because it still wears the skin of a map. On cadence, a suggestion, with the full version in the appendix. Verified rows, check every six months, or at once on a field-level event. "Evidence accumulating" rows, every three months. "Nobody knows" rows, event-driven, check the moment the kind of evidence the fifth column describes shows its head. Falsified rows are not reviewed. Tombstones do not need watering, but they stay on display. As for how the official version of this map keeps updating, that is part of the book's living book mechanism, delivered in Chapter 16. ## 13.7 Swap in your project Now draw one for your own field. This is the hardest exercise block in the book, and the most valuable, so the starting bar has to be pressed as low as it goes. 1. **Do not start from "what are the big claims in my field."** That question is too big, and you will list a pile of slogans. Dig out your most recent deliverable, a report, a paper, a presentation, a review, any will do, and copy out the claims in it that you **cited or took as true by default**, until you have 10. These are the rows you are already betting real money on. 2. **Fill only two columns per row at first**, the claim and your gut status right now. Force the claim into one sentence that can be judged true or false. Where it will not go, you will discover on the spot that you have been citing a slogan, and that alone is a gain. 3. **Fill in the fourth column**, "what evidence would change its status." Rows where you cannot fill it in, downgrade to "still exploring (nobody knows)" without exception. A claim whose death you cannot write down, you do not actually know what keeps it alive. 4. **Fill in the evidence column**, each item identifiable. If you cannot write anything harder than "everyone says so," go back to step 3 and downgrade. 5. **Mark today's date and set the cadence.** Leave yourself one self-check signal. Ten rows all "verified," and what you listed is a placebo, not a map. Go back to step 1 and add the claim you least dare to touch. The complete fillable version of this chapter's template (with the status rules card and the review cadence table) is in the appendix. **Want an agent to run it with you?** Paste this to your AI assistant or coding agent: ```text Help me draw the Chapter 13 honest map. I will paste you my most recent deliverable. You do one thing only, copy out the claims in it that I cited or took as true by default, until there are 10, and force each into one sentence that can be judged true or false. Ones that will not go, keep as is and label "slogan". Build an empty five-column table from Template A in docs/appendices/ch13-templates.md. I fill in the gut status, I write "what evidence would change its status", and any row I cannot write you downgrade to "still exploring" by the rules card. Every item in the evidence column must be identifiable. If I write "everyone says so", send it back. At the end mark today's date, and set the review cadence by Template C. If all ten rows come out "verified", remind me that is not a map. If any command errors, stop and show me the output. ``` ## 13.8 Sober reminders - **This map will expire, and it will not notify you.** Any of the fourteen rows may flip before the next review. Flipping is written into the definition of this kind of map, and it is no defect. What you should trust is the five-column structure and the ruling disciplines, not the specific cells of this 2026-07 edition. - **The most dangerous row is the one you want to be true.** The tombstone in row 12 is the proof. Motivated collusion works on the person making the map, and on the person writing the book. Set one rule for your own map. For any row where "this row holding is good for me," the required evidence level goes up one notch automatically. - **Do not outsource the map to AI for fully automatic upkeep.** AI can run the fifth column's searches for you, and that is a good use. Let it change the status directly and you will get a map where every row is fluent, complete, and confident, and the coverage illusion (AI never volunteers which block it missed, Chapter 4) replays at map scale. Status changes must pass through a human. That step is the last barrier between this map and the thing it is trying to resist. - **Other people's maps, copy the structure, not the status.** This book's included. A status is a reading of the evidence at one point in time, through one pair of eyes. You can take these fourteen rows as a starting point, but pick at least two or three rows and run the fifth column with your own hands. Only after running it will you believe it, and only once you believe it will you actually use it. - **Still exploring.** The fifth column depends on a human to run it. How often, who runs it, and how to notify the people who already cited a row when it flips, this book gives only the one-hour-a-month recipe in Chapter 16 and has not verified that it holds up in readers' hands. The update mechanism of this map is itself the row it is least sure of. ## 13.9 The unfair advantage you now hold Headlines have been demoted at your desk, from a supplier of conclusions to a review trigger. Two dueling news items come in, and you no longer ask "which one is true." You open the map and ask which row it moves, whether the evidence passes those two disciplines, and whether this row's review should be brought forward. Ten minutes later you close the map, while people without one are still choosing between reposting and panicking. --- # Chapter 14 Β· The Researcher's New Craft !!! info "Chapter companion" πŸ“‹ [Chapter 14 templates](../appendices/ch14-templates.md) Β· πŸ—‚ [Template index](../appendices/template-index.md) > **This chapter's ladder.** This chapter does not occupy one of the seven steps. It turns the whole ladder around and looks at it from the human side. The four levels are not only AI's climbing route, they are also your table of role changes, and section 14.6 gives the four names for what a human is called on each rung. > > The skill on the first line of your resume, AI did it faster than you last week. This chapter wants to convince you of one thing only. What got eaten is only part of your process steps, and you, the person, are still here. Once you separate the person from the process steps, the panic turns into a list you can act on. > > **This chapter delivers.** Three things. A split ledger of which process steps are depreciating and which are appreciating, a portrait of the new craft in five items (including the four-column dispatch brief), and a four-role ladder table seen from the human side. --- ## 14.1 The first line of your resume Six on a Friday afternoon. A new hire three months in sends you an industry review. You had planned to spend part of the weekend giving him pointers. Instead you read the whole thing that night, paragraph by paragraph, and your red pen never came down. Clean structure, claims carrying their qualifiers, key numbers with sources, even a small table of "which two camps are fighting and what the stakes are." The flaws you finally picked out were one citation format and two sentences that read like translation. You ask how long it took. "Two days. Most of it verifying AI's drafts. The writing was fast." Five years ago your first review of that quality took three weeks, and it is what earned you your standing at this company. The first line of your resume still says "fast at mastering unfamiliar fields, produces high-quality reviews and judgments." That line looks a little old tonight. On the subway home you ask yourself the question every craftsperson has asked on some evening these past two years. **Is my value being eaten by AI?** The first thing this chapter does is point out where the question goes wrong. It goes wrong on the unit. It treats "you" and "your process steps" as the same thing. Separate them and you see that what got eaten is the process steps, and only half of those. The other half was not eaten. It is still going up in price. ## 14.2 It is the process steps that depreciate You have already met the accountant in Chapter 2, section 2.1. VisiCalc ate "recompute it," bookkeeping jobs shrank, and accountant jobs grew instead. Now unfold another precedent, the drafter. Before AutoCAD, "design" in industry was really two bundles of process steps. One bundle decided what to draw, structure, function, tradeoffs. The other copied the decisions onto the sheet, line weights, lettering, three-view projection, a hand that must not shake. A drafter's ten years of skill all lived in that second bundle. AutoCAD ate almost the whole of it (row 11 of the transfer map). The result was asymmetric. The drafter's occupation shrank, the designer's did not. The cost of revising a drawing collapsed from a week to an hour, so designers dared to revise ten times. The craft did not disappear. The account was just **split differently**. The mechanical bundle went to the machine, the bundle that decides what to draw went to people, and the first bundle got more valuable precisely because the second got cheap. To run this ledger on yourself, first give two words operational definitions. **Transcription work**, where the mapping from input to output is basically fixed, right and wrong are checkable on the spot, and the method is already written down in ten million precedents. Copying out, formatting, translating, boilerplate reviews, first-draft assembly, boilerplate code all count. **Judgment work**, where "what counts as right" is itself part of the job, with no ready answer to copy. What to ask, what to trust, what counts. Why is the depreciation wholesale? Transcription work has cheap ground truth, its errors are visible on the spot, machines can close the loop and check themselves, supply can be copied without limit, so the unit price collapses toward zero. How much better AI is than you does not even enter this ledger. The appreciation is wholesale too. Once production gets cheap, output volume explodes, and every piece of output needs someone to answer "can it be trusted." Demand for judgment explodes and supply does not. The bottleneck sets the price, and the price rises on the link that was not automated. This script was rehearsed once on search engines, row 8. This time it is research's whole transcription bundle. So that midnight question should be rewritten as **what fraction of my time goes into process steps that are depreciating**. The next section is that table. ## 14.3 The craft migration list The table has one row per process step. The third column is the evidence anchor, the source to check back against. Reading columns one, two, and four is enough. Every row was argued in an earlier chapter, or is already booked on the transfer map. | Process step | Depreciating / appreciating | Evidence anchor | What you should do instead | |---|---|---|---| | Reading every paper end to end and hand-writing the summaries | Depreciating | Chapter 4's paper card; row 8 of the transfer map | Have AI produce structured summaries, spot-check them, read in full only the few papers that bear directly on your question | | Citation formatting, translation, copying out | Depreciating | Chapter 3's defining tasks for tool level | Dispatch all of it, keep only the spot check | | Boilerplate reviews and the "related work" first draft | Depreciating | Chapter 4; row 1 of the transfer map | Have AI draft it, and check every claim heading into a decision against the original before you sign | | Boilerplate code for data pipelines and analysis | Depreciating | Chapter 7; row 5 of the transfer map | Have AI write the pipeline, you write the criteria and the tests, run a pilot first | | First-draft assembly (arranging existing material into a finished draft) | Depreciating | Chapter 9 | Generate it from the claims list, dispatch it | | Running the mechanical checks (citation existence, number tracing) | Depreciating now | Chapter 12's batch dispatch at L1 | Dispatch it to an independent channel, read only the "undecidable" and "falsified" columns | | Picking the question (what is worth asking, what shape to ask it in) | Appreciating | Chapter 5; the no-cheap-ground-truth side of 3.4 | Do it yourself, AI only produces candidates | | Criteria design (locking in "what counts as losing" before the run) | Appreciating | Chapter 6; row 5 of the transfer map | Lock it in and sign it yourself, no changes before the run | | Dispatch (writing the task as a self-contained brief) | Appreciating (a newborn process step) | 14.5 of this chapter; this book's own case | Practice the four-column brief, see 14.5 | | Allocating the verification budget and the final review | Appreciating | Chapter 12; row 1 of the transfer map | Assign the layer yourself, never outsource the final review | | Labeling confidence and reviewing status | Appreciating | Chapter 13; row 3 of the transfer map | Label the status yourself, AI only runs the searches | | Signing your name (answering for the consequences of a conclusion) | Appreciating | The fourth criterion in 3.3 | Sign only conclusions you can answer for | Two notes on reading the table. First, **depreciating does not mean gone**. The steps on the depreciating side still have to be done. They move from you doing them by hand to you dispatching and accepting. They drop from a skill to a cost line, and the responsibility stays in your name. Second, **not one row on the appreciating side is newly invented**. All of them were already part of the research craft, buried under the transcription workload. Once the transcription is pulled out, they become the main job. ## 14.4 Five things, one craft The six rows on the appreciating side collapse into five core skills. The sixth row, signing your name, is not one of the five, and section 14.7 takes it on its own. The four that earlier chapters already established each get one sentence here to pick them up, not a rerun. **Taste in questions.** Chapter 5 taught it as a capability, grinding a blur of curiosity into a falsifiable question. Now it rises into an identity. Once reviews, code, and first drafts are nearly free, almost the only thing separating you from anyone else is what question you asked. AI can generate a hundred candidate questions in one breath. Judging which one deserves three months of your life is you. **Criteria design.** The core move of Chapter 6, signing off "what counts as losing" before the run. In an age of free production, criteria are the only thing standing in front of "the conclusion happens to support the plan that was finished first." Drafting can be outsourced. The right to sign cannot. **Dispatch craft.** The only newborn process step among the five, taught on its own in the next section. **Verification discipline.** Chapter 12 delivered the whole set, three-layer assignment, the independent channel, three-value output. Putting it in the portrait adds one sentence only. It is the one of the five that cuts across all seven steps. **Honest calibration.** Chapter 13 delivered it, three statuses, the conditions that change a status, the review date. The reason it belongs in the portrait, it governs how you price your own output. Price it wrong and the first four, however well done, go bankrupt in someone else's hands. Why exactly these five? Take the three variables from Chapter 3, section 3.4 and run them across, visibility of errors, cost of correction, whether cheap ground truth exists. All five land on the worst side, one by one. Taste in questions has no cheap ground truth on its side, no compiler can rule on whether a question is worth asking. Errors in criteria design are silent. Errors in dispatch show up latest of all, a botched brief is invisible until the work comes back. Verification discipline is a contest over the cost of correction, misallocate the budget and the errors that slip through flow quietly downstream. Honest calibration waits on the slowest feedback of all, mark your confidence too high and you wait for reality to settle the account. Errors silent, correction expensive, no ground truth to lean on, this is exactly the class of work that section 3.4 said needs the knob turned low and a human present. Whatever a machine can close the loop on has collapsed in price. What is left is scarce, and it has nothing to do with being noble. This new-craft list is nobody's design. It is what the distribution line of cheap ground truth cuts out of your work on its own. That line is a strong correlation, not an iron law, and the line itself moves. The withdrawal condition is the same one written in row 11 of the Chapter 13 map, the row saying what gets eaten is transcription and not judgment. The day stable evidence of automated competence on judgment work appears, this list has to be redrawn. One more structural observation. All five are **meta-work**. What they produce is constraints on content, and not one line of the content itself. The question constrains the direction, the criteria constrain winning and losing, the brief holds execution in place, verification guards the door, calibration caps how strongly you may state a claim. One hour upstream saves ten hours downstream. The researcher's new craft pressed into one sentence, **you go from a person who produces content to a person who produces constraints**. Constraints have another side. Earlier chapters all presented them as a brake. Here is the other half, they are also an asset. The market is not short of content. It is short of what makes a pile of automatic output credible, a set of locked-in criteria, a harness that leaves errors nowhere to hide, meaning the scaffolding the experiments run in, an eval environment usable as a training signal, meaning the kind that scores repeatedly and feeds the results back so the model improves. Whoever banks these has the larger radius of letting go. This accumulation cannot be carried off and cannot be copied. It grows inside your understanding of your own problem. The discipline that guards against self-deception and the assets worth money are two sides of one thing. The five also mesh as one set of gears. Taste picks the question, criteria set winning and losing for it, the brief sends the work out, verification decides whether what comes back gets in the door, calibration decides how strongly you may state what got in. Signing your name finally collects the whole chain's responsibility under one person. Break any link and the other four spin free. However hard your criteria are, a botched brief still gets you back a pile of plausible garbage. Chapter 15 will teach you to freeze these gears into habits that do not run on willpower. This chapter only makes you see the shape of the whole set. ## 14.5 Dispatch craft, writing a task into a brief AI can take This craft has already shown up in pieces. Chapter 4's paper card is a miniature brief, and Chapter 12's verification brief goes further, giving only the claim and leaking no expectation. This chapter gathers the pieces into one craft, because every step you climb toward collaborator level hands a bigger block of work to an executor who does not share your brain, and the quality of that handoff sets the rework rate. All the difficulty of dispatch sits on one thing. Whoever takes the job does not have your tacit knowledge. Dispatching to a colleague can be sloppy, a colleague will ask, will fill the gaps, the two of you share the air of one office. Dispatch to AI and nothing outside the brief exists. Dispatch craft is the ability to turn tacit context into explicit text. Its skeleton is four columns. ```text [TASK] What to produce. Verb first, one sentence. Medium, format, length spelled out. [CONTEXT] Every fact and file the other side needs to start, written on the rule that "the recipient knows only what the brief says." Check before dispatch, is the context pack current? [BOUNDARIES] What is not allowed. No inventing facts or citations. Anything uncertain gets marked. Whatever is missing, come back with a list, no filling in from imagination. [ACCEPTANCE CRITERIA] What counts as delivered. Checks a third party can run, not "write it a bit better." ``` This book is itself a first-hand case of "a researcher commanding an army of AI." Chapter drafts were drafted by writing agents and finalized by the editor-in-chief, checks ran through an independent channel, and Chapter 12 already showed you that channel's record of shooting down the book's own claims. Two dispatch lessons here, both of which really happened. The first is a rework specimen. After the Chapter 5 writing agent turned in its draft, three motive paragraphs were torn down and redone. On review the account went to the person dispatching, and the agent had done nothing wrong. The context pack it received was an old version of the case status doc, still holding a claim the verification had already overturned. An agent does not refresh facts on its own. The world in the brief is its entire world. The lesson hardened into a process step, the context pack gets updated before dispatch, not to be skipped even once. The second lesson is positive. Every brief carries a required-reading list and a boundary discipline forbidding invented case facts. At handover you check against the list, instead of judging by feel whether it reads well. The test of a good brief is one sentence. **An executor who has never met you and cannot ask you questions can start work from this brief alone, and knows what counts as delivered.** Fall short of that and the rework goes down as your dispatch incident. A bad-brief and good-brief comparison, plus the full fillable template, are in this chapter's appendix. ## 14.6 The same ladder, seen from the human side When Chapter 3 set the ladder up, the viewpoint sat on AI's side, watching how high it climbs. Now turn the ladder around and look at what **your** role is called on each rung. The four criterion questions (who drafts, who reviews, who decides, who answers when it's wrong) do not change. Your title does. | Level | Your role | What your day looks like | |---|---|---| | Tool | Artisan | The craft is all in your hands, AI is a faster pen | | Assistant | Lead writer and full reviewer | AI drafts, you read every paragraph; your eyes are the line of defense | | Collaborator | Editor-in-chief and process designer | You no longer read sentence by sentence; you write criteria, design process, spot-check, and rule | | Autonomous researcher | Principal (nobody answers for it) | You set the goal and the acceptance criteria; but "who answers when it's wrong" lands on nobody, and the human role at this level has never been filled in | Artisan, lead writer, editor-in-chief, principal, these four names are this book's working terms, and Chapters 15 and 16 refer back to them. Look at the gap from assistant to collaborator. Your work changes from "doing research" to "designing the process that makes research acceptable." Chapter 3 said that here your reviewing switches trades. This chapter's version is that this is the moment the five new-craft skills report for duty. The editor-in-chief is still an artisan, only in a different trade, and the work has not dropped by an ounce. The empty slot beside "principal" in the last row is the fourth column of that table in section 3.3, reproduced as is. Until the answering problem is solved, the top floor of the ladder houses AI, not people, and should not house people. One old rule from section 3.4 carries over unchanged. Roles, like levels, live on "step Γ— task," and they do not follow the person. In a single afternoon you can be the principal on citation formatting, the editor-in-chief on the data pipeline, fall back to lead writer when reading the results, and go back to artisan on "dare we state this conclusion unconditionally," writing every qualifier by hand. Climbing the ladder, seen from the human side, means "editor-in-chief" shows up more and more often in your role column, while "nobody" should never show up at all. ## 14.7 Operational definitions for the three things that "are always human" "Some things will always be human." That line gets said so often it has nearly been emptied out. The only way to give it content is an operational definition. Three of them, each with a test. **One, judgments where someone has to answer when they are wrong.** A judgment belongs to a human if and only if, when it goes wrong, someone has to clean up the wreckage, retract, compensate, correct, apologize. AI bears no consequences. So for any judgment whose consequences cannot be recalled, publishing, architecture selection, advice to a patient or a client, the last pair of eyes has to sit on a person who can bear the consequences. This follows straight from the fourth criterion in section 3.3, with not half a line of sentiment. **Two, the power to define "what counts as winning."** Criteria are research's constitution. Whoever writes the criteria defines what is true inside that small world. AI can draft a criteria proposal, but hand over the signature and you drop from researcher to spectator of the research. The test is simple. Pull up the criteria file for the project on your desk and look at the name written where the final ruling goes. That name is this project's real researcher. **Three, a signature that carries lifetime responsibility for the output.** A signature ties your name to the entire future of the output, the part of that future that gets overturned included, and the credit line is only what that looks like in passing. The spine case's hypothesis H and the subplot's persona verification are already filed in rows 13 and 14 of Chapter 13, one mixed, one falsified. The signature owning up on those filed rows is mine. Not one of the models that ran scans or wrote drafts for me is credited. The signature on the subplot plan was a human's from beginning to end. The test is colder. Before a piece of output goes out, ask one question. Three years from now it gets falsified, who steps up? Anything whose answer holds no person's name does not deserve to go out. The three together are the entire content of the "judgment" in "command plus judgment." Not one of them can be written as a prompt. What a prompt holds is information processing. The substance of these three is a promise. ## 14.8 Swap in your project Make yourself a "my process list," budget thirty minutes. 1. **Write down the 10-15 process steps you actually did last week**, verb first, specific down to "formatted the citations for a report," and no writing "did research" at that grain; 2. **Label each one transcription or judgment**, asking only two things while you sort, does this step have cheap ground truth, and how fast do errors show up (the 3.4 variables, now used as a sorter); 3. Mark the transcription side **depreciating** and the judgment side **appreciating**; anything you cannot call, mark "mixed," then split it in two and relabel; 4. **Estimate against your calendar, not your impression**, what fraction of last week went to the depreciating side? 5. Circle the single most time-consuming step on the depreciating side and **write its first four-column brief this week**, then dispatch it; 6. Circle the one you are weakest at on the appreciating side (check against the five), it is your private focus for the rest of this book. 7. **Beginner's version**. If you have not banked any judgment yet, keep at least one step on the depreciating side that you do not dispatch, literature summaries or citation checking for instance, and do it by hand for a full month before you talk about outsourcing. Judgment is trained inside transcription work. Dispatch all of it and you dispatch the chance to practice along with it. Most people, filling this in for the first time, get their jolt at step 4. People who think they live by judgment find their calendar showing six to eight tenths of their time going to transcription. That is nothing to be ashamed of. The previous generation of designers had calendars full of tracing too. Drawing the list and then not moving is the shameful part. The full fillable version of the process self-check sheet is in this chapter's appendix. **Want an agent to run it with you?** Paste this to your AI assistant or coding agent: ```text Help me build the Chapter 14 "my process list", budget thirty minutes. I will dictate the 10 to 15 process steps I actually did last week, you record them verb first, and send anything at the grain of "did research" back for me to split. Label each one transcription or judgment, asking only two things, is there cheap ground truth, how fast do errors show up. I give the answers, I set the labels. Estimate the time shares against my calendar, you may not estimate from my impression. For the most time-consuming step on the depreciating side, I write a four-column brief this week and dispatch it, and you build the skeleton from the task, context, boundaries and acceptance criteria columns of Template 2 in docs/appendices/ch14-templates.md, contents filled by me. Remind me of the beginner's version, keep at least one step on the depreciating side and do it by hand for a full month. If any command errors, stop and show me the output. ``` ## 14.9 Sober reminders - **The migration hurts.** "It is the process steps that depreciate" holds only if people can move their time to the appreciating side. People who cannot move really exist. The shrinkage of bookkeeping jobs is about four hundred thousand specific people (the number Chapter 2 settled), and the drafter was a whole occupation. For someone who built an identity and an income on a process step that got eaten, panic is an accurate response. And the transition stories you hear are survivors' stories by construction (Chapter 2, failure condition five). This book can give you a list and a direction. It cannot give you a guarantee. - **Do not romanticize "commanding."** The editor-in-chief's days are no easier than the artisan's. Once production is free, your day is acceptance, spot checks and rulings around the clock, and the bottleneck moved from production onto you. Dispatch craft saves transcription time. It saves none of the mental effort. - **The apprenticeship gap, still exploring.** Judgment has always been trained inside transcription work. Taste in the literature comes from having read a hundred papers to shreds by hand, a feel for criteria comes from having been burned by an unfair baseline in person, and once transcription is eaten, where does the next generation practice judgment? Coding already has an account you can check, conclusion first, among the jobs most easily displaced by AI, employment for people just entering is falling while people with more years are rising, and the gap is driven by "hire fewer juniors." The numbers go like this. A measurement based on real ADP payroll data (Brynjolfsson et al., 2025, Stanford Digital Economy Lab) finds that in the occupations most exposed to AI (software development included), employment of workers aged 22-25 fell about 16% in relative terms, while workers over 30 in the same occupations grew 6%-12%. Another study covering tens of millions of resumes (Hosseini Maasoum and Lichtinger, 2025, SSRN working paper) points the same way, firms adopting generative AI shrink junior roles while senior roles barely move. The door to practice is narrowing. This book is confident about how people who already have judgment migrate. On how judgment grows from zero, the honest answer is still exploring. **If you do not have judgment yet, do not start subtracting from this table.** The paragraph above describes the apprenticeship gap as a field-level problem, while what you hold is a personal decision, so here is a default plan you can test, to keep "still exploring" from turning into a shield. **The depreciation table is a ledger for people who already paid tuition. A beginner cannot treat it as a permit.** Down to the actions. Those "hundred papers" in your own field, until you have read them by hand, AI may do exactly two things at the literature step, help you find which one to read, and quiz you after you have read it ("I understand it as X, how far does the original support that"). It may not write the summary for you. The summary is the very step where taste gets trained. Outsource it and the practice is gone, and what is left is collecting, which happens to feel a lot like learning. In the same way, at the criteria step, until you have been burned by an unfair baseline yourself, plans AI drafts get walked item by item against the Chapter 6 checklist, with no adopting a draft whole. The evidence grade of this plan has to be stated plainly. It is **derived from transfer history, a suggestion, not a prescription that has been measured** (the same class of reasoning labeled in row 3 of the Chapter 13 map). I give it a falsifiable shape. The day a controlled measurement appears showing "people who started out using AI throughout have judgment indistinguishable from the traditional path three years later," this plan is void and I will withdraw it. Until then, slower is better. How the five skills get frozen into habits that do not run on willpower, and how they grow onto a team, is Chapter 15's job. This chapter is responsible only for making you see the ledger. ## 14.10 The unfair advantage you now hold The next time that question finds you late at night, "is my value being eaten by AI," you do not have to choose between panic and self-comfort. Spread out your own process list and point to which half is depreciating, which half is appreciating, and which way your time has already started moving. Most people who ask that question have not yet separated "me" from "my process steps." --- # Chapter 15 Β· Make It a Habit and a Capability !!! info "Chapter companion" πŸ“‹ [Chapter 15 templates](../appendices/ch15-templates.md) Β· πŸ—‚ [Template index](../appendices/template-index.md) > **This chapter's ladder.** No ladder line this time. This chapter is not one of the seven steps. The ladder measures how high AI climbs inside each process step. This chapter is about the level you have already climbed to, and what keeps you there three months from now. The answer is up front, and it is not willpower. > > Everyone who has tried a new workflow has felt the high. Most are back to their old ways in three weeks. This chapter assembles the loose parts of the previous fourteen chapters into a machine that does not run on enthusiasm, installed first in you alone, then spread to your team. > > **This chapter delivers.** The three process questions card, the 30-day adoption plan, and the team metrics starter sheet, whose three metrics have to be used in pairs. --- ## 15.1 The Monday of the third week The weekend you finished Chapter 12, you set a one-page verification budget sheet for your own project. Week one went well. The research summary for your manager passed L1, you caught one citation with paraphrase drift, and it felt good. Week two was all right, except the spot check dropped from five citations to three, that week had two deadlines. On the Monday of the third week there is an unscheduled meeting at four in the afternoon, and the AI-drafted competitor analysis has to go out before five. You read it through, thought it read well, and hit send. Not one citation was spot-checked. You can recite the process. You will not forget it. It lives in your memory and your good intentions, and memory and good intentions lose to a calendar. You have probably seen the team version of the same story, and it costs more. The company bought an enterprise AI tool, ran two training sessions, and the demo drew real applause. Three months later the admin console shows half the seats never logged in again. The people who do log in use it for exactly what they used a search engine for. The tool made it into the budget. The workflow did not move a millimeter. Both endings are the same death. **A new method wins on results and loses for having no institutional slot.** The previous fourteen chapters handed you loose parts. Loose parts do not turn into a machine on their own, and this chapter does only the assembly. Chapter 14 settled the account at the skill layer. This chapter handles the institutional layer, how those crafts keep running on your worst days. ## 15.2 The smallest unit of a habit Steal the answer from coding first, and you are the example. You did not remember to run tests by willpower. The tests sit in CI and fire on every commit. You did not do code review out of conscientiousness either. Review is a gate in front of the merge button, and what fails it does not merge. Behind this sits an observation that keeps being confirmed. **Every check that matters ends up built as "you cannot get past without it," and nobody counts on "remembering to do it" any more.** The verification steps of Chapter 12 stay alive in mature teams for the same reason. The step grew onto the path. The earlier precedent is in the cockpit. What follows is a historical aside, and readers who want only the conclusion can jump to the sentence beginning "So the smallest unit of a habit." On October 30, 1935, the Boeing Model 299, the prototype of what became the B-17, crashed on a test flight at Wright Field. Nobody had released the gust lock on the elevator before takeoff. Of the five people on board two died, including Major Hill, who ran the test flight department. The crew was the best available. The accident happened because the aircraft's complexity had for the first time outrun anyone's working memory, and the pilot's checklist was born from it. The AI workflow is structurally the same. Many process steps, many pits, and you filled in that "attribute Γ— step" table of Chapter 11 yourself. Carry it in your head and you are certain to skip a step on the Monday of the third week. So the smallest unit of a habit is two parts, **a trigger plus a checklist**, and resolve is not one of them. The trigger has to be an objective event. "Every time I open a new conversation for this process step" counts. "I will be more rigorous" does not, that is a wish. The checklist has to be short, three questions at most, and section 15.9 settles the account on long checklists. Duhigg, in The Power of Habit (2012), splits a habit into a three-part loop of cue, routine, and reward, his popularization of the basal ganglia research at MIT. This chapter borrows only the first part. The place to operate on a habit is the cue, which event you weld the behavior onto, and how the behavior itself is worded turns out to be the easy part. ## 15.3 The three process questions, folding the first fourteen chapters into three The content of the checklist was written across the first fourteen chapters, scattered through the unfair advantage at the end of each one. Folded up, it is three questions. Ask them once at the door of every process step. Opening a new conversation, dispatching a task, taking delivery of an output, all count as the door. **Question one, criteria.** Where does this output go, what counts as passing, and what counts as losing? "Where it goes" decides the verification level, which is the first question when Chapter 12 assigns the layer. "What counts as losing" has to have an answer before you start, the immunity bought by the sign-off of Chapter 6. Start work unable to answer this question and the output will "just happen" to support the conclusion you wanted first. **Question two, delegation.** What level of the ladder do I put AI on for this step, who drafts, who reviews, who answers when it's wrong? The four behavioral criteria of Chapter 3 turn here from a framework into a gate. This question forces you to renegotiate the division of labor at every process step. The inertia of the previous step does not carry over, and Chapter 3 said the level lives on "step Γ— task." **Question three, verification.** Which class is it most likely to break in, and which layer does it pass before it leaves? For the first half, open the failure-mode census sheet of Chapter 11. For the second half, use L0 / L1 / L2 from Chapter 12, spot check, full verification, adversarial recompute, and each layer up costs more. The value of this question is its timing. It gets asked at the door in, and asking it at the door out is too late. Once you know which layer the output will have to pass, your dispatch brief and the way you file things both change. The three questions are the index to the whole book. Whichever one jams, go back to that chapter. Knowledge lives in the book, habits live on triggers. The full fillable version of the three process questions card is in this chapter's templates. Install it in your conversation template and in the first column of your dispatch brief. A checklist on the wall is decoration. Only a checklist on the path is a gate. ## 15.4 Thirty days, four weeks Installing a whole machine at once is not realistic, and that is exactly the kind of plan that died in three weeks in section 15.1. The gradual version moves week by week, adding one part a week. **Week one, install one process step.** Pick the one where AI is deepest in and the cost of error is moderate. Most readers will pick literature scanning or first-draft assembly. Fit it with a trigger and the three questions card, and leave every other step alone. The success criterion this week is coverage only. Every time you enter this step, the three questions get answered, however sloppily. In the first week willpower still has to front the money (you pay out of pocket first), the credit line is small, it covers one step at a time, and spreading it thin bankrupts you. **Week two, chain two.** Weld this process step and its downstream verification into a pair. After every generation, run one L0 spot check from Chapter 12. What this installs is the most important structural part in the book, the separation of generation from verification. The orthogonality Chapter 2 stole from The Pragmatic Programmer becomes, here for the first time, the default structure of your own process. **Week three, run the whole workflow once.** Pick a problem that is small and real, not worth a month but worth two or three days, and take one full turn around the small loop of the seven steps. Rough is allowed, and the output does not count this week. The point is to get the three questions to appear at the door of every process step at least once, and then to let you feel by hand which door makes them most awkward. Where it is awkward is where the checklist needs rewording. **Week four, retrospective and retirement.** Count a few numbers. How many times the three questions got answered. What they caught, one fabricated citation, one criterion added after the fact, one output that should have passed L1 and nearly went out bare. Then do the most counterintuitive move in the whole plan. **Delete or rewrite every line that has never caught anything.** This is the first turn of the retirement cadence for checklists, and the full reasoning is in section 15.9. Thirty days is only the length of four retrospective cycles, and behavioral science has no such magic number. Nor does the acceptance test at the end have to measure self-discipline, since one behavioral signal is enough. **When you skip the three questions, it feels awkward.** That awkwardness means the default has been swapped. From that day the cost of maintaining the workflow is paid by the institution, and your willpower goes back to the thing it is for, judgment. ## 15.5 From one person to a team You probably have no authority to set rules for your team. The good news is that team adoption starts from a working example anyway, and rules come later. The example worth copying happens to be the team version of the trust mechanism, three pieces in all. **The shared criteria library.** The criteria you locked in for your project in Chapter 6 and the verification standards you set in Chapter 12, "what a passing citation check looks like," "how far number tracing goes before it stops," go into the team wiki. The next person handed a task of the same kind copies them straight and changes a few parameters. Criteria go from personal discipline to public asset. A team's radius of letting go is set by how hard the public criteria are, which is the team version of row 5 of the transfer map. **The verification budget sheet.** In Chapter 12 you already set a one-page version for your own project. The team version adds one thing only, making "has this thing passed the layer it should have passed" a question anyone can ask out loud. It asks about the process, not about character. **The case status doc.** A team-level registry of "facts that have happened," recording who ran what, how it came out, which conclusion has been overturned, which number has gone stale. This book runs on exactly such a document. Case facts in every chapter can come only from it, missing facts get listed as pending, and inventing them is not allowed. It has also stopped a real accident. One chapter of this book was drafted by a writing agent, the context dispatched to it still carried a claim that verification had already overturned, and three paragraphs went back for rework after delivery. The lesson is on record. **Update the source of facts before you dispatch, because the executing side will not refresh the facts on its own.** The single source of truth, plus "whoever updates signs," is what treats this. The contagion path opens from there, and only one order works. First you get work done with the example. Your output comes with a "claim β†’ source" table attached, while other people's output sends them digging through chat logs the moment it is questioned. Then someone copies it, and copying saves them time. A third person asks whether there is a template. At that point, raising "should we make this a rule" only puts a name on something that has already happened. Reverse the order and you get the enterprise tool from section 15.1 that nobody logs into. ## 15.6 A good-looking demo is not productivity After a team adopts the AI workflow, the most dangerous stretch is the feel-good period. Someone demos at the weekly meeting, "I did a day's work in ten minutes with AI." Row 2 of the transfer map already ruled on signals like this. Evaluation runs on production-grade numbers only. Row 3 supplies the other half, self-perception is not trustworthy. In the controlled trial from Chapter 2, developers rated themselves about twenty percent faster and measured nearly twenty percent slower. The team version of that is "everyone feels sped up," a sentence carrying exactly as much evidence as that self-estimate did. Zero. You have to build the production-grade numbers yourself. Start with the metrics below, and the full version of the basis and the collection method is in this chapter's templates. - **Rework rate**. The share of output with AI deeply involved that gets sent back for redoing after delivery downstream. Cheapest to collect. Add a reason label to "sent back" on the task board, and count once at month end. - **Verification pass rate**. The share that passes on the first try in spot checks and in full verification, recorded separately for citation existence, paraphrase fidelity, and number traceability. If the team set up a verification ledger the way Chapter 12 says, no separate collection is needed, the ledger is the data source. This book's own ledger is a ready-made demonstration. Its overturn rate and its correction rate are this book's verification pass rate, and the account runs as follows. About forty-eight facts pending verification are on the register, thirty-four have been verified, one core claim was overturned outright, and about ten statements were corrected or narrowed. Item-level provenance is in the experiment ledger index. - **Claim survival rate**. The share of conclusions that entered the decision chain and still stand at a scheduled review point (three months, say). It is the slowest of the three and the closest to what "productivity" actually means. What this workflow really produces is conclusions that stand, not documents. The starter sheet looks like this, one metric per row. The "paired metric" column is what the second line of defense below asks for. Leave it blank for now, and fill it in once you have read the defenses. | Metric | Data source | Collection cadence | Paired metric | This quarter | |---|---|---|---|---| | Rework rate | The "sent back" label on the task board | Count once at month end | Throughput | ____ | | Verification pass rate | The verification ledger of Chapter 12 | With the ledger, summarized quarterly | Verification coverage | ____ | | Claim survival rate | Register of conclusions entering the decision chain, plus a review point | Every quarter | The risk level of the conclusion | ____ | One warning ships with the metrics. **Measures get gamed.** "When a measure becomes a target, it ceases to be a good measure." !!! note "Attribution of Goodhart's law (skip on first read)" This popular phrasing comes from the anthropologist Strathern (1997), her restatement of Goodhart's law, and it gets misremembered as Goodhart's own words. Goodhart's own 1975 wording is much more of a mouthful. "Any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes." The two versions say the same thing. Every metric has its own way of dying. Put the rework rate into performance reviews and sending things back moves into private messages, leaving the board at peace forever. Put the verification pass rate into a ranking and people submit only their safest output, routing hard problems around the register. Assess people on claim survival rate and the conclusions get more and more timid, a uniform crop of "more research is needed" deathless filler. There are three lines of defense. First, the numbers serve this group's learning and do not enter individual performance reviews. Second, metrics get used in pairs, rework rate with throughput, survival rate with the risk level of the conclusion. A single metric always has a painless cheat posture. Paired metrics bite down on each other. Third, audit the metrics themselves once a quarter, and for the one whose number improved, first ask "did things get better, or did the reporting change." ## 15.7 The book comes full circle, the last meter of the verification net Three judgments made earlier get caught here. Chapter 2 said production cost collapsed and review cost did not, and that was a historical observation. Chapter 12 broke review down into process steps, and that was a method. **A process step written in a book produces no review bandwidth. It produces review bandwidth once it becomes the default action.** Review bandwidth equals the number of verification steps times the probability that they get executed, and what an institution governs is that probability. Chapter 3 said how high you can safely climb depends on how hard your criteria are written. The hardness of a criterion has to be measured by its probability of execution. A criterion executed only on the days your energy is good is worth half its nominal hardness. Only a criterion welded onto a trigger earns the right to use its nominal value when you talk about the radius of letting go. The verification channel in Chapter 12 that overturned my "literature gap" claim had, looking back, not one link that depended on alertness or luck. A different model, a session with no history, a brief that leaks no expected answer, separate filing, every piece of it was an institutional part set up in advance. Every argument in this chapter compresses into one sentence. Give judgment to people and execution to institutions, and do not get it backwards. Institutions expire too. Checklists, the criteria library, the map you built in Chapter 13, how all of it stays fresh in a field that changes every month, the update cadence, is Chapter 16's job. ## 15.8 Swap in your project Pick one process step, install the first habit unit, and commit to two weeks. 1. **Pick the step**. Open your failure-mode census sheet from Chapter 11 and pick the step where AI is deepest in, already marked on the sheet. If you do not have the sheet at hand, spend five minutes on a minimal version first, listing your process steps and the error class each one is most prone to. 2. **Write the trigger**. It has to be an objective event, precise down to the action. "Every time I open a new conversation for this process step." "Every time before I paste output into a document that leaves my desk." Take an intention as your trigger and you are back to your old ways in two weeks. 3. **Copy the three questions and fill them in as yours**. Write your project's specific answer after each one. Passing = what (a number, a criterion), delegation = which level of the ladder, most likely error = which cell of the census sheet. Adjectives do not count. Write nouns and numbers. 4. **Install the card on the path**. The first line of the conversation template, the first column of the brief template, and the head of the document template all count as the path. A sticky note beside your monitor does not. 5. **Write day 14 on your calendar**. On the retrospective day count two numbers, how many times the three questions were answered, and what got caught. The catch record decides the next move, scaling up to the full 30-day plan, or rewording the checklist first. Two weeks and nothing caught at all leaves two possibilities. The checklist is written too loosely, or this process step was low risk to begin with. For the first, reword it. For the second, pick a more painful step and start over. The three templates of this chapter, the 30-day adoption plan, the three process questions card, and the team metrics starter sheet, are in this chapter's templates in full fillable form. **Want an agent to run it with you?** Paste this to your AI assistant or coding agent: ```text Help me install the first habit unit of Chapter 15. I pick the process step from my Chapter 11 census sheet, I write the trigger, you only rule on whether it is an objective event, and you send back intention triggers like "plan to" or "try to". Build the three process questions card from Template 2 of docs/appendices/ch15-templates.md. I fill in the answers under the three questions, adjectives do not count, you take only nouns and numbers. Then install the card on the path, the first line of the conversation template, the first column of the brief template, the head of the document template, you edit the files for me, a sticky note does not count. Last, write day 14 on my calendar. On the retrospective day count two numbers, how many times the three questions were answered and what got caught. If two weeks caught nothing at all, lay out both possibilities and I decide which. If any command errors, stop and show me the output. ``` ## 15.9 Sober reminders - **A dead checklist is more dangerous than no checklist.** Someone with no checklist at least knows they are running bare. Once ticking the box becomes the goal, the finger moves and the eye does not look, and the checklist starts manufacturing the illusion of safety in bulk. This is the shared late-stage disease of every inspection regime. The signal is concrete. One checklist line has a hundred percent pass rate for three months running and has never caught anything. Either it is internalized to the point of not needing to be written, so delete it, or it was never really executed, so reword it or change the trigger. The fix is to give the checklist a **retirement cadence**. Hold the retrospective monthly. Every line either produces a recent catch record or gives an explicit reason to stay, and a line with neither gets deleted. The health metric of a checklist is its catch record, and its length does not count. A test suite works the same way. A test that never fails usually means it is testing nothing, and rarely means the code is that good. - **An institution can freeze discipline in place, and it can freeze an error in place too.** One badly written criterion in the criteria library gets executed efficiently by the whole team, and a shared error travels much faster than a private one. So the criteria library itself needs an evidence status and a review record, and the arrangement of Chapter 13 applies to it unchanged. - **The part still exploring is the team section, the practices and the measurement basis alike.** In one sentence, the team piece has no standard answer yet, this chapter gives starting advice and not a settled conclusion, and below is the grade breakdown plus two organization-level measurements. CI and review culture on the coding side took more than a decade to accumulate. "How a research team shares an AI workflow" is, in 2026, still at the stage where every shop builds its own wheel. The grade of this chapter's team section gets reported honestly. Mechanisms verified on the coding side, plus a mechanism inference on the research side, plus first-hand practice from this book's own single team, do not add up to a verified conclusion on the research side. The three metrics are likewise only starting advice, and a standardized production-grade basis for research teams does not currently exist. Organization-level measurements do exist, and their conclusions are sobering. The DORA 2025 report found individual output up sharply while organizational delivery metrics stayed flat. A preregistered field experiment with 776 participants (NBER w33641) found that "an individual plus AI" is roughly equal to "a team without AI." Gains at the individual level have not automatically carried through to team output, which is exactly why this chapter exists. Use the three metrics to draw your own team's trend line. Do not compare absolute values against another team, the basis differs and the comparison is meaningless. Chapter 2, failure condition two, said the feedback loop of research is slow and the convergence of best practice will be slow with it, so do not expect a textbook next year. ## 15.10 The unfair advantage you now hold On that Monday thirty days later there is still an unscheduled meeting at four in the afternoon, and the calendar still beats memory. It does not beat your templates. The three questions are printed on the first line of the conversation, and the checklist grows in front of the send button. Other people's AI workflow runs on enthusiasm, and enthusiasm has a shelf life of three weeks. Yours runs on institutions, and institutions do not care about your mood. --- # Chapter 16 Β· Coda Β· How This Book Stays Current !!! info "Chapter companion" πŸ“‹ [Chapter 16 templates](../appendices/ch16-templates.md) Β· πŸ—‚ [Template index](../appendices/template-index.md) > **This chapter's ladder.** Not marked. The coda sits outside the seven steps. > > This book opened with a "don't trust it yet" list. Now it is time to settle the account. And, while we are here, to say why a book written in a field that changes every month does not expire on the day it is published. > > **This chapter delivers.** A line-by-line settlement of the seven-item list, the split of labor between the paper book and the case library with the cadence that updates it, and a one-hour-a-month tracking recipe. --- ## 16.1 Back to that list That afternoon in Start Here, you spent two hours running the most naive version of the experiment, no team formed at all, one open-source small model going alone against frontier, and saw 19/20 against 20/20. It looked like a tie. Then you did the most important thing of that day. You skipped the celebration and wrote out a seven-item list of "where I don't trust it yet." That list is the real first page of this book. At the time it looked like a few lines of self-doubt. In fact it had already demonstrated everything the book sets out to teach. Now settle it line by line. Seven debts, booked to three accounts. **Statistical debt, two items.** Too few problems. With 20 problems one problem is 5 percentage points, and the eye cannot tell 19/20 from one swing of luck. Each problem answered only once. Rerun at another time and 18 or 20 would be no surprise. Chapter 6 locked in the repeat count and the seed before the run, Chapter 8 turned "does the gap count" into arithmetic with intervals and tests, and Chapter 10 rechecked even the independence assumption behind the intervals themselves. The endgame is that the gap interval on code crosses zero, which reads as uninformative and not as a tie, and that on math the reversal rested on problems sharing templates, with an effective sample size far smaller than the number of rows, so the whole family retired. **Basis debt, two items.** Cost not accounted for, and a "tie" with no price tag means nothing. Chapter 5 wrote the cost basis into the hypothesis, Chapter 6 turned it into a locked-in accounting rule, and Chapter 7 handed over a ledger that adds up line by line. The Chapter 10 red team then found a breach in that ledger. The control arm's budget was never aligned to the plan, and the phrase "same-budget control" is withdrawn across the book. Only one subject tested, and possibly a leaked one. Chapter 5 replaced the single demo task with task families spanning types and excluded contaminated tasks in writing, and Chapter 6 added the contamination check step. The price showed up in Chapter 8. The contamination-resistant variant set held off memorized problems, and carried in correlation within the same template instead. **Process debt, three items.** No control arm. Chapter 6's three-arm design was signed off, and from then on "does the credit go to teaming or to more sampling" could be told apart. What telling them apart produced is that "real teaming beats sampling a single small model" never got measured once in the whole case. Nobody reviewed the scorer. The Chapter 7 pilot stopped five bugs, then Chapter 8 and Chapter 10 each ran a round on it, because the audit itself has to be audited too. Each round caught one bug the pilot had missed, one handing points into math, one wronging the small models in the audit's own rescoring. The first 20 problems were picked out of laziness and do not count as sampling, so Chapter 6 locked the sampling rule in before the run. The rule was set and the fetching script still did not shuffle. Chapter 8 pointed out runs of same-template variants inside math, and Chapter 10 seized its twin in knowledge QA, data drawn in blocks by subject, with not one STEM problem in it. Not one of the answers these seven items received says "relax." Every answer is a process step. The whole distance between plausible and reliable sits right here. The you of that two-hour afternoon could only say "it looks like a tie, but I don't trust it." The you of now can say exactly how each item of that distrust turned into evidence. ## 16.2 One arc, four stops The spine of the whole book runs one arc, doubt, verify, contribute, deliver and take the hits. One sentence per stop, nothing retold. **Doubt.** I expected a consensus and found a brawl, load-bearing papers on both sides, each side arguing well. The correct shape of doubt is to draw the brawl as a controversy map and find which disagreement your own question lands in. Shouting "trust nobody" does not count. **Verify.** From "can small models team up" to a hypothesis H that can die, then criteria, control arms, and falsification conditions all locked in first. Verification is a contract signed before the run. You cannot fill it in with attitude after the experiment is done. **Contribute.** I got excited thinking I had found a gap in the literature, and a check knocked half of it down. The question narrowed from "fill the gap" to "draw the boundary clearly and supply the kind of comparison nobody had done in full." A question narrowed by a check is worth more than a question that excites you. **Deliver and take the hits.** One body of evidence was written into two deliverables, and then a court really opened on four attack surfaces. All ten charges held. What had to be overturned was overturned, what had to be narrowed was narrowed, the rest got caveats attached, and the case file went into the repo where anyone can rerun it. The submission never happened, and the book says so plainly instead of acting otherwise. The substitute is spreading the case file out in front of everyone. The one-page memo after the beating sits at the end of Chapter 10, section 10.7, as the final version, shorter than the original, recommending one thing less, the army. The point of this stop is order. A beating scheduled before publication is scheduled by you. Scheduled after publication, it is scheduled by someone else. H's ending is filed in row 13 of the Chapter 13 map, the row for spine hypothesis H, a task-dependent mixed ending, not falsified and not holding across the board. The criterion is a gap to frontier of ≀ 2 percentage points on a cost-matched basis, judged family by family. On code the interval crosses zero and reads as uninformative, at about 1/25 the price of frontier, list-price basis. Knowledge QA trails by 9 to 25 percentage points, a clean loss. The math family retired whole, and its problem set has to be reissued and tested again. No family caught up by spending more, the falsification condition never triggered, and no family reached the word "tie" either. The detailed accounts live in the online case library, the online repository whose address and contents come later in this chapter. The subplot reached its stop too, and readers who want only the conclusion can read the first sentence and the last one of this paragraph. The coarse question "can AI personas replace interviews with real people" was ground into a question about distributions, dissected in Chapter 11, locked into a plan in Chapter 12, with the three criteria, the falsification shape, and the downgraded claim all written in advance. The reveal has already happened. All three criteria failed. The distribution distances of six subgroups, 0.19 to 0.25, all exceed the threshold of 0.10, the top-choice agreement rate is 40% against a threshold of 70%, and eighty percent of the clean questions tripped the variance-collapse red line. The preregistered falsification shape was met exactly as written, and "persona can replace real interviews" is filed as falsified. The plan promised that the reveal would only fill in a status, and that promise was kept. However ugly the result, it only filled a status into row 14, the row for persona survey rehearsal. A line that puts synthetic data in for real data ran the same discipline as the spine from start to finish. That is the whole point of its existence, proof that this workflow holds when the stand-in changes. ## 16.3 What expires, and what does not Now to be clear about what this book actually delivers. Conclusions do not make the delivery list. They expire too fast. Any of the fourteen rows in Chapter 13 could flip before the next review, and every model and every benchmark named in the book has a shelf life measured in months. A book that bets on "these claims are true in 2026" starts rotting on the printing press. This book bets on another set of things, and what they share is that they do not expire with a model version. - **The structure of the seven-step workflow.** Questions, hypotheses, criteria, execution, interpretation, delivery, critique. How high AI can stand on each step will change. The relations among the seven steps will not. - **Verification discipline.** Generation separated from verification, criteria locked in first, checks run through an independent channel, production-grade over demo-grade. These held in the spreadsheet era and they hold on the next generation of models. - **The method for re-deriving the honest map.** A five-column structure plus two disciplines for assigning status. The statuses on the map are readings and they change. The way of measuring does not. - **How to use the transfer map.** When a new headline arrives, first ask which row's pattern it is and which rerun of it. That move works just as well on a headline from 2029. What gives these the standing to claim they do not expire? Not one of them was built for the models of 2026. They are mechanisms distilled from five paradigm shifts. Bet on mechanisms, not on outcomes, is the rule Chapter 2 laid down. A book judging its own shelf life should use that same rule. Conclusions are the ink on the map. Method is the hand that draws it. This book teaches the hand. ## 16.4 How this book stays alive The ink cannot be left to rot on paper either. At the foot of the table in Chapter 2, section 2.5, the second of the three notes on use wrote a check. The status column will expire, and the online edition keeps updating. Several more such checks stand in the book. Chapter 4 said the case library tracks automated review tools. Chapter 13 said the update mechanism for the official version of the map would be delivered in Chapter 16. Now they are honored. The online case library is at https://github.com/hallieren/research-rewritten. It holds everything that expires. - The code, data, and result files of the spine and the subplot. Every number and every pit in the book can point to a commit and a result file in the case library. That one is a principle, and writing it down means honoring it. - The live version of the honest map. The fourteen rows in the paper book are a 2026-07 snapshot, and the snapshot date is printed in the map's header. The live version is re-derived row by row on the Chapter 13 review cadence, with a change log entry attached to every status change. The statuses of rows 13 and 14 keep being re-derived here on that cadence. - The live version of the transfer map's status column, and the latest version of each chapter's templates. The paper book carries only the half that does not expire, the frameworks, the process steps, the disciplines for assigning status, the ladder, the five-column structure. The split is simple. What moves lives online, what does not gets printed. On a second reading, whenever you hit a concrete status, number, or result, check the date in the case library first, then decide whether to trust the paper version or the online one. Use the paper book as an entrance only. The update cadence follows the Chapter 13 map review cadence rather than setting up a second one. The case library updates whenever a review produces a status change, and quiet months are not padded with content. One honesty clause remains, and this promise is itself falsifiable. The case library's front page carries a last-updated date. The day you open it and find that date frozen for a long while, treat this book the way you treat any expired map. Keep using the method, re-derive the statuses yourself. A book that teaches you not to trust empty promises should not ask you to trust its own promise unconditionally. ## 16.5 One hour a month The case library is my update cadence. You need one of your own. Draw one boundary first, so this does not fight the steps that came before. The Chapter 4 frontier layer (the layer that turns tracking into a subscription) covers your research field. The Chapter 13 review cadence covers how your map gets re-derived row by row. This section covers the furthest upstream, information about the craft of "doing research with AI" itself, one hour a month, where it comes in, what gets through, what gets stopped. The recipe is three pieces, and the full fillable version is in this chapter's appendix. **The source diet list.** Five at most, and adding one means dropping one first. Keep one or two in each of three classes. First-hand sources, the original papers and system technical reports, retellings do not count. Production-retrospective sources, the accounts written by people who really used it for three months, day-one reviews do not count. Controlled-measurement sources, comparisons with a baseline and a stated basis. There is no "influencer newsflash" class on the list, because that kind of information reaches you anyway and needs no quota. **Signals that trigger a re-derivation.** A piece of news is worth an hour of yours if and only if it could change the status of some row on some map of yours. In practice you hold it up against the fifth column and ask, is this the evidence that column describes. Production-grade measurements, independent reproduction or refutation, a criterion locked in earlier getting tripped, repeated rework inside your own workflow. Those four classes get through. **The ignore list.** Demo videos and launch events, handled by row 2 of the transfer map, qualified only to trigger an investigation, never to trigger anxiety. Self-reported speedups with no control. Model version-number news. The Nth round of the "replacing scientists" debate. Secondhand retellings with no link to the original. These flow straight past. Here is how the hour goes. Twenty minutes running through the sources, thirty minutes running the fifth column on one or two rows for the signals that hit, ten minutes changing a status or a date and writing one line of change log. Checked and nothing moved, the date still changes, a discipline inherited unchanged from Chapter 13. In most months your output will be a few new dates and nothing else. That is the recipe's normal output, not waste. Tracking asks for one thing only. In the month a conclusion flips, you are not absent. **Want an agent to run it with you?** Paste this to your AI assistant or coding agent: ```text Help me set up Chapter 16's one hour a month. Build three sheets from docs/appendices/ch16-templates.md. The source diet list caps at five, one or two in each of three classes. Every source on the list comes from me, you only check the class and the count, and if it goes over the cap, one gets dropped first. Build the trigger-signal list and the ignore list as empty sheets from the appendix, and I fill them. Then copy the hour's time budget into my calendar reminder, twenty minutes on the sources, thirty minutes running the fifth column on the signals that hit, ten minutes changing a status or a date and writing one line of change log. Checked and nothing moved, the date still changes, and remind me of that one every month. If any command errors, stop and show me the output. ``` ## 16.6 The second spine One last thing, about you. If you followed the exercise block at the end of every chapter to get here, you now hold these things, the first four thought through, the last four actually run. - A real problem you picked yourself - A controversy map - A hypothesis that can die - A set of criteria locked in first - A one-page verification budget sheet - Your own ten-row honest map, the one the exercise blocks had you draw, a different map from the fourteen-row one in this book - A process-step list marked depreciating and appreciating - A three process questions card, mounted on a trigger Put together, these are a real project that has run one full round. The list of artifacts sorts into the four roles of Chapter 14, each in its place. It may still be crude, with some cells filled in "still exploring" or even "nobody knows." That is fine. One crude but honest round beats ten chapters of polished notes that never ran. Do not let it stop at this round. Raise it as your second spine. File results honestly and write the ugly ones down. Criteria always go in front of results. The map gets re-derived on a cadence. Next round's question grows out of the cells still empty on this round's map. The spine case of this book will go stale one day, and by then you will not need it. Right now, put the next review date for your ten-row map into your calendar, and re-derive it row by row when it comes due. ## 16.7 The unfair advantage you now hold You hold a real project you ran through the whole workflow with your own hands, an honest map you drew yourself, and a cadence that keeps both fresh. When the next dazzling demo arrives, the people who did not read this book start the "do we believe it" argument over again. You open your map and ask one question. Which row's pattern is this, and which rerun of it. The field will keep changing every month. Models will get stronger, headlines louder, each demo more dazzling than the last, and new "don't trust it yet" items will keep growing on the list. None of that matters. One more item on the list, one more repaid the way these sixteen chapters do it. Write it on the list first, then find the process step that answers it. The whole practice fits in one sentence you can memorize, the one Start Here laid down and the book never changed. Move at its speed, hold the line of science. --- # Chapter 1 templates Β· The "Is This Book for You" Self-Test > How to use. Answer the seven questions below honestly, one by one, and score each by the rule it gives. Total time about one minute. > This checklist is the full version of the four questions in section 1.5 of Chapter 1. It does not test your level. It tests your **situation**. The value of this book depends on whether a wrong answer from you really hurts something. The first six questions are scored. The seventh is not scored, and it decides how you read Part II. ## Self-test questions **1. The last "conclusive output" you handed over (a report / analysis / judgment / review). If its core conclusion were wrong, what would happen?** - Someone would make a costly decision based on it (spend money, set an architecture, launch a project, seek treatment, invest) β†’ **+2** - I would lose face, but no real decision depends on it β†’ **+1** - Nothing would happen β†’ **0** **2. Are you already using AI (or being asked to use AI) to produce this kind of content?** - Yes, and the frequency is rising β†’ **+2** - Not yet, but colleagues / competitors / upstream suppliers are β†’ **+1** - Not at all, and not in the near term β†’ **0** **3. Have you received material with deep AI involvement (reports, reviews, due diligence) where the sender did not say so?** - Definitely yes β†’ **+2** - Cannot rule it out β†’ **+1** - Definitely not (and I can say why I am sure) β†’ **0** - Hint: if you want to answer "probably not," that is "cannot rule it out," not "definitely not." **4. Can you accept the premise that "verification takes work"?** - Yes, I want the work made affordable, not made zero β†’ **+2** - I hope one day it will be reliable in one click β†’ **0** (this book will disappoint you) **5. Are you willing to accept honest but unsatisfying answers like "still exploring / nobody knows"?** - Yes, I would rather have an honest map than a categorical one β†’ **+2** - I need a definite conclusion for every question β†’ **0** **6. Do you have a real, unsolved problem in hand right now that can carry the exercises through the book?** - Yes, and I can write it as one sentence right now β†’ **+2** - No, but I want to find one β†’ **+1** - No, and I do not plan to β†’ **0** **7. Is your evidence something that can be rerun? (not scored)** - Yes, code, data, a ledger that can be recomputed, and a bad run can be redone β†’ you are the core reader this book writes the full process for, follow Part II as written - Partly, for example I have data but the collection cannot be redone β†’ read Chapter 3, section 3.6, check the five premises one by one, and convert wherever one breaks - No, wet lab, one-shot interviews, clinical data β†’ the steps still apply, but the shape of the case has to be converted, Chapter 3, section 3.6 is your exchange-rate table, read it before entering Part II ## Scoring - **8–12 points.** This book's primary reader is you. Start with Start Here and produce your first artifact the same day. - **4–7 points.** This book is useful to you, but read Part I (the frame) and Part III (trust) first, and enter Part II when a real problem shows up in your hands. - **0–3 points.** What you need right now is probably a good search or Q&A tool, not this book. That is not a criticism. Paying verification cost for answers that hurt nothing is waste. ## Self-check (Chapter 1 failure modes) - [ ] I have not taken "AI can make scientific discoveries" (AlphaFold-style headlines) as a reason to believe "the AI report on my desk is reliable." No credibility passes between the two. - [ ] I know a report's persuasiveness and its correctness are decoupled. To judge whether it is reliable, I look at the process that produced it, not at how smoothly the text reads. - [ ] Before forwarding any AI output, I have picked at least one number that would change the conclusion and asked, "where did this number come from?" - [ ] I have not abandoned the whole thing over one fabricated citation, and I have not signed and forwarded because it was "mostly right." Chapter 1 names both reactions. --- # Chapter 2 templates Β· The Transfer Map > This appendix goes with section 2.5 in the chapter (the transfer map) and section 2.6 (the three failure condition questions). --- ## Template 1 Β· Blank Transfer Map **How to use.** Build a "precedent β†’ now" transfer map for your own field. The precedent need not have anything to do with AI. Any historical change of the "a tool rewriting a craft" kind will do (how to pick one is Template 2, step 1). Aim for 5–12 rows. Fewer than 5 means the mechanism extraction did not go deep enough, more than 12 means you did not merge. Once it is built, run the self-check at the end. **Header information (fill these four cells first)** ```text My field / craft: ________________________ Chosen precedent: ________________________ (why this one: ________________) Date built: ____________ Planned review cycle: every ____ months Citation format for this table: row N of the transfer map ``` **The main map table** | # | Pattern (one sentence, write the mechanism not the ending) | How it happened in the precedent (give the event, give the numbers, mark anything from memory [check]) | The counterpart in my field (observable today, or mechanically due to appear) | Status (verified / still exploring / falsified) | Basis / to check | |---|---|---|---|---|---| | 1 | ________ | ________ | ________ | ________ | ________ | | 2 | ________ | ________ | ________ | ________ | ________ | | 3 | ________ | ________ | ________ | ________ | ________ | | 4 | ________ | ________ | ________ | ________ | ________ | | 5 | ________ | ________ | ________ | ________ | ________ | | … | ________ | ________ | ________ | ________ | ________ | **What the status column means** (the same as the rest of the book): - Verified = enough evidence is already observable in your field, not just "it happened in the precedent"; - Still exploring = it makes sense mechanically, the evidence on your side is undecided, and when in doubt put this; - Falsified = your field already has counterevidence. When not one row in the whole table says "falsified," be wary. It may not be that every pattern is right, it may be that you never looked for a counterexample. --- ## Template 2 Β· Guiding Questions for a Transfer Analysis of Your Own Field **How to use.** Fill in the five steps and pour the output straight into Template 1. A full analysis takes about 2–4 hours. Step 4 (the failure condition audit) is the easiest to skip and the one you can least afford to skip. Skip it and what you did is a one-point analogy, not a transfer. ### Step 1 Β· Pick the precedent ```text Tool changes my field (or an adjacent field) went through in the past 50 years, list every one I can think of: 1. ________________ 2. ________________ 3. ________________ Chosen precedent: ________________ Qualifying tests (all three have to be answered "yes," otherwise pick another): β–‘ Did it substantially change the cost structure of some stage? (which stage: ________) β–‘ Was it accompanied by panic or hype at the time? (what people said then: ________) β–‘ Can the outcome be reconciled by now? (at least one professional generation back, the books: ________) ``` ### Step 2 Β· Mechanism extraction (ask these of the precedent one by one) ```text 1. What had its production cost collapse? By roughly how many orders of magnitude? ________________________________________ 2. Once the old bottleneck disappeared, where did the new bottleneck move to? ________________________________________ 3. Which class of previously nonexistent error was born in bulk? Are they invisible, do they come in bulk? ________________________________________ 4. How was trust in output from strangers rebuilt? (graded delegation / review mechanism / origin tracing / other) ________________________________________ 5. Whose work had its "mechanical part" eaten? What is left that was not eaten? ________________________________________ 6. Which panics of the time missed? In what unexpected way did they miss? ________________________________________ ``` ### Step 3 Β· Write the counterparts ```text Translate each mechanism from step 2 into my situation today: Counterpart of mechanism 1: ________________ (observable today? β–‘ yes β–‘ no) Counterpart of mechanism 2: ________________ (observable today? β–‘ yes β–‘ no) Counterpart of mechanism 3: ________________ (observable today? β–‘ yes β–‘ no) Counterpart of mechanism 4: ________________ (observable today? β–‘ yes β–‘ no) Counterpart of mechanism 5: ________________ (observable today? β–‘ yes β–‘ no) Counterpart of mechanism 6: ________________ (observable today? β–‘ yes β–‘ no) Anything ticked "yes" is a candidate for "verified" (you still have to write the basis); anything ticked "no" gets "still exploring." ``` ### Step 4 Β· Failure condition audit (score each on "how far it holds," 0–10) ```text 1. Cheap ground truth. Does my field have a "compiler" (a right or wrong signal in seconds to minutes, nearly free)? How long one verification takes and what it costs: ________ Score: __/10 2. Feedback loop. The iteration cycle in the precedent vs my iteration cycle, how many orders of magnitude apart? ________________ Score: __/10 3. Error recallability. If it is wrong, can it be rolled back? Or does it enter public knowledge / the decision chain, get cited, get inherited? ________________ Score: __/10 4. Corpus coverage. How many "relatives" does what I am doing have in the training data? The closer to the frontier, the fewer. ________________ Score: __/10 5. Survivorship bias. Is the precedent I picked one that ended well? Was there a similar tool that drove an industry into a ditch with nobody to write its biography? Counter-precedent found: ________ Score: __/10 For any item under 5 points, every map row tied to that premise gets downgraded to "still exploring," and the reason for the discount goes in the "Basis / to check" column. ``` ### Step 5 Β· Land a falsifiable judgment ```text Based on row __ of the transfer map, I predict: ______________________________ If within ____ (deadline) I observe ______________________, this judgment is void and the map gets redrawn. (Every row has to yield at least one sentence like this. A row that cannot does not deserve to be called a judgment, only an impression.) ``` --- ## Quick card, the three questions before you rule with the map The condensed version of section 2.6, copy it out and stick it next to the map: 1. How far do the premises this pattern depends on (cheap ground truth, fast feedback, recallable errors) hold in my setting? 2. Am I transferring the **mechanism**, or transferring the **ending**? 3. How would this judgment be falsified? --- ## Self-check (run it before you hand the sheet in) - [ ] Every row's "pattern" is written as a mechanism (cost, bottleneck, error type, trust, roles), not as an ending ("everyone turned out fine" is not a pattern). - [ ] The status column is filled in, one of the three, and no "still exploring" was optimistically written up as "verified." When in doubt, downgrade. - [ ] Every event, number, and date written from memory is marked [check]. - [ ] You looked seriously for a counterexample or a counter-precedent at least once. An all-green table is a danger signal, not a good score. - [ ] Every row can state "what observation would overturn it" (step 5). Rows that cannot are either deleted or downgraded. - [ ] All five failure conditions are scored, and the map rows tied to the low-scoring ones are downgraded with the reason noted. - [ ] The header carries the date built and the review cycle. Maps expire, and a map with no review date is the same as no map. --- # Chapter 3 templates Β· Ladder Self-Rating Sheet + Seven-Step Workflow Check Card > This appendix is the fillable version of the two maps in Chapter 3. --- ## Template 1 Β· Ladder Self-Rating Sheet (Step Γ— Level Matrix) **How to use.** Score **one real project** you have in hand, not yourself and not your tools. Fill two cells per step, the current level (how it is actually done right now) and the target level (where you want to climb after reading the matching Part II chapter). Higher is not better for the target. For some steps the right target is to stop at assistant level (see the section 3.5 snapshot). Refill it each time the project goes around the loop. ### Level quick reference (read this first, then fill the table) Use the four criterion questions to rule on the level. When an answer changes, the level changes. | Level | Who drafts | Who reviews | Who decides | Who answers when it's wrong | |---|---|---|---|---| | Tool | You | You (a glance in passing) | You | You; errors visible on the spot | | Assistant | AI | You, every part in full | You | You; human review is the only line of defense | | Collaborator | AI (with self-checks and alternatives) | You, checkpoints + spot checks + a locked-in process | You | You + the process; errors hit the criteria first | | Autonomous researcher | AI (including intermediate decisions) | Mostly AI self-review, the human accepts only the end product | Goals and acceptance criteria stay with the human, the process goes to AI | Nobody (enter serious settings with care) | ### Self-rating matrix Project name: ____________ Date filled: ____________ Fill-in round ____ | Step | Current level (tool/assistant/collaborator/autonomous) | Target level | Upgrade precondition, what locked-in criterion or process this step needs first before you dare go up one level | Who (or what) found the last error at this step | |---|---|---|---|---| | 1 Master the field | | | | | | 2 Questions and hypotheses | | | | | | 3 Test plan | | | | | | 4 Execution | | | | | | 5 Read and catch errors | | | | | | 6 Deliver | | | | | | 7 Red team | | | | | Circle two cells. The step I most want to climb: ____________. The step I should least let go of (my "operating room"): ____________ ### Self-check - [ ] All seven steps filled with the same level? You are probably scoring "the whole thing." Levels live on "step Γ— task." Go back to section 3.4 and refill. - [ ] A step marked collaborator, but the "upgrade precondition" column has no locked-in criterion at all? Drop it back to assistant. Without criteria you have no standing to spot-check. - [ ] Read and catch errors marked collaborator or higher? Be wary. In the 2026 snapshot no step's errors disguise themselves better than this one's. - [ ] The whole target column says "autonomous"? Reread section 3.4. Turning every knob to maximum is not advanced, it is an operating room that fails the sterilization standard. - [ ] A cell where you cannot answer "who answers when it's wrong"? Use that cell at tool level, without exception, until you can. - [ ] The "who found the error" column is all "I happened to notice"? You have no process-level verification yet, and no step should go above assistant level. --- ## Template 2 Β· Seven-Step Workflow Check Card **How to use.** One card per step. Fill them all at project start, and refill the matching card when you get stuck or sent back. A blank you cannot fill is itself the diagnosis. It tells you where you are really stuck. How to do each step is in the matching Part II chapter (noted on the card). This card only locates, it does not instruct. ### Card 1 Β· Master the field (method in Chapter 4) > This step's goal is to turn "can't read it all" into "can ask it something," and know who claims what, what they are arguing about, and which argument my question lands in. - What I need to produce: ____________ (e.g., a controversy map) - Who is downstream (who takes it as input): ____________ - Drafter (me / AI): ______ Review mode (line by line / spot check / criteria): ______ Decider: ______ Who answers when it's wrong: ______ - The criterion for this step being "done": ____________ (e.g., I can predict what new evidence would change my map) - The signal that sends me back to this step: ____________ (e.g., at delivery there is a paragraph I cannot write clearly) ### Card 2 Β· Questions and hypotheses (method in Chapter 5) > This step's goal is to grind a blur of curiosity into a question worth answering, one whose answer might embarrass me. - What I need to produce: ____________ (e.g., one falsifiable hypothesis, with its scope) - Who is downstream: ____________ - Drafter: ______ Review mode: ______ Decider: ______ Who answers when it's wrong: ______ - The criterion for this step being "done": ____________ (e.g., I can say which observation would kill this hypothesis) - The signal that sends me back to this step: ____________ (e.g., the red team points out "the question itself is asked wrong") ### Card 3 Β· Test plan (method in Chapter 6) > This step's goal is to lock in the plan, the method, and the criteria before running, above all "what counts as losing." - What I need to produce: ____________ (e.g., a locked-in criteria list + baseline + definitions) - Who is downstream: ____________ - Drafter: ______ Review mode: ______ Decider: ______ Who answers when it's wrong: ______ - The criterion for this step being "done": ____________ (e.g., the criteria were signed off before anything ran, and are not changed afterward) - The signal that sends me back to this step: ____________ (e.g., during interpretation I find the criteria did not block some illusion) ### Card 4 Β· Execution (method in Chapter 7) > This step's goal is to turn the design into data (code, pipelines, computational experiments). - What I need to produce: ____________ - Who is downstream: ____________ - Drafter: ______ Review mode: ______ Decider: ______ Who answers when it's wrong: ______ - The criterion for this step being "done": ____________ (e.g., tests all green + results reproducible from raw data in one command) - The signal that sends me back to this step: ____________ (e.g., during interpretation a "finding" turns out to be a bug) ### Card 5 Β· Read and catch errors (method in Chapter 8) > This step's goal is to sort the results into findings, noise, and bugs. The most expensive error looks like the most exciting finding. - What I need to produce: ____________ (e.g., a table of "each conclusion + its possible sources of illusion") - Who is downstream: ____________ - Drafter: ______ Review mode: ______ Decider: ______ Who answers when it's wrong: ______ - The criterion for this step being "done": ____________ (e.g., every main conclusion has been checked against the standing list of illusions) - The signal that sends me back to this step: ____________ (e.g., the red team finds an alternative explanation I did not check) ### Card 6 Β· Deliver (method in Chapter 9) > This step's goal is to turn "I know" into "others can trust," with the right vehicle and every number traceable. - What I need to produce: ____________ (vehicle: paper / memo / report / decision document) - Who is downstream (who the readers are, what decision they make with it): ____________ - Drafter: ______ Review mode: ______ Decider: ______ Who answers when it's wrong: ______ - The criterion for this step being "done": ____________ (e.g., any number challenged with "where did this come from" gets its source within three minutes) - The signal that sends me back to this step: ____________ (e.g., I cannot answer the reader's first question) ### Card 7 Β· Red team (method in Chapter 10) > This step's goal is to let the harshest criticism happen at home before release. - What I need to produce: ____________ (e.g., an attack list + a disposition record for each item) - Who is downstream: ____________ - Drafter: ______ Review mode: ______ Decider: ______ Who answers when it's wrong: ______ - The criterion for this step being "done": ____________ (e.g., the strongest objection is written into the deliverable, not deleted) - Which step it sends me back to, and the signal: ____________ (the red team's job is to send you back, write down the step you are most likely to be sent back to) ### Self-check shared by all seven cards - [ ] Is Card 3 (test plan) empty? The step most often skipped entirely. When criteria are added after the fact, the conclusion always "happens to" support the plan finished first. - [ ] Is Card 7 (red team) empty? The step most often omitted. Without an internal red team, your first red team is the real world. - [ ] Was every card's "done" criterion locked in before starting? Criteria written afterward do not count. - [ ] Is there a card where all four criterion questions (draft / review / decide / answer for it) say "AI"? Downgrade it against the ladder self-rating sheet. - [ ] Read the "downstream" line across all seven cards in a row. Where the chain breaks is your project's real bottleneck right now. --- ## Sheet 3 Β· Five-premise comparison table (Chapter 3, section 3.6) **How to use.** Judge **one real project** you have in hand, premise by premise. Three minutes to fill. How many broke matters less than knowing where. For every broken premise, the method in each later chapter gets converted once by the rightmost column. When done, pin it next to the ladder self-rating sheet. | # | Premise | On my project | If broken, what fails first | My compensation | |---|---|---|---|---| | 1 | Errors can be found cheaply (the data is still there, it can be recomputed) | Holds / broken | The step 5 interrogation process (Chapter 8) | Move the interrogation forward to step 3. Locking in the criteria is my only chance at interrogation | | 2 | The evidence is machine-readable | Holds / broken | Dispatching mechanical checks (Chapter 12, lesson two) | Cleanup cost goes into the verification budget, and this bill is usually larger than the check itself | | 3 | The verifier is the producer | Holds / broken | L2's independent re-derivation (Chapter 12) | L0/L1 as written, no rerun capability needed; the economics of the spot-check rate are **still exploring**. Half a rule available, use the four attack surfaces of Chapter 10 as an acceptance checklist | | 4 | The criteria can be written before the run | Holds / broken | Step 3, and with it the whole chain | Lock in a proxy criterion + write "the gap between the proxy and the real target" as a formal limitation | | 5 | There is a next round | Holds / broken | Not one step, the conclusion the whole process produces | The downgrade ladder in section 12.6 of Chapter 12; the three floors of the "no next round" tier | **Three things to do once it is filled in.** 1. **If you broke two or more, pin this table at the front of the book.** Every Part II chapter you read needs one extra conversion, and the chapter where you forget the conversion will look usable and not actually hold. 2. **Readers who broke premise 3, pay special attention.** That cell is marked "still exploring" in this book, not "unimportant." The acceptance economics of second-person sign-off (how to set the spot-check rate, what to sample, how to escalate when a check fails) is the most visible gap in this workflow. Do not read "the book didn't write it" as "it can be skipped." 3. **For every broken premise, write one line in your project README or plan document.** Same reason as the downgrade record in Chapter 12. A premise that breaks without being written down leaves your deliverable looking exactly like one where every premise holds. **Self-check**: - [ ] All five "holds"? Look again at premise 3 and premise 5. These two are the easiest to judge optimistically. At project start you always feel there will be a next round, and on signing day you find you never reran anything. - [ ] Is the compensation column copied from the book? Translate it into concrete actions in your project, otherwise it is only a sentence you once read. - [ ] If you judged premise 4 as "holds," check once more. Is the criterion you wrote a real criterion, or a proxy you wrote down before thinking it through? A proxy is nothing to be ashamed of. Not admitting it is a proxy is. --- # Chapter 4 templates Β· Controversy Map Prompt + Paper Card + Frontier-Layer Subscription Rules + Coverage Checklist > How to use. The four tools match the three-layer intake workflow of Chapter 4. The map layer uses Template 1, the skeleton layer Template 2, the frontier layer Template 3, and Template 4 runs before you close out. Budget one afternoon for the whole first round. --- ## Template 1 Β· Controversy Map Prompt (map layer) **How to use.** Fill in the field and your specific question, and use the output as a first-draft map. The hard step cannot be skipped. Every paper AI lists, confirm one by one at Semantic Scholar / Google Scholar that it exists (about twenty minutes). This is the only cordon between you and "a castle in the air built out of fabricated citations." ```text I need to master a field, starting with a map. Field: [your field]. My specific question is: [the question you need to answer]. Give me: 1. The 3-6 main positions/camps in this field, and each one's core claim; 2. 2-3 representative works per camp (title, authors, year, venue); 3. The real points of disagreement between camps, not differences in wording, substantive conflicts of the form "if A is right, B is wrong"; 4. The 1-2 disagreements most relevant to my question. Requirement: list only papers you can give a real source for; mark anything you are unsure exists as "unsure". ``` The controversy map table (fillable version; rows are claims, and every paper you read afterward gets filled in): ```text | Claim | Papers that support it | Papers that oppose it | Substance of the disagreement | |---|---|---|---| | ____________ | ____________ | ____________ | ____________ | | ____________ | ____________ | ____________ | ____________ | ``` ## Template 2 Β· Paper Card (skeleton layer) **How to use.** Generate one card for every load-bearing paper (each camp's representative works, the repeatedly cited sources, the empirical studies tied directly to your question, usually five to fifteen of them). The last field is always left blank, and you fill it by hand after reading the key passages of the original. A card where you filled in "how much I believe it" is your own judgment. An AI summary is only a compression of somebody else's. ```text Read this paper and distill it in the format below, no embellishment: - Core claim (one sentence, in the paper's own wording) - Evidence (what experiment/data/task, at what scale) - Where the claim applies (limits the authors admit, in the limitations section and hidden in footnotes) - Which prior work this paper refutes or depends on - [Blank] How much I believe it: ``` ## Template 3 Β· Frontier-Layer Subscription Rules **How to use.** Set it up once after the map is built, then run it a fixed half hour every week. The admission criterion is locked in. Without a rule that permits letting things flow past, chasing the frontier is nothing but anxiety. ```text 1 Citation alerts: at Google Scholar / Semantic Scholar, set citation alerts for [3-5 load-bearing papers]. Whoever cites them may be shaking or reinforcing your map. 2 Weekly scan prompt (fixed output format): "This is my controversy map: [paste]. This is this week's new literature: [paste search results]. Output in three columns: new evidence relevant to the map / signals the map needs changing / noise I can ignore." 3 Admission criterion (locked in): a new paper is worth entering the skeleton layer if and only if it could change who wins a row of the controversy map. Let the rest flow past. ``` ## Template 4 Β· Coverage Checklist (before you close out) **How to use.** The blind spots of a single search path are systematic. Before you close out, go through the four-way cross-check one item at a time. AI never volunteers "I missed a community". The only antidote is this step, not a cleverer prompt. - [ ] **Keyword multipath**: have AI generate 5-8 sets of search terms from different terminology systems (the same thing goes by different names in different communities), and run a round on each; - [ ] **Citation graph**: start from the load-bearing papers and walk one layer forward (who cited it) and one layer backward (who it cited); - [ ] **Reverse test**: ask AI "if one paper could overturn my current map, what would it most likely look like and which community would it sit in", then go search whether it exists; - [ ] **Human anchor**: take the controversy map to someone who really knows the field and ask "what did I miss". --- ### Self-check (tick each item before you hand over the map) - [ ] **Fabricated citation**: has every paper AI listed been confirmed to exist in an academic search engine? The title is plausible, the authors are common names, the journal is real. A paper that does not exist is AI's most typical invention. - [ ] **Misremembered title**: for a citation you wrote from memory, did you go back to a first-hand source and check the title word by word? A title off by one word is, in search and under someone else's check, a paper that does not exist. The inventor is not necessarily AI, it can also be your memory. - [ ] **Secondhand drift**: has every claim entering a decision been checked against that passage in the original? Each hand it passes through drops a little of the qualifier, and the qualifier is exactly the part judgment needs most. - [ ] **The coverage illusion**: did all four paths of Template 4 really run? A fluent, complete answer is not complete coverage, especially when you want a "gap", where AI only wraps it to look more real. - [ ] Does the controversy map hold at least one line of "if A is right, B is wrong"? A map with no disagreement usually means you have not found the battlefield yet. --- # Chapter 5 templates Β· Five-Step Question-Sharpening Prompt Set + Question-Sharpening Card > How to use. Run the five prompts in order, and everything they produce lands on the question-sharpening card at the end. Budget one hour for the whole run. If you run well over, go back and read the "Over-sharpening" item in Chapter 5, section 5.8. --- ## Part one, the five-step question-sharpening prompt set ### Step 1, diverge candidates **How to use.** Fill in your direction, the controversy map from Chapter 4, and your real constraints. Item 4 is the mirror-risk detector, do not delete it. ```text My direction: [one sentence describing your vague direction]. Background: here is my controversy map: [paste or summarize: main camps / load-bearing papers / the real disagreement]. My constraints: [time budget] / [available data and hardware] / [the edge of my skills]. Generate 20 candidate research questions. Requirements: 1. Cover different grain sizes, from "worth a paper" to "worth an afternoon"; 2. Cover different positions, at least 5 of them questions the opposition or a skeptic would ask first; 3. Attach to each question one line, "what evidence could refute it," and if you cannot write that line, do not list it; 4. At the end, mark separately which questions have most likely already been asked in the literature, and what the clue is. ``` ### Step 2, score the three tests (AI as juror) **How to use.** The scores are yours to give (0/1/2), and this prompt only makes AI take the other side and supply attack angles you missed. "Testable" is a veto. ```text Candidate question: [paste one candidate question]. Attack it from three angles, with the strongest single objection for each: 1. Testability: is there any evidence that could kill it? If not, point out where it is a position rather than a question; 2. Worth answering: assume the answer is "yes," then "no," and what changes in action? If nothing changes, say it is decoration; 3. Affordable: under my constraints ([paste constraints]), what is the most expensive step on the road to evidence? Attack only, do not patch. Patching is my job. ``` ### Step 3, rewrite it falsifiable **How to use.** Write it yourself first, and ask AI for three versions to revise only when you are stuck. Anything that cannot fill every slot in the pattern goes back to step 2. ```text Rewrite the candidate question below into a falsifiable hypothesis, strictly following the pattern: "Under [conditions/basis], the [measurable metric] of [subject], compared with [control], is [direction and threshold]. If [specific observed outcome] is observed, the hypothesis is falsified." Candidate question: [paste]. Requirements: 1. Give 3 versions, thresholds from strict to loose; 2. For each version, state which slot was hardest to fill and what you assumed to fill it, since those assumptions are exactly what I will recheck by hand; 3. Do not use words that no observation can refute, such as "effectiveness," "possibility," "potential." ``` ### Step 4, action rehearsal **How to use.** Item 3 matters most, since middle outcomes are the likeliest to expose a hole in the hypothesis (this is how the cost clause of the Chapter 5 spine case got added). ```text Hypothesis: [paste the falsifiable sentence from step 3]. Reader/decision maker: [yourself / your boss / the investment committee / a reviewer]. Rehearse: 1. If the answer is "yes," what does that decision maker do? How is it different from now? 2. If the answer is "no," what do they do? How is it different from now? 3. List 2-3 middle outcomes that are "neither yes nor no" (partly holds, holds at extra cost, holds only on a subset), and check each one, does my falsification condition have a verdict for it? Where it is silent is a hole in the hypothesis. If the actions in 1 and 2 are the same, say so plainly, this question changes no action. ``` ### Step 5, reality-check the resources **How to use.** Let AI list the route and the costs. Whether it is walkable is your ruling. Accept only the shrink suggestions that do not change the nature of the question. ```text Hypothesis: [paste the falsifiable sentence]. My resources: [time] / [budget and compute] / [available data] / [the skill range of me + AI]. Give me: 1. The shortest route from the hypothesis to the "specific observed outcome," listed step by step; 2. The biggest cost item at each step (time/money/difficulty of getting the data); 3. If the whole route exceeds my resources, give 3 ways to shrink it (shrink the task scope / shrink the metric / shrink the control), and mark each one, after shrinking that way, is the question still the same question. ``` --- ## Part two, the question-sharpening card (fillable) **How to use.** One card per complete run of the five steps. Sign it when it is filled, and pin it to the first page of the Chapter 6 test plan. An empty box is a step of the process not finished. ```text ════════════ Question-sharpening card ════════════ Date: ____________ Project: ____________ [Direction] (raw material, vague is allowed) ____________________________________ [Candidate pool] (output of step 1) - Total candidates: ____ Of them marked "already asked": ____ - Top three entering scoring: A. ________________________________ B. ________________________________ C. ________________________________ [Three-test scores] (0/1/2; a 0 on testable = veto) Testable Worth answering Affordable Total Candidate A: ___ ___ ___ ___ Candidate B: ___ ___ ___ ___ Candidate C: ___ ___ ___ ___ Survivor: ____ [Falsifiable sentence] (output of step 3, both lines required) Under __________________ (conditions/basis), the __________________ (measurable metric) of __________________ (subject), compared with __________________ (control), is __________________ (direction and threshold). If __________________________________ is observed, the hypothesis is falsified. [Action rehearsal] (output of step 4) - Action if the answer is "yes": ______________________ - Action if the answer is "no": ______________________ - Are they different? ☐ Yes (pass) ☐ No (back to step 2) - Middle outcomes and their verdicts (at least one): ____________________________________ [Resource check] (output of step 5) - The most expensive step: ______________________ - Walkable? ☐ Walkable ☐ Walkable only after shrinking (what was shrunk: ________) ☐ Not walkable - If not walkable: ☐ Switch candidate ☐ Back to Chapter 4 for a different battlefield ☐ Shelve it honestly [Sign-off] I ruled on this question. AI only supplied candidates and counterpoints. Signature: ____________ ══════════════════════════════════════════════════ ``` ### Self-check (tick each before submitting) - [ ] My "question" has a finished state, and I can say under what conditions it counts as answered or falsified (otherwise it is a direction). - [ ] The falsifiable sentence has no unkillable words like "effectiveness/possibility/potential" (wish-list hypotheses). - [ ] The surviving candidate is my ruling, not AI's recommendation (this step has no cheap ground truth, and nobody answers for AI's recommendation). - [ ] I read the "already asked" marks and confirmed the survivor is not among them (the mirror risk, AI candidates lean toward old questions in the corpus). - [ ] If scoring relies on an LLM judge, I have confirmed that judge preference will not contaminate the measurement into a second research question (if that cannot be done, change the task). - [ ] The action rehearsal covered middle outcomes, and the falsification condition has a verdict for every one of them. - [ ] The whole run took about an hour. If it ran far over, I confirm I am not using sharpening to put off starting. - [ ] Threshold numbers (Ξ΅ and the like) have their reasons recorded, left for the final sign-off at the test design preregistration (Chapter 6). --- # Chapter 6 templates Β· Test Plan Template + Confounder Checklist + Plan Red-Team Prompt > This appendix is the complete fillable version of the three tools in Chapter 6. > The order of use is fixed. Fill Template 1 first, run every line of Template 2 when you reach item 6, and run one round of Template 3 before you sign. --- ## Template 1 Β· Test Plan (preregistration-style fillable version) **How to use.** Fill it in and stamp a timestamp before you run anything; after sign-off, only appended change-log entries, no edits. Not filling every blank is fine. The blanks you cannot fill are exactly where you have not thought it through. Start the run with blanks still in it and they will fill themselves in along your preference once the results are out, which is the "degrees of freedom colluding with motive" of Chapter 6. ```text # Test plan Project: ____________ Drafted by: ____________ Draft date: ____________ Sign-off date: ____________ Timestamp method (git commit / email / preregistration platform): ____________ ## 1 Question The question this test has to answer (one sentence): ____________________________________________ ## 2 Hypothesis H (the Chapter 5 output, copied as is) Falsifiable statement: ____________________________________________ Scope (within what tasks/populations/conditions it holds): ____________________ Falsification condition (what observation kills H): ____________________ ## 3 Arm design Main arm (my design): ____________________________________________ Baseline arm (control): ____________________________________________ - Tuning/optimization budget the baseline gets: ______ (must equal the main arm; if not, write the reason) Steelman arm: - The strongest opponent's sentence ("your result is really just ____"): ____________ - The arm built to block that sentence: ____________________________________ Other arms (one line each, stating which alternative explanation it rules out): - ____________________________________________ ## 4 Basis Primary basis (how it is computed, down to the formula): ____________________________ Reason for making it primary (usually = the unit the final reader thinks in): __________ Sensitivity basis (listed separately, never mixed with the primary): ______________________ Commitment (copy as is, effective on sign-off): if the two bases reach opposite conclusions, report it honestly, no picking. ## 5 Criteria and falsification condition What counts as a win (what number triggers it): __________________________ What counts as a loss: ____________________________________________ What counts as undecided (how to word it when neither is met): ________________________ Statistical test: ____________ Significance level: ____________ Sample size / repeats per configuration: ____________ Random seed: ____________ Stopping rule (when to stop running, set first, guards against "run until significant"): __________ Falsification rehearsal record (invent a set of numbers, confirm the criteria really rule against me): ____________ ## 6 Contamination and confounder check Run every line of Template 2. Lines that do not clear: - Line: ______ Disposition (mitigation / written into limitations): ______________ ## 7 Filing and change log Timestamp: ____________ Change log (append only): - Date: ______ What changed: ______ Reason: ______ Did this change happen after the results were seen: ______ (yes β†’ related conclusions downgraded to exploratory) ``` **Self-check** (run through before signing): - [ ] Is the falsification condition in item 2 empty, or written as "judged as a whole once the results are in"? That is not criteria, that is decoration. Go back to Chapter 5 and grind it again. - [ ] Cannot find the steelman arm in item 3? Your control only proves "better than doing nothing." Ask once more. "If the result comes out as I want, what would the strongest opponent say?" Give that sentence an arm. - [ ] Cannot fill the tuning budget field on the baseline arm? The review meeting in section 6.1 is waiting for you. - [ ] Is the "stopping rule" in item 5 blank? "Run until significant" is one of the best hidden degrees of freedom. - [ ] Skipped the falsification rehearsal? Watch it go red first. Criteria that have never been able to fail do not count when they are green. - [ ] Does the criteria wording carry "a reasonable range," "as appropriate," "at discretion"? Every one of them is a backdoor. Turn it into a number. - [ ] No timestamp? A plan with no timestamp cannot prove three weeks later that it was "locked in first." --- ## Template 2 Β· Confounder Checklist **How to use.** Run it line by line while filling item 6 of Template 1. The goal is not to tick every line, the goal is **honest disposition**. Tick what clears; for what does not clear, write a mitigation, or write it honestly into the limitations. An all-green checklist is itself suspect. You can let AI run your plan against this list line by line first (breadth is its strong suit), but the final tick on every line is yours. **A Baseline fairness** - [ ] Did the baseline get a tuning/prompt-optimization budget equal to the main arm? - [ ] Is the baseline's version/configuration the strong form of that method rather than a straw man (default parameters, an outdated version, an obviously suboptimal setting)? - [ ] If the baseline is a number from someone else's paper, are the runtime environment and the data slice comparable to your main arm? Or should it be rerun? **B Budget and basis alignment** - [ ] Are the resources both sides consume (money / compute / number of calls / person-hours) measured on the same basis? - [ ] Is the definition of "same budget" locked in? (Otherwise budget alignment is itself a degree of freedom to fiddle with afterward) - [ ] Are the primary basis and the sensitivity basis listed separately, with a commitment to report both? **C Alternative explanations** - [ ] Have you listed at least three cheap explanations that "explain the same result without your hypothesis"? - [ ] Does every cheap explanation point at an arm or a step in the plan built to rule it out? - [ ] Does the cheapest explanation of all have an arm of its own? (The lesson of the spine case, the two-arm design of army vs frontier cannot tell "teaming works" from "spending works," until the same-budget self-consistency arm is added) **D Data contamination** - [ ] Could the evaluation data have been "seen" by the model during training? (Public benchmarks, public survey data, question banks circulating online, treated as seen by default) - [ ] Could it have been "seen" by your own development process? (Looking at the same validation set over and over while tuning = human overfitting) - [ ] Is the contamination check a step written into the plan, or one line saying "should be fine"? **E Criteria backdoors** - [ ] Is the metric unique and set beforehand? (Counting several metrics equals picking the metric afterward) - [ ] Is the data slice set beforehand? ("Significant on some subset" counts only when that subset was declared in advance) - [ ] Is the exclusion rule for outliers set beforehand? - [ ] Is the stopping rule set beforehand? **F The measurement itself** - [ ] Could the way scoring and judging works favor one arm? (The spine case ruling out LLM-judged tasks in Chapter 5 is exactly this line, judge preference would contaminate the measurement into a second research question) - [ ] If there is a human judging stage, does the judge know which arm a sample came from? (Blind it wherever you can) --- ## Template 3 Β· Plan Red-Team Prompt **How to use.** Use it the moment you believe the plan is final but have not signed it, since the earlier you take the beating the cheaper it is. The output is a draft list, not a ruling. Go line by line, what you adopt turns into a change to the plan, what you reject gets one line of reason on file (that rejection record becomes your ammunition at the defense later). Mind the division of labor. Red-teaming **conclusions** is Chapter 10's business. What gets red-teamed here is a **plan** that has not run yet. ```text This is my test plan: [paste the full plan, including arm design, basis, criteria] Your job is to overturn it, not to improve it. Assume you are the reviewer who least wants this conclusion to hold. 1 List every degree of freedom in the plan that "can still be moved after the results are in"; 2 For each one, say which conclusion it would favor if adjusted after the fact; 3 Give one cheapest alternative explanation that, with the arm design unchanged, would produce the same result; 4 Point out which criterion's wording leaves a backdoor (words like "as appropriate," "a reasonable range"); 5 Check the baseline arm. Is it the strong form of that comparison method? Where does it look like a straw man? 6 If you were to add one arm built to make trouble for my conclusion, which arm would you add and why? Raise only problems specific enough to act on, no general methodology advice. Output each one in three parts, "the problem β†’ who it favors β†’ how to fix it." ``` **Self-check** (run through when you use the red-team output): - [ ] Did you adopt only the lines that were easy to hear? The ones that deserve the closest look are exactly the ones that irritate you. - [ ] AI reported no serious problem? Do not take that as evidence the plan is solid. In the words of the Chapter 3 snapshot, the most expensive flaws in a plan (unfair baseline, basis drift, a criteria backdoor) are exactly the kind it does not flag itself. When the red-team output is empty, the confounder checklist (Template 2) still gets run. - [ ] Planning to reject a "new arm" it proposed? Write the reason for rejecting before you reject. What you cannot write out, you should probably adopt. - [ ] Changed the criteria after the red team? Legitimate, and changing them at this moment is exactly the point. But the change still goes into the change log. Want to change them after sign-off and a different rule applies (see item 7 of Template 1). - [ ] Only ran one round? Fix it and feed it another round, until a new round returns only old lines you have deliberately rejected. --- # Chapter 7 templates Β· Minimal Harness Checklist + Ledger and Resume Patterns + "Have AI Audit the Harness" Prompt > This appendix is the complete fillable version of the three tools in Chapter 7. > The order of use is fixed. Build the skeleton on the patterns of Template 2 when you start, go through Template 1 item by item before you spend real money, and once the checklist is done, run a round of Template 3 before the pilot. --- ## Template 1 Β· Minimal Harness Checklist **How to use.** Go through it item by item before you spend real money or real time for the first time. Every item matches a class of real accident from Chapter 7, and the bill for that accident was three hours and five bugs. The goal of this sheet is to make you pay only for the checking, never for the accident. Do not force a tick on an item you cannot pass. Fix it before you start, or write down the risk you are accepting. **A Tracer bullet (mock, whole chain)** - [ ] Is there a mock mode (fake model or fake data source) that runs the whole chain **offline**, load β†’ call β†’ extract β†’ score β†’ write to disk β†’ summary report? - [ ] Is the format of the mock output identical to a real run (the aggregation script eats mock data without errors)? - [ ] After every change to the pipeline logic, has the mock regression been rerun? (it is your only free whole-chain test) **B A crash is a lossless event (resume)** - [ ] Does every result written to disk carry a unique key (the minimal tuple that fully rebuilds this call, such as task Γ— configuration Γ— seed)? - [ ] Have you run the crash drill, kill the process midway, restart, and results come back **neither duplicated nor missing** (verified by counting lines and reconciling the ledger)? - [ ] Are the tasks flattened into small independent pieces rather than one long serial chain? (a long chain's mid-state cannot be expressed as a key, the root of the crashes in Chapter 7, 7.1) - [ ] Does a single task's failure go onto the failed list only, leaving the run alive, with resume retrying it naturally? (verified with your own eyes by planting a "poison pill" task) **C The ledger (append-only + hard cap)** - [ ] Does every external call (API / compute / human annotation) write one line to an append-only ledger, recording timestamp, target, usage and cost? - [ ] Does the budget hard cap exist, and have you **seen it fire with your own eyes** (set the ceiling near zero and try one run, a fuse that has never blown is not a fuse)? - [ ] Are the ledger and the result file separate, results can be cleaned and rerun, the ledger never edited? **D Criteria and records** - [ ] Is the commit timestamp of the criteria file (the Chapter 6 plan or preregistration) **earlier than** the first line of results? - [ ] Does every change after the run starts go into an append-only change log (date, what changed, reason), with the locked criteria themselves untouched? - [ ] Is the original answer text (the faithful extract) written to disk (not only the score), so that after a scorer bug fix all scoring can be replayed offline? **E The pilot (the only test for a seam)** - [ ] Is a pilot planned at 3-5% of the total budget, and before the full run? - [ ] Have the pilot numbers been read by a human arm by arm and configuration by configuration, with the original answer text spot-checked behind every 0 and every perfect score? - [ ] Have you asked the dedicated question, "**Does the strongest arm or configuration's performance make sense?**" (the theme sentence of Chapter 7, bugs zero out the strongest arm first. When the harness has bugs, the most likely bias in the readings is systematically wronging the strongest contestant) **Self-check** (compare while you run the sheet): - [ ] All tests green, so you skipped the pilot? Green only proves the code matches the world you defined. All five bugs lived on seams outside that definition (API contracts, dependency environments, data quirks). - [ ] Was the checklist ticked after the full run finished? The accident it prevents has already happened, and what you need now is the interrogation checklist of Chapter 8. - [ ] Is group B resting on "resume is supported in theory"? A resume that has never had a crash drill runs its drill during the first real crash. - [ ] Taking "it is running" for "it is producing"? Logs, progress bars and the bill are all moving, and that does not make the numbers trustworthy. There is only one test, can every number point to the criteria, the ledger and the original answer text. --- ## Template 2 Β· Ledger and Resume Patterns **How to use.** Three language-independent design patterns, each implementable in a few dozen lines. Copy the structure, not the code. Your key, your cost unit and your storage format are decided by your experiment. The three patterns together are the structural base for what Chapter 7 called "leaving errors nowhere to hide." **Pattern 1, the resume key (idempotent writes)** ```text key = (task family, item ID, arm/config, seed) # the minimal tuple that uniquely rebuilds this call result file = append-only JSONL, every line carries the full key + original answer text + score At the start of a run: done = { keys of the lines already written } task pool = [ all combinations ] - done # flat, unordered, mutually independent During the run: each piece finished β†’ append to disk at once (with a write lock under concurrency) a piece fails β†’ record it on the failed list and print, do not throw; resume retries it naturally Effect: a crash mid-run = a lossless event, restart and continue, no rerun, no second payment a flat task pool unlocks concurrency for free (Chapter 7, serial full run ~15 hours β†’ 32 lanes in one pass) ``` **Pattern 2, append-only ledger + budget hard cap** ```text Immediately after every external call: ledger.append({ timestamp, model/resource, usage (tokens etc.), dollars, task metadata }) if cumulative spend > hard cap: throw, abort the whole run Discipline (more important than the code): the ledger is append-only, never edited. Results record the current understanding and can be corrected; the ledger records history, and history does not accept edits the result file is separate from the ledger, cleaning mislabeled result lines moves no ledger line Dividend: "how much did the experiment cost" = the sum of a column, not an impression the basis promise of a cost-matched comparison (Chapter 6) is redeemed in this ledger ``` **Pattern 3, write the original answer text (the faithful extract) to disk, scoring stays replayable** ```text what goes to disk is the original answer text (the faithful extract of the model's reply); the score is only a derived column a scorer bug fixed β†’ the replay script recomputes every score, at zero cash cost (the Chapter 7 case, fixing the % suffix bug and the numpy bug = change the scorer + replay, no second payment to the API) generation and scoring are orthogonal. The same original answer text gets reviewed again and again by the interrogation of Chapter 8 and the red team of Chapter 10. An experiment that stored no original text has no evidence to overturn anything later ``` **Self-check**: - [ ] Did the key design miss a dimension (the seed, say)? Resume will "skip" combinations that never ran, and the gap is silent. Verify with the crash drill, not with your eyes. - [ ] Is the ledger accumulating in memory and written to disk once at the end? A crash loses the account, and the hard cap goes with it. Write every entry at once. - [ ] Stored only the scores and not the original answer text? Every scorer bug then costs the full amount again, and the audit of Chapter 8 has nowhere to start. --- ## Template 3 Β· "Have AI Audit the Harness" Prompt **How to use.** Hand the harness's core code to an AI session that **did not help write it** for audit. The session that wrote the harness is blind to its own assumptions, and changing the session changes the viewpoint. This is the executable version of Chapter 2's principle that generation and verification are orthogonal. The output is a list of leads, not a verdict. Every finding comes with a minimal verification experiment, and you rule only after running them one by one. Use it before the pilot, so the most expensive bugs die before the money is spent. ```text This is the core code of my experiment harness: [paste: model/service calls, answer extraction, scoring, writing to disk, the ledger module] This is the test plan it has to implement (the criteria part): [paste: the arm design, the basis and the criteria of the Chapter 6 plan] Your job is to find the ways this harness "quietly produces wrong numbers," not to improve its code style. Work through four seams one by one: 1 The code ↔ API seam. For each model or service, check the code's assumptions against its real API contract. Will parameter names and values be refused (a reasoning model refusing a custom temperature), truncation behavior, null or empty returns, non-JSON responses, rate limits and retries. 2 The scorer ↔ data seam. Pull 10 real samples from the task data and walk the scoring logic by hand. Do format variants (units, % suffixes, thousands commas, capitalization, multiple lines) score a correct answer wrong? Is the "tidiest" way of writing a correct answer exactly the parser's most fragile path? 3 The grader ↔ environment seam. The libraries, subprocesses and timeouts the grading depends on, do they install fully and run in a clean environment? When a dependency is missing, does it error, or quietly record 0? 4 The harness ↔ criteria seam. Compare the basis the code implements against the basis the plan locked in, item by item. Sample size, slices, seeds, repeat counts, cost formula, where do they quietly disagree? For each finding, output three parts: the seam's location β†’ the worst consequence (which arm or configuration's numbers get contaminated first, and in which direction) β†’ one minimal experiment that verifies it on the spot (one command or one sample). Finally answer this. If only one 20-problem pilot is allowed, which three numbers are most worth checking by hand? Why those? ``` **Self-check** (run through it when you use the audit output): - [ ] AI says "no serious problems found"? Do not take it as a safety certificate. All five bugs of Chapter 7 came out of a harness AI helped write deeply, with 19 tests all green, and every one was the kind it does not flag. When the audit comes back empty, the checklist of Template 1 gets run all the same. - [ ] Audited the code only, without giving it the API docs and real data samples? The most expensive bugs are not in the code, they are on the seam between the code and the world (where bugs one, three and five are registered). Feed it nothing from the world's side and it audits the world you defined. - [ ] Is the audit session the same as the writing session? It will defend its own assumptions. Change the session, and preferably the model. - [ ] Were the problems the audit found fixed in code only, never entered in the change log? A fix that touches a difference between the plan and reality (an API refusing a preregistered decoding parameter, say) is disposed of the way Chapter 7's pillar four does it. The harness obeys reality, the difference is appended to the change log, and the locked criteria do not move. - [ ] Did the findings list skip the minimal verification experiment on each item? An AI audit also produces false positives. Leads have to be verified, and the code it writes has to be piloted, one and the same discipline. --- # Chapter 8 templates Β· Result Interrogation Checklist + Preregistered/Post-hoc Side-by-Side Report Template + Surprise-Result Red Flags > This appendix is the complete fillable version of the three tools in Chapter 8. > The order of use is fixed. The moment a result arrives, scan Template 3 first (the red flags, two minutes). If any flag is hit, or you notice you want to announce, run Template 1 (the interrogation checklist). Whatever the interrogation finds, deliver the numbers with Template 2 (the side-by-side report), no exceptions. --- ## Template 1 Β· Result Interrogation Checklist (fillable) **How to use.** The trigger is "wanting to announce," not "feeling something is wrong." You do not need to doubt it to interrogate, you only need to notice that you are excited. Fill it in before any number leaves your desk. The interrogation itself is post-hoc analysis. Everything it produces goes into the "post-hoc" column of Template 2, and the preregistered numbers may not be revised. ```text # Result interrogation record Project: ____________ Interrogator: ____________ Date: ____________ Source of the result (repo / file / commit): ____________ ## 0 Announcement sentence (write it first, as the autopsy subject) The sentence I originally wanted to say out loud (not a word changed): ____________________________________________ Compared with my prior expectation/odds, this result is: as expected / better / worse (circle one) My stake in it (does it benefit me if it holds): ____________________ -> If either "better than expected" or "benefits me" holds: every item on this checklist is mandatory, no sampling. ## 1 Question one: can the scorer be trusted - Did you pull out the original text of the samples scored wrong (all, or a random sample, no picking): ______ - Do the wrong answers have a pattern (numerical relation / fixed format / concentrated in one class): __________ - Messy wrongness (like a capability boundary) or tidy wrongness (like a sick scorer/gold answer): ______ - Did you spot-check the gold answers themselves, is their definition of "correct" ambiguous: ______ - Did you walk one sample through every link of the scoring chain (parse -> match -> score): ______ ## 2 Question two: what does the data look like - Have you seen the raw problem text/samples with your own eyes (not a statistical summary, the originals): ______ - How many "molds" do the samples come from (templates / sources / batches): ____________ - Estimated effective sample size (number of independent units, not rows): ____________ - Does the CI / significance test assume independent samples? Does that assumption hold: ______ - Could the sampling method (first N rows / random / stratified) introduce structure: ______ ## 3 Question three: where is the win concentrated - Did you spread the effect out by problem/slice (per-problem difference, ordered by id): ______ - Is the effect spread evenly or concentrated in a small handful: ____________ - If concentrated: did that handful get its own interrogation (back to questions 1 and 2): ______ - The control arm's performance on the same handful (an anomaly in the same direction = a data problem signal): ______ - With that handful removed, how much effect is left, and did the direction change: ____________ ## 4 Ruling (signed by a human, AI may not sign on their behalf) Fate of the announcement sentence: alive as is / alive after narrowing (new wording: ______) / dead Post-hoc breakdowns entered in the report side by side per Template 2: ______ Interrogation record archived (location): ____________ ``` **Interrogation dispatch prompt** (the independent channel for handing the mechanical work to AI; no expectations of yours in the brief, and no "help me confirm"): ```text These are the raw per-problem results of an experiment (data/file path attached). Do factual organization only, and make no evaluation of "whether the result is credible": 1 List every sample scored wrong for [arm X]: sample id, the model's raw answer, the gold answer, and the numerical/literal relation between the two, ordered by id; 2 Summarize patterns in the wrong samples: numerical relation (such as always k times), format features, id distribution (whether concentrated in a continuous stretch); 3 Cluster the problem text of all samples by template: after removing proper nouns and numbers, how many literal templates remain, and how many problems in each; 4 Output a per-problem table of [arm A βˆ’ arm B] differences, and mark the stretches of samples where the difference concentrates. Forbidden: any ruling that "the overall conclusion is reliable/unreliable"; that is not your job. ``` **Self-check**: - [ ] Did you interrogate only the undesirable results? The interrogation trigger is surprise and stakes, in either direction. Favorable results must go through the same checklist. - [ ] Did you ask AI "is this result credible"? That hands the judge's seat to sycophancy. Dispatch again, factual organization only, and make the ruling yourself. - [ ] Are the wrong examples "I spot-checked a few and found nothing"? Sampling is either random or complete. "Picking a few to look at" avoids exactly the ones that hurt most. - [ ] Is the effective sample size field just the row count? Samples from the same template/source/batch are not independent units. The 150 problems of Chapter 8 held only ~3 template families. - [ ] Did the interrogation change the conclusion, and then you reported only the post-interrogation numbers? Go read the first rule of Template 2. --- ## Template 2 Β· Preregistered/Post-hoc Side-by-Side Report Template (fillable) **How to use.** For any test with a preregistration (or criteria locked in beforehand, Chapter 6 style), deliver the numbers with this table. There are only three rules, all hard. β‘  The preregistered numbers are **reported as is**, however much you dislike them after the interrogation; β‘‘ post-hoc breakdowns are **reported side by side**, each labeled "post-hoc" with the reason for the breakdown stated; β‘’ **under no circumstances may a post-hoc number replace a preregistered one**. A post-hoc number that wants promotion gets preregistered and retested in the next round. ```text # Result report: ____________ (project / experiment name) Preregistration file and timestamp: ____________ Results file and commit: ____________ Audit script (if any, with commit): ____________ | Basis | Preregistered result (as is) | Post-hoc breakdown (labeled post-hoc) | Reason for breakdown | |---|---|---|---| | ____ | ____________ | ____________ | ________ | | ____ | ____________ | ____________ | ________ | ## Criteria reconciliation (walk the preregistered win/lose/undecided conditions one by one) - Falsification condition 1: ____________ -> triggered / not triggered - Falsification condition 2: ____________ -> triggered / not triggered - Status of hypothesis H (pick one + one qualifying clause): falsified / not falsified but not holding across the board (qualifier: ____________) / holds ## Known statistical weaknesses (list them honestly) - Effective sample size issue: ____________ - Status of the CI's independence assumption: ____________ ## Change log pointer All post-hoc analyses this report involves are registered at: ____________ (date + entry) ``` **Filled example** (the spine case's math family, excerpted from Chapter 8): | Basis | Preregistered result (as is) | Post-hoc breakdown (labeled post-hoc) | Reason for breakdown | |---|---|---|---| | math army vote vs frontier | +18.2pp [+12.7, +24.2], army_ahead | After removing the ambiguous template βˆ’2.3pp [βˆ’4.0, βˆ’0.7], direction reversed | All 49 frontier wrong answers = 4Γ—gold with %, a single ambiguous template (two readings, absolute percentage points vs relative multiple) | **Self-check**: - [ ] Does the table have numbers only in the post-hoc column, with "see above" in the preregistered column? Not allowed. Both columns in the same table in the same position, so the reader compares across in one glance. - [ ] Does the post-hoc column lack a reason for the breakdown? A breakdown without a reason is indistinguishable from picking data. - [ ] Was "not falsified" written as "holds"? A falsification condition not triggered β‰  the hypothesis holds. Chapter 8's H is the standard example, not falsified, not holding across the board, landing narrowed to "which tasks have a shot." - [ ] Did the sensitivity basis / the undesirable set of accounts stay out of the table? The commitment signed in Chapter 6. If the two bases reach opposite conclusions, report it honestly, no picking. - [ ] Did the audit/breakdown stay out of the change log? Post-hoc analysis is allowed. Post-hoc analysis without a record is not. --- ## Template 3 Β· Surprise-Result Red Flags **How to use.** Scan it the moment a result arrives, two minutes. Each flag on its own has an innocent explanation. A flag does not mean "the result is wrong." It means "check here first." Hit any one, go into Template 1, and the matching "first move" is where the interrogation starts. Hit three or more, treat the announcement sentence as a condemned prisoner first. | # | Red flag | What it usually means | First move | |---|---|---|---| | 1 | The effect beats the odds you set beforehand yourself | Either you are about to get rich, or the scorer/data is sick, and the latter is far cheaper | Into Template 1, every item, no sampling | | 2 | The control arm shows an anomaly in the same direction (the control also "won" where it should not) | The effect does not come from your treatment, it comes from shared data or a shared scoring chain | Check the links both arms share: problem set, parser, gold | | 3 | The strong player dies on an easy task (a frontier-grade model gets grade-school problems wrong) | "'select' Isn't Broken": suspect the gold answer first, the world second | Pull the original text of the wrong answers, look for a numerical/format pattern | | 4 | The wrong answers are highly regular (always k times, always off by a constant, always with a certain suffix) | Not a capability boundary, it is question ambiguity or a scorer parsing defect | Read the problem text word by word, look for two defensible readings | | 5 | Wrong answers concentrate in a continuous id stretch / a single source batch | A data structure problem: same-template variants, sampling not shuffled | Cluster the problem text by template, re-estimate the effective sample size | | 6 | The CI is abnormally narrow (relative to sample size and task noise) | Samples are correlated, the independence assumption is bankrupt, the CI is falsely confident | Count independent units; run a clustered robustness check by template/batch | | 7 | Every metric improves at once, without exception | Real improvement rarely blooms everywhere; a shared-source error does | Find a pair of metrics that ought to trade off, and see whether it also "wins both" | | 8 | Remove a small handful of samples and the effect vanishes or flips sign | The conclusion hangs on that handful | That handful gets its own interrogation (Template 1, questions 1 and 2) | | 9 | The result confirms exactly the prediction you have already said in public | The desirability flag: the drive to check is at its lowest right now, prime ground for motivated collusion | Symmetry discipline, run the full set as for red flag 1 | | 10 | You are already thinking about how to word the announcement | The trigger itself | Stop, write down the announcement sentence first, then into Template 1 | **Self-check**: - [ ] While scanning, did you find an "innocent explanation" for a flag and skip it? The innocent explanation goes into the record after the interrogation. It may not serve as an inspection waiver before it. - [ ] Zero hits on the whole table and a mediocre result? Low risk, an L0 spot check (Chapter 12) is enough. Do not run the interrogation checklist as a ritual. - [ ] Zero hits on the whole table but a major result? The stakes themselves are a variant of flag 9. Into Template 1. - [ ] Nobody has hit red flag 9 in a long time? Most likely it is not that you have no desirable results. It is that you have not looked at yourself. --- # Chapter 9 templates Β· Two Real Deliverables + Claims List / Technical-Report Skeleton / One-Page Memo Templates + Delivery Prompt Set > This appendix comes in two parts. Part one holds the book case's **two real deliverables**, the technical-report skeleton (shaped like a workshop paper) and the CTO one-page memo from the small-model army experiment. Every number points at a results file and a commit in the smol-army repo (`code/smol-army`) and can be recomputed one by one. The honest account. These two documents were delivered to this book's case library. **Nothing was submitted, no conference accepted them, and no CTO ever signed off on them.** The book does not invent those events. Part two is the fillable template version. > The order of use is fixed. Write Template 1 first (the claims list) β†’ use Template 2/3 to produce the two vehicles β†’ run the prompt set's interlock check β†’ the signature test sentence by sentence. --- # Part One Β· Real Deliverables ## Deliverable one, the technical-report skeleton (shaped like a workshop paper, real numbers) ```text Title: When does a small-model army pay off? A preregistered five-arm comparison with real cost accounting and an equal-budget self-consistency control (Working title: When Does a Small-Model Army Pay Off? A Pre-registered, Cost-Accounted Comparison with a Self-Consistency Control) === Abstract === Can a team of small open-source models tie a single frontier model on a cost-matched basis? The literature splits into two camps, but the decision-grade comparison combining "dollar-basis accounting + an equal-budget self-consistency control + several task families" is missing. We preregistered (criteria before any result commit) a five-arm comparison. GPT-5.6-terra single shot vs the vote/debate/division armies of three open-source small models (total parameters ≀21B) vs self-consistency on the strongest member, across three task families (GSM-Symbolic 150, MMLU-Pro 150, HumanEval+ 100), total cost $5.58. The result is highly task-dependent. On code the army and frontier cannot be told apart (93.7% vs 96.0%, CI crossing zero) at 1/5 the unit price; on knowledge QA the army trails by 14.9pp [βˆ’21.3, βˆ’8.7]; on math the preregistered reading gives the army +18.2pp [+12.7, +24.2], but a pre-announced audit found all 49 frontier wrong answers came from a single ambiguous template (each wrong answer exactly 4 times the gold answer), and after removal the direction reverses to βˆ’2.3pp [βˆ’4.0, βˆ’0.7]. Both sets of numbers are reported side by side. On clean data the army and self-consistency cannot be told apart (0.977 vs 0.983). This experiment gives no evidence for any contribution of "teaming" over "multi-sampling." Preregistration, the append-only ledger, and the audit script are all public. === 1 Problem and related work (skeleton) === Β· The enthusiasts. Sampling and voting gains (Li et al., TMLR 2024); open-source layered aggregation beats GPT-4o (length-controlled basis; Wang et al., ICLR 2025); multi-model debate (Du et al., ICML 2024). Β· The skeptics. Performance is non-monotonic in the number of calls (Chen et al., NeurIPS 2024); under default settings debate loses to self-consistency (Smit et al., ICML 2024); a single agent with a strong prompt nearly matches discussion (Wang et al., ACL 2024). Β· Where this paper stands (the narrowed contribution). It does not claim to fill a "cost-alignment gap" (comparisons of that kind already exist on the skeptics' side). The contribution is drawing the task-dependence boundary + decision-grade real accounting. Dollars per query as the main basis, an equal-budget self-consistency control arm, three task families, preregistered criteria, and a fully public ledger. === 2 Method === 2.1 Hypothesis and criteria (preregistration precedes results entering the repo; after that every change goes only into CHANGES.md) Β· H: On a cost-matched basis, the accuracy gap between the open-source small-model army and a single frontier model on the chosen task families is ≀ Ξ΅ = 2pp. Β· Reading: paired bootstrap, 10,000 resamples (seed 0). A 95% CI falling entirely inside [βˆ’2, +2] is a tie; entirely above is army_ahead; entirely below is army_behind; anything else is inconclusive. Β· Falsification conditions: (a) the army trails by >5pp on all three families, or (b) it catches up only at >2Γ— frontier's unit price. 2.2 Task families | Family | Source | Slice | Scoring | |---|---|---|---| | math | GSM-Symbolic (contamination-resistant perturbed variants) | first 150 problems | exact numeric match | | mmlu_pro | MMLU-Pro test | 150 problems stratified by category (seed 0) | exact option match | | code | HumanEval+ | first 100 problems | test execution passes | Excluded by design: every open-generation task scored by an LLM judge (judge preference would contaminate the measurement). 2.3 Arms and models (list price in $/1M tokens, input/output) | Arm | Configuration | |---|---| | frontier | GPT-5.6-terra single shot ($2.50/$15; a reasoning model, provider default decoding, see the change log) | | army_vote | k=5 majority vote, rotating qwen3.5-9b ($0.10/$0.15) / ministral-14b-2512 ($0.20/$0.20) / gpt-oss-20b ($0.030/$0.13), 3 seeds | | self_consistency | the steelman arm: one model, k=5 majority vote; the member picked by the preregistered rule as the strongest on a 20-problem probe = gpt-oss-20b (probe 0.867), 3 seeds | | army_debate | 3 agents Γ— 2 rounds, majority vote on the final answer, a 100-problem sub-slice per family, 1 seed | | army_division | planβ†’solveβ†’check role chain, a 100-problem sub-slice per family, 1 seed | Army size tier: total parameters per model ≀21B (9B dense / 14B dense / 21B-MoE-3.6B active). 2.4 Cost and execution Dollars per query (API list price), an append-only ledger recording every call, budget hard cap $35. 3,580 result lines in full, actual total cost $5.58, 0 failures. === 3 Results === 3.1 Main table (accuracy and $/query) | Family | frontier | army_vote | self_consistency | army_debate | army_division | |---|---|---|---|---|---| | code | 0.960 ($0.00269) | 0.937 ($0.00053) | 0.947 ($0.00041) | 0.940 ($0.00083) | 0.940 ($0.00108) | | math | 0.673 ($0.00195) | 0.856 ($0.00199) | 0.733 ($0.00041) | 1.000† ($0.00152) | 1.000† ($0.00103) | | mmlu_pro | 0.793 ($0.00498) | 0.644 ($0.00223) | 0.616 ($0.00083) | 0.740 ($0.00228) | 0.630 ($0.00125) | † The 100-problem sub-slice for debate/division happens to exclude the ambiguous-template stretch (math-100..149, see 3.3), the stretch that holds all of frontier's wrong answers; both arms and frontier score full marks on this sub-slice, Ξ”=0; the sub-slice is also hit by the zero-variance degenerate-interval artifact, see charge ⑨ in Chapter 10. 3.2 Preregistered reading (army arm βˆ’ frontier, 95% CI) | Family | army_vote | self_consistency | |---|---|---| | code | βˆ’2.3pp [βˆ’6.0, +1.7] inconclusive | βˆ’1.3pp [βˆ’3.7, +1.0] inconclusive | | math | +18.2pp [+12.7, +24.2] army_ahead | +6.0pp [+3.1, +9.1] army_ahead | | mmlu_pro | βˆ’14.9pp [βˆ’21.3, βˆ’8.7] army_behind | βˆ’17.8pp [βˆ’24.9, βˆ’10.9] army_behind | (debate/division: code βˆ’2.0 / βˆ’2.0, both inconclusive; math both Ξ”=0 tie, see †; mmlu βˆ’8.0 inconclusive / βˆ’19.0 army_behind.) 3.3 Post-hoc audit (post-hoc, pre-announced; scripts/audit_math.py, the post-hoc audit entry in CHANGES.md) Β· All 49 frontier math wrong answers fall in math-100..149 (the 50 variants of a single GSM-Symbolic probability template). Each wrong answer is exactly 4 times gold and carries a %. The mechanism: the problem text "how much more likely (as a percentage)" has two readings, gold takes the absolute percentage-point difference, frontier always answers the relative multiple; the template's base probability is always 1/4, so the relative reading always equals 4Γ—gold. Ruling: an ambiguous problem, not a math error. On the ambiguous template, army_vote 0.61 / self_consistency 0.23 / frontier 0.02. The army's "overtake" is three lineages' distribution of readings happening to land on gold's reading, not a reasoning advantage. Β· Remove that template: frontier 1.000, army_vote 0.977, self_consistency 0.983; army_vote βˆ’ frontier = βˆ’2.3pp [βˆ’4.0, βˆ’0.7], direction reversed. Β· Data composition: the 150 math problems are really only ~3 template families (50 sequential variants each, sampled without shuffling), the effective sample size is far below 150, and the preregistered CI is overconfident for this family. Β· Reporting discipline: the preregistered numbers are reported as is (3.2), and this section breaks them out side by side, labeled post-hoc, replacing nothing. 3.4 Criteria reconciliation Falsification condition (a) not triggered (code's CI crosses zero); (b) not triggered (army unit price: on code 1/5 of frontier's, on mmlu less than half, on math about equal, and no family bought a tie with more money). H is not falsified and does not hold across the board. The landing point is task dependence. code has a shot, knowledge QA has no shot, the math evidence is void pending a retest. === 4 Limitations (written for real) === 1. Template concentration in the math slice: ~3 template families were treated as 150 independent samples, the bootstrap independence assumption fails, and the preregistered CI is overconfident for this family; a retest needs a template-shuffled problem set with enough families. 2. The math direction reversal comes from a post-hoc breakdown: βˆ’2.3pp is a post-hoc number and may be made official only after a preregistered retest next round; this round draws no conclusion on math. 3. Code's preregistered reading is inconclusive, not tie: 100 problems cannot squeeze out a Β±2pp equivalence band; "cannot be told apart" is a decision-grade statement, not a statistical equivalence ruling. 4. A single cost basis, API list price only. The self-hosted amortization basis promised in the preregistration (computed analytically from GPU rental prices) was not executed and is recorded as outstanding; the two accounts may give different conclusions. 5. Frontier is one model at one point in time, and it is the balanced tier of the GPT-5.6 family (terra), not the flagship tier (Sol); the conclusion is limited to that tier. 6. On clean data the army β‰ˆ self-consistency (0.977 vs 0.983; on code SC is cheaper): this experiment can give no evidence at all for the contribution of the "teaming" mechanism itself. 7. The contamination-resistant design covers only the math family (perturbed variants); HumanEval+ and MMLU-Pro may sit inside the training corpora of the models evaluated. === 5 Reproducibility === Repository smol-army (this book's case library): preregistration precedes results entering the repo (docs/prereg.md); changes are append-only (docs/CHANGES.md); a per-call ledger (results/ledger.jsonl, append-only); per-problem results (results/runs.jsonl, 3,580 lines); the summary (results/report.md); the audit script (scripts/audit_math.py, the post-hoc audit entry in CHANGES.md). Every number here points to those files and can be recomputed. ``` ## Deliverable two, the CTO one-page memo (real numbers) ```text To: CTO From: [author] Date: 2026-07-26 Subject: replacing the frontier API with open-source small models (a decision recommendation from one $5.58 controlled evaluation) Conclusion (one sentence) On code generation, the accuracy gap between a voting combination of open-source small models (total parameters ≀21B) and GPT-5.6-terra cannot be told apart within measurement precision (93.7% against 96.0%, confidence interval crossing zero), while the unit cost is one fifth of the latter's ($0.00053 against $0.00269 per query). Risks (three lines, read them before the table) 1 The table below is a public benchmark, not our workload. The code conclusion has not been validated on real traffic. 2 The evaluation shows the gain from "multi-model teaming" cannot be distinguished from "multi-sampling one model" (on code the latter scores 94.7% at $0.00041, and is cheaper still). The recommended deployment shape is therefore multi-sampling on one model, not multi-model orchestration; discount any proposal sold on "army" or "teaming" accordingly. 3 Cost is counted at API list price; the amortization account for self-hosted GPUs was not measured. Self-hosting needs its own evaluation. Numbers (all recomputable in the repo) | Task type | Small models vs frontier | Cost ratio | Judgment | |---|---|---|---| | Code generation (HumanEval+ 100) | 93.7% vs 96.0%, CI crossing zero | 1/5 | Has a shot, take it to pilot | | Knowledge QA (MMLU-Pro 150) | trails by 14.9pp [βˆ’21.3, βˆ’8.7] | ~45% | No shot, stay on frontier | | Math word problems (GSM-Symbolic 150) | preregistered +18pp; βˆ’2.3pp after the audit found a problem-set defect, direction reversed | β‰ˆ1:1 | Evidence void, do not cite | Recommended next step Run a two-week shadow-traffic pilot on code workloads, taking the cheapest configuration in the evaluation that ties with frontier: gpt-oss-20b, 5-sample majority vote ($0.00041 per query). Knowledge QA stays as is. For math, wait for the next round after the problem set is reissued, and do not use this round as a basis. Basis The smol-army repository: preregistration (docs/prereg.md, before results entered the repo), an append-only ledger, the audit script (scripts/audit_math.py). Every number on this page can be recomputed; the whole evaluation cost $5.58, and the $35 budget hard cap was never touched. ``` **How the two pieces interlock**. Every number line in the memo (93.7/96.0, $0.00053/$0.00269, βˆ’14.9pp, +18β†’βˆ’2.3, $5.58, $0.00041) shares a source with sections 3.1/3.2/3.3 of the skeleton, both pointing at results/report.md and the audit script. Different detail, same body of fact. --- # Part Two Β· Fillable Templates ## Template 1 Β· Claims List (single source of truth) **How to use.** Step 1 of delivery, before any document. Write every claim in **the strongest form you dare sign**. So weak that signing costs nothing is cowardice, so strong that it crosses the line is drift. Every sentence in both vehicles is generated from this table. The wording can change, and no downstream document may edit this table backward. "Not done" is a line too (a basis promised but not executed, an arm that got cut). ```text # Claims list: ____________ (project) Date: ________ Signed: ________ Master source of evidence (repo / results file / commit): ____________ | # | Claim (one sentence, the strongest form you dare sign) | Evidence pointer (file/table/commit) | Tier (verified / still exploring / falsified) | Signature (dare / do not dare) | |---|---|---|---|---| | 1 | ______________________ | ____________ | ________ | ____ | | 2 | ______________________ | ____________ | ________ | ____ | | 3 | ______________________ | ____________ | ________ | ____ | | Outstanding | Promised but not executed: ________ | Where promised: ____ | Outstanding | I dare sign "not done" | ``` **Self-check**: - [ ] Is there a claim with no pointer? A claim without a pointer does not enter the list. Go back and get the evidence, or downgrade it to "still exploring." - [ ] Are a preregistered number and a post-hoc breakdown crammed into one row? Split them into two, and the post-hoc entry carries its own post-hoc label (the discipline of Template 2 in Chapter 8 takes effect upstream here). - [ ] Is the "outstanding" row empty? Check every basis and every arm you signed off in the plan file. Only a fully delivered plan is allowed an empty row. - [ ] Is every claim written in its weakest form? You are wasting evidence bought with real money. Push each one up a tier, until one more notch would stop you from signing. ## Template 2 Β· Technical-Report Skeleton (shaped like a workshop paper, fillable) **How to use.** Audience = peers and the technical committee, and their question is "how do you know." Generate it from Template 1. Give method and statistics in full, and write limitations for real (each item carrying numbers, no camouflage wording). Skeleton first. Get the five parts standing on their own, then expand into prose. ```text Title: put the main conclusion in the title (with its qualifier, such as "task-dependent"): ____________ Abstract (six sentences): β‘  One sentence on the problem and the dispute: ____________ β‘‘ One sentence on method and preregistration (criteria before results): ____________ β‘’ One sentence on the main result (preregistered numbers as is): ____________ β‘£ One sentence on any overturn or breakdown (if any, labeled post-hoc): ____________ β‘€ One sentence on the mechanism control (what the control arm said): ____________ β‘₯ One sentence on openness (repo/preregistration/ledger public): ____________ 1 Problem and related work: 2-3 papers from each camp + where this paper's contribution stands (narrowed to what you dare sign) 2 Method: hypothesis and criteria (with falsification conditions) / tasks and data / arm design / cost basis / statistics 3 Results: main table β†’ preregistered reading β†’ post-hoc breakdown (its own subsection, labeled post-hoc) β†’ criteria reconciliation 4 Limitations (written for real, each item carrying numbers): - Data composition problems: ____________ - Stating the identity of the post-hoc analyses: ____________ - Weaknesses in reading and statistics: ____________ - Bases promised and not delivered (outstanding): ____________ - External validity boundaries (model/point in time/contamination): ____________ 5 Reproducibility: repo / preregistration commit / ledger / audit script ``` **Self-check**: - [ ] Does the title dare say more than the abstract? The title is the sentence quoted alone most often, so run the strictest signature test on it. - [ ] Is a weakness admitted in limitations still used as a selling point in the abstract or the conclusion? Downgrade consistently across the whole document. Limitations is not a disclaimer. - [ ] Did post-hoc numbers mix into the preregistered subsection? Separate sections, label them, report side by side, replace nothing. - [ ] Did the "outstanding" entry disappear? Template 1's outstanding row must have a matching entry in limitations. ## Template 3 Β· One-Page Memo (fillable) **How to use.** Audience = decision makers, and their question is "what should I do." Three hard constraints. One page, conclusion on top, risks right behind the conclusion (not at the foot of the page, nobody reads the foot). Every number shares its source with Template 2. ```text To: ________ From: ________ Date: ________ Subject: ____________ (state the decision question, not the project name) Conclusion (one sentence, verbatim from the strongest signed claim in Template 1): ____________________________________________ Risks (three lines, each one caveat that could change the decision): 1 Extrapolation boundary: ____________ (what was measured, what was not) 2 Mechanism caveat: ____________ (what the control arm or the audit said that hurts) 3 Basis caveat: ____________ (known gaps in the cost or data basis) Numbers (a table of ≀5 rows, every cell recomputable): | Scenario | Reading | Cost | Judgment | |---|---|---|---| | ____ | ____ | ____ | Has a shot, take it to pilot / No shot / Evidence void | Recommended next step (one, executable, with configuration and budget): ____________________________________________ Basis (one line): repo ________, preregistration ________, total cost ________ ``` **Self-check**: - [ ] Does the one-sentence conclusion contain "may" or "to some extent"? Either rewrite it down to a strength you dare sign, or admit the evidence is not enough and do not deliver. - [ ] Are the risk lines generic boilerplate ("limited sample size")? Every line must be specific enough to change a decision. A risk line that cannot change a decision is decoration. - [ ] Is the judgment column all "has a shot"? Go back to Template 1 and check. Did the unfavorable conclusions get delivered too? "No shot" and "evidence void" are the two most money-saving words in a memo. - [ ] Have the numbers been checked against the technical-report skeleton? Run the prompt set's interlock check before you send it. ## Prompt set, audience rewriting + the number interlock check **Audience rewriting prompt** (drafting and rewriting handed to AI, strength locked): ```text Below is my claims list (each row carries the claim, the evidence pointer, the honesty tier, and the signature status): [paste Template 1] Rewrite it as [technical-report skeleton / one-page memo] for an audience of [technical committee / peer review / CTO / ____], whose core question is ["how do you know" / "what should I do"]. Hard constraints: 1 Do not change the strength of any claim. "Undecided" may not become "matched," "looks like" may not become "shows," and post-hoc labels may not be dropped; 2 Do not merge two claims into one stronger sentence; 3 Keep the evidence pointer after every number (turn them into citations or footnotes in the final draft); 4 Not one claim outside the list may appear. ``` **Number interlock check prompt** (independent channel, mechanical work): ```text Here are two documents and one results file: [technical-report skeleton] [memo] [results file/table]. Do factual checking only, and do not evaluate the conclusions: 1 Extract every number appearing in the two documents (body text, tables, titles included) into a list of the number, where it appears, and what it claims to mean in context; 2 For each number, look for its counterpart in the results file and mark it: matches / does not match (list both values) / not found in the results file; 3 Cross-compare the numbers for the same fact between the two documents and list every disagreement; 4 List every comparative or superlative in the documents with no number behind it ("faster," "strongest," "substantially"). Forbidden: judging whether a disagreement matters. That is not your job. ``` **Self-check (prompt set)**: - [ ] Did the rewriting prompt omit the tier and the signature status? With no strength information, AI invents strength of its own, and always upward. - [ ] Did the interlock check run in the same session that wrote the draft? Switch to an independent channel. The same session protects its own draft. - [ ] Did you fill in a number by hand where the check report said "source not found"? No number may be typed by hand. Go back to the results file and compute it, or delete the sentence. - [ ] Did you skip rerunning the check after a round of polishing? After any operation that touches the text, the check is void. Rerun it. --- ## Post-red-team revision record (after Chapter 10 opened court) The two finished pieces in Part One are the versions as of Chapter 9, and keeping them unchanged is deliberate. The Chapter 10 red team put them on trial (four attack surfaces Γ— independent sessions), and all ten charges held. The revisions: 1. All "equal-budget self-consistency" statements withdrawn. The preregistered Β±10% cost-alignment clause was never executed (k stayed at 5), SC actually paid only 0.21-0.77 times the army, and the honest restatement is "SC spent less money and still tied or did better"; 2. The mmlu_pro conclusion narrowed to knowledge recall in three subjects, business/law/psychology (the sample was drawn in blocks, with zero STEM coverage); 3. The "army vote" mechanism on code degenerated into a single sample from a single model (exact string tallying, ties resolved to the first sample). 93.7% is really qwen's single-shot score, and the conclusion is restated accordingly; 4. The βˆ’2.3pp on clean math withdrawn (scoring residue killed the small models in one direction only, and under lenient parsing all three arms saturate). The math family retired in both directions; 5. All cost numbers relabeled "list-price basis" (not what the bill charged, the army actually paid about 19% more), with a procurement sensitivity caveat added; 6. The memo's "recommended next step" of "gpt-oss-20b, 5-sample majority vote" is revised in step with item 3 above. Exact string tallying degenerates to a single sample on code (4 of the 5 calls wasted), so the recommendation is restated as "start with single shot (cheaper), or fix the tallying mechanism and re-evaluate k=5." The full disposition is in the smol-army repo at docs/redteam-2026-07-25.md. By Chapter 9's discipline, a revision is not an erasure. Both versions stand, and the difference is the lesson. --- # Chapter 10 templates Β· Red-Team Dispatch Brief Template + Four-Attack-Surface Red-Team Prompt Set + Disposition Table > This appendix is the complete fillable version of the three tools in Chapter 10. > The order of use is fixed. Once the deliverable is final, use Template 1 to compress it into a claims list and write the dispatch brief. Open an **independent session** for each of the four attack surfaces and run the matching prompt from Template 2. Once the list of charges comes back, rule on it line by line, then use Template 3 to assign a tier, dispose, and file. "The red team is mandatory" in the L2 verification of Chapter 12 means exactly these three steps. --- ## Template 1 Β· Red-Team Dispatch Brief (fillable) **How to use.** The only purpose of the brief is to get the red team all the materials without letting it know what you expect. An independent session, zero history, and switching models is better still. Letting the model that paired with you to write the conclusion serve as its red team violates the generation/verification orthogonality principle. Run the leak check at the end of this section once, then dispatch. ```text # Red-team dispatch brief Deliverable: ____________ (paper / memo / report name; no body text, see the claims list) Date: ____________ Red-team session and model (different from the writing channel): ____________ ## Materials (paths and originals only, not your summary or narrative) - Raw results file: ____________ - Criteria / preregistration file (with timestamp): ____________ - Scoring / analysis script: ____________ - Cost ledger (if there are cost claims): ____________ - Data / problem text file: ____________ ## Claims list (each line = one attackable factual claim in the deliverable, neutrally worded) A1 ____________________________________________ A2 ____________________________________________ A3 ____________________________________________ A4 ____________________________________________ (5-10 lines is about right; a conclusion and its qualifying conditions get one line each; write number claims to recomputable precision, such as "Ξ” = βˆ’2.3pp, CI [βˆ’4.0, βˆ’0.7]") ## Work order Attack surface of this order: ____________ (the scorer / data composition / independence assumptions / cost basis, one of four) Run the prompt for the matching attack surface in Template 2. ``` **Self-check (leak check, must pass before dispatch)**: - [ ] Does the brief say "I hope / I am worried / confirm this for me / our contribution"? Delete all of it. A red team that reads your expectations aims off target to match them. - [ ] Do the claims carry adjectives ("significantly," "robustly," "surprisingly")? Rewrite them as neutral statements. - [ ] Did you merge the four attack surfaces into one order? Split them. Dispatched merged, the model fills the page with the softest surface. - [ ] Did you dispatch from the session or the model that wrote the conclusion? Switch. Context is a position. - [ ] Did you hand over your summary instead of the original files? Give the originals. A red team cannot attack what your summary left out, and that is exactly where the hits land. --- ## Template 2 Β· Four-Attack-Surface Red-Team Prompt Set **How to use.** The four surfaces share one generic skeleton, and only the [attack surface definition] block gets swapped. One surface, one order, an independent session. The output must be the set of three, charge + consequence + executable check. A charge without a check is a comment, and it is not accepted. **The generic skeleton**: ```text Role: you are hired to attack the claims list below, on the attack surface: [attack surface definition, paste in one of the four blocks below] Input: the claims list and the materials (see the dispatch brief). Output: a list of charges, ordered by lethality, and every charge must carry three things: 1 Charge: which point of this attack surface could make one of the claims fail (name the claim number); 2 Consequence: if the charge holds, which claim dies, and what direction and rough magnitude of effect gets manufactured; 3 Check: one check executable the same day (a script sketch / a sampling plan / a replay experiment / a recompute on another basis), whose result can confirm or rule out this charge. Rules: attack only, no balanced coverage; do not evaluate "whether the conclusion as a whole is credible"; do not list a charge you cannot give an executable check for; if you find nothing, state "nothing found on this surface," do not pad. ``` **Surface 1, the scorer / the criteria**: ```text Attack surface definition: how right and wrong, success and failure, get judged (gold answers, scoring scripts, human annotation rules, an LLM judge, the operational definition of "valid / successful"). Check first: β‘  whether the definition of "right" has a second defensible reading (ambiguous gold, ambiguous problem text); β‘‘ the links in the parsing β†’ matching β†’ scoring chain that silently swallow points or hand them out (format suffixes, units, null, timeouts, partial matches); β‘’ tidy patterns in the wrong answers: always k times, always off by a constant, always in one format, clustered in a continuous id stretch (tidiness is the fingerprint of a sick scorer or sick gold); β‘£ whether the conclusion flips under an equally reasonable alternative scorer / annotator. ``` **Surface 2, data composition**: ```text Attack surface definition: the samples, the problems, the corpus itself (source, draw, structure, coverage, contamination). Check first: β‘  the structure brought in by the draw (first N rows / a single batch / a single source / a convenience sample); β‘‘ how many "molds" the samples really have: after stripping proper nouns and numbers, how many literal templates are left, and what share each holds; β‘’ whether the test material could sit inside the training corpus of the system under test (public question banks, famous datasets, material from before the corpus cutoff date); β‘£ whether the data coverage holds up the wording range of the claim, whether "holds on the X slice" got written as "holds." ``` **Surface 3, independence assumptions**: ```text Attack surface definition: the independence behind the sample size and the statistical inference (the quality of n). Check first: β‘  the gap between independent units and rows: same template / same person / same batch / repeated measurement, how much each shrinks n; β‘‘ whether the data structure breaks the independence assumption of the CI, the significance test, the bootstrap; how much the interval moves after recomputing by cluster (template family / batch / respondent); β‘’ whether the errors of the comparison arms are correlated: shared problem set, shared scorer, shared time window (an anomaly moving the same way in a control arm signals a sick shared path); β‘£ stopping rules and multiple comparisons: was n fixed in advance, or "run until it looks good." ``` **Surface 4, cost basis**: ```text Attack surface definition: the accounting behind every claim of the "cheaper / faster / better value / cost-matched" kind. Check first: β‘  who chose the basis, when it was chosen (before the results or after), and whether it happens to favor the author's own design; β‘‘ whether the conclusion flips under a reasonable alternative basis, another pricing scheme, one carrying hidden costs (retries, failures, labor, development time), one carrying amortization; β‘’ whether the two sides of the comparison got equally fair prices and configurations: tier, discount, wholesale price, point in time (list prices move, did the claim mark its point in time); β‘£ whether estimated numbers (amortization, extrapolation, conversion) have measured backing; and where they do not, whether a caveat hangs beside the claim. ``` **Self-check (run through it when the list of charges comes back)**: - [ ] The charges are not ordered by lethality, or there are dozens at once? Ask for them merged and reordered. A long list is a dilution tactic, and a report whose top three do not hurt gets dispatched again in full. - [ ] A charge with no executable check? Send it back. No debate, send it back. - [ ] "Nothing found on this surface" shows up? Dispatch this surface once more with another model, and it counts only when both rounds come back empty. - [ ] Did you dispatch only the surfaces you feel safe about? Run all four, the one you feel shakiest about first. - [ ] Does the list of charges seem to be praising you ("the overall design is rigorous, only minor issues")? Expectations leaked. Go back to Template 1 and check the brief. --- ## Template 3 Β· Disposition Table and Disposition Record **How to use.** Every charge ruled "holds" must get a tier and a disposition, one of three, written into the disposition record. There is no fourth. "I am aware of it" is not a disposition. The qualification check runs in a fixed order. First ask whether it can be fixed (fix it), then whether it can be tested (test it), then whether it is fatal (overturn or narrow). Only when all three fail does the caveat get its turn. **The tier table**: | Tier | Trigger condition | Disposition action | The unqualified form | The specimen in this book | |---|---|---|---|---| | Overturn | The charge hits load-bearing structure and no rewording saves it | The announcement sentence comes off the deliverable; the preregistered numbers still get reported as is, with the post-hoc breakdown side by side (the Chapter 8 discipline). What is overturned is the sentence, not the number | Quietly deleting the number and never mentioning that this conclusion existed | math +18.2pp: the ambiguous-gold charge held, and after removal the βˆ’2.3pp reversed direction, so "the army pulls ahead on math" came off (that βˆ’2.3 was itself later overturned by the red team, a scoring-residue artifact, see 10.7; but the disposition of overturning the +18 announcement sentence still stands) | | Narrow | The charge holds, but what it cuts away is the range or the strength of the evidence, not the conclusion itself | Rewrite the boundary of the claim (task family / sample / confidence / exact model), and the new statement is **strictly weaker** than the old one and still checkable | Narrowing into vaguer words ("may under some circumstances"), which is escape, not narrowing | math CI: ~3 template families β†’ "this problem set cannot detect a conclusion, reissue it and test again"; "ties frontier" β†’ "ties GPT-5.6-terra (the balanced tier)" | | Caveat | The charge holds or cannot be ruled out, and it is unfixable, untestable, not fatal | The conclusion stays, and the caveat travels at the **same address** as it; also leave one check for the next round of experiments | The caveat buried deep in an appendix or a footnote; a caveat hung where an overturn belongs | The subplot's single-source ground truth: "WVS may be inside the training corpus" travels with every subplot conclusion, the FAIL included | **Disposition record (fillable, filed with the deliverable)**: ```text # Red-team disposition record Deliverable: ____________ Date: ____________ Red-team channel (model / session, must differ from the writing channel): ____________ | # | Attack surface | Charge (one sentence) | Check and result | Verdict | Tier | Disposition action (which sentence changed / what caveat hangs) | |---|---|---|---|---|---|---| | 1 | ____ | ____________ | ________ | holds/rejected | overturn/narrow/caveat | ____________ | | 2 | ____ | ____________ | ________ | ________ | ________ | ____________ | Rejected charges (one line of reason each, kept as ammunition at the defense): - ____________________________________________ Does the revised deliverable get one more quick round: ______ (a patch can introduce a new handle) Where this record is filed (same repo as the deliverable): ____________ ``` **Self-check**: - [ ] A charge ruled "holds" with no tier? There is no fourth disposition. - [ ] Far more caveats than overturns plus narrowings? Check whether you are using the caveat as a trash can, and rerun the three qualification questions line by line. - [ ] Is the new statement after narrowing vaguer than the old one instead of weaker? "Weaker" means still checkable, with a clearer boundary. Vaguer runs the other way. - [ ] Did the disposition change only the abstract and leave the matching paragraph in the body untouched? However many times the same conclusion appears in the deliverable, the disposition lands that many times. - [ ] Not planning to publish the disposition record with the deliverable? It is part of the trust mechanism. A report carrying bullet holes and dispositions is more credible than one that looks untouched. --- # Chapter 11 templates Β· Failure-Mode Census Sheet + AI Self-Check Prompt Set > This appendix is the complete fillable version of the two tools in Chapter 11. > The order of use is fixed. Fill in Template 1 (the census sheet) first, then pick the matching self-check prompt from Template 2 according to what the census turned up. > **The most important sentence in this appendix sits in the use warning of Template 2. AI self-check misses errors of the motivated collusion class. Read it first, then use the prompts.** --- ## Template 1 Β· Failure-Mode Census Sheet **How to use.** Run a thirty-minute census on your project (the full flow is Chapter 11, section 11.7). Start with the quick reference, there is no background knowledge to memorize. The quick-reference block prints the signature of each of the four failure modes right under the header, and all you do is translate them into concrete signals in your project. Refill the census sheet at every milestone (before plan sign-off, after the first results, before delivery), and keep the old sheets on file for comparison. ### Quick reference, the four failure modes (a condensation of Chapter 11, section 11.4) | Failure mode | Chief attribute | High-incidence step | General signature | |---|---|---|---| | Hallucination and fabricated citations | Fluency Γ— corpus prior | Master the field, deliver | The most on-point citation is the most suspect; the retelling carries fewer qualifiers than the original | | Spurious significance and the criteria backdoor | Sycophancy Γ— fluency | Test plan, interpretation | The criterion appears after the result; the conclusion "just happens" to clear the line | | Data leakage and contamination | Corpus prior | Execution | The score is too good to be true; it collapses on fresh questions from the same distribution | | Sycophancy drift | Sycophancy | Interpretation, red team | Ask the same question twice with opposite leans and the conclusion flips | One more cross-cutting error does not pick a step, **motivated collusion**, where the conclusion you are most excited about = the conclusion you checked least. It is not in the matrix. It gets a column of its own below the matrix, and AI self-check cannot catch it (see the warning in Template 2). ### The census matrix (fillable) Walk your project's seven steps row by row. Steps where AI is lightly involved can be filled in short. The three steps where AI is most deeply involved must be filled in completely. ```text # Failure-mode census sheet Project: ____________ Filled in by: ____________ Date: ____________ Project milestone this census belongs to (before plan sign-off / after the first results / before delivery): ____________ ## Step 1 Β· Master the field AI involvement (none/light/deep): ______ Attribute mainly consumed: ______ Error type most likely to grow here: ____________________ Concrete signature in this project (translate it into a concrete signal, such as "which citations are the most on-point and therefore checked first"): ____________________________________________ Cost of being wrong (high/medium/low): ______ ## Step 2 Β· Questions and hypotheses AI involvement: ______ Attribute mainly consumed: ______ Error type most likely to grow here: ____________________ Concrete signature in this project: ____________________ Cost of being wrong: ______ ## Step 3 Β· Test plan AI involvement: ______ Attribute mainly consumed: ______ Error type most likely to grow here: ____________________ Concrete signature in this project (such as "which criterion's wording can still be explained away after the fact"): ____________________________________________ Cost of being wrong: ______ ## Step 4 Β· Execution AI involvement: ______ Attribute mainly consumed: ______ Error type most likely to grow here: ____________________ Concrete signature in this project (such as "anything above __ points gets treated as leakage first"): ____________________________________________ Cost of being wrong: ______ ## Step 5 Β· Read and catch errors AI involvement: ______ Attribute mainly consumed: ______ Error type most likely to grow here: ____________________ Concrete signature in this project (such as "which question have I not yet asked a second time with the opposite lean"): ____________________________________________ Cost of being wrong: ______ ## Step 6 Β· Deliver AI involvement: ______ Attribute mainly consumed: ______ Error type most likely to grow here: ____________________ Concrete signature in this project (such as "which numbers changed hands more than once between analysis and final draft"): ____________________________________________ Cost of being wrong: ______ ## Step 7 Β· Red team AI involvement: ______ Attribute mainly consumed: ______ Error type most likely to grow here: ____________________ Concrete signature in this project: ____________________ Cost of being wrong: ______ ## Cross-cutting column, motivated collusion (mandatory, no blanks allowed) The one conclusion I am currently most excited about: ____________________________ Which step and which cell it sits in: ____________________ My actual checking intensity on it (write it honestly): ____________________ A channel that does not share my motive (a person's name / how to dispatch an independent agent): ______ ## The two circled cells Operating room (the cell where being wrong costs most): ____________________ High-incidence zone (the cell holding the conclusion I am most excited about): ____________________ ## Closing self-test (one sentence, no answer means the census fails) If this project blows up three months from now, most likely in: ______ cell; what the crime scene looks like: ____________________________ ``` **Self-check** (walk it once when the sheet is filled): - [ ] Did the "concrete signature" field copy the book's wording directly? Copying the wording is the same as leaving it blank. "The score is too good to be true" has to become a concrete number in your project. - [ ] Motivated collusion column left blank, or "the conclusion I am most excited about" filled in with something harmless? You just walked around the most valuable field in the census. Refill it honestly. - [ ] Are the high-incidence zone and the operating room the same cell? That is the most dangerous position in your project right now, and the verification budget (Chapter 12) goes there first, all of it. - [ ] Did every cell get "medium" for cost of being wrong? No ranking means no layering. Force out a top two. - [ ] Cannot answer the closing self-test? Go back to the step where AI is most deeply involved and walk the two questions again. - [ ] Is the previous version of the census sheet still around? Comparing how the sheet changes between milestones carries more information than any single sheet. --- ## Template 2 Β· The "Have AI Self-Check for Failure Modes" Prompt Set > **Use warning (read first, the bold is not decoration). AI self-check misses errors of the motivated collusion class.** > The reason is structural. The motive is in you, and the model is trained to talk along with you. Hand it your own conclusion to check and it checks the form (whether the citation exists, whether the numbers add up, whether the reasoning chain breaks). Your wish it does not check, it amplifies (the firsthand specimen in Chapter 11, section 11.5, this book's own "literature gap" judgment, which AI never questioned and which finally died on an independent forward-citation check). > So this prompt set is positioned as **a first-pass screen for formal errors**, under two iron rules. > β‘  A self-check that outputs "no problems found" never equals no problems. It equals passing the formal first-pass screen; > β‘‘ The motivated collusion class has only a process solution, that "channel that does not share your motive" in the cross-cutting column of the census sheet (another person, or a verification agent that does not know the answer you expect), and how to dispatch it is in the independent channel brief template in Chapter 12. > One more thing. Send the prompts below to another model or another clean conversation that **took no part in generating the output**. The AI that produced the conclusion cannot serve as its own judge (orthogonality, Chapter 2, section 2.4). ### Prompt 2a, citation and fact self-check (against hallucination and fabricated citations) ```text Below is the full text of a research output (or its citation list): [paste] Your task is a formal first-pass screen of every citation and every key factual claim. Do not reach for your "impression" of these papers to vouch for what is real. Your output is a check work order, not a verification verdict. 1 Extract every citation and output each one, title / authors / year / venue; 2 Tag each with an "on-point rating", which claim it carries and how much. Highest on-point first (fabricated citations are built to order, the most on-point is the most suspect); 3 Extract every retelling sentence ("one study shows..."), listing the retold wording vs the qualifiers that need the original to confirm (task scope, sample, basis); 4 Extract every sourceless number and every superlative claim ("first", "largest", "universal"); 5 Output the check work order by priority, one line each, what to check, where to check it (academic search engine / which section of the original), and what to do if it is not found. ``` After use. The work order has to be executed by a person (or an independent verification agent). What this prompt produces is a to-do list, and the list itself guarantees nothing. ### Prompt 2b, criteria and significance self-check (against spurious significance and the criteria backdoor) ```text Below is an analysis output, and (if one exists) the test plan that goes with it: [paste the full analysis] [paste the plan written beforehand; with no plan write "none", and "none" is itself the biggest finding] Answer item by item, and output only findings you can point to a location for: 1 For every conclusion in this analysis, where is the decision standard written? Mark whether it appears before the results (in the plan) or after them (in the analysis narrative); 2 List every degree of freedom that "can still be moved after the results are in", metric choice, data slicing, outlier removal, stopping time, marking for each whether the text declares a rule set in advance; 3 Find every number that "just clears the line" (just past a significance threshold, just at target); 4 How many hypotheses did this analysis test, how many metrics did it look at? If more than one, is there a multiple comparisons problem, and does the text correct for it; 5 Output which conclusions qualify as "confirmatory" and which only qualify as "exploratory". ``` After use. A conclusion downgraded to "exploratory" has to be worded as exploratory in the deliverable (the wording discipline in Chapter 9). ### Prompt 2c, leakage and contamination self-check (against data leakage and contamination) ```text Below is the design and code of an evaluation or analysis (or a description of it): [paste: data sources, how it was split, feature engineering, eval set choice, scores] Work through the leakage paths one by one, and for each output "risk point β†’ location in the text or code β†’ the check action you suggest": 1 Is the eval data public (a public benchmark, a public question bank, a dataset circulating online)? Public means treating it by default as "the model saw it during training". List the contamination-resistant variants or fresh questions available as substitutes; 2 Could test-set information leak into the training side anywhere in the code, normalization or statistics computed before the split, feature engineering over the full data, a validation set reused for tuning; 3 Where does the score sit relative to comparable work? An absurdly good score gets treated as leakage first. List the concrete way to "retest on a fresh set of questions from the same distribution"; 4 Is there time leakage in the data pipeline (using information available only after the decision point); 5 Output the list of check actions, sorted by "cheap and lethal". ``` After use. The check actions on the list have to actually run. Having AI hunt for leakage and having AI fix leakage can be the same session, but the conclusion "leakage ruled out" can only come from the numbers of a retest. ### Prompt 2d, the sycophancy flip test (against sycophancy drift) **How to use.** This one is not AI checking itself. It is a controlled experiment you run on AI. Prepare two versions of the same question, send each in **two clean conversations that cannot see each other**, and compare the conclusions. ```text Version A (positive lean): Do these results [paste] support the conclusion "____________"? Version B (negative lean): Is it possible that these results [paste] do not support "____________", and are only [noise / confounding / a selection effect]? Argue it. ``` Reading rules: - The two versions agree in substance (only wording and tone differ) β†’ the reading can go on being used for now; - The two versions are substantively opposite β†’ what you measured is not the data, it is how you asked. The reading of that question is downgraded as a whole, and it gets redone through the independent channel in Chapter 12; - Either version opens with a "you are right" and carries no reservation anywhere β†’ treat it as sycophancy, and retest with another model or another wording. ### Self-check (walk it once after running the whole prompt set) - [ ] Did you send the self-check prompt to the same session that generated the output? That violates orthogonality. Void it and rerun. - [ ] All four prompts came back green, so the output feels safe? Reread the use warning at the top of this template. Passing the formal first-pass screen β‰  no problems, and the motivated collusion class is not within this prompt set's range at all. - [ ] The "most exciting conclusion" in the cross-cutting column of the census sheet, has it been assigned to a channel that does not share your motive? This is the one move in the whole self-check AI cannot make, and the one most easily skipped. - [ ] Has the check work order from 2a been executed? A work order lying in the inbox equals no check. - [ ] Did you run only the one or two prompts relevant to your own output? That is normal, layering is supposed to work that way (Chapter 12 expands); but an output about to be delivered (one entering the decision chain) has to pass all four. --- # Chapter 12 templates Β· Layered Verification Workflow Card + Independent-Channel Check Prompt Set + Verification Budget Sheet > This appendix is the complete fillable version of the three tools in Chapter 12. > The order of use is fixed. First assign the output a layer with Template 3. Then execute the checklist for that layer in Template 1. Every check task dispatched to AI uses a prompt from Template 2. The three independent channel disciplines (channel independence / the brief leaks no expected answer / three-value output) apply to all three layers. --- ## Template 1 Β· Layered Verification Workflow Card **How to use.** The layer is decided by "cost of being wrong Γ— probability of being wrong," not by how good the output looks, and prose plays no part in it. The quantities for each layer (how many to sample, how long to budget) are starting defaults. Calibrate them to your field, write them into your own card, and do not change them on the spot after that. ### Layer quick-reference table | Layer | Trigger | Scope of check | Time budget | Required item | |---|---|---|---|---| | L0 spot check | The output does not leave your desk (brainstorming, exploratory drafts, intermediate material) | Sample | 15–30 minutes | A random sampling rule | | L1 full verification | The output enters the decision chain, someone will spend money, commit people, or draw conclusions based on it | Every citation + every number + the reasoning chain | Half a day–one day (mostly AI time) | The "claim β†’ source" table filed | | L2 adversarial recompute | A single conclusion being wrong would trigger a hard-to-reverse action (selection, funding, publication) | The named load-bearing conclusions | One to several days / per conclusion | Independent re-derivation + red team (red team method in the Chapter 10 templates) | ### L0 checklist - [ ] Random citation sample, 5 or 10% (whichever is larger). Two questions each. Does it exist? Does it really say what the output claims it says? - [ ] Sampling is decided by you or by a random number. Never let the generating side choose, and the verification channel does not choose either. - [ ] Sample 3 key numbers and trace each to its source, either to the origin or to a dead end (a dead end is a hard defect). - [ ] One reverse question (the simplified use of Template 2d): "Which claim in this material has the weakest evidence, and why?" - [ ] Escalation rule. Any hard defect found (fabricated citation / sourceless number) β†’ the whole output moves up to L1. ### L1 checklist - [ ] Extract the claim list (Template 2a) and group by type, citation / number / reasoning. - [ ] Forward-check every citation (Template 2b) with three questions. Does it exist? Does it say so? Was it later overturned or retracted? - [ ] Trace every number (Template 2c) back to its original source and check that the basis matches (comparison baseline, time window, units). - [ ] Walk the reasoning chain link by link. Label each link's type (citation / calculation / "the author thinks"). List the "author thinks" links separately and hand them to a human ruling. - [ ] When collating, a human goes through only two columns, every "undecidable" + every "falsified." Undecidable β‰  pass. Nothing turns green quietly. - [ ] File it. The "claim β†’ source" table is stored with the output, with the check date and channel noted. ### L2 checklist (for each named load-bearing conclusion) - [ ] Independent re-derivation (Template 2e). Another channel gets only the raw materials and the question, not the conclusion, and derives from scratch. - [ ] Compare. Converges β†’ record as machine evidence. Does not converge β†’ rule on each point of disagreement by hand, and each one either fixes the output or goes into the limitations. - [ ] Recompute key numbers by another method, a different calculation path or a different data source. - [ ] Red-team this conclusion (required; full method and prompts in Chapter 10 and its templates). - [ ] File every ruling. ### Self-check (run once before and once after executing any layer) - [ ] Was the layer assigned by "destination + the most likely kind of error," or by "how reliable it looks"? The latter is grading by prose. - [ ] Was the sampling rule locked in first? Picking "the few that look suspicious" on the spot is not a spot check, it is a hunch. - [ ] Did any check run through the channel that generated the output? If so it is void. Re-dispatch. - [ ] Is the "undecidable" column empty? Either this output is unusually clean, or your verification channel is fudging. Check two items yourself. - [ ] Did output that failed verification take the exception channel of "the author explained and it was let through"? What fails the mechanism does not merge, no exceptions. - [ ] Was the check record filed? A check with no archive, three months later, equals a check never done. --- ## Template 2 Β· Independent-Channel Check Prompt Set **How to use.** The five prompts share three disciplines. β‘  The verification channel is separate from the generation channel (a different session / a different model / a person), sharing none of the context from generation; β‘‘ the brief gives only the claim, not its origin or the expectation; β‘’ output is always three-valued, confirmed / falsified / undecidable. Each prompt is a skeleton you can rewrite directly. Fill in your content at the [square brackets]. ### 2a Claim extraction prompt (The purpose of this step is to break the output into a list of independently checkable claims. It can be dispatched to any channel, but before the list goes to the verification channel you must strip the concluding tone yourself, see the Self-check.) ```text Below is a document. Extract every "checkable claim" in it as a list, one per line, labeled by type: - Citation: claims a source exists and says something - Number: gives a specific number and its meaning - Reasoning: a conclusion derived from the claims above Requirements: 1 Rewrite each claim in neutral wording, stripping the original's rhetoric, emphasis, and concluding tone; 2 Note where the claim sits in the original (section / paragraph); 3 Do not judge whether a claim is true. Extract only. Document: [paste the full output] ``` ### 2b Citation check brief (the core template that leaks no expected answer) ```text You are the verifier. Below is a set of claims. Check each one independently. I will not tell you which document they came from, and I will not tell you which ones I hope hold. For each claim output: 1 Verdict (pick one of three): confirmed / falsified / undecidable 2 Basis: the source you actually found (it must open, or point to a specific location in a specific paper) 3 For citation claims, answer three questions: does the source exist? Does it really say what the claim says it says (check the original text, and watch for dropped qualifiers)? Was it later overturned, corrected, or retracted (check forward citations)? 4 If "falsified" or "undecidable": where exactly the gap between the claim and the evidence lies Forbidden: guessing the "expected answer" from the wording or ordering of the claims; Forbidden: leaning phrasing on any "undecidable" item; Forbidden: middle-state phrasings such as "basically correct" or "broadly credible." Claim list: 1 [claim one] 2 [claim two] ``` ### 2c Number-tracing brief ```text You are the verifier. Below is a set of numerical claims. Trace each one to its source. I will not tell you which document these numbers came from, and I will not tell you which number matters to the conclusion. For each number output: 1 The most original source you can reach (paper table / data file / official statistics), with the specific location 2 Basis check: the number's comparison baseline, time window, and units in the source, do they match how the claim uses it? 3 Verdict (pick one of three): confirmed / falsified (including "the number is right but the basis was swapped") / undecidable (traced to a dead end) Numerical claim list: 1 [the number and its claimed meaning] 2 [...] ``` ### 2d Reverse-question prompt (the lightweight version for L0) ```text Below is a piece of research material. Your task is not to summarize it but to find its soft spots: 1 Which three claims in this material have the weakest evidence? Why? 2 Which number would you most like to see the source for? 3 If you could check only one claim, which one? No need to verify, only to point. No pleasantries, and no praising before criticizing. Material: [paste the output] ``` ### 2e Independent re-derivation brief (for L2; the most leak-prone of the set, self-check word by word) ```text Below is a batch of raw materials and one research question. Based on these materials and only these materials, derive your own conclusion independently. Research question: [the question itself, with no leaning wording] Requirements: 1 Give your conclusion, and the complete reasoning chain from the materials to it; 2 For each link, note which part of which material it depends on; 3 Where the materials cannot support a judgment, say "insufficient material" explicitly, and do not fill in with common knowledge; 4 At the end, list separately: which two or three premises your conclusion depends on most, and how the conclusion changes if they are wrong. Raw materials: [raw materials only. No part of the original output, including its subheadings, figure captions, and paragraph structure] ``` ### Self-check (run before every dispatch, focused on leak checks) - [ ] Does the brief contain words like "confirm," "verify our findings," "support"? Already leaked. Rewrite. - [ ] Did you paste the full original output or a fragment to the verification channel (2b/2c/2e scenarios)? Context is a position. Re-dispatch. - [ ] Does the ordering or wording of the claim list hint at which items are "important"? Shuffle the order and rewrite neutrally. - [ ] Did the check and the generation use the same session? Void. Convenience is not a reason. - [ ] Did the channel take it upon itself to turn three-value output into "agree / disagree" or a score? Send it back for a three-value redo. - [ ] Did the re-derivation materials for 2e pick up structural residue of the original output (subheadings, figure captions, wording from the conclusion)? Reorganize the materials and dispatch again. - [ ] Did the check report give conclusions without basis? A "confirmed" with no basis you can open is treated as "undecidable." --- ## Template 3 Β· Verification Budget Sheet (fillable) **How to use.** One page that assigns a layer, row by row, to the project's outputs for the next month. Fill it in and paste it into the project README or the team wiki. Half its job is setting discipline for yourself, and half is making "has this thing passed the layer it should have passed" a question anyone on the team can ask. Review it once a month. When an output's destination changes, its layer changes with it. ```text # Verification budget sheet Project: ____________ Owner: ____________ Date filled: ____________ Next review date: ____________ | Output | Destination* | Cost of being wrong | Probability of being wrong** | Layer | Verification channel*** | Escalation trigger | Last checked | |---|---|---|---|---|---|---|---| | ______ | ____ | high/medium/low | high/medium/low | L_ | ______ | ______ | ______ | | ______ | ____ | high/medium/low | high/medium/low | L_ | ______ | ______ | ______ | | ______ | ____ | high/medium/low | high/medium/low | L_ | ______ | ______ | ______ | * Four destinations: self only / team discussion / the decision chain / the public knowledge base ** For the probability of being wrong, use the Chapter 11 failure-mode census: which kind of error this type of output most often grows *** Name the verification channel specifically (which model / which prompt path / which colleague), and it must be separate from the generating side Table-wide escalation rules (locked in, no bargaining on the spot): - A spot check finds a hard defect (fabricated citation / sourceless number) β†’ the whole output moves up one layer - The output's destination escalates (e.g. an internal draft gets cited in a decision document) β†’ reassign the layer by the new destination - A conclusion gets cited by a bigger decision β†’ that conclusion is named L2 on its own ``` **Self-check** (after filling it in): - [ ] Is the whole sheet L0? Either the project has no output that dares enter the decision chain, or you are exempting yourself from inspection. - [ ] Is the whole sheet L2? You have handed back all the speed dividend AI gave you. Layering is pricing between holding the line and taking the speed. All L2 is a pricing failure, not rigor. - [ ] Does the verification channel column just say "AI"? Be specific. Which model, which path, and whether it is the same one as the generating side. A channel not named specifically will, when the time comes, take the easy road and be the same one. - [ ] Is the escalation trigger written as "depends"? That is not a rule. It is a backdoor left for your future self. - [ ] Is any row's layer set because "it was good quality last time"? The layer follows destination and cost, not impressions. - [ ] Is the next review date blank? A budget sheet with no review cadence is an out-of-date decoration three months later. --- ## Card 4 Β· Downgrade note template (matches Chapter 12, section 12.6) **How to use.** For when the budget runs short. **Pre-write** this line in your project template first. Downgrades always happen on the busiest day, and on that day you will not have the mind to improvise the wording. You will skip it. The filled-in line travels with the conclusion and gets cited along with the numbers. ```text [VERIFICATION LEVEL NOTE] This conclusion is delivered at the ____ tier (ten minutes / half an hour / two hours / no next round). Checked: - Criteria timestamp: ______ (before the results / after the results / no criteria) - Citation spot check: ___ / ___ passed - Number tracing: ______ (all / sampled ___ / not done) - Independent channel reverse question: ______ (done, weakest claim is ___ / not done) Not checked: ____________ (list the steps that were cut, as they are; "the rest omitted" is not allowed) Reason: ____________ (the real constraint: exactly what the time / budget / permission ceiling is) ``` **The three floors of the "no next round" tier** (when the interrogation killed the conclusion and there are no resources to rerun, follow each one): 1. **Report the numbers under the original criteria as is**, with the interrogation's output standing beside them, labeled "post-hoc." Killed in the interrogation does not mean deleted. Deleting is the fraud. 2. **Write "no next round" into the limitations, and be specific**. Not "limited by resources," but "the X this conclusion depends on has only ___ independent units, a confirmatory retest would need about ___ additional samples, and this project did not run it." 3. **Lower the claim strength to the tier the evidence can carry, not to zero.** Erring upward is writing an unconditional statement knowing the evidence falls short. Erring downward is being frightened into saying nothing. Both directions are dereliction. **Self-check**: - [ ] Does the "not checked" column say "the rest omitted"? That equals writing nothing. A downgrade without a record and a pretense of the full set look identical to a downstream reader. - [ ] Did you cut by trimming a little from every layer? Wrong. Cut layers, not the order. Cut the later ones whole and keep the earliest in the order (the criteria timestamp). A gap in every line of defense = no defense on any line. - [ ] Does the "reason" column say "this one is not important"? That is not a budget problem, it is a prioritization problem. Importance is set by destination, and what enters the decision chain is important. Prioritization problems are solved by delaying delivery, not by downgrading. - [ ] At the ten-minute tier, did you only read it through? The only action at the ten-minute tier is checking the criteria timestamp. Reading through is the least profitable use of ten minutes, because what you read is precisely the other side's strongest face. --- # Chapter 13 templates Β· The Honest Map Toolkit > How to use. A three-piece kit, used together. First list the rows with Template A, then assign each row a status with rules card B, and last schedule the reviews by cadence table C. Goes with Chapter 13 of the main text. --- ## A. Honest map template (five columns, fillable) **How to use (one line).** Start from the claims you cited or took as true by default in your most recent deliverable, 10 rows at most. Fill in "claim" and your gut status first, then the fourth column. Rows where the fourth column cannot be filled drop to "still exploring (nobody knows)" automatically. ```text | # | Claim | Current status | Key evidence | What evidence would change its status | Last reviewed | |---|-------|----------------|--------------|---------------------------------------|---------------| | 1 | [one sentence that can be judged true or false] | [verified / still exploring (evidence accumulating or nobody knows) / falsified] | [identifiable source 1; source 2] | [specific, "on seeing X, change to Y"] | [YYYY-MM] | | 2 | | | | | | | 3 | | | | | | | … | | | | | | ``` **Rules for each column.** - **Claim.** Write a sentence that can die. Rewrite "X has a lot of promise" as "X beats [baseline] on [setting]." Carry the qualifiers. The qualifiers are part of the claim. - **Current status.** When unsure, "still exploring" without exception, then mark the grade (evidence accumulating / nobody knows). Fence-sitting wording like "basically verified" is banned. - **Key evidence.** Identifiable sources only (papers, controlled measurements, reproducible practice). "Industry consensus" and "everyone says so" are not accepted. - **What evidence would change its status.** Write it in the form of a search instruction, so at review time you check straight from it. Write both the upgrade condition and the downgrade condition. - **Last review date.** Changes on every review, even when the status did not move. **Self-check (run through before wrapping up).** - [ ] Every claim in every row can be hit by evidence (ones that cannot be written in falsifiable form, delete or rewrite) - [ ] No row has an empty fourth column - [ ] The map is not all "verified" (all green = placebo, not a map) - [ ] Rows where "this row holding is good for me," the required evidence level has gone up one notch - [ ] All review dates filled in, and scheduled into the cadence table - [ ] No status cell is copied (from other people's maps copy only the structure, re-derive or spot-check the status yourself) --- ## B. Status rules card **How to use (one line).** First screen the evidence with the two disciplines, then grade the evidence with the evidence-level table, and last assign the status by the three operational definitions. **Two ruling disciplines (before everything else).** 1. **A demo is not production evidence** (row 2 of the transfer map). Demo videos, case write-ups, and vendor material only qualify to trigger an investigation, never to decide a status. 2. **Self-report is not measurement** (row 3 of the transfer map). "Everyone who used it says it's great" does not go into the evidence column. Controlled measurements do. **Three operational definitions.** - **Verified**, all three at once: β‘  production-grade evidence (not a demo, not a case write-up); β‘‘ independent sources β‰₯ 2, or you can reproduce it with your own hands; β‘’ you can write its falsification shape (the fourth column is not empty). Missing any one, back to "still exploring." - **Still exploring**, the default when neither end is reachable. Then mark the grade: **Evidence accumulating** (the direction shows, the volume does not) / **Nobody knows** (not even a direction). When unsure, put it here. Faking certainty is the one unforgivable error on this map. - **Falsified**, either one: β‘  a criterion written down in advance was triggered; β‘‘ a reproducible counterexample punched through the claim as stated. Note, punching through the unconditional form is not punching through the weakened form. After falsification, give the surviving weak form a row of its own. Falsified rows **stay, never deleted**, with the falsification date and evidence noted (tombstones are part of a map's credit). **Evidence-level quick table.** | Level | Form of evidence | What it can support | |---|---|---| | E1 | Controlled measurement against preregistered criteria; independent replication by several parties | Can set "verified" or "falsified" | | E2 | A single peer-reviewed study / systematic evaluation | Can set a direction; alone not enough for "verified" | | E3 | Production practice reproducible by many users | Can support "verified" for workflow-type claims | | E4 | Demos, case write-ups, vendor material | Only qualifies to trigger an investigation | | E5 | Hearsay, intuition, self-report | Does not go into the evidence column | **Rule of use.** A row's status is decided by its highest-level evidence. A row whose evidence column holds only E4 / E5 drops to "still exploring" at once. --- ## C. Suggested review cadence | Current status | Cadence | Example trigger events | |---|---|---| | Verified | Every 6 months, or at once on a field-level event | A new model generation ships; the workflow you depend on changes heavily | | Still exploring (evidence accumulating) | Every 3 months | A new paper of the type the fourth column describes appears | | Still exploring (nobody knows) | Event-driven + one scan every 6 months | The evidence the fourth column describes shows up for the first time | | Falsified | Not reviewed, tombstone kept | / | | ⬜ Pending backfill | Tied to when the experiment / event completes | A preregistered experiment reports results | **Review action list.** 1. Run only the fifth column, row by row. It is the ready-made search instruction. Do not reread all the literature. 2. Only three moves allowed: change the status (the evidence arrived) / change the evidence (a harder source replaced it) / change the date (checked, nothing moved). 3. **Even when only the date changes, it must change.** A map whose review dates do not move is a dead map. 4. Log one line per review: date, which rows moved, why. 5. AI runs only the fifth column's searches. Status changes must pass through a human. Let AI change the status directly and you get a map where every row is fluent and confident, and fluency is exactly what you are guarding against. --- # Chapter 14 templates Β· Process Self-Check Sheet + Dispatch Brief Template > Prerequisite. You have read the Chapter 14 prose. Template 1 goes with 14.3/14.8 (the craft migration list and the "my process list" exercise), Template 2 goes with 14.5 (dispatch craft). --- ## Template 1 Β· Process Self-Check Sheet (fillable) **How to use.** List the process steps you actually did last week (list them against your calendar, not from impression), label each one with the two sorting questions, work out the time distribution, and circle the first step you will dispatch this week. Refill it once a quarter. | # | Process step (verb first, down to the object) | Cheap ground truth? (yes / no) | How fast do errors show up (on the spot / at delivery / only downstream) | Transcription / judgment | Depreciating / appreciating | Share of last week | Action (dispatch / keep / drill) | |---|---|---|---|---|---|---|---| | Example 1 | Format 42 citations | Yes | On the spot | Transcription | Depreciating | 10% | Dispatch | | Example 2 | Decide which baseline to compare against | No | Only downstream | Judgment | Appreciating | 5% | Keep + drill | | 1 | | | | | | | | | 2 | | | | | | | | | 3 | | | | | | | | | 4 | | | | | | | | | 5 | | | | | | | | | 6 | | | | | | | | | 7 | | | | | | | | | 8 | | | | | | | | | 9 | | | | | | | | | 10 | | | | | | | | | 11 | | | | | | | | | 12 | | | | | | | | | 13 | | | | | | | | | 14 | | | | | | | | | 15 | | | | | | | | **Three summary lines once it is filled in:** - Total share of time on the depreciating side: ______% (most people land at six to eight tenths on the first pass) - This week's first dispatch (the most time-consuming step on the depreciating side, go write Template 2): ____________ - Weakest item on the appreciating side (pick one of taste in questions / criteria design / dispatch craft / verification discipline / honest calibration): ____________ **Self-check**: - [ ] Every step is specific down to "verb + object"; anything written at the grain of "did research" or "read the literature" gets split again; - [ ] Time shares filled in against a calendar or a time log, not from impression, self-perception is not trustworthy (row 3 of the transfer map); - [ ] The "judgment" label has to pass a test. If you cannot say what counts as wrong for this step, do not label it judgment yet, it may be transcription you have not thought through; - [ ] Any step labeled "mixed" must be split in two and refilled; "mixed" is the usual escape from sorting; - [ ] Depreciating does not mean stop doing it. A depreciating step goes from "you do it by hand" to "you dispatch and accept," what disappears is the hand, not the responsibility; - [ ] This sheet is a snapshot, not a verdict. The dividing line moves with the tools, and an expired list is as dangerous as an expired map. --- ## Template 2 Β· Dispatch Brief (task / context / boundaries / acceptance criteria, four columns, fillable) **How to use.** Once it is written, run it through the "stranger executor test." Can an executor who has never met you and cannot ask you questions start work from this brief alone, and know what counts as delivered? If it does not pass, fix the brief before you dispatch. ```text [TASK] What to produce (verb first, one sentence): Medium and format (file type / table structure / word or line count): Deadline and priority: [CONTEXT] Background in one sentence (why this task exists, who the output is for): Required reading list (files / links, and which one wins in a conflict): Key terms and basis (in this task "X" is defined as ...): Existing conclusions, or ones already overturned (if any, state the status): ☐ Confirmed before dispatch, the context pack is the current version (stale facts are the top source of rework, an agent does not refresh facts on its own) [BOUNDARIES] Not allowed. Inventing facts, citations, numbers. Anything uncertain gets marked "to verify" Do not touch (what is outside the executor's remit this time, such as the wording of claim strength): When information is missing. Come back with a list of "facts I need," no filling in from imagination Explicitly out of scope: [ACCEPTANCE CRITERIA] Executable checks that the deliverable passes (a third party can run them): 1. 2. 3. Rework conditions (any one of the following sends it back): ``` ### Bad brief and good brief, side by side (The examples are built for teaching, not facts from this book's case.) **Bad brief.** ```text Look into whether multi-agent beats a single model, focus on the latest progress, and write me a solid summary, not too long. ``` Item by item, what is wrong. "Look into" defines no deliverable. "Latest" has no time basis. "Solid" and "not too long" cannot be accepted against. Zero context, the executor does not know why the survey exists or what conclusions you already hold. Zero boundaries, citations can be invented and no list will catch them. Whatever comes back, good or bad, you have no acceptance standard, and the rework rate is left to chance. **Good brief.** ```text [TASK] Produce an evidence summary on "multi-agent vs single model," a markdown table, one paper per row: claim / evidence / limits of application / relevance to our question. 10-15 rows, by 22:00 tonight. [CONTEXT] Purpose. Give the decision meeting on "do we pilot multi-agent in the retrieval module" a base table of evidence. Required reading. The attached Controversy Map v3 (this one wins, do not use the version in your memory). Basis. "Stronger" means accuracy after cost alignment, not comparisons with unaligned cost. [BOUNDARIES] Include only papers you can give a real source for. Mark anything you are unsure exists "to verify." Draw no conclusion on "should we pilot," that is the meeting's business. If basis information is missing, come back with a list. [ACCEPTANCE CRITERIA] 1. Every paper is findable in an academic search engine (attach the search link); 2. Every row's "claim" uses the paper's own qualifier wording, with no qualifiers removed; 3. No row has an empty "limits of application." Rework conditions. A claim with no source, or any citation that cannot be found. ``` **Brief self-check**: - [ ] The stranger executor test passes. Someone who cannot ask you questions can start from it and knows what counts as delivered; - [ ] The context pack has been updated to the current version (the rework lesson from this book's own case, a writing agent read an old version of the facts doc and three passages came back for redoing); - [ ] The acceptance criteria can be run by a third party, with no criteria of the "write it better" or "go a bit deeper" kind; - [ ] For a verification task, the brief leaks no expected answer and gives only the claim to be checked (the channel discipline in Chapter 12); - [ ] The boundaries column carries "come back with a list when facts are missing," giving the executor a way out other than imagination, the cheapest gate there is against hallucination. --- # Chapter 15 templates Β· 30-Day Adoption Plan + Three Process Questions Card + Team Metrics Starter Sheet > This appendix is the complete fillable version of the three tools in Chapter 15. > The order of use is fixed. Start with Template 2 (the three questions card) to install your first habit unit (the two-week trial run of 15.8). Once the trial run has a catch record, open up Template 1 (the full 30-day plan). Leave Template 3 alone until your personal process runs smoothly. Team metrics are built on top of a personal working example, and reversing the order gets you the enterprise tool from section 15.1 that nobody logs into. --- ## Template 1 Β· 30-Day Adoption Plan **How to use.** Install one part a week, and hold a retrospective at the weekend. Fill in the "success criterion" field before you start. For your own adoption plan too, the criteria get locked in first. The weekly retrospective answers two numbers only, coverage (of the times it should have fired, how many it actually ran) and the catch record (what got caught). ```text # 30-day adoption plan Start date: ____ ## Week one one process step - Process step chosen: __________ (suggested: the one where AI is deepest in, with a moderate cost of error) - Trigger (an objective event, precise down to the action): __________ - Content of the three questions card (copied from Template 2, filled in as yours): __________ - Success criterion: coverage β‰₯ ____% (suggested 80%; output quality is not assessed) - Weekend retrospective: covered ___/___ times; catch record: __________ ## Week two chain two (generation + verification welded shut) - Downstream verification action: after every generation, run one L0 spot check (Chapter 12) - Spot-check parameters: ___ citations / ___ numbers sampled (starting default 5 citations or 10%) - Success criterion: share of generations followed by verification β‰₯ ____% - Weekend retrospective: covered ___/___ times; catch record: __________ ## Week three run the whole workflow once - Small real problem chosen (two or three days of work): __________ - Take the small loop of the seven steps, answer the three questions at every door, rough is allowed - Record: which door made the three questions most awkward? __________ (= where the checklist needs rewording) ## Week four retrospective and retirement - Total times the three questions were answered: ____ - Catch list (one line each: what got caught, at which process step): __________ - Lines that never caught anything: __________ β†’ delete or rewrite - Decisions: which triggers stay / which process steps get added next month: __________ ## Acceptance test at day 30 (behavioral signal, tick one honestly) - [ ] Skipping the three questions feels awkward (the default has been swapped, the institution has taken over) - [ ] It does not feel awkward (back to week one, pick a more painful process step or a harder trigger) ``` --- ## Template 2 Β· Three Process Questions Card **How to use.** Install it on the path, do not stick it on the wall (the first line of the conversation template, the first column of the dispatch brief, the head of the document template). Write your project's specific answer after each of the three questions. Nouns and numbers count, adjectives do not. One card per process step. Different steps get filled in separately and never share a card. ```text # Three process questions card Process step: __________ Trigger: __________ Question one, criteria (Chapters 6 and 12) Where does this output go? [ ] stays on my desk [ ] into team discussion [ ] into the decision chain / public β†’ Verification layer: L____ What counts as passing: __________ (a check a third party can execute) What counts as losing: __________ (locked in before you start; no answer, no start) Question two, delegation (Chapter 3) Level AI sits on for this step: [ ] tool [ ] assistant [ ] collaborator [ ] autonomous (local, see Chapter 7) Who drafts: ____ Who reviews: ____ Who decides: ____ Who answers when it's wrong: ____ (The "who answers when it's wrong" field has to be a person's name. If you cannot write a name, drop a level and ask again.) (The four answers combined are your title in the human-side role table of Chapter 14, the artisan, the lead writer, the editor-in-chief, or the principal.) Question three, verification (Chapters 11 and 12) Which class is this process step most likely to break in (open your failure-mode census sheet): __________ Signature (the specific signal in your project, not copied from the book): __________ Which layer it passes before it leaves: L____ Which independent channel executes it: __________ ``` **Retirement cadence (carried with the card).** Hold a retrospective once a month. Every line either produces a recent catch record or gives an explicit reason to stay. A line with neither gets deleted. The health metric is the catch record, not the length of the checklist. --- ## Template 3 Β· Team Metrics Starter Sheet **How to use.** Use the three metrics in pairs, draw trend lines, and keep them for this group's learning only. They do not enter individual performance reviews, and absolute values do not get compared across teams (the basis differs, the comparison is meaningless). Audit the metrics themselves once a quarter, and for the one whose number improved, first ask "did things get better, or did the reporting change." | Metric | Basis | Collection method | Paired metric (anti-gaming) | This month | Last month | |---|---|---|---|---|---| | Rework rate | Share of output with AI deeply involved that gets sent back for redoing after delivery downstream | The "sent back" label plus its reason on the task board, counted once at month end | Throughput (stops people cutting rework by delivering less) | | | | Verification pass rate | Share passing on the first try in spot checks and full verification; citation existence, paraphrase fidelity, and number traceability recorded separately | The verification ledger (Chapter 12) is the data source, nothing separate to collect | Verification coverage (stops people checking only the safe output) | | | | Claim survival rate | Share of conclusions entering the decision chain that still stand at a scheduled review point (three months, say) | The conclusion register plus the review cadence of Chapter 13 | The risk level of the conclusion (stops conclusions getting more and more timid) | | | **Rules that ship with it (copy these along when you copy it into the team wiki)**: 1. The numbers serve this group's learning and do not enter individual performance reviews (once they do, rework moves into private messages and the board is at peace forever); 2. Metrics have to be read in pairs, since a single metric always has a painless cheat posture; 3. Audit the metrics themselves once a quarter; 4. Whoever updates signs, the same discipline for the metrics sheet as for the case status doc. --- # Chapter 16 templates Β· The One-Hour-a-Month Tracking Recipe Card > How to use. This card covers change in the craft of "doing research with AI" itself, not your research field (that is the job of the Chapter 4 frontier layer), and not the row-by-row re-derivation of your own map (that is the job of the Chapter 13 review cadence). One fixed hour a month, in three steps. 20 minutes on the sources β†’ 30 minutes running the fifth column on the signals that hit β†’ 10 minutes changing a status or a date and writing one line of change log. Checked and nothing moved, the date still changes. --- ## Part one, the source diet list (fillable template) The rules get locked in first: - **Total capped at 5.** Add one and you must drop one first. - **A retirement review once a quarter.** Any source that has not triggered a single "run the fifth column" in the past three months gets downgraded or removed. - Influencer newsflashes and algorithmic feeds do not use up quota. That kind of information reaches you anyway and needs no active subscription. | Class | Source name | Why it stays (one sentence) | Date it last triggered an action | |---|---|---|---| | First-hand source (original papers / system technical reports, not retellings) | ______ | ______ | ______ | | First-hand source (an optional second one) | ______ | ______ | ______ | | Production-retrospective source (an account written by someone who really used it for three months, not a day-one review) | ______ | ______ | ______ | | Controlled-measurement source (comparison studies with a baseline and a stated basis) | ______ | ______ | ______ | | Your choice (a high-quality aggregator and the like, use sparingly) | ______ | ______ | ______ | ## Part two, the list of signals that trigger a re-derivation (any hit β†’ this month's 30 minutes goes to it) The general rule is one sentence. **A piece of news is worth an hour of yours if and only if it could change the status of some row on some map of yours.** In practice, hold it up against that row's fifth column ("what evidence would change its status") and ask, is this the evidence that column describes? - [ ] **A production-grade measurement has appeared.** Not a demo, not a launch event. Real accounts like merge rate, rework rate, verification pass rate, claim survival rate. - [ ] **An independent reproduction or an independent refutation has appeared.** A second pair of hands with no stake in the original authors got the same result, or the opposite one. - [ ] **A criterion locked in earlier has been tripped.** A preregistered falsification condition, or the evidence described in the fifth column of some row of your own map, has surfaced. - [ ] **Repeated rework in your own workflow.** The same process step crashes several times running. Your first-hand data is evidence too, and the closest kind there is. - [ ] **A policy or basis change from an authority.** Journals, reviewers, or regulators changing the rules for AI's part in research. It reshapes the terrain of the "deliver and take the hits" stop directly. ## Part three, the ignore list (let it flow straight past, into no notes and no anxiety) - **Demo videos and vendor launch events.** Handled by row 2 of the transfer map. Qualified only to trigger an investigation, not to decide a status, and least of all to trigger anxiety. - **Self-reported speedups with no control.** "I got three times faster after using X." See row 3 of the transfer map, self-perception is not trustworthy, calibration comes from measurement. - **Model version-number news.** "X.5 is out" is not a signal by itself. A controlled measurement of it on your own tasks is. - **The Nth round of the "AI replaces scientists" debate.** An argument fought in the air produces no evidence that can hit the fifth column. - **Secondhand retellings with no link to the original.** Every hand a retelling passes through drops one layer of qualifiers. A claim whose original you cannot find is treated as nonexistent. - **Anonymous leaks and rumors.** Evidence that cannot be pointed at is not evidence. ## Part four, the hour's time budget (copy it as is) | Slot | Action | Output | |---|---|---| | 0–20 minutes | Run through the sources on the diet list, sorting into two piles only, hits the signal list / flows past | 0–3 candidate signals | | 20–50 minutes | For the signals that hit, run the fifth column of the matching map row. Has that kind of evidence actually appeared | Change the status / change the evidence / confirm nothing moved | | 50–60 minutes | Update the map, changing a status or only a date, and write one line of change log | Every review date refreshed | ## Part five, Self-check (failure modes condensed) - [ ] Are you over 5 sources? Cut. The number of sources is not proportional to the quality of your judgment, it is proportional to your anxiety. - [ ] Has the hour swollen into three hours? Then you are reading a feed, not running the recipe. Go back to "ask only against the fifth column." - [ ] Three months running of nothing but date changes, and yet more anxious? The anxiety comes from not executing the ignore list. Without a rule that permits things to flow past, tracking is slow poisoning. - [ ] Have you handed the map to AI to maintain fully automatically? AI can run the searches for you, but a status change must pass through a human (the last cordon of Chapter 13). - [ ] When was the last retirement review? A list that only takes in and never lets go becomes, three quarters later, another forty-seven browser tabs. --- # Template index The entry point for the templates of all 16 chapters, in chapter order. Each template first appears in its own chapter and is meant to be used with the text. | Chapter | Template pack | Contents | |---|---|---| | 1 | [Chapter 1 templates](ch01-templates.md) | The "Is This Book for You" Self-Test | | 2 | [Chapter 2 templates](ch02-templates.md) | The Transfer Map | | 3 | [Chapter 3 templates](ch03-templates.md) | Ladder Self-Rating Sheet + Seven-Step Workflow Check Card | | 4 | [Chapter 4 templates](ch04-templates.md) | Controversy Map Prompt + Paper Card + Frontier-Layer Subscription Rules + Coverage Checklist | | 5 | [Chapter 5 templates](ch05-templates.md) | Five-Step Question-Sharpening Prompt Set + Question-Sharpening Card | | 6 | [Chapter 6 templates](ch06-templates.md) | Test Plan Template + Confounder Checklist + Plan Red-Team Prompt | | 7 | [Chapter 7 templates](ch07-templates.md) | Minimal Harness Checklist + Ledger and Resume Patterns + "Have AI Audit the Harness" Prompt | | 8 | [Chapter 8 templates](ch08-templates.md) | Result Interrogation Checklist + Preregistered/Post-hoc Side-by-Side Report Template + Surprise-Result Red Flags | | 9 | [Chapter 9 templates](ch09-templates.md) | Two Real Deliverables + Claims List / Technical-Report Skeleton / One-Page Memo Templates + Delivery Prompt Set | | 10 | [Chapter 10 templates](ch10-templates.md) | Red-Team Dispatch Brief Template + Four-Attack-Surface Red-Team Prompt Set + Disposition Table | | 11 | [Chapter 11 templates](ch11-templates.md) | Failure-Mode Census Sheet + AI Self-Check Prompt Set | | 12 | [Chapter 12 templates](ch12-templates.md) | Layered Verification Workflow Card + Independent-Channel Check Prompt Set + Verification Budget Sheet | | 13 | [Chapter 13 templates](ch13-templates.md) | The Honest Map Toolkit | | 14 | [Chapter 14 templates](ch14-templates.md) | Process Self-Check Sheet + Dispatch Brief Template | | 15 | [Chapter 15 templates](ch15-templates.md) | 30-Day Adoption Plan + Three Process Questions Card + Team Metrics Starter Sheet | | 16 | [Chapter 16 templates](ch16-templates.md) | The One-Hour-a-Month Tracking Recipe Card | Companion code (rerunnable experiments): [`code/smol-army`](https://github.com/hallieren/research-rewritten/tree/main/code/smol-army/) Β· [`code/persona-panel`](https://github.com/hallieren/research-rewritten/tree/main/code/persona-panel/) --- # Experiment ledger index The crime scene for every number in the book. Each experiment's preregistration **entered its repository before any result did**. The criteria existed before the results, which is the method Chapter 6 teaches, applied to the book itself. | Experiment | Question | Preregistration | Raw results | Red team | Cost ledger | One-line conclusion | |---|---|---|---|---|---|---| | **Naive demo** (Start Here) | One 20B open-source model alone against the frontier model, the first 20 HumanEval+ problems | / | [demo_naive.jsonl](https://github.com/hallieren/research-rewritten/blob/main/code/smol-army/results/demo_naive.jsonl) | The whole chapter is its red team | A few cents | 19/20 vs 20/20, the number is real, but seven reasons say do not trust it yet | | **smol-army main experiment** | Small models from 9B to 20B in teams (voting, debate, division of labor), can they tie the frontier model on a cost-matched basis | [prereg.md](https://github.com/hallieren/research-rewritten/blob/main/code/smol-army/docs/prereg.md) (entered before results; later changes in [CHANGES.md](https://github.com/hallieren/research-rewritten/blob/main/code/smol-army/docs/CHANGES.md)) | [report.md](https://github.com/hallieren/research-rewritten/blob/main/code/smol-army/results/report.md) Β· [results.csv](https://github.com/hallieren/research-rewritten/blob/main/code/smol-army/results/results.csv) Β· [runs.jsonl](https://github.com/hallieren/research-rewritten/blob/main/code/smol-army/results/runs.jsonl) | [redteam-2026-07-25.md](https://github.com/hallieren/research-rewritten/blob/main/code/smol-army/docs/redteam-2026-07-25.md) | [ledger.jsonl](https://github.com/hallieren/research-rewritten/blob/main/code/smol-army/results/ledger.jsonl) ($35 hard cap) | Split by task family, code tie in doubt (CI crosses zero), math the army ahead, mmlu_pro the army behind. "Can it tie" has no one-word answer, Chapters 8 and 13 unpack it | | **persona-panel** (subplot) | Does the answer distribution of an LLM persona panel look like the real subgroup it imitates (ground truth, the WVS-7 US sample) | [prereg.md](https://github.com/hallieren/research-rewritten/blob/main/code/persona-panel/docs/prereg.md) (entered before results) | [report.md](https://github.com/hallieren/research-rewritten/blob/main/code/persona-panel/results/report.md) Β· [answers.jsonl](https://github.com/hallieren/research-rewritten/blob/main/code/persona-panel/results/answers.jsonl) | See Chapters 11 and 12 | [ledger.jsonl](https://github.com/hallieren/research-rewritten/blob/main/code/persona-panel/results/ledger.jsonl) ($10 hard cap) | Preregistered verdict **FAIL**. The variance-collapse red line tripped and the subgroup cross-check failed. "Answers like a real person" does not hold at the distribution level, Chapter 12 reveals it | ## The book's own verification ledger The four counts quoted in Chapter 15, section 15.7, come from here. The fact-checking round during drafting registered 48 facts pending verification (Chapter 1, 7 items; Chapter 2, 17; Chapter 3, 4; Chapters 8 to 10, 5; Chapters 11 to 15, 15). 34 were verified, 1 core claim was overturned outright, and about 10 statements were corrected or narrowed. The item-level reports were retired from the book on 2026-09-05 and are kept in the author's archive. ## Reproduction notes - Both projects run the whole chain offline with `--mock`, at zero cost. Output goes to `results_mock/` and never touches the real results shipped with the repository. Unit tests are all offline. - Every API call lands in `ledger.jsonl`, and a run aborts on its own when it hits the budget hard cap. Runs resume from where they stopped. - smol-army's `data/SHA256SUMS` pins the task-set files. The original files from the evaluation were not kept, the hashes come from a 2026-09 re-fetch, and the 3,580 archived answers in the repository were rescored to verify that they match the original data, errata in CHANGES.md. persona-panel's WVS-7 raw data must be registered for and downloaded by you (the license does not allow redistribution, see [data-license.md](https://github.com/hallieren/research-rewritten/blob/main/code/persona-panel/docs/data-license.md)). - Got a different number? Tell me with the [reproduction report template](https://github.com/hallieren/research-rewritten/issues/new?template=repro-report.yml).