Preface · This Book Was Put on Trial
Chapter companion
You have probably just read a report written by AI, thought it looked better than the last one you wrote by hand, and stopped with your hand over the forward button. That one second of stopping is the whole problem this book deals with. Before saying how it helps you, first what it did to itself.
For a book that teaches you not to trust AI output too readily, the most embarrassing way to die is for its own claims to fail its own process. Before writing I set one house rule. Every step of the process this book teaches gets run on the book itself first. This preface puts the bill in front of you. The bill is also the first piece of evidence for the book's claim.
The claim first. AI drove the cost of producing plausible to the floor. A report with a clean structure, full citations, and confident wording went from three weeks to three hours. The cost of reliable barely moved. The crack that opened between them is the whole content of this book. The craft of research has a set of steps whose only job is to turn plausible into reliable. Sharpen the question, lock in the criteria, run it so it can be reproduced, interrogate the results, deliver honestly, red-team your own work. A criterion is the pass-or-fail line written in advance, and it cannot be changed afterward. Red-teaming means picking holes in your own claim before anyone else does, the same move as red-teaming a model in AI safety. As long as a wrong answer means something real gets hurt, these steps are your protection, and academic etiquette is the least important of their identities. This book writes out the full process for one kind of person, the software or AI engineer who has to do research at work. You are asked to sign off on a technical judgment (can a small model replace the large one, should the retrieval layer be swapped, can the numbers in this report be trusted), and the evidence is a computational experiment anyone can rerun. The signature is yours. If you are not that person, you can still read it. Researchers, analysts, and people who make decisions use the same process, with the shape of the case converted using the conversion table in Chapter 3, section 3.6, and the paragraphs below make that clear.
Now the bill. This book took several cuts of its own, each one on record.
A claim in the first draft of Chapter 4 was shot dead on the spot by the book's own process. I wrote, with some excitement, "cost-matched comparisons are a gap in the literature." The wording was pretty, and it happened to raise the value of my own experiment. An independent channel check, which is to say opening a brand-new session and handing the question to a different model that did not know which answer I was hoping for, traced the forward citations, the papers that later cited it, and came back. The comparisons exist, and they cluster on the skeptics' side. The claim died. Its corpse is pinned to the map in Chapter 13 that labels the status of every claim.
Two experimental lines were run with real money, and one of them did not survive to the finish. The book's spine case (a team of small open-source models vs a single frontier model, where frontier means the strongest commercial large models of the moment) and its subplot case (can AI personas replace interviews with real people) are both preregistered real experiments. The criteria went into the repo first, and after that only a change log could be appended. Which line died, how it died, and how a pretty number I almost celebrated early got overturned in the interrogation before it was announced, I leave to the main text. Only one thing can be said now. I did not know any of these endings before the runs started, and the timestamps on the criteria will testify to that.
That sentence itself needs a discount. When this book started its runs, there was no third-party preregistration. Chapter 12 will teach you to check the timestamp on the criteria before accepting any output, but that check has a hole of its own. A git timestamp can be forged. You only need to change the local clock. So "the timestamps will testify," in front of a reader who really presses, counts as self-attestation and falls short of evidence. To become evidence it needs a third party I do not control to stamp the time, and the common practice on the academic side is a preregistration hosting platform. When writing this book I did not do that from the start. This is a real omission. The remedy and the current status are recorded in the notes of the online case library, the part of this book that keeps updating on the web, with the address in Chapter 16. I did not delete that sentence, and this paragraph stays in the preface, because it demonstrates the book's basic move, take apart your own chain of evidence first, and write down the holes you find.
Every number points to its crime scene. Every accuracy, cost, and confidence interval in the book ends up pointing at specific files and commits in two public repositories (smol-army and persona-panel, under https://github.com/hallieren/research-rewritten/tree/main/code). You do not need to believe me. You can rerun it.
One more entry has to go in the preface, because by this book's own discipline, hiding it would be cheating. AI was used deeply in producing this book. Chapter drafts, literature checks, code scaffolding, a large share of the steps were executed by AI. That is exactly why I use the verification process in this book every day to protect myself first, and only then recommend it to you. The independent channel check has overturned my claims. The interlock check, which is to say making the same number in different chapters answer to each other, caught number drift in the drafts. In the subplot case's interviews there was also a survey question I misremembered, and only the preregistration discipline caught it. That story is in Chapter 7. A book with AI this deep in it is the first test bed for whether "move at its speed, hold the line of science" holds up. Whether it does, you judge for yourself when you finish.
What this book does not cover is also stated here. It is not an encyclopedia of cases by discipline, not a tool manual, and not a job-hunting guide. One more point, and it matters more. It is for people who are not the core reader. Every first-hand case in the book is a computational experiment that can be rerun, which is what my craft covers. If your evidence is wet lab, fieldwork, one-off interviews, or clinical data, you will find the process in this book and not a case of your shape, and forcing one would only produce cliches. The process sharpens how judgment is done right, and it does not care about the shape of the case, so the process still works for you, but the conversion is your job. Chapter 3, section 3.6 gives a conversion table with five premises. People for whom all five hold are the core reader described above. No need to memorize the five now. Check them one by one when you get there, and wherever you break one, patch yourself at that spot. There is one corner that is an exception. The output you sign off on came from someone else's hands and you never ran it yourself, which is Premise 3 in Chapter 3, section 3.6. Spot checks and full verification you can still do, and Chapter 12 is written for exactly that situation. How much to sample before it is enough, how to escalate when a check fails. This book has no answer. I labeled them "still exploring" there, and did not talk my way past them.
How to read, four paths, depending on who you are.
- The impatient enter through Start Here, the piece right after this preface. Two hours later you will have a number that makes your heart race, and a list that shuts it up.
- People with a real project to use it on right now, the route is Start Here to pick your question → Chapter 3 (especially the five premises in 3.6, which decide whether each later chapter needs converting) → Chapter 6 (lock in the criteria first) → Chapter 12 (acceptance) → Chapter 11 (failure modes). Those four chapters together are the most compressed part of this book.
- People who mainly accept other people's output, whose inbox holds more AI reports than experiments they ran themselves, the route is Start Here for the seven-item list only → Chapter 3, section 3.6 → Chapter 11 (failure modes) → Chapter 12 (the two-hour acceptance workflow) → Chapter 10 (the four attack surfaces as an acceptance checklist).
- People deciding whether this book is worth finishing go straight to Chapter 13. That is the book's master ledger, and the one place where I put my own claims and other people's claims on the same table and label their status, including two rows downgraded by my own rules and two tombstones of my own. If that table does not convince you, the twelve chapters before it probably will not either.
The map of the book's chapters is in Chapter 1, section 1.5.
One last ugly sentence, the one the whole book keeps repeating. Every claim in this book carries an evidence-status label, verified, still exploring, or falsified, and each one records "what evidence would change its status." This preface included.