Skip to content

Chapter 16 · Coda · How This Book Stays Current

Chapter companion

📋 Chapter 16 templates · 🗂 Template index

This chapter's ladder. Not marked. The coda sits outside the seven steps.

This book opened with a "don't trust it yet" list. Now it is time to settle the account. And, while we are here, to say why a book written in a field that changes every month does not expire on the day it is published.

This chapter delivers. A line-by-line settlement of the seven-item list, the split of labor between the paper book and the case library with the cadence that updates it, and a one-hour-a-month tracking recipe.


16.1 Back to that list

That afternoon in Start Here, you spent two hours running the most naive version of the experiment, no team formed at all, one open-source small model going alone against frontier, and saw 19/20 against 20/20. It looked like a tie. Then you did the most important thing of that day. You skipped the celebration and wrote out a seven-item list of "where I don't trust it yet."

That list is the real first page of this book. At the time it looked like a few lines of self-doubt. In fact it had already demonstrated everything the book sets out to teach. Now settle it line by line. Seven debts, booked to three accounts.

Statistical debt, two items. Too few problems. With 20 problems one problem is 5 percentage points, and the eye cannot tell 19/20 from one swing of luck. Each problem answered only once. Rerun at another time and 18 or 20 would be no surprise. Chapter 6 locked in the repeat count and the seed before the run, Chapter 8 turned "does the gap count" into arithmetic with intervals and tests, and Chapter 10 rechecked even the independence assumption behind the intervals themselves. The endgame is that the gap interval on code crosses zero, which reads as uninformative and not as a tie, and that on math the reversal rested on problems sharing templates, with an effective sample size far smaller than the number of rows, so the whole family retired.

Basis debt, two items. Cost not accounted for, and a "tie" with no price tag means nothing. Chapter 5 wrote the cost basis into the hypothesis, Chapter 6 turned it into a locked-in accounting rule, and Chapter 7 handed over a ledger that adds up line by line. The Chapter 10 red team then found a breach in that ledger. The control arm's budget was never aligned to the plan, and the phrase "same-budget control" is withdrawn across the book. Only one subject tested, and possibly a leaked one. Chapter 5 replaced the single demo task with task families spanning types and excluded contaminated tasks in writing, and Chapter 6 added the contamination check step. The price showed up in Chapter 8. The contamination-resistant variant set held off memorized problems, and carried in correlation within the same template instead.

Process debt, three items. No control arm. Chapter 6's three-arm design was signed off, and from then on "does the credit go to teaming or to more sampling" could be told apart. What telling them apart produced is that "real teaming beats sampling a single small model" never got measured once in the whole case.

Nobody reviewed the scorer. The Chapter 7 pilot stopped five bugs, then Chapter 8 and Chapter 10 each ran a round on it, because the audit itself has to be audited too. Each round caught one bug the pilot had missed, one handing points into math, one wronging the small models in the audit's own rescoring.

The first 20 problems were picked out of laziness and do not count as sampling, so Chapter 6 locked the sampling rule in before the run. The rule was set and the fetching script still did not shuffle. Chapter 8 pointed out runs of same-template variants inside math, and Chapter 10 seized its twin in knowledge QA, data drawn in blocks by subject, with not one STEM problem in it.

Not one of the answers these seven items received says "relax." Every answer is a process step. The whole distance between plausible and reliable sits right here. The you of that two-hour afternoon could only say "it looks like a tie, but I don't trust it." The you of now can say exactly how each item of that distrust turned into evidence.

16.2 One arc, four stops

The spine of the whole book runs one arc, doubt, verify, contribute, deliver and take the hits. One sentence per stop, nothing retold.

Doubt. I expected a consensus and found a brawl, load-bearing papers on both sides, each side arguing well. The correct shape of doubt is to draw the brawl as a controversy map and find which disagreement your own question lands in. Shouting "trust nobody" does not count.

Verify. From "can small models team up" to a hypothesis H that can die, then criteria, control arms, and falsification conditions all locked in first. Verification is a contract signed before the run. You cannot fill it in with attitude after the experiment is done.

Contribute. I got excited thinking I had found a gap in the literature, and a check knocked half of it down. The question narrowed from "fill the gap" to "draw the boundary clearly and supply the kind of comparison nobody had done in full." A question narrowed by a check is worth more than a question that excites you.

Deliver and take the hits. One body of evidence was written into two deliverables, and then a court really opened on four attack surfaces. All ten charges held. What had to be overturned was overturned, what had to be narrowed was narrowed, the rest got caveats attached, and the case file went into the repo where anyone can rerun it. The submission never happened, and the book says so plainly instead of acting otherwise. The substitute is spreading the case file out in front of everyone. The one-page memo after the beating sits at the end of Chapter 10, section 10.7, as the final version, shorter than the original, recommending one thing less, the army. The point of this stop is order. A beating scheduled before publication is scheduled by you. Scheduled after publication, it is scheduled by someone else.

H's ending is filed in row 13 of the Chapter 13 map, the row for spine hypothesis H, a task-dependent mixed ending, not falsified and not holding across the board. The criterion is a gap to frontier of ≤ 2 percentage points on a cost-matched basis, judged family by family. On code the interval crosses zero and reads as uninformative, at about 1/25 the price of frontier, list-price basis. Knowledge QA trails by 9 to 25 percentage points, a clean loss. The math family retired whole, and its problem set has to be reissued and tested again. No family caught up by spending more, the falsification condition never triggered, and no family reached the word "tie" either. The detailed accounts live in the online case library, the online repository whose address and contents come later in this chapter.

The subplot reached its stop too, and readers who want only the conclusion can read the first sentence and the last one of this paragraph. The coarse question "can AI personas replace interviews with real people" was ground into a question about distributions, dissected in Chapter 11, locked into a plan in Chapter 12, with the three criteria, the falsification shape, and the downgraded claim all written in advance. The reveal has already happened. All three criteria failed. The distribution distances of six subgroups, 0.19 to 0.25, all exceed the threshold of 0.10, the top-choice agreement rate is 40% against a threshold of 70%, and eighty percent of the clean questions tripped the variance-collapse red line. The preregistered falsification shape was met exactly as written, and "persona can replace real interviews" is filed as falsified. The plan promised that the reveal would only fill in a status, and that promise was kept. However ugly the result, it only filled a status into row 14, the row for persona survey rehearsal. A line that puts synthetic data in for real data ran the same discipline as the spine from start to finish. That is the whole point of its existence, proof that this workflow holds when the stand-in changes.

16.3 What expires, and what does not

Now to be clear about what this book actually delivers.

Conclusions do not make the delivery list. They expire too fast. Any of the fourteen rows in Chapter 13 could flip before the next review, and every model and every benchmark named in the book has a shelf life measured in months. A book that bets on "these claims are true in 2026" starts rotting on the printing press.

This book bets on another set of things, and what they share is that they do not expire with a model version.

  • The structure of the seven-step workflow. Questions, hypotheses, criteria, execution, interpretation, delivery, critique. How high AI can stand on each step will change. The relations among the seven steps will not.
  • Verification discipline. Generation separated from verification, criteria locked in first, checks run through an independent channel, production-grade over demo-grade. These held in the spreadsheet era and they hold on the next generation of models.
  • The method for re-deriving the honest map. A five-column structure plus two disciplines for assigning status. The statuses on the map are readings and they change. The way of measuring does not.
  • How to use the transfer map. When a new headline arrives, first ask which row's pattern it is and which rerun of it. That move works just as well on a headline from 2029.

What gives these the standing to claim they do not expire? Not one of them was built for the models of 2026. They are mechanisms distilled from five paradigm shifts. Bet on mechanisms, not on outcomes, is the rule Chapter 2 laid down. A book judging its own shelf life should use that same rule.

Conclusions are the ink on the map. Method is the hand that draws it. This book teaches the hand.

16.4 How this book stays alive

The ink cannot be left to rot on paper either. At the foot of the table in Chapter 2, section 2.5, the second of the three notes on use wrote a check. The status column will expire, and the online edition keeps updating. Several more such checks stand in the book. Chapter 4 said the case library tracks automated review tools. Chapter 13 said the update mechanism for the official version of the map would be delivered in Chapter 16. Now they are honored.

The online case library is at https://github.com/hallieren/research-rewritten. It holds everything that expires.

  • The code, data, and result files of the spine and the subplot. Every number and every pit in the book can point to a commit and a result file in the case library. That one is a principle, and writing it down means honoring it.
  • The live version of the honest map. The fourteen rows in the paper book are a 2026-07 snapshot, and the snapshot date is printed in the map's header. The live version is re-derived row by row on the Chapter 13 review cadence, with a change log entry attached to every status change. The statuses of rows 13 and 14 keep being re-derived here on that cadence.
  • The live version of the transfer map's status column, and the latest version of each chapter's templates.

The paper book carries only the half that does not expire, the frameworks, the process steps, the disciplines for assigning status, the ladder, the five-column structure. The split is simple. What moves lives online, what does not gets printed. On a second reading, whenever you hit a concrete status, number, or result, check the date in the case library first, then decide whether to trust the paper version or the online one. Use the paper book as an entrance only.

The update cadence follows the Chapter 13 map review cadence rather than setting up a second one. The case library updates whenever a review produces a status change, and quiet months are not padded with content.

One honesty clause remains, and this promise is itself falsifiable. The case library's front page carries a last-updated date. The day you open it and find that date frozen for a long while, treat this book the way you treat any expired map. Keep using the method, re-derive the statuses yourself. A book that teaches you not to trust empty promises should not ask you to trust its own promise unconditionally.

16.5 One hour a month

The case library is my update cadence. You need one of your own.

Draw one boundary first, so this does not fight the steps that came before. The Chapter 4 frontier layer (the layer that turns tracking into a subscription) covers your research field. The Chapter 13 review cadence covers how your map gets re-derived row by row. This section covers the furthest upstream, information about the craft of "doing research with AI" itself, one hour a month, where it comes in, what gets through, what gets stopped.

The recipe is three pieces, and the full fillable version is in this chapter's appendix.

The source diet list. Five at most, and adding one means dropping one first. Keep one or two in each of three classes. First-hand sources, the original papers and system technical reports, retellings do not count. Production-retrospective sources, the accounts written by people who really used it for three months, day-one reviews do not count. Controlled-measurement sources, comparisons with a baseline and a stated basis. There is no "influencer newsflash" class on the list, because that kind of information reaches you anyway and needs no quota.

Signals that trigger a re-derivation. A piece of news is worth an hour of yours if and only if it could change the status of some row on some map of yours. In practice you hold it up against the fifth column and ask, is this the evidence that column describes. Production-grade measurements, independent reproduction or refutation, a criterion locked in earlier getting tripped, repeated rework inside your own workflow. Those four classes get through.

The ignore list. Demo videos and launch events, handled by row 2 of the transfer map, qualified only to trigger an investigation, never to trigger anxiety. Self-reported speedups with no control. Model version-number news. The Nth round of the "replacing scientists" debate. Secondhand retellings with no link to the original. These flow straight past.

Here is how the hour goes. Twenty minutes running through the sources, thirty minutes running the fifth column on one or two rows for the signals that hit, ten minutes changing a status or a date and writing one line of change log. Checked and nothing moved, the date still changes, a discipline inherited unchanged from Chapter 13. In most months your output will be a few new dates and nothing else. That is the recipe's normal output, not waste. Tracking asks for one thing only. In the month a conclusion flips, you are not absent.

Want an agent to run it with you? Paste this to your AI assistant or coding agent:

Help me set up Chapter 16's one hour a month. Build three sheets from docs/appendices/ch16-templates.md. The source diet list caps at
five, one or two in each of three classes. Every source on the list comes from me, you only check the class and the count, and if it goes
over the cap, one gets dropped first. Build the trigger-signal list and the ignore list as empty sheets from the appendix, and I fill them. Then copy
the hour's time budget into my calendar reminder, twenty minutes on the sources, thirty minutes running the fifth column on the signals
that hit, ten minutes changing a status or a date and writing one line of change log. Checked and nothing moved, the date still changes, and remind me of that one every month. If any command errors, stop and show me the output.

16.6 The second spine

One last thing, about you.

If you followed the exercise block at the end of every chapter to get here, you now hold these things, the first four thought through, the last four actually run.

  • A real problem you picked yourself
  • A controversy map
  • A hypothesis that can die
  • A set of criteria locked in first
  • A one-page verification budget sheet
  • Your own ten-row honest map, the one the exercise blocks had you draw, a different map from the fourteen-row one in this book
  • A process-step list marked depreciating and appreciating
  • A three process questions card, mounted on a trigger

Put together, these are a real project that has run one full round. The list of artifacts sorts into the four roles of Chapter 14, each in its place. It may still be crude, with some cells filled in "still exploring" or even "nobody knows." That is fine. One crude but honest round beats ten chapters of polished notes that never ran.

Do not let it stop at this round. Raise it as your second spine. File results honestly and write the ugly ones down. Criteria always go in front of results. The map gets re-derived on a cadence. Next round's question grows out of the cells still empty on this round's map. The spine case of this book will go stale one day, and by then you will not need it. Right now, put the next review date for your ten-row map into your calendar, and re-derive it row by row when it comes due.

16.7 The unfair advantage you now hold

You hold a real project you ran through the whole workflow with your own hands, an honest map you drew yourself, and a cadence that keeps both fresh. When the next dazzling demo arrives, the people who did not read this book start the "do we believe it" argument over again. You open your map and ask one question. Which row's pattern is this, and which rerun of it.

The field will keep changing every month. Models will get stronger, headlines louder, each demo more dazzling than the last, and new "don't trust it yet" items will keep growing on the list. None of that matters. One more item on the list, one more repaid the way these sixteen chapters do it. Write it on the list first, then find the process step that answers it. The whole practice fits in one sentence you can memorize, the one Start Here laid down and the book never changed.

Move at its speed, hold the line of science.