Skip to content

Chapter 9 · Delivery

Chapter companion

📋 Chapter 9 templates · 🗂 Template index · 💻 code/smol-army

This chapter's ladder. At the delivery step AI sits steadily at assistant level, and this is one of the steps whose price collapsed hardest in the whole book. First drafts, audience rewrites, figure captions, all can be handed off. There is no half-open door here. Responsibility for claim strength and wording cannot be delegated. One sentence gets pulled out and quoted on its own. Do you stand behind it? No model can answer that for you.

Spine update. The numbers have been interrogated. In this chapter I load the same batch of evidence into two completely different heads, a technical-report skeleton and a one-page CTO memo, and every number in the two documents has to interlock and point back to the repo.

This chapter delivers. A claims list, which is the single source of truth. Templates for the technical-report skeleton and the one-page memo. Plus a prompt set for audience rewriting and the number interlock check.


9.1 A twelve-page document, dead in thirty seconds

Friday evening. You turn three weeks of experiments into a twelve-page document. Method, every results table, sensitivity analysis, plus four appendix links. You send it to the CTO and copy the team. This feels like the easiest step of the three weeks. The work is done, and all that is left is writing it up.

Monday morning, the CTO replies with one line. "So, should we switch or not?"

Your first reaction is that this is unfair. The answer is on page 7, table 3, spelled out. The second reaction is the right one. He never reached page 7, and probably did not finish page 1. That is not his fault. His job is to make a decision in fifteen minutes, and reading your twelve pages was never in his job description. Your document did not answer his question. It dumped the raw material for answering it on his desk.

The academic mirror takes one sentence. A reviewer asks "why did you not control for variable X," the answer is in appendix C, and reviewers never read appendix C. Whichever page the objection rises on, the answer has to be buried on that same page. One page late counts as unwritten.

Two scenes, one structure. The quality of the evidence and the quality of the delivery are two independent variables. Three weeks of evidence can die inside thirty seconds of reading, and die silently. Nobody will tell you "your conclusion was right, I just never got to it." The delivery step loads evidence into the audience's decision loop. Writing down what you did does not finish it. Different audiences, different loop shapes. One batch of evidence should almost never have only one vehicle.

9.2 Evidence does not speak for itself

Start with what got cheap, and this time it got cheap all the way down.

First drafts got cheap. The blank page used to be delivery's first wall. Going from a claims list to a structurally complete first draft is now a matter of minutes.

Audience rewriting got cheap. This is the chapter's real new dividend. "Write another version for management" used to mean half a day. Now the same batch of evidence gets rewritten into a paper, a memo, a slide script, an email summary, in batch. "Reorder the detail by audience" went from luxury to default move.

Then the part that did not get cheap, and it decides why this step's ladder stops at assistant level. Which sentence you dare sign has not dropped a cent in price. The signature test has a plain definition. A sentence leaves your document, gets quoted on its own, and travels with your name on it. Do you stand behind it? "Proved," "looks like," "undecided," every choice among the three strength tiers is a signature. Row 1 of the transfer map reads like this at this step. The production cost of text collapsed, the review cost of claims did not.

AI also brings this step a new risk of its own. Claim strength drifts quietly upward during drafting and polishing. The shortest path to fluency is confident wording. "The results show" reads far smoother than "under condition X it looks like." Qualifiers shed a layer with every round of AI rewriting. Chapter 4 said secondhand retelling drifts. Here is the home-brewed version of it. The model did not lie. This is a legitimate side effect of optimizing wording, and it happens to land on the step where errors are silent.

So this chapter's workflow has exactly one core move. Lock the facts first, then let go of the wording.

While you are here, give the step an operational definition so it can be accepted. In a deliverable that passes, the audience arrives with their own question and hits the answer on the first screen. Any number gets challenged, and you point to its source within fifteen seconds.

9.3 Lessons stolen from PR descriptions

You write three kinds of text for one diff every day, the commit message, the PR description, the release note. Three lessons carry straight over. The only thing to watch is where they break on research delivery.

Lesson one, a vehicle can wear many faces, a fact gets only one body. The three texts differ in detail, terminology, and length, and they all point at the same diff. The diff is the single source of truth. Research's counterpart sits right there. The report, the memo, and the slides can look completely different, but every number has to point at the same results file in the same repo. The counterexample is the norm. One number in the report, another in the slides, a third in the memo, all three "roughly right." Three roughly right numbers add up to a system with zero credibility. Here is where it breaks. The diff is machine-generated, so the three texts cannot lie against it. A results file will not jump up and reconcile itself, which is why the interlock check in step 4 of section 9.4 has to be dispatched on purpose.

Lesson two, the most valuable field in a PR description is "what this change does not do." Seasoned engineers state boundaries and known defects up front, for a cold reason. A reviewer who finds a pit you did not declare costs you ten times as much, and from that moment he rereads all your code with "what else is hidden" in his eyes. Research's counterpart is limitations, the section of a paper that states its own limits. Do not treat it as a humility ritual. It is a fight over who speaks first. A defect you declare is called rigor. A defect someone else finds is called an incident.

Lesson three, once documents generate in one click, "written well" stops being evidence of "done right." Row 6 of the transfer map, lowering the bar to operate is lowering the bar to abuse. AI turned "looking rigorous" into a single generation. A fully formatted limitations section, a respectable confidence interval, modest restrained wording, all of it can decouple completely from the actual evidence. The coding side's answer is to grow documents out of the code. Interface docs come from types, changing the code changes the docs, and decoupling becomes mechanically hard. Research's counterpart, no number may be typed by hand. Every number in both vehicles is quoted from the results file, then interlocked against the other. What readers take as a signal is changing generations too. A number that can point to a repo is replacing "properly written" as the new mark of trust.

9.4 One source of truth, two vehicles

Here is the workflow you can copy straight out. Five steps, half a day.

Step 1, write the claims list, not the document. One line per claim, four fields. The claim in one sentence, the evidence pointer (file/table/commit), the honesty tier (verified / still exploring / falsified), and the signature field (dare / do not dare). There is a trick to writing claims. Write the strongest form you dare sign. Turning "seems somewhat useful" into a conclusion is cowardice. Write it up to the point where one more notch of strength would stop you from signing. On the same batch of evidence, "tie" has at least four ways to be written.

  • "Proved a tie"
  • "Looks like a tie"
  • "Undecided, but no gap shows in direction"
  • "A tie under these three premises"

Which tier you pick is the signature itself. This list is the single source of truth, and both vehicles get generated from it.

Step 2, vehicle one, the technical-report skeleton. The audience is peers and the technical committee, and their question is "how do you know." This skeleton is shaped like a workshop paper. Question, method, results, limitations, reproducibility. Anyone submitting can expand it directly. Detail tilts toward method and statistics, with criteria, arm design, and test methods given in full. Honesty layering, which is to say the tier column in the previous step's list, lands in this vehicle as the limitations section, the section that states your own limits, and it has to be written for real. Every item the Chapter 8 interrogation turned up, template concentration, the post-hoc breakdown, the overconfident confidence interval, goes in, written with numbers, not in camouflage wording like "certain limitations may exist."

Step 3, vehicle two, the one-page memo. The audience is decision makers, and their question is "what should I do." The structure is fixed. The one-sentence conclusion on top, three lines of risk right behind it, then a number table of at most five rows, one next-step recommendation, and one line of evidence pointers. Decision makers do not read the limitations section, so honesty layering lands in this vehicle as those three risk lines. The same fact changes shape across the two vehicles. The template concentration problem in the math problem set is a paragraph of statistical discussion with numbers in the report, and one line in the memo, "the math evidence is void, do not cite it." Tune the depth of the layering to the audience. Never tune its honesty.

Step 4, the interlock check. Dispatch AI to pull every number out of both documents and check them one by one against the results file. Any disagreement among the three means at least one of them is lying. This is pure mechanical work, so hand it off. The prompt is in the appendix.

Step 5, close with the signature test. Go through both documents sentence by sentence and ask of each one, if this sentence gets pulled out on its own with my name on it, do I stand behind it? A sentence you dare not sign has two roads only, lower the strength or delete it. This is the one step in the whole flow that may not be outsourced.

Charts get their own item, and both vehicles use it. Before drawing any chart, answer one question. What judgment do you want the reader to make? Generating charts and captions can be handed to AI. Two things stay with a person, the error bars and the axes. A "large lead" drawn with the y axis starting at 0.9 is the graphical cousin of wording drift. The numbers are right, the strength lies.

The AI division of labor fits in one line. First drafts, rewrites, number extraction, checking, hand them off. Step 1 and step 5, fixing the facts and signing, stay with you. The full fillable version of this chapter's templates is in the appendix.

9.5 If you are submitting, three corners of the same discipline

Most readers deliver memos and technical reports. This section is for people submitting to a venue or writing a proposal. Everyone else can skip to section 9.6.

Choosing where to submit is also an audience judgment. The shape of this batch of evidence, a mixed ending, one post-hoc overturn, decision-grade accounting, fits a workshop, the venue for trading half-finished work and negative results. It cannot carry the packaging of a main-conference claim. Upgrade the vehicle and readers automatically upgrade how they read your claim strength. Inflating a workshop skeleton into a conference paper means you completed the wording drift yourself. In an engineering setting this is the decision of who sees the study first, an internal team review or a direct report to management. The point-by-point response letter to reviewers works the same way. Let AI draft the point-by-point replies. How far each point concedes, and which comment is worth resisting, is a negotiation over claim strength and stays with a person. Review comments are also a free round of red team, which Chapter 10 opens up. The "preliminary results" section of grants and proposals is the disaster zone for strength drift, where pilot data grows into "an established method" under AI polishing. The discipline does not change. Numbers point to the repo, strength matches the evidence, and you submit only what you dare sign. That "preliminary results" paragraph in a budget request is its engineering cousin.

9.6 Spine update · writing −2.3 and 1/5 into two documents

Now I run the workflow in front of you. Chapter 8 closed owing two debts. How the sturdiest conclusion in the whole case, the code one, gets into the first line of the CTO memo, and how the math overturn gets into the technical report's limitations.

First the lazy route, having AI draft the report abstract straight from the results file. Feed it only the preregistered results table, one call, raw output stored as is in the repo at results/abstract_draft_raw.txt (that run was in Chinese, so the specimen stays Chinese and the quotes here are translated). The first few sentences of the draft are actually quite restrained. The code differences are "all not significant," cost 15%~40%, all correct. The closing sentence gives it away. "Overall, an army of open-source small models can match or surpass frontier models on structured reasoning tasks at lower cost, though its gains are clearly task-dependent…"

Every word of that sentence has a source, and not one word can be signed. "Surpass" comes from the math preregistered table, where the verdict column really does read army_ahead, and the middle sentence honestly reported "army vote leads significantly by 18.2 percentage points." But that +18 already died in the Chapter 8 audit. The model did not know. What it was handed was the preregistered table. "Match" is stamped on code's CI crossing zero. Read strictly by the preregistration, crossing zero is called "undecided," not "match." As for that handsome class name "structured reasoning tasks," it is an honorary title invented on the spot for one ambiguous-template artifact, an artifact being an illusion the measurement process manufactures itself, not a real capability gap.

AI fabricated not one number, and it even carried the "task-dependent" caveat for me. Inside the legitimate wording space, it parked the summary sentence in the most flattering corner. Strength drift raises no alarm. This is what it looks like. Not one number is wrong. What crossed the line is the nerve of the half sentence after "overall."

So the process backs up to step 1, write the claims list first. The tier column gained one tier beyond the three, outstanding, meaning the parts the preregistration promised and this round did not do.

Claim (the strongest form I dare sign) Evidence pointer Tier Signature
On the code task the gap between the army and frontier cannot be told apart within measurement precision (93.7% vs 96.0%, CI crossing zero), at 1/5 the cost the code row of report.md Verified (decision-grade) Dare
On knowledge QA the army trails by 15pp [−21.3, −8.7]; the SC (self-consistency) arm trails as well, teaming cannot rescue a weak knowledge base the mmlu row of report.md Verified Dare
Trust neither side of the math evidence. Preregistered +18.2pp, post-hoc removal of the ambiguous template −2.3pp, direction reversed; the problem set holds only ~3 template families, the CI is overconfident audit_math.py, the post-hoc audit entry in CHANGES.md Verified (for "the problem set is sick") All I dare sign is "this exam is void"
This experiment cannot detect any contribution of "teaming" over "multi-sampling" (clean subset 0.977 vs 0.983; on code SC is cheaper still) report.md + the audit Verified Dare
The self-hosted amortization basis, promised in the preregistration, never run in the repo the cost section of prereg.md Outstanding I dare sign "not done"

The last row is worth a stop. The Chapter 6 plan signed off on two cost accounts, API list price as the main one, self-hosted amortization as the secondary. The repo delivered only the first. The amortization account has to be estimated from GPU rental prices, and its slot died on budget and time. The temptation at a moment like this is silence, since nobody remembers a Chapter 6 promise. The discipline goes the other way. What you promised and did not do goes into limitations, recorded as outstanding, and not one number gets invented.

Then the two vehicles take shape. The technical-report skeleton puts the task-dependent framing "when does it pay off" straight into the title, so the abstract does not get read halfway. The abstract reports numbers under the Chapter 8 side-by-side discipline. Seven limitations, each carrying numbers. The three ugliest, the math slice's ~3 template families, the direction reversal from the post-hoc audit, and the outstanding amortization basis, happen to be the three most informative. Writing that section I noticed one thing. It is the only section in the whole document that felt steadier the longer I wrote.

The memo's first line is the claims list's first line verbatim. Three risk lines follow in order. This is a public benchmark, not your workload. The gain from "teaming" cannot be told apart from "multi-sampling on one model," so the deployment recommendation gets simpler instead, one model plus majority vote. Cost is counted at API list price only, and the self-hosted amortization account was never measured.

The "no shot" on knowledge QA trailing by 15 points goes into the judgment column of the number table. "No shot" and "evidence void" are the two most money-saving words in a memo. The second risk line's origin is worth noting. The most valuable line of advice in the memo came from the finding most damaging to the army narrative. Once the SC arm audited "team magic" away, what was left was a recommendation with half the ops, half the story, and twice the credibility. Honesty often improves the recommendation itself. It is far more than a moral pose.

Close with the interlock check. Dispatch AI to extract every number in both documents and match them against report.md. All of them matched. Catching nothing is still worth it. The value of a step is that it runs every time.

The subplot got delivered along the way. The variance-collapse red line in the text below is one of the three criteria. It means the spread of answers within one persona group is far smaller than among real people, like a room of recitation machines. The text reads, "The persona panel cannot replace real interviews for a decision-grade rehearsal. All six subgroup distributions fail the threshold, and the variance-collapse red line is tripped. 10,800 interviews, $0.47, bought one clean negative." The most common delivery failure for a negative result is burial, spreading a FAIL across ten pages of "worth further study." A negative put on top in one sentence is how a team saves itself one more $0.47. All the numbers and caveats get on the table only in Chapter 12. Here only the conclusion sentence is delivered.

Finally, the honest identity of these two finished pieces. They sit in full in this chapter's appendix, every number pointing at a specific file and commit in the smol-army repo. Nothing was submitted, no CTO ever signed off on them, and the book does not act it out. Right now they have passed only their own signature test. Chapter 10 will let them take a real beating.

9.7 Swap in your project

Dig out your most recent batch of evidence that is "done but not delivered." Budget half a day.

  1. Write the claims list. One sentence per claim (the strongest form you dare sign) + evidence pointer + honesty tier + signature field. Write what you did and what you did not, and "not done" is a line too;
  2. Pick two real audiences. One wants "how do you know" (peers, reviewers, the technical committee), one wants "what should I do" (your boss, a client, a funder);
  3. Have AI produce two drafts from the list. The technical-report skeleton + the one-page memo, with "do not change the strength of any claim" locked into the prompt (templates in the appendix);
  4. Reconcile the strength. Compare sentence by sentence against the claims list and mark every upward drift. You will usually find more than three. That is normal, not an incident;
  5. Interlock check. Dispatch AI to extract every number in both documents and match it to the source. Release only when all three agree;
  6. Signature test. Ask of each sentence "pulled out on its own, do I stand behind it," and for anything you dare not sign, lower the strength or delete it.

The acceptance criterion in one sentence. Anyone points at any number in either document and asks "where did this come from," and you point to the source file within fifteen seconds.

Want an agent to run it with you? Paste this to your AI assistant or coding agent:

Help me run the Swap in your project of Chapter 9. First I write the claims list with Template 1 of docs/appendices/ch09-templates.md, one sentence per claim
plus an evidence pointer and an honesty tier, "not done" is a line too, and you only build the table, you never write my conclusions. Once the list is final,
use the prompt set to produce two drafts from it, a technical-report skeleton and a one-page memo, with "do not change the strength of any claim" locked into
the prompt. Then I reconcile the strength, and you mark every wording stronger than the list. Then open a separate session, give it only the two documents and
the list, run the number interlock check, and list every disagreement among the three for me to rule on. I do the signature test sentence by sentence, and lower the strength or delete anything I dare not sign. If any command errors, stop and show me the output.

9.8 Sober reminders

  • One draft sent to everyone is the number one failure mode. Sending the same full text to every audience outsources the reordering of detail to the reader. Readers do not reorder. Readers just do not read.
  • Polishing is not a free operation. Every "smooth out the tone for me" reshuffles claim strength again. Fix the rule. After polishing you must rerun the strength reconciliation.
  • Limitations is not body armor. Once you have written template concentration and an overconfident CI, the abstract may not cite that CI as a victory. Honesty layering is a downgrade consistent across the whole document, and a disclaimer at the end cannot stand in for it. Using limitations to insure an overclaim in the main text is worse than writing no limitations at all.
  • Do not invent a recipient. A deliverable's credibility comes from being recomputable, and the social proof of "already adopted" does not help with that. If it was not submitted, say it was not submitted. A pilot is called a pilot.
  • Still exploring. Fully automatic tools from results to paper are iterating fast. As of this writing, my observation is that the draft structure is usable and the claim-strength calibration is not. Their default wording setting is selling, not reporting. The book's online case library tracks it.

One step remains before delivery. Kill your own conclusion before someone else does. Chapter 10, red team.

9.9 The unfair advantage you now hold

With a batch of evidence in hand you can produce a technical-report skeleton and a one-page memo in half a day, numbers interlocked, every number traceable to the repo in fifteen seconds. What most people deliver is a twelve-page document that answers nobody's question and earns a one-line "so what" on Monday morning.