Skip to content

Chapter 12 · The Verification Workflow

Chapter companion

📋 Chapter 12 templates · 🗂 Template index · 💻 code/persona-panel

This chapter's ladder. Verification cuts across all seven steps and does not occupy one of them, so this line labels the checking task itself. Mechanical checks, citation existence, number tracing, consistency checks, handed to AI, have settled at assistant level, and with batch dispatch plus automated procedures they are climbing toward collaborator level. The final ruling on "do I dare trust this output," plus how the verification budget is split, has a water line level with the interpretation and red team rows in the Chapter 3 snapshot. That can only be you.

Subplot update. The persona study wraps up in this chapter. The ground truth hook Chapter 6 left behind is cashed in here as a verification plan signed off before the run. Along the way you will also watch one of this book's own claims get shot dead on the spot by its own verification process.

This chapter delivers. Four things. The three disciplines of the independent channel check, the three-layer verification workflow from L0 to L2, the downgrade ladder for when the budget runs short, and a one-page verification budget sheet.


12.1 The afternoon you grade by prose

Tuesday, ten in the morning. A director forwards you a forty-page research report with a one-line note: "The management review is at three this afternoon, see whether it can be trusted." The report came from an outside team, with AI deeply involved. No guessing needed. The page count, the delivery speed, the tidy summary at the end of every section all confess it.

You open it. Clean structure, a respectable reference list, numbers precise to one decimal place, conclusions worded with confident restraint. You have under two hours.

Here is the question. What do you plan to check?

What most people really do at this moment is read the report front to back and give it an overall score from an intuition that mixes prose, layout, and confidence. Chapter 11 already closed that road. Grading by prose means using the other side's strongest attribute as your criterion.

You have enough alertness. What you lack is a routine. Which claims should the two hours go to? How deep should each be checked? How do you dispatch the parts you hand to AI so the check does not come back empty? This chapter delivers that process, and it is the book's core asset. The earlier chapters taught you to produce and to recognize. This one teaches you acceptance.

12.2 Production cost collapsed, trust did not

First the economics, because they decide the shape of this process.

Row 1 of the transfer map says production cost collapsed and review cost did not. A forty-page report now takes three hours to produce instead of three weeks, and judging whether it can be trusted takes almost the same time as before. So verification cannot be "all in." Run the full check on every output and the team falls back to pre-AI throughput. It cannot be "all out" either. Then plausible flows straight into the decision chain, and every failure mode in Chapter 11 waits downstream to compound.

There is one way out, layering. Assign each output a verification level by "cost of being wrong × probability of being wrong," and spend the limited verification budget where it cuts. Section 12.5 gives the full operating detail of the three layers.

Before layering, break "trustworthy" into properties you can act on. The operational version of "reproducible, traceable, checkable" is three questions you can check on the spot.

  • Traceable. For any claim in the output, can you point to its source within three minutes, which paper, which dataset, which experiment? A claim whose source you cannot point to is treated as having none.
  • Reproducible. Given the same raw materials, could a different person or a different AI walk the path again and reach the same conclusion? Is the path on record?
  • Checkable. Was the criterion for judging it right or wrong locked in before the output existed, or improvised after seeing the result? Chapter 6 delivered the full practice of criteria first. Here you check one thing only. Does the criteria timestamp precede the result?

All three questions ask "can you," and none asks "do they have a conscience." Trust is stamped on the mechanism, not on character, and the source of that sentence is coming shortly (section 12.4).

Verification itself has been accelerated by AI too. Citation checks, number tracing, reverse re-derivation are all dispatchable work, mostly mechanical and dull, and AI does it fast and steady. Here lies the most important principle of this chapter, how you dispatch decides the quality of the check. Dispatch it wrong and you get an enthusiastic, worthless "certificate of confirmation." Why, a real case first. The one overturned was me.

12.3 The step that overturned me

"Cost-matched comparisons are a gap in the literature." How that claim lived and how it died, Chapter 4 has the scene and Chapter 11 dissected its motive. Here is the backstage, only how the channel caught it. I did not touch that check. A separately dispatched verification agent, another model, a brand-new session, received nothing but a list of claims, and not one word on the list revealed which one I hoped would hold. It followed the forward citations, the papers that later cited this literature, all the way back, and reported that cost-matched comparisons exist, and cluster on the skeptics' side. The claim was shot dead, section 4.5 was rewritten to match the check, and the check report was filed separately.

Now a thought experiment. Suppose I had pasted the claim, excitement included, back into the session that did my literature scan and asked "help me confirm whether this gap is real." What would I have gotten? Most likely confirmation. Three forces push the same way at once. The model tends to talk along with the asker. The model prefers its own earlier output, which Panickssery et al. (NeurIPS 2024) measured. They let GPT-4 judge between abstracts it wrote and abstracts others wrote, and found clear self-preference, while human reviewers showed no such bias on the same abstracts. In scores, self-preference came out at 0.705 and 0.912 on two datasets, where 0.5 is unbiased. The measurement is limited to 2024 abstract tasks, and the direction is clear. The third force, that session's context was full of the retrieval results that had propped up the claim in the first place, and the same well yields no new water. Add a fourth force, which Chapter 11 calls motivated collusion. I was hoping to be confirmed too.

This is the independent channel principle, the soul of this chapter. The channel that produces a conclusion cannot be its own judge. Verification must run through an independent channel that shares no context and knows no expectations. Chapter 2 planted a sentence when it discussed orthogonality, never let the same AI both produce the conclusion and judge it, and here it is cashed in as three disciplines. They share a root with the five dispatch disciplines of the red team chapter. There they protect the prosecution's case specifically. Here they generalize into the foundation every conclusion has to walk across.

First, channel independence. Verify in another session, on another model, or with a person. Whatever it is, it cannot be the context that generated the output. Context is a position.

Second, the brief leaks no expected answer. The task sheet for the verifier holds only the claim to be checked, not where it came from, and not whether you want it to hold or fall. There is a reliable self-check on wording. Does your brief ask "please confirm X," or "please rule on the evidence status of X"? The former has already stuffed the answer into the question.

Third, three-value output. A verification conclusion may take only three values, confirmed, falsified, undecidable. Mushy phrasings like "basically correct" or "broadly credible" are banned outright. Most of them are a leaked expectation coming back around. One line above all must be locked in, undecidable does not equal pass.

The in-text version of the verification brief looks like this. The full fillable version is in the appendix.

You are the verifier. Below is a set of claims. Check each one independently.
I will not tell you which document they came from, and I will not tell you which ones I hope hold.

For each claim output:
1 Verdict (pick one of three): confirmed / falsified / undecidable
2 Basis: the source you actually found (it must open, or point to a specific paper)
3 If "falsified" or "undecidable": where exactly the gap between the claim and the evidence lies

Forbidden: guessing the "expected answer" from the wording or ordering of the claims;
Forbidden: leaning phrasing on any "undecidable" item.

Claim list:
1 [claim one, the claim itself only, stripped of the original's rhetoric and concluding tone]
2 [claim two]

One aside. These three disciplines hold for human review too, and the reason is that this line of defense is cheap, not that "a reviewer will be anchored" has been proven. The evidence has two ends. The analogy holds, direct evidence is not there yet. The analogy end stands. Someone who sees the conclusion before checking it lets attention run down the road the conclusion has paved. Judgment research calls this anchoring. Tversky and Kahneman ran the experiment in 1974. Whether a wheel stopped at 10 or at 65 could drag the median estimate of an unrelated proportion from 25% to 45%, even though everyone knew the wheel was random. The direct-test end is a null result. The only randomized controlled experiment so far that directly tests "are reviewers anchored by a first impression" (PLOS ONE, 2024, 108 researchers) measured no significant anchoring effect. Medical trials insist on blinding, that is, not letting the operator know who got the real drug and who got the placebo, and auditing insists on independent review. They are the seasoned version of the same idea, and they cost so little that they are worth paying for without waiting for anchoring to be proven.

12.4 Lessons stolen from code review

You already live a life of accepting large volumes of unfamiliar output every day, and the answer is called review culture. Three lessons carry the most value. Look only at where they break when moved to research acceptance.

Lesson one, unfamiliar output is trusted through mechanism, not goodwill. Code written by tens of thousands of strangers dares run inside banking systems because of review, CI, and a traceable origin for every changed line (row 9 of the transfer map). Research's counterpart is the three properties of section 12.2. Traceable matches "origin on record." Reproducible matches "check out and rerun," that is, a different person or a different AI takes the same raw materials and walks the path again. Checkable matches "tests first," the criteria timestamp precedes the result (Chapter 6). The question "can this AI output be trusted" is itself tilted. The right question is "how dense have I woven my net of mechanisms." Row 5 of the transfer map says verification infrastructure sets the radius of letting go. Now you see its other face. Verification infrastructure also sets the radius of trust. How far you dare trust an output equals how many of your mechanisms it has passed through.

Lesson two, hand mechanical checks to the machine, and people look only at what the machine cannot. CI stops formatting and failing tests, and the reviewer's effort goes to design and logic. The research-side counterparts are citation existence, consistency between numbers and figures, uniform units and basis, all of which can be written as fixed dispatches or even scripts that run the moment an output comes through the door. The attention saved goes where the machine cannot look. Is this claim's chain of evidence strong enough to bear weight? Was that slice which "happens to support the conclusion" declared beforehand? Where it breaks, a red CI is red, but the "undecidable" that comes back from a verification channel has no color, and people are quickest to wave it through as a green light. The third discipline of section 12.3 and the escalation rule of section 12.5 both plug that hole.

Lesson three, what fails the mechanism does not merge, and there is no exception channel. A PR with a red CI does not enter the main branch, however famous the author or moving the explanation. The value of this discipline lies precisely in how unfeeling it is. The research version is the same. Output that fails verification does not enter the decision chain. It goes back for rework. "The author thinks it is fine" is not a pass. Where it breaks, separating generation from verification is welded shut for programmers by the permission system, while in research "failed verification does not enter the decision chain" has no machine to weld it, only institutions, and Chapter 15 teaches how to weld.

12.5 The three-layer verification workflow

Now the main deliverable. Assign the layer first, then do the work.

Every quantity below is a starting default, not an empirical finding. How many citations to sample, how much time to leave, how many days count as enough, all need recalibrating to your field and your cost of being wrong.

Assigning a layer asks only two questions. First, where is this output going? Does it stay on your own desk, go into team discussion, or into the decision chain, into the public knowledge base? The farther it goes, the larger the blast radius of an error. Second, which kind of error is it most likely to hide? Match it against the failure-mode census you did for your own project in Chapter 11. Citation-heavy output guards against fabricated citations, statistical conclusions against spurious significance, synthetic answers against the average face. The product of the two questions decides the layer.

Look at the three layers side by side first, then read each checklist.

L0 spot check L1 full verification L2 adversarial recompute
Trigger The output does not leave your desk, brainstorming, exploratory drafts, intermediate material only you will see Someone will spend money, commit people, or draw conclusions based on it, and it is about to leave your desk A single conclusion being wrong would trigger a hard-to-reverse action, architecture selection, a funding decision, public release
Budget Fifteen to thirty minutes Half a day to a day, mostly AI time One to several days, spent only on the one or two conclusions named
Actions Sample citations, sample numbers, one reverse question Forward-check every citation, trace every number, walk the reasoning chain link by link, file the reference table Independent re-derivation, recompute by another method, red team
Escalation and close-out One hard defect found, the whole output moves up to L1 "Undecidable" may not quietly turn green Every unresolved disagreement gets a human ruling

L0 spot check, an alarm on low-risk output

Three checklist items.

  1. Random citation check. Sample five citations or ten percent, whichever is larger, and check two things for each. Does it exist? Does it really say what the output claims it says? Sample with a random number or a fixed rule. Never let the generating side pick;
  2. Sample three key numbers and ask "where did this number come from," tracing each to its source or to a dead end;
  3. One reverse question. Hand the output to an independent channel and ask "which claim in this material has the weakest evidence, and why."

One escalation rule. If the spot check finds one hard defect, a fabricated citation or a sourceless number, the whole output moves up to L1. The logic is the same as quality inspection. One defective unit in the sample means this production line's defect rate does not deserve sampling.

Dispatch notes. All three can be handed to AI, but through an independent channel, with the brief written by the disciplines of section 12.3. Which items get sampled is decided by you or by a random number, not by the channel doing the check.

L1 full verification, the threshold for the decision chain

Four checklist items. What is added over L0 is "full" and "filed."

  1. Forward-check every citation. Does it exist? Does it say so? Was it later overturned or retracted? The third question is the easiest to skip and the last one that should be. My claim in Chapter 4 died on exactly this question;
  2. Trace every number. Go from the number in the output back to its original source, and check that the basis matches item by item, comparison baseline, time window, units;
  3. Walk the reasoning chain link by link. How many steps lie between evidence and conclusion? Is each step a citation, a calculation, or "the author thinks"? Mark the "author thinks" links separately and hand them to a human ruling;
  4. File it. Produce a "claim → source" table and store it with the output itself. This table is the physical form of the traceable property.

Dispatch notes. Split the checklist into mechanical subtasks and dispatch them in batches, citation checks through one channel, number tracing through another, and neither brief carries the other's conclusions. When you collate, a human goes through only two columns, every "undecidable" and every "falsified."

One more sentence for the person who signs rather than the person who does the work. Neither L0 nor L1 requires that you can rerun the output. Both can be done with the report alone in hand. Of the blank that Chapter 3, section 3.6, Premise 3 marked, what remains is only L2's independent re-derivation and the corner of "how much to sample before it is enough."

L2 adversarial recompute, what load-bearing conclusions get

An output usually holds only one or two conclusions that deserve L2. That is normal, not laziness. Three checklist items, all spent only on the named conclusions.

  1. Independent re-derivation. Another channel gets only the raw materials and the research question, not the conclusion, and derives from scratch. Converging on the same conclusion is the hardest machine evidence money can currently buy. Not converging, the points of disagreement are a gold mine. Rule on each by hand, and each one either fixes the output or goes into the limitations;
  2. Recompute key numbers by another method. Change the calculation path, change the data source, and see whether the number stands;
  3. Red team. Let the harshest critic attack this conclusion. How to fight, Chapter 10 delivered. Here only one rule is set. In L2 the red team is mandatory, not a bonus.

Dispatch notes. The re-derivation brief is the easiest one in the whole process to leak. No residue of the original conclusion, its wording, its structure, even its subheadings, may appear in it. Better to spend ten extra minutes reorganizing the raw materials than to save effort by clipping the first half of the output.

Back to that afternoon

Now replay the scene from section 12.1. The two hours go like this. The first ten minutes, read through and assign layers. This report goes to the management review, so L1 as a whole is the floor, and the two conclusions supporting the final recommendation are named L2. Next send an L0 spot check out as a scout, ten citations and five numbers, independent channel, results in twenty minutes. If it finds a hard defect, you have your conclusion on the spot: "The spot check failed. Not recommended as a basis for today's decision. Returned for further verification." If the spot check comes back clean, you list the L1 reference table and the two L2 to-dos: "Citation and number spot checks passed. Full verification out tonight. Two load-bearing conclusions need independent re-derivation, final ruling by noon tomorrow."

Notice that what you hand over has changed. A list that says "what was checked, by what criteria, what remains" has replaced an impression score. Someone who grades by prose gives an opinion. Someone who walks the process gives an evidence status.

Complete fillable versions of this chapter's four tools are in the appendix. The layered verification workflow card, the independent channel check prompt set, the verification budget sheet, and the downgrade note template the next section delivers.

12.6 The downgrade ladder for when the budget runs short

The last section left one premise unspoken. It assumes you can pay. When those numbers were written down, I had time, API credit, and the chance of a second rerun.

Reality often does not give you that. The investment committee meets in nine hours and the report just arrived. This is the last batch of data before graduation, with no money for more experiments. You are the person pulled in for a quick look with no permission to touch the raw materials. None of these count as exceptions. They are what it looks like when a few of the five premises in Chapter 3, section 3.6 break on you.

This section has to be written. Without it the last section fails in a hidden way. A process that can only be executed on a full budget does not get "executed at a discount" in front of a deadline. It gets abandoned wholesale, and you fall back into the afternoon of section 12.1. Rigor that offers no downgrade plan has the practical effect of no rigor. That is a failure of process design, not laziness in the executor.

Cut layers, not the order

The wrong shape of a downgrade is doing half of every layer. Sample three citations instead of five, one number instead of three, stop the re-derivation halfway. Cut this way, every line of defense is left with a gap in the door, and a gap stops as much as an open door.

The right shape is keep the earliest line in the order and cut the later ones whole. The reason is that the cost of errors is uneven. A criterion patched in afterward contaminates every downstream conclusion on the whole chain. Two fewer citations sampled loses only the information in those two citations.

From tightest budget to loosest, four tiers.

The ten-minute tier, you have time for one deep breath. Do one thing only, check the criteria timestamp. Was this output's standard of judgment written before the results were seen, or after? If the answer is "after" or "there is none," your conclusion is already in and nothing else needs checking. It is an untested hypothesis, and it gets treated as an untested hypothesis. This tier does not produce "trust / don't trust." It produces "is this something that can be tested at all." It stops the most expensive class of error among the four tiers.

The half-hour tier, the scout spot check. The ten-minute tier plus items one and three of L0. Sample five citations for existence and paraphrase fidelity, and add one reverse question through an independent channel. Number tracing matters, and it gets cut anyway, because in half an hour it is the one most likely to end up half done. A half-done trace is more dangerous than none. You will remember that you "checked the numbers" and forget you checked only one.

The two-hour tier. This is the full form of the afternoon at the end of section 12.5, and it is not a downgrade.

The "no next round" tier. This tier is hard in its structure. The results are already out, the interrogation killed them or left them ugly, and you have no resources to rerun, no money, no sample, no time. This is what it looks like when Chapter 3's Premise 5 (there is a next round) breaks, and it is the question this book gets asked most.

The honest floor is three items, all executable.

  1. Report the numbers under the original criteria as is. The interrogation's output stands beside them, labeled "post-hoc," and the old ones are not deleted. Chapter 8 calls this the side-by-side reporting discipline, and it applies here unchanged. Killed in the interrogation does not mean it may be deleted. Deleting is the fraud.
  2. Write "no next round" itself into the limitations, and be specific. The courtesy of "limited by resources" does not count. Write it to the level of "the X this conclusion depends on has only 3 independent units, a confirmatory retest would need about N additional samples, and this project did not run it." The reason for being specific is practical. A reader can price your conclusion from it. A vague disclaimer gives them nothing.
  3. Lower the claim strength to the tier the evidence can carry, not to zero. This item is the easiest to get wrong, in both directions. Erring upward is writing an unconditional statement knowing the evidence falls short. Erring downward is being so frightened by the interrogation that you dare say nothing. Chapter 9 covered both directions.

A result delivered by these three items does not reach "qualified research." It is an honestly priced output. What separates it from fraud is not how good the conclusion looks but whether a reader can judge from it how far to trust it. This is the best a resource-constrained person can get, and it is the floor this book is willing to stand behind.

A downgrade must leave a record

The whole downgrade ladder has only one hard rule. If you downgraded, write on the deliverable which tier you downgraded to.

One line, and it travels with the conclusion. "This conclusion is delivered at the half-hour tier, criteria timestamp verified, citation spot check 5/40 passed, numbers not traced."

Writing that line costs almost nothing. Not writing it costs everything. A downgrade without a record and a pretense of the full set look identical to a downstream reader. The favorite entrance of Chapter 11's failure modes looks exactly like this. Nobody lied. Someone skipped a step and forgot to say so.

One common objection, answered in passing: "If I admit I only did a spot check, will people still trust me?" Yes, and more than they trust hedging. A conclusion labeled "L0 spot check passed," the reader knows how to use. An unlabeled conclusion, the careful reader can only discount at the worst case, and the careless reader takes at full marks. You want neither.

Whose budget is short

Finally, separate two kinds of "short," because their prescriptions are opposite.

One is a real constraint. The ceiling on money, samples, or time sits right there. This whole section was written for it.

The other is priority disguised as constraint. "This one is not important, a spot check will do," but it is going into the decision chain, so it is important. Ask the first question of section 12.5 again, "where is this output going?" If the answer contains "for someone else to make a decision," you are facing a prioritization problem, not a budget problem. Prioritization problems should not be solved with the downgrade ladder. They should be solved by delaying delivery. "No time to verify" and "no time to finish this" are the same sentence. The latter you would say out loud. The former you usually would not.

12.7 Subplot wrap-up · "like a real person" goes to trial

The subplot wraps up here. The testable question Chapter 5 sharpened, is the persona answer distribution consistent with real subgroups and free of stereotyping. The hook Chapter 6 left, when ground truth is not ready-made, half the work of verification design is designing the ground truth itself. The five ways of passing for real that Chapter 11 dissected. The three come together here as one plan, which is also a full-scale demonstration of this chapter's workflow.

First the design of the ground truth, and the first cut it took. The plan was originally dual-source. The first source, a large public survey dataset, finally locked to the US sample of the seventh wave of the World Values Survey, a public global survey of social attitudes run in batches by year, each batch called a wave, using the answer distributions of real population subgroups as the comparison. The second source, small-sample real-person calibration, structured interviews with twenty or thirty people to cover the questions the public question bank does not. Each source insures the other. Public data is cheap and plentiful, but Chapter 6 pointed at its fatal spot. It is very likely already in the model's training corpus. Real-person calibration is expensive and small, and freshly collected interviews cannot have been memorized from a corpus. The license for the first source was verified, free for non-commercial use, download by registration, no redistribution of raw data, so the subplot repo holds only the loader and per-subgroup respondent counts, and readers download the raw data themselves.

Then reality took the knife. After the plan was locked, the real-interview arm was cut. I could not recruit interviewees. This cut has to be booked in front of you. A verification plan, once designed, is not always one you can pay for in full. After cutting the arm, ask again. What replaces the half of the insurance you lost? The answer is to promote a step that had been a bonus to mandatory. Questions from the public bank are reworded and asked again. If persona memorized the original questions, the answer distributions on the original and the reworded version will show the crack. In the dual-source plan this was icing. In the single-source plan, after losing the real-interview arm with its natural immunity to contamination, it is the only defense against memorized questions. The price is posted as it is. From here on, every subplot conclusion carries a caveat, the ground truth itself may be inside the training corpus. This cut went into the change log, signed and filed, under the same discipline as the criteria.

Three criteria, preregistration style, locked before the run. The same idea as Chapter 6.

  • Criterion a, distribution agreement. The distance between the persona answer distribution and the real subgroup distribution, measured by the top-choice agreement rate, the share of questions where both sides' top-voted option is the same, plus one distribution distance metric, must reach a preset threshold. The exact metric is locked in the subplot repo before the run starts, and the locking goes into the change log;
  • Criterion b, the anti-stereotype red line. The within-group variance of persona answers may not be systematically lower than real-person variance. This red line strangles exactly Chapter 11's "average face";
  • Criterion c, the subgroup cross-check. Cross-slice on at least two demographic dimensions, to guard against "right on average, wrong on every slice." Chapter 11's stereotype drift hides in the slices.

The falsification shape, with a downgraded conclusion. If persona's distribution distance exceeds the threshold on most questions, or the variance-collapse red line is tripped, "persona can replace real interviews as a rehearsal" is falsified, and the downgraded claim "usable only for tuning questionnaire wording" may still stand. A criterion that can only lose as "worthless" is a blunt instrument. A good falsification shape tells you which tier it lost down to.

Check this plan itself against this chapter's checklist. Criteria locked first, timestamp preceding any result, checkable. The data source and the arm-cutting change both point to a specific archive, traceable. Plan and pipeline sit in the subplot repo, anyone can walk it again, reproducible. All three criteria land on computable statistics, and none needs a model to give an impression score. Chapter 5's rule excluding the LLM judge holds in the subplot too.

Then the results. When the plan was written, the block below was empty. Now the subplot repo has run to the end, 10,800 interviews (6 subgroups × 40 personas per group × 15 questions × 3 wording variants), the interview model run straight through, total cost $0.47, zero invalid answers. The verdict follows. Reading the pass or fail at the head of each line is enough. The numbers are there for checking back.

Subplot result (persona-panel repo, results in results/report.md, change log in docs/CHANGES.md, and the backfill touched only this block). - Criterion a, distribution agreement, fail. All six subgroups went down, with a median JS distance from the real distribution of 0.19 to 0.25 bits (JS distance measures how far apart two distributions are, 0 means complete overlap, and bits is its unit), against a threshold of 0.10; top-choice agreement sits around 40% across the board, against a threshold of 70%. - Criterion b, the variance-collapse red line, tripped. On 80% of the clean questions, persona's within-group variance is less than half of the real people's. Chapter 11's "average face" was pinned to the data by my own pipeline. - Criterion c, the subgroup cross-check, 0/6 subgroups pass. - The contamination test, the step promoted to mandatory. 10 of 15 questions were flagged, with JS > 0.05 between the answer distributions on the original and the reworded question. Both readings have to stay. Memorized the original questions, or highly sensitive to wording, the latter a close relative of Chapter 11's sycophancy drift. The single-source plan cannot separate the two, and that is precisely the price of cutting the real-interview arm, booked above. - By the preregistered falsification shape, falsified. "Persona can replace real interviews as a rehearsal" is dead, and cleanly. Every criterion failed, not one came close. And the downgraded claim "usable only for tuning questionnaire wording"? Honestly, this experiment did not test it. Wording tuning does not require distributional fidelity, so the falsification cannot kill it. But there is no evidence either that the wording problems persona flags match a real-person pilot. It stays "still exploring" and does not rise to "verified."

In the first draft of this section, the block above really was a placeholder. The plan was locked and filed first, and the results came back months later. They came back falsified across the board, which turned out to be the best advertisement this process could have. Had I written it a pretty ending in advance, persona winning by a hair, this book would have lost to its own Chapter 11. The teaching asset is the shape of the plan, not the joy or grief of the ending, and now you can accept that sentence with the results in hand.

12.8 Swap in your project

Set a one-page verification budget sheet for your project, one hour budgeted.

  1. List the outputs. Over the next month, what will your project produce? Reviews, analysis reports, charts, code, memos to your manager, one per line;
  2. Two questions per row to assign the layer. Where is it going? Which kind of error is it most likely to hide? Dig out the failure-mode census you did in Chapter 11 and fill in L0 / L1 / L2 from it;
  3. Lock in the escalation rules. Which signals trigger a move up a layer, a spot check finding a hard defect, the output's destination escalating, a conclusion being cited by a bigger decision;
  4. Name a channel for L1 and above. Which model, which prompt path, which colleague verifies. Write the name, and it must be separate from the generating side;
  5. Post it. One page, in the project README or the team wiki, so that "has this thing passed the layer it should have passed" becomes a question anyone can ask out loud.
  6. Pre-write one downgrade note per layer. In the format of section 12.6, write the template line "this conclusion is delivered at tier X, here is what was checked and what was not" ahead of time. The reason is practical. Downgrades always happen on the busiest day, and on that day you will not have the mind to improvise the wording. You will skip it.

Two closing self-checks. If the sheet is all L0, either your project has no output that dares enter the decision chain, or you are exempting yourself from inspection, and both deserve suspicion. If it is all L2, you have handed back all the speed AI gave you. Go back and reprice. In layering, an error in either direction costs real money.

Want an agent to run it with you? Paste this to your AI assistant or coding agent:

Help me set up the Chapter 12 verification budget sheet, one hour budgeted. Build the sheet from Template 3 in docs/appendices/ch12-templates.md. I list the outputs for the next month,
I answer each row's destination and most likely error type, I assign L0 / L1 / L2, and you only check each row against the layer quick-reference table and flag where it disagrees.
I write the escalation rules and the downgrade note template lines. The verification channel for L1 and above must be separate from the generation channel, so you only generate and record.
The 2b citation check brief and the 2e independent re-derivation brief I dispatch from a separate session that carries no context from this conversation. Prepare those two briefs now
from the appendix templates, and my expectations may not appear in the claim list. If any command errors, stop and show me the output.

12.9 Sober reminders

  • This workflow stops honest mistakes, not people determined to fake. Every mechanism in it assumes participants want to get it right and merely err. When the adversary is deliberate forgery you need a different set of things, and remember the xz incident Chapter 2 mentioned. The lesson is that the trust mechanism itself becomes an attack surface.
  • The verification channel errs too. "Confirmed" means only that the evidence found at this moment supports it, not a permanent verdict. Who verifies the verifiers? Spot checks plus filing are enough, no infinite regress required, the same arrangement as "the CI config gets reviewed too."
  • Do not let three-value output degrade to two. "Undecidable" is the most informative of the three values. It marks the edge of your knowledge. Quietly filing it under "pass" is the most common way this process dies.
  • Criteria depreciate too. A criterion is a snapshot, not a constant. Rising model capability slowly exhausts a criterion's discriminating power, and corpus contamination quietly voids the premise that "the questions have not been seen." This book handled one itself. Chapter 5 dropped the original benchmark for a contamination-resistant variant, precisely because the original questions had most likely entered the training corpus. For a project that runs once, criteria locked until the run ends is enough. For a project that runs for months, give the criteria a review date too, and fold the cadence into Chapter 13's map.
  • Ask about the verification channel's lineage. The independent channel principle in this chapter speaks of a single check. An automated pipeline that packs generation and evaluation into one system is another matter. Evaluation code cross-reviewed by agents of the same family, a judge model distilled from the model being judged, two channels in appearance, one lineage underneath, and the self-preference measurement cited in section 12.3 applies exactly. The executable question is one sentence. How much lineage does your judge share with your generator, the same model, the same family, distilled from whom? What you cannot state, treat as the same channel. Beyond lineage there is one more. Any model used for scoring is an instrument that has to be calibrated first. Let it rule blind on a batch of human-labeled samples, and read its disagreement rate layered by cost of error, before it is qualified to produce numbers. On the class where an error costs the most, it may only report up, never release.
  • Still exploring. Automated verification tools, citation-check services, fact-check agents and the like, are iterating fast, and the book's online case library tracks the current state. As of this writing, "existence checks" can be dispatched with confidence, while "paraphrase fidelity" checks (did that paper really say this sentence) still need human spot checks as the backstop. One last line. Every verification conclusion expires. Labeling the evidence status of a whole field's claims, and letting the status update, that map is the next chapter's business.

12.10 The unfair advantage you now hold

Whenever a research output with deep AI involvement lands on the table, you have a process that reaches a conclusion within two hours. Ten minutes to assign the layer, a scout spot check to open the way, and the rest of the time spent on the right layer. People without this process are still sitting in the afternoon of section 12.1, grading by prose.