Skip to content

Chapter 15 · Make It a Habit and a Capability

Chapter companion

📋 Chapter 15 templates · 🗂 Template index

This chapter's ladder. No ladder line this time. This chapter is not one of the seven steps. The ladder measures how high AI climbs inside each process step. This chapter is about the level you have already climbed to, and what keeps you there three months from now. The answer is up front, and it is not willpower.

Everyone who has tried a new workflow has felt the high. Most are back to their old ways in three weeks. This chapter assembles the loose parts of the previous fourteen chapters into a machine that does not run on enthusiasm, installed first in you alone, then spread to your team.

This chapter delivers. The three process questions card, the 30-day adoption plan, and the team metrics starter sheet, whose three metrics have to be used in pairs.


15.1 The Monday of the third week

The weekend you finished Chapter 12, you set a one-page verification budget sheet for your own project. Week one went well. The research summary for your manager passed L1, you caught one citation with paraphrase drift, and it felt good. Week two was all right, except the spot check dropped from five citations to three, that week had two deadlines. On the Monday of the third week there is an unscheduled meeting at four in the afternoon, and the AI-drafted competitor analysis has to go out before five. You read it through, thought it read well, and hit send. Not one citation was spot-checked.

You can recite the process. You will not forget it. It lives in your memory and your good intentions, and memory and good intentions lose to a calendar.

You have probably seen the team version of the same story, and it costs more. The company bought an enterprise AI tool, ran two training sessions, and the demo drew real applause. Three months later the admin console shows half the seats never logged in again. The people who do log in use it for exactly what they used a search engine for. The tool made it into the budget. The workflow did not move a millimeter.

Both endings are the same death. A new method wins on results and loses for having no institutional slot. The previous fourteen chapters handed you loose parts. Loose parts do not turn into a machine on their own, and this chapter does only the assembly. Chapter 14 settled the account at the skill layer. This chapter handles the institutional layer, how those crafts keep running on your worst days.

15.2 The smallest unit of a habit

Steal the answer from coding first, and you are the example. You did not remember to run tests by willpower. The tests sit in CI and fire on every commit. You did not do code review out of conscientiousness either. Review is a gate in front of the merge button, and what fails it does not merge. Behind this sits an observation that keeps being confirmed. Every check that matters ends up built as "you cannot get past without it," and nobody counts on "remembering to do it" any more. The verification steps of Chapter 12 stay alive in mature teams for the same reason. The step grew onto the path.

The earlier precedent is in the cockpit. What follows is a historical aside, and readers who want only the conclusion can jump to the sentence beginning "So the smallest unit of a habit." On October 30, 1935, the Boeing Model 299, the prototype of what became the B-17, crashed on a test flight at Wright Field. Nobody had released the gust lock on the elevator before takeoff. Of the five people on board two died, including Major Hill, who ran the test flight department. The crew was the best available. The accident happened because the aircraft's complexity had for the first time outrun anyone's working memory, and the pilot's checklist was born from it. The AI workflow is structurally the same. Many process steps, many pits, and you filled in that "attribute × step" table of Chapter 11 yourself. Carry it in your head and you are certain to skip a step on the Monday of the third week.

So the smallest unit of a habit is two parts, a trigger plus a checklist, and resolve is not one of them. The trigger has to be an objective event. "Every time I open a new conversation for this process step" counts. "I will be more rigorous" does not, that is a wish. The checklist has to be short, three questions at most, and section 15.9 settles the account on long checklists. Duhigg, in The Power of Habit (2012), splits a habit into a three-part loop of cue, routine, and reward, his popularization of the basal ganglia research at MIT. This chapter borrows only the first part. The place to operate on a habit is the cue, which event you weld the behavior onto, and how the behavior itself is worded turns out to be the easy part.

15.3 The three process questions, folding the first fourteen chapters into three

The content of the checklist was written across the first fourteen chapters, scattered through the unfair advantage at the end of each one. Folded up, it is three questions. Ask them once at the door of every process step. Opening a new conversation, dispatching a task, taking delivery of an output, all count as the door.

Question one, criteria. Where does this output go, what counts as passing, and what counts as losing? "Where it goes" decides the verification level, which is the first question when Chapter 12 assigns the layer. "What counts as losing" has to have an answer before you start, the immunity bought by the sign-off of Chapter 6. Start work unable to answer this question and the output will "just happen" to support the conclusion you wanted first.

Question two, delegation. What level of the ladder do I put AI on for this step, who drafts, who reviews, who answers when it's wrong? The four behavioral criteria of Chapter 3 turn here from a framework into a gate. This question forces you to renegotiate the division of labor at every process step. The inertia of the previous step does not carry over, and Chapter 3 said the level lives on "step × task."

Question three, verification. Which class is it most likely to break in, and which layer does it pass before it leaves? For the first half, open the failure-mode census sheet of Chapter 11. For the second half, use L0 / L1 / L2 from Chapter 12, spot check, full verification, adversarial recompute, and each layer up costs more. The value of this question is its timing. It gets asked at the door in, and asking it at the door out is too late. Once you know which layer the output will have to pass, your dispatch brief and the way you file things both change.

The three questions are the index to the whole book. Whichever one jams, go back to that chapter. Knowledge lives in the book, habits live on triggers. The full fillable version of the three process questions card is in this chapter's templates. Install it in your conversation template and in the first column of your dispatch brief. A checklist on the wall is decoration. Only a checklist on the path is a gate.

15.4 Thirty days, four weeks

Installing a whole machine at once is not realistic, and that is exactly the kind of plan that died in three weeks in section 15.1. The gradual version moves week by week, adding one part a week.

Week one, install one process step. Pick the one where AI is deepest in and the cost of error is moderate. Most readers will pick literature scanning or first-draft assembly. Fit it with a trigger and the three questions card, and leave every other step alone. The success criterion this week is coverage only. Every time you enter this step, the three questions get answered, however sloppily. In the first week willpower still has to front the money (you pay out of pocket first), the credit line is small, it covers one step at a time, and spreading it thin bankrupts you.

Week two, chain two. Weld this process step and its downstream verification into a pair. After every generation, run one L0 spot check from Chapter 12. What this installs is the most important structural part in the book, the separation of generation from verification. The orthogonality Chapter 2 stole from The Pragmatic Programmer becomes, here for the first time, the default structure of your own process.

Week three, run the whole workflow once. Pick a problem that is small and real, not worth a month but worth two or three days, and take one full turn around the small loop of the seven steps. Rough is allowed, and the output does not count this week. The point is to get the three questions to appear at the door of every process step at least once, and then to let you feel by hand which door makes them most awkward. Where it is awkward is where the checklist needs rewording.

Week four, retrospective and retirement. Count a few numbers. How many times the three questions got answered. What they caught, one fabricated citation, one criterion added after the fact, one output that should have passed L1 and nearly went out bare. Then do the most counterintuitive move in the whole plan. Delete or rewrite every line that has never caught anything. This is the first turn of the retirement cadence for checklists, and the full reasoning is in section 15.9.

Thirty days is only the length of four retrospective cycles, and behavioral science has no such magic number. Nor does the acceptance test at the end have to measure self-discipline, since one behavioral signal is enough. When you skip the three questions, it feels awkward. That awkwardness means the default has been swapped. From that day the cost of maintaining the workflow is paid by the institution, and your willpower goes back to the thing it is for, judgment.

15.5 From one person to a team

You probably have no authority to set rules for your team. The good news is that team adoption starts from a working example anyway, and rules come later.

The example worth copying happens to be the team version of the trust mechanism, three pieces in all.

The shared criteria library. The criteria you locked in for your project in Chapter 6 and the verification standards you set in Chapter 12, "what a passing citation check looks like," "how far number tracing goes before it stops," go into the team wiki. The next person handed a task of the same kind copies them straight and changes a few parameters. Criteria go from personal discipline to public asset. A team's radius of letting go is set by how hard the public criteria are, which is the team version of row 5 of the transfer map.

The verification budget sheet. In Chapter 12 you already set a one-page version for your own project. The team version adds one thing only, making "has this thing passed the layer it should have passed" a question anyone can ask out loud. It asks about the process, not about character.

The case status doc. A team-level registry of "facts that have happened," recording who ran what, how it came out, which conclusion has been overturned, which number has gone stale. This book runs on exactly such a document. Case facts in every chapter can come only from it, missing facts get listed as pending, and inventing them is not allowed. It has also stopped a real accident. One chapter of this book was drafted by a writing agent, the context dispatched to it still carried a claim that verification had already overturned, and three paragraphs went back for rework after delivery. The lesson is on record. Update the source of facts before you dispatch, because the executing side will not refresh the facts on its own. The single source of truth, plus "whoever updates signs," is what treats this.

The contagion path opens from there, and only one order works. First you get work done with the example. Your output comes with a "claim → source" table attached, while other people's output sends them digging through chat logs the moment it is questioned. Then someone copies it, and copying saves them time. A third person asks whether there is a template. At that point, raising "should we make this a rule" only puts a name on something that has already happened. Reverse the order and you get the enterprise tool from section 15.1 that nobody logs into.

15.6 A good-looking demo is not productivity

After a team adopts the AI workflow, the most dangerous stretch is the feel-good period. Someone demos at the weekly meeting, "I did a day's work in ten minutes with AI." Row 2 of the transfer map already ruled on signals like this. Evaluation runs on production-grade numbers only. Row 3 supplies the other half, self-perception is not trustworthy. In the controlled trial from Chapter 2, developers rated themselves about twenty percent faster and measured nearly twenty percent slower. The team version of that is "everyone feels sped up," a sentence carrying exactly as much evidence as that self-estimate did. Zero.

You have to build the production-grade numbers yourself. Start with the metrics below, and the full version of the basis and the collection method is in this chapter's templates.

  • Rework rate. The share of output with AI deeply involved that gets sent back for redoing after delivery downstream. Cheapest to collect. Add a reason label to "sent back" on the task board, and count once at month end.
  • Verification pass rate. The share that passes on the first try in spot checks and in full verification, recorded separately for citation existence, paraphrase fidelity, and number traceability. If the team set up a verification ledger the way Chapter 12 says, no separate collection is needed, the ledger is the data source. This book's own ledger is a ready-made demonstration. Its overturn rate and its correction rate are this book's verification pass rate, and the account runs as follows. About forty-eight facts pending verification are on the register, thirty-four have been verified, one core claim was overturned outright, and about ten statements were corrected or narrowed. Item-level provenance is in the experiment ledger index.
  • Claim survival rate. The share of conclusions that entered the decision chain and still stand at a scheduled review point (three months, say). It is the slowest of the three and the closest to what "productivity" actually means. What this workflow really produces is conclusions that stand, not documents.

The starter sheet looks like this, one metric per row. The "paired metric" column is what the second line of defense below asks for. Leave it blank for now, and fill it in once you have read the defenses.

Metric Data source Collection cadence Paired metric This quarter
Rework rate The "sent back" label on the task board Count once at month end Throughput ____
Verification pass rate The verification ledger of Chapter 12 With the ledger, summarized quarterly Verification coverage ____
Claim survival rate Register of conclusions entering the decision chain, plus a review point Every quarter The risk level of the conclusion ____

One warning ships with the metrics. Measures get gamed. "When a measure becomes a target, it ceases to be a good measure."

Attribution of Goodhart's law (skip on first read)

This popular phrasing comes from the anthropologist Strathern (1997), her restatement of Goodhart's law, and it gets misremembered as Goodhart's own words. Goodhart's own 1975 wording is much more of a mouthful. "Any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes." The two versions say the same thing.

Every metric has its own way of dying. Put the rework rate into performance reviews and sending things back moves into private messages, leaving the board at peace forever. Put the verification pass rate into a ranking and people submit only their safest output, routing hard problems around the register. Assess people on claim survival rate and the conclusions get more and more timid, a uniform crop of "more research is needed" deathless filler. There are three lines of defense. First, the numbers serve this group's learning and do not enter individual performance reviews. Second, metrics get used in pairs, rework rate with throughput, survival rate with the risk level of the conclusion. A single metric always has a painless cheat posture. Paired metrics bite down on each other. Third, audit the metrics themselves once a quarter, and for the one whose number improved, first ask "did things get better, or did the reporting change."

15.7 The book comes full circle, the last meter of the verification net

Three judgments made earlier get caught here.

Chapter 2 said production cost collapsed and review cost did not, and that was a historical observation. Chapter 12 broke review down into process steps, and that was a method. A process step written in a book produces no review bandwidth. It produces review bandwidth once it becomes the default action. Review bandwidth equals the number of verification steps times the probability that they get executed, and what an institution governs is that probability.

Chapter 3 said how high you can safely climb depends on how hard your criteria are written. The hardness of a criterion has to be measured by its probability of execution. A criterion executed only on the days your energy is good is worth half its nominal hardness. Only a criterion welded onto a trigger earns the right to use its nominal value when you talk about the radius of letting go.

The verification channel in Chapter 12 that overturned my "literature gap" claim had, looking back, not one link that depended on alertness or luck. A different model, a session with no history, a brief that leaks no expected answer, separate filing, every piece of it was an institutional part set up in advance. Every argument in this chapter compresses into one sentence. Give judgment to people and execution to institutions, and do not get it backwards.

Institutions expire too. Checklists, the criteria library, the map you built in Chapter 13, how all of it stays fresh in a field that changes every month, the update cadence, is Chapter 16's job.

15.8 Swap in your project

Pick one process step, install the first habit unit, and commit to two weeks.

  1. Pick the step. Open your failure-mode census sheet from Chapter 11 and pick the step where AI is deepest in, already marked on the sheet. If you do not have the sheet at hand, spend five minutes on a minimal version first, listing your process steps and the error class each one is most prone to.
  2. Write the trigger. It has to be an objective event, precise down to the action. "Every time I open a new conversation for this process step." "Every time before I paste output into a document that leaves my desk." Take an intention as your trigger and you are back to your old ways in two weeks.
  3. Copy the three questions and fill them in as yours. Write your project's specific answer after each one. Passing = what (a number, a criterion), delegation = which level of the ladder, most likely error = which cell of the census sheet. Adjectives do not count. Write nouns and numbers.
  4. Install the card on the path. The first line of the conversation template, the first column of the brief template, and the head of the document template all count as the path. A sticky note beside your monitor does not.
  5. Write day 14 on your calendar. On the retrospective day count two numbers, how many times the three questions were answered, and what got caught. The catch record decides the next move, scaling up to the full 30-day plan, or rewording the checklist first.

Two weeks and nothing caught at all leaves two possibilities. The checklist is written too loosely, or this process step was low risk to begin with. For the first, reword it. For the second, pick a more painful step and start over. The three templates of this chapter, the 30-day adoption plan, the three process questions card, and the team metrics starter sheet, are in this chapter's templates in full fillable form.

Want an agent to run it with you? Paste this to your AI assistant or coding agent:

Help me install the first habit unit of Chapter 15. I pick the process step from my Chapter 11 census sheet, I write the trigger, you only rule on whether it is an objective event,
and you send back intention triggers like "plan to" or "try to". Build the three process questions card from Template 2 of docs/appendices/ch15-templates.md. I fill in the answers
under the three questions, adjectives do not count, you take only nouns and numbers. Then install the card on the path, the first line of the conversation template, the first column
of the brief template, the head of the document template, you edit the files for me, a sticky note does not count. Last, write day 14 on my calendar. On the retrospective day count
two numbers, how many times the three questions were answered and what got caught. If two weeks caught nothing at all, lay out both possibilities and I decide which. If any command errors, stop and show me the output.

15.9 Sober reminders

  • A dead checklist is more dangerous than no checklist. Someone with no checklist at least knows they are running bare. Once ticking the box becomes the goal, the finger moves and the eye does not look, and the checklist starts manufacturing the illusion of safety in bulk. This is the shared late-stage disease of every inspection regime. The signal is concrete. One checklist line has a hundred percent pass rate for three months running and has never caught anything. Either it is internalized to the point of not needing to be written, so delete it, or it was never really executed, so reword it or change the trigger. The fix is to give the checklist a retirement cadence. Hold the retrospective monthly. Every line either produces a recent catch record or gives an explicit reason to stay, and a line with neither gets deleted. The health metric of a checklist is its catch record, and its length does not count. A test suite works the same way. A test that never fails usually means it is testing nothing, and rarely means the code is that good.
  • An institution can freeze discipline in place, and it can freeze an error in place too. One badly written criterion in the criteria library gets executed efficiently by the whole team, and a shared error travels much faster than a private one. So the criteria library itself needs an evidence status and a review record, and the arrangement of Chapter 13 applies to it unchanged.
  • The part still exploring is the team section, the practices and the measurement basis alike. In one sentence, the team piece has no standard answer yet, this chapter gives starting advice and not a settled conclusion, and below is the grade breakdown plus two organization-level measurements. CI and review culture on the coding side took more than a decade to accumulate. "How a research team shares an AI workflow" is, in 2026, still at the stage where every shop builds its own wheel. The grade of this chapter's team section gets reported honestly. Mechanisms verified on the coding side, plus a mechanism inference on the research side, plus first-hand practice from this book's own single team, do not add up to a verified conclusion on the research side. The three metrics are likewise only starting advice, and a standardized production-grade basis for research teams does not currently exist. Organization-level measurements do exist, and their conclusions are sobering. The DORA 2025 report found individual output up sharply while organizational delivery metrics stayed flat. A preregistered field experiment with 776 participants (NBER w33641) found that "an individual plus AI" is roughly equal to "a team without AI." Gains at the individual level have not automatically carried through to team output, which is exactly why this chapter exists. Use the three metrics to draw your own team's trend line. Do not compare absolute values against another team, the basis differs and the comparison is meaningless. Chapter 2, failure condition two, said the feedback loop of research is slow and the convergence of best practice will be slow with it, so do not expect a textbook next year.

15.10 The unfair advantage you now hold

On that Monday thirty days later there is still an unscheduled meeting at four in the afternoon, and the calendar still beats memory. It does not beat your templates. The three questions are printed on the first line of the conversation, and the checklist grows in front of the send button. Other people's AI workflow runs on enthusiasm, and enthusiasm has a shelf life of three weeks. Yours runs on institutions, and institutions do not care about your mood.