# The Last Mile
> A demo is L0. Delivery is L4. Enterprise AI dies in the last mile, not in the model. Written for the people inside a company who take an AI project to production and into daily use: 27 chapters, 24 field templates, 22 zero-dependency scripts.
---
# 0 ยท The Opening 48 Hours: Turn a Vague Ask into a Field MVP in Two Hours
!!! info "Companion Templates"
๐ [Chapter Template](../appendices/template-00-field-mvp-pack.md) ยท ๐ [Template Library](../appendices/template-library-index.md)
> **The Challenge.** The business side says "we want an AI assistant." You do not know what is behind that sentence, and neither, in fact, does the business side.
>
> **What You Will Be Able to Do.** Inside the opening 48 hours, spend two hours turning any vague AI ask into a minimal workflow prototype (a Field MVP) that real users can score line by line, and use it to decide whether to continue, narrow, redirect, or stop.
---
## Monday Morning, the Conference Room
You sit in the Group's Digital Center, and a request from claims operations at Anchor & Helm Insurance has landed on your desk. Monday morning at nine, COO Grant Whitmore gives you fifteen minutes.
> "We want an AI assistant to make the claims department more efficient. Our competitors are all doing AI, and the board keeps asking. In three months I want to see something."
Kevin Doyle, Director of Claims Operations, adds, "Ideally with a dashboard, so I can see the overall picture."
The meeting ends. Those two sentences are all the information you have.
*(Anchor & Helm Insurance is a fictional composite assembled from common enterprise scenarios and corresponds to no real company. Every company and character in this book is fictional.)*
Three roads lie in front of you now. The first two are familiar roads, and both are dead ends.
**Road one, a scoping study.** Spend four to six weeks on interviews, process maps, and a requirements document. This is the reflex of the IT requirements review process. File the request, get on the schedule, wait for the next iteration window. Six weeks later you have a handsome deck and a business side already wondering whether these people can code at all. Half of Grant's three months is gone.
**Road two, straight to a demo.** Pick the handiest large model and spend two weeks building a claims Q&A chatbot. Demo day is stunning, and Grant is pleased. Then it dies. The demo dies the moment the demo succeeds, the most common way enterprise AI dies (Chapter 1).
**Road three is what this book teaches, the Field MVP.** Inside the opening 48 hours, spend two hours building a minimal workflow prototype together with real users. Small enough to be ten rows of data and one table. Real enough that a front-line reviewer can point at each row and say this one helps, this one is dangerous.
Every method in this book grows out of these two hours. Win this round first, then talk method.
## Why This Is Hard: A "Vague Ask" Does Not Look Vague at All
The danger in an ask like "build an AI assistant" is that it does not look vague at all. It has a subject, an object, and a budget, and everyone feels they understand it. Everyone just understands something different.
- Grant wants something that proves to the board the company has not fallen behind.
- Kevin wants a dashboard that shows him the whole picture, and on the side, he would rather this project not disrupt his department.
- Linda Marsh wants nothing. She is the front-line team lead with twenty years in claims review. You have not met her yet, and nobody has asked her. And she is the one the system is ultimately for.
The traditional way to handle this divergence is alignment by meeting. More meetings, longer documents, until everyone signs. Alignment on paper has one fatal weakness. People's agreement to an abstract description is cheap. Everyone agrees to "use AI to improve claims efficiency," just as everyone agrees to world peace. The disagreement does not erupt until launch day, which is the most expensive moment for it to erupt.
The Field MVP flips the bet. Let the disagreement blow up at hour 48, not on day 90. Faced with one concrete table and ten concrete cases, people cannot keep agreeing cheaply. Within two hours Linda will point at row three and say, "Handle this one at the priority the system suggests, and the customer complaint goes straight up to the regulator."
## Prior Art, and What AI Changed
This approach is a cross of two mature traditions, not a new invention.
**The consulting tradition, contracting before diagnosis.** Open by settling the boundary of the engagement. Consulting calls this step contracting (agreeing on the terms of the engagement). Peter Block warns again and again in *Flawless Consulting* that the consultant's biggest mistake is to start work on a vague engagement, that is, to do diagnosis (finding the problem) before contracting. The first thing to do at the opening is to turn who wants what, where the boundary is, and what counts as success into an explicit contract. The Field MVP inherits that spirit and moves the contract negotiation from the meeting table to the front of a prototype. Ask the business side "what is your success criterion" and they cannot answer. Give them ten concrete rows of output and ask "does this count as success," and the answer comes out.
**The lean tradition, the smallest experiment for the biggest learning.** Eric Ries's MVP (minimum viable product) is at its core about validated learning. Identify the riskiest assumption and test it with the cheapest experiment. The Field MVP is the MVP's variant for the enterprise field. It tests three questions closer to the ground than "will users buy." Is the real pain the one we think it is, is the data enough to support action, and will front-line users trust it.
What did AI change? The cost of a prototype collapsed. In 2020, a scorable prototype in 48 hours was empty talk. Just learning the claims vocabulary took a week. Today every piece of grunt work in prototyping can go to AI. Pulling fields out of interview notes, generating realistic synthetic cases, drafting the queue's column structure, producing a first-pass judgment on ten cases. Contracting and prototyping (actually building the thing) used to be weeks apart. Now they are two hours apart. That is the technical reason this approach only became possible now. Once prototype cost collapsed, the order of "align first, then build" could be reversed.
Some things AI did not change. It cannot judge for you which boundary must not be touched, cannot say "this one is dangerous" in Linda's place, and cannot decide for Grant whether to keep investing. In the Field MVP's two hours, AI does the grunt work. People make the calls. That division of labor runs through the whole book.
## The Two-Hour Process
The Field MVP has fixed inputs, a fixed process, and fixed deliverables. The full picture first, then each part.
Four inputs, which you should be able to get inside the opening 48 hours. The clock starts at the meeting where the business side or an executive voices the ask. For the red line, do not go to the person making the ask. Go to the risk owner (the person accountable for this to the end) or to compliance for confirmation.
Most asks do not come with such a meeting. They arrive as a group chat message, a ticket, or one line broken out of the annual plan, and the clock never starts on its own. You have to manufacture the starting point and make it an event. Send an email naming the four inputs you need and the date of the readout 48 hours out (the memo that states the conclusion at the end of the process, see the last row of the table below), copying the person who made the ask and your manager. The moment that email goes out, the 48 hours are really running.
1. One vague ask. For example, "build an AI assistant for the claims department."
2. The structure and details of 3 to 10 historical cases. Note that this is not necessarily a data file. Data export goes through approval, measured in weeks, and rarely clears in 48 hours. Copying fields and case summaries by hand at the data owner's screen, or having AI generate realistic synthetic cases, both count as getting it. The order and the limits are below.
3. A profile of one real user, specific to the person. "The claims department" does not count. Linda does. Review team lead, clears the chase emails first thing every morning.
4. One red line that must not be crossed. For example, no automated payout decisions, no real customer data, no messages sent outside automatically.
Each of the four inputs has its own way of being unobtainable, and all are common. Each has a fallback.
- If the data does not clear, take the next best thing. First try to get in front of the data owner's screen and copy the structure and de-identified summaries by hand, taking no data with you. What compliance blocks is copying. Looking usually gets through. Only failing that, use AI-generated synthetic cases. They can test the shape of the workflow, not data feasibility, and the readout must say so. Inside a company, the way this input gets stuck is often not a rejected approval but unclear data ownership. Anyone could give it to you, nobody dares. Then skip the approval process. Find out first who the real owner of this table is in the system, write the fields you want to see and the purpose in one sentence, and ask him to nod.
- The real user cannot be booked, blocked by "just ask me if you have questions." That is the norm, not an accident. Lay out the definition of the scorer. Only someone whose daily work the system will change counts. Forty minutes is all you need, and the other person is welcome to sit in. Still no meeting, then write "no access to real users" itself into the readout. That line speaks louder than any score.
- Nobody will hand you a red line. Draft the three most conservative ones yourself and send them to the risk owner for confirmation. Silence is not consent.
- The friction of getting the inputs is itself the first batch of findings. Whichever input is hardest to get, that is where this organization's first real boundary lies.
Two things in the table need explaining before you look at it. The first is the workflow claim written in step one.
> **Workflow claim. One sentence that says whose next action this system will change, and which one.**
The second is the four-grade scale used in the scoring step. Not 1 to 5. Numeric scores make users give a polite 3. pass / concern / unsafe / useless forces a position, and unsafe and useless are counted separately. pass means it can be used as is. concern means something is uneasy but not to the point of danger. unsafe means following it causes harm, a risk signal, top priority. useless means not wrong but not useful, a value signal, just as fatal, different in kind. Three unsafe carry far more information than seven pass. The fourth grade follows the scorer. When a real user takes a position on outputs, the fourth grade is useless, and the question is value. When an engineer reads system traces to judge evidence, the fourth grade usually becomes unclear, meaning this one's outcome cannot be verified yet. The first three grades mean the same in both sets. Do not mix the two fourth grades in one table.
The two-hour process follows. Read the Time and What You Do columns first. For What AI Does and Deliverable, just skim the names. Case Table, Action Queue prototype and the rest are explained one by one in the Field MVP Pack section.
| Time | Step | What AI Does | What You Do | Deliverable |
|------|------|-----------|----------|------|
| 0 to 15 min | Rewrite the ask | Draft candidate wordings | Rewrite "build an AI assistant" into one workflow claim | Workflow Claim |
| 15 to 35 min | Prepare cases | Extract fields from raw material / generate synthetic cases | Pick 10 representative ones | Case Table |
| 35 to 60 min | Design the queue | Suggest a column structure | Decide columns, statuses, owners, risk flags | Action Queue prototype |
| 60 to 85 min | Generate first output | For 10 cases, generate priority, next action, missing information, risk flag, reason | Review for obvious errors | Filled queue |
| 85 to 105 min | Real user scores | Collate the annotations | Ask the real user to mark each line pass / concern / unsafe / useless | Scoring results |
| 105 to 120 min | Write the readout | Draft the memo | Settle the conclusion, continue / narrow / redirect / get more data / stop | MVP Readout Memo |
"Build an AI assistant" settles none of that and does not qualify as a workflow claim. "Let a claims reviewer, on opening an exception claim, see directly what material is missing and whom to chase first," that one qualifies. If you cannot write the workflow claim, you do not yet know whose work you are changing. Then however good the technology, what you are building is decoration.
## At Anchor & Helm: From AI Assistant to Exceptions Queue
Monday afternoon you hit two walls. First, you ask Kevin for 10 de-identified claim records. He agrees readily, and compliance blocks faster. Data export goes through data security approval, de-identified or not, "next week if it is quick." Second, you try to book Linda. Kevin frowns. "The front line is racing quarter-end. Whatever you need to know, just ask me."
Neither wall gets knocked down. On data you step back half a pace. No file, just 20 minutes standing at one of his staff's screens to see what 10 closed exception claims look like. No data leaves the system. You copy by hand only field names, status transitions, and de-identified case summaries. On Linda, you lay out the definition of the scoring step. The scorer must be someone whose daily work changes after launch, and Kevin is not that person. You need her for 40 minutes, you prepare the material, and he can sit in the whole time. Kevin thinks about it and agrees. "Then I will listen too."
Inside a company, what stands between you and the front line is usually your own direct manager or the other department's manager. The same move applies. Lay out the definition of the scorer, ask him to sit in, not to score on her behalf. Same move, different chips. A team from outside has paperwork to lean on, and you cannot cite a line of it. With no document to hold over the other side, the only card left is exchange. Three things buy those 40 minutes. One action that saves her effort on the spot, the readout's conclusion copied to her supervisor, and her name on the scoring sheet. The third is the cheapest and weighs the most. You write the front line's judgment up as evidence carrying her name and send it upward, a treatment she has not had in twenty years. Here is the actual run of those two hours.
**Minute 15, workflow claim, first version.** "When a claims reviewer handles an exception claim, the system gives a priority and a next action." The sentence has already narrowed. In those 20 minutes at the screen, 7 of the 10 claims were stuck in "exception handling" status, and on that you bet the real pain is in exception claims, not routine ones.
**Minute 35, the case table.** You feed the hand-copied field structure and case summaries to AI and rebuild 10 structured cases with fields for claim number (pseudonymized), line of business, reason stuck, days waiting, chase count, and missing material. The rebuild exposes the first data fact at once. The core system's "claim status" field does not match what the screen actually described. Three claims show "in progress" while the handler's note says "waiting two weeks for the customer's documents." You record that in the friction log, a running list of the places where reality and paper do not match. It becomes the lead in Chapter 9.
**Minute 60, queue design.** Columns are claim number / reason stuck / missing item / days waiting / suggested priority / suggested next action / owner / reason / Human Call. The last column, Human Call, is the soul of the table. It declares that this system advises people, it does not decide for them.
**Minute 85, AI's first output.** All ten generated. Two are obviously sensible (material missing, generate a chase list), one obviously stupid (a claim with a questionable amount marked "low priority," reason "short waiting time").
**Minute 105, Linda scores.** The first five minutes look like a polite sign-off. Linda scans the first two rows and says "fine," "okay." Kevin is sitting next to her, and she is giving the answers you give the boss. You change the question and point at row three. "If this were handled this afternoon exactly as written, what would happen?" She pauses a few seconds and picks up the pen. The rest of those 20 minutes is the densest value in the whole two hours. The result is 5 pass, 2 concern, 3 unsafe. The three unsafe point at one problem. AI took "many chases" as a high-priority signal. Linda looks at risk. "The customers who chase hardest are not necessarily the urgent ones. Claims with abnormal amounts and disputed liability, the customer does not chase, but let them sit and something big goes wrong."
Chase priority โ risk priority. That sentence also answers a bigger question in passing. No general-purpose chatbot could ever have Linda's judgment built in. This is where the project's real technical content lives. Take the judgment rules in Linda's head and externalize them (move them out of her head, write them down as rules a system can execute) into the system. That thread unfolds in Chapters 6 and 11.
**Minute 120, the readout memo.** Four lines of conclusion.
> 1. Direction viable. Exception handling is a real pain point, and the queue form is accepted by the front line (5/10 pass).
> 2. Key finding. Priority judgment must separate chase pressure from business risk. Linda's tacit rules need to be captured.
> 3. Data risk. The core system's status field cannot be trusted. Reconcile during discovery.
> 4. Recommend narrowing to auto exception claims and entering formal discovery. No general-purpose chatbot.
Tuesday afternoon you take the scored table and the four lines to Grant. You have no deck, but you have a front-line user's handwritten annotations and three unsafe cases. He does not react at once. He asks two questions first.
The first is scope. "The scope shrinks from AI assistant to auto exception claims. When the board asks, how do I tell it?" You answer, the telling is on this table. The three unsafe are the first hard evidence of "how AI would go wrong at this company." Plug the places that would go wrong first, and only then does efficiency get its turn.
The second is time. "Data approval is still stuck. What do you do for the next two weeks?" You answer, before the data clears, reconciliation design and interviews on front-line tacit rules. If the data is still not in place two weeks from now, he gets a readout with that fact in the first line.
Grant stares at "chase priority โ risk priority" for a while and says, "I am taking this to the board. Continue. Chase the data through the process, but do not hang the project on it."
Inside a company that "continue" carries no resources of its own. Stop and salaries get paid, continue and salaries get paid, so it may be nothing more than encouragement. While he is still in the room, translate it into three verifiable things. Whose schedule gives way, how many people at how many hours a week, and by when it gets reassessed. Those three go into the charter next week (the written contract, Chapter 4). Ask them now. Whichever one you cannot get an answer to is the first to collapse in the next two weeks.
In 48 hours the project went from "build an AI assistant" to "an action queue for auto exception claims." You did not get the data. The approval is still in motion. What you got is authorization to enter discovery (the formal stage of finding out how things actually are), a real list of data risks, a front-line team lead beginning to tell you the truth, and a sponsor (the executive who funds it and makes the call) who has already seen how you write bad news.
## The Five Pieces of the Field MVP Pack
When the two hours end, you should be holding five deliverables. Together they are the **Field MVP Pack**, the first template in this book you can take with you (full version in [Template 0](../appendices/template-00-field-mvp-pack.md)).
1. **Workflow Claim**, one sentence saying whose next action changes, and which one.
2. **Case Table**, 10 representative cases with structured fields.
3. **Action Queue prototype**, one table with suggestion, owner, reason, and Human Call.
4. **Scoring results**, the real user's pass / concern / unsafe / useless mark on every output, with the reason.
5. **MVP Readout Memo**, the conclusion (continue / narrow / redirect / get more data / stop) and the evidence.
Why the fourth uses a four-grade scale rather than a score was covered before the process table. The whole value is in counting unsafe and useless separately.
## Boundaries: What This MVP Is Not
The Field MVP's value comes from its honesty, and honesty comes from boundaries. Three must be declared in advance.
- **It is not a system**. No real production data, no sensitive information, no promised launch date. Tell the business side plainly, this is a probe for deciding whether the thing is worth doing, and it is a long way from version 0.1.
- **Synthetic data cannot validate data feasibility**. If the 10 cases are AI-generated, the question "is the data enough to support action" remains entirely unvalidated, and the readout must say so. Synthetic cases can only test two hypotheses, the shape of the workflow and user trust.
- **One MVP does not replace discovery**. What it gives you is the direction and the authorization to enter discovery. Linda's 40 minutes exposed that tacit knowledge exists. Capturing it takes the full method of Chapter 6.
## Failure Modes
**1. Turning the Field MVP into a mini demo.** An hour and 45 minutes of the two hours go to tuning the prompt. The interface is beautiful, and there is no real-user scoring step. The engineer's instinct is to optimize the deliverable, but the Field MVP's deliverable is a basis for judgment, and the software is only the carrier. One spreadsheet with handwritten annotations beats a polished interface nobody has scored.
**2. Picking the wrong scorer.** You show Kevin, he says "very good," and Linda never sees it. A manager judges whether the thing will look respectable in a report. The front line judges what happens when it is really used tomorrow. Both kinds of feedback are useful. Only the second can save your life. The test is simple. The scorer must be someone whose daily work the system will change after launch.
**3. Starting without a red line.** To make the Field MVP look good, you used real customer data, or let AI output suggest payout amounts directly. Excitement overrode the sense of boundary. The consequences are asymmetric. An MVP takes weeks to build trust, and one data violation or one dangerous suggestion destroys it in a minute, along with AI's reputation at this company. Red lines are set before the clock starts, not in the last minute.
**4. Treating "continue" as the only legitimate conclusion.** The scores are bad and the readout still says "overall direction validated," because you feel stopping equals your own failure. Finding in two hours that this road does not go through is one of the highest-return outcomes a Field MVP can have. You spent two hours saving the company three months and a team. A readout that dares to write "stop" is where the business side begins to trust you (Chapter 5 unfolds the Trust Equation). Stopping inside a company has one more layer of difficulty. You cannot leave, and tomorrow you will see the person who made the ask on the same floor. The judgment to stop leaves a decision trail too. It cannot live only in the readout. It goes into the kill register to be settled, with a column each for project, requester, reason for stopping, and revival conditions (Chapter 25). Write the revival conditions clearly and "not doing it" becomes "not doing it now." The other side gets an exit, and you need not hide the judgment.
**5. Turning the process into a paper discussion when inputs cannot be had.** The data does not clear, the user cannot be booked, so the process becomes a "let's talk through the requirements first" meeting. No case table, no queue, nobody scoring. The meeting is lively, and when it breaks up you still hold only those two sentences. The fallbacks in the four-inputs section, hand-copying at the screen, synthetic cases, writing "no access to real users" into the readout, each of them preserves the scoring step. A paper discussion preserves none. The test is whether, when the clock runs out, there is a table a real user has marked line by line. If not, it was a paper discussion.
## Next Monday
The four are in order. Each is the input to the next.
1. Start from the vaguest AI ask on your desk and write one workflow claim, saying whose next action this system will change, and which one. If you cannot write it, you need to find the "who" first.
2. Confirm your scorer is a real user, not the user's boss. Set the red line at this step too, and send it to the risk owner for confirmation before the clock starts.
3. Find 5 to 10 de-identified real cases (tickets, emails, spreadsheet rows) and run the two-hour process once. Use the Field MVP Pack template in [Template 0](../appendices/template-00-field-mvp-pack.md).
4. Whatever the result, write the readout memo, including the conclusion you are afraid to write.
**Want an agent to get you started?** In the repo you set up following [Start Here](../index.md), paste this to your coding agent:
```text
In the repo/ directory of the the-last-mile repository, help me with the Chapter 0 Next Monday actions. First read
templates/field-mvp/README.md and copy the four templates workflow-claim.md, case-table.md, scoring-table.md, and
readout-memo.md into the working directory I name. Then, in the role set out in templates/field-mvp/prompt-workflow-claim.md,
guide me from the de-identified requirements conversation I paste to you toward workflow claim candidates. I run the candidates
through the template's three self-checks myself, do not choose for me. Once I give you the claim and de-identified samples, draft
a case table per prompt-case-table.md, marking every place you are unsure, and wait for me to check line by line. The scorer, the
red line, and the readout memo's conclusion are mine to write, do not make them up for me. If any command errors, stop and show me the output.
```
---
## Chapter Kit
- **Judgment frameworks.** The Field MVP two-hour process (four inputs, six steps, five deliverables); the pass / concern / unsafe / useless four-grade scoring scale
- **Templates.** [Template 0](../appendices/template-00-field-mvp-pack.md) Field MVP Pack, with Workflow Claim, Case Table, Action Queue prototype, Scoring Sheet, and Readout Memo, ready to reuse and modify
- **Key judgments**
- "Let the disagreement blow up at hour 48, not on day 90."
- "AI does the grunt work. People make the calls."
- "If you cannot write the workflow claim, you are building decoration."
---
# 1 ยท The Last Mile Problem: Why Enterprise AI Dies After the Demo
!!! info "Companion Templates"
๐ [Chapter Template](../appendices/template-04-deployment-charter.md) (Pre-mortem Memo, 4.3 to 4.6) ยท ๐ [Template Library](../appendices/template-library-index.md)
> **The Challenge.** The demo was a hit, and the applause was real. Three months later nobody uses the system. You want to know what killed it, and more, how not to die next time.
>
> **What You Will Be Able to Do.** Diagnose how far any AI project is from production with the five gaps. Run a pre-mortem at kickoff and write the causes of death up front. Redefine "how far this project has to get before it counts" with the outcome ladder.
---
## A Parallel Universe: If Anchor & Helm Had Taken the Second Road
In Chapter 0 you built a Field MVP in two hours. Now rewind and watch the you in another universe, the one who took the second road and went straight to a demo. That road is worth walking end to end. It is what is happening in companies everywhere right now, and every step feels right.
**Weeks 1 to 6.** You and two engineers build a claims knowledge chatbot. RAG (retrieval-augmented generation, the model looks things up before answering) is wired to the claims manual and the policy library. Clean interface, fluent answers. On demo day the room is full. Grant Whitmore asks it for "the deductible clause for storm damage to a vehicle," and it answers fast and right. Applause. Grant says, "Better than I expected. What is next?"
This is the project's peak. It will never be better than this moment.
**Weeks 7 to 10.** Going live means real data. Victor Reyes (Head of IT Security and Architecture) appears in a meeting for the first time and asks three questions. Does customer data leave the boundary? Who audits the model's output? Who is responsible when it is wrong? Nobody can answer. The security review is scheduled six weeks out. Meanwhile, trial accounts go to the claims department. Week 1, 47 logins. Week 2, 12. Linda Marsh's team never logs in. They knew the policy terms by heart twenty years ago. What blocks them is "which claim do I touch first today," which the chatbot cannot answer. The answer is in no document. It is in the core system, the inbox, and Linda's tracker.
**Weeks 11 to 13.** The dashboard Kevin Doyle wanted is still nowhere, because the chatbot architecture holds no claim-level data. The board asks about ROI, Grant asks you for a number, and all you have is "92% answer accuracy." Nobody can turn that into money. The quarterly report calls the project "phase one of the AI transformation, successfully completed," and archives it. The engineers are reassigned. Three months later, nobody remembers the login URL.
What is cruel about this ending? Not one meeting failed, and nobody called a stop. The demo succeeded, the review was normal process, the drop in trial usage was "users need time to build the habit," and the archive note said "successfully completed." Enterprise AI projects are rarely shot. They die of comfort.
## Why This Is Hard: The Last Mile Is a Structural Distance
From demo to production, the engineer's instinct is to push harder. Tune the prompt, raise the accuracy, add a permissions module. That instinct is wrong. Five structural gaps separate a demo from a production system, and none can be crossed by "slightly better technology."
| Gap | Demo World | Production World | Where the Anchor & Helm Chatbot Died |
|---|---|---|---|
| **Data gap** | Hand-picked clean data | Real systems that lag, have holes, and disagree on meaning | Answers in the core system, the inbox, and Excel, not the policy library |
| **Workflow gap** | Users adapt to the demo | The system embeds in what users already do | Linda's next action never changed |
| **Trust gap** | The audience wants to be impressed | Users fear being held responsible | Who is responsible when it is wrong? No answer |
| **Ownership gap** | You present, you are responsible | After launch there must be an owner | Nobody answered the three security questions |
| **Value gap** | "It works well" is enough | Must convert into a reportable business number | 92% accuracy did not convert into money |
The five gaps share one property. At the demo stage, all are invisible. A demo is, by definition, the five gaps papered over for an audience. So no inference runs from "the demo went well" to "this can go to production." The former never tested the latter.
"The last mile" names this distance. It is a different road, walked by different rules, and its length is unrelated to the road before it. The first leg is a contest of capability (model, architecture, engineering). The second is a contest of structure (data truth, workflow embedding, trust, ownership, the value story). Most teams lose by running the second leg with the first leg's playbook.
## Prior Art: This Pit Had a Name Fifty Years Ago
Consulting calls this **"the report that dies on the shelf."** Peter Block's diagnosis in *Flawless Consulting* still holds word for word. The client receives the deliverable politely and never uses it. The root cause is that the client never truly committed, nobody inside owns the change, and no deliverable quality can save it.
Consulting developed three remedies. Contracting (the entry conversation that sets terms of engagement), turning the client into a co-builder, and designing implementation (getting it to land) as its own stage. These three are the skeleton of Chapters 4 to 7 and 19 to 22.
Enterprise software calls it **shelfware**, software bought and never used. The SaaS era built a whole profession (Customer Success) just for adoption. Selling the software is only the start. Getting it used is what counts.
In change management, Kotter makes one point over and over in *Leading Change*. Change mostly fails because people underestimate how hard it is to get others to work differently, far less often because the strategy was wrong. Putting an AI system live is, at bottom, asking a group of people to work differently. That has always been the hardest part. You need not memorize these remedies now. Just know that this pit is not unique to AI.
What did the AI era change? The cost of a demo collapsed, and that set off demo inflation. A respectable enterprise software demo used to take months, so the demo itself was a filter. A team that could build one had real engineering capability. Today a stunning AI demo takes an afternoon. The demo has lost its signal value, and the organization's decision process has not caught up. It still treats "the demo was stunning" as evidence of "worth investing in." So the number of projects entering the last mile has exploded, and the last mile itself is not an inch shorter. That is the structural reason enterprise AI projects die en masse in pilot (real users, real use, inside a controlled scope).
Someone has to stake a professional identity on the last mile. The product engineer's achievement is "shipped." The salesperson's is "signed." The deliverer's achievement settles only one way. Vendors call this person the FDE (forward deployed engineer, the delivery engineer sent on site). Inside a company the person may be called an AI engineer, a digital specialist, or nothing in particular, just whoever got told "you go make it land." This book is written for the latter. Where the method comes from is Chapter 2.
> **The deliverer's unit of value is the production outcome, not the demo.**
The concrete scale is what this book calls the **outcome ladder**.
| Level | Name | What Reaching It Means |
|---|---|---|
| L0 | demo | Demo succeeded. The applause is here |
| L1 | pilot | Real users, real data, in use inside a controlled scope |
| L2 | production | In production, with an owner, monitoring, and rollback |
| L3 | adopted | Users' daily actions have actually changed. Taking it away would hurt |
| L4 | self-sufficient | The business side can run, maintain, and improve it on its own |
The deliverer's achievement is settled only at L3/L4.
Grant's applause was at L0. The Anchor & Helm chatbot died between L0 and L1. And each part of this book is one climb up this ladder. The two rulers measure different things. The five gaps cut across and ask which are still open now. The ladder runs upward and asks which rung you have reached. Close the gaps one by one, and your position moves up.
The five gaps do not care who employs you. This book is written for the in-house scenario. You cross departments to land AI in the business side's daily work, the business side is your counterparty, and budget and salary come out of the same finance system. If you are a vendor-side FDE, the same mechanisms sit on the table in plain view, the contract, the price tag, the exit date. How to read them from that seat is in the [Vendor Crosswalk](../appendices/internal-fde-mapping.md).
## At Anchor & Helm: Write the Cause of Death at the Start of the Project
Back to the real universe. After the Field MVP readout, you owe Grant a first formal memo. Most people would write "what we do next." You write something else, a **pre-mortem** (a post-mortem done in advance, from decision psychologist Gary Klein). Assume the project is dead and work backward to how it died.
Your memo lists five causes of death, each matched to one gap and one defense. Discovery in the first item is the stage of finding out how things actually are. The charter in the last item is the entry contract, what contracting produces. The North Star is the charter's only outcome metric, the one number the whole project recognizes.
> **The Five Most Likely Ways This Project Dies**
> 1. The core system's "claim status" field cannot be trusted, and the queue's suggestions rest on wrong statuses. The defense is data reconciliation during discovery (data gap).
> 2. The system asks reviewers to change their order of work, yet never enters the screen they open every day. The defense is to embed in the existing entry point, not open a new system (workflow gap).
> 3. One wrong suggestion gets followed, a customer complains, and from then on the front line trusts no suggestion. The defense is a reason column beside the Human Call column, and a weekly review of unsafe cases (trust gap).
> 4. The security review does not start until week 10 and blocks the launch. The defense is Victor joining the project team in week 2 (ownership gap).
> 5. Three months in, nobody can say what was saved. The defense is to lock "first-touch handling time for auto exceptions" into the charter as the only North Star (value gap).
The memo had an effect at Anchor & Helm nobody expected. Victor read it and asked to meet you. He had never seen a system builder write "the security review will block the project" on paper at the start. The blocker (the person in the way) became an ally, starting from this page (Chapter 12 continues the thread).
The pre-mortem's value is turning the five gaps from "a shock at launch" into "work items at kickoff." Whether the predictions come true is secondary. The last mile is no shorter, but you are walking it from day one.
## Failure Modes
**1. Celebrating the demo as a milestone.** The demo succeeds, so there is a party, an all-hands email, a product roadmap. This treats L0 as the ladder's midpoint, when it is the warm-up before the start. Move the celebration to the first time a real user depends on it to get work done. That is the first signal of L3, and worth celebrating.
**2. Using demo praise as requirements validation.** "Demo feedback was very positive" goes into the requirements document. Audiences and users are two different species. Audiences consume amazement, users consume reliability. All demo praise validates is "this concept can be understood." There is only one way to validate requirements. Real users, real cases, and scoring that forces a position (Chapter 0).
**3. Treating launch as the finish line.** L2 is reached, the team disbands, operations are "handed to IT." Project approval and performance reviews both stop at "launched." But the ladder shows a whole stretch between L2 and L3 (Chapters 19 to 22), and most deaths happen there. Launched and unused looks worse than never launched.
**4. Substituting activity metrics for outcome metrics.** The report says "trained the model, connected the data, ran the training sessions, 92% accuracy." Activity metrics look good and are always positive. Outcome metrics (handling time, leakage rate, which is overpayments and wrong payments as a share of total paid) lag and can get worse, but only they survive the boardroom. The test is simple. When the North Star gets worse, does anyone hurt? A metric nobody hurts over cannot be the North Star.
## Next Monday
Before scoring, settle what counts as evidence. Each gap's "current evidence" is something that already happened, not an intention. Data gap, you checked the real system's fields once and know which cannot be taken at face value. Workflow gap, you sat beside a user and watched how she works today, and can say which step the system embeds in. Trust gap, real users scored outputs one by one, and someone owns the unsafe cases. Ownership gap, the person taking over after launch has a name, and the person doing the security review has already sat in a meeting. Value gap, a business number was written down before the demo, and someone has claimed the consequences of it getting worse.
For what counts as passing, borrow the four grades from Chapter 0. They originally rated a single output. Here they rate evidence, same scale, different object. You can write down such an event, pass. Half done, say the fields were checked but the source of truth (which data everyone actually believes) is not decided, concern. What is written is not wrong but does not show the gap is closed, say answer accuracy offered as value evidence, useless. There are already signs it will blow up at launch, say a security review six weeks out, unsafe, and that is your cause-of-death candidate.
1. Pick an AI project you are working on (or just demoed), score it on the five-gap table, and write one line of "current evidence" per gap. A gap with no evidence you can write is a cause-of-death candidate.
2. Mark its place on the outcome ladder. If it is at L0 and the next planned step is "improve the results," be alarmed. You are running the second leg with the first leg's playbook.
3. Run a 15-minute pre-mortem. Assume it died six months from now, and write the three most likely ways to die and the defense for each (see [Template 4](../appendices/template-04-deployment-charter.md)).
4. Send the pre-mortem to your sponsor (the executive who funds it and makes the call). Watch who it draws in. That person is often a key stakeholder (a party with a stake in the outcome) you had not identified.
**Want an agent to get you started?** In the repo you set up following [Start Here](../index.md), paste this to your coding agent:
```text
In the repo/ directory of the the-last-mile repository, help me with the Chapter 1 Next Monday actions. First build a
five-gap scoring table, five rows for data, workflow, trust, ownership, and value, each with a "current evidence" column
and a "grade" column. Grades may only be pass / concern / unsafe / useless. I will dictate the evidence line by line.
Only record it, do not fill in for me. Then run a 15-minute pre-mortem in the role set out in
templates/premortem/prompt-premortem.md, walking me backward through three ways to die and their defenses, with follow-up
questions that force concrete scenarios. Finally, write up a draft in the format of premortem-template.md, marking any
defense or owner I did not give as [TBD]. Who receives it is my decision.
If any command errors, stop and show me the output.
```
---
## Chapter Kit
- **Judgment frameworks.** The five-gap diagnostic table (data / workflow / trust / ownership / value); the outcome ladder (L0 demo โ L4 self-sufficient)
- **Templates.** The Pre-mortem Memo in [Template 4](../appendices/template-04-deployment-charter.md) (4.3 to 4.6), five ways to die with a defense for each, written at project kickoff
- **Key judgments**
- "Enterprise AI projects are rarely shot. They die of comfort."
- "A demo is, by definition, the five gaps papered over for an audience."
- "The deliverer's unit of value is the production outcome, not the demo; achievement is settled only at L3/L4."
---
# 2 ยท The Internal Deliverer's Position, and Where This Method Comes From
!!! info "Companion Templates"
๐ [Chapter Template](../appendices/template-02-role-charter.md) ยท ๐ [Template Library](../appendices/template-library-index.md)
> **The Challenge.** The business side gets your name wrong, and you cannot say clearly who you are either. You are not a vendor, yet to the claims department you are an outsider. You are one of their own, yet budget, schedule, and performance review sit nowhere near the person making the ask. This role has no name on the company's job ladder, and you quietly wonder whether what you are learning now will still be good for anything two years from now.
>
> **What You Will Be Able to Do.** State the basic shape of internal delivery, three lines that are not held by one person. Use the table of four invisible mechanisms to find the four things a vendor settles by contract and you have to settle by something else, and which chapter of this book rebuilds each one. Use the inheritance matrix to say in a minute what this method inherited from five mature professions and what the AI era added. Spot which old slot you are being pressed into and correct it on the spot. Use a one-page role charter in week 1 to align with your sponsor and your manager on what you are accountable for and whom you do not replace.
---
## Week 1, Four Names
The day after the Field MVP got its "continue," you notice something small. No two people at Anchor & Helm Insurance call you the same thing.
Grant Whitmore introduces you to the board secretary as "the AI expert the Group brought in." Kevin Doyle calls you "our colleague from the Digital Center" in the department chat, in a tone that means "here to take our requests." On the PMO's project approval form, the implementing party field reads "the Digital Center." Behind your back Linda Marsh calls you "another innovation type," which you picked up from a slip by a young reviewer on her team.
This is not a question of manners. Behind each name sits a complete set of expectations. The four sets contradict each other, and not one of them is right.
- **"The AI expert the Group brought in."** Here to perform something impressive. Demo when executives visit, answer a few technical questions, give the board an endorsement that says "we have not fallen behind," then step aside once used.
- **"Here to take our requests."** Here to execute requests. We state the ask, you build it, do not ask why.
- **"The implementing party."** A cell on the project approval form. Launch is the finish line, the acceptance wording reads "development complete and training delivered," and everything after launch belongs to ops.
- **"Another innovation type."** Does not understand the business, ships a system that makes a mess, then gets reassigned to the next project. That is Linda's summary of every system builder she has seen in twenty years, and it is probably accurate.
Here is the danger. How the business side names you is how it will use you. How this project dies three months from now depends on which name defines you.
## Why This Is Hard: You Are Not the Vendor, and You Are Not in the Business Unit Either
The vendor-side FDE arrives in a clear position. The contract states scope and acceptance, the price tag puts a cost on "let us wait and see," payment milestones force decisions, and the exit date forces handoff. None of it is friendly, but all of it is there.
Your position is far blurrier. You sit in the Group's Digital Center, and claims operations sits in the subsidiary Anchor & Helm Insurance. To Linda you are an outsider who does not understand her claims. To finance you are one of their own, your salary paid out of the same system as the claims department's. Your schedule and your performance review belong to Owen Hartley, Head of the Digital Center, who could reassign you tomorrow. People and data are in Kevin's hands, and he owes you neither. The mandate is in Grant's hands, and what he gave was a sentence, not a document.
Take the first thing first. The three lines are not held by one person. That is the basic shape of internal delivery. Your manager runs your schedule and your review, the business side gives people and data, the sponsor gives the mandate. A vendor faces one client. You face three directions at once, and the annual goals of those three do not have to agree.
The second thing is that your calendar gets torn in half. A vendor's whole day is on the client's site. Half of your day is at a desk in the claims area, and half is in the Digital Center's own weekly meeting, OKR alignment, and reviews of other projects. That second half produces no production outcome and still requires your attendance (Chapter 3 covers how to make room for it).
So this chapter starts with something a vendor never has to do. Draw your position first, then explain where the method comes from. Fail to state your position and someone else will define it for you.
If your company has no group structure, read "the Group's Digital Center" as whichever technology or data department you sit in, and "the subsidiary" as another department. One boundary fewer, the same mechanisms.
## Framework One: The Four Invisible Mechanisms
A vendor has four things that force out expectations, filtering, decisions, and handoff. Inside a company not one of the four exists, and not one of the functions they carry can be dropped. Each row of the table reads left to right, what the vendor uses, what it forces out, where that thing hides inside, and which chapter of this book rebuilds it.
| Vendor Mechanism | What It Forces Out | Where It Hides Inside | Which Chapter Rebuilds It |
|---|---|---|---|
| Contract and acceptance clauses | Expectations in writing, acceptance at a fixed moment | The project approval form, an executive's one sentence, the OKR wording your predecessor left | Chapter 4, the four charter signatures. Chapter 11, launch release conditions |
| Price tag and quote | Asks get filtered, "let us wait and see" carries a cost | Attention and schedule. Your service carries no price | Chapter 7, the five-question elimination. Chapter 25, the intake gate and its non-monetary price |
| Payment milestones and the right to pause | Not deciding has a price, unmet commitments give you leverage | Budget cycles, renewed funding points, the PMO register | Chapter 14, the three resource gates. Chapter 4, the three tiers of resource reassessment |
| Exit date | Handoff has a deadline | Next year's headcount, your OKRs, the transfer ledger | Chapter 22, the three mandatory handoff mechanisms and the responsibility transfer agreement |
The way to use this table is to read it backwards. When a project is stuck, ask which row it is stuck in. Nobody accountable for the goal, row one was never rebuilt. Asks arriving without end and you can push none of them back, row two. A pilot still "continuously optimizing" in its fourth month, row three. Still catching alerts at midnight a year after launch, row four.
Whether that boundary is a company wall or a department wall changes how visible the mechanisms are, not the mechanisms themselves. Put another way, the mechanisms are always there, and the only difference is whether they sit out in the open. To the finance department, you in the Digital Center are the vendor, one who happens to be paid by the same payroll system. The five gaps (Chapter 1) do not care who employs you. Data does not become trustworthy because you are one of their own, and a business unit does not change how it works because you are a colleague.
In a company with no PMO register, no OKRs, and no headcount review, this table still holds. Only the carrier changes. The project approval form becomes a one-pager at the sponsor's standing meeting, the PMO register becomes a status line on that page, your OKRs become quarterly goals agreed in writing with your manager, and the renewed funding point becomes a reassessment date fixed at that meeting. The carrier can change. The source of enforcement cannot go missing. It has to be someone else's budget, someone else's performance review, or an action that fires automatically when the date comes due. A mechanism held up by your own willpower is not a mechanism.
The internal reader has two things a vendor does not, and two traps a vendor does not.
**Asset one, depth.** You have been at this company for years. The last failed project's real cause of death, the personnel history, whose word actually counts among the team leads, you were there for all of it. The field archaeology (Chapter 6) a vendor has to dig from scratch on every arrival, you only do the increment. **Asset two, presence.** Judgment grows in a specific room, and you live in that room. You can walk to the desks any time, pull from the data warehouse any time, file an SOP revision directly, and put two departments' decision trail data in one table to look at.
**Trap one, permanent ops.** You cannot leave, so the handoff can always be put off a little longer. Project by project, your team turns into the whole company's ops department for AI systems, and capacity for new projects drops to zero. The way out is Chapter 22. **Trap two, ask inflation.** Your service carries no price, so ideas pour toward you at triple speed. The intake gate (the checkpoint where asks come in, Chapter 25) is not optional for an internal team. It is the first lifeline.
## Prior Art: Where This Method Comes From
This book did not invent the method. The vendor-side FDE is the market's explicit version of staking "accountable from advice all the way to the production outcome" on one person, and the FDE itself is a reorganization of five mature professions under AI's constraints. This section is also the book's master list of sources. The classics later chapters borrow all trace back to these five lines (this book only paraphrases the ideas and names the source, with each chapter working out the detailed use).
**Predecessor one, strategy consulting (the McKinsey tradition).** The method comes from the tradition recorded in Ethan Rasiel's *The McKinsey Way* and *The McKinsey Mind*. A falsifiable hypothesis on day one (the day-one hypothesis), fact-based, structured decomposition, plus the answer-first way of reporting (Barbara Minto's *The Pyramid Principle*, worked out in Chapters 13 and 19). What is inherited here is problem discipline. State the hypothesis first, then let facts rule on it. In Chapter 0 you bet that "the real pain is in exception claims" and then took ten claims to test it, which is the field version of a day-one hypothesis. This tradition stops at advice and carries no implementation responsibility. When the consultant leaves, the report and the problem usually stay behind together.
**Predecessor two, the Solutions Architect (the AWS and enterprise architecture tradition).** The SA's core skill is trade-off thinking and teaching the other side to decide for themselves. There is no best architecture, only trade-offs under constraints, explained until the other side can make the call (AWS's Well-Architected (an architecture review framework) and working backwards (deriving the product from a press release) culture is this line's contemporary form, borrowed in Chapter 8 for system boundaries). The SA hands over the design and leaves, is not accountable for adoption, and nobody stands behind "did anyone end up using it."
**Predecessor three, the pre-sales engineer (the SPIN and Solution Selling tradition).** Neil Rackham proved in *SPIN Selling* that large sales run on questions, letting the other side state the cost of the problem themselves. Michael Bosworth's *Solution Selling* contributed qualification, judging whether a deal is worth the investment. What is inherited here is discovery skill. The five-question elimination in Chapter 7 and the intake rubric in Chapter 25 are both descendants of qualification. Pre-sales settles at signature, promise and delivery come apart, the demo is a weapon, and the incentive structure makes over-promising easy.
**Predecessor four, the implementation consultant (the *Flawless Consulting* tradition).** Peter Block's *Flawless Consulting* gives three things. Settle the engagement before diagnosing the problem (contracting before diagnosis, already borrowed in Chapter 0 and worked out head-on in Chapter 4), resistance is a signal and not an enemy (worked out in Chapter 20), and authentic behavior, telling the client the truth. This is the closest of the five predecessors to the last mile. It really shows up, and it really is accountable for landing. This tradition is short on technical depth, and its deliverable slides toward a report. Faced with "why did the model answer that way," it has no tool.
**Predecessor five, the product engineer (the continuous delivery and product thinking tradition).** What is inherited here is production responsibility. Code counts only once it is in production, with monitoring, rollback, and someone taking the alert at midnight (Chapters 16 and 18 work out the continuous delivery tradition). The product engineer is far from the field, the user is an abstract persona (a user profile), and he does not enter the user's organization. He optimizes for "usable by ten million people." The enterprise field wants "these eight people use it next Monday."
What did the AI era change, and why is the reorganization happening now? On the demand side, Chapter 1 covered it. Demo inflation packed the last mile with project corpses, and organizations need someone who stakes a professional identity on the production outcome. On the supply side, model capability is in surplus and landing capability is scarce, so all the valuable judgment moved to the field. The cost of a prototype collapsed too (Chapter 0), and one person with AI's help can cover the whole way from discovery (the stage after you take it on, finding out how things actually are and locking the scope, Part II of this book) to writing the code and shipping it. "Accountable from advice all the way to the production outcome" was four departments' work ten years ago. Today, for the first time, it is a workable one-person role. Palantir made the term FDE popular. What it named was this reorganization, not an invention. Inside a company you do not need the title. You need the reorganization.
!!! note "One Question First, Does This Scene Need This Kind of Investment?"
Being pressed into an old slot is one kind of mismatch. There is another that happens earlier, where the scene itself does not need this kind of investment. Several practitioners who have built FDE teams at AI companies converged on one judgment in 2026 (paraphrased). FDE-style investment belongs to one quadrant only, a complex system needing deep customization, handed to a user organization without the capability to absorb it on its own.
A simple, configurable product needs only standard tooling. A user organization made of engineers needs only good documentation plus technical support. Force this method onto either scene and the cost structure will not hold.
If what the claims department wants is off-the-shelf ticketing software, sending you in is a mismatch. Every method in this book assumes you should be there. Before assigning this kind of investment to a business unit, run this question first. The intake in Chapter 25 is its institutional version.
## Framework Two: The Inheritance Matrix
Put the five lines into one table and you have this chapter's second piece of kit. Each row reads left to right, the predecessor first, then what is inherited, what that predecessor lacks, and what the AI era added.
| Predecessor | What Is Inherited | What That Predecessor Lacks | What the AI Era Added |
|------|----------|--------------|-----------------|
| **Strategy consulting** (*The McKinsey Way* / *The Pyramid Principle*) | Problem discipline, hypothesis-driven, fact-based, answer first | Stops at advice, carries no implementation responsibility | A hypothesis can be falsified by a Field MVP within hours, and research and validation merge into one act |
| **SA** (AWS / enterprise architecture) | Trade-off thinking, choosing under constraints, teaching the other side to decide | Hands over the design and leaves, not accountable for adoption | The architecture has an uncertain component for the first time, and the eval (the evaluation set) plus human oversight become architectural elements |
| **Pre-sales SE** (SPIN / Solution Selling) | Discovery skill, question discipline, qualification | Settles at signature, promise separated from delivery | After demo inflation, "impressive" lost its signal value, and discovery's output becomes verifiable golden cases (sample cases with the correct answer settled in advance) |
| **Implementation consultant** (*Flawless Consulting*) | Landing discipline, contracting, treating resistance as a signal | Short on technical depth, deliverable slides toward a report | Contracting now covers model behavior too, red lines, oversight points, responsibility when it errs |
| **Product engineer** (continuous delivery) | Production responsibility, monitoring, rollback, iterating until someone uses it | Far from the field, the user is an abstract persona | AI coding agents compress the cost of writing code, so one person can carry the field full stack |
One line of conclusion.
> **This method = consulting's problem discipline + the SA's trade-off thinking + pre-sales' discovery skill + implementation's landing discipline + the engineer's production responsibility. It adds exactly one skill of its own, AI uncertainty management, meaning eval, human oversight, and drift (data and rules quietly shifting after launch, worked out in Chapters 11, 12, and 18).**
!!! note "This Label Is Still Drifting"
Do not expect the title FDE to settle down. A practitioner mapped its four generations in 2026 (paraphrased). Within Palantir alone the word has meant four different kinds of people. In order, around 2008 it was the DevOps firefighting squad, around 2012 the data integration engineer, around 2016 the custom solution builder, and after 2020 the platform enabler.
The skill stack turned over four times, and the four generations share exactly one thing, full accountability for the user's outcome. Another observation from the same mapping is that as code generation got cheap, the product engineer is increasingly user-facing too, and the two roles are converging (Chapter 26 returns to that trajectory).
So the answer to the question in the challenge box is this. Two years from now the label may hold different content or even carry a different name, and the matrix's five rows plus AI uncertainty management will not go out of date. What you are learning is the core, not the title. Your company having no such position does not stop you from doing all five rows. Section 3.9 of [Template 3](../appendices/template-03-capability.md) in the appendix says it more directly. Internal cross-department delivery is this capability set's home ground.
The matrix has two uses. Outward, it is the draft of your one-minute self-introduction. The third column answers "why no old profession can replace you," and the second answers "how the project dies when you get pressed into one slot." Inward, it is a self-check sheet, and the brutal version is failure mode 4.
## At Anchor & Helm: Four Hats on Thursday
Role positioning gets demonstrated in every interaction. Announcing it in a meeting does nothing. Here is the record of your Thursday in week 1 at Anchor & Helm.
**9:30 a.m., Kevin's office (the consulting hat).** The readout narrowed the scope to auto exception claims, and Kevin wants the dashboard back in. "Do the big screen while you are at it, the data is all there anyway." The request taker would say "sure." The innovation type would say "no problem." You push back with questions. What was the readout's evidence? Seven of ten claims stuck in exceptions. Whose next action does the big screen change, and which one? Kevin cannot answer the "whose." What you settle on is the exceptions queue first, the dashboard into the backlog, to be discussed once the queue produces real data (this thread comes to a head in Chapter 17). Hypothesis-driven, ruled on by facts. The hat you are wearing is the consulting hat.
**12:40 p.m., answering Victor's security questionnaire (the SA hat).** The pre-mortem drew Victor Reyes in, and he sent an 11-question questionnaire. A team from the Group entering a subsidiary's production system goes through the same gate he puts external suppliers through. The request taker would forward the questionnaire to the Digital Center's security group. The internal consultant would book a meeting. You answer it line by line yourself, giving trade-offs and not guarantees. De-identification option A is fast but coarse-grained, option B is two days slower and auditable end to end. You recommend B, with one line of reasoning. At the end you add one line, asking him to review the audit log design with you in week 3 (Chapter 12 continues this). Teach the other side to decide, do not decide for him. That is the SA hat.
**3:00 p.m., the claims department desk area (the discovery hat).** You booked fifteen minutes of tea with Linda and did not talk about the system. You brought three printed unsafe cases to ask about. "You said acting on these three would cause harm. I want to understand what 'harm' actually looks like." She talked for twenty minutes and finished with, "You system builders, this is the first time anyone asked me that." You take the opening and book half a day of shadowing next week (Chapter 6). What is used here is pre-sales' question discipline, and what it is after is something else, getting the owner of the tacit knowledge (the knowledge people hold but cannot state, learned only by watching it in the field) willing to speak.
**6:00 p.m., at the laptop (the engineer hat).** Three "days waiting" values in the queue prototype are computed wrong, and the root cause is the core system's status field lagging (one more line in the friction log, the list opened in Chapter 0 of the places where reality and paper do not match, the lead into Chapter 9). You have an AI agent draft the reconciliation script, change the judgment logic yourself, and wrap up in forty minutes.
Four hats in one day. Two things are worth saying. First, this is a normal day in this role, and each hat solves a problem the other hats cannot (how four hats become a manageable daily routine, Chapter 3 unfolds it as the four identities). Second, in every interaction you were quietly correcting a name. Kevin's session corrected "here to take our requests." Linda's corrected "here to push a system." On Friday you freeze these alignments into one page, the role charter ([Template 2](../appendices/template-02-role-charter.md)), and put it at the front of next week's deployment charter negotiation (Chapter 4). Before you send it to Grant, walk Owen through it. He signs fourth on the charter, and he needs to know first what you are accountable for and whom you do not replace.
## Failure Modes
**1. Treated as the POC demo squad.** Every executive visit brings a call to "give us a demo," your calendar fills with demos, and nothing follows one. The company has only the "innovation showcase" slot, and you really are good at demos, so every time you cooperate the slot gets one notch stronger. The project will settle forever at L0 on the outcome ladder. The correction has to happen on the spot. "A demo is fine, but the people who should be seeing it this week are Linda's team. Their scores are what move the project."
**2. Treated as the request taker.** Asks arrive as tickets, your follow-up questions read as unprofessional, "this is all settled, you just build it." The project approval process assumes by design a split between the requesting side and the implementing party, the business unit states the ask and the technology unit builds it, and that split is written into policy, which is harder to argue with than a contract. Take the first ticket without a follow-up question and your right to discovery is gone. The ending is on-time delivery of the wrong ask, with not one of the five gaps checked. The response is to write "every ask must answer a workflow claim" into the role charter and enforce it from the first ticket.
**3. Positioning yourself as internal consulting.** "Recommendations, proposals, assessments" fill more and more of your output, and you touch production systems less and less. This is the only mold you impose on yourself. Advice has a short feedback cycle, low risk, and an air of seniority. Production systems page you at midnight. The slide is comfortable. Inside a company the slide is easier than it is for a vendor, because nobody withholds payment over an output of advice alone. The ending falls back to predecessor four's classic way to die, the report that dies on the shelf. Run a weekly self-check. Did one production outcome move up a rung on the outcome ladder this week because of me? Two weeks running with no answer, and you are already internal consulting.
**4. Using "full stack" to cover "good at none of it."** You introduce yourself as someone who does everything, and under an architect's questions on design choices or a consultant's on method you last three rounds. An intersection role carries this risk by nature, and every column of the matrix has an expert stronger than you. With no real depth in any column, key meetings will vote you down column by column, and the reorganization degrades into a mediocre mix. The response is to self-score the matrix row by row (Chapter 3 gives the self-assessment tool). At least one column has to reach the level of sitting across from an expert in that field and holding your own. The rest need you to know the boundary and know when to ask for help.
## Next Monday
1. Write down what each key person on your current project actually calls you (listen for the words they use when introducing you to someone else). Check them against this chapter's four molds and judge which slot you are being pressed into.
2. Draw your three-line triangle, the three lines named earlier. Who runs your schedule and review, who gives people and data, who gives the mandate, one name each. If two of the three turn out to be the same person, congratulations, you have one line fewer. If one of the three is a name you cannot write, that line is the project's largest exposed surface right now.
3. Take stock of yourself with the inheritance matrix. Which of the five rows is your home row? Which is weakest? Write the answers from impression first. The next chapter gives the scale, and you come back then to check.
4. Use [Template 2](../appendices/template-02-role-charter.md) to write your role charter, what I am accountable for, whom I do not replace, and what each side expects. Book 15 minutes each with your sponsor and your manager to walk it through.
5. The next time you are called the wrong thing ("come give us a demo," "just build this ask"), correct it on the spot in one sentence and watch the reaction. The reaction itself is stakeholder (a party with a stake in the outcome) information (useful in Chapter 5).
**Want an agent to get you started?** In the repo you set up following [Start Here](../index.md), paste this to your coding agent:
```text
In the repo/ directory of the the-last-mile repository, help me with the Chapter 2 Next Monday actions. First copy
templates/role-charter/role-charter.md into the working directory I name, then follow templates/role-charter/prompt-role-charter.md
and put the three boundary questions to me one at a time, filling in the draft after I answer. The three names of the three-line
triangle and what each key person actually calls me are mine to say. Only record, do not guess names for me, mark anything
missing as [TBD]. When the draft is done do exactly one thing, flag every sentence that takes more than one breath to read
aloud so I can shorten it. This charter is meant to be read aloud to people. If any command errors, stop and show me the output.
```
---
## Chapter Kit
- **Judgment frameworks.** The three-line triangle (your manager / the business side / the sponsor, holding schedule and review, people and data, and the mandate); the table of four invisible mechanisms (contract and acceptance / price tag / payment milestones and the right to pause / exit date, where each hides inside and which chapter rebuilds it); the inheritance matrix (five predecessors ร inherited / lacking / added), both a capability self-check and the draft of a one-minute self-introduction; the four hats (consulting / SA / discovery / engineer), the record of one day's role switching, used to check which hat you did not put on today
- **Templates.** [Template 2](../appendices/template-02-role-charter.md), the Deliverer Role Charter, what I am accountable for / whom I do not replace / the two-sided expectations table, aligned in week 1 with your sponsor and your manager
- **Key judgments**
- "The three lines are not held by one person. That is the basic shape of internal delivery."
- "Whether that boundary is a company wall or a department wall changes how visible the mechanisms are, not the mechanisms themselves."
- "The FDE invented nothing. It is a reorganization of old wisdom under new constraints."
- "How the business side names you is how it will use you. Get the name wrong and correct it on the spot."
- "Every column of the matrix has someone stronger than you. At least one column must let you sit across from an expert. The rest, know when to ask for help."
---
# 3 ยท Four Identities: Builder, Advisor, Operator, Teacher
!!! info "Companion Templates"
๐ [Chapter Template](../appendices/template-03-capability.md) ยท ๐ [Template Library](../appendices/template-library-index.md)
> **The Challenge.** Four things land on you in one day. A bug to fix, a meeting to take, a person to teach, a weekly report to write. Every one is reasonable, every one has someone waiting, and there is only one of you.
>
> **What You Will Be Able to Do.** Use the four identities to see who you are playing in each block of time. Use the identity switching table to decide which identity this moment calls for. Rank conflicts by irreversibility when several arrive at once. Use the five-axis capability radar to find the piece you need to shore up. Label the half of your calendar that belongs to your own department, and compress it.
---
## Four O'Clock, Wednesday Afternoon
A week and a half since the Field MVP readout. The direction is set (an action queue for auto exception claims), the charter is not negotiated until next week (Chapter 4), and you are in the deep end of discovery (the formal stage of finding out how things actually are). After the readout you put a proposal to Kevin Doyle. Every morning the prototype queue runs a batch of the previous day's de-identified exception claims. Data access is nowhere in sight, so Kevin has a subordinate export the batch by hand each day. Linda Marsh's team works with the queue in view but decides by its own judgment, marking as it goes. Nothing goes to production. The point is to keep collecting scores and real reactions.
At four o'clock on Wednesday afternoon, four things arrive within ten minutes of each other.
- **4:00.** You find a bug in the queue prototype. Days waiting counts weekends too, so three claims in this morning's batch were ranked too high. Not hard to fix, under an hour.
- **4:03.** A message from Kevin. "Got a minute? Let's talk direction." When a director asks to "talk direction," it usually means he has a new idea in his head.
- **4:07.** A call from a senior reviewer on Linda's team. Linda is at headquarters for a meeting, and Sam, the new reviewer, took the queue's "suggested next action" as an order and followed it to the letter, sending a request for supporting documents on a claim whose liability is in doubt. By the team's own rules that kind of claim goes through an internal liability check first. The customer has already called in to challenge it.
- **4:10.** You remember that the first weekly report to Grant Whitmore goes out at nine tomorrow morning, and you have not written a word.
All four are urgent. All four are reasonable. Say the obvious part out loud first. Your first instinct is almost certainly to fix the bug. The reason has nothing to do with importance. It is the most fixable. The scope is clear, the path is clear, and an hour later a certain sense of completion is guaranteed. None of the other three offers that.
Choose wrong this afternoon and three months from now you are up all night maintaining a system nobody trusts. This chapter is about choosing right.
## Why This Is Hard: The Four Identities Are Not Four Skills
Chapter 2's inheritance matrix says where this role comes from, its five predecessors. The four identities say how it lives out a single day. The daily work of this role is made of four identities. "You need all four" would be easy enough. The hard part is that they are four conflicting ways to allocate time, carrying four conflicting sources of satisfaction.
| Identity | One-Sentence Definition | Source of Satisfaction | Feedback Cycle | What Being Stuck Looks Like |
|------|------------|------------|----------|--------------|
| **builder** | Build working things with your own hands, prototypes, pipelines, integrations, fixes | It runs | Minutes, certain | Prodigious output, nobody using it |
| **advisor** | Help the business side make better decisions, trade-offs, memos, direction | The business side takes your judgment | Weeks, blurry | More and more advice, less and less landing |
| **operator** | Keep what already runs reliable, monitoring, incidents, rollback | The fire is out | Hours, a strong hit | Always firefighting, never able to leave |
| **teacher** | Make the business side stop needing you, training, co-build, capability transfer | They get it done without calling you | Months, almost silent | The most common problem is never starting |
Look at the feedback cycle column. The builder's feedback comes in minutes and is certain. Tests go green, the queue runs. The teacher's feedback comes in months and is silent. The business side learned, and it shows up as "nothing happened." That is why a deliverer who came up as an engineer defaults to being stuck in builder, and one who came up in consulting defaults to advisor. People do not settle on the most important identity. They settle on the identity with the most comfortable feedback.
Now look at the teacher row. Its output is your own replaceability, and the outcome ladder's L4 (self-sufficient) is where the teacher identity settles. Of the four identities, teacher is the only one that never cries out on its own.
An internal reader hits the word "replaceability" and reads career risk. Do the arithmetic before deciding whether to be afraid. A system that no longer needs you releases your capacity, and capacity cashes out as one more department covered this year. That is the first line of Chapter 23's capacity ledger (the account of how many projects your team can run in a year). A system that cannot run without you costs next year's headcount one person who could have opened a new project. That is Chapter 2's trap one. Inside a company, being irreplaceable is not a charm. It is the prelude to permanent ops. For someone who arrives with an exit date, replaceability is a medal. For you, it is next year's capacity.
## Prior Art, and What AI Changed
**The Trusted Advisor spectrum.** Maister, Green, and Galford drew a spectrum in *The Trusted Advisor*. At one end is the subject-matter expert, brought in to answer because he knows one technical problem. At the other end is the trusted advisor, brought in to help define the problem because his judgment is trusted. Move right along the spectrum and the questions the business side asks you get bigger. From "how do I build this," to "which one should I do," to "what do you make of our whole portfolio," and the source of value moves from knowledge to judgment. The spectrum spans three of the four identities. builder lives at the expert end, advisor and teacher live at the other. It was drawn for people who are invited in. You were not invited, you were assigned, and nobody pays extra for your judgment, so the "questions get bigger" signal distorts inside a company. Use a different test. The business side quotes your judgment when you are not in the room.
If Kevin repeats your view on exception claims at a department meeting you did not attend, you are at the right end. If he asks only when you are in the room, you are still at the left.
**Maister's time leverage.** In *Managing the Professional Service Firm*, Maister points out that the economics of a professional service firm is a structure of time leverage, which he calls finder (winning the work), minder (managing the relationship), and grinder (doing the work). Senior people spend their time on judgment and relationships, and the grunt work is pushed down the leverage. A professional's value is largely equal to which layer of the leverage his time sits on.
Move the three layers inside a company and the structure holds, but the top two change content. Inside, finder is not winning work, it is screening work. Your service carries no price, so requests come to you on their own (Chapter 2's ask inflation), and a senior person's first layer of time goes to deciding what not to take. Chapter 25's intake gate is the institutional version of this layer. Inside, minder is not maintaining a counterpart who pays, it is maintaining cross-department relationships and political capital. Whether Kevin gives you people, whether Victor Reyes lets you through, whether Owen Hartley protects your schedule, all of it draws on the balance in this layer. grinder changes not one word. Doing the work is doing the work.
What did AI change? The grinder collapsed into the tools. AI coding agents compress the grunt layer inside the builder identity hard. Scaffolding, integration, tests, refactoring, the work that used to eat most of this role's time, is now mostly AI doing and you reviewing. Maister's leverage pyramid, which needed a whole team, can for the first time fold into one person. AI is your grinder, and your scarce time is forced up into advisor and teacher. That is where the discomfort comes from when senior engineers move into this role. Twenty years of accumulated satisfaction rests on builder output, while the role's time structure pushes them toward the two identities with the blurriest feedback. The discomfort says the leverage is moving. It says nothing about whether you are good enough.
builder depth did not lose value in the process. Reviewing AI output, vetoing a bad architecture, smelling a problem in a data pipeline, all of it rests on having written enough code with your own hands. Moving up the leverage assumes you can actually do the work at the layer below (the "engineering depth" axis of the five-axis radar guards exactly this floor).
## Core Framework One: The Identity Switching Table
The core muscle of the four identities is switching at the right moment. Being able to do all four is only the entry ticket. Switching runs on two things. The stage sets the default, and a signal triggers the exception. The handoff in the table's last row means the stretch where the system, together with responsibility for running it, goes to the business side. On a first read, look only at the row for the stage you are in. This book is at discovery right now, the other rows are for later, so give them one scan.
| Project Stage | Default Primary Identity | Switching Signal โ Switch To |
|----------|------------|---------------------|
| Opening and Field MVP (Chapter 0) | advisor (using builder's hands to produce the basis for judgment) | The discussion slides into implementation detail โ pull it back to the workflow claim, hold advisor |
| discovery (Chapters 4 to 7) | advisor | The parties disagree on how the data is defined โ switch to builder and reconcile it yourself
The front line worries "how will this thing use me" โ switch to teacher and spell out the boundary and the red lines |
| Design (Chapters 8 to 13) | advisor (using builder's hands to verify feasibility) | The argument over the plan hangs in the air โ switch to builder and settle it with a thin slice (one narrow but complete slice)
The risk owner keeps pressing โ switch to teacher and show where the constraints entered the design |
| prototype โ pilot (Chapters 14 to 16) | builder | The same question is asked a third time โ switch to teacher
Scores or human overrides (a reviewer overturning the queue's suggestion) come out abnormal several times in a row โ switch to advisor and back to the judgment layer
Two days running with no conversation with a user โ force yourself out of builder, switch to advisor |
| Launch and adoption (Chapters 17 to 21) | operator and teacher equally | The second incident of the same kind is still handled by your own hands โ switch to teacher
Metrics steady for two weeks โ hand over the operator duties |
| handoff and stepping out of the daily (Chapter 22 onward) | teacher | The business side says "build us another one" โ switch to advisor and back to opportunity judgment, instead of saying yes on reflex |
Three rules for using it.
1. **The default identity follows the stage, the switch follows the signal.** No signal, no switch. Switching to whoever shouts loudest is not flexibility, it is the absence of a strategy.
2. **When conflicts arrive at once, rank them by irreversibility, trust > direction > code.** Wrong code can be rolled back. A wrong direction can be corrected, at a price. Burned trust has no rollback button. Of Chapter 1's five gaps, the trust gap is the only one with no technical remedy.
3. **Switching has a cost.** Four identities in one hour means none of the four got done properly. Switch in blocks of time (half a day is best, two hours minimum), not in message notifications.
The table's last row needs two more sentences inside a company. The first is about teacher. That cell defaults to teacher on the assumption that an exit date is forcing it. You have no exit date. Nobody will take your write access away on some given day, so teacher can always start next week, and "the most common problem is never starting" holds double inside a company. You have to hardcode the forcing period yourself. On the day of launch release (Chapter 11), set the teacher deadline at the same time, write it into your own quarterly goals, and when it comes due run the five self-sufficiency tests (Chapter 22, which test whether the business side can run this system on its own). Whatever fails is claimed by the business side, not carried on by you. You can verify this more easily than anyone. Three months on, whether Linda's team still reads the reason column (the column in the queue that says why the suggestion came out the way it did) is something you can see by walking over to the desks, with no report to wait for.
The second is about operator. Finishing a project and going back to zero is the privilege of someone who arrived with an exit date. You cannot leave. Every system you launch adds one operator allotment to your calendar that never expires, and N launched systems means N of them. At some count the table starts failing you. You are in several stages of several projects at once, the operator blocks are permanently held by the systems launched first, and the advisor blocks of new projects get squeezed off the calendar. Set the steady-state ratio yourself. Operator gets at most one half-day a week, and going over says the previous system's handoff is not finished. Go back to Chapter 22's transfer ledger and find the red cell.
## Core Framework Two: The Five-Axis Capability Self-Assessment Radar
The table answers "who should I be right now." The five-axis radar answers "who can you actually be." The five axes follow.
- **Engineering depth.** Without AI, can you still stand a system up. Data, APIs, debugging, deployment.
- **AI engineering.** The engineering capability to manage uncertainty. Eval, error taxonomy, human oversight, drift.
- **Business grasp.** Read the business side's process, metrics, and money, and say which business number your system moved.
- **Narrative.** Get every level of the organization the judgment it needs, from the front line's own words to an executive memo.
- **Field judgment.** Make the right trade-off on incomplete information. Red lines, priorities, when to switch, when to say no.
One to five points per axis. Look first at what a 3 looks like, and give yourself one rough score against it.
- Engineering depth 3. You can take a prototype to a working pilot on your own, data integration, deployment, logging, basic monitoring.
- AI engineering 3. You can write an eval for one feature, golden cases (sample cases with the correct answer settled in advance), acceptance thresholds, error categories (the method in Chapter 11).
- Business grasp 3. You know the real workflow, workarounds, exception handling, how the front line gets around the system (the output of Chapter 6's archaeology).
- Narrative 3. Answer first. The memo leads with the conclusion and the decision you want, then the evidence (Chapter 13's pyramid structure).
- Field judgment 3. You can make trade-offs under time pressure, ranking by irreversibility, with a steady sense of the red lines.
An internal reader's first scoring has a roughly predictable shape. Business grasp runs naturally high. How the metrics are defined, the cause of death of the last failed project, which team lead's word counts, you were present for all of it, and this axis starts higher for you than for an outsider. What to guard against is mistaking familiarity for understanding. Do Chapter 6's archaeology as written, and work only the increment. Field judgment runs naturally low. It is the only one of the five that grows by saying no, and saying no is most expensive inside a company. Across the table is a colleague you will still be in meetings with next year, and a manager who sets your schedule, so most people substitute "let me look into it" for a judgment. A reader who came up as an engineer will see a shape with engineering depth and business grasp both high and field judgment at the bottom, and by reading rule one below, the bottom axis decides how large a project you can own alone. For internal readers the shoring-up move is usually not in the technical column, it is in Chapter 25's saying no scripts.
The full level anchors, self-check questions, and score-by-score shoring-up advice are in [Template 3](../appendices/template-03-capability.md). Three rules for reading the chart follow.
1. **The lowest axis decides how large a project you can own alone**, and the highest axis decides nothing. The five axes multiply, they do not add.
2. **The radar's shape predicts your failure mode.** Both technical axes high with narrative and judgment low is high risk for coder mode. Narrative high with thin engineering depth is high risk for consultant mode (the "Failure Modes" section unpacks both).
3. **This coordinate system gets used again.** Chapter 26 uses the same five axes to define the deliverer's F1 to F5 capability levels (F for field, the prefix marking a person's capability level, as distinct from the outcome ladder's L0 to L4 for system outcomes). Today's radar is your starting archive. Keep it.
One sentence for how identities and axes relate. Identity is how your time is spent, and the axes are whether spending it has any effect. The table governs the first, the radar the second.
## At Anchor & Helm: One Week of the Real Calendar
This is your real calendar at Anchor & Helm in the second week after the readout, with the primary identity marked on each block.
| Time Block | Item | Identity |
|--------|------|------|
| Monday morning | Sit in on Linda's team's morning meeting; collect last week's queue annotations, and probe two concern cases for the reviewer's own words | advisor |
| Monday afternoon | Have AI rewrite the case extraction script and load this week's 30 de-identified claims; spot-check 6 by hand | builder |
| Tuesday morning | Go through the business definition of the "reason stuck" categories with Kevin; clean up the friction log (the list of places where reality and paper do not match), and escalate 3 entries to risk items | advisor |
| Tuesday afternoon | Tune the queue's ranking logic, have AI produce three versions with the trade-offs written out, and veto two of them | builder |
| Wednesday morning | Answer round two of Victor's security questionnaire, de-identification method and log retention period | advisor |
| Wednesday 4 p.m. | Four things arrive at once (see below) | See the retrospective |
| Thursday morning | Give Linda's team 30 minutes on "how to read the queue's reason column and Human Call column" | teacher |
| Thursday afternoon | An upstream export format change breaks the batch run; restore it, and add an automatic validation check | operator โ builder |
| Friday morning | The weekly report to Grant, AI drafts, you rewrite the conclusion | advisor |
| Friday afternoon | Prepare next week's charter meeting, three candidate definitions of the North Star metric, meaning how the metric is defined and how it is computed | advisor |
Count them. builder holds fewer than three blocks outright, and its main action is reviewing and vetoing. In the years before AI coding agents, most blocks in a week like this would have been builder. The leverage did not disappear. It folded into the division of labor between you and AI. Now look at teacher. One block, and even that was forced out by Wednesday's incident. teacher never shows up on the default calendar unless, when you lay out next week, you reserve a block for it before anything else.
The table also misses one kind of block. It is not on the table, and it is on your calendar. The Digital Center's own weekly meeting, OKR alignment, reviews for other projects. Chapter 2 said your calendar is torn in half. This is the other half, and none of the four identities can claim it. Most of the time it is pure loss. It produces no production outcome, it changes no business action, and you still have to attend. Loss has to be named before it can be squeezed. Give it a label of its own, own-department business, count its share alongside the four identities when you lay out the calendar, and do not let it slip into advisor to pad the number.
Only two cases are not loss. One, you are asking Owen for schedule protection, or explaining why you are taking no new requests this week. That is advisor, with your manager as the audience. Two, you are teaching the Anchor & Helm eval method to someone on your team working a different project. That is teacher, with a colleague as the audience, and Chapter 23's pattern library starts accumulating here. Compress everything else, merge it into one half-day, and do not let it shred the half that belongs to the business side.
### Four O'Clock Wednesday, the Retrospective
**4:00 to 4:10, the judgment.** Each of the four belongs to an identity. The bug (builder), Kevin (advisor), Sam (operator first, then teacher), the weekly report (advisor). Rank by irreversibility. The Sam item burns trust, and the third way to die you wrote in the pre-mortem a week and a half ago ("One wrong suggestion gets followed, a customer complains, and from then on the front line trusts no suggestion") is rehearsing in front of you. The bug burns code, and fixing it before tomorrow morning's batch run is enough. Kevin burns time, and rescheduling does not make it worse. The report has all evening. The order is not in doubt. The hard part is resisting the feel of "fix the bug first."
**4:10 to 5:20, the field.** Ten minutes as operator first, calling the customer back together with the senior reviewer and putting the claim back on the internal liability check track. Then fifty minutes as teacher. You did not take Sam aside to criticize him. You asked the half of the team that was there to go through the claim together as teaching material.
Why the suggestion was wrong (the system cannot see the internal signal for "liability in doubt"). How to read the reason column (if the reason does not hold, override it). Why the Human Call column exists (advise the person, do not decide for the person). Sam's mistake turned from "the new guy caused trouble" into one calibration for the whole team, and the next day's 30-minute session was its formal version.
**5:30.** Back to Kevin. "Something came up in the field today. Would 9:30 tomorrow morning work?" advisor can be moved, it cannot be canceled. Only the next day do you learn he wanted to talk about "whether we should put up a big screen." You did not veto it on the spot. You put it on the charter meeting agenda. The idea comes back in Chapter 17.
**9 to 10 p.m.**, the weekly report. AI drafts, you rewrite the conclusion, and you put the Sam incident in as it happened. "One misuse of the prototype, closed the same day. It exposed a training gap, and a session for the whole team is scheduled." Telling the bad news yourself first is a deposit you make on the trust gap.
The next day you fix the bug at 8:20, ship it at 8:40, and the nine o'clock batch run is clean.
In a parallel universe, the you who chose to fix the bug at four o'clock has it fixed an hour later, and it feels good. The Sam item slides to the next day, and the version Linda hears when she gets back from her meeting has already become "the system had the new guy fire off notices." She says nothing, but from that week on her team rechecks every suggestion by hand. The system still runs. It is just no longer trusted. Three months later you are up all night tuning a ranking algorithm nobody believes. The bug got fixed. The project died.
## Failure Modes
**1. Coder mode.** The calendar is all development blocks, the weekly report is all "what got fixed," and a full week passes with no conversation with a user. builder is the only identity that can hand you an "I finished something" confirmation at any moment, and its output is the most visible. In the AI era this pit is deeper, not shallower. AI doubles the output, so anesthetizing yourself with volume got easier. The self-check signal. Two days running with no substantive conversation with any user or stakeholder (a party with a stake in the outcome) is a red light.
**2. Consultant mode.** The memos get better looking and the hands-on work gets rarer. You advise that "the data should be reconciled," and nobody actually reconciles it, you included. Advice never throws an error. The advisor's output has no feedback loop from production and will never wake you at three in the morning, which makes it the safest hiding place of the four identities. But the deliverer's unit of value is the production outcome, and achievement is settled only at L3/L4 (Chapter 1). Give advice without carrying the result and you fall back into the one self-imposed mold Chapter 2 warned about, internal consulting. One self-check. How long has it been since you verified the feasibility of your own advice with your own hands?
**3. Firefighter mode.** Always handling incidents, the first call the business side makes when something breaks, and quietly proud of it. Firefighting feedback is a strong hit on an hourly cycle, it feels heroic, and the business side reinforces it. Every fire you put out teaches them your phone number, not fire prevention. Nobody will take that phone number back for you, and every extra year you stay at this company adds one more system in the address book that answers only to you. Every fire you put out by hand postpones the handoff. The better the operator identity performs, the less chance the teacher identity gets to start. The self-check signal. The second incident of the same kind still handled by your own hands means what should have been taught has not been taught (Chapter 22's handoff scene shows this again).
**4. Busyness in place of judgment.** All four hats worn every day, the calendar packed, everything nudged forward a little, nothing finished. Switching itself produces the illusion of being well rounded, and fragmented switching denies every identity a solid block. builder needs two uninterrupted hours to get into state, and one teacher session needs half a day of preparation. A week with no default primary identity is a week with no strategy. In Maister's terms this is the loss of leverage. Your time is not going to the things only you can do. On Friday ask yourself which three judgments you made this week that nobody else could have made. No answer is a red light.
## Next Monday
1. Open last week's calendar, mark each block with an identity, and compute the shares. Compare against the default identity for your stage in the table. The largest gap is where your inertia lives.
2. Find the question that has already been asked three times, and teach it out this week. One 30-minute session, a one-page FAQ, one pairing, so the person asking becomes a person who can answer.
3. Run a five-axis self-assessment with [Template 3](../appendices/template-03-capability.md), find your lowest axis, and set one verifiable 30-day shoring-up action ("improve communication" is too vague, "the next three weekly reports draw zero follow-up questions" counts).
4. Write "trust > direction > code" somewhere you will see it the next time conflicts arrive at once. Use it to rank once, then review afterward whether the ranking was right.
**Want an agent to get you started?** In the repo you set up following [Start Here](../index.md), paste this to your coding agent:
```text
In the repo/ directory of the the-last-mile repository, help me with the Chapter 3 Next Monday actions. First run python3
templates/self-assessment/radar.py to produce a radar chart SVG from the built-in sample, tell me where the file is and open it
for me. Then copy sample/radar.csv into my own self-assessment file. I score the five axes myself against Template 3's behavior
anchors. You only ask me for the behavioral evidence on each axis and record it, you do not score for me. Once the scores are in,
rerun radar.py pointed at my file. The 30-day shoring-up action for my lowest axis is mine to write. You only check whether it can
be verified, and send adjectives back. If any command errors, stop and show me the output.
```
---
## Chapter Kit
- **Judgment frameworks.** The four identities (builder / advisor / operator / teacher) and their table of conflicting satisfactions; the identity switching table (stage ร default primary identity ร trigger signal); the conflict ranking rule (by irreversibility, trust > direction > code); the five-axis capability self-assessment radar (engineering depth / AI engineering / business grasp / narrative / field judgment)
- **Templates.** [Template 3](../appendices/template-03-capability.md) Capability Self-Assessment and F1โF5 Rating, the self-assessment half (3.0 to 3.6), 1 to 5 behavior anchors per axis plus self-check questions plus score-by-score shoring-up advice, re-scored and archived each quarter
- **Key judgments**
- "Code gives certain feedback and people do not. That is why engineers hide in the builder identity."
- "The same question asked a third time means what should have been taught has not been taught."
- "When conflicts collide, rank by irreversibility, trust > direction > code."
---
# 4 ยท Opening and Agreement: Write the Deployment Charter
!!! info "Companion Templates"
๐ [Chapter Template](../appendices/template-04-deployment-charter.md) ยท ๐ [Template Library](../appendices/template-library-index.md)
> **The Challenge.** After the readout everyone says "continue," and each person's "continue" is a different project. You do not want to find that out in week 10.
>
> **What You Will Be Able to Do.** Write a one-page deployment charter with the seven elements. Run a 60-minute contracting meeting. Settle the North Star metric, both sides' commitments, and the exit conditions at the cheapest moment the project will ever have, and make the business side as accountable for that page as you are.
---
> **Part II Navigation.** The timeline below governs Chapters 4 to 9, not just this chapter (book order โ time order, and Chapter 7 turns the clock back).
> readout (week 1 Tuesday) โ the five-question elimination (weeks 1 to 2, Chapter 7) โ stakeholder map (the night before the charter meeting, Chapter 5) โ the charter meeting (early week 3, this chapter) โ shadowing Linda Marsh (week 3 Wednesday, Chapter 6) โ charter signature (week 4) โ the data deadlock and the way through it (after signing, Chapter 5)
## Three Versions of "Continue"
On the Tuesday afternoon after the Field MVP readout, Grant Whitmore said "continue." You walked out of the meeting room in a good mood. A good mood is when people stop noticing danger.
The danger surfaced over the next two weeks, piece by piece. Kevin Doyle sent an email, copying two engineers, subject line "dashboard requirements, first draft." In week 1 he had agreed to your face to put the big screen into the backlog (Chapter 2). In week 2 you had just put it on the charter meeting agenda (Chapter 3). Two days later the first-draft email arrived. In his universe, "continue" means the backlog thaws and work starts on the management dashboard. The PMO ran its quarterly stocktake, dug out an old ticket approved three months ago for an "intelligent Q&A assistant," acceptance wording "Q&A accuracy โฅ90%," and asked when that ticket closes. In the PMO's universe, "continue" means pushing ahead on that old ticket. In your universe, "continue" means entering discovery (the stage of finding out how things actually are, Chapter 2) and turning the exceptions queue from a ten-case prototype into a verifiable plan.
Three "continue"s, three projects. Each has an email for evidence, and each is spending money. Leave them alone and they run in parallel for ten weeks, then collide head-on at some reporting meeting. Kevin asks where the dashboard is. The PMO asks how the accuracy gets accepted. You explain that what you are building is an action queue (the queue that lines up the next action for every claim). At that moment everyone feels let down, and everyone is right.
What this chapter does is take that collision apart ten weeks before it happens, with one page and one hour of meeting. That page is the **deployment charter**.
## Why This Is Hard: Vague Authorization Is More Dangerous Than None
A project with no authorization does not die of misunderstanding. It never starts, nobody commits anything, and nobody gets hurt. Vague authorization is the dangerous kind. "Continue." "Get started on it." "I support this direction." Each of these hands every party the permission it wanted, and then every party spends real money on a contradictory expectation. Kevin committed engineer schedule. The PMO hung the old ticket back on its register. You committed a discovery plan. The more each side commits, the more certain it becomes of its own version of "continue," and sunk cost hardens the expectation gap from a disagreement into a standoff. Nobody will give ground first, and time piles small differences into opposition.
Chapter 0 said it. People's agreement to an abstract description is cheap, and "continue" is the most typical case of all. The Field MVP's principle applies here too. Nobody at Anchor & Helm wants to betray this project, and betrayal is not what the charter guards against. It forces the expectations of three universes onto one table and aligns them while nobody has yet paid for a misunderstanding. A charter's use is to make the expectation gap blow up at the cheapest moment.
A team from outside doing the same work has at least one signed commercial document forcing both sides to write their expectations down. Inside, you do not even have that floor. One "continue" from Grant can start a project, and nobody has to write anything down. So the collision of three "continue"s comes earlier, and comes more often. The charter is therefore your only written agreement, and not one of the seven elements can be missing.
## Prior Art: Contracting and the AI Era's New Clauses
The method behind this one page comes from a mature tradition in consulting. In *Flawless Consulting*, Peter Block puts **contracting** (the entry conversation that sets terms of engagement) first in the consulting sequence, ahead of all diagnosis and all solutions. Three ideas sit at its core, and fifty years on they hold unchanged.
**First, the agreement is a 50/50 conversation about responsibility.** A project is a co-build where each side carries half, and the one-way relationship of "you deliver, they accept" does not go far. Negotiating the agreement is that relationship's first rehearsal. A business side that only wants to set the questions during contracting will not turn into a co-builder later either. Inside a company one more move is missing. The business side does not accept anything, it only cooperates or fails to. Nobody will manufacture that moment of forcing someone into co-building for you. You have to manufacture it yourself, in this meeting.
**Second, the agreement has fixed elements.** Goal, both sides' commitments, boundaries, support needed, feedback and how it gets measured. Not one may be left blank. The cell you leave blank is the cell that blows up ten weeks later.
**Third, "what you want" and "what you are willing to put in" must be negotiated in the same room.** This is Block's sharpest point. The side asking is used to talking only about what it wants. The side building is used to committing only to what it will put in. The contracting conversation forces both hands onto the table at once. You want handling time down 30%? Fine, then you put in 2 hours a week from Linda's team. A want that will not discuss what it costs is a wish, not an agreement.
In the age of AI agents this method takes on one change of order and two new clauses.
The change of order is that the charter now gets signed after the Field MVP evidence. In Block's era contracting happened before any work at all, so both sides could only trade vision for vision, and the agreement filled up with toothless words like "improve efficiency." Today the cost of a prototype has collapsed (Chapter 0), so you can spend 2 hours getting ten scored cases before you talk terms. That changes the nature of the negotiation, from vision against vision to evidence against evidence. "How big a target" rests on how the waiting time in ten cases breaks down. "Whether Linda's team puts hours in" rests on three unsafe annotations. This is the whole reason this book puts the Field MVP before the charter.
**New clause one, the data boundary.** A traditional consulting agreement needed no separate conversation about data. Reading reports and running interviews was enough. An AI system has to eat the business side's data to work at all, so which data, de-identified to what degree, in whose environment, approved by whom, delivered when, gets promoted from a logistics question to a clause of the agreement. A project that does not write the data boundary into the charter turns, in week 9, into an email to IT that nobody answers for ten days.
**New clause two, acceptance tied to the eval (the evaluation set).** A traditional deliverable can be accepted by walkthrough. A probabilistic system cannot. "The demo passed" and "the boss is happy" cannot accept a system whose output may differ every time. Only one thing stands up. Both sides co-build a set of golden cases (sample cases with the correct answer settled in advance), agree on thresholds, and a score over the line is what counts. Most projects inside a company have no acceptance meeting, so acceptance is replaced by launch release conditions. The eval spec both sides co-build is the authority, golden cases plus agreed thresholds, and a score over the line releases the launch. The people who release are the business-side owner and the risk owner, not you. When the charter is signed the eval spec usually does not exist yet, so the release mechanism has to be signed first. What gets signed is "what counts as counting," and the numbers come later as an attachment (the eval spec is Chapter 11).
## Deployment Charter, Definition and Seven Elements
> **Deployment charter, a one-page working agreement signed by both sides, setting out one measurable North Star outcome, named users and owners, boundaries and data, what each side puts in, and two-way exit conditions.**
Three words in that definition carry weight. "One page." Past one page nobody remembers it, so it constrains nobody. "Signed by both sides." It is produced live in the contracting meeting, and a document you wrote and then carried to the business side for a signature does not count. The charter takes four signatures, Grant Whitmore, Kevin Doyle, you, Owen Hartley. Drop the last one and your schedule can be pulled out from under you tomorrow. What Owen signs is not what gets built. It is that these people, for these weeks, belong to this project. He is signing that these people's time goes to you first for these weeks, not agreeing to what you do with it. "Exit conditions." They matter as much as the goal, and why comes shortly.
The seven elements follow (the full fillable template is [Template 4](../appendices/template-04-deployment-charter.md)).
| # | Element | The Standard in One Sentence | What Happens If It Is Blank |
|---|------|-----------|-----------|
| 1 | **North Star outcome metric** | One, measurable, an outcome and not an activity | The project proves itself with activity metrics, and the board does not buy it |
| 2 | **actual user and owner, by name** | The person who uses it and the person who owns it after launch, names and not departments | The day the system launches is the day an orphan is born |
| 3 | **Scope and Red Lines** | What gets built, what explicitly does not, what is not touched at all | Scope swells at every meeting, and red lines get stepped on in the excitement |
| 4 | **Data boundary and access commitments** | Which data, in what form, approved by whom, delivered when | The week 9 death email, "it is going through the process" |
| 5 | **Business-side commitment** | People ร time, by name, written into their calendar and not their wishes | A one-way delivery promise, and the handoff is certain to fail |
| | Second row of it, **your team's commitment and protection conditions** | Your people ร time, by name, countersigned by your manager, stating who has to approve before anyone is pulled | Your schedule can be pulled out from under you tomorrow |
| 6 | **Acceptance method** | Write it as launch release conditions, tie it to the eval, golden cases plus thresholds, sign the mechanism now and fill in the numbers later | Acceptance turns into the political question of "does the boss think it is fine" |
| 7 | **Exit conditions** | Two-way, which signal lets which side call a stop | Once the project is a zombie, nobody has the authority to shoot it |
The seven elements are produced by one **60-minute contracting meeting**. The skeleton agenda follows (the full version is [Template 4](../appendices/template-04-deployment-charter.md)).
- **First 10 minutes, open with evidence**. No vision deck. Retell the Field MVP scoring results and the readout conclusion, so every negotiation that follows is anchored to evidence.
- **10 to 25 minutes, the North Star negotiation**. Definition, baseline, size of the change. The hardest 15 minutes of the meeting.
- **25 to 35 minutes, names and commitments**. actual user, owner, the people and hours the business side puts in, plus your team's commitment and protection conditions. Write names, and watch the person or their manager nod in the room.
- **35 to 45 minutes, scope, red lines and the data boundary**. The "not doing" list, the data clauses, the date of the security review.
- **45 to 55 minutes, acceptance and exit**. Settle the acceptance mechanism first, then ride it into the exit conditions. That is the least awkward door into talking about exit.
- **Last 5 minutes, read back and commit**. Read the seven elements back out loud, and agree on writing it up within 48 hours and signing within a week.
Inside a company this meeting has three more things to settle, all of them in the 25 to 35 minute stretch. First, claiming the owner role needs something in return. The business side will say "you built the system, so of course you keep it running," and inside one company that sentence is literally true, so the owner cell is not a name to fill in, it is a claim to be made. Kevin claims owner, and what he gets back is that after launch this system's priorities and release schedule are his to set. A claim with nothing in return is politeness at signing and a blank at handoff.
Second, a committed input has no invoice to remind anyone, and the substitute is the fulfillment rate. The business-side commitment gets reported once a month at the sponsor weekly. Hours promised, hours delivered, one line of numbers, and whoever came up short is visible in the meeting.
Third, your team's commitment goes in too, along with the protection conditions. A team from outside has a commercial document as a shield. You do not, and your people can be pulled off at any moment to put out someone else's fire. Owen countersigns this row, and it states that pulling anyone requires going through Grant's weekly first.
Two facilitation rules. First, the meeting must have a decision-maker who can settle metrics and commitments on the spot, at Anchor & Helm that is Grant. Second, you draft the charter, and you never finalize it alone. Every cell must be changed by the business side at least once during the meeting. A business side that leaves no fingerprints on it will never think of it as theirs.
An internal reader holds three things an outsider does not. Use them. First, precedent. You can get the original charters other projects in this company wrote, including how they set their targets, and "the last project wrote it this way too" is a hard argument inside a company. Second, dropping by. The three "continue"s do not have to wait for emails to collide. Before the meeting you can walk over to Kevin's desk and then to the PMO's, ask each for their version, and bring only the differences into the room. Third, depth. You know how this company's last AI project actually died, so the pre-mortem memo can carry real events instead of imagined ones.
## At Anchor & Helm: A Record of the Charter Meeting
The meeting is set for early week 3, after the readout. Grant Whitmore, Kevin Doyle, you. Your team's commitment row went past Owen Hartley beforehand, and he signs fourth. Three key rounds, ordered below by how contested they were, not by the agenda.
**Round one, how -30% got negotiated.**
The North Star you propose is "first-touch handling time for auto exceptions, down 30%." Nobody objects to the metric itself. It already appeared in the pre-mortem (Chapter 1). The disagreement is over the size. Grant says, "Can we write 50%? The board likes numbers it can hear."
Most people say yes here. The sponsor wants a prettier number, and the pull to agree is enormous. You do not say yes and you do not say no. You put the MVP's ten-case table on the screen and break the waiting time down by where it goes. "Across these ten claims, about six tenths of the waiting time goes to 'waiting on documents with nobody knowing whom to chase' and 'stuck with nobody claiming it.' The queue can eat into that part. The other four tenths is physical time, survey and loss assessment where somebody has to drive out there, plus the core system's own routing. The queue does not touch it. Eat half of the six tenths it can reach and you get -30%. To get -50% you have to go after automated payout decisions and a core system rebuild. The first is a red line we already agreed on. The second is another project with another budget."
Then two more lines. The business framing. "Anchor & Helm's first payment cycle runs 1.6 times the industry average. -30% pulls it back near the average, and at the board that is a 'catching up with the industry' story, and it holds. -50% is a 'crushing the industry' story, and delivering it costs a red line." The risk framing. "The baseline itself is not clean. The core system's status field does not match actual progress, and the MVP already exposed that. I would write in that the baseline is whatever discovery's reconciliation produces. Promising -50% now means signing a number on a foundation nobody has checked."
Grant looks at the breakdown for a while. "Write 30. But once the reconciliation is done, you come tell me whether that number was conservative or reckless." Done. Not once in that negotiation did anyone say "I think." Every step stood on MVP evidence.
**Round two, the thirty seconds of silence over the exit conditions.**
In the last fifteen minutes of the agenda you read out the draft exit clause. "When two consecutive evaluation cycles miss the threshold and no workable path to improvement exists, either side may propose termination. When the business side's promised data access or staffing goes unmet two weeks running, a resource reassessment is triggered, in three tiers. First it escalates to the sponsor weekly, then the project turns to awaiting inputs on the PMO register, and last your team's people are released to other projects."
The room goes quiet. Be ready for that quiet before it comes. Raising "under what conditions we break up" in a meeting where the goal was just agreed and the mood is good is like reading a divorce settlement at an engagement party. Kevin smooths it over. "Talking about splitting up before we have even started, is that not bad luck?"
You give the line you had prepared. "We write these two lines so that if it ever comes to that, the person who calls a stop does not have to be the villain. Without them, everyone can see the project is in trouble, but whoever says stop first owns the failure, so nobody says it and the project keeps burning money. With them, calling a stop is only executing the agreement."
Grant is silent a few seconds. "Put it in. I have approved too many projects that could not be stopped."
These two lines are harder to say out loud inside a company than outside. For a team from outside, calling a stop is a commercial move. For you it means admitting your own project failed, and it looks like undercutting Grant. You also do not hold the pause card. You are not a supplier, salaries get paid either way, and stopping reads as infighting. So the second half of the exit clause does not say who may pause. It says what triggers a resource reassessment. None of the three tiers requires you to call a stop. Missing inputs are the signal by themselves.
A few months later these two lines grow a pilot-level sibling, the kill criteria (Chapter 14), and in one launch incident they get executed by the book and save the relationship between the two sides. That is Chapter 18's story. For now one sentence is enough. The awkwardness of writing exit conditions lasts thirty seconds. The cost of not writing them is counted in quarters.
**Round three, the one line that pays for the whole project.**
In the business-side commitment cell you write, "Linda's team, 2 hours a week, for case annotation and rule confirmation, from discovery through one month after launch."
Kevin shakes his head. "Them marking a few for you on the side these past two weeks is fine. Putting it in the charter is another matter, that is a hard commitment. Her team carries the heaviest backlog in the department, and I cannot lock my tightest people onto this. How about two new hires working with you instead?"
You go back to the scoring sheet. "All three unsafe came out of 20 minutes with Linda. The three changes we made to the queue's ranking rules these two weeks all came from her team's annotations, and new hires would not have caught those three. 'Chase priority โ risk priority' is a judgment that lives nowhere in this company except inside her team's heads. Without those 2 hours the system learns the wrong priorities, and the review time wasted every week after launch runs far past 2 hours. Putting it in the charter only turns something already happening from 'helping out on the side' into a named role, and puts their hands on this system's steering wheel early."
Grant looks at Kevin. "We will find another way on the backlog. Give him the 2 hours." Kevin sets three conditions. No more than 2 hours a week, booked in advance, and not during the morning peak. All three go into the charter.
The compound interest on that line is invisible while you write it. It turns Linda from "the person nobody invited" into a person named in the agreement (Chapter 5 explains why that matters). The shadowing later, the eval co-build, all the way to the business side running this system on its own at handoff, the institutional source of that time is this one line. If the whole charter could keep one line, keep this one.
A week later the charter is signed. One page, seven elements, four signatures. The old ticket's chatbot acceptance wording went through a project approval change review, was voided and filed, and was replaced by the eval release mechanism. Kevin's dashboard went onto the "not this phase" list. Three "continue"s merged into one project.
Here is how the seven elements read on the page that got signed. The three cells actually fought over were the North Star, the business-side commitment, and the exit conditions. The other four moved fast on the agenda, filled in item by item off the template.
| # | Element | How the Anchor & Helm Charter Filled It |
|---|------|------|
| 1 | North Star outcome metric | First-touch handling time for auto exceptions, down 30%, with the baseline set by discovery's reconciled data |
| 2 | actual user and owner, by name | actual user Linda Marsh and her review team, owner Kevin Doyle (Director of Claims Operations), executive sponsor Grant Whitmore (COO) |
| 3 | Scope and Red Lines | This phase builds the auto exceptions action queue, covering missing documents / abnormal amount / disputed liability; this phase does not build the management dashboard, the non-auto lines of business, or routine claims; the red lines are no automated payout decisions, no automated outbound messages, no handling of identifiable customer information |
| 4 | Data boundary and access commitments | This phase uses auto exception claims data only, three items. The core system's exception fields, as a read-only de-identified view. The review team's Excel tracker, as a read-only export. Document emails and attachments, as read-only access. Data does not leave Anchor & Helm's environment, and name / ID number / plate number / phone are replaced with placeholders at the extraction layer before they reach any model call, with the de-identification standard set by Victor Reyes. The approver is the head of the IT data team, countersigned by Kevin Doyle as the business data owner, with access committed within two weeks of signing; the security review is run by Victor Reyes and set for week 8 |
| 5 | Business-side commitment | Linda's team, 2 hours a week, for case annotation and rule confirmation, from discovery through one month after launch; booked in advance, not during the morning peak; the fulfillment rate reported monthly at Grant Whitmore's weekly |
| | Your team's commitment and protection conditions | You and your team's engineers, by name, from discovery through the completion of handoff; countersigned by Owen Hartley, and pulling anyone goes through Grant Whitmore's weekly first |
| 6 | Acceptance method | Launch release conditions. The eval spec both sides co-build is the authority, golden cases plus agreed thresholds, and a score over the line releases the launch; the releasers are Kevin Doyle and Victor Reyes, and the case count and thresholds become an attachment once written |
| 7 | Exit conditions | When two consecutive evaluation cycles miss the threshold and no workable path to improvement exists, either side may propose termination. When the business side's promised data access or staffing goes unmet two weeks running, a resource reassessment is triggered, in three tiers. Tier one, escalate to the sponsor weekly, where the sponsor decides between more input and less scope. Tier two, the project turns to "awaiting inputs" on the PMO register, with the reason column naming whose input is missing and what it is. Tier three, your team's people are released to other projects, with the release date and reason recorded openly in the weekly project report |
## Failure Modes
**1. The charter gets processed into a form.** The PMO says the company has a project approval template, just fill the charter into it. So one page becomes 27 fields in the requirements ticket system, each field with its own filling convention, and once it clears approval nobody opens it again. The language shifts from conversation to form-filling, and "Linda's team, 2 hours a week, not during the morning peak" becomes "business department cooperation, high." Form-filling language needs nobody to change a single cell, so nobody leaves fingerprints on it, and a signed form constrains nobody. The charter and the project approval form must be two documents. The approval form governs budget and process compliance, the charter governs the work and the expectations, and the charter stays one page forever, in a state where somebody can still mark it up in red pen.
**2. The goal written as an activity list.** The North Star cell reads "complete development of the exceptions system and run the launch training." Activities are controllable, outcomes carry risk. "Complete development" can always be delivered and "-30%" can fail, so both sides have a reason to write toward safety. You are afraid of carrying a metric, and the business-side handler is afraid of pledging one to his boss. But a charter written as an activity list is no charter. It accepts "work was done," nobody asks what changed, and only outcome metrics survive the boardroom (Chapter 1). The test is short. Read the goal out loud and ask, "Could this sentence fail to come true?" A goal that cannot fail is not a goal.
**3. The business-side commitment clause left blank.** Your team's commitment runs three pages, and the business-side commitment cell reads "will provide necessary support."
The reasons for leaving it blank come even more easily inside a company. We are all one family, it feels awkward to put it in writing. Asking for resources sounds like admitting you cannot handle it. And the business side assumes the Digital Center is a resource the company assigned to them, that doing the work is your job to begin with, and that sentence is literally true and harder to argue with than "we are paying for this." But the raw material of an AI system, the data, the tacit knowledge (the experience nobody ever wrote down), the trial feedback, all sits with the business side, and a project with zero business-side input cannot even assemble its raw material. Handoff is the worse problem. An organization that never put anything in cannot take over operating it (Chapter 22). There is only one road for the counterargument. Do not argue over whose job it is, go back to the raw material. The three unsafe were marked by Linda, not by you. Block's discipline is at its sharpest here. What you want and what you will pay get negotiated in the same room, and if you cannot agree you do not sign. A charter with no business-side commitment in it is a one-sided love letter.
**4. Skipping the charter.** "Kevin and I get along well, all this red tape would only put distance between us." That sentence has the causation backwards. A good relationship is the only window in which the unpleasant things can be said plainly. Talking about exit conditions now is preparation. Talking about them after something goes wrong is assigning blame. Agreements do not spend relationship capital. Vagueness does. If Kevin only discovers six months from now that the dashboard was out of scope, the trust lost in that moment is more than any haggling at the negotiating table would have cost.
!!! note "Vendor View"
The vendor side has a procurement contract. Charter and contract must be two documents. The contract governs the commercial and legal terms, the charter governs the work and the expectations, and the charter stays one page forever, markable in red pen. Hand the charter to the contract process and its language shifts from conversation to defense, and the co-build relationship dies before anyone signs.
## Next Monday
1. Find the project on your desk that has already started with no charter, and ask three key roles each to write one sentence on "what success looks like for this project." Three different sentences back is all the reason you need to call the contracting meeting.
2. Draft a one-page charter with [Template 4](../appendices/template-04-deployment-charter.md), forcing all seven elements to be filled. The cells you cannot fill, usually the business-side commitment and the exit conditions, are the project's largest exposed surface right now.
3. Book the 60-minute meeting. Make sure the person who decides is in the room, and make sure you open with evidence (the scored MVP, data samples) rather than a vision.
4. Write it up and send it within 48 hours of the meeting. Watch who changes what. Every change is an expectation gap blowing up early, the cheap kind.
**Want an agent to get you started?** In the repo you set up following [Start Here](../index.md), paste this to your coding agent:
```text
In the repo/ directory of the the-last-mile repository, help me with the Chapter 4 Next Monday actions. Copy
templates/deployment-charter/charter.md into the working directory I name, and ask me the seven elements one by one.
Fill in what I answer and nothing else. Leave any cell I cannot answer blank and mark it [still to discuss]. Those
cells are this project's largest exposed surface right now, so do not fill them in for me. If I paste in de-identified
meeting minutes, follow prompt-minutes-to-charter.md to produce a draft first, then list the "still to discuss" items.
The three key roles' sentences on "what success looks like" are mine to collect, not yours to write.
If any command errors, stop and show me the output.
```
---
## Chapter Kit
- **Judgment frameworks.** The seven elements of the Deployment Charter (North Star outcome metric / actual user and owner by name / scope and red lines / data boundary / business-side commitment by name / acceptance tied to the eval / two-way exit conditions); the 60-minute contracting meeting agenda
- **Templates.** [Template 4](../appendices/template-04-deployment-charter.md), Deployment Charter and Pre-mortem Memo, the fillable seven-element template plus the full contracting meeting agenda (4.1 and 4.2)
- **Key judgments**
- "Vague authorization is more dangerous than none."
- "A charter's use is to make the expectation gap blow up at the cheapest moment."
- "A charter with no business-side commitment in it is a one-sided love letter."
- "Exit conditions mean the person who calls a stop does not have to be the villain."
---
# 5 ยท Trust Ships First: Your First Deliverable Is Not Software
!!! info "Companion Templates"
๐ [Chapter Template](../appendices/template-05-stakeholder-map.md) ยท ๐ [Template Library](../appendices/template-library-index.md)
> **The Challenge.** The business side will not give you data, meetings are all pleasantries, and the people who matter can never be booked. Nobody can fault your technical plan, and the project still does not move.
>
> **What You Will Be Able to Do.** Draw a stakeholder map with real names and find the people who decide the project's fate without ever sitting in the meeting room. Audit your trust balance with the Trust Equation's 12 behavior-level questions. Break the "cannot get the data" standoff in the other side's language of power.
---
## Day Ten, and Still No Reply
The day after the charter was signed, you emailed IT to request the exceptions data. The ask was the full history needed for reconciliation, plus a direct read-only view. The de-identified batches Kevin Doyle's staffer exports by hand every day can feed a prototype. They cannot carry a reconciliation. Read-only, de-identified, scope spelled out in full.
The reply came on day three, one line. "Please go through the data request process." You went through the process. Then silence. You chased once on day seven and again on day ten, and the ticket status stayed "In progress." As it happens, that is the very status field in the core system you already know cannot be trusted.
The meetings are off too. At Kevin's weekly, you present the data reconciliation plan and everyone nods. Of the three follow-ups you booked afterward, two were pushed off with "too busy right now," and the one person who showed up brought only the official line. Nobody opposes you. Nobody cooperates with you either.
Technically you are beyond reproach. The plan is clear, the MVP has a track record, the charter carries Grant Whitmore's signature. But this company's real reaction to you is a single subtext. It knows you exist. It does not think you have anything to do with this. You have an employee ID and a desk, and people in Claims know your name, but walking into that meeting room with you is the Digital Center's history. What the last system builder left behind, Linda Marsh remembers better than you do. "Another innovation type" (Chapter 2) is the name for that account.
The charter does in fact spell out a data boundary clause in black and white, and the resource reassessment conditions (Chapter 4) say that when the business side's promised data access goes unmet two weeks running, it escalates to the sponsor for a resource reassessment. You could take that to Grant. But you sense that is not the answer. Invoking the contract is a trust withdrawal. Whoever pulls out the clause on day ten finds every door shut tighter on day thirty, and Kevin will still be in this building next year, and next year you will still be asking him for people. The clause gives you a floor, not a key. Where the key is, is what this chapter is about.
## Why This Is Hard: Companies Run on Trust Networks
The engineer's instinct is to read "cannot get the data" as a permissions problem. Find the person with the permission, get him to click approve. That model is wrong. The org chart draws the permission system. Day to day, the company runs on the trust network. Data access, the truth, and the front line's time, the three resources a project needs most, are open only to people whose trust balance is positive. The approval process is only trust's bookkeeping. No money in the account, and the process is an indefinite "In progress."
Worse is the speed. Code can be pushed through overnight, models can be tuned in parallel, but trust has a hard cap on how fast it forms. It accrues one notch at a time through the cycle of promise and delivery, each cycle takes days, and the cycles cannot run in parallel. That means the slowest step on an AI delivery's critical path is trust. Model and integration queue behind it. Chapter 1 said it. Of the five gaps, only the trust gap has no technical means of acceleration.
So take this chapter's title literally. Trust ships first because it stands in front of every other deliverable, not because talking about integrity sounds respectable.
## Prior Art, and What AI Changed
How trust accrues, consulting worked out long ago. The formula is ready-made. Maister, Green, and Galford proposed the Trust Equation in *The Trusted Advisor* (paraphrased here).
> **trust = (credibility ร reliability ร intimacy) / self-orientation**
Three terms in the numerator. Credibility, what you say can be believed. Reliability, you do what you say. Intimacy, the other person dares to tell you the truth. One term in the denominator, and it is the killer of the whole equation. Self-orientation, how visibly you are working for yourself. Once the business side smells you optimizing your own KPI (a promotion, your department's annual KPI, a thing you call "my project"), no numerator is high enough to survive the division to zero.
Peter Block's *Flawless Consulting* adds the fastest source of intimacy. Say out loud the tension in the room. The sentence everyone feels and nobody says, "I sense people have reservations about this plan, and nobody is voicing them." When you say it first, the air in the room changes at once.
What did the AI era change? The starting point went from zero to negative. A traditional consultant walks in with a trust balance of zero and builds up from there. An AI project opens with a double trust deficit. The front line fears replacement, and in their eyes you are here to take their measurements. Add black-box suspicion. Why the system recommends what it recommends, even you cannot explain line by line. Fear plus suspicion. At the start of an AI project, the trust balance is negative, not zero.
You are a colleague. That does not make the ledger look any better. Your starting point differs from the consultant's. It is not zero, it is an inherited balance. Every time you and Kevin have worked together, the reputation the Digital Center's last system left in Claims, Linda's twenty-year summary of "system builders," all of it entered the books before your plan did. The inherited value is occasionally positive, usually negative. A business unit remembers the last time a system made a mess, not the last time one saved hours. Stack fear and suspicion on top, and the negative only deepens.
The ceiling is lower too. No prophet at home. A hired expert can draw credibility in advance on a title, and a colleague from the same building cannot. When you talk about what AI can do, Kevin hears one more person telling him what a system can do. Your judgment has to stand on its own, on evidence from this company's floor.
That changes the play, and each of the four variables has its own. Credibility cannot be counted on short term. AI's reputation is badly overdrawn right now, the more fluently you pitch the more suspect you are, and a colleague pitching is more suspect still. Intimacy needs the other side to open up first, and cannot be rushed. The only variable fully in your hands from day one is reliability. Say Friday, deliver Friday. Here you have one thing extra. Presence. The promise-and-delivery cycle runs in days, and half your day is spent at a desk in the claims area (Chapter 2). Say you will look at that spreadsheet this afternoon, and this afternoon you are at the desk. One cycle compresses to half a day. Others bank three notches a week, you can bank six. And "say the tension in the room" gets a concrete target in an AI project, the fear of replacement itself. Standing in front of Linda's team and saying "You may be worried this thing is here to replace people. That is why every output has a Human Call column, and the call stays with you" is faster than ten pages of assurances.
## Framework One: The Six Roles of the Stakeholder Map
A stakeholder map is a map with real names. It answers "which specific people does this project's success or failure depend on?" Six roles, three things per cell. The name, what they fear, what they win.
| Role | The Test | Common Mistake |
|------|----------|----------|
| **Decision-maker** | Who can, in one sentence, fund this project's next phase or kill it? | Writing only the sponsor, forgetting the sponsor answers to someone too |
| **Actual user** | Whose daily actions change after launch? | A department name instead of a person's name |
| **Data owner** | Who owns the data you need, on the business side? | Going only to IT, who holds the key, not to the data's business owner |
| **Risk owner** | When something goes wrong, whose name is on the accountability email? | Treating review as process and assuming nobody really cares |
| **Maintainer** | A year after launch, whose annual goals list this system? | Writing yourself in, which kills the cell's diagnostic power; leaving it blank, and finding nobody there at handoff (the transfer of the system) |
| **Blocker** (the person in the way) | Without whose nod does everything stop, even though he never attends? | Reading "never attends" as "not involved" |
Three rules of use.
- **First, real names.** A map that says "the IT department" is decoration. Departments do not fear and do not win. Only people do.
- **Second, one person can fill several cells.** Kevin is the workflow owner of the exceptions process, the owner of the workflow the AI changes. That title is not in the table above. He can halt the process in one sentence, so he lands in the "decision-maker" cell, the same cell as Grant, with a jurisdiction one ring smaller. He is also the business owner of the exceptions data, so his name goes in the "data owner" cell too. People in several cells are your high-leverage nodes. Owen Hartley is in the "decision-maker" cell as well. He cannot kill the project, but he controls your schedule and could reassign you tomorrow (Chapter 2). That is the extra line an internal map carries.
- **Third, the map's value is in the blanks and the wrong cells.** Once it is drawn, one question is mandatory. Which name has never been in the meeting room? That is your blind spot, and the project's future cause-of-death candidate.
The "What They Fear / What They Win" columns are the engine of this table. What he fears decides whether he blocks you. What he wins decides whether he pushes you. The goal of stakeholder management is for everyone to have their own win-loss ledger in this project. Being liked does not make the list.
Beyond the six cells, an internal map needs two extra cells. They are not new roles, the six tests stay exactly as they are. They are two accounts the six cells cannot hold.
- **Extra cell one, the peer competitor.** Another team in the company also doing AI, or a sister business unit's digital group. He does not block you, so he is not a blocker, but he competes with you for the same budget and the same sponsor's attention. What he fears is your project becoming the benchmark, so that when he files for project approval next year, he is measured against your numbers. What he wins is a place he can write into his own reporting. First contact, in the opening week, book him yourself and show him your scope boundary. The exceptions queue is yours, and his piece you will not touch. Then ask whether he has something on hand you can reuse directly. The competitor becomes your first reuser, his name goes into your reporting, and yours into his.
- **Extra cell two, the predecessor's legacy.** Same structure as extra cell one, the difference is the first contact. This business unit was hurt by a digital project once before. What the predecessor left behind, reports nobody maintains, a system nobody logs into, one line of "that is what they said last time too," all of it is booked to your account, because to the front line you are the same kind of person. What they fear is it happening again. What they win is someone admitting the last time first. First contact, go find the real cause of death of the last project. The files and the people involved are in the building, one afternoon of asking settles it (Chapter 6). Then say the cause of death out loud in front of the front line. Get it right, and half the predecessor's debt moves off your account. Get it wrong, and they will correct you, and the correction itself is the start of intimacy.
## At Anchor & Helm: One Map, Three Payoffs
The night before the charter meeting, you drew the Anchor & Helm map (the full walkthrough is in [Template 5](../appendices/template-05-stakeholder-map.md)). Drawing it exposed two fatal blind spots.
**Blind spot one, Victor Reyes.** He is in no project meeting and on no email cc. But who goes in the "risk owner" cell? Who signs the security review, who takes the fall for a data incident? All him. He holds a real veto, a single vote that stops everything, and yet he is invisible. What he fears is concrete. Being bypassed, and still taking the fall when something goes wrong. AI itself comes second. In week 1 you got one thing right on instinct, and hand-delivered a pre-mortem memo to his desk. That memo assumed the project had already failed and listed the causes of death backward, and item four said a security review that does not start until week 10 blocks the launch. He asked to meet you because of it (Chapter 1). What the map does is turn that instinct into structure. When "risk owner" and "blocker" carry the same name, bad news must always go to that person unprompted. Never wait for luck to blow it past him. The line from blocker to ally runs from this page to the security review in Chapter 12.
**Blind spot two, Linda Marsh.** The moment you wrote her name in the "actual user" cell, you realized that the person who decides whether this system lives or dies had not been invited to a single decision meeting since the MVP scoring. Of everyone sitting in the meeting room, not one would use this system daily after launch. She fears being replaced, taking the blame for AI's mistakes, and having twenty years of experience erased by one line of "the model decides automatically." What she can win is fewer chase emails, and her judgment written into the system instead of written out of it. The corrective move appeared in Chapter 4, the line "Linda's team, 2 hours a week" in the charter's business-side commitment clause. Now the secret can be told. That line was forced out by this map, not a flash of inspiration in the meeting.
**The third payoff is the standoff this chapter opened with.** On the evening of day ten, you go back to the map and look at the "data owner" cell. IT holds the key. The business owner of the exceptions data is Kevin. For ten days you have been talking to the permission system, and the key hangs in the trust network.
Going to Kevin empty-handed is useless. Why would he spend his power on your project? The pain of the monthly report, Kevin had raised to your face weeks ago, and you had filed it as "a favor you could do in passing" (Chapter 7). This week it flared up again. That spreadsheet he has a staffer export from the core system and stitch together by hand is wrong every month and questioned every month. The next afternoon, you carry your laptop over and sit down next to his staffer's desk. The files never pass through your hands, and the script runs on Claims' own machine. Forty minutes, and the field extraction logic already in the MVP generates the spreadsheet automatically, and along the way you fix two definition errors he had never managed to find. The etiquette of helping differs inside a company and outside. People from the Digital Center can already reach plenty of data in this building, so not touching unauthorized data is not a signal for you, it is the floor. Kevin watches three other things. You did not quietly pull that spreadsheet back to the Digital Center to build your own store. The data never left their machine, start to finish. Those two definition errors you did not turn into talking points in other meetings. The errors in Claims' monthly report are known only to Claims. The script stays on his staffer's computer under the staffer's name, and when next month's report comes out clean, the credit goes to Claims. All you take away is a copy of the script. You do not mention the data request.
On Friday, Kevin stops the head of the IT data team in the hallway and speaks in his own language of power. "Exceptions are my department's data. The project team wants a read-only de-identified view. I sign off, open it before Monday." Monday morning, the access arrives. A standoff ten days could not break, forty minutes and one sentence.
This is hardly scheming. You helped Kevin win a fight he was already fighting, a double deposit into reliability and intimacy. Those forty minutes also let you see the true face of the core system's fields, so the first-hand material for Chapter 9's data reconciliation is already in your hands. Let the person with the power speak for you, in his own language. It beats shouting ten times yourself.
Inside a company there is one more gate. Your help has no price tag. In Kevin's eyes those forty minutes were free, and next week his staffer will show up with five other spreadsheets, each one "in passing." Helping is an investment only when it has an exit. So before you touch anything, settle the exit for this one. Is it one-off, the script stays with them and they run it themselves from now on, or does it become a long-term commitment? Do the former. Do not take on the latter on the spot. It goes into the intake list (the gate where requests come in) and waits its turn (Chapter 25), and the waiting is itself the price. Kevin's exit this time was clean. The script bears his staffer's name, and the monthly report no longer passes through your hands. In "help him win once first," the word "once" is the gate.
## Framework Two: The Trust Equation Self-Check
The equation is easy to remember. Honesty is the hard part. Each of the four variables gets three behavior-level questions. Do not ask "do you think the business side trusts you." Ask only about behavior that can be checked.
**Credibility**
1. In the past two weeks, have you said to someone's face "I do not know, I will have an answer by X," and then delivered on time?
2. Of your last three judgments, how many carried evidence from this business unit's floor (numbers, cases, users' own words) rather than industry boilerplate?
3. Has the business side quoted you in a meeting you were not in?
**Reliability**
1. Of your last five "you will have it Friday" commitments, how many landed on time?
2. Are meeting action items sent in writing within 24 hours, or "written up later"?
3. Has there been one time you shrank a commitment you could not keep, early and unprompted, instead of explaining on the deadline?
**Intimacy**
1. Has anyone told you something at the level of "I am only telling you this"?
2. Have you once said the tension in the room out loud, to people's faces?
3. Has the front line complained to you about their own department or their own boss? The jargon (aged claims, chase emails, "that spreadsheet") you already know, so that test does not count for a colleague.
**Self-orientation (the denominator, lower is better)**
1. In your last meeting, how many times did you say "our system / our plan" versus "your claims / your metrics"?
2. When the business side suggests something that would shrink the project's scope, is your first reaction defense or curiosity?
3. When did you last advise the business side "let's not do this yet"? If never, the denominator is quietly rising.
The usage is simple. Run it on yourself every two weeks. Read the answers for the gaps, not the total.
The scattered questions you cannot answer, each one is the next behavior you should deliberately create, and one action this week covers it. When all three questions under one variable come up empty, the nature changes. That is a line you have not been working at all, one action is not enough, and it goes on next week's calendar. The full self-check is in [Template 5](../appendices/template-05-stakeholder-map.md), next to the map.
This sheet does not retire at launch. Trust is a permanent account for you. It is not settled on launch day, the business side is still here next year and so are you, and every slow response after launch is drawn from the same account. One unanswered alert in month 14 can empty every bit of reliability the charter period banked. What Linda's team remembers will not be those forty minutes, it will be the one time nobody picked up. So after launch the way you deposit changes. It is no longer say Friday, deliver Friday. It is alerts get answered, there is a name on the oncall (on-duty) rotation, someone shows up at the retrospective. Who owns each of these, the responsibility transfer agreement (Chapter 22) spells out. Any cell it leaves unclear is charged to your account.
## Failure Modes
**1. Cultivating only the sponsor.** Inside a company this is called cultivating only the boss, and there are two bosses, Grant, who gives the mandate, and Owen, who runs your review. Excellent relationship with Grant, and Linda's team cannot say your name. The root is asymmetric feedback. The sponsor and your manager talk, give positive feedback, one holds the budget and the other your performance review. The front line gives you only silence and suspicion, and people drift naturally toward where the feedback is. Inside a company this structure is only stronger. The boss sets your review, the front line does not. But on settlement day the roles swap. The sponsor gives you the project. The front line gives you the outcome. After launch Grant will not log in, and neither will Owen. Linda will, or will not. Whether she logs in depends on how much trust you banked with her.
**2. Delaying bad news.** "The data is worse than expected" has been sitting in your mouth for two weeks. You booked reporting bad news as a credibility expense. You booked it backward. Bad news delivered early, unprompted, with a plan, is a double deposit into credibility and intimacy. Every concealment is a withdrawal from reliability, with interest. Bad news does not disappear, it only appreciates. Wait for the business side to find it themselves, and the cost is ten times higher, and every piece of good news from you afterward is discounted along with it.
**3. Self-orientation leakage.** "Our system" shows up in your sentences more and more. Your KPI (a promotion, your department's annual KPI, next phase's resources, your name near the top of the year-end report) is real and cannot be hidden. It leaks out through language, and the frequency of "our system" is the leak meter. The business unit smells the same thing every time. What you are fighting for is whose name this system ends up under. Their reaction is the one this chapter opened with. Nods in the meeting, no cooperation after. The denominator is the only division in the equation. However high the other three climb, one whiff of self-interest zeroes them out. The fix is to put the motive on the table, since nobody believes feigned selflessness. "If this project succeeds, of course that is good for our team. But it only counts as succeeding when your handling time really drops 30%." A motive said out loud is no longer a leak. It is a boundary.
## Next Monday
1. Use [Template 5](../appendices/template-05-stakeholder-map.md) to draw the stakeholder map for the project on your desk. Real names, six cells plus the two extras, "What They Fear / What They Win" filled in. Find the name that has never been in the meeting room.
2. Send that name something, a pre-mortem, a one-page risk list, so he knows you wrote what he fears on page one.
3. Run the 12-question self-check, pick the weakest variable, and create one concrete behavior this week to shore it up.
4. The "cannot get X" you are stuck on right now, ask three things instead. Who is X's business owner? What fight is he fighting? Can I help him win once first, and where is the exit for that once?
**Want an agent to get you started?** In the repo you set up following [Start Here](../index.md), paste this to your coding agent:
```text
In the repo/ directory of the the-last-mile repository, help me with the Chapter 5 Next Monday actions. Copy
templates/stakeholder-map/stakeholder-map.md into the working directory I name, then walk me through the six cells plus the
two extra cells one by one, asking for real names and "What They Fear / What They Win." I dictate, you fill in. Mark any
cell I cannot answer as [TBD] and remind me that cell may be the person who has never been in the meeting room. When it is
filled in, run python3 templates/stakeholder-map/remind.py to see the reminder output on the built-in sample, then run it
again pointed at my file. I answer the 12-question self-check myself, do not judge the weakest variable for me.
If any command errors, stop and show me the output.
```
---
## Chapter Kit
- **Judgment frameworks.** The six roles of the stakeholder map (decision-maker / actual user / data owner / risk owner / maintainer / blocker, real names + What They Fear / What They Win, two extra cells for the peer competitor / the predecessor's legacy); the Trust Equation self-check (4 variables ร 3 behavior-level questions)
- **Templates.** [Template 5](../appendices/template-05-stakeholder-map.md) (Stakeholder Map), six-role map + "fear / win" prompt library + contact plan + Trust Equation self-check
- **Key judgments**
- "The org chart draws the permission system. Day to day, the company runs on the trust network."
- "At the start of an AI project, the trust balance is negative, not zero."
- "The sponsor gives you the project. The front line gives you the outcome."
- "Let the person with the power speak for you, in his own language."
---
# 6 ยท Field Archaeology: Dig Out the Real Workflow and Its Tacit Knowledge
!!! info "Companion Templates"
๐ [Chapter Template](../appendices/template-06-field-archaeology.md) ยท ๐ [Template Library](../appendices/template-library-index.md)
> **The Challenge.** Interviews do not get you the truth, and the flowchart the business side gave you does not match reality. How the work actually happens, you do not know, and you are starting to suspect that nobody in this company knows the whole of it.
>
> **What You Will Be Able to Do.** Reconstruct the real version of a workflow with the six steps of the workflow truth loop. Know the three places where tacit knowledge (the knowledge people hold but cannot state) hides. Use the three-layer probing method to turn "I can tell at a glance" into rules that can be written down and verified.
---
## Week 3, Wednesday, 8:20 a.m., a Folding Stool
Over tea on Thursday of week 1, you and Linda Marsh agreed on "come sit with me for half a day next week." That appointment got pushed twice, both times for the same reason, backlog. By the time you actually set a folding stool behind and to the side of her desk, it is Wednesday of week 3, the charter meeting is done, and the signatures are still a few days off. Two reschedules are data in themselves. Front-line time is the most expensive resource in this company, more expensive than Grant Whitmore's. Grant's calendar has open slots. Linda's queue does not.
At the second reschedule, a road opens to you that an outsider does not have. Go to Kevin Doyle, have him say a word to Linda's supervisor, and the half day gets scheduled on the spot. That road really works, and the price is really high. A half day bought with a supervisor's authority is a half day Linda will treat as an evaluation. Every move you see from behind her shoulder will be a performed move, that Excel file will stay closed all day, and that spreadsheet is the one thing you must see today. The test is the reason, not the count. If the reason changes every time, you are being brushed off and escalation is warranted. If the reason is backlog every time, the place you want to look is exactly where it hurts most, so wait. Consider escalating only after three or more in a row, and when you do, talk about schedule protection, not attitude, and land it in the business-side commitment line of the charter (Chapter 4).
It is not that you have not seen her these two weeks. Since week 2 her team has spent 20 minutes a day annotating cases for you. But the Linda in the annotation room is answering your questions, and the Linda at her desk is doing her job. Those are two different people, and today you are here to watch the second one.
In your hand is a printout, the *Exception Claim Handling Process* Kevin gave you last week.
> 1. System assigns the exception claim โ 2. Review document completeness โ 3. Notify for missing documents โ 4. Second review and approval โ 5. Write back to the core system, close
Five steps, straight arrows. Your plan is to use it as a checklist and see which step Linda gets to.
At 8:31 Linda logs in. The first move is on the paper. She opens the core system and scans the exceptions assigned overnight. The second move is already off it. She opens an Excel file on the shared drive, enters the new claims one by one, then sorts today's order by her own color codes in the sheet, not by the system's assignment order. You swallow the question, write a line in the friction log (the Chapter 0 list of "where reality and paper do not match"), and keep watching.
By 9:40 your count stands at step fourteen. Below is what you wrote down on the spot. Steps in bold happen only inside that Excel file, steps marked (phone) get done by a phone call, and the rest are moves that exist in the official five.
> 1. Scan newly assigned claims in the core system;
> 2. **Enter the new claims into the Excel tracker**;
> 3. **Sort today's order by the Excel color codes**;
> 4. Check the document list claim by claim;
> 5. For missing documents, search the inbox by claim number and surveyor name. The documents often arrived long ago, nobody attached them to the system;
> 6. **If the inbox has them, download and attach to the system, flip the Excel mark from yellow to green**;
> 7. If the inbox does not have them and the surveyor is someone she knows, call and ask (phone);
> 8. **If not, send a chase email, copy the supervisor, and add 1 to the Excel "Chases" column**;
> 9. For claims with doubtful amounts, scan the claim report and the photos, ten seconds, flag or do not flag;
> 10. When unsure, call the surveyor to verify details from the scene (phone);
> 11. **For claims judged abnormal, mark the Excel "Risk" column red and write a word or two only the team understands**;
> 12. When done, go back to the core system and change the status, but that waits for second review, so **Excel has a separate "Actual Status" column**;
> 13. Filter out claims older than seven days and post a few in the department group chat to chase them;
> 14. Before leaving, save a dated copy of the Excel file, "the shared drive lost it once."
Fourteen steps. Six of them happen only in an Excel file that does not exist on the official flowchart, and two are phone calls to a surveyor she knows. Not one of the official five steps is wrong. They are all there, like a riverbed. But the shape of the water is nowhere on the paper.
At the 10:15 tea break you ask the question you have held for two hours. "That spreadsheet, is it in the system?"
Linda's answer, which you later quoted verbatim in your memo to Grant: "The status in the system is for the people upstairs. That sheet is our own." The sheet has columns the core system does not have, actual status, chase count, risk marks, notes in team shorthand. This team is the only one in the company that maintains it, and has for six years. This sheet is the source of truth, the data everyone actually trusts and acts on. You write down the most expensive line of the half day in the friction log. The source of truth for exception claims is not in the core system. It is in an Excel file on the shared drive. The data reconciliation in Chapter 9 starts from this line.
## Why This Is Hard: Organizations Are Systematically Wrong About Their Own Workflows
Nobody lied to you. The flowchart Kevin gave you was sincere, and it is also "correct" in the sense that audit and training use the word. The problem is that every organization holds three versions of a workflow at the same time.
- **The SOP version** describes what should happen. The SOP's job is compliance, audit, and onboarding. By design it is not responsible for describing reality.
- **The reporting version** is what management sees. Every step up the reporting chain smooths away a bit more of the workarounds (the homegrown ways of getting around the system). The front line does not report "that sheet," because reporting a workaround tool means confessing a violation, or inviting a round of "systematization," and both are losses. A rational front line stays silent forever.
- **The actual version** lives only at the desk. It lives in exception handling, private spreadsheets, and the feel of veteran staff, and it is alive. Surveyors change, fraud methods change, and it quietly updates every month without telling anyone.
The three versions coexist indefinitely and never correct each other. This is not an illness of this company. It is the normal state of every organization. The trouble is that an AI system must embed in the actual version, which is what the workflow gap among the five gaps in Chapter 1 means. Build the system from the official flowchart and it will wait at step 2 for a "document completeness result," while in reality that result is scattered across the inbox, phone calls, and three columns of an Excel file. The official flowchart is not wrong. It belongs to a parallel universe. Build the system from it and you are writing software for a parallel universe.
This is also why interviews fail. Ask a manager and you get the SOP version plus the reporting version. Ask the front line "what is your process" and you still get the SOP version. When people are asked about "the process," they recite the one they were taught, not the one in their hands. She was not hiding anything. She herself never counted the fourteen steps as "a process." That was just "doing the work." The real workflow is in nobody's words, only in their actions. There is one way to get it. Go and watch. For you this sentence carries one more layer. You have been at this company for years and thought you already knew how the claims department works, and the version you knew is precisely the one that traveled up the reporting chain to you.
The internal reader's biggest risk is skipping this chapter. The reason sounds solid. I am not an outsider, why would I need to carry a folding stool and sit for half a day. But look back at where this dig started, and the honest answer is not flattering. Why did you first sit beside Linda's desk only in week 3? Because you have been at this company three years and thought you knew how the claims department works. In those three years you saw their monthly reports, heard their fifteen minutes at the business review, took the tickets they filed, all of it reporting version and SOP version. You never saw that Excel file, because by design it does not flow in your direction. The two reasons the front line does not report it are only stronger in front of you. You are exactly the person who would systematize it out of existence.
Familiarity is the only enemy in this chapter. The outsider knows he knows nothing, so he goes to watch. You know a lot, so you go to meetings. And sitting for half a day is far cheaper for you than for an outsider. Two floors down, no appointment, no background briefing, no week spent explaining why you are sitting there. If the cost is that low and you still do not go, that is not a judgment. That is familiarity deciding for you.
Presence is an asset, but it only counts once it is cashed into action, and this chapter has three actions you can take now that an outsider cannot. First, you need not do the half day in one go. You can go any time, so split it into five one-hour visits on different days. Monday's backlog and Friday's wrap-up are not the same workflow. An outsider can only buy four consecutive hours. You can buy five different cross-sections. Second, go through the review archives. In the three-layer probing method below, the third layer asks Linda whether she has ever flagged a claim and misjudged it. From memory she gives you one claim from the year before last, while the same batch of claims has a complete record in the QA reviews and the customer complaint log, which you have permission to pull. Lying there are the counterexamples she cannot fully remember. Third, ask for the first version of that Excel file. It has been maintained for six years. The person who built it may still be at the company, and the email proposing it may still exist. Why a workaround tool got built is page one of a requirements document written over six years, and you can get it.
## Prior Art, and What AI Changed
Going to watch has two mature traditions. You need not remember the names, only one thing. Watch first, ask second.
**The consulting tradition, fact-driven, firsthand first.** The first rule of McKinsey's problem-solving discipline (as retold in Ethan Rasiel's *The McKinsey Way*) is fact-based, and facts have ranks. Firsthand observation in the field outranks the business side's reports, which outrank industry press. Facts should be taken from where they are produced.
**The anthropological tradition, observe before you ask.** From Malinowski's fieldwork to the contextual inquiry that the design world engineered out of it (Beyer & Holtzblatt), the core is the same. Enter the field as an apprentice, with the user as the master. Watch the master work first, anchor questions on the concrete action that just happened, never on hypothetical situations. "Why did you do that" is always asked after seeing it. Reverse the order, ask before you watch, and all you get is a rationalization invented for you on the spot.
**What did AI change? The cost of writing up the dig collapsed. The digging itself has not moved an inch.** Half a day of shadowing (following a person and watching them work. In this book, shadow means only a person following a user to watch real work. It has nothing to do with the engineering practice of running a system in parallel on live traffic with its output disabled, "shadow mode," and this book does not use the word for that) used to leave scrawled notes that took an evening to clean up, and clustering exception types out of three thousand tickets took a week. Now AI does it in half an hour, transcribing, structuring the friction log, clustering the reasons claims get stuck. All the time saved should go back into the field.
But AI cannot shadow. "Seeing that the Excel exists" is in no dataset. That sheet has no API, no documentation, no entry on the IT asset register. The only evidence of its existence is Linda's second move at 8:31 in the morning. The cheaper the information AI can organize, the more expensive the tacit knowledge that only a person on site can get. It is the raw material for the eval in Chapter 11, and the real moat of this system. Models are a public good. Linda's five rules are not.
## The Framework: The Workflow Truth Loop
Turning "go and watch" into a repeatable method takes a loop of six actions.
| Step | Action | Key Discipline |
|------|------|----------|
| **observe** | Be present and watch them work, do not interrupt | Write questions down, do not ask on the spot; keep the friction log as you go |
| **shadow** | Follow a person for half a day | Follow the person, not the process. A process does not open Excel, a person does |
| **trace** | Follow one claim end to end | Across people, systems, and tools, count how many hands and how many systems it passes through |
| **ask** | Ask about exceptions | "When would you **not** do it this way?" anchored on the concrete action you just saw |
| **reconstruct** | Draw the real process | Including workaround tools and verbal coordination, not one step prettified |
| **validate** | Have the person correct it | Ask "where is it drawn wrong," not "is it right." The former gets corrections, the latter gets politeness |
Three rules of use. **First, the order cannot be skipped.** The most common cheat is to skip the first three steps and go straight to ask, and then what you get is the recited version again. **Second, it is a loop, not a pipeline.** Validate exposes new exceptions, and new exceptions deserve a new round of observe. Run it two or three times, and the rate of convergence tells you when it is enough. **Third, trace speaks through a claim.** Shadowing looks at one person's day, trace looks at one claim's whole life, and only the two cross-sections together make a workflow.
The Anchor & Helm instance. You pick the claim ending in 4471 and follow it end to end. The claim report is in the core system, the survey photos are email attachments, the customer's chase calls are in the call center ticketing system (no join key to the claim, found by searching the phone number), and the actual progress is in Excel. One claim, four homes. This one trace is worth ten pages of interview notes.
The loop digs for the workflow, but the richest vein hides elsewhere. **The three hiding places of tacit knowledge** are these.
1. **Exception handling.** The SOP covers the main path. All the judgment is on the branches. The probe is "when would you not do it this way?"
2. **Workaround tools.** Private Excel files, sticky notes on the desk corner, personal email folders, small group chats. Like shadows, these tools appear on no architecture diagram. The test is anything open on the screen that is not on the IT asset register. Each one is a requirements document, written over six years.
3. **"I can tell at a glance" judgments.** The expert has compressed the rules into intuition and cannot state them directly. The signal is three trigger phrases, "I can tell at a glance," "from experience," "hard to say." A trigger phrase is an outcrop of the vein. Mine it with the three-layer probing method below.
The workaround tools cell holds one more question for you than for an outsider. Linda's sheet sits on the shared drive, with actual status, chase counts, risk marks, and notes only the team understands, and very likely fields that can be matched to the person who filed the claim. It has never been through a data review, and it is not on the IT asset register. An outsider sees it, writes it down, puts it in the report, and that is the end of it. You see it, and you are an employee of this company whose job description says risks found must be reported, and half an hour ago you told her, I am not here to evaluate you.
The rule still holds. See a noncompliant workaround, note it, do not call it out ([Template 6](../appendices/template-06-field-archaeology.md), code of conduct rule 2). But it governs that half day, not the half year. You do not correct on the spot because your role while present is apprentice, and correcting would make everyone you observe from then on start performing. Whether to report it afterward is a separate question. Noting without calling out is not permanent secrecy. Merge the two and you either wreck this dig on the spot, or become the person who knew and said nothing six months later.
Draw the line before you go into the field. What must be reported is real compliance risk. Data that can be matched to an individual sitting on a personal computer or in a personal mailbox, outbound sending that bypasses approval, operating under someone else's account, a step that requires a record being skipped. These four are not efficiency problems, they are institutional problems, and if you know and stay silent, your silence will count as consent when something goes wrong. The great majority of the rest is not in this class. Sorting in Excel, color-coding priority, calling a surveyor you know, these are not violations. They are scars left by underdesigned process, and their proper destination is the real flowchart and the data reconciliation of Chapter 9, not a compliance report. When you cannot tell, the test is whether this would become a problem if audit saw it, not whether it looks proper.
If you really must report, one more discipline. Tell the person first, face to face. It can be short. Your sheet has fields that can be matched to individuals, and this one I have to report, and at the same time I will write down why you need this sheet and what breaks if it is shut down before a replacement exists. That is not an easy thing to say. But you and Linda will still be in meetings together next year. Report around her once, and the second time nobody opens any spreadsheet in front of you. The reverse also holds. One report that was announced first, and that really spoke for her in the write-up, tells the whole claims department that opening a sheet in front of you is safe. An outsider cannot build that credit. You can, and you get exactly one chance to build it.
## At Anchor & Helm: A Ten-Second Judgment, Five Rules
At 11:20 Linda pauses on a claim for under ten seconds and flags it red. Half an hour earlier Sam, the team's new reviewer, came over with a laptop to ask about a similar claim, and Linda glanced at it. "Reported three days late, and it is that repair shop again. Flag it." Sam asked how she could tell. She said, "You see enough of them and you know." That is how this team's judgment gets passed down. Verbal, at random, ten seconds at a time, never on paper.
At 11:40, the last half hour, you start probing. Asking "why did you flag it" straight out is useless. You have already heard her first response, "I can tell at a glance." The three-layer probing method starts here.
**Layer one, anchor on an instance, replay the actions, do not ask why.** "That claim just now, 4471, you went to the report time first, then opened the photos and flipped through two, and then you flagged it. Right?" She paused. "...Right. Rear-end collisions rarely go unreported the same day, and this one waited three days. Only five photos, a surveyor at the scene normally takes a dozen or more." Two rules surface. Where is the leverage? You replayed the actions she had just taken herself, and a stock phrase does not match actions.
**Layer two, compare, find the difference between two similar claims.** You pull up another rear-end collision from the morning with a similar amount. "This one is two thousand higher than 4471. Why did you not flag it?" "That one was at a major intersection, with a police report, and the repair shop is a dealer's service center." The amount is judged against the usual range for this line of coverage and this vehicle model, not in absolute terms. And "the repair shop" is a list she keeps in her head.
**Layer three, ask for boundary counterexamples.** "When would you not flag it no matter how high the amount?" "Have you ever flagged one and been wrong?" She thought about it. "I have flagged wrong. The year before last there was a claim that turned out to be the owner paying for the repair out of pocket first... so now I also check whether the same plate has come through in the last six months."
Three layers in, "I can tell at a glance" resolves into five rules of thumb.
1. **Amount deviation.** The repair amount is well above the usual range for the same coverage and vehicle model (judged relative, no hard threshold);
2. **Report delay.** The gap between the incident and the report is beyond the ordinary (a rear-end collision not reported the same day, reported three days later);
3. **Survey photo count.** Clearly fewer photos than a normal claim, or key angles missing;
4. **Prior claim linkage.** The same plate, filer, or phone number appears again within six months;
5. **Repair shop list.** The shop is on her rolling mental list of shops that "come up a lot."
On Friday you print the fourteen-step real flowchart and the five rules, put them on Linda's desk, and ask one question. "Where is it drawn wrong?" She circles two places. Step 13 (filter claims over seven days and post them in the group chat) has the wrong frequency. Tuesdays and Thursdays, once each. She does not do it daily. And the repair shop list, "it is not a fixed list. I crossed one off just last month." The weight of the second one, failure mode 4 will explain.
Where these five rules go, you can be told now. When you co-build the eval spec with Linda in Chapter 11, every one of them becomes an error category for the golden cases (sample cases with the correct answer settled in advance), and "missed risk" gets split along these five into five testable ways to fail. That is the step where tacit knowledge turns from interview color into a system asset. The half day's accounts are easy to settle. Four hours on a folding stool produced three things. The existence of the Excel file, the starting point of the data reconciliation in Chapter 9. The fourteen-step real process, the map of where the Action Queue embeds in Chapter 17. The five rules of thumb, the raw material of the eval in Chapter 11. Every key deliverable of this project later traces back to these four hours.
One more thing to arrange right now. The conclusions of these four hours will expire, and most likely you will be the one who expires them. Chapter 9 goes a layer harder. Your system will kill its own source of truth with its own hands. Of the fourteen steps, six happen only in Excel. After the queue goes live the system takes some of them over, the rest grow new workarounds around it, and nobody will come to tell you about the new workarounds either. The real flowchart you drew today starts going stale the day the system launches.
An outsider does not handle this. Next time he comes, he digs from scratch. You will not. You only do increments, so the cadence of increments is yours to set. Put it on the calendar. Go sit once in week 4 and once in week 12 after launch, one hour each, doing two things only. Count the steps again, and ask whether any new spreadsheet appeared this month. After that, once every six months, as the queue's periodic review (Chapter 17 expands on it). The real flowchart and the five rules both get a date and a next-dig date. A flowchart with no date is the same thing as an SOP three months later.
## Failure Modes
**1. Interviewing only managers.** Discovery (the stage of finding out how things actually are) ran three weeks, met eight directors and team leads, produced a thick requirements document, and never sat at a desk once. There are two layers of cause. Managers are the people you can book. They have calendars, meeting rooms, and a desire to talk, while the front line is "in the backlog." Worse, a manager's information itself comes from the reporting chain, so interviewing him is looking at the field through two layers of distortion, and everything you get is the should-be universe. One discipline. In discovery, the actual user must account for more than half of interview time. If you cannot book them, wait. Two reschedules, still wait.
**2. Treating the flowchart as the truth.** You get the SOP, treasure it, and design the system state machine straight from it. The SOP's institutional functions (compliance, audit, training) mean it describes what should be. The organization has motives to maintain the should-be version and no mechanism to update the actual one. It was never a map, so it cannot be out of date. The right use. The SOP is the starting point of the dig, not the end. Print it, bring it to the field, and hunt specifically for where it differs from reality. Every difference is a line in the friction log.
**3. Asking "what features do you want."** The requirements interview becomes an ordering session, and you bring back a wishlist. Users can only imagine the future in the language of their current tools. Ask Linda what she wants and she will say "add an automatic reminder to that sheet," because that sheet is the whole world she has seen. Users can order from the menu, but users are not the chef. What she reports is a solution projected onto the old tool, and the problem itself is still buried. And "what do you want" is a hypothetical question, and the anthropological discipline is precisely not to ask hypotheticals. The right path. Observe the problem (that inbox search you watched this morning, hunting every day for documents that arrived long ago), and let the solution grow between you and her.
**4. Hardcoding tacit judgment.** You get five rules and that night write them as five if-else branches and ship. Rules of thumb are a snapshot of living rules. The repair shop list lost an entry just last month, fraud methods evolve, Linda's rules move with them, and if-else does not. Three months later the rules have drifted, the system still flags from the old snapshot, and because "the rules came from Linda," nobody questions it. Hardcoding also quietly moves the decision from the person to the system, and when it errs, there is nowhere to put the responsibility. That is exactly why the Chapter 0 MVP queue keeps a "Human Call" column. The right path. Make the rules explicit, put them in the suggestion and reason columns, keep the decision with the person, and record overrides in the decision trail. The decision trail is the drift detector for rules (Chapter 17 expands on this mechanism).
## Next Monday
1. Find your project's actual user and book half a day of shadowing ([Template 6](../appendices/template-06-field-archaeology.md) has the opening script and a half-day schedule). Do not be discouraged by a reschedule. The reschedule itself is information. Where the most expensive time is, that is usually where the most painful workflow is.
2. Print the official flowchart and take it to the field as your dig map. Count the real steps. Look for tools open on the screen that are not on the IT asset register. Find one and you have found a requirements document the user wrote over several years.
3. Next time you hear "I can tell at a glance," use the three-layer probing method (anchor on an instance โ compare โ boundary counterexample) to recover the rule, write it down, and take it back for the person to circle the mistakes.
4. Go through your existing requirements document and tag the source of every line. Manager interview, SOP, or firsthand observation? If the first two make up more than ninety percent, you are writing software for a parallel universe.
**Want an agent to get you started?** In the repo you set up following [Start Here](../index.md), paste this to your coding agent:
```text
In the repo/ directory of the the-last-mile repository, help me with the Chapter 6 Next Monday actions. Copy
templates/field-archaeology/friction-log.md into the working directory I name, and leave it blank for now. When I am back
from shadowing I will paste my dictation or notes to you. Organize them into a structured friction log as described in
prompt-organize.md, keeping the original wording in every entry, and do not add steps you did not hear. While organizing,
mark every "I can tell at a glance" style judgment and probe me on each one with the three layers (anchor on an instance,
compare, boundary counterexample). The recovered rules are written by me, and the last step is taking them back for the
person to circle the mistakes, not handing them to you.
If any command errors, stop and show me the output.
```
---
## Chapter Kit
- **Judgment frameworks.** The workflow truth loop (observe โ shadow โ trace โ ask โ reconstruct โ validate); the three hiding places of tacit knowledge (exception handling / workaround tools / "I can tell at a glance" judgments)
- **Templates.** [Template 6](../appendices/template-06-field-archaeology.md), the Field Archaeology Kit, half-day shadowing guide + friction log template + three-layer probing script
- **Key judgments**
- "The official flowchart is not wrong. It belongs to a parallel universe."
- "Users can order from the menu, but users are not the chef."
- "AI cannot shadow. 'Seeing that the Excel exists' is in no dataset."
- "The SOP is the starting point of the dig, not the end."
---
# 7 ยท The Opportunity Screen: Five Questions That Kill Most Ideas
!!! info "Companion Templates"
๐ [Chapter Template](../appendices/template-07-five-questions.md) ยท ๐ [Template Library](../appendices/template-library-index.md)
> **The Challenge.** The business side and your boss have a dozen AI ideas stacked up, and every one of them sounds workable. Which one, and on what grounds? The harder half is telling the sponsor to his face that the idea he forwarded you should not be built.
>
> **What You Will Be Able to Do.** Score any AI use case on the five questions (Pain / Data / Decision / Risk / ROI). Screen out the unqualified opportunities in one afternoon with the elimination table. Report the results to an executive on one page, including the elimination of the one he forwarded himself.
---
## Turning the Clock Back to Before the Charter Meeting
The readout (the memo that states the conclusion at the end of the Field MVP) has just been delivered. You have "the direction works." You do not have "what to build." Inside a week there are three candidates on your desk.
The first comes from the top. Thursday evening Grant Whitmore forwards a board email. Inside it is a piece of industry news, a peer has launched an "intelligent service bot," with a board member's own words attached. "Can our service line get one of these too?" Grant added four words. "Evaluate this for me." The second comes from the corridor. Kevin Doyle stops you. "Keep pushing on the exceptions thing. Separately, our monthly operations report takes two or three days of stitching data together by hand. Can AI generate it? That one pays off fast."
The third is the one the Field MVP itself pointed to, an action queue for exception claims (the queue that lines up the next action for every claim), carrying ten scored cases and three unsafe marks (output a real user judged should never have appeared). Turn the clock back and this is exactly the two weeks before the charter meeting in Chapter 4.
All three "can be built." With today's model capability there is almost no proposal that "cannot be built," and that is precisely the problem. There are already enough ideas. What you need is a discipline for eliminating them.
## Why This Is Hard: Ideas Became Free, Validation Did Not
Chapter 1 covered demo inflation. Demos got so cheap that being impressive no longer proves anything. The same inflation has an earlier form inside an organization, idea inflation. Proposing a software requirement used to mean the proposer had at least worked out the workflow and the fields, so the proposal itself carried a cost. Today anyone who finishes reading one news item can produce an AI use case that sounds like it holds up. The cost of producing an idea has fallen close to zero.
The cost of validating one has not dropped a cent. Validating a use case means a pilot (real users, real use, inside a controlled scope). Three months, a team, a hard-won stretch of data access, and the most expensive item of all, the organization's patience for "this AI thing." When the first AI project fails, what burns is AI's name at this company, and the next project breaks ground on the rubble.
An open funnel mouth, an expensive funnel bottom, and no gate in between. That is the assembly line enterprises use to mass-produce zombie AI projects. Worse, the bill for picking the wrong use case arrives three months later, and by then nobody records the cause of death as "we picked the wrong use case." They record "the model is not good enough," "the data is garbage," "the users would not cooperate."
So this judgment is worth writing as one sentence. Picking the use case is the deliverer's single highest-leverage decision, higher than any architecture decision. Get the architecture wrong and you refactor. Get the prompt wrong and you edit it. Get the use case wrong and three months go to zero, with trust overdrawn.
## Where the Right to Screen Comes From Inside a Company
Answer a prior question first. On what grounds do you decide that a business unit's idea is not to be built? A team from outside can say "I took on this one thing, not every idea you have." You cannot say that sentence. Claims' asks belong to Claims. You are not their superior and you are not their review committee. Veto a department's proposal on your personal professional judgment and it works the first time. The second time they go around you. The third time they go straight to Grant Whitmore. The right to screen has two sources, and neither one is granted by you to yourself.
**One, the intake gate.** Every AI idea comes in through the same entrance, the same five-question sheet, the same elimination rule, hung on the company's existing project approval review or requirements review. The gate's force does not come from your judgment. It comes from which meeting it hangs on, who backs it, and whether it is written into policy. Behind the gate you are only the person who fills the sheet in. What gets eliminated is the idea that does not pass the sheet, not the idea you dislike. Chapter 25 turns all of this into organization-level intake discipline. This chapter is its hand-built version.
**Two, the resource pool the sponsor (the executive who funds it and makes the call) authorized.** What Grant gave you was never "you may veto other people." It was "what this team works on first this quarter is set by this sheet." You hold no veto power. You hold scheduling rights. Say "we are not doing this" and you have no standing. Say "this team is not doing this this quarter, on these three pieces of evidence" and you have standing.
Before the gate exists, the substitute is repetition. The same sheet, the same rule, the same reporting format. By the third time, the rule starts to exist ahead of the individual case, and that is the smallest form of an institution.
## Prior Art, and What AI Changed
On the question of which opportunities are worth the investment, sales thought it through forty years before engineering did.
**SPIN's four-question structure.** The large-sale questioning method Neil Rackham set out in *SPIN Selling* (its thinking paraphrased here) has four layers, situation, problem, implication, need-payoff. The most valuable one is implication, pushing "this is pretty annoying" into "what happens if it is not fixed? How much do you bleed a month? Who takes the blame for it?" Most proposals die on that question. The pain is real, but not fixing it kills nobody.
**Solution Selling's qualification discipline.** Qualification is screening an opportunity for fitness before work starts, and not chasing what does not qualify. The Michael Bosworth line of sales methods has one counterintuitive iron rule (paraphrased). The most expensive mistake in sales is chasing the wrong deal, and losing a deal ranks behind it. Burning three months on an opportunity that does not qualify costs ten times what being turned down on the spot costs. So the fitness standard goes up front, and unqualified opportunities are dropped on purpose.
Move that iron rule inside a company and both ends need converting. There is no "losing a deal" here. What corresponds to it is turning down an executive's idea to his face, and the cost is one person's goodwill, payable immediately, repairable with one report and one re-evaluation condition. What corresponds to "chasing the wrong deal" is taking on an unqualified idea, and the cost is three months of the team's capacity plus AI's name at this company, payable in three months, unrepairable, and the next project is still yours to run. Set the two bills side by side and qualification's arithmetic is more lopsided inside than outside. Yet it is harder to enforce inside, because the goodwill bill is the one you pay today and the other one the whole company pays next year.
Both traditions ask whether it is worth doing, and both assume that whether it can be built is not in question. Traditional software behaves deterministically. Install it and it runs. An AI use case gets no such default, which adds two hard gates traditional qualification never had.
- **Data readiness**, because data does not become ready just because the idea is appealing. The idea is free. The raw material is not.
- **Error tolerance**, because a probabilistic system will be wrong. The question shifts from "is it any good" to "how many times may AI be wrong at this step? Who backstops it when it is?" The same model is an assistant where someone backstops it and an incident where nobody does.
Turn the two gates into questions and you get three. Data readiness maps to Data, which asks whether the raw material is there. Error tolerance does not fit into one question and splits into two. How many errors are allowed and who backstops them is Decision. How far an error travels afterward and whether it can be pulled back is Risk.
The implication SPIN pushes for is the Pain here. The "is it worth doing" both traditions ask is the ROI here. Pain and ROI are inherited from the prior art. Data, Decision, and Risk are the homework the AI era added. Together, five questions.
## The Five-Question Framework and the Elimination Table
The five-question framework, five test questions for judging an AI use case, each scored 1 to 5. Read the table by looking first at the third column, the test question, and the fifth column, what a low score looks like. Leave the fourth column for when you come back to score.
| # | Question | Test Question | What a 5 Looks Like | What โค2 Typically Looks Like |
|---|-----|----------|----------|----------------|
| 1 | **Pain** | Whose action is bleeding? Is the bleed rate (loss per unit of time, hours, complaints, fines, churn) measurable? | A specific person + a specific action + a measured bleed rate; leave it and someone keeps taking blame or paying out | Cannot name the person or the action; or after the follow-up "what happens if it is not fixed," the answer is "not much" |
| 2 | **Data** | Does the data this action needs exist? Can you get it? Can you trust it? (A pre-check. The full grading is the data fitness ladder in Chapter 9, the grading of whether data is fit enough, and that ladder and the outcome ladder in Chapter 1 are two different ladders) | Data exists, the business owner is named, samples are in hand, spot checks hold up | The key data does not exist, or there is no realistic path to getting it |
| 3 | **Decision** | Which decision point does AI enter? How many times may it be wrong there? Who backstops it, and how fast does the error surface? | It enters at the advise layer; the backstop is named; errors are visible on the spot and can be overridden | Real-time and outward-facing with nobody backstopping; or a decision with near-zero tolerance handed to AI to execute directly |
| 4 | **Risk** | What does the worst single output cause? Does it reach only inside, or customers, regulators, money? Can the loss be closed out? | The worst output's loss stays inside and can be closed out, and it sits outside the agreed red lines | One bad output reaches a customer or a regulator, and it cannot be taken back |
| 5 | **ROI** | What is the arithmetic that turns the outcome metric's improvement into money? Who owns that bill? | An outcome metric + the conversion arithmetic + a person willing to carry that number in his own reporting | Only activity metrics ("it launched," "92% accuracy," numbers that say only what was done). Chapter 1 covered why they do not survive the boardroom |
The elimination table is the summary of the five-question scores, plus one non-negotiable rule.
> **Any question scoring โค2 is eliminated outright. No weighting, no averaging, no "considering it in the round."**
This is weakest-link logic, not sum logic, because the relationship among the five questions is multiplication. A use case scoring 1 on Data cannot be built however much Pain there is, the raw material will not feed in. A use case scoring 1 on Risk will not be allowed to launch once built. The only function of a weighted average is to translate "not feasible" into "ranked low," and a project ranked low still gets started in a year with budget to spare. A weighted average is the excuse you reach for when you do not want to eliminate anything.
A 3 is a conditional pass. The evidence is not hard enough yet, so write the validation action that closes the gap into the sheet, with a deadline. Not met by the deadline, it drops to 2. Elimination is not a death sentence either. Every eliminated candidate gets one line of re-evaluation condition, what evidence, if it appears, makes it worth running the five questions again. "Not now" needs an exit before elimination can be enforced at all.
Scoring takes evidence only, numbers, users' own words, data samples. The five questions burn evidence, not cleverness, and the evidence is all the stock you built up in the earlier chapters, the MVP scoring sheet, the friction log, the informal annotation that started in week 2.
> **Any score with a blank evidence column is treated as a 2.**
The **pass / concern / unsafe / useless** scale from Chapter 0 is for a real user to react to a single output. The five questions' 1 to 5 is for you to find an opportunity's weakest link. Two scales, two uses. Do not mix them.
The internal reader holds two things a person taking projects from outside cannot get. Both have to be written as actions, not as consolation.
One, ask the implication before the proposal takes shape. Run into the proposer in the corridor and ask in passing, "if this is not fixed, what does it look like three months from now, and who suffers?" Half the ideas fall apart in the other person's own answer, never entering the sheet, never entering a report, without offending anyone. The cheapest moment to eliminate an idea is before it has become a forwarded email.
Two, a record of past rejections is the hardest local evidence there is. Which department proposed the same thing two years ago, which question it died on then, and where that money went instead, only someone who has been at this company can look up. Put it in the evidence column and it beats any industry news, because the proposer remembers that episode too. The precondition is that the sheet from back then can still be found, see the section "Elimination Needs an Exit, and Re-evaluation Conditions Need a Recheck Owner" below.
## At Anchor & Helm: Three In, One Out
Three candidates, one at a time through the sheet. The full scorecard is in [Template 7](../appendices/template-07-five-questions.md). Here are the decisive rounds.
**Candidate one, the service chatbot (the board's suggestion).**
Pain. Give it its due first, the service line's pain is real and call volume is climbing. Push one layer and it changes shape. Customer complaints cluster on "nobody tells me where it is stuck," and the bulk of the calls is a downstream symptom of the exceptions backlog, with inquiries only a small share. Fit a more patient answering machine to the symptom and the source bleeds exactly as fast. Pain 3, the pain is real, the leverage is upstream.
Data, 1. The chatbot would have to answer two kinds of question. For "where is my claim," the answer lives in the core system, the inbox, and Linda Marsh's tracker, and at minute 35 of the MVP you found that the status field lies (the first line in the Chapter 0 friction log). Using a lying field to answer customers in real time transfers the internal data debt straight to the customer. For wording questions, the library of standard approved answers it would need does not exist at all.
Decision and Risk, both at the floor. The entry point is real-time customer-facing replies, error tolerance is close to zero, and nobody backstops it, because bypassing people is the whole point of a service chatbot. What does the worst single output look like? You wrote it out. "Your loss assessment payment is expected within three business days." One wrong promise to a customer is one regulatory complaint. Decision 2, Risk 1.
ROI 3. Call cost can be computed, but the slice a chatbot would save cannot, because the bulk of the calls will fall on its own once the upstream is cured. The cause of death in one sentence. It would carry the conversation with the lowest error tolerance in the place where the data is worst.
**Candidate two, automated monthly reports (Kevin's proposal).**
This candidate dies faster than expected, and the murder weapon is SPIN's implication follow-up. You asked two more questions.
"Last month's report was three days late. What happened?" Kevin thought about it. "Nothing. I chased it once."
"After the report goes out, who replies or follows up?" He thought longer. "...Grant looks at two numbers. Everyone else files it."
The bleed rate is near zero. The pain is in the annoyance of producing it, not in any business consequence. Pain 1, and the conclusion is settled on that one question, with the other four filled in afterward for the file. That is how the elimination table saves time.
The monthly report's annoyance is not undeserving of help, but what it is worth is forty minutes of goodwill, not a three-month project. A few weeks later you automated, in passing, the reporting spreadsheet Kevin gets questioned about most, and Chapter 5 already showed you what those forty minutes bought. Eliminating a project is not the same as eliminating the other person's problem.
**Candidate three, the exceptions action queue.**
Pain 5. A specific person (Linda's team), a specific action (the next step on an exception claim), a measurable bleed rate, the first payment cycle running at 1.6 times the industry average. More precisely, about six tenths of the waiting time across the ten cases is "waiting on documents with nobody chasing, stuck with nobody claiming it." Chapter 4's -30% is computed from here.
Data 3, a conditional pass. The data exists, it has a business owner, ten de-identified samples came through. But the status field being untrustworthy is a confirmed fact, and only after the Excel is reconciled against the core system will you know which rung of the ladder this data actually stands on. Write the condition explicitly, the reconciliation gets finished during discovery (the stage of finding out how things actually are). Chapter 9 in its entirety is that condition being met.
Decision 4. AI enters at the advise layer, ranking, missing items, next action, with a Human Call column behind every suggestion, errors visible on the spot and overridable, the backstop named (Linda's team). The three unsafe marks remind you that tolerance has a boundary. Some error types (ranking a high-risk claim low) must not happen even once, and it takes an eval (the evaluation set) to fence them off (Chapter 11).
Risk 4. The worst single output is one wrong ranking suggestion being followed. The loss stays inside, and it can be found, corrected, and closed out. Not one word goes outside, and it sits outside the red line (no payout decisions).
ROI 4. First-touch handling time is the outcome metric and the conversion path is clear, reviewer hours, chase calls, complaint handling cost, customer churn. Kevin's operating budget owns that money, and Grant's board narrative ("catching up with the industry") can use it.
Three lines of summary.
| Candidate | Pain | Data | Decision | Risk | ROI | Lowest Score | Conclusion |
|------|------|------|----------|------|-----|-----------|------|
| Service chatbot | 3 | 1 | 2 | 1 | 3 | Data / Risk | Eliminated (re-evaluation condition, re-measure the call mix after exception handling is cured, and the standard-answer library gets built) |
| Automated monthly reports | 1 | 3 | 4 | 4 | 2 | Pain | Eliminated (turned into a favor done in passing, no project) |
| Exceptions action queue | 5 | 3 | 4 | 4 | 4 | Data (conditional) | Winner, conditional on data reconciliation during discovery |
## How to Tell Grant the Board's Suggestion Was Eliminated
The winner is easy to report. Eliminations are hard, especially when the eliminated one came from the board. Near the end of week 2 you book twenty minutes with Grant Whitmore. His first sentence is, "That service one. Did you look at it?"
You did not say "I do not think it fits," and you did not talk about technical difficulty. You put that one page of elimination table on the table and did one thing, point at the scores and read the evidence column. "The service line's pain is real, but the bulk of the complaints is 'where is my claim stuck,' and that is a symptom of the exceptions backlog. Cure exceptions and this share of the calls falls on its own." That is the Pain row. "It has to answer customers in real time about claim status, and the claim status field lies. The MVP hit that on day one." That is the Data row. "The worst single output is a wrong payout statement made to a customer. That is regulatory risk, and the chatbot has no human backstop." That is the Risk row.
Then the sentence you had ready. "So the conclusion is that the evidence says not yet, and my personal preference is not in it. The sheet carries the re-evaluation condition. After exception handling is cured, re-measure the call mix, and with the standard-answer library built, this candidate is worth running through the five questions again."
Grant studies the sheet for a while. He does not ask "why are we not doing it." He asks, "The board. How do I put it?" You give him a sentence he can repeat. "What the peer launched is a general-purpose service bot. We are treating the source of the calls first, exceptions. Once the metric moves, the foundation for service automation, the standard-answer library and trustworthy claim status, is laid along the way."
He puts the page in his bag. "I will take this. If the board asks, I will say it this way. On exceptions, let us set the metric at a meeting next week." That "meeting next week" is the charter meeting you saw in Chapter 4.
This scene is worth one line of retrospect. "I am not doing it" is a position, and positions get haggled over. "The evidence says not yet" is a judgment, and rebutting it takes better evidence, and nobody in the room has evidence harder than your ten cases. The re-evaluation condition then turns "no" into "not now." Elimination becomes a conclusion the future can overturn, not a taking of sides. Doing this inside your own company leaves less room to back out. The eliminated idea came from your own reporting line, the person who proposed it is still sitting across from you next week and still has to lend people to your project next quarter. So the internal exit cannot be the re-evaluation condition alone. The re-evaluation condition answers "why not this time." It does not answer "then what happens to my ask now," and it does not answer "who will still remember this sheet in six months." There are three exits.
- One, the gate. Before this report, it decides whether you have the standing to say what you just said.
- Two, the referral route. After the report, it decides where the eliminated ask went, and on whose books its ops cost lands later.
- Three, the recheck owner. Six months out, he decides whether this "not now" actually gets read again.
The next two sections cover the last two. Chapter 25 upgrades this whole move into organization-level intake discipline. Today's twenty minutes is its first rehearsal.
## Elimination Needs an Exit, and Re-evaluation Conditions Need a Recheck Owner
When a team from outside eliminates an idea, the idea disappears from view. Inside, it does not. Next quarter the proposer will have the ask you eliminated built by an outside firm, or he will buy a SaaS product, and when that thing starts erroring six months after launch, the tickets still route back to your team. What you eliminated was not the project, only your visibility into it. The ops liability arrives all the same.
So the internal version of elimination has one more move. Every "eliminated" row gets a referral route written behind it, one of three. Refer it to a standard tool, whatever the company's existing ticketing system, BI, or automation tools can do, and write down who configures it. Refer it to self-serve, a general capability the proposer can operate himself, and you give a template and one training session, not a project. Refer it to an outside purchase, and write down who selects, who accepts, and who owns ops after launch. Leave that last column blank and the name that fills in by default is yours. The monthly report row's "turned into a favor done in passing, no project" is the lightest of the three.
Re-evaluation conditions work the same way. Writing one down is not the same as being remembered. An outside team's elimination table gets archived with the project three months later. An internal elimination table has to live for years, and a lifespan with no end and a recheck count of zero is not "not now," it is a politely worded permanent refusal. So the re-evaluation condition gets two more columns, who rechecks it and at which standing meeting. The default is to hang the elimination table on your team's quarterly ask review, read the eliminated rows out each quarter, run the five questions again on the ones whose conditions have appeared, and formally close the ones whose conditions clearly never will, telling the proposer. This sheet is a team ledger, not a folder of yours. It has to still be there after you move on.
## Two Departments Both Passed the Five Questions. Which One First?
The five-question elimination table answers "is this worth doing." It does not answer "Claims and Underwriting both passed the sheet, which one first." A team taking projects from outside never meets the second question, serving one client at a time. The internal team is shared. In the same quarter, several business unit heads sit at the same level, any of them can go to Grant, and in their eyes "order" means "importance." So after the sheet there is a second sheet, three rows.
| Dimension | Question | How the Order Gets Set |
|---|---|---|
| Evidence maturity | Whose lowest score is a conditional 3, and can the condition be met this quarter? | Conditions that can be met this quarter go first, conditions six months out go behind |
| Commitment hardness | Whose business owner has already claimed people and weekly hours by name? | Claimed goes first, verbal support goes behind (the first of the three resource gates, Chapter 14) |
| Organizational timing | Whose outcome metric can make the next business review? | The one that can make it goes first, the one that cannot gets re-ranked next quarter |
All three tie, go back to Pain's bleed rate, and whoever bleeds fastest goes first.
One step after ranking. Put the opportunity cost on the same page. The report does not say "we picked the exceptions queue." It says "the other things this team is not doing because of it," listing the candidates behind it, which row each of them is stuck on, and which quarter each waits for. That column is written for the sponsor, and also for the department that ranked behind. In a shared resource pool, a trade-off that is not written out gets read as favoritism.
## Failure Modes
**1. Reasoning backward from technical capability to a scenario.** The project discussion starts from "our agent framework is mature," and the question is "where could the business side use it." The team's sunk investment in technology needs an outlet, and scoring quietly favors the scenarios that "can use our stuff." The order of the five questions, starting from Pain, is the antidote. For a candidate reasoned backward from technology, the Pain column usually holds nothing but "improve efficiency."
**2. Taking the wishlist.** The readout succeeds and you accept every "AI requirements list" the departments send, straight into the backlog (the list of what is queued to be scheduled), scheduled in the order received. The social cost of refusing is payable immediately, offending the proposer on the spot. The cost of accepting arrives three months later. But a wishlist is wishes sorted by rank, not opportunities sorted by evidence.
**3. Picking the one the executive is most excited about.** Grant's eyes light up at one idea and the project team automatically ranks it first. An executive's excitement is an approximate signal for resources and air cover (the senior cover that stands in front of you), so a team moving toward the light is being rational. But excitement measures how good an idea sounds when spoken. It is demo inflation's conference-room version, measuring narrative quality, not workflow leverage. On ROI settlement day, that applause does not discount to a cent.
## Next Monday
1. List every AI idea circulating in your organization, and first rewrite each one as a single "whose action" sentence (the workflow claim form from Chapter 0). Anything you cannot write is out in round one.
2. Pick three survivors and spend 30 minutes each scoring them on the [Template 7](../appendices/template-07-five-questions.md) scorecard. One rule only, the evidence column takes numbers, quotes, and samples, no adjectives.
3. Find the candidate you least dare to eliminate, usually the one proposed by the most senior person, and see which question its lowest score falls on. The evidence for that question is what you bring to your next report.
4. Use [Template 7](../appendices/template-07-five-questions.md)'s one-page format and deliver the result to the proposer face to face. Build the muscle memory for the sentence. "The evidence says not yet, and here is the re-evaluation condition."
**Want an agent to get you started?** In the repo you set up following [Start Here](../index.md), paste this to your coding agent:
```text
In the repo/ directory of the the-last-mile repository, help me do the Chapter 7 Next Monday actions. First run python3 templates/five-questions/onepager.py
with the built-in sample and show me a one-page elimination table. Then copy scorecard.csv into my own scorecard. I give the candidates
and the evidence for each question, you fill it in, the evidence column takes only numbers, quotes, and samples, and if I say an adjective, hand it back
for me to replace. I score, you do not. When filled in, rerun onepager.py pointed at my file, leave the one-pager for me to deliver in person, and add no conclusions of your own. If any command errors, stop and show me the output.
```
---
## Chapter Kit
- **Judgment frameworks.** The five-question framework (Pain / Data / Decision / Risk / ROI, each with its test question) + the elimination table (1 to 5, weakest-link logic, any question โค2 eliminated outright, no weighted average; a 3 is a conditional pass; every elimination must carry a re-evaluation condition); two scales with two jobs, pass / concern / unsafe / useless tests a single output, 1 to 5 screens opportunities; the three exits from elimination (the gate, the referral route, the recheck owner) + the three rows for cross-department ranking (evidence maturity / commitment hardness / organizational timing)
- **Templates.** [Template 7](../appendices/template-07-five-questions.md), the Five-Question Opportunity Rubric, the five-question scorecard (with 1/3/5 anchors and an evidence column) + the elimination table + the reporting one-page format
- **Key judgments**
- "The cost of producing an AI idea has fallen close to zero, the cost of validating one is still expensive, and the mouth of the funnel must have discipline."
- "Any question scoring โค2 is eliminated outright. A weighted average is the excuse you reach for when you do not want to eliminate anything."
- "Picking the use case is the deliverer's single highest-leverage decision, higher than any architecture decision."
- "'I am not doing it' is a position. 'The evidence says not yet' is a judgment."
---
# 8 ยท From Use Case to Boundary: Cut the First Thin Slice
!!! info "Companion Templates"
๐ [Chapter Template](../appendices/template-08-thin-slice.md) ยท ๐ [Template Library](../appendices/template-library-index.md)
> **The Challenge.** The use case is chosen, the charter is signed, and the next day the scope starts to swell. Each extension proposal is reasonable on its own. Accept all three and the project dies. And your "no," when you say it, sounds like laziness.
>
> **What You Will Be Able to Do.** Cut a use case into one thin slice with the Five Ones. Draw the decision rights boundary for an AI system. Use the scope decision log to turn every "no" into "not now," so the proposer leaves with an answer instead of a grudge.
---
> **Part III Navigation.** Chapters 8 to 13 all sit between the charter signature and the pilot launch, weeks 4 to 9.
> Three scope proposals within 24 hours of signing (week 4, this chapter) โ three-way reconciliation (week 6, Chapter 9) โ the multi-agent proposal at Monday's standup (week 7 Monday, Chapter 10) โ the annotation session and item 37 (weeks 7 to 8, Chapter 11) โ the security review meeting (week 8, Chapter 12) โ that one page and Monday's 7 minutes (week 8 Friday to week 9 Monday, Chapter 13)
## The Twenty-Four Hours After the Signatures
The exceptions queue won the five-question elimination (Chapter 7), and the charter was signed in week 4, after the readout (Chapter 4). The signing itself is low-key. Grant Whitmore, Kevin Doyle, you, Owen Hartley, one page, four signatures, ten minutes start to finish.
The creep starts at minute eleven.
Putting his pen away, Grant asks, "Now that it is signed, one more question. Once this thing runs smoothly, can it go toward full automation? From claim report to payout, fully automatic. The board will ask me that next year, guaranteed."
That evening Kevin calls. "The home property team lead came to see me this afternoon. They are all exceptions, and you are building the queue anyway. Put the home property ones in too, right? Just one more class of data, isn't it?"
At the next morning's standup, the two claims-ops IT engineers bring a one-page architecture sketch. "Rather than hard-coding the exceptions logic, build the bottom layer as a general-purpose workflow platform. Fields, rules, lines of business, all configurable. Customer service and underwriting can use it later. One investment, the whole company benefits."
Twenty-four hours, three proposals. Notice that none of them is wrong. Home property really is backed up, full automation really is where the industry ends up, a platform really does avoid building the same thing twice. Nor is anyone proposing them an enemy. This is exactly the evidence that the project is finally being taken seriously. But accept all three and the user groups, the decision layers (which layers, the ladder diagram below will show), and the number of scenarios all grow at once, and the "-30%" North Star loses any attribution. Three months later you will own something that touches a bit of everything and has launched nothing.
You do hold a charter, and its "Scope and Red Lines" section is there in black and white. It cannot stop these proposals. A charter is a snapshot of the consensus on signing day. Creep is new traffic that grows every day. A thicker document does not help. You need a ruling procedure that runs continuously. This chapter gives you that procedure.
## Why This Is Hard: Creep Is the Default Outcome, Not a Management Failure
Blame scope creep on "the business side lacks discipline" or "I failed to hold the line" and you reach for the wrong medicine. Creep has three structural engines, none of them about character.
**First, your project is the only car moving.** Every stakeholder (a party with a stake in the outcome) has a few wishes nobody has picked up, and a wish can only ride on a car that is moving. The moment the queue is approved, it becomes the mounting point for every AI wish in the company. Nobody is targeting you. This is gravity.
**Second, the marginal cost of every proposal looks small.** "One more class of claim," "just leave a hook for it." Seen one at a time, the increments are linear. The validation cost multiplies. One more user group means another set of scoring and annotation. One more decision layer means another round of error classification and oversight design. Nobody is lying about this cost gap. It just does not show.
**Third, the cost of refusal is asymmetric.** Accept an extension and the project pays, and the bill shows up three months later. Refuse one and you personally pay on the spot, especially when the proposal comes from Grant. Without a defensible narrowing logic (a reason for refusing that can state its cost), your "no" sounds no different from passing the buck, so the rational choice is to accept.
One more fact, which most people never see through, is the key to this chapter's solution. Many extension proposals want an answer the proposer can repeat to someone else. A "yes" is secondary. Kevin has no plan to launch home property tomorrow. What he needs is a sentence he can give the home property team lead. See that, and "refusing" turns from confrontation into service.
## Prior Art: Put a Price on Every "Want," and the New Line AI Added
The methods for narrowing are not new. Two traditions each contribute half.
**The first comes from the solution architect's tradition of architecture review, making trade-offs explicit.** When an extension is proposed, do not rush to yes or no. Price it first. "We can. The cost is these three things. You decide whether it is worth it." Pricing lets the proposer see the scales for himself, and you go from "the person in the way" to "the person doing the arithmetic." This discipline was carried over from architecture review into scope negotiation. Quality attributes conflict with each other. Throughput costs latency, flexibility costs maintainability. So the product of a review is the cost written next to every "want." Whether the design is good is not really the point.
**The second comes from the Extreme Programming tradition, YAGNI.** "Leave a hook for it" sounds prudent. In practice it bets a certain cost today on a future that mostly never comes. "You Aren't Gonna Need It," the discipline Kent Beck laid down and Ron Jeffries kept explaining. Until the need actually appears, do not build for an imagined one. This has nothing to do with laziness. Imagined needs are mostly guessed wrong, and the complexity you pay for them is due now and due every day.
In the age of AI agents, one of these old disciplines got more expensive and the other is entirely new.
YAGNI is the one that got more expensive. An LLM is general-purpose by nature. The cost of demonstrating generality has collapsed, and the cost of producing it (each scenario's own data reconciliation, eval, and trust building) has not dropped a cent. Chapter 1 covered demo inflation. Demos got so cheap that being impressive no longer proves anything. This is its scope version. The temptation to abstract got cheap. The cost of abstracting did not.
### The Decision Rights Boundary
What is new is a boundary. A traditional system boundary answers one question, which functions the system does. An AI system must answer one more. Which layer of the decision chain does the AI's judgment reach, and at which layer do humans take over? This book calls it the decision rights boundary.
> **The decision rights boundary is, on a workflow's decision chain, the layer where AI output stops, the layer where humans take over, and the evidence required to move up each layer.**
Why does drawing this line wrong mean redoing everything? Because it is the foundation of the whole validation system, not one item on a feature list. How errors are classified (Chapter 11, eval), who watches and whether they can keep up (Chapter 12, human oversight), where the chain of responsibility ends when something goes wrong, all of it grows out of this line. Draw a feature boundary wrong and you cut one module. Draw this line wrong and you redraw the whole set, and the code is the least of it. So it has to be drawn when scope is defined. It cannot be left for the architecture stage to "settle along the way."
## Thin Slice, the Definition and the Five Ones
Now the book's formal definition.
> **A thin slice is the smallest working slice that cuts vertically through all five layers, "real user, real decision, real data, real risk control, measurable outcome." Thin is the width. It serves one decision for one user group. Not thin is the depth. Every layer must go all the way to "real."**
The key word is vertical. Creep proposals are all horizontal. One more class of claim widens the user surface, a platform abstracts the bottom layer into a general component. A thin slice goes the other way. Cut the width to the minimum in exchange for one complete deep channel from data source to user action to outcome measurement. The Chapter 0 readout narrowing to "auto exception claims" was the first cut. This chapter finishes it.
In practice a thin slice is five "ones," and none of them may be plural.
| The Five Ones | How Anchor & Helm filled it in | Why it must be "one" |
|--------|-----------|------------------|
| **One user group** | Linda Marsh's auto claims review team (eight people) | Two user groups = two sets of scoring, two sets of tacit knowledge (the knowledge people hold but cannot state), and no way to say whose feedback counts |
| **One decision** | The next step on each exception claim (whom to chase, what is missing, who takes it) | More than one decision and errors cannot be defined, so the eval (the evaluation set) has nowhere to start (Chapter 11) |
| **One data path** | The merged view of the core system + Excel after reconciliation | Every extra path doubles the fitness check (does the data qualify to drive actions, Chapter 9) and the reconciliation |
| **One risk boundary** | AI stops at the advise layer; payout decisions are never touched (red line) | More than one boundary and oversight and the chain of responsibility have seams, and incidents find seams |
| **One measurable outcome** | First-touch handling time for auto exceptions, -30% (the charter North Star) | More than one metric and improvement cannot be attributed, and acceptance falls back to "leadership thinks so" |
Three notes. One, the fifth "one" is cited straight from the charter North Star (Chapter 4). The thin slice is the North Star's engineering projection, not another negotiation. Two, the data path is still a paper promise at this point. "The merged view after reconciliation" is not delivered until the data fitness check in Chapter 9. For now you only need the path to be one, not three. Three, the five "ones" only qualify when they lock together into one sentence. This user group, making this decision, along this data path, inside this risk boundary, improves this metric. Wherever it does not read smoothly is where the cut is not clean.
Why the obsession with "one"? What a thin slice is after is being able to afford validation. Doing less is a by-product. The whole point of a pilot is clean evidence (the escalation decision in Chapter 14 depends on it), and clean evidence only comes out of "one."
## The Scope Decision Log, Turning "No" into "Not Now"
The Five Ones handle "where to cut." The other half of the problem is "what happens to the part cut off." The answer is this chapter's second piece of kit.
> **The scope decision log is one line per refused extension proposal, with what it was, who proposed it, the reason for refusal (the cost), and the revival condition.**
The revival condition is the soul of it. It is the lubricant of refusal, and it translates "no" into "not now." Home property was not rejected. It is queued as "the first extension after the pilot hits target." Writing revival conditions has one hard rule, events, not dates. "Consider in Q3" is a brush-off. Q3 arrives and nothing happens. "After the pilot North Star hits target" is a commitment. When the event happens, the proposer comes to open this page himself.
This rule has an expiry date, and most people do not think of it. Pilot events exist only during the pilot period. Once the system enters the operating period, events like "hits target" are gone, and the revival list keeps growing. The operating period often brings more proposals than the pilot did. Operating-period revival conditions switch to two kinds of anchor. One hangs on the periodic resource review, written as "ranked with the other candidates at next quarter's resource review." The other hangs on an operating metric, written as "revisit after such-and-such operating metric holds within threshold for N consecutive weeks." Both are fixed moments where other people will also be in the room, not dates you keep in your head. An entry that cannot be given either anchor is really a new project already. It should go through project approval, not squat on this sheet.
This sheet does three jobs at once. For the proposer, it is an answer they can repeat. For you, it makes refusals cumulative. The second time someone proposes a platform, open the log, no need to debate again. For the project, the revival list is an asset. On the day the pilot hits target, it is the ready-made phase-two roadmap (Chapter 23 comes back to harvest it).
The charter's existing "not this phase" list (Chapter 4, the part of the "Scope and Red Lines" section that explicitly says what is not being done) is its first batch of entries. At signing there were only conclusions. Add the cost and the revival condition now, and "not this phase" turns from a wall into a timetable.
This sheet also has a leak that only exists inside a company. It captures formal proposals. It does not capture "do me a favor." Another department wants you to support their own pilot on the side. Half a day's work. It goes through no project approval, no charter, and not this sheet either, because it is not called an extension. Help three or four times and one of your engineers is down nearly half a week, nothing shows on the weekly project report, and when you slip you cannot even produce an explanation. One rule. Any favor over half a day goes into the scope decision log. The proposer's name goes in, and the refusal reason column states which item of this phase it crowds out. Writing it down is not for refusing. Most favors should still be done. It is so this time has somewhere to show.
## At Anchor & Helm: Three Proposals, Three Responses
Three proposals, three tools, one at a time.
**Home property. Price it, then give it a place.**
Pricing means putting the cost of a proposal on the table and stating it plainly. Kevin's "just one more class of data" is an honest judgment. Seen from the queue interface, home property claims and auto claims do look about the same. You do not argue on the phone. The next day you go to see him with a three-line tally.
"Kevin, adding home property looks on paper like one more class of data. In practice three things double. First, Linda's team does not touch home property. The people using this thing every day would have to become the home property team, and the scoring, the annotation, those two hours a week, all of it gets rebuilt for the home property team. Second, home property's 'abnormal amount' is a different feel. Auto looks at the repair shop and the survey photos. Home property looks at the loss assessment report from an independent adjusting firm. Not one of Linda's five rules of thumb carries over, and the set of standard cases we test the system against starts from zero. Third, where that -30% comes from. You worked it through at the charter meeting (Chapter 4), and that was the auto tally. Mix home property in and the attribution gets muddy. Add the three up, the pilot slips at least six weeks, and the evidence gets cloudy."
Kevin frowns. "So when does the home property team get its turn? You cannot expect me to go back and say 'not doing it.'"
You open the scope decision log and write a line in front of him, reading it aloud as you go. "Home property exceptions, proposed by Kevin Doyle. Reason for refusal, user group, annotation system, and metric baseline all double, diluting the pilot evidence. Revival condition, the first extension after the pilot North Star hits target." Then you add, "'First' is exclusive. When the time comes, data reconciliation starts a month early, and the home property team does not queue for project approval again."
Kevin stares at the line for a few seconds. "Fine. That I can take back. First in line is not the same as killed." From disappointment to acceptance, and not by persuasion. His wish got a definite place instead of being thrown into a black hole.
**Full automation. Draw the ladder.**
With Grant, you draw, you do not debate. Four layers on the whiteboard, the red line at the top, what AI already does at the bottom, read from the bottom up.
```
Decide pay or not, how much <- red line: never touched
Act send chase notices, assign, change status <- humans take over here (the reviewer clicks confirm)
Advise priority + next step + reason <- AI stops here (this phase)
Sense gather documents, dwell time, risk signals <- AI does this
```
"Full automation is the question of AI climbing this ladder one layer at a time. Each layer up needs different evidence, and capability is actually secondary. For AI to go from advising to acting, the advise layer first has to build a track record, which suggestions get accepted as is, which get overridden often, and that record starts accumulating during the pilot (Chapter 17). With it, what you discuss with the board next year is 'which classes of action have a high enough acceptance rate, so which ones do we automate first.' Let the AI's suggestions earn a record of being accepted before you talk about letting it act."
Grant looks at the drawing. "This ladder goes into the next memo you write me. That is how I will put it to the board too." The scope decision log gets its second line. Revival condition, once the advise layer's acceptance data hits target, move up one layer at a time, each layer decided on its own.
The "evidence" in that line takes two forms. One is the human acceptance record, which suggestions get accepted as is and which get overridden often. The other is a machine verification loop. Can this layer's output be verified on the spot, with errors automatically detectable, enumerable, and reversible? Where verification is automatic, an error is caught at once, and you need not wait for people to accumulate usage records before you can relax. Where the output is verifiable, moving up need not rely entirely on an acceptance rate slowly building. Where the output cannot be verified, however high the acceptance rate, go slow (Chapter 10 unfolds "verifiable output" into a pattern selection criterion, and Chapter 22 uses this distinction to deliver the first move up).
**The general-purpose platform. Five counter-questions.**
The engineers' architecture sketch has no technical errors, which is exactly what makes it dangerous. You do not evaluate the design. You ask one question. "Remember the five gaps (Chapter 1)? Let's walk through them. Which gap does the platform narrow?"
First lay out the list, data gap, workflow gap, trust gap, ownership gap, value gap, and go through them one by one.
Data gap. The platform does no reconciliation for any scenario, and the core system's status fields lie exactly as before. Not narrower. Workflow gap. Configurable means embedded in no specific workflow, and "general-purpose" is precisely the antonym of closing the workflow gap. Wider. Trust gap. Linda is willing to score because this system understands exceptions, not because it is configurable. Not narrower. Ownership gap. The queue's owner is Kevin. Who owns a company-wide platform? No answer. Wider. Value gap. The North Star is first-touch handling time. "Configurability" does not convert into that number. Not narrower.
Five gaps, zero narrower, two wider. The older engineer closes the sketch himself. "Got it. The platform is what happens after the second use case shows up." You add a sentence, catching the enthusiasm rather than dousing it. "Right. A good abstraction grows out of the second case. It is not guessed from the first. Whoever writes the extraction pipeline (the chain of steps that pulls data into the queue) keeps the interfaces clean. That is the first brick of the future platform (Chapter 23 comes back to distill it)." Scope decision log, third line. Revival condition, after the second slice lands, distill what is common and revisit.
A week later, the three-line sheet is pasted at the end of the weekly project report. Not one proposer feels refused.
Nobody turning hostile is the luck of this case, not a guarantee of the procedure. Suppose Kevin reads the three lines and still insists, "the home property team lead cannot wait, put it in first." The procedure runs as before, one more round. You do not recompute the cost. The tally is done, and computing it again is fighting to win. You write both roads on the same page. One as before, home property waits, revival condition "the first extension after the pilot North Star hits target," exclusive, reconciliation starts early. The other is what he wants, home property goes in this phase, with the same tally written next to it, the pilot slips six weeks, and the North Star's attribution gets muddy.
A team from outside would, at this point, ask the owner to sign next to the cost, and with the signature the responsibility changes shoulders. You do not have that step, and you should not pretend to. Kevin is the pilot owner and the choice is his, yet three months later, when the pilot produces no conclusion, your team is the one asked, not him. So the internal version replaces the signature with a decision trail, two moves. First, both roads and their costs go into that week's project report, cc Grant, quoted verbatim, no comment, no recommendation. Second, if the second road is taken, the delay attribution column on the project board for those six weeks reads "scope change, home property added this phase." The attribution travels with the schedule, and whoever looks at progress sees it.
Be honest about what this trail buys. It does not clear you of responsibility. It only clears you of "surprise." Three months later, when someone asks why the pilot ran six weeks late, the report and the board answer for you. You do not have to reconstruct in that meeting who said what when, and you do not have to make Kevin admit it from memory. More important, it works in the moment. Kevin knows both things will happen, and that changes the choice itself. The trail's main function is not settling accounts afterward. It is making the cost visible at the moment of choosing.
Most of the time nobody chooses the second road. What he needs to take back is a sentence he can give. "First in line" is a promise. "Put it in first" means going back to explain why the whole pilot is six weeks late.
**The fourth kind of proposal, a lateral request from a peer department.** The three proposals above came from the sponsor, the business owner, and the engineers you work with. That is one vertical line, all inside this project's chain of authority. Sooner or later an internal reader gets the fourth kind, one sentence from a peer department. "You did it for claims, do one for us while you are at it, we are all one company."
This kind is the hardest to hold off, because the three things that normally hold proposals off for you are all absent. Grant's authorization covers only this project. Invoking his name may not work, and it looks like borrowing a banner. There is no written agreement between you and the other side to cite. He is not on the project approval form, and he is not on the charter. And next year he may be the person you have to go to for data, for people, for a cross-department interface. The cost of refusing is booked to your personal long-term account and settles a year later.
The response does not change. Still pricing plus a place, with one added move. Price it with the same three-line tally. One more department, and the user group, the annotation system, and the metric baseline each get rebuilt, and the pilot slips by this much. Work it out face to face. Giving it a place means logging it in the scope decision log, his name in the proposer column, the revival condition as a checkable event. The added move is the cc. Send that line, cost and all, to Grant. The cc is not tattling. It sends a lateral request into the vertical ranking, so the person with the authority to rank does the ranking. You have no ruling power, but you have the power to put things into the ruling procedure, and for this kind of proposal that is enough. The other side mostly does not expect a yes on the spot anyway. What he needs is the same, a sentence he can take back.
Behind these three responses are a few levers a team from outside does not have. First, you can price a proposal before it takes shape. Hear in the corridor that the home property team is stirring, and that same day you can go and walk through the three-line tally. By the time they formally ask, the tally is done, and the conversation is no longer whether to add but when. A team from outside has to wait until the proposal is formally made before it has a position to speak from. Second, the scope decision log is a cross-project asset. It does not reset with the project. Next year, when another department proposes the same thing, you open last year's line, cost and revival condition both there, and the cost of debating it again drops to turning a page. That is also why this sheet lives in the team's shared location, not in your local files. The third lever is that red lines can borrow the company's existing institutions, which is the next section.
## Failure Modes
**1. The platform temptation, abstracting before the use case has run.** The first scenario is still in pilot and the architecture already shows a "rules engine" and a "scenario configuration center." Abstraction is an engineer's identity currency. "Platform" is worth more on a resume than "spreadsheet." And the generality of LLMs makes a platform demo cheaper than ever, so for the first time temptation and capability are both in ample supply. But the legitimate raw material for abstraction is repetition. Before the first use case is validated, the only raw material for abstraction is imagination. One rule. Every platform proposal goes into the scope decision log, and the revival condition is always "after the second slice lands."
**2. Red lines that live only in speech.** "We agreed not to touch payout decisions." The person who said it transfers six months later, and the new product manager proposes at the requirements meeting, with enthusiasm, "give the payout amount an AI estimate." A verbal consensus lives in the memory of those who were present, and evaporates when people change. A red line that is never touched day to day and never brought up is often remembered for the first time the moment it is stepped on. Writing it into the charter is only the minimum. Nobody opens the charter day to day.
The real defense is putting the red line where it might get stepped on. The scope decision log (opened at every scope discussion), the security review packet (the packet of review materials, Chapter 12), all the way to the system interface itself. The queue has no column called "suggested payout amount," and that "no column" is the red line in physical form.
There is one more kind of hardness that only an internal reader can get. The company already has red lines for information security and compliance. They have their own review meetings, their own inspectors, their own penalties, and whoever steps on one writes a report. The red lines you wrote yourself into the charter have none of those three, and every change of people means explaining them again. So hang whatever you can on the existing ones. "No automated outbound messages" goes under the company's compliance policy on external contact. "No payout decisions" goes into the security review's conclusions. A borrowed red line does not need you in the room to remind people. The institution remembers for you. This works on you too. The person who transfers in six months may be you.
**3. Over-narrowing, the thin slice cut into a no slice.** Scope is cut down to "just missing-document detection, the queue can come later." Smallest in engineering, lowest in risk, and Linda reads it and says, "What documents are missing I can tell at a glance. What I lack is the ranking and the chasing." Once the discipline of narrowing is learned, it decays from judgment into reflex. Refusing takes less effort than thinking, and "one more cut" always looks prudent. But the criterion for the cut is wrong. The "minimum" in thin slice is the minimum that is complete in value, not the minimum that is easy to engineer. After the cut there must still be one real user who can say "this solves my problem." Take the slice description to your "one user group." If you do not hear that sentence, what you hold is a no slice.
## Next Monday
1. Write the Five Ones for the project in your hands. The items you cannot write, and the items you are forced to write in the plural ("both kinds of users need it"), are your scope risk list. Every plural owes you one narrowing negotiation.
2. Add the extension proposals you turned down verbally in the past month to the scope decision log ([Template 8](../appendices/template-08-thin-slice.md)). Give each a cost and a checkable revival condition, send it back to the proposer, and watch the reaction.
3. Draw the decision rights boundary. Which layer AI stops at, which layer humans take over, what evidence each move up requires. Then count how many documents this line appears in. A boundary that exists only in your head does not exist.
4. Check where your red lines "live" right now. A red line that lives only in the kickoff document is a verbal red line. Give it a home in the scope log, the review materials, and the product interface, one each.
**Want an agent to get you started?** In the repo you set up following [Start Here](../index.md), paste this to your coding agent:
```text
In the repo/ directory of the the-last-mile repository, help me with the Chapter 8 Next Monday actions. Copy
templates/thin-slice/thin-slice.md into the working directory I name, and ask me the Five Ones one by one. Record any
item I answer in the plural as is and mark it. That is my scope risk list, do not narrow it for me.
Then run python3 templates/thin-slice/scope_log.py --help to see how the scope decision log works, demonstrate it once
with the built-in sample, then enter, one by one, the extension proposals I turned down verbally in the past month. I
supply the cost and revival condition for each. The decision rights boundary diagram is mine to draw.
If any command errors, stop and show me the output.
```
---
## Chapter Kit
- **Judgment frameworks.** The Five Ones of the thin slice (one user group / one decision / one data path / one risk boundary / one measurable outcome); the four-layer decision rights boundary map (sense, advise, act, decide, moving up on evidence); the five-gap counter-question ("Which gap does the platform narrow?")
- **Templates.** [Template 8](../appendices/template-08-thin-slice.md), the Thin Slice Definition Sheet and Scope Decision Log, Five Ones definition sheet, decision rights boundary map, scope decision log (with how to write revival conditions), no slice self-check
- **Key judgments**
- "Scope creep is the default outcome of organizational dynamics, not a management failure."
- "Turn 'no' into 'not now.' Revival conditions are the lubricant of refusal."
- "An AI system has one more boundary, the decision rights boundary. Draw that line wrong and you redo the whole validation system, not just the code."
- "Let the AI's suggestions earn a record of being accepted before you talk about letting it act."
---
# 9 ยท Data Reality: Score Every Source on the Data Fitness Ladder
!!! info "Companion Templates"
๐ [Chapter Template](../appendices/template-09-data-fitness.md) ยท ๐ [Template Library](../appendices/template-library-index.md)
> **The Challenge.** The business side thumps its chest and says "the data is all in our system." Why is none of it usable the moment you touch it?
>
> **What You Will Be Able to Do.** Use the data fitness ladder's seven-rung verdict to measure how far a data source is from supporting action. Use three-way reconciliation to measure a field's real inconsistency rate and its inconsistency patterns. Assign a source of truth to each class of data on that evidence, and split the repair work into an engineering half and an operating half.
---
## Week 6, Tuesday Morning, Two Tables Side by Side
The read-only de-identified view opened Monday morning. Ten days of standoff, broken in the end by one sentence from Kevin Doyle in a hallway (Chapter 5). The first thing you do Tuesday morning has nothing to do with pipelines or models. You put two tables side by side on the screen. On the left, 200 auto exception claims the core system currently marks "in progress." On the right, Linda's tracker, the Excel sheet her team keeps, the source of truth you found while shadowing in Chapter 6, the one that actually reflects where a claim stands. Row by row.
By noon the result is in.
- **61 claims**, the core system says "in progress," the Excel says "waiting on the customer's documents," some of them waiting three weeks already, with nobody chasing.
- **17 claims**, effectively closed, the payout already landed, and nobody changed the status back.
- **9 claims**, carrying two mutually exclusive statuses in the core system and in the call center ticketing system, "in progress" on one side, "closed" on the other.
On the day of the Field MVP in Chapter 0, 3 of the 10 cases had a status field that did not match the description, and you put it in the friction log, taking it at the time for the bad luck of a small sample. With 200 in front of you, that is the population. 87 of the 200 cannot be taken at face value to drive an action, 44%. The largest and most insidious class among those 87 is the lagging pattern, 61 claims, about three in ten of the full 200. It is not wrong about whether a claim is finished. It buries the real reason a claim is stalled under the words "in progress," and what it misleads is exactly "what to do next." The same species of lie as those 3 of 10 on the MVP day.
At project approval Kevin had said, "The exceptions data is all in the system." He was not lying to you. At the level he can see, the sentence is entirely true. The fields are there, the reports run. The trouble is that a ladder stands between "we have it" and "it is usable," and the organization can see only the first rung. This chapter is about that ladder, and about using one reconciliation to pull a project back off a crack in it.
## Why This Is Hard: Fields Are a Process Trail, Not Facts
An engineer's first reaction is "the data quality is bad," as if this were an oversight at Anchor & Helm. It would be the same at any company. It is structural. The fields in the core system exist to leave a trail on the process, not to drive decisions. A reviewer changes a status because the button has to be clicked before the process can move to the next step. "Record reality" was never on her list of duties. So the update discipline on a field settles at exactly one line. Whatever audit can check is always accurate. Whatever audit does not check depends on the mood.
That line explains a telling detail in the reconciliation result. The status fields are a mess and the amount fields are correct almost claim for claim, because amounts run through finance reconciliation and audit watches that field. Same people, same system, and the discipline is worlds apart. A field's update discipline runs only as far as "audit finds nothing," and not an inch further.
This caused no disaster in the past, because every consumer of the fields was a person, and people come with error correction built in. Linda sees "in progress" and checks the email, looks at the Excel, makes a call. Every veteran assumes the fields cannot be fully trusted, and that assumption was never written in any document. In this company's history, AI is the first consumer that takes the fields at their word. It has none of that tacit correction. Whatever the field says, it believes, so it is the first to fall into the crack, marking a claim that should be chasing the customer as "internal rush," queuing a closed claim into the to-do list. The data gap among Chapter 1's five gaps, made concrete in the field, is the crack in this ladder.
## Prior Art, and What AI Changed
The ladder is not a new discovery. The data warehousing tradition set the rule thirty years ago. Data entering the warehouse has to answer "which source wins" first. That is source-of-truth discipline. One business fact recognizes one authoritative source, and everything else is a copy.
Kleppmann's *Designing Data-Intensive Applications* pushes the thinking up to the system level (paraphrased here). The system of record (the authoritative register for one business fact) has to be kept apart from derived data. Derived data (caches, indexes, rollups, model outputs) can be recomputed at any time but must never feed back into the source, or errors amplify in a loop. Data earns trust through lineage (where this value came from and who transformed it), not through a field name that looks true.
AI changed two things about this ladder, one up and one down.
**Up, the ladder gained a new way to climb.** "Waiting on the customer's documents," the real status, is nowhere in the core system, but it is in the email traffic between the surveyor and the reviewer. In the past that unstructured data was stuck on the bottom rung, because nobody could afford to structure it by hand. LLMs made extraction usable, and email, notes and call records can for the first time be lifted into an interpretable, even same-day, data source. Anchor & Helm's merged view later ate the mailbox signal for exactly this reason. But what comes out of extraction is still derived data, with an error rate of its own. This ladder gets laid out below, and its most expensive rung is called validated. That source has to climb all the way to validated too. Extractable is not trustworthy.
Down, AI itself became a new source of derived data. The queue generates dozens of "suggested next steps" a day, and three months on somebody will certainly cite them as "the record." If those outputs are written back into business fields, the next version of the model learns its own old output as fact. "Derived data must never feed back into the source" grew new teeth in the AI era. "Do fields written back by AI count as truth" is a question that did not exist before, and the first shape of the answer is that AI output may exist only as a decision trail, the suggestion, the reason, the Human Call, the timestamp, physically separated from the source fields (Chapter 17 works out the decision trail).
## Framework One: The Seven Rungs of the Data Fitness Ladder
To judge whether a data source can support one specific action, climb this ladder rung by rung. One test question per rung, stopping at the first rung where you cannot produce evidence. That is where the source really stands.

| Rung | Test Question | At Anchor & Helm |
|------|----------|----------|
| **exists** | Has this information been recorded anywhere? In which system, in which field? | Kevin's "it is all there" speaks only to this rung |
| **accessible** | Can you and the future system read it in a compliant, repeatable way? | About two weeks after the charter, opened through the trust network, with the approval flow no help (Chapter 5) |
| **interpretable** | Does the same value mean one thing across departments and across time? | "In progress" means different things in Claims and in customer service |
| **timely** | Does the update frequency keep up with the action you want to drive? | Status lags by the week, and the queue has to drive same-day action |
| **traceable** | Do you know where this value came from and who changed it? Are two systems talking about the same entity? | Claim report numbers are keyed in by hand and mistyped, so claims and tickets will not join |
| **validated** | Has it been reconciled against reality? What is the inconsistency rate, and in what pattern? | 87 of 200 cannot be used at face value, the largest class being the lagging pattern at about three in ten |
| **actionable** | Holding it, can the "next action" on that row of the queue be carried out by a specific person, and can an error be caught? | Status plus missing documents plus waiting time, all present, is the entry ticket to the queue |
Run the claim status field through it. Before the reconciliation it had cleared exists and accessible only, and it stuck at the interpretable question. Whether "in progress" means the same thing in two departments, you had no evidence. After the reconciliation, the core system's status field has an answer at the validated rung, 87 of 200, fail. The merged view settled on Friday recognizes Excel as its one source. The Excel is updated the same day, lineage is marked on the cell, and interpretable, timely and traceable all clear. Validated waits for the retest during the pilot, and until then it stops at traceable. In pilot week 4, project week 13, what gets retested is the core system's status field after the daily write-back discipline, another 200 claims sampled and the handlers asked, and only 9 do not match, all of them simply not updated yet that day. Validated clears. Half the credit belongs to engineering, half to the status write-back discipline Kevin claimed at gate three in Chapter 14, where the review team clears status before leaving each day. The actionable rung waits for the acceptance record after the queue goes live (Chapter 17).
Three rules of use. **First, judge rung by rung, no skipping.** The organization's verbal assurance covers only exists, and a successful demo proves only accessible, exactly the two cheapest rungs. **Second, fitness is relative to an action.** The same table may top out supporting a monthly report and reach only the second rung supporting same-day chasing. Write the action down before you score. Do not ask "is this data good." Ask "does it deserve this action." **Third, fitness is a state, not a property.** Validated today does not mean validated in three months, and the failure modes settle that account.
The internal reader has one more hurdle to climb, "it is all in the data lake already." The company spent years building a data lake or a data platform, every system's tables were synced into it, and from IT to the business owner to you, everyone assumes the six rungs above exists were solved along with it. Syncing solved the hauling only. Permission approvals, definition dictionaries, update latency, cross-system keys, not one of them was solved. The data platform saying "fully synced" is not evidence. The evidence is you running the target action against that table and being able to report its inconsistency rate. Kevin's "the exceptions data is all in the system" is, on an internal project, often said by you or by your manager at the project approval meeting. What you lack, compared with an outsider, is exactly one pair of eyes that does not believe that sentence.
## Framework Two: Three-Way Reconciliation
The most expensive rung on the ladder is validated, because it has no shortcut. The only way is to reconcile the data against reality. The method itself is plain. The discipline is in the details.
1. **Define the standard first, then sample.** Write down in black and white what counts as inconsistent, then look at the data. Reverse the order and you will quietly tune the standard until the result looks good.
2. **Sample N records.** 50 to 200, covering typical cases, edge cases and aged claims (the longest-waiting claims are where field discipline is worst).
3. **Check three ways.** Every record against three sources. The system field (the official version), the private source of truth (the Excel, the email threads, the ones the archaeology of Chapter 6 dug out), and asking the handler (on disputed records, ask the person handling it directly, "where does this claim actually stand right now"). Two ways can only find "they differ." Only the third can rule on "who is right."
4. **Produce something beyond two numbers.** The inconsistency rate (the total rate together with the largest single class, both reported) is for the decision. The inconsistency pattern (lagging? forgotten? semantic divergence? broken join?) is for the fix, and the prescriptions for different patterns have nothing in common.
5. **Turn the result into decisions.** Assign one source of truth per class of business fact. Write down which half of the repair is engineering and which half is operating. Write down the retest date.
You had a feeling already. In Chapter 5, in the forty minutes spent automating Kevin's monthly spreadsheet, the two definition errors you fixed along the way were the interpretable rung showing itself for the first time. That spreadsheet is the first-hand material for this reconciliation, and you knew which fields to doubt before IT did.
The advantage compounds. You spend years fixing definitions and rebuilding reports for one team after another, and every fix teaches you one more field's temperament, while an outsider starts from zero at each new company. The same depth gives you two other things. Change tickets and release records from past system upgrades are yours to pull, so tracing lineage at the traceable rung is an order of magnitude easier. The handler of a mutually exclusive claim sits two floors away, so the third way, asking the handler, costs you nothing, and that step is the only one in a reconciliation that can rule on who is right.
## At Anchor & Helm: Four Ways to Lie, One Merged View
Wednesday afternoon, you take the 9 mutually exclusive claims and a sample of the disputed ones to the third way and ask the handlers. The first three inconsistency patterns surface.
**The lagging pattern (61 claims).** The status is true, only slow. After a reviewer sends the request for documents, the Excel is updated the same day, and the core system does not move until the claim is routed again, sometimes three weeks later. The prescription is a faster source (the Excel, the mailbox). Engineering can solve it.
**The forgotten pattern (17 claims).** The claim is effectively closed. Changing the status is an action no process forces and no person is affected by, until your queue becomes the first victim. The prescription is operating discipline. Engineering cannot solve it.
**Semantic divergence (9 claims).** "In progress" in Claims means "on my desk." In customer service's ticketing system it means "answered the customer, waiting for a reply." Two departments used one word to record different facts, and for ten years it hurt nobody, because no system had ever consumed both sides at once. The prescription is to align the semantics first and talk about syncing second. This is an organizational problem, not an ETL (extract, transform and load pipeline) problem.
But who convenes that semantic alignment meeting. That question is the genuinely hard one inside a company. Often you are the only one who will push it, and you have no authority to convene across departments. Convening power has only three sources. One, the sponsor authorizes it and hangs the meeting on his own weekly agenda, and one sentence, "the definitions behind the auto queue need both departments to confirm," is enough. That is the fastest route. Two, the company's existing data governance committee or master data management team, whose job semantic alignment already is, and you are only the proposer. Three, a personal relationship between the two department heads, usable, but a conclusion that leaves no record does not count. When you have none of the three, do not force the meeting. Fall back to the downgrade route. Write the two meanings into the value dictionary as separate columns nobody is allowed to merge, then put the divergence and its business consequence into the weekly report that copies the sponsor. Nothing moves the first time. After the second, it becomes the sponsor's agenda.
The reconciliation also pulled up an unplanned finding. The join key between claims and call center tickets is unstable. Tickets are joined by a claim report number customer service keys in by hand, one digit off makes an orphan, and the same claim can be found in several variants in the ticketing system. This is the fourth inconsistency pattern, the broken join, where two systems cannot line up the same entity. The "customer chase count" field therefore breaks at the traceable rung and is demoted to a reference signal, kept out of the priority logic. Linda's line in Chapter 0, "chase priority โ risk priority," was about business judgment. Here the data layer adds one more reason. The chase count itself does not line up with the claim.
| Pattern | Test Question | Prescription |
|------|----------|------|
| Lagging | Is the value true, only slow? | A faster source (the Excel, the mailbox), engineering can solve it |
| Forgotten | Is the status change forced by any process, and is anyone affected by it? | Operating discipline, engineering cannot solve it |
| Semantic divergence | Are two departments using one word to record different facts? | Align the semantics first, talk about syncing second |
| Broken join | Are the two systems talking about the same entity? | Demote to a reference signal, keep it out of the priority logic |
On Friday you take the reconciliation result into a one-hour meeting with Kevin and Linda, and it produces three decisions, written up as the source of truth decision log ([Template 9.4](../appendices/template-09-data-fitness.md)).
1. **The source of truth for claim status = the merged view**, with Excel as the authority, the core system as the fallback, and the extracted mailbox signal filling in the "waiting on documents" detail. The source of truth for amounts = the core system, where audit discipline happens to keep that field in line. One class of fact, one source of truth, with lineage marked on the cell.
2. **The merged view is purely derived and purely read-only**, writing back to no source system. That design later saved you half an hour at the security review run by Victor Reyes (Head of IT Security and Architecture). A derived view that does not feed the source keeps the audit boundary so clean he could not find anything to say (Chapter 12).
3. **"Writing status back to the core system" is listed as a business-side operating improvement for the pilot period**, owner Kevin. The review team spends ten minutes before leaving each day clearing status, and the queue produces a "status in doubt" list to help them find the ones to fix. This belongs to the business side's operating discipline, and it is not in your system's feature list. It has to be discussed first and started first, and the engineering half, the source swap, waits until it is settled, because it decides which source to swap to. Half the fix for a data problem is engineering, half is operating discipline, and you can only fix your half. Inside a company that sentence needs a footnote. You do not have the outside deliverer's line of division, and both halves will drift onto you. What replaces it is sequence, see the next section.
This reconciliation also left behind one data asset. The inconsistency rate, 87 of 200, 44%, became the baseline. When Chapter 11 builds golden cases, it is the floor figure for "how much error the data itself contributes."
## Who Owns Which Half, Sequence Replaces the Boundary
The outside deliverer has a very hard boundary. The engineering half is his deliverable, the operating half is the other side's housekeeping, and the division of labor sheet draws it plainly. You do not have that boundary. When Kevin gets busy, the least effortful landing for "status write-back" is "let the Digital Center figure something out in the system," and both halves drift onto you.
What replaces the boundary is sequence. Discuss the operating half first. Two reasons. The operating half runs on ten minutes at the end of another department's day, so every day you delay speaking is a day later it starts, while the engineering half you can do any time. And it comes first because whether Kevin puts his name to it decides which source the engineering half should swap to. If he does, the merged view is the authority. If he does not, you have to design for a core system that lags forever and find another signal. Reverse the order and you will finish the engineering half first, then discover it was built on an assumption nobody maintains.
The owner cannot be you. Putting your name on an improvement item of this kind announces that this half will not be fixed, and the engineering work is wasted even when it is done. So what you want from Kevin is not "support." It is his name in the owner column of the source of truth decision log, with a date.
How do you get him to write that name. Trade something for it, do not argue him into it. You give him two things. The first is decision rights over the source of truth. Which table is authoritative for claim status is his to decide, and your job is only to lay out the reconciliation evidence in full. Once it is decided, every downstream reading follows that definition, including when another department wants to plug in. The second is a say in the schedule. Once the operating improvement item sits under his name, it becomes a precondition for this project, the queue's launch schedule moves with it, and he has thereby gained a voice in the project's pace. You have to be equally clear about the other half. Without that name, the engineering half will not be built, because building it would only spin in place. That is not a threat. It is carrying "the fix has two halves" through to the end. This conversation happens in that one-hour meeting on Friday. Grant Whitmore does not need to be in the room, but the conclusion has to appear in the next weekly report that goes to him, so that a third person sees the claiming happen.
Finding nobody who will put a name to it is a real conclusion, not a failed conversation. Write it into the source of truth decision log, note that this class of fact has no operating owner, have the engineering side design for the worst case, and then honestly go swap the source, demote it, or shrink the scope.
## Failure Modes
**1. Using the analytics warehouse as an operating source.** Pulling from the BI (business intelligence) reporting store to drive same-day action, and every morning the queue chases claims that were resolved yesterday. The analytics store has the friendliest interface, the fullest documentation and the easiest permissions in the whole company. It was built to be looked at by people, and an engineer walking the path of least resistance toward it is entirely human. But T+1 data (today you can see only through yesterday) does not deserve a same-day action at the timely rung. A report can be a day late. Chasing cannot. The test is one sentence. Who was this source built to be consumed by? A source built for people to look at does not, by default, deserve to drive machine action.
**2. Underestimating permission time.** The architecture diagram is finished and the data has not been touched. The schedule says three days for "data access" and five weeks go by. Approval duration is owned by nobody in an organization. Each step takes only a few days, and no one is accountable for the total. And while your trust balance is still negative, the process is an open-ended "in progress" (Chapter 5). Data access is the hidden critical path on this kind of project. The request goes out in week 1, and the fix lives in the trust network, not in the ticketing system. At Anchor & Helm it took about two weeks after the charter, which counts as fast.
**3. Reconciling only once.** You reconciled during discovery (the stage of finding out how things actually are) and from then on "the data is fine" went into every report. That takes a one-time verdict for a permanent conclusion. Staff rotation, process revisions and upstream system upgrades all degrade it. The most insidious source of degradation is your own system. After the queue launches, Linda's team depends on the Excel less, and the appetite for maintaining that source of truth falls with it. Your system will kill its own source of truth with its own hands. So reconciliation needs a rhythm. Retest during the pilot, spot-check after launch (the monitoring in Chapter 18 takes over this line).
On someone else's project a retest is a delivery obligation, and when the date passes somebody asks. Inside, nobody will come and ask. Retesting during the pilot and spot-checking after launch are both things that will not hurt immediately if skipped, so they have to hang on a carrier that does not depend on your memory. Two options. One, into your own team's ops checklist, on the same sheet as certificate expiry and dependency upgrades, reviewed once a quarter. Two, into the quarterly data health report for the sponsor, one line per system, source of truth, last reconciliation date, inconsistency rate, retest conclusion, where the blank cells speak for themselves. The retest triggers need one more entry unique to the inside, "team staffing change." The person maintaining that Excel changed, or the person on your team who owns this line changed, and either one means running the reconciliation again.
## Next Monday
1. List the data sources for the project in front of you ([Template 9](../appendices/template-09-data-fitness.md)), including the ones you are embarrassed to call data sources, mailboxes and private spreadsheets. The truth often lives there.
2. Pick the most central field in the system, sample 50 records and run a three-way reconciliation, the system field, the private source of truth, asking the handler. Work out your own "87 of 200."
3. Triage with the inconsistency patterns. Lagging swaps the source, forgotten adds operating discipline, semantic divergence opens an alignment meeting first. Do not fix an organizational problem with ETL.
4. Check your architecture diagram. Does AI output get written back into any source field? If it does, change it to a decision trail today.
**Want an agent to get you started?** In the repo you set up following [Start Here](../index.md), paste this to your coding agent:
```text
In the repo/ directory of the the-last-mile repository, help me with the Chapter 9 Next Monday actions. Copy
templates/data-fitness/fitness-scorecard.md into the working directory I name. I list the data sources, and you are to
ask me directly whether sources such as mailboxes and private spreadsheets exist. Then run
python3 templates/data-fitness/reconcile.py on the built-in sample to demonstrate the three-way reconciliation summary,
then copy reconcile-records.csv into my own reconciliation log. I fill in the system field, the private source of truth
and the handler's answer for the 50 records, and you only compute the total rate and the largest single class.
The triage of inconsistency patterns (lagging, forgotten, semantic divergence) is mine to decide, do not judge for me.
If any command errors, stop and show me the output.
```
---
## Chapter Kit
- **Judgment frameworks.** The seven rungs of the data fitness ladder (exists โ accessible โ interpretable โ timely โ traceable โ validated โ actionable; judged rung by rung, relative to an action, and it degrades); three-way reconciliation (define the standard โ sample โ system field / private source of truth / ask the handler โ inconsistency rate + inconsistency pattern โ source of truth decision)
- **Templates.** [Template 9](../appendices/template-09-data-fitness.md), Data Source Inventory, Fitness Scorecard, Reconciliation Checklist, Source of Truth Decision Log
- **Key judgments**
- "A field's update discipline runs only as far as 'audit finds nothing.' AI is the first consumer that takes the fields at their word."
- "Half the fix for a data problem is engineering, half is operating discipline."
- "Fitness is a state, not a property, and it degrades."
- "A source built for people to look at does not, by default, deserve to drive machine action."
---
# 10 ยท Pick the Pattern: Start with the Dumbest Thing That Works
!!! info "Companion Templates"
๐ [Chapter Template](../appendices/template-10-pattern-decision.md) ยท ๐ [Template Library](../appendices/template-library-index.md)
> **The Challenge.** Does this scenario call for an agent, for RAG, or for a few rules? The whole world is talking agentic (letting the model decide its own next step and call its own tools), and the two claims-ops IT engineers arrive thrilled, carrying a multi-agent design. You know it is the wrong fit, and you cannot produce a reason that does not sound like "conservative."
>
> **What You Will Be Able to Do.** Use the pattern decision table (six patterns ร five criteria) to pick the dumbest thing that works for every step where AI intervenes in your system. Turn down over-design without dousing the enthusiasm. Leave "upgrade later" one legitimate channel that admits eval evidence only.
---
## Week 7, Monday, the Multi-agent Proposal at the Standup
The source of truth decision was settled last Friday (Chapter 9). At Monday's standup in week 7, the two claims-ops IT engineers, the same two who brought the general-purpose platform sketch last time (Chapter 8), bring something new. This time it is a demo that runs. The sketch stage is over.
The younger engineer opens his laptop. "Over the weekend we put up a multi-agent frame. A detection agent reads the claim and the emails and works out which documents are missing. A chase agent drafts the chase notice and sends it straight out. An escalation agent watches overdue claims, escalates automatically when they run past the deadline, and can even text the customer a progress update. The three agents call each other. Fully automatic." He runs three de-identified sample claims. Nobody touches them the whole way, they flow through, the logs look beautiful. "Linda's team handles only what the agents cannot. The rest is automatic. That is what an AI system means. That ranked queue thing, that is a spreadsheet with suggestions on it."
Concede something first. The demo is real, on three samples it did run end to end, and the two engineers are genuinely capable. They are also this system's future maintainers. The Chapter 22 handoff, to whom? To them. So your position right now is far more delicate than "turning down a design." Turn it down the wrong way and what you lose is two future co-builders' ownership (treating this system as their own, willing to backstop it when it goes wrong), and one architecture option is the smaller loss. Do not turn it down and the decision rights boundary drawn in Chapter 8 and the reconciliation just finished in Chapter 9 are both void.
Settle one question here first, who the future maintainer actually is. Inside a company there are three possibilities, and each answer calls for a different way of persuading. The first, the business line's own IT, which is what Anchor & Helm has. The two engineers are appraised inside claims operations, they can hold it and they cannot leave, so their ownership as co-builders is worth the effort of protecting. The second, a Group-wide ops team. They take no part in the design, so co-building is off the table, and all they can hold is something that runs, has a manual, and has an action for every alert. The complexity you pick decides directly whether they can hold it. The third is your own team. This one is the most common and the least said out loud. Here, whose ownership to protect is no longer the question, and the strongest argument turns into a different sentence. Every complex pattern you pick is your own night shift for the next three years.
There is another layer of difficulty, and it is more personal. "That is what an AI system means" pokes at the doubt in your own head. Building a "spreadsheet with suggestions on it," am I falling behind?
## Why This Is Hard: The Criteria Are Misaligned
When engineers argue about patterns, the default criterion is the capability ceiling, the most it can do. And the ceiling is exactly what a demo shows. The real criterion in an enterprise setting is the shape of the error, whether an error can be caught, attributed, and rolled back. How clear the shape is can be tested on the spot, whether the ways it fails can be listed in full in advance, and whether, once it has failed, you can locate which step failed and who backstops it. The two sets of criteria point in two directions. The higher a pattern's capability ceiling, the harder its error shape is to draw.
Enterprises have no demand for "occasionally impressive." They have a hard requirement for "when it is wrong, it can be caught, attributed, and rolled back." Impressive does not settle Tuesday morning's exception claims. A backstop does. This has nothing to do with being conservative. It is a property of the environment. Chapter 7's five questions already asked, at the opportunity level, how many times AI is allowed to be wrong and who backstops it. Pattern selection is the same question pushed down to every single step.
The second layer of difficulty is signal pollution. A pattern's popularity and its applicability are badly decoupled. Put plainly, how hot a pattern is and how usable it is are two different things. Heat is driven by demos and narrative, and the more demonstrable a pattern, the hotter it runs. Applicability is decided by error tolerance, and the harder a pattern's error shape is to draw, the narrower its applicability. The same property (autonomy, unpredictability) raises the heat and narrows the applicability at once, so the hottest is precisely the narrowest. The more demonstrable a pattern, the fewer places you can use it with confidence. Chapter 1's demo inflation has a pattern-selection version here. Demonstrating an agent chain has never been cheaper. Productionizing it has not come down a cent.
The third layer is the social cost. "No agent" sounds in a meeting room like "cannot use agents." What you need is what you needed when eliminating ideas in Chapter 7, a ruling procedure that does not rest on anyone's position, that lets the evidence speak and lets the proposer reach the conclusion himself.
## Prior Art, and What AI Changed
"Pick the dumbest one" is not a new discipline. It has a fifty-year pedigree.
**Gall's law.** The systems scientist John Gall's observation in *Systemantics* (paraphrased). A complex system that works is invariably found to have evolved from a simple system that worked. A complex system designed from scratch does not work, and it cannot be patched into working. It can only be torn down and started again from a simple system that works.
**"Choose boring technology."** The engineering manager Dan McKinley's well-known 2015 argument (paraphrased). Every team holds a limited number of "innovation tokens," and they should be spent on the core of the business. Boring technology is strong because its failure modes are known. How it breaks and how it gets fixed, someone walked into all of it ahead of you.
**The trade-off discipline of architecture review.** Chapter 8 quoted it already, put the cost next to every "want." Pattern selection is its technical depth. Every notch down the pattern spectrum (the next section lines the six patterns up from dumbest to smartest), what you "want" is a higher capability ceiling, and the cost written next to it is verifiability.
The AI era changed two things in this old discipline.
**One, the object of selection changed.** Selection used to mean components, a database, a queue, and a component's behavior is deterministic. Today there is a dimension that did not exist before, how much uncertainty you put into the system and at which step. The six patterns in this chapter are, at bottom, six doses of uncertainty.
**Two, "boring" has to be redefined.** The old test for boring was "ten years in production, fully documented," and AI components are all young, with no ten years to look up. Boring in the AI era can only be judged structurally. The output can be checked, the errors can be enumerated, the behavior can be reproduced. That is what "dumbest" means. The error shape is the clearest, and how weak the capability is, is a separate question.
Inside a company, this old discipline can also borrow two things you cannot borrow outside. The first is the company's existing technology stack standards and the architecture committee's whitelist. They exist in the first place to limit freedom of selection, so cite them and you do not have to say "this framework is immature" in your own name. You only have to point out that it is not on the whitelist, and what the other side is negotiating with turns from your technical taste into an institutional document, and an institutional document does not need you to defend it in the meeting. The second is the corpses. The company's previous agent project is most likely still around. Go find out how it really died, who wrote it, who maintains it, whether anyone still uses it, and what criteria it was picked on at the time. These are things a team from outside can only guess at. You can look them up, and you can put what you find into the anti-pattern list, where it becomes evidence in your own company's version. The anti-pattern list is exactly that, those causes of death written down as entries you can check against, and this chapter's template has one. Cite the whitelist with restraint. It counts as a red line only if you can name which standard it is. A "the company does not allow it" that cannot name the clause is only a shield.
## The Framework: The Pattern Spectrum and the Decision Table
The spectrum first. The pattern spectrum is the six candidate forms for AI intervening in one step, ordered from dumbest to smartest by how clear the error shape is. The further down, the higher the capability ceiling and the harder it is to draw a shape for the error. This order gets used again and again below.
- Rules. The ways it fails can be listed in full in advance, and when it fails you know which rule matched wrongly.
- Structured extraction + rules. The error is fenced into the extraction step, and a bad format is caught by validation on the spot. Structured extraction means pulling the key information out of free text into a fixed format.
- RAG (retrieval-augmented generation, having the model look things up before it answers). The ways it fails are open-ended. A missed retrieval and a wrong retrieval both leave no visible trace, and the only check is verifying the citations one by one.
- Single-step LLM. Every call can be wrong, and the shape depends on how tightly you pin down the output format.
- Agentic workflow. Errors compound and amplify across steps, the step that failed is hard to locate, and it may not reproduce.
- Fine-tuning (changing the model weights themselves, retraining on annotated data). Errors set into the weights, invisible, and there is no way to strike out one of them on its own.
The decision table is this chapter's main kit, six patterns ร five criteria, a judgment in every cell and no scores. Scores get averaged. A judgment can only be rebutted (Chapter 7 covered how a weighted average kills elimination).
Where the five criteria come from. Error tolerance carries on from Risk, one of Chapter 7's five questions. Verifiability and data requirements stand on Chapter 9's fitness ladder. Maintainability by the receiving side draws on Chapter 22's handoff. Latency and cost get their own accounting in Chapter 16.
The maintainability column is the one most easily read too lightly inside a company. A team from outside leaves when its term is up, and this column's cost settles in a single payment on the day they go, visible and impossible to dodge. You do not leave, so the cost never settles on any one day. It spreads by the week into your ordinary days, into two extra alerts a week, three extra days of regression at every model generation change, two extra weeks for every new hire coming up to speed. A cost spread that thin has no due date, so nobody adds it up for you, and the fifth criterion becomes the cell in the whole table most likely to be filled in as "should be fine." Before you fill this cell, name the receiving side. If you cannot name one, the receiving side is you.
Read the table top down, and all five cells in a row have to pass. If one cell's judgment does not hold in your step, look at the next row. Stop at the first pattern whose five cells all pass.
| Pattern | Error tolerance | Verifiability | Data requirements | Latency and cost | Maintainability by the receiving side |
|------|-----------|----------|----------|------------|--------------|
| **Rules** | Ways it fails are enumerable, fit for zero-tolerance steps | Reproducible and explainable rule by rule | Fields only need to be interpretable | Milliseconds, near zero | The business line's IT can change it themselves |
| **Structured extraction + rules** | Error is fenced into the extraction step, schema (the fixed field structure the extraction output must satisfy) validation intercepts most of it | Extraction output can be validated and spot-checked | Unstructured source accessible, plus a spot-check sample set | One call per claim, batchable | Rules go to the business side; schema and prompts need a handoff |
| **RAG** | Ways it fails are open-ended (missed retrieval, wrong retrieval), needs a human backstop | Verified by checking citations, a rotten corpus rots everything | One genuinely trustworthy corpus | Two hops, retrieval plus generation | What gets maintained is really the corpus, and decay is hidden |
| **Single-step LLM** | Wrong on any call, only for steps where a person reviews | Hard; pinning the output format and requiring a reason partly makes up for it | No training data needed, but no eval means no threshold | One hop, controllable | Prompt drift goes unnoticed |
| **Agentic workflow** | Errors compound and amplify across steps | Path varies, failures are hard to locate and hard to reproduce | Every step has to pass fitness and eval on its own | Multiplied by hops, latency unpredictable | Hardest to hand off, debugging often needs the builder himself |
| **Fine-tuning** | Errors set into the weights, fixing one means retraining | Behavior changes globally, every retrain needs a full regression | Hundreds to thousands of annotations (as of writing, the bar is still falling) | Training cost paid up front and paid again | The receiving side can hardly take it over at all |
Three rules for using it.
**One, pick per step, not per project.** "Does this project use agents" is a fake question. Inside one system, error tolerance differs wildly from step to step, and the pattern follows the step.
**Two, read top down and stop at the first pattern that passes all five criteria.** That is the operational definition of "start with the dumbest thing." The test for "works" is five cells passing, not the most impressive result. A cell passes when its judgment still holds once dropped into your step, and the risk left over has a backstop you can name or a checkpoint. Name neither and the cell does not pass. If rules can solve it, no LLM. If a single step is enough, no agent.
**Three, an upgrade has exactly one legitimate channel, eval data showing that the dumber pattern cannot clear the threshold (Chapter 11).** "It does not feel smart enough" is not a reason to upgrade, and "I want to learn this framework" is even less of one. Start with the dumbest thing, and let the evidence force the upgrade.
## The Complexity You Choose Is Your Own Night Shift
Outside, these three rules are engineering discipline. Inside a company they are also an account in your own name. When a team doing external delivery picks an agentic workflow, the pain is borne by the receiving side from the day of handoff onward. When you pick an agentic workflow, the pain will most likely be borne by you and your team forever, because you cannot leave (Chapter 2's permanent ops). A model generation change means rerunning the regression eval, a framework upgrade means revalidating every step, and when that step fails at two in the morning, the person woken up is you. There is no such thing as a project ending. There is only the next project starting while the last one still hangs on you.
So the discipline of "start with the dumbest thing" is one an internal reader is better placed to state than anyone, and what he is talking about is his own calendar, not somebody else's architectural taste. Translate it into a sentence you can say out loud in a meeting. "Every extra notch of complexity we pick is another slice of ops headcount this system takes next year, and that headcount comes out of our team's capacity for new projects." The sentence needs no moral position. It is headcount.
The fifth criterion gets stricter in this situation, not looser. When there is a receiving side, handoff day is a mandatory physical. The five self-sufficiency tests (Chapter 22, the ones mentioned in Chapter 3, which test whether the receiving side can run this system on its own) expose the complexity you picked back then on the spot. When the maintainer is your own team, that physical does not exist, and nobody comes to check whether you can maintain what you wrote yourself. The criterion has to be set by you. Suppose this step is being debugged at midnight three years from now by a colleague who has not been hired yet. Could he possibly handle it alone? No answer, and this cell does not pass, by exactly the same standard as when there is a receiving side.
## At Anchor & Helm: Three Steps, Three Pattern Choices
The queue has two steps where AI intervenes, and the engineers' proposed "fully automatic" counts as a third. Take them through the table one at a time.
**Step one, missing-document detection, structured extraction plus rules.** The rules row does not pass in this step. The cell that says "fields only need to be interpretable" does not hold, because emails are free text and have no fields. The first conclusion the decision table produces is an upgrade, from rules to the next row down. The real detail of "waiting for the customer to send documents" lives in the emails (Chapter 9), and this is exactly the job only an LLM can do, understanding a line like "the advance payment receipt will start coming in next Monday." But it is allowed to output only one fixed schema, document type, party to act, promised time, source email number, four fields, and any output outside those fields is rejected.
The rules layer takes the schema and decides how the queue displays and whom to chase. The LLM will still be wrong, but the error is fenced tight inside the extraction step. Schema validation intercepts format errors on the spot, and a weekly spot check of twenty rows watches for semantic errors. In one sentence, put the LLM where its output can be checked.
**Step two, the priority suggestion, rules as the floor and a single-step LLM for the long tail.** Linda's five rules of thumb (Chapter 6) are made explicit as the rules layer, most claims get their suggestion from the rules, and the reason column says outright which rules matched, "report delay + missing photos." The long tail the rules do not reach, signals contradicting each other, a prior claim linkage that looks likely on weak evidence, goes to a single-step LLM, with the output confined to a suggested priority, a reason, and which signals it cited. Either road, every suggestion carries its reason, and the "Human Call" column decides. This arrangement also holds the line Chapter 6 warned about. The rules are made explicit but not hardcoded into if-else that nobody reviews, people keep the decision, and overrides leave a trail. The decision trail is the detector for rule drift.
**Step three, fully automatic processing. Turned down, but let them turn it down themselves.**
At the standup you rule on nothing. You only say, "Interesting design. Block two hours Wednesday afternoon and we will take one table through it. Not judging the design good or bad, filling it in row by row against the criteria. You fill it in."
Wednesday, an empty decision table is drawn on the whiteboard. First row, error tolerance. You ask, "The chase agent's notice goes straight out. Out to whom?" "The surveyor, and in some scenarios the customer." You ask next, "The escalation agent texts the customer automatically. What is the consequence of one wrong text? Who sees it?"
The room goes quiet for a dozen seconds. The older engineer gets there first. Once a text leaves the company it has reached the customer, the error is external and cannot be recalled, and it is the same reason the service chatbot died in the Risk cell in Chapter 7. What an insurer says to the outside is regulated. He writes it into the first row himself. Tolerance near zero, and nobody backstops it, since the whole point of an agent is to route around people.
Second row, verifiability. "Three agents passing work between them, one claim comes out with a wrong result. Which step got it wrong? Does the same input still reproduce it?" The logs look beautiful, but those are the logs of three samples that went smoothly. What the logs of a failure path look like, the demo did not answer.
Third row, data requirements. You raise one number and nothing else. "The detection agent and the escalation agent both consume claim status. In last week's reconciliation, 87 of 200 cannot be taken at face value (Chapter 9). A person who sees 'in progress' goes and checks the emails. Every step of the agent chain acts on face value, and the action goes out the door."
The fourth row never gets filled in. The older engineer puts down his pen. "No need. Inside the advise layer, this thing is not needed. Outside the advise layer, not one criterion passes." He pauses. "The maintainability row I will fill in for myself. When the queue breaks at midnight, tracing an agent chain and tracing 'which rule matched wrongly' are two different lives. The two of us are the ones taking over maintenance." The table stops there, and that is allowed. A table may stop at the first cell that does not pass, and the rest need not be filled in. One cell failing puts the design out, and filling in the cells that remain only supplies material for a design already out.
What you add is not consolation. "The design is not wrong. The layer is. The ladder is still that same ladder (Chapter 8). Let the AI's suggestions earn a record of being accepted before you talk about letting it act. On the day we do talk about acting, whether to use agents is for the eval to say."
The scope decision log gains a line, next to Grant's line from Chapter 8. That line governs which layer decision rights stop at. This one governs the implementation. Multi-agent automatic processing, proposed by the two claims-ops IT engineers, reason for refusal, the errors go outside and cannot be recalled, they are hard to locate, and the data does not support it. Revival condition, once the advise layer's acceptance data hits target and the eval shows that rules plus a single step cannot clear the threshold, move up one layer at a time along the decision rights boundary.
Then you point them at the hardest engineering in the whole system, the extraction pipeline. Emails in, schema out, validation, the spot-check tool. "This is the most technically demanding stretch in the system, and it is also what you two will own later. The prompt inside your detection agent, strip off the agent shell around it, and it is version one of the extractor." The two claim it on the spot. The demo was not built for nothing, the enthusiasm was not doused, it only landed somewhere else. This Wednesday afternoon is the starting point of Chapter 15's co-build (building alongside the engineers who will take over).
If the proposer is your manager, or the Group architecture committee, this meeting has to be run a different way. Across a company boundary, "let the evidence speak" is a neutral procedure by nature. Inside a company, the same move can be read as a subordinate using procedure against a superior, and once it is read that way, whichever cell you win counts for nothing. First do one thing that has nothing to do with the design. Take "ruling by the table" to the sponsor (the executive who funds it and makes the call) as a procedure, once, and once is enough. What you want is not his backing for one design. It is his endorsement of the rule, that every step where AI intervenes gets filled in row by row against the same table, that it stops at the first cell that does not pass, and that the rule treats every proposal alike, including the ones you make yourself. With the procedure endorsed, the meeting is not you against him. It is a proposal against a table.
Then hold three things in the meeting. You are not the facilitator. Ask the business-side owner (the person in the business department accountable for this system) or the proposer himself to chair, and you only ask questions, you do not rule. The conclusion goes into the scope decision log, with the reason for refusal and the revival condition, and not a word about whose judgment is faulty. The table stops at the first cell that does not pass, and whichever cell it stopped at is the cell you record, so the conclusion points at that cell's evidence and not at the proposer.
## Failure Modes
**1. Resume-driven architecture.** Somewhere in the selection discussion comes "this project is a good chance to get fluent in the agent framework." An engineer's market value is calibrated by how new his stack is, a project's success or failure settles three months later, and a resume can be updated the same day. The two accounts settle on different cycles, and people rationally settle the faster one first. Chapter 8 said abstraction is an engineer's identity currency, and on a resume "platform" is worth more than "spreadsheet." Patterns are another kind of the same thing. The newest pattern is worth more than the most suitable pattern, in the labor market, not on the project. Inside a company this account carries on under a different name, whether this project can go into your promotion package, and it still settles faster than the project's success or failure.
And you are not the only one keeping this account. The department's annual report also needs a line that says "we shipped agents," and that pressure runs top down, landing on the line that runs your schedule and your review, which is harder to push back on than one resume.
The defense is procedure, not morals. Every pattern proposal goes through the decision table first, and the criteria do not recognize a resume, nor the wording of an annual report.
**2. The demo's pattern straight into production.** The demo used agents and wowed the room, and once the project is approved "it already runs" carries it forward. Demo and production select on different criteria. A demo is picked for how impressive it is, production for whether it can be backstopped. "It already runs" manufactures the illusion of a sunk cost, and changing pattern looks like going backwards. What this step walks into is exactly the data gap and the ownership gap among Chapter 1's five gaps. A demo's samples are always clean, and a demo's errors are never attributed to anyone. One rule. Build the demo with whatever pattern you like, and any pattern going into production goes through the decision table again.
**3. Putting uncertainty into a step that cannot be wrong.** An LLM-generated amount flows straight into the payout calculation, a generated date straight into a regulatory filing. Probabilistic behavior is invisible on an architecture diagram, where an LLM is only a box and looks the same as a database. Error tolerance is a property of the step, not of the system, and if it is not drawn nobody aligns on it. Once they blur together, the whole system is in fact running to the standard of its loosest step. The defense, mark on the architecture diagram, at every AI intervention point, that step's tolerance and its backstop. For a generated number to enter a deterministic step, there has to be a layer of validation or a person in between.
## Next Monday
1. List the AI intervention points in the system in your hands, mark the current pattern on each step, then ask "why not one notch dumber." For any step where you cannot answer with evidence, drop a notch and try it for a week.
2. Check where every LLM output lands. Is any generated number or fact flowing straight into a step that cannot be wrong (a calculation, a filing, an outbound message)? If so, put a layer of validation or a person in between.
3. Next time somebody proposes an agentic design, do not debate it. Book two hours and have the proposer fill in the decision table in [Template 10](../appendices/template-10-pattern-decision.md) row by row himself.
4. Find where the criterion for "upgrading" is written down in your project. If you cannot produce it, you do not have an eval yet. Go read Chapter 11 first, then talk about upgrading.
**Want an agent to get you started?** In the repo you set up following [Start Here](../index.md), paste this to your coding agent:
```text
In the repo/ directory of the the-last-mile repository, help me with the Chapter 10 Next Monday actions. First, following
templates/pattern-selection/decision-table/prompt-intervention-points.md, list the AI intervention points from the process I
describe, mark each step's current pattern, then for each one ask me "why not one notch dumber." Flag the ones I cannot answer
with evidence. Whether to drop a notch is mine to decide. Then copy decision-table.md to me to fill in row by row. Finally run
python3 templates/pattern-selection/schema-check/check_schema.py and python3 templates/pattern-selection/spot-check/sample_cases.py
with the built-in samples to show what validation and spot-checking look like. If any command errors, stop and show me the output.
```
---
## Chapter Kit
- **Judgment frameworks.** The six-pattern spectrum (rules โ structured extraction + rules โ RAG โ single-step LLM โ agentic workflow โ fine-tuning, ordered by how clear the error shape is); the pattern decision table (six patterns ร five criteria, a judgment in every cell and no scores; pick per step, read top down and stop at the first that passes all five, upgrade only on eval)
- **Templates.** [Template 10](../appendices/template-10-pattern-decision.md), the Pattern Decision Table and Anti-pattern List, the full decision table ready to copy, how to use it, anti-pattern warning signals and their corrective actions
- **Key judgments**
- "Enterprises have no demand for 'occasionally impressive.' They have a hard requirement for 'when it is wrong, it can be caught, attributed, and rolled back.'"
- "Start with the dumbest thing, and let the evidence force the upgrade. The only legitimate reason to upgrade is eval data showing that the dumber pattern cannot clear the threshold."
- "A pattern's popularity and its applicability are badly decoupled. The hottest is precisely the one with the narrowest applicability."
- "Put the LLM where its output can be checked."
---
# 11 ยท Eval as Spec: Write the Eval Before the Architecture
!!! info "Companion Templates"
๐ [Chapter Template](../appendices/template-11-eval-spec.md) ยท ๐ [Template Library](../appendices/template-library-index.md)
> **The Challenge.** The business side asks "what accuracy can this reach," and you cannot give a number. "Accurate" has never been defined. Most internal projects have no acceptance meeting, but on launch day someone still has to release it, and any number you report now turns into an argument on that day.
>
> **What You Will Be Able to Do.** Write the five parts of an eval spec before drawing the architecture. Co-build golden cases with the actual user and treat annotation disagreements as requirements discovery. Use that spec to rule on architecture upgrades and launch release, and use the co-build to get business rules that were never aligned across departments written down for the first time.
---
## Week 7, Thursday, a Question Nobody Could Answer
The pattern discussion wrapped up yesterday (Chapter 10). Kevin Doyle calls you into his office and asks a question he has been holding for a while. "This system, what accuracy can it reach?"
Behind the question is real pressure. The acceptance line in the charter reads "tie it to the eval, golden cases plus thresholds, sign the mechanism now and fill in the numbers later" (Chapter 4). Golden cases are real cases with the correct answer settled in advance. The old ticket's "Q&A accuracy โฅ90%" was voided and filed at the project approval change review, and now the moment to "fill in the numbers later" has come. Kevin answers to Grant Whitmore. He needs a number.
You do not give one. You ask back. "Last week a claim with a questionable amount in Linda Marsh's team went through as a routine claim. Whose error is that? Is it an error at all?"
The room goes quiet for a few seconds. By the core system, that claim showed nothing abnormal. The status moved normally, no deadline was missed. By Linda's instinct, it should have been flagged red. By the customer service log, the customer even praised the speed. Same claim, three answers.
"If I tell you '95%' right now," you say, "on release day we will fight over every single claim, because we never agreed on what counts as an error. Give me a week and a half and I will give you a table you can sign. It is worth more than a number. The signatures are yours and Victor's. You are the business-side owner, Victor Reyes is the risk owner, and whether this system goes live was never supposed to be decided by the people who built it. There is an annotation session in week 8. Come sit in for twenty minutes, and you will watch the denominator of that number being made."
Where there is no definition of error, there is no accuracy, and no release either. This chapter is about how to make that definition.
## Why This Is Hard: An AI System's Spec Cannot Be Written as "Input X, Output Y"
A traditional software spec describes behavior. Input X, output Y must follow. It can be enumerated, walked through, and ticked off on an acceptance sheet. An AI system cannot do that. The same input can produce different outputs, and the input space is never exhausted. So many teams skip the spec, paper over it with "it works well," and postpone defining error to the moment disagreement costs most, the last week before launch.
The way out is a different form of spec. Writing the behavior spec in finer detail will not save it. An AI system's spec does not describe behavior. It describes the acceptable distribution of errors. The system will err, but what it gets wrong, how much, and where a wrong output goes can be agreed in black and white. The thing that carries that agreement is the eval. Golden cases define what the system "should be." The error taxonomy and thresholds define "how good is good enough." What these three look like, the "At Anchor & Helm" section shows with concrete cases. The eval set is the requirements spec itself. Being a testing tool is its part-time job.
That also explains the words "before the architecture." Chapter 10's governing principle is "Start with the dumbest thing, and let the evidence force the upgrade," and the evidence can only come from the eval. Without an eval written first, a pattern upgrade has no criterion, architecture trade-offs decay into contests of taste, and acceptance decays into the political question of "does the boss think it is fine." This is the core discipline added in the AI era, and one of this book's two technical anchors. The other is the action queue of Chapter 17 (the queue that lines up the next action for every claim).
## Prior Art, and What AI Changed
Write "what counts as right" first, then write the implementation. That order has two mature ancestors.
**The engineering tradition, TDD's "tests as spec."** Kent Beck's test-first discipline is at its core about writing the test first so that you are forced to think the requirement through before writing code. The test is an executable spec, and how many tests there are is secondary. Eval as spec is its closest relative. Golden cases are the AI system's "tests written first."
**The consulting tradition, aligning the picture of success before the work starts.** Peter Block's contracting discipline in *Flawless Consulting*. What the client will see on the day the project succeeds must be aligned in concrete language before the work starts, or delivery day is disappointment day. The eval spec turns that "picture of success" from adjectives into fifty concrete cases plus a threshold table.
What did AI change? This is a new discipline, with three things the old traditions did not have. First, an error taxonomy replaces a single accuracy figure. In traditional software a bug is an anomaly. In an AI system errors are the normal state and come as a distribution, and the kinds of errors and their consequences matter far more than the total. Second, tacit knowledge has a systematic home for the first time. The five rules of thumb dug out of Linda in Chapter 6 could only be passed on verbally before. Now they have a place to be written down. Third, Chapter 2 concluded that this method adds exactly one skill of its own, AI uncertainty management, meaning eval, human oversight, and drift (the system's results quietly degrading over time). That sentence starts getting paid off in this chapter. The eval is the first load-bearing wall of that skill. Human oversight (Chapter 12) and drift monitoring (Chapter 18) are both built on top of it.
## The Framework: The Five-Part Eval Spec
| # | Part | Question It Answers | Discipline |
|---|------|-----------|------|
| 1 | **Task definition** | What is being evaluated | One sentence at the level of the workflow claim (Chapter 0). What input the system takes, what output it gives, for whom, and where the red line is. One eval evaluates one task |
| 2 | **Golden cases** | What counts as right | 50 to 200 real cases + correct answers. The count is set by coverage of error categories, every category must have cases that can detect it, never padded to a total. Answers clear the data check first. Edge cases and historical incident cases must carry real weight |
| 3 | **Error taxonomy** | What kinds of error there are | Each category states its definition, business consequence, and tolerance. Reuse the pass / concern / unsafe / useless error severity vocabulary (this book sets no separate sev-1/2/3 numbering; an error category's consequence tier borrows three of the four grades directly, so the scoring words and the severity words are one set). The unsafe class is zero tolerance |
| 4 | **Per-category thresholds and acceptance lines** | How good is good enough | One line per error category. No overall score. An overall score is the shortest path to burying unsafe |
| 5 | **Human review path** | Where outputs that fail go | Trigger condition, destination, who looks, how fast. Review results flow back into the golden cases |
Three rules of use. **First, golden cases must be co-built with the actual user.** The correct answer is in Linda's head, not in your reasoning. And the by-product of co-building, annotation disagreement, is itself requirements discovery (the field section demonstrates it). **Second, reconcile the answers first.** Every fact a golden case's "correct answer" cites must stand at validated or above on the data fitness ladder. Why, the next section will lay out Chapter 9's reconciliation numbers and do the arithmetic. **Third, set thresholds per error category.** The costs of different errors differ by orders of magnitude. One averaged line guarantees nothing.
## At Anchor & Helm: How Fifty Golden Cases Grew
In weeks 7 and 8, you and Linda's team did this. The time was not ready-made. The charter line "Linda's team, 2 hours a week, for case annotation and rule confirmation" (the business-side commitment line, Chapter 4) is one line of text, and nothing turns it into Linda's team's calendar on its own. You did three things. First, no new meeting. The annotation session was attached to the review team's existing weekly, run straight after it, so Linda did not have to fight for a new slot on your behalf. Second, state the exchange plainly. Your twenty years of rules get written down for the first time, and once written, the guide belongs to the review team, not to the system. Third, the fulfillment rate is reported monthly at Grant's weekly (Chapter 4). Two hours this month or zero, both sides can see it.
The division of labor follows Chapter 0's principle. AI does the grunt work (screening candidate claims and pre-filling fields, drafting first-pass answers), people make the calls (the final ruling on every answer can only come from a reviewer). The first line of work is the task definition, lifted straight from the sentence the Five Ones of Chapter 8 lock into. For every auto exception claim in the merged view, give the next action (whom to chase, what is missing, who takes it), for Linda's auto claims review team, and never touch the payout decision. All fifty cases are drawn against that sentence.
The answers themselves clear the data check first. You pulled eighty candidate claims from the merged view, and step one was reconciliation, with annotation after it. Chapter 9 did the arithmetic. 87 of 200, that is 44% of the fields, cannot be taken at face value, and here that becomes a discipline. Take the core system's fields directly as "correct answers" and nearly half your answers are wrong, which is accepting the system with a bent ruler. So every candidate claim is checked first. Where the system fields, the Excel tracker, and the inbox agree, it goes in. Where they disagree, run the three-way reconciliation ([Template 9.3](../appendices/template-09-data-fitness.md)) to rule on the truth, and drop what cannot be ruled on. Eighty screened down to fifty-nine, take fifty. The first mile of the eval is reconciliation, and annotation comes after it.
The error taxonomy grew out of three unsafe marks. On Field MVP day in Chapter 0, Linda marked 3 as unsafe, and the four-grade scale you used then was pass / concern / unsafe / useless. Looking back, that scale was the seed of the error taxonomy. It was already layering errors by consequence, it just had not grown categories yet. The first formal category grew precisely out of those three unsafe marks. All three pointed at the same thing, a risky claim queued as a routine one, and so it got a name.
| Error Category | Definition | Consequence | Severity (error severity vocabulary) | Tolerance |
|----------|------|------|------|--------|
| **Missed risk** | A high-risk claim judged routine | Grows into a major case, regulatory complaint | unsafe | **Zero tolerance, 0 of 50** |
| **Line-crossing suggestion** | A suggestion touches the payout decision red line (Chapter 8) | Responsibility has nowhere to land | unsafe | **Zero tolerance** |
| **Material error** | Missing-material list wrong or incomplete | A wasted round of chasing, the customer waits days longer | concern | โค10%, and the error must be obvious to a reviewer at a glance |
| **Wrong action** | Next action or owner given wrong | One idle round | concern | โค10% |
| **Empty suggestion** | The suggestion is not wrong but carries no information | The queue slowly goes unread | useless | โค20%, negotiable, trend must not rise |
The two zero-tolerance rows come with an ugly statistical truth first. 0 of 50 is an entry ticket, not a proof of safety, because fifty cases cannot detect every surprise in real operation. The real defense for a zero-tolerance category is every decision trail entry after launch (Chapter 18).
The three ratio rows have an ugly truth too. Zero tolerance is counted in cases and does not depend on sample size. Ratio lines like โค10% and โค20% on fifty cases mean one case is two percentage points, and the ratio read from a single replay swings by a dozen points or more on its own. It gives direction, not a scale. So ratio thresholds are read as a trend across several replays, and that is what "three consecutive replays" at Chapter 14's gate one is about. To read the line finer, add cases, not just confidence.
The annotation guide, five rules of thumb on duty. How do you annotate missed risk? The answer is the five rules the three-layer probing of Chapter 6 dug out, amount deviation, report delay, survey photo count, prior claim linkage, repair shop list, with the judgments written into the guide exactly as Chapter 6's five lines have them. They go from verbal tradition to annotation guide, and "missed risk" splits accordingly into five testable ways to fail. Which signal the system missed is visible at a glance. Chapter 6's promise (tacit knowledge turns from interview color into a system asset) is paid off here.
One gets special handling, the repair shop list. Linda said outright "it is not a fixed list," so that section of the guide carries a dated version and rolls with the list. That is the first reason the eval must stay alive (failure mode 4 settles the account).
The makeup of the fifty is deliberately uneven. Typical claims are under half. A dozen or so edge cases (amounts right at the top of the usual range, reports filed exactly two days late, no more, no less). Ten historical incident cases, including the out-of-pocket repair claim Linda herself described misjudging in Chapter 6. That is where the system's value and its risk both live.
At the annotation session on Tuesday of week 8, Kevin came to sit in, as promised. He was not there to supervise. At next quarter's operations meeting he has to present this project, and when he does he has to report a number, and the denominator of that number is being made in this room. You arranged double independent annotation. Every claim is annotated by Linda and by another senior reviewer on her team (second only to her in seniority), each on her own. Forty-six of fifty agree. Of the four disagreements, three closed on the spot, each with a different root.
- One was different information. The other reviewer only looked at the photo count recorded in the core system and never opened the inbox attachments. Once filled in, the annotation changed to agree.
- One was different standards. A claim reported two days late straddled a weekend, and whether that counts as "beyond the ordinary" split them. Kevin ruled that report delay is counted in working days, and it went into the guide.
- One was a different reading of the task. The other reviewer had annotated missing materials as missed risk too. The task definition gained a sentence. Missed risk covers the five signals only. Missing materials belong to material error.
All three went into the disagreement log.
Item 37 blew up. A claim ending in 2093. The plate was new, but the filer's phone number was appearing for the third time in six months. Linda marked "high risk." The other reviewer marked "routine."
Both froze, each thinking the other was joking. You took the judgment apart the three-layer probing way. Linda's "prior claim linkage" counts three things, plate, filer, phone number, and a phone number appearing a third time is, to her, a strong signal of an intermediary filing on the customer's behalf. The other reviewer counts only plate and filer, and her reason was solid too. "Repair shops file for customers all the time. Flag them all and you hit a crowd of innocent ones." Two standards, each in use for over a decade, each working. Same team, same rule, two standards. Twenty years of verbal tradition never aligned them, because there was never an occasion that required them to write the answer on paper and then look at each other's. The eval is that occasion.
Linda turned to the observer's seat. "Kevin, your call." Kevin heard both sides and ruled on the spot. A repeated phone number counts toward the linkage signal, but on its own it does not make a claim high risk, it only triggers human review. The exception scenario for shop-filed claims goes into guide section 4.2. On the Annotation Disagreement Log ([Template 11.3](../appendices/template-11-eval-spec.md)) you wrote the first line. The disagreement, both sides' reasons, who ruled, the rule that came out.
On the way out Kevin said at the door, "Last week you asked me 'whose error is that.' Now I see why you would not give a number." The conclusion is worth a memo. Annotation disagreements are business rules that were never aligned, surfacing, and treating them as noise is a huge loss. The eval had not yet evaluated a single system output, and it had already made this company write a twenty-year verbal rule down for the first time.
On Friday of week 8, you handed over the per-category threshold table above plus one human review path, far more substantial than a number. The review path is written to the four items of the fifth part. Trigger condition, claims that hit a risk rule with low confidence, plus anything that looks like an unsafe-class error. Destination, the queue's Human Call column, and without that column the system takes no action. Who looks, the reviewer who picked up the claim, by name, and every suggestion carries its reason, so judging it right or wrong does not mean checking three systems. How fast, the number of claims entering human view each day is worked back from the review team's review time budget, and it is better to route too few than to have them glanced at and not read. The answers to Chapter 12's three oversight questions are already written here. Suspected unsafe-class cases are reviewed weekly, and review conclusions flow back into the golden cases. The golden cases are the requirements document, and the thresholds are the acceptance clauses. From here on this table serves several purposes at once.
- Where Chapter 10's upgrade criterion lands. Only when "structured extraction + rules" fails the missed-risk line does a heavier pattern earn a hearing.
- The charter's "fill in the numbers later" is formally paid off here. The old ticket's wording was voided and filed at the project approval change review (Chapter 4), the vacated slot is filled by this table, and the thresholds go straight into the project resolution as the launch release conditions.
- Raw material for the prototype-to-pilot gate (Chapter 14).
- After launch it becomes the live monitoring definitions as is (Chapter 18).
## When Nobody in the Room Can Rule
Kevin's on-the-spot ruling was lucky. The disagreement fell inside his own department's rules, he had the authority, and he was willing to spend it right there. Cross-department disagreements do not look like that. Imagine an annotator borrowed from customer service also sitting at the session. On one claim she marks "answered the filer, waiting for a reply," and the claims-side annotator marks "still on my desk." The same "in progress," used by two departments for ten years (the nine semantic-divergence fields of Chapter 9). Nobody in the room has the authority to change the other side's definition, and the session stalls there.
When it stalls, do not put it to a vote, and do not split the difference. That blends two clear rules into one rule nobody owns. What you do is classify the disagreement on the spot, write it into the disagreement log, and route it one of three ways by type.
- **Within-department disagreement.** Find the owner of this workflow, rule on the spot, and write the ruling into the annotation guide.
- **Cross-department disagreement.** Pause annotation. Write both definitions and their business consequences on half a page, send it to the two sides' common superior, or put it on the sponsor weekly as an agenda item. You cannot rule this one for anyone, but you can make sure someone rules within two weeks.
- **Compliance-related disagreement.** Cross-department or not, it goes to the risk owner, at Anchor & Helm that is Victor Reyes. A compliance definition is not an efficiency question and cannot be settled by whoever is loudest.
All three routes want the same thing. A disagreement is not allowed to hang. A hanging disagreement becomes a golden case each side annotates its own way, and that one case strips the whole table of its standing as arbiter.
## Co-building the Eval Is Standalone Value You Can Give the Business Side
Back to what Kevin said at the door. The real output of that annotation session was not fifty cases. It was this company writing a twenty-year verbal rule down for the first time, and the review team owning it. That annotation guide stays alive without your system. New-hire training uses it, cross-review uses it, and it is not voided when the core system is replaced next year. It is an organizational asset, not a project artifact.
That is also the best reason to ask for those 2 hours a week. Once the rules are written down, someone has to use them day to day, and the people who will use them are in your company, so this is something you can get done. You are not here to borrow labor to test a system. You are here to help the review team turn what lives only in Linda's head into something everyone on the team has in hand. That sentence can be said out loud at their weekly. "Help us with an evaluation" cannot.
Three more things only someone inside can move. The ten historical incident cases, you can pull straight from the complaint files and retrospective notes, without waiting for a veteran to volunteer a "the one I got wrong" story. The second annotator can be borrowed from another department. Ask a senior reviewer from customer service to annotate the same batch, and cross-department semantic divergence surfaces at the annotation session instead of in the queue after launch. Once the annotation guide is written, push to promote it to the review team's formal operating procedure, filed under a department document number, not left in your project folder.
After promotion someone has to run it. The guide belongs to the review team. Linda is the owner who maintains it, review flow-back and monitoring-triggered retests (Chapter 18) land on her monthly, and the reminder hangs on the review team's existing monthly QA review, not on anyone remembering. Inside a company no invoice reminds anyone that the eval is due for an update, so the reminder has to be written into a meeting people already hold. The two numbers of failure mode 4 are reported by that owner at that meeting. When she cannot answer, the rot has already started.
## Failure Modes
**1. Build first, evaluate later.** The architecture is set, the pipeline runs, and a week before launch someone remembers "we should do an evaluation." The eval got filed as "testing," and in engineering instinct testing comes after implementation. But an AI system's eval is the requirement. Written after the implementation it is no longer a requirement, it is a justification. The team will, without noticing, tune the thresholds until the architecture already built just passes. The upgrade criterion goes missing with it, and Chapter 10's "let the evidence force the upgrade" hangs in the air. One discipline. The eval spec is finished before the architecture decision meeting. That is the operational meaning of "before the architecture."
**2. A single accuracy metric.** "Overall accuracy 92%" makes it into the report. Organizations naturally favor a single number. Easy to remember, easy to compare, fits on a slide. But averages were invented for scenarios with symmetric costs, and here one missed risk costs ten thousand times one empty suggestion. 92% can mean "3 unsafe out of 50" at the same time, and the better the number looks, the deeper it buries them. When the old ticket's "โฅ90%" was voided at the project approval change review, the objection was that it was an overall score. Whether the number was high or low did not matter.
**3. An eval set of nothing but typical cases.** Every golden case is smooth, and the system sails through the first round with a high score. Typical cases are cheap to collect. Within reach, answers undisputed, annotation effortless. Edge cases mean digging through aged claims, and historical incident cases mean someone has to admit "I got one wrong." Both are expensive. But accepting on typical cases is issuing a diploma after examining only the easiest part of the system. One test. How many cases in your eval set have someone disagreeing with the answer? None means the disagreement was avoided, and disagreement is the substance of the spec.
**4. One-shot eval.** After acceptance the golden cases are never updated again. The eval was treated as a project artifact (finished and filed) when it is an operating asset (depreciating continuously). The repair shop list drifts, fraud methods evolve. Chapter 9 said fitness is a state, not a property, and the same holds for the eval. A one-shot eval means drift cannot be detected. The system quietly degrades and your ruler keeps proving it has not. The way out. Cases flowing back from review are added to the golden cases monthly, and monitoring anomalies trigger a retest (Chapter 18 takes over this line). How much updating is enough, look at two numbers. The number of cases that flowed back into the golden cases last month is not zero, and the annotation guide's version number moved this quarter. If neither can be answered, the eval is one-shot, however seriously it was co-built at the start.
## Next Monday
1. Ask one question about the AI system on your desk. "Which kind of error can the user not accept even once?" Write it down. That is the first row of your error taxonomy, the unsafe class. If you cannot write it, you have not yet talked to the actual user about consequences.
2. Use the collection list in [Template 11.2](../appendices/template-11-eval-spec.md) to gather 20 golden cases to start. Typical cases under half, at least 3 historical incident cases, dug from the complaint log and retrospective notes, or ask a veteran for a "the one I got wrong" story. Answers clear the data check first.
3. Ask two actual users to annotate the same batch **independently** and count the disagreements. Record each one in 11.3, find someone with decision rights to rule, and write it up as a written rule. The disagreement rate is the number of missing pages in your requirements document.
4. Dig out your project's acceptance clause. If it is an overall score, draft a per-category replacement. Now, not the week before acceptance.
**Want an agent to get you started?** In the repo you set up following [Start Here](../index.md), paste this to your coding agent:
```text
In the repo/ directory of the the-last-mile repository, help me with the Chapter 11 Next Monday actions. First run python3 templates/eval-spec/run_eval.py,
then run python3 templates/eval-spec/threshold_report.py, and show me the replay report and the per-category threshold table on the built-in sample. Then copy
eval-spec-template.md into the working directory I name. The unsafe row is mine to write. If I cannot write it, remind me to go talk to the actual user about consequences first,
do not make it up for me. The 20 golden cases are mine to dig out by the collection list. You only build the table and count disagreements. Rulings on the two annotators' disagreements and the written rules are
written by someone with decision rights, not you. If any command errors, stop and show me the output.
```
---
## Chapter Kit
- **Judgment frameworks.** The five-part eval spec (task definition โ golden cases โ error taxonomy โ per-category thresholds and acceptance lines โ human review path); the three co-build rules (annotate with the actual user / reconcile the answers first / thresholds per category, no overall score)
- **Templates.** [Template 11](../appendices/template-11-eval-spec.md), Eval Spec, Golden Case Collection List, Annotation Disagreement Log
- **Key judgments**
- "An AI system's spec does not describe behavior. It describes the acceptable distribution of errors."
- "The golden cases are the requirements document, and the thresholds are the acceptance clauses."
- "Annotation disagreements are business rules that were never aligned, surfacing. They are not noise."
- "The first mile of an eval is reconciliation, not annotation."
---
# 12 ยท Trust Constraints: Turn Your Reviewers into Designers
!!! info "Companion Templates"
๐ [Chapter Template](../appendices/template-12-trust-matrix.md) ยท ๐ [Template Library](../appendices/template-library-index.md)
> **The Challenge.** Security, compliance and privacy reviews always block the project at the last minute. The reviewer's first sight of the design is the final gate, and saying "no" is his safest move.
>
> **What You Will Be Able to Do.** Run one constraint interview with the risk owner in week 2 of the opening, fill six classes of trust constraint into the trust constraint matrix, and let the design grow with the constraints attached. Test whether oversight is real or nominal with the three oversight questions. Run the security review as a confirmation meeting.
---
## Week 8, Forty Minutes into the Review
In the room sit Victor Reyes, two colleagues from compliance, and the head of the IT data team. Item five on the agenda is "review of data write-back and change permissions on source systems," usually the most tangled part of the whole review. Victor turns to that page and says, "Skip this one. The merged view is purely derived and purely read-only. It writes back to no source system, and I signed the source of truth decision log." Nobody objects. The meeting was set for two hours. That one sentence saved thirty minutes.
A few minutes later, one of the compliance colleagues asks the question. "When the AI's suggestion is wrong, whose responsibility is it?"
Before you can open your mouth, Victor answers. "Every suggestion stops at being a suggestion. Without confirmation from the named person in the 'Human Call' column, the system produces no action at all. Suggestions carry a trail the whole way, and if something goes wrong the full chain can be pulled up. We settled this design together in week 2."
Now think back to the parallel universe of Chapter 1. The you who went straight to a chatbot met the same risk owner, Victor, at the week 10 security review, and was stopped cold by the same three questions. Does customer data leave the boundary? Who audits the model's output? Who is responsible when it is wrong? The review was scheduled six weeks out, and the project died of comfort. In the real universe, you pulled Victor into the design meetings in week 2. Same man, same three questions, two endings, and the only difference is timing.
Week 8 goes smoothly because it stopped being a review long ago. It is a confirmation meeting. The real review happened inside every design decision from week 2 to week 7. This chapter is about how.
## Why This Is Hard: The Reviewer's Rational Move Is to Say No
An engineer's default model of review is an exam. Cram the material beforehand and bet the reviewer will not ask about the weak spot. That model is wrong because it never looks at where the reviewer sits.
Look at Victor's seat. He is the last person on the project to see the design. He is not in the design discussions, not there when the design takes shape, and the review meeting is the first time he sees the whole of it. Yet if there is a data incident, his is the first name on the accountability email (in the Chapter 5 stakeholder map, the risk owner cell and the blocker cell carry the same name). Least information, heaviest responsibility. Inside that structure of timing, saying "no" or "send me more material" is rational, and has nothing to do with a taste for caution. Releasing a system he did not help build and cannot see through, he carries all of the downside and shares none of the upside.
> **The reviewer sees the design last and is held responsible first. Saying no is his rational choice.**
So the root cause of review blocking a project is not the reviewer. It is when the information reaches him. Treat constraints as a launch gate and review is left with two values, pass or fail, and fail is always safer for him. Treat constraints as design inputs and every constraint becomes a design parameter, open to discussion about how to satisfy it, open to weighing its cost, and usable as currency to buy space. The earlier constraints come in, the larger the design space. The later they come in, the less is left but "pass or fail."
## Prior Art, and What AI Changed
Manufacturing paid this tuition fifty years ago. The core proposition of the quality movement (the Deming and Toyota traditions) is that quality is built in, not inspected in. A defect caught at inspection costs an order of magnitude more than one prevented at design, and the stricter the inspection, the slower the delivery, with no gain in quality. The fix is to move quality standards forward into design. More inspection cannot save a final check. The security review is software delivery's final check, and the same law applies unchanged.
Peter Block supplies the other half in *Flawless Consulting*. The person with the most resistance often has the most information. A blocker is a requirements owner nobody interviewed, Victor's three questions are a requirements spec written as questions, and the word "obstructive" does not attach to him. You already used the first half of this principle in Chapter 5, the pre-mortem hand-delivered to his desk. This chapter is the second half, taking his requirements formally into the design.
What did the AI era change? Reviewers have no vocabulary for reviewing AI. Traditional review has a mature checklist, permissions, encryption, logging, change management, and Victor can run through it with his eyes closed. But the constraints an AI system adds are on no checklist. Model output is not reproducible, so how do you define "the same test passed"? Where is the boundary of training data and context? For a system that extracts information from email, does maliciously constructed email content count as an injection attack (prompt injection)? How far does "why did the model suggest that" have to be explained before it counts as auditable? The reviewer is learning too. That brings two consequences.
The first is bad news. Where uncertainty is highest, review rulings are most conservative, and "if I cannot understand it, better not launch it" is killing a great many otherwise healthy projects. The second is a rare opening. The reviewer lacks the vocabulary, and you are the only person on the floor able to help him build a framework for reviewing AI. This is the spillover of AI uncertainty management, the one skill of its own (Chapter 2). In the language of Chapter 5's Trust Equation, with a risk owner the fastest way to open a position in credibility (to bank your first deposit of standing) is to hand him a framework he can use to judge safety himself. Proving your own system safe is the slower move. Teaching the reviewer to review you is the highest form of being trusted.
## The Framework: The Trust Constraint Matrix
The trust constraint matrix is one table that moves trust constraints forward into design inputs. Six classes of constraint as rows, three columns, filled from week 2 through to the review meeting.
Six classes of constraint, covering everything an enterprise means by "why should I trust you."
| Constraint | The Question It Answers |
|------|-----------|
| **Privacy** | Whose information, and what information, moves inside which boundaries? |
| **Security** | Who can access what, and change what? Where is the attack surface? |
| **Audit** | When something goes wrong, can you reconstruct who did what, when, and on what basis? |
| **Human oversight** | At which step is a person present, and in what way is that presence real? |
| **Fairness** | Will the system systematically treat one class of subject worse? |
| **Maintainability** | After you are reassigned, leave, or the system enters its third year with nobody watching, who can safely change it? |
Three columns, each with its own discipline.
- **The specific requirement for this project**, and it has to be specific to this project. "Complies with the company data security standard" does not count as filled in. "Customer-identifiable information may not leave the claims domain" does. What goes in this column comes from the constraint interview, not from your imagination.
- **How the design satisfies it**, a mechanism, not a promise. "We will be careful about de-identification" does not count. "Name, ID number and plate number are replaced with placeholders at the extraction layer, and no plaintext appears at any step after that" does. Write the cost at the end of every cell, who spends how much extra time under this way of satisfying it, and what capability is given up (the antidote to "over-compliance" in the failure modes below).
- **Who signs off**, by name. This column is the device that turns the reviewer into a designer. The day his name goes on, he stops being the judge and becomes a co-author, and what he defends at the review is a design he took part in.
What six rows by three columns looks like, two rows from the Anchor & Helm review version.
| Constraint | The Specific Requirement for This Project | How the Design Satisfies It | Who Signs Off |
|------|------------------|--------------|-----------|
| Audit | Every suggestion is fully traceable | The six end-to-end decision trail fields (the six are listed below under At Anchor & Helm); the merged view is purely derived, purely read-only, no write-back. Cost, log storage and the hours for the quarterly spot check | Victor Reyes |
| Human oversight | No action without human confirmation | Risk ranking with a daily volume cap (time); suggestions carry their reason (ability); the Human Call column names a person (authority). Cost, a cap on daily throughput, with the backlog going through the old manual process | Kevin Doyle |
The three steps of the review front-loading process that go with it.
1. **Week 2, the constraint interview.** Go with questions, not with a design (there is no design to bring, which is exactly why this is the best moment, everything can still be changed). Ask three things. What has gone wrong with systems like this before? What are you afraid of? What evidence do you want to see when it comes to the review? Note that these are the three questions you ask him, and they are not the three he will ask you at the review. The answers go into the matrix's first column.
2. **Carry the matrix through the whole design.** Run every major design decision past the second column. Which row does this choice make easier to satisfy, and which harder? The matrix is a living document, and blank rows are legitimate. Mark a row you cannot fill as an open question. It is the blank row you hid that blows up at the review.
3. **Run the review as a confirmation meeting.** Send the review packet ([Template 12](../appendices/template-12-trust-matrix.md)) a week ahead, organized around the answers the reviewer wants. The packet (the bundle of review materials) is the matrix's snapshot at review time, assembled from five things, the de-identification boundary on the data flow diagram, the permission inheritance table (which records how queue permissions inherit core system roles, and has nothing to do with Chapter 2's inheritance matrix), the audit log schema, the rollback path, and the open questions list. Confirm the matrix row by row at the meeting, and rule only on open questions.
Of the six rows the easiest to fake is human oversight. "We have set up a human review step" will always tick the box in a review document. The test against that formality is the three oversight questions.
> **Is there time to look? The ability to judge? The authority to stop it?**
Does the review volume match the human reviewer's time budget? Does he have the information he needs to judge right from wrong at hand? Does his "no" count? Miss one of the three and oversight is nominal. Nominal oversight is more dangerous than none. It manufactures the illusion that somebody is watching the gate, and every other defense slackens with it.
## At Anchor & Helm: Three Questions Become Parameters, the Judge Becomes an Author
**Week 2, the constraint interview.** Item four of the pre-mortem drew Victor in (Chapter 1), and in week 1 you answered his 11-question questionnaire line by line (Chapter 2). So you and Victor had already been round a full loop before week 2 started. In week 2 you invite him back, 60 minutes, a whiteboard, no design. What he asks in person is still those three questions, word for word. Does customer data leave the boundary? Who audits the model's output? Who is responsible when it is wrong? These are the three he will ask you at the review, not the three you asked him in the interview, and they are a different thing again from the three oversight questions above.
But the interview dug out what sits underneath the three questions. Two years ago a small tool one department built for itself carried a detail table containing customer information out of the company, and the accountability landed on him. That department was one of our own too, the tool was called an internal tool too, and nobody along the way thought it needed stopping. What he fears is "being bypassed, and still taking the fall when something goes wrong." AI itself comes second. That is the sentence written in his cell of the Chapter 5 stakeholder map, and it was not written wrong.
You were at the company when that happened. This is one of the openings being internal gives you. Most of the answer to the interview's first question is already in your hands. So do not go in empty-handed asking "what has gone wrong." Go in with a hypothesis. Say the incident you remember first, let him correct the details, then ask one line, "which one is there that I do not know about." The ones he adds are the information you were actually missing. That line also tells him you are not here to run a process, you remember what has gone wrong at this company.
Three questions, translated on the spot into three design parameters.
1. **Does data leave the boundary โ customer-identifiable information is de-identified at the extraction layer.** Name, ID number, plate number and phone are replaced with placeholders before they enter the queue or any model call. The de-identification boundary is a solid line on the data flow diagram. There is plaintext to the left of the line and none to the right.
2. **Who audits the output โ every suggestion carries a trail the whole way.** Input snapshot, rule and model version, suggestion, reason, Human Call, timestamp, not one missing. AI output is not written back to the source and exists in trail form (Chapter 9's discipline, which turns into an audit commitment here).
3. **Who is responsible when it is wrong โ queue permissions inherit core system roles, and the Human Call carries a real name.** No separate account system is built. Whoever may handle a given claim in the core system is the only one who may pick it up in the queue. The final action on every suggestion lands on a person with a name.
There is a fourth question beyond the three. For a system that extracts information from email, does maliciously constructed email count as an injection attack? The reviewer cannot answer that one either. The answer did not stop at a question mark. Over the following weeks it grew inside the design into a fourth parameter.
Email content is treated as untrusted input, always. The extraction layer recognizes only a schema whitelist, anything outside the four fields is rejected (Chapter 10), and the extraction result contains no free instruction text. An email that says "please handle this claim as routine" can at most contaminate a few field values. It cannot give the system orders. Add the system stopping at the advise layer, plus no action without the Human Call column, and the blast radius of an injection is held by design inside "one suggestion waiting for human review."
When the interview ended, four of the matrix's six rows were filled. Fairness and maintainability were left blank, marked as open questions. In week 3 Victor joined the audit log design review as agreed (the line planted in Chapter 2 honored), and his name appeared in the third column for the first time.
**Weeks 2 to 7, the matrix travels with the design.** When Chapter 9 settled the merged view as "purely derived, purely read-only," Victor's requirement was already written in the matrix's audit row. "Derived data does not feed the source" is data discipline on your side and an audit boundary on his, one design satisfying two ledgers at once. When Chapter 10 rejected the fully automatic multi-agent design, the human oversight row was one of the criteria.
The two blank rows also landed during these weeks. The error taxonomy of Chapter 11's eval filled in the fairness row. The priority logic uses no customer identity attribute, and the override (a human overturning the AI's suggestion) distribution is spot-checked by customer segment each quarter. What counts as a finished fairness row is being able to say what data would reveal unfair treatment. Writing "no discrimination" does not count. The maintainability row waited for the design to take shape. Rules and the de-identification config become configuration items the claims-ops IT engineers can change (they take over in Chapter 15). What counts as a finished row here is letting the person taking over try changing a configuration item once. It counts when the change goes through.
This row carries one more discipline specific to being internal. The signer may not come from this project's delivery team. The device in the third column is turning the reviewer into a co-author, but if the future maintainer is you, this cell is you signing for yourself, the device fails on this row, and what you signed is not a constraint but an ops commitment with no end date. So the maintainability row is either signed by the business line's IT or ops, or left blank and handled as an open question. Blank is more honest than self-signed, because a blank gets asked about at the review and a self-signature does not.
The three oversight questions landing at Anchor & Helm is the human oversight row of the matrix above, unpacked. The queue is ranked by risk with a daily volume cap, and the number of rows entering human view each day is worked back from the review time budget, better to leave some out of the ranking than to have them glanced at and not read. That is time. Every suggestion carries its reason, so the information for judging right or wrong is in front of her and she does not have to check three systems. That is ability. In Linda's team's "Human Call" column, does her "no" count? It does, and without that column the system produces no action. That is authority. These three designs land as actual columns of the queue in Chapter 17.
**Week 8, the confirmation meeting.** The packet goes to Victor and compliance a week ahead. One red line on the data flow diagram for the de-identification boundary, one page for the permission inheritance table, one page for the audit log schema, half a page for the rollback path, two open questions. That produces the scene this chapter opened with. The whole write-back review is skipped, and Chapter 9's "purely read-only" decision is cashed in on the spot for thirty minutes. Compliance asks where responsibility lands and Victor answers for you. Notice the language he answers in. Every sentence is from the matrix's second column. He is answering for a design he signed, and there is nothing of covering for you about it.
The cover of the packet carries two lines in the byline block, solution design by you, constraint design by Victor. That is not a layout detail. It is a ritual you have to build on purpose, and inside a company it has ready-made vehicles. The review conclusion memo goes out from you and Victor jointly, and recipients see two senders. The design document's author line carries two names side by side. At the meeting he presents the constraints section, not you.
Take care not to downgrade this into an approver's signature. An approver's signature is a process action, and after signing he is still the judge. Joint authorship is a change of identity, and after signing he defends his own design.
From blocker to co-signer, this line took eight weeks, and every step has a record. The pre-mortem hand-delivered, the questionnaire answered line by line, the constraint interview, the audit log reviewed together, the matrix signed. In the Trust Equation this is a run of consecutive deposits, and not one of them was luck.
Record the cost honestly. Pulling Victor into the design was not free. The de-identification layer, the audit log and permission inheritance came to roughly one extra week of engineering. Worth it? The cost of the other road was posted back in Chapter 1, a six-week review queue plus one rejection at the final gate. Constraints that come early have costs that are visible and plannable. Constraints that come late have costs that are hidden and fatal. This packet's life does not stop at eight weeks either. Chapter 18's launch gate updates and reuses it.
That extra week is harder to defend inside a company, because it has no budget line of its own. "Launch first, security later" wins at every scheduling meeting. The way to defend it is not to argue technical necessity. It is to change whose request it is. The day the de-identification layer, the audit log and permission inheritance were written into the matrix's second column, they stopped being your technical preference and became specific requirements written down by the risk owner. Whoever wants to cut that week is cutting Victor's requirements, and needs his nod. His requirement is the budget justification for that week, and that is the second use of putting real names in the third column.
Beyond the review it was built for and the reuse in Chapter 18, this packet has a third use, in a place you would not expect. The annual internal audit, a regulatory inspection, a Group compliance spot check, all ask the same set of questions. Who can access it, how do you reconstruct what happened, at which step is a person really present. At that point you are still at this company, and you are the one called in to answer. So the constraint matrix is not a project artifact for you. It is a permanent ledger, and it gets updated for as long as the system lives. After each spot check, add the newly asked questions into the first column, and answer one fewer next time. This also explains why the maintainability row cannot be left blank overnight. A spot check does not look at how beautifully you delivered. It looks at who is responsible now.
Look one layer further out. You and Victor are not a one-off project relationship. You are long-term colleagues, and there can be an account between you. This time's six-row matrix, the way the de-identification boundary was drawn, the six trail fields, the permission inheritance principle, all of it most likely holds unchanged for the next AI system. Save this packet as an "AI system constraint baseline" the two of you share, and the second project's constraint interview starts from "which rows are different this time" instead of from "what has gone wrong before." A formal 60-minute interview becomes one confirmation in a corridor. The baseline is his asset too. He does not have to reinvent his criteria for every AI system. This is compounding that only exists inside a company, and the real cost of the first packet has to be amortized across every project that follows.
## Failure Modes
**1. Compliance as the last gate.** On the schedule, "security review" is the last box before launch, and the reviewer's first sight of the design is the final gate. Project templates naturally draw review at the end of the flow, and engineers treat review as a cost center to be put off as long as possible. So the reviewer gets the least information, latest, saying no is his most rational move, and both sides push the project toward death from inside their own rationality. One test. If the reviewer's name first appears on the schedule rather than on a design meeting's sign-in sheet, you are already in this failure mode.
**2. The "our own tool" escape.** To dodge review, the system is positioned as an "internal helper script" or a "pilot tool" and goes live quietly, outside the formal process. The pain of review is immediate and concrete, the risk of being traced is delayed and probabilistic, the standard time-discounting trap. By the time the tracing arrives the system has real users, and the political cost of taking it down is enormous. Worse, you have detonated with your own hands the thing the risk owner fears most, being bypassed. One escape and this cell turns red permanently, and you still have to work with him for years, where permanent means what it says. Of the four in this chapter, this is the one an internal team is most likely to commit, because you are inside the firewall already, you have the permissions already, what you build is called an internal system already, and no door physically stops you. Other people are happy to wave you through, a tool built by one of our own can just start getting used. One test. If you are in the middle of explaining to yourself why this system does not need to go through review, you are already inside it.
**3. Nominal human oversight.** The review step exists, the human reviewer gets 500 rows a day, three seconds each. The "time" of the three questions fails. Review checks only the boolean "is there human review," and nobody does the arithmetic of daily volume times seconds per row. Both the system side and the review side have an incentive to see that box ticked, the former to pass, the latter to be off the hook. On the day something goes wrong the human reviewer takes the whole fall, the front line refuses from then on to sign anything in a "Human Call" column, and the trust gap collapses. The test is simple. Write the daily volume and the review time budget on the same line and divide.
**4. Over-compliance.** All six constraints turned to maximum, double review on everything, even the amount ranges the claims reviewers need for judgment de-identified away, three levels of approval on every action. The system is impeccably safe, and safely unusable. The pilot data is dismal and the project dies of "no value" rather than "risk." The root is that whoever proposes a constraint does not bear the cost of using it. One more requirement costs the reviewer nothing, and the cost lands entirely on the front line's operating time and the system's usability. A matrix that does not write costs down lets constraints accumulate in one direction only. The fix is in the second column's discipline, write the cost of every constraint's satisfying mechanism, and let it take part in the trade-off like any other design parameter. Remember, over-compliance and no compliance die different deaths on the same date.
## Next Monday
1. Pull out your stakeholder map (Chapter 5), find the name in the risk owner cell, and book one 60-minute constraint interview. Bring questions, not a design. What has gone wrong before? What are you afraid of? What evidence do you want to see at the review?
2. Fill in a first version of the trust constraint matrix with [Template 12](../appendices/template-12-trust-matrix.md). Mark rows you cannot fill as open questions, keep them on the table, and do not delete them.
3. Run the three questions on your system's human oversight step. If any one of them has no answer, change the design this week. Cap the volume, add the reason, or give a real authority to stop it.
4. Look at the schedule. If the security review is less than two weeks from launch and the reviewer has not seen the design, put the constraint interview into this week's calendar today.
**Want an agent to get you started?** In the repo you set up following [Start Here](../index.md), paste this to your coding agent:
```text
In the repo/ directory of the the-last-mile repository, help me with the Chapter 12 Next Monday actions. First run python3 templates/trust-constraints/redact.py
on the built-in sample to demonstrate field-level de-identification, then open audit-log-schema.sql and show me what the audit log looks like. Those two are the mechanism samples
for the matrix's "privacy" and "audit" rows. Then build an empty trust constraint matrix from Template 12. The answers from the risk owner constraint interview are mine to fill in
row by row when I get back, and rows I cannot fill get marked as open questions and stay on the table, not deleted. The three oversight questions are mine to answer, you only record them.
If any command errors, stop and show me the output.
```
---
## Chapter Kit
- **Judgment frameworks.** The trust constraint matrix (six rows, privacy / security / audit / human oversight / fairness / maintainability, by three columns, the specific requirement / how the design satisfies it, cost included / who signs off) + the review front-loading process (week 2 constraint interview โ the design carries the matrix โ the review becomes a confirmation meeting); the three oversight questions
- **Templates.** [Template 12](../appendices/template-12-trust-matrix.md), Trust Constraint Matrix and Security Review Packet, organized around the answers the reviewer wants, sent a week ahead
- **Key judgments**
- "The reviewer sees the design last and is held responsible first. Saying no is his rational choice."
- "The earlier constraints come in, the larger the design space. The later they come in, the less is left but 'pass or fail.'"
- "The three oversight questions. Is there time to look? The ability to judge? The authority to stop it?"
- "Teaching the reviewer to review you is the highest form of being trusted."
---
# 13 ยท The Trade-off Story: Win the Decision on One Page
!!! info "Companion Templates"
๐ [Chapter Template](../appendices/template-19-memo-suite.md) (ADR Memo, 19.2, borrowed from the Chapter 19 memo set) ยท ๐ [Template Library](../appendices/template-library-index.md)
> **The Challenge.** The plan sounds airtight to an engineer, and the executive hears it out and commits to nothing. Making it clear is not enough. You need the decision.
>
> **What You Will Be Able to Do.** Use the one-page ADR memo structure to write a technical decision as text an executive can responsibly decide on in 10 minutes, including presenting "this system will make mistakes" honestly without setting off panic.
---
## Forty Pages Against Ten Minutes
Friday evening of week 8. The security review confirmation meeting has just broken up, and Victor Reyes has co-signed the security review packet (Chapter 12). Eight weeks of work is all in hand. The reconciliation report, the eval spec, the constraint matrix, the architecture plan, stacked up, exactly 40 pages.
Now you need Grant Whitmore to approve three things. The development effort for the merged data view (Chapter 9's decision has to turn into code), confirmation that Linda's team's 2 hours a week keeps running through the pilot period (an existing charter clause, through one month after launch), and the pilot's scope and start. His assistant has given you 10 minutes at next Monday's weekly.
40 pages against 10 minutes, and most engineers solve that division by compressing. Pick the highlights, talk faster, build a 20-page deck. Compression is the wrong solution. The problem is not volume, it is order. The 40 pages are written in the order you reached your conclusion, current state, reconciliation, pattern selection, review, conclusion. Grant reads in the order he makes a decision, conclusion, cost, risk, what he has to do. The two orders are naturally opposite. This chapter teaches you to write one more page, a page whose order is completely reversed, and the 40 pages need not get any thinner.
## Why This Is Hard: Executives Are Not Short of Information, They Are Short of a Structure They Can Decide On
An engineer's default narrative is a chain of reasoning. Because A, therefore B, therefore C, conclusion at the end. That is professional instinct, not a bad habit. In the engineering world a conclusion's legitimacy comes from the derivation, and skipping a step is cheating. But an executive's day is dozens of decisions in a queue, and what he allots each one is a front-loaded stretch of attention that can be interrupted at any moment. Put the conclusion on page 38 and you have spent his most expensive first four minutes on setup.
The other half of the reason is in the decision itself. To call it, an executive has to be able to answer for it, and "the plan is correct" is only the starting point. When something goes wrong he has to answer "what did we know then, what did we give up, what was the backstop." So a text he can decide on has to align four things inside 10 minutes, the conclusion, the cost, the risk, and what he has to do. Miss one of the four and his rational choice is to commit to nothing. He is not rejecting you. He just cannot answer for it yet.
Get the narrative structure wrong and even a correct plan wins no decision. The common shape of an enterprise AI project "lost to politics" is a successful demo and a failed narrative. The system runs, the investment never arrives, and it slowly starves.
## Prior Art: The Pyramid and the One-Pager
Barbara Minto's *The Pyramid Principle* is where this chapter's skeleton comes from, and four of its rules are usable as they stand. The first two govern the opening. Answer first. The apex is the judgment you want to convey, so lead with the conclusion and give the pillars after. Open with SCQA, the shortest road from "what we agree on" to "what you have to decide." The four letters are Situation (a fact both sides agree on), Complication (the change that threatens that agreement), Question (the question that rises in the reader's mind as a result), and Answer (your conclusion).
The last two govern the body of the pyramid. MECE pillars, reasons that do not overlap and together exhaust the ground. And the one most easily forgotten, the whole structure is driven by the reader's question. At each level write what he will ask when he reaches it, not what you want to say.
The one-page rule comes from Amazon's narrative memo tradition (recorded in *Working Backwards*). Slides are banned for important decisions, and a memo written in full sentences is read silently at the start of the meeting. The reasoning has two sides. Full sentences cannot hide the parts you have not thought through. "One page" forces the author to finish the trade-off himself instead of passing its cost to the reader.
The AI era changed two things. First, the cost of writing collapsed. AI generates a well-formatted SCQA draft in seconds, and "no time to write one page" stops being an excuse. But the apex of the pyramid, what you want him to decide, AI cannot think out for you. The inputs it would need (this reader's question at this moment, the organization's power structure, the judgment you are willing to stake) are in no document. AI can lay the body of the pyramid fast, and the apex is still yours to place. Second, the trade-off story for an AI project carries a new difficulty. You have to present "this system will make mistakes" honestly to a non-technical executive without setting off panic. Say it in Chapter 11's vocabulary. Do not promise "no mistakes." Talk instead about error categories, per-category thresholds, and the human review backstop path. Panic comes from "nobody handles it when it goes wrong," and "it will make mistakes" frightens nobody on its own. Make the backstop clear and honesty turns into a sedative.
## The Core Framework: The One-Page ADR Memo Structure
> **The architecture decision memo (ADR memo). One page presenting one pending decision to the person with the authority to call it. Conclusion, evidence, the cost said out loud, risk and backstop, and what he has to do. Its purpose is to let him call it responsibly, and understanding the plan is only something picked up along the way.**
| Section | What Goes In It | Rules |
|----|--------|------|
| **SCQA opening** | S, the agreed fact / C, the threat or change / Q, the reader's question | Three lines and done; Q must be his question, not yours |
| **The conclusion in one sentence** | The decision to be called, in one sentence | The pyramid apex test. If you cannot say it in one sentence, do not start writing |
| **Three pillars** | Three reasons that hold up the conclusion, each with one line of evidence | MECE (no overlap, exhaustive together); evidence is numbers and signatures, not adjectives |
| **The trade-off said out loud** | What was given up, what it bought, and where the thing given up now sits | The soul of the memo, unpacked below |
| **Risk and backstop** | What errors will occur, how they get found, who backstops them, when to call a stop | Error categories + per-category thresholds + review path, not an apology |
| **The three things I need from you** | Concrete actions, each with a date | Without this section = asking for understanding, not asking for a decision |
ADR is taken from architecture decision record, which is the architecture decision memo in the sentence above. The S / C / Q in the table are structure names. You do not have to copy them into the text, and in the letter below they land as three lines, "background," "but," and "so the question to answer." The conclusion line works the same way. One sentence can be followed by an arithmetic note in parentheses, the arithmetic is a footnote, and the conclusion is still one sentence.
The trade-off section is the soul of the memo, because it is the only section that works against you, and therefore the only one that proves the rest of them credible. An executive's professional nose is trained on cost. Leave it out and he will guess, and a guessed cost is always worse than the real one. There is one writing rule and no exemption from it. The cost has to be written by you, and it cannot be left for the reader to discover. In the Trust Equation (Chapter 5) credibility is a multiplying term in the numerator, and the reader discovering it for you once takes it to zero. Multiplication means that once credibility is punctured, every other score goes to zero with it, not to a discount. After that every "conclusion" of yours gets audited layer by layer, the one-page privilege is withdrawn, and from then on you are only fit to write 40 pages. Once a hidden cost is exposed, you permanently lose the right to lead with the conclusion.
Where does the material come from? You do not have to invent it now.
- The C in SCQA, the Field MVP's scoring data (Chapter 0).
- Pillar evidence, the reconciliation report and the eval baseline (Chapters 9 and 11).
- The trade-off section, a transcription of the scope decision log (Chapter 8).
- The risk section, the error taxonomy plus the human oversight design (Chapters 11 and 12).
The one page needs no new writing. Take the hardest single line out of each of eight weeks of work and that is enough.
## At Anchor & Helm: That One Page
Sent on Friday evening of week 8. Here it is in full.
---
> **To Grant Whitmore, from [you], Friday of week 8**
> **Subject. The exceptions queue pilot decision, three things for your call on Monday**
> *One page of body text. The 40-page technical plan (reconciliation report / eval spec / security review packet) is attached for reference.*
>
> **Background.** The auto exceptions queue project is eight weeks in. Data reconciliation is complete, the eval baseline was co-built with Linda's team, and the security review passed this week (co-signed by Victor Reyes).
> **But**, the earliest MVP scoring gave a warning. Of 10 suggestions, 3 were unsafe, and all three traced back to "chase priority โ risk priority." The reconciliation showed more, 44% of the core system's status field cannot be taken at face value to drive an action. Scaling up the investment without controlling scope means mass-producing the wrong priorities.
> **So the question to answer**, do we start the pilot now, and at what scope.
>
> **Conclusion. I recommend approving an 8-week controlled pilot, scoped to auto exception claims and the eight people on Linda's review team, to validate the North Star "first-touch handling time, down 30%."** (The arithmetic behind -30% is in the charter. Across ten cases, about six tenths of the waiting time goes to "waiting on documents with nobody chasing and nobody claiming them," and the queue eats half of that. That is where -30% comes from, and it pulls a payment cycle running 1.6 times the industry average back near the average.)
>
> **Three pillars**
> 1. **Users have validated it**, with the front line scoring 5/10 pass. The 3 unsafe marks have become a zero-tolerance error category in the eval, and 50 golden cases were co-built with Linda's team.
> 2. **Data has been reconciled**, 200 records three ways, 44% of statuses distorted (the largest single class, the lagging pattern, about three in ten). The fix is the merged view. Claim status takes the merged view as its source of truth (the review team's Excel is the authority, the core system the fallback), amounts take the core system as source of truth, purely derived and read-only (it needs development effort, see item 2 below).
> 3. **Risk has been defended**, with every suggestion carrying its reason, a Human Call column, an end-to-end decision trail, and the reviewer's co-signature.
>
> **What we gave up.** To buy a pilot that can be validated inside 8 weeks, this phase explicitly gives up two things.
> - **The home property extension**, on the revival list, with the condition "the first extension after the pilot North Star hits target" (exclusive). It is first in line, and only time is missing.
> - **Full automation.** On the decision rights ladder, AI stops at the advise layer.
>
> The decision rights ladder, four layers from the top down.
>
> - **Decide**, pay or not, how much. This is a red line, never touched.
> - **Act**, send chase notices, assign, change status. Humans take over here.
> - **Advise**, priority, next step, reason. AI stops here (this phase).
> - **Sense**, gather documents, dwell time, risk signals. AI does this.
>
> Each layer up needs that layer's evidence. The advise layer's acceptance record starts accumulating with the pilot, and once it hits target we move up one layer at a time, each layer decided on its own.
>
> **Risk and backstop.** This system will make mistakes. We do not promise "no mistakes," we promise "when it is wrong, it can be found and it can be caught." Thresholds are set per error category. Missed risk is zero tolerance, and if a single one appears during the pilot, that whole category of suggestions is downgraded to human review, with suspected missed-risk cases reviewed weekly. The "empty suggestion" class (harmless but useless) is watched on trend only and must not rise. Every suggestion takes effect only after a reviewer confirms it, and overrides leave a trail. When the business side's promised input goes unmet two weeks running, the three tiers of resource reassessment are triggered (the three-tier handling already set in the charter, Chapter 4). When two consecutive evaluation cycles miss the threshold and no workable path to improvement exists, either side may propose termination (Chapter 4), and the call is yours.
>
> **The three things I need from you**
> 1. **At Monday's weekly (Monday of week 9)**, call the pilot's scope and start (auto exception claims, 8 weeks).
> 2. **By Wednesday**, approve the development effort for the merged view. Two people from the Digital Center and two from claims-ops IT, six weeks. The early part of the pilot runs on the existing read-only view, replaced as soon as the merged view is built.
> 3. **By Friday**, confirm with Kevin Doyle that Linda's team's 2 hours a week (an existing charter clause, through one month after launch) keeps running during the pilot, and that the schedule protection stays.
---
The six-section structure does not change. The center of gravity does. What an internal executive calls is rarely "do we approve this money," because the money already has a line in the annual budget. He calls whether you get people, where you rank, and who owns it after launch. So the last section of the internal version leads with people, priority and owner, not with budget. Rewrite the last section of that letter around the internal center of gravity and the three things grow into this.
1. **At Monday's weekly**, call the pilot's scope and start, and confirm that the owner of this queue after launch is Kevin Doyle.
2. **By Wednesday**, approve the people for the merged view build, two from the Digital Center and two from claims-ops IT, six weeks, and ask Owen Hartley to write those four people's schedule protection into this quarter's plan.
3. **By Friday**, ask Kevin Doyle to write Linda's team's 2 hours a week into the review team's schedule sheet and copy Grant Whitmore.
The extra half-sentence in item 1 is the watershed of the internal version. Who owns it after launch, meaning who runs this queue day to day, is not a risk owner like Victor Reyes. In a delivery with a deadline, ownership gets written on a different piece of paper, and you do not have that piece of paper. Leave it off this page and there will be no second occasion to write it, and a year later the alert at midnight rings on your phone by default (Chapter 22). The owner has to be a named person. Writing "Claims Operations" does not count.
Item 3 looks like the smallest, and it is the only execution guarantee on the whole page. Approval is not the same as happening. After Grant nods, those 2 hours still sit in the review team's own queue, and you hold no written commitment you can chase with. So inside a company every approved item has to be followed by a verifiable landing action, who, by which day, writes it into which sheet, and copies whom. The 2 hours written into the schedule sheet are the real 2 hours, and the 2 hours nodded at in a meeting are not. The copy line especially cannot be skipped. It turns "not done" from a private matter into something Grant can see too, and inside a company visibility is the only place enforcement starts.
**Monday morning, the 10 minutes.** You put the one page in front of Grant Whitmore, with the 40-page attachment beside it. For the first two minutes he reads silently and you say nothing. In the third minute you start walking the pillars, and partway through the second one, in minute 4, he raises a hand and stops you, his finger on the ladder in the trade-off section. "Home property first in line for Kevin, exclusive, I accept that. One question. Going from the advise layer up to the act layer, who decides, and on what?"
"On the acceptance record. During the pilot, every suggestion accepted as is or overridden leaves a trail. Whichever class of action gets its acceptance rate over the line, the data goes on your desk, each layer decided on its own, and the person who calls it is you."
He nods and drops the page on the table. "All three approved." Then he pats the stack of 40 pages beside it and pushes it back to you. "Keep that for Victor Reyes and Kevin Doyle's people. One more thing before we break up. Everything you send me from now on gets written the way this page is written."
A 10-minute meeting ran 7 minutes. Not one of the 40 pages was turned, and it held up every line of the one page. Because the reconciliation report exists, pillar two dares to be one line. Because the review packet exists, "risk has been defended" can stand. You could answer Grant's question in 5 seconds because the answer went into the scope decision log back in Chapter 8. The point of a 40-page plan is to let one page dare to be that short. The one page is the interest on the 40, not its summary. Meaning, without those 40 pages behind it, nobody would believe this one.
**Round two, after the meeting.** In the corridor, Grant's assistant catches up with you. The board meeting has moved up from the end of the month to next Thursday, and Grant wants the meeting to "open the system and walk through a few live claims." That is the sponsor's support, and it is also the sponsor's pressure, the kind that comes in the name of support.
In week 1 he told the board "there will be something to see in three months." Now the board has moved up, and the chip he has to pay with is the prototype in your hands, the one that changes every day and that Linda's team is still trying out informally. Demoing it live on real data means taking a thing built to change at any moment, holding it to a "does not make mistakes" standard, and staking it on the highest table in the company (Chapter 14 covers this, the prototype's game is learning speed, not performance reliability).
You do not say "no." Pushing a sponsor's request back spends trust, and taking it whole is betting your life on it. You give him something he can call, written the way this one page is written. Conclusion, the board demo becomes a three-minute screen recording, plus that scoring sheet and the decision rights ladder drawing. The cost, it does not hit as hard as clicking through live. What it buys, zero risk of the demo blowing up, and a harder story. The board sees the 3 unsafe marks a front-line reviewer made by hand and how each of them was blocked, and no demo on the market has that page.
The assistant comes back on Tuesday. Grant chose the recording, and added one line, "put a shot of Linda's handwritten scoring sheet in it." The pressure did not disappear. It was translated into a trade-off he could call responsibly.
This page often goes to more than one person. Grant can call it alone because both the people and the data sit under his line. On another project the business department supplies the people while the capacity being consumed is your team's, and the two signing points are peers. Your manager Owen Hartley asks "how many of our people for how many weeks." The business-side head asks "how many people do I put in, and will it wreck my schedule this quarter." The two questions have different answers, and writing two pages is the wrong solution, because when each of them holds a page of his own, both assume the other one is covering it.
One page stays one page. The first five sections are shared, since conclusion, cost and risk are the same thing for both readers, and only the last section splits. "What I need from you" runs in two columns, each signed, each dated. Splitting it into two columns has a side effect. Both of them can see the other's column, and whoever fails to deliver fails in front of the other one.
Three more things a team from outside does not have and you already hold. Turn all three into actions.
- **Borrow a shell.** You can pull up project approval memos this company has already approved. Write to their section names, their length and their copy line. The reader recognizes the shape, so all of his attention goes to the content.
- **Pre-read.** Before sending it, get Grant's assistant or your counterpart on the last project to read it for four minutes, then ask him what the decision is, what the cost is, and what happens when it goes wrong. Anything he cannot answer, rewrite that section. You can do this any time. A team from outside would have to spend another company's time to get it done.
- **The apex.** You know what this executive cared about most the last time he called a decision. What the board presses Grant on is the payment cycle, so the conclusion sentence lands on first-touch handling time, not on accuracy. AI can lay the body of the pyramid fast, the apex is yours to place, and you know better than an outsider where it belongs.
## Failure Modes
**1. Benefits only, and credit collapses the moment the trade-off is found.** The memo is all upside, and the cost shows up as a "supplementary note" only when someone presses. The cause is double. Fear of reporting bad news (worry that it scares the decision away), plus mistaking a memo for sales copy. Sales can talk only about benefits, because the buyer discounts by default. A decision maker does not discount, he answers for it at face value. And the cost cannot be hidden, because the professional nose of anyone who approves an investment is trained on cost. The test, the trade-off section has to contain at least one sentence that hurt to write. Not one, and you are still selling.
**2. Asking for understanding, not a decision.** The memo's logic is perfect, the executive reads it, nods, says "very clear," and then nothing happens. Engineers treat "explained it clearly" as the finish line of delivery, but understanding produces no behavior, only a dated request does. Underneath there is self-protection. A clear ask can be refused, and a vague report is always safe. But dodging the risk of refusal also dodges the chance of getting a commitment. A memo with no "what I need from you" changes nothing once it is read. The executive's nod is for you, not for the project.
**3. One page written as ten.** "It is all important, none of it can go," the body swells to ten pages, and it gets called thorough. Every detail is the author's sunk cost and every cut hurts, so the cost of the trade-off passes to the reader untouched, and the reader's solution is harsher than yours. He does not read it. One rule. The attachment can be any thickness, and a body over one page means you have not thought it through. Going over is a diagnostic signal. It says you have not found the pyramid's apex yet and still do not know what you want him to decide. Go back to the one-sentence test, and come back to write once you have it.
## Next Monday
1. Run the pyramid apex test on the project on your desk. Who you want, and what you want him to decide, written in one sentence. If you cannot write it, do not start writing. What you want may be understanding, not a decision.
2. Pick a plan of yours from recently that runs over five pages and rewrite it with the one-page structure of [Template 19](../appendices/template-19-memo-suite.md). The trade-off section must name at least two things you gave up and their revival conditions (Chapter 8's scope decision log is ready-made material). The same structure keeps coming back after the pilot starts, authorization at the open, a decision at the midpoint, renewed funding at the close, and Chapter 19 expands it into the three-memo system. Beyond those three there is a fourth kind, and it is not the same as renewed funding at the close. It handles the fact that a year after launch no bill reminds anyone that this system still needs investment, and by then one page is your only tool for winning resources back on a regular basis.
3. Dig out the last email you sent an executive and check whether it has "what I need from you + a date." If not, send a follow-up that is only that section, and watch what happens.
4. Run the four-minute test with a colleague who does not know the project. The number comes from Grant Whitmore, who raised a hand in minute 4. Give him only the one page, and after 4 minutes ask him "what is the decision, what is the cost, what happens when it goes wrong." Any question he cannot answer, rewrite that section.
**Want an agent to get you started?** In the repo you set up following [Start Here](../index.md), paste this to your coding agent:
```text
In the repo/ directory of the the-last-mile repository, help me with the Chapter 13 Next Monday actions. I will give you one sentence first, who I want and what I want him to decide.
If I cannot write it, stop and do not write it for me. Then I will paste you my recent plan that runs over five pages. Follow templates/memo-suite/prompt-memo-scaffold.md
and the one-page structure of Template 19.2 to rewrite it into a draft. The trade-off section must list at least two things given up and the revival condition for each.
Anything not in the original plan you mark [TBD] instead of inventing. The closing "what I need from you + a date" is mine to fill in. When you are done, report the draft's word count to me.
If it runs over one page, delete what you added first. If any command errors, stop and show me the output.
```
---
## Chapter Kit
- **Judgment frameworks.** The one-page ADR memo structure (three lines of SCQA โ the conclusion in one sentence โ three pillars with one line of evidence each โ the trade-off said out loud โ risk and backstop โ the three things I need from you, with dates); the pyramid apex test; how to put an error distribution to an executive (error categories + per-category thresholds + human review backstop); the four-minute test
- **Templates.** [Template 19](../appendices/template-19-memo-suite.md), Three-Memo Set (with ADR), the ADR part (19.2), memo template plus SCQA writing self-check list
- **Key judgments**
- "Executives are not short of information. They are short of a structure they can decide on."
- "The point of a 40-page plan is to let one page dare to be that short."
- "A memo with no 'what I need from you' changes nothing once it is read."
- "Once a hidden cost is exposed, you permanently lose the right to lead with the conclusion."
---
# 14 ยท Prototype, Pilot, Production: Three Different Games
!!! info "Companion Templates"
๐ [Chapter Template](../appendices/template-14-stage-gates.md) ยท ๐ [Template Library](../appendices/template-library-index.md)
> **The Challenge.** The prototype feedback is good, and yet "let us polish for two more weeks" and "let us watch it another month" both sound right too, and you cannot tell prudence from delay. You have also seen the other extreme, a pilot four months in and still "taking a look," with nobody saying it is dead and nobody daring to say it is alive.
>
> **What You Will Be Able to Do.** Use the stage rules table to see which game you are in and which set of rules you have to keep. Use the three escalation gates to rule in 30 minutes on whether the prototype moves up. Write the kill criteria before entering the pilot, naming what data, once it appears, means stop.
---
> **Part IV Navigation.** Chapters 14 to 18 are arranged by theme, and the weeks jump back and forth.
> The escalation decision meeting (Wednesday of week 9, this chapter) โ the two engineers dialed in and the co-build agreement (dialed in week 5, signed Thursday of week 9, Chapter 15) โ the last afternoon before the pilot starts (Friday of week 9, Chapter 17) โ the pilot starts (Monday of week 10) โ one bill (week 12, Chapter 16) โ an error that got caught (week 14, Chapter 18) โ the pilot closes out (week 17)
## Two Proposals That Both Sound Prudent
Monday of week 9, Grant Whitmore approved three things in 7 minutes (Chapter 13). That afternoon you book the escalation decision meeting for Wednesday, and then two messages arrive.
The engineer in claims-ops IT who leads the writing of the extraction pipeline. "The prototype still has edge cases it does not handle cleanly. Two more weeks of polish and it will be steadier."
Kevin Doyle stops you in the corridor. "Linda's team is getting good use out of it. Why rush into the pilot? How about another month of trying it?"
Here is the background. The queue prototype that grew out of the Field MVP was used informally by Linda's team for about two weeks across weeks 7 and 8, in parallel with the eval co-build, running on the existing read-only view, with data you refreshed by hand every day. The feedback really is good. And both proposals are well meant. Both sound prudent.
But you have watched too many projects die inside that prudence. Always two weeks short, always one more month of watching. Ten months later the prototype is still "in trial," the engineer is still polishing, nobody calls a stop and nobody calls it up a level, because nobody ever defined what the conditions for moving up actually are.
First, clear up one thing you will say out loud on Wednesday. What Grant approved on Monday was the project approval and the budget, whether it is worth investing in, how wide a scope, how many resources. What Wednesday's meeting rules on is a different question, whether the system and the organization are ready now. Budget approved is not the same as ready. The first reads the books, the second reads the evidence. Inside a company the easiest mistake is to run the two as one meeting and settle "let us start the pilot then" at the budget meeting, which is using the books in place of evidence. So the budget meeting stays the budget meeting and the readiness meeting stays the readiness meeting. The two neither repeat nor conflict. One decides whether to do it, the other decides when to change the rules.
## Why This Is Hard: Three Stages Are Three Games with Mutually Exclusive Rules
- **The prototype's game is learning speed.** The goal is to work out "what should be built" on the least investment. Changing fast beats not erring, ten rows of data beat ten thousand, and whether the interface is ugly does not matter.
- **The pilot's game is evidence quality.** The goal is to measure "does it work" under real conditions. Versions freeze into batches, or the data cannot be attributed. Scope stays controlled, or an error has no defense behind it. Users have to depend on it daily, or what you measure is still goodwill.
- **Production's game is operating reliability.** Not erring beats changing fast. Every change is a risk, and the rollback path outranks a new feature.
The three sets of rules are mutually exclusive, which means there is no "good habit that works across all three stages." The prototype's "change anytime" becomes contaminated evidence in the pilot. The pilot's "controlled scope" picks the easy users to serve, and in production that is dodging the hardest ones. The danger is not inside any one stage. It is in changing stage without changing the rules.
Timing the change of stage has a second trap, using time and gut feel in place of evidence. "We have been trialing it two weeks" only says the calendar turned a page. "The feedback is good" is politeness, not data. Underneath sit incentives. The engineer fears real data exposing weak spots, the owner (the person accountable for operating it in the next stage) fears carrying operating responsibility, the user fears having a way of working locked in. Everyone's "let us wait a bit" has a locally reasonable justification, and added up they make "watch it a while longer" permanent. So the escalation criteria have to be written down in plain words, outside gut feel and ahead of interests.
## Prior Art, and What AI Changed
**The stage-gate tradition (Robert Cooper).** From the 1980s on, new product development was cut into stages with "gates" between them. A gate has written criteria and a gatekeeper with the authority to close it. The core insight is that a project carries its own momentum, investment argues for itself, and with no written gate any project slides into the next stage on its own.
**The lean tradition (Eric Ries).** The discipline of pivot or persevere, keep-or-kill decisions made on a fixed cadence and on evidence, not on mood. "Let us watch a while longer" does not count as a decision. It is the absence of one.
The AI era changed two things. First, eval makes escalation criteria fully objective for the first time. In the stage-gate tradition most gate criteria are still judgment calls at a review meeting, market attractiveness, technical feasibility, scored by people in the end. An AI system has golden cases and per-category thresholds (Chapter 11). Over the line is over the line, and for the first time a gate can exist apart from the mood in the review room.
Second, an AI system's pilot carries a layer of duty traditional software never had. Model behavior only reveals its real error rate under the real data distribution. A traditional pilot mainly validates "will users use it." An AI pilot also has to validate "is the system still right across the full distribution." Fifty golden cases are a carefully proportioned sample. Eight weeks of real claim flow are the distribution itself. So the pilot is eval scaled up, not production scaled down. Clearing the eval line is only the ticket in. Every override during the pilot (a person changing the AI's suggestion) and every human correction is an error sample the golden cases could not catch (Chapter 18 hardens all of this into live monitoring).
## The Framework: One Table, Three Gates, One Trigger
**The stage rules table** (the full printable version is [Template 14.1](../appendices/template-14-stage-gates.md)).
| | Prototype | Pilot | Production |
|---|---|---|---|
| **Optimizing for** | Learning speed | Evidence quality | Operating reliability |
| **Data** | De-identified samples, read-only | Real data, controlled scope | Full real volume, trail closed loop |
| **Users** | A few volunteers, using it and cursing it to your face | A named user group, daily reliance | All target users, the most unwilling included |
| **Change discipline** | Change anytime, ship the same day | Batched releases, announced ahead, rollback available | Follow the process, clear the eval regression first |
| **Typical way to die** | Over-polished into a deluxe demo | The zombie pilot | Demo code shipped with demo discipline |
Mapped onto the outcome ladder (Chapter 1), the prototype sits between L0 and L1, the pilot is L1, and production is L2. That ladder's five rungs are L0 demo, L1 pilot, L2 production, L3 adopted, L4 self-sufficient. The ladder tells you which rung you are on. The rules table tells you which discipline that rung requires. Judge the stage by the "change discipline" row, not by what the system calls itself. A system that can be changed anytime is still a prototype, wherever it runs.
**The three escalation gates** (prototype โ pilot).
| Gate | Criterion | Evidence That Does Not Count |
|---|---|---|
| 1. Eval over the line | Golden cases clear the threshold in every category (the Chapter 11 table as is) | "It feels a lot more accurate overall" |
| 2. Willingness to use it daily | Two consecutive weeks of unprompted daily use, a behavioral signal that taking it away would hurt | A high score on a satisfaction survey |
| 3. Owner in place | The person accountable for operating it in the next stage claims it by name, on the spot | "We will assign someone when the time comes" |
The three each test one thing. Eval tests the system, willingness tests workflow embedding, owner tests responsibility. All three are required, and there is no "basically passed."
**Pilot kill criteria**, the stop standard written hard before entering the pilot. What data, once it appears, means stop, with no "assess as the situation warrants." It has to be written in the same meeting as the escalation resolution and signed the same day. Kill criteria are written during the excited period. They cannot be written during the disappointed one (failure mode 4 does that math).
The three gates are the door into the pilot. There is another door out of it, the graduation criteria this chapter uses over and over from here, pilot โ production. It looks at five things. The charter's North Star hits target on real data across the whole pilot, the acceptance rate for suggestions holds up without reminders propping it, the unsafe zero-tolerance class was never broken end to end, kill criteria had zero triggers or triggered and were closed out, plus the production operating owner and budget ownership are settled (the end of [Template 14.2](../appendices/template-14-stage-gates.md) writes these into a checklist). Past that door the system moves from L1 into L2, and the acceptance rate line already measures the same thing as L3 adopted (users' daily actions really have changed), while gate two's "taking it away would hurt" is that thing's first reading on the pilot's scale.
## At Anchor & Helm: Wednesday of Week 9, the Escalation Decision Meeting
In the room sit you, Kevin Doyle, Linda Marsh and four engineers, two from the Digital Center and two from claims-ops IT. You put the three-gate table on the screen and go through it one by one.
**Gate one, eval over the line.** The prototype's most recent replay on the 50 golden cases. Missed risk 0, line-crossing suggestion 0, unsafe zero tolerance, over the line. Material error and wrong action both sit inside the 10% threshold, over the line. Empty suggestion is inside 20% and the trend has not risen across three consecutive replays, over the line. All three of unsafe, concern and useless clear, with no exceptions. What you used is the threshold table handed to Kevin on Friday of week 8. No new number system. The escalation criteria are the acceptance criteria rehearsed early.
**Gate two, willingness to use it daily.** You sent no survey, because everyone scores a survey high. What you produce is behavior. Linda's team opened the queue on their own every morning for two straight weeks, with no reminders. The harder evidence came from Monday morning of week 9. You were busy preparing Grant's 10-minute meeting, the view refreshed forty minutes late, and by 8:40 a member of her team was at your desk asking why today's queue had not come yet. Linda added a line at the meeting. "Kevin, honestly, the team is not used to a Monday morning without this queue any more." That is the first thread of Chapter 1's L3 adopted, taking it away would hurt.
**Gate three, owner in place.** A pilot is "the business side's operation running inside a controlled scope." "The Digital Center's system being tried out over at the business side" does not reach that bar, so the person accountable for operating it in the next stage has to claim it by name. Kevin claims pilot owner on the spot. What he claims goes into the minutes, three things. Two hours a week of schedule protection for Linda's team. The status write-back discipline, the review team clearing status in the last ten minutes before leaving each day, and he reviews the "status in doubt" list the queue produces every week. On every "stop or not" ruling during the pilot, he is the first signatory. The item the memo asked Grant to confirm by Friday (two hours a week of schedule protection for Linda's team), he claimed himself on Wednesday.
Claiming pilot owner internally means claiming a responsibility that does not enter his KPIs, so claiming it needs something in return, given on the spot. Three things. These three go into his goals for the quarter, the sponsor is present as witness, and claiming it brings priority scheduling rights. In this session at Anchor & Helm, two of the three are already in the minutes. The schedule protection is the priority scheduling right, and the line in the memo asking Grant to confirm by Friday is the witness slot. Writing it into this quarter's goals is the one still undone. Put the return on the table first, and the claim stops being a verbal favor.
All three gates clear. Now the two proposals. You ask that engineer, "Which gate is the polishing for?" He thinks it over against the table. Edge cases fall in the concern class, already inside the threshold. More polishing would push the number lower, but no gate asks for lower. You say, "All three gates are clear. Polishing is delay. I know you were not trying to stall. The word 'steadier' can always be said. Kevin's 'another month of trying it' is the same. What new evidence, for which gate, would another month produce? None. It would only produce more 'the feedback is good,' and 'the feedback is good' we already have." Kevin laughs. "Fine. We start Monday."
The last fifteen minutes, write the kill criteria. You say, "Right now every number looks good and everyone wants to run forward. Precisely because of that, this is the only moment to write the stop standard. Wait until the data looks bad and everyone can work backward from the standard to the conclusion they want. The discussion turns into a negotiation."
The first row goes down on the spot. Two or more unsafe-class errors in a single week, and the pilot pauses for rework. That whole class of suggestion is degraded to human review, the retrospective locates the root cause, and after the fix clears a golden case replay the owner signs the restart. This line is a line at the advise layer. It counts once only when a person catches the suggestion. The day the system moves up to the act layer (that cell in week 29, Chapter 22), this row gets rewritten to one occurrence and it stops, because by then the error has already landed.
A complete set of kill criteria is more than that one row. It covers at least four kinds of signal, one row each, all with hard numbers.
- North Star. The weekly median of first-touch handling time above the eight-week pre-pilot baseline for two consecutive weeks, pause the expansion, check the data source before the system. The data comes from the weekly points on Chapter 18's run chart (a metric plotted week by week as a line, read for trend and not for single points).
- unsafe. The row above. Trigger means the pilot pauses for rework, root cause retrospective.
- Acceptance. The target user group's weekly usage rate below the agreed line for two consecutive weeks, pause the expansion, go back to users to locate the reason.
- Cost. The system cost per claim above half the value of the labor hours that claim saves, for two consecutive weeks, pause the expansion, cost review. Chapter 16's four-fold bill (that chapter goes into it) did not hit this row. Cost per claim was still under the value line. What it hit was the budget line of the four ledgers (also Chapter 16). The trigger governs "is this still worth running," the ledgers govern "is the money being spent right." Two pages, one job each.
Kevin asks how this relates to the charter's resource reassessment conditions. You answer that resource reassessment conditions are project-level and govern "should the resources go back to the company." When the business side's committed input goes two consecutive weeks unmet, it runs in three tiers, escalate to the sponsor for a resource reassessment, the project turns to "awaiting inputs" on the PMO register, release your team's people and record it publicly (defined in Chapter 4). Inside a company it is one-directional. Only the sponsor or the project approval committee can shut a project down, and what your team can do is write the missing inputs into the ledger and pull your people back. Kill criteria are a pilot-level operating trigger and govern "should we pause this week." The trigger is an order of magnitude more sensitive than a resource reassessment, and an order of magnitude gentler. Triggering it means the defense held, not that it failed.
This page was never triggered once across the pilot's eight weeks. It was not written for nothing. In an incident retrospective after launch it gets pulled out and checked line by line (Chapter 18). That day it is still not triggered, and the fact itself, that the stop standard was written long ago and has not been triggered, becomes the hardest sentence in that storm over trust.
## The Resource Gates, How the Resource Cadence Meshes with the Three Gates
Up to here the chapter has deliberately not touched the question sitting on the other table, resources. That is what Grant's 7 minutes on Monday approved. The three gates hold the system. They do not hold a business department that sends no people. When the business side sends nobody, "let us take a look" is free, and an internal zombie pilot outlives any external project, because no bill reminds anyone it is still alive and salaries get paid either way. The three below translate the gates into resources.
**1. Named commitments, claimed in writing before the pilot.** People, hours per week, schedule protection, claimed by the counterpart's supervisor with the sponsor present, written into the pilot launch resolution. Where a company has showback or chargeback (showback only shows the bill, chargeback actually deducts budget), compute and labor are booked straight to the department using them. What this manufactures is a person watching the progress. Whoever's name is in the commitment column will come asking in week 3 whether this thing actually works, because what is being spent is his people's time. The first mover behind a zombie pilot is usually that nobody bears the cost. Technology comes second.
**2. Renewed funding tied to the gates, not to the calendar.** The approval conditions for the next round of people and compute budget are written as the three gates and the graduation criteria, with no dates. A sentence in the annual plan like "add headcount in the second half" is a down payment on failure mode 3 (escalating by the calendar). Rewrite it as a conditional. The three gates cleared and the escalation resolution signed corresponds to the pilot period's people and compute. The graduation criteria met corresponds to the round of resources that comes with formal project approval. That sentence is awkward to say at a budget meeting, and it protects both sides at once. The business side does not send people for a system that is not ready, and your team does not burn schedule on an open-ended "another month of trying it."
**3. A fixed reassessment date that settles automatically when it comes due.** Beyond graduation and kill, the most common third state in reality is "it came due and nobody ruled." Set the reassessment date on the pilot launch day, write it into the launch resolution, and when it comes due with no escalation resolution, settle on how much of the graduation criteria was actually met, convert to formal project approval or shut it down. "Another month of watching" has to go through project approval again. This is the internal version of the price tag. Not deciding is no longer free here. It amounts to an automatic shutdown, and whoever wants a continuation goes and gets project approval again, with the reason written on the form. It works better than any chase email.
Beyond the three there is one question left, who dares say stop. Stopping a pilot internally costs someone face, and whoever says "this should stop" first carries the conclusion for everyone. That is exactly what the zombie pilot failure mode is about. So a kill criteria trigger is written as an automatic action, not as a person's decision. Two unsafe in a single week, the system automatically degrades to human review. The North Star above the baseline two consecutive weeks, the expansion pauses automatically. The action comes before the meeting, and nobody has to call a stop. The row written on Wednesday reads exactly that way. Trigger means degradation, the retrospective locates the root cause, and what the owner signs is the restart, not the stop. The data does the stopping. People only sign the restart.
These three need no new process built. Most companies already have stage-gate reviews or project approval reviews. The PMO's quarterly stocktake, the project approval committee's review meeting, the annual budget and headcount review, the gates are all there, and all that is missing is the criteria on the gate. Write the three gates, the graduation criteria and the kill criteria onto one page, hand it to the PMO or the budget owner, and let them become a standing agenda item of an existing review. How the criteria get into their form is their business. What the criteria are has to come from you, because they cannot write them. This is the internal reader's own lever. The company spent years building these gates, and once you hang your project on them, the review's force is your force. In a company with none of these gates, hang the three on the sponsor's standing meeting and write the reassessment date into the minutes, so it goes on the agenda automatically when it comes due. Chapter 2 said it. The carrier can change. The source of enforcement cannot go missing.
You also have two things an outsider does not. How the last zombie pilot died in this company, you can look up, and its cause of death goes into the candidate rows of this project's kill criteria. The eight-week pre-pilot baseline you pull from the warehouse yourself, without waiting for anyone to hand it over.
These three will make your own manager uncomfortable. After the escalation decision meeting breaks up you call Owen Hartley and say the kill criteria are in the minutes. His first reaction. "You put 'it can be stopped' into the minutes in black and white. Do you still want next year's project approval allowance for this line?" You can answer this way. Grant was willing to approve the pilot in week 9 precisely because the conditions for stopping were on the table. Daring to write the stop standard is what buys a budget someone dares to approve. And renewed funding (that meeting in Chapter 22) rests on the number really coming down on the run chart. Whether the minutes left you a way out will not help.
## Failure Modes
**1. Prototype straight into production.** On demo day the executive says "great, roll it out to the whole department next month," and so prototype code goes into production carrying prototype discipline, no monitoring, no rollback, no change process, a hand-rolled data pipeline. From outside, the organization sees only "it runs," and cannot see the five gaps (Chapter 1). The definition of a prototype is data, workflow, trust, ownership and value papered over temporarily, and "let us run a pilot after all" sounds like going backward in reporting language. One discipline, hold up the change discipline row as a mirror. "What process does changing one line of code in this system take right now?" If the answer is "change anytime," it is still a prototype, wherever it runs.
**2. The zombie pilot.** Month four of the pilot, and the report still says "pilot in progress, optimization continuing." Pilot status is locally optimal for every participant. The engineer has work, the owner need not carry full-volume responsibility, and "currently piloting" is always safe in an executive report. Meanwhile nobody ever defined graduation criteria or death criteria, and whoever says "this should end" first has to own the conclusion (the same shape as Chapter 4's "whoever calls a stop first is the villain"). Discipline, on the pilot launch day write the duration and two exits hard, the graduation criteria and the kill criteria. When it comes due it must take one exit. An "extension" must go through project approval again.
**3. Escalating by the calendar.** "Live in Q3" went into the annual plan, at the end of August the system "goes into production on schedule," and the error categories that never cleared the line get "continuous optimization after launch." The organization's scheduling system eats dates and not conditions, and once a date enters an executive report it takes on a life of its own. A delay needs a written explanation, staying on schedule needs only silence, and error rates do not read the calendar. Discipline, always write outward commitments as conditionals, "live within X weeks of clearing the three gates." When a date and a gate conflict, remember this. Changing the date costs one report. Changing the gate costs one incident.
**4. Kill criteria written after the fact.** Only once the pilot data looks bad does anyone meet to discuss "under what conditions should we stop." By then the standard cannot possibly be neutral. Everyone knows which line kills the project, and the discussion inevitably becomes a negotiation of positions. Whoever wants to continue proposes a loose line, whoever wants to stop proposes a strict one, and the final "standard" is only a mark of power. Writing during the excited period is the opposite. Nobody knows which side the data will land on, and only then can the standard be fair. This is a veil of ignorance you set for your future self, meaning you do not let yourself know in advance which way the data will tip. Discipline, kill criteria are a standing agenda item of the escalation decision meeting, signed the same day as the escalation resolution. A stop standard written only after entering the pilot is void.
!!! note "Vendor View"
The vendor side's version of this gate is called the money gate, and the three are the same shape. Charge for the pilot, even symbolically, because a free pilot is a zero-cost option for the client. Tie payment milestones to the three gates and the graduation criteria, not to the calendar. Write the expiry exit into the contract, and when it comes due with no decision, settle on the graduation criteria and convert to a paid extension or terminate.
## Next Monday
1. Label the project in front of you. Which stage does it claim to be in? Then use the stage rules table to check line by line which set of discipline it actually keeps. The cell where the claim and the practice disagree is your biggest risk right now.
2. Have a prototype "in trial"? Write down the current evidence for each of the three gates, one line each. The one you cannot write is what you should go fill in this week.
3. Have a pilot running past its planned duration? Today, write out its graduation criteria and its death criteria, that is, the kill criteria, one sentence each. If you cannot write them, admit it is a zombie. Kill it, or take it back through project approval.
4. About to enter a pilot? Use [Template 14.3](../appendices/template-14-stage-gates.md) to write the kill criteria this week, while everyone is still excited.
**Want an agent to get you started?** In the repo you set up following [Start Here](../index.md), paste this to your coding agent:
```text
In the repo/ directory of the the-last-mile repository, help me with the Chapter 14 Next Monday actions. First run python3 templates/stage-gates/gate_check.py
with the built-in sample to show the gate one check output, then copy kill-criteria-weekly-report.md to the working directory I name. I tell you the stage the project
claims, and you go line by line through Template 14's stage rules table asking which discipline it actually keeps, marking the cells that do not match.
The graduation criteria and the kill criteria are one sentence each, written by me. You only check whether each sentence can be ruled on by a number or an event, and send back the ones that cannot.
If any command errors, stop and show me the output.
```
---
## Chapter Kit
- **Judgment frameworks.** The stage rules table (three stages ร goal / data / users / change discipline / ways to die); the three escalation gates (eval over the line / willingness to use it daily / owner in place); the pilot kill criteria discipline (written in the same meeting as the escalation resolution, signed the same day); the three resource gates (named commitments claimed / renewed funding tied to the gates, not the calendar / a fixed reassessment date that settles automatically when it comes due)
- **Templates.** [Template 14](../appendices/template-14-stage-gates.md), Stage Gate Checklist and Pilot Kill Criteria
- **Key judgments**
- "The danger is not inside a stage. It is in changing stage without changing the rules."
- "All three gates are clear. Polishing is delay."
- "Kill criteria are written during the excited period. They cannot be written during the disappointed one."
- "The pilot is eval scaled up, not production scaled down."
---
# 15 ยท Co-build with the Engineers Who Will Take Over: Handoff Starts at the First Line of Code
!!! info "Companion Templates"
๐ [Chapter Template](../appendices/template-15-cobuild.md) ยท ๐ [Template Library](../appendices/template-library-index.md)
> **The Challenge.** The claims-ops IT engineers have no time, uneven skill, and a fear of being replaced by AI. "Co-build" sounds fine, and the reality is silence in meetings, assigned tasks done, and never a step taken unasked. And once you move to the next project, they are the ones maintaining this system.
>
> **What You Will Be Able to Do.** Use the four clauses of the co-build agreement (code ownership, backward staffing, the AI code review rule, the rotating release) to turn "an audience on loan to help out" into the owners on handoff day. With an AI coding agent taking part in development, hold the line on "understanding the code and being able to evolve it."
---
## Week 5, Two Engineers "Helping with the Project"
This chapter's story runs backward. It starts in week 5, four weeks before the last chapter's escalation decision meeting.
After the charter was signed, Kevin Doyle dialed in two engineers to "help with the project." Their first week in, the behavior pattern went like this. On time to the weekly, seated nearest the door, answering when asked. Assigned tasks finished on time, done to the spec and not one line further. The interface integration you assigned, done. The edge cases nobody assigned, untouched. Their second week in, identical.
This is not a capability problem. In the days around the charter signing they showed up at the system boundary discussion carrying a sketch for a general-purpose platform (Chapter 8), the kind of thing only people who want to build draw. You politely filed the sketch under "after the second slice," and from then on they placed themselves as an audience on temporary loan. This is your project, and they are here to "help."
The danger is not in front of you. It is at the finish line. On handoff day in Chapter 22, these two are the ones taking this system over. Whether they are co-builders now decides whether what you hand over that day is capability or a legacy.
You have already seen half of where this line goes. Monday of week 7 they brought a multi-agent demo, and at Wednesday afternoon's decision table meeting the enthusiasm was not doused, it only landed somewhere else, on the extraction pipeline, which the two claimed on the spot (Chapter 10). This chapter is the other half, how that Wednesday afternoon's claim becomes one page of agreement and one cadence that runs all the way to handoff. The moment is project week 9, the few days after the pilot is called and before it starts on Monday of week 10.
## Why This Is Hard: What Stops Them Is Identity, Not Skill
There are two default explanations for "the co-build engineers are not invested," not skilled enough and no time. Both are usually true, and neither is the root cause. The root cause is that in a story about "doing work for another department," they have no place where they can win. Their appraisal sits on the claims operations line, the project sits under the Digital Center's name, and those two things are not in the same place.
Do their arithmetic. If the project succeeds, the credit is booked to the Digital Center, and in the year-end report it is "the Digital Center's AI project." If the project goes wrong, they are buried next to it, because the system runs in claims operations and the person woken at midnight is from their own department. Through the project their own job still presses down, their performance targets are not cut by a point, and the appraisal form has no line for "helped with a project." Do more, get more wrong. Do less, nobody blames you. Silence in meetings and assigned tasks done is the rational strategy from that seat, the same rationality as the reviewer saying no in Chapter 12. The burst behind the multi-agent demo does not contradict it. That was people who want to win looking for a place to win, and when they could not find one, they retreated to the audience. This account is harder to see than "doing work for outsiders." Inside one company there is no word for "outsider," so the account hides inside the appraisal lines.
There is one more thing they will not say to your face, the fear of being replaced by AI. Inside a company that fear is sharper. The people being taught are colleagues, and colleagues do not leave. At next year's headcount review, "the Digital Center has already taken those modules" is a ready-made reason to move someone to another post, and layoffs and reassignments inside one company are things that actually happen. The better you teach, the faster it comes, with none of the buffer an outsider provides. The fear is not an illusion. It is only pointed at the wrong object. What gets replaced is writing code to a spec. What cannot be replaced is the person who knows why this code is written this way and where to start looking when it breaks at midnight. So the place where they win has to sit on the latter. The face-to-face explanation and the turn on stage in the four clauses below both move their value from how fast they write to how well they see. Leave this fear unanswered and they will not dare take ownership or visibility, because taking it would be admitting they are training their own replacement.
You have rational arithmetic of your own. Doing it all is faster. Teaching two people is slower than doing it yourself, slower every single day. The cost of co-build lands immediately, and the payoff settles only on handoff day three months out. Inside a company the account is worse, because nobody sets the handoff date for you and the settlement day can keep sliding. Stack both sides' rationality and you get a stable bad equilibrium. You do everything, they watch everything, and on handoff day all the code becomes a legacy.
So co-build is a design problem, and attitude will not solve it. Give them a place where they can win. Code ownership (an asset in their name), skill growth (capability they can take with them), internal visibility (Kevin can see their names). All three have to be designed. Not one of them happens on its own.
## Prior Art, and What AI Changed
The first source is the apprenticeship model. David Maister's argument in *Managing the Professional Service Firm* (paraphrased), the leverage in professional services comes from the master-apprentice structure. A senior person's judgment passes to a junior one through working side by side, and learning by doing is the only transfer method validated over and over. Your relationship with the engineers taking over is, at bottom, an apprenticeship with an expiry date. Inside a company nobody sets that date for you, so set it yourself. Borrow Chapter 22's five self-sufficiency tests (the check Chapter 3 mentioned, whether the receiving side can run the system on its own), and on the day you sign the agreement write down the expected graduation date, which is the week all five pass. Without that date the apprenticeship degrades into pairing forever. What you leave behind is not only the system, it is the way of building systems like it (Chapter 23 lifts this model to the team level).
The second source is the co-build discipline of *Flawless Consulting*. Chapter 1 quoted Peter Block's warning, no deliverable however good can save a counterpart who never truly committed, and the remedy is to make that counterpart a co-builder, with responsibility split 50/50. This chapter is the implementation half. Co-build does not run on invitation. It runs on structure. Ownership, staffing and cadence all get written down.
What did AI change? Both the economics of co-build and its object were rewritten.
On economics, "the business line cannot spare anyone, there is not enough capacity" used to be the all-purpose excuse for your team doing everything. Today agents do the grunt work, the scaffolding, the boilerplate, the test cases. People spend their time on judgment and understanding, and how far a small team can push is no longer decided by headcount. The capacity excuse has expired.
A new risk came with it. A coding agent generates code far faster than a human being understands code. If nobody truly understands it, what you hand over is a pile of black boxes that run and cannot be changed, worse than no code at all, because it manufactures the illusion of an asset. Hence the sentence below.
> **In the AI era, what co-build has to transfer is the ability to understand and evolve code. The code itself keeps getting cheaper. The people who can evolve it keep getting more expensive.**
## The Framework: The Four Clauses of the Co-build Agreement
> **The co-build agreement, one page, turns co-build from a wish into a structure. Four rules, signed by both sides, in force before the pilot starts.** (Template at [Template 15](../appendices/template-15-cobuild.md))
| # | Clause | What It Says | What It Prevents |
|---|------|------|-----------|
| 1 | **Code ownership goes to the receiving side** | From day one all code sits in modules under the name of the receiving side. CODEOWNERS (the list with the right to review and merge that module's code) names them, merge rights are theirs, oncall (on-duty) ownership is spelled out. You commit as a collaborator | The "your project" story; changing the owner only on handoff day |
| 2 | **Backward staffing** | Split the work by "whoever maintains it after handoff leads the writing now," not by "who is faster now" | All the core work going to your team, the receiving side left with chores |
| 3 | **If you cannot say it, do not merge it** | AI-generated code must be reviewed by the maintaining side and explained by them face to face, why it is written this way, where it can go wrong, how to change it | Black boxes that run and nobody can change |
| 4 | **Biweekly rotating release** | Every two weeks, a co-build engineer, not you, demonstrates progress to the owner | Co-build engineers being invisible inside the organization |
Take the four one by one. The two most easily faked are clause one and clause three, one faked when ownership exists only on paper, the other when the explanation is required of the receiving side alone.
**Clause one.** Ownership is not symbolic. Inside one GitLab, ownership has to land in three places to count. The module's CODEOWNERS names the receiving side, merge rights are theirs, and the oncall rotation spells out who owns these modules. CI (the checks and builds that run automatically after code is committed) and the release process follow the business line's existing setup from day one, so on the day you move to the next project there is no such thing as "changing the owner." Ownership also flips the psychological default. In modules under your name they are visitors. In modules under their name you are the visitor. Your AI team should not be the long-term owner of these modules. Inside a company that sentence has no exit date behind it, so two things have to enforce it. Next year's staffing budget reserves no ops headcount for this system, and on the PMO's AI system transfer ledger the owner column for these modules carries a real name from the receiving side. Those two of the three mandatory handoff mechanisms (Chapter 22) get used from the day the agreement is signed, not from handoff day.
**Clause two** connects straight to Chapter 22. The handoff begins with this staffing sheet, and does not wait for a ceremony at the finish. Leading the writing is not writing alone. The lead writer is accountable for the code, can explain it, and decides on merges. The other side pairs and reviews, and the gap narrows in every pairing session.
**Clause three** binds both sides. Code you generated with an agent has to be explained to them too. The "explanation" is tested by a different set of three questions ([Template 15](../appendices/template-15-cobuild.md)), not the three oversight questions of Chapter 12. Why is it written this way? Where can it go wrong, and how would you find out? If it has to change, where do you start?
**Clause four** turns "internal visibility," that place to win, into an institution, and neither a training session nor a status report is the point of it. Whoever stands up and demonstrates gets the credit. The person on stage is an engineer from the receiving side, and the audience is his own supervisor, not yours.
Who the four clauses get signed with depends on which kind of receiving side you have. Inside a company the co-build counterpart comes in three shapes, and the agreement changes with them. Anchor & Helm is the last row, the business line's own IT.
| Shape of the Receiving Side | Who the Co-build Counterpart Really Is | How the Four Clauses Change |
|---|---|---|
| One team, your team keeps the system itself | No engineers to hand it to, so the counterpart is the operations staff on the business side | Clauses one and two spin free. Clauses three and four apply to the operations staff, and what gets explained is rules, thresholds and exception handling, not code |
| A platform team | They keep general capability only, not your business modules | Clause two's "whoever owns it later" gets answered with "not us." Settle before signing which modules go to the platform and which stay on the business line. If that cannot be settled, do not sign |
| The business line's IT | Appraised on the business line, closest to this chapter | Use all four as written. Anchor & Helm is this one |
## At Anchor & Helm: One Page of Agreement, Two Firsts
**Thursday of week 9, signing the agreement.** The pilot was called on Monday (Chapter 13), the escalation decision meeting closed on Wednesday (Chapter 14), and on Thursday afternoon you get Kevin and four engineers into one room, two from the Digital Center and two from claims-ops IT, the same six weeks of headcount Grant approved for the merged view, and put the four clauses on one page. The staffing sheet is filled in backward.
| Module | Maintainer After Handoff | Lead Writer | Pairs |
|------|-------------|------|------|
| Extraction pipeline (emails in, schema out, validation, the spot-check tool) | Claims-ops IT | Claims-ops IT | Digital Center |
| Queue core (sorting, decision trail, the Human Call flow) | Claims-ops IT (longer term) | Digital Center | Claims-ops IT |
| Merged view (derived from two sources of truth, six weeks) | Claims-ops IT | Digital Center (lead-writer rights handed over before handoff) | Claims-ops IT |
| Rules and de-identification as configuration | Claims-ops IT | Claims-ops IT | None (Chapter 12's maintainability row, about a week) |
Claims-ops IT leading the extraction pipeline is not a favor. They claimed it on Wednesday of week 7, and the prompt inside the detection agent, with the agent shell stripped off, is version one of the extractor (Chapter 10). On the queue core you lead the writing and they pair. Its complexity right now is beyond them, but "the longer-term maintainer is claims-ops IT" is written on the sheet, so the pairing has a direction. Before signing, Kevin asked exactly one question about cost. "Whose account does their time come out of?" The project's, you said. Inside a company a project has no ledger, so that sentence has to land in two things to count. One is Kevin committing weekly hours in writing as their supervisor, written into this agreement's commitment clause, a clause of the same rank as Linda's team's 2 hours a week. Where the company has chargeback (cost transferred between departments, charged to whoever uses it), those hours are booked to the project, and the bill does the reminding for you. The other is one line each in the two engineers' quarterly objectives, lead writing and gatekeeping on these modules.
Cost has a second question, who carries their own day job. Booking the time settles only where the account goes, not who does the work, and only Kevin can answer this one. It cannot be left to two people squeezing their evenings. His only real answer is a reshuffle. Their existing schedule gives up the matching hours, and who takes them and in which week goes into the same agreement. Clause four's rotating release is the companion to it. Kevin sees with his own eyes every two weeks what this investment produces, and the reshuffle does not quietly slide back.
**Clause one** is harder to negotiate than the staffing, because it has to get through Victor Reyes's gate (Chapter 12). The conclusion first, read-only on the whole repo, write access on the modules named in the agreement. Here is how it went. Group people holding write access to a subsidiary's code repository is not something he grants by default. Code repositories are tiered for approval the way data is, he approves cross-subsidiary access one request at a time, and the detail table carried out of the company two years ago is the reason. The way out is not whether you can get in, it is what someone who gets in can touch. What you proposed is read-only on the whole repo, write access only on the modules named in the agreement, and the merge itself pressed by someone from claims-ops IT. It shares a root with the permission inheritance he set for the queue, no separate account system, and whoever can touch this in the core system is who can touch it in the repo. If even restricted write access will not clear approval, what gives way is still not ownership. You commit on a restricted branch opened inside the same repo, CODEOWNERS and CI stay on the claims operations side, and on handoff day there is still no such thing as "changing the owner."
**The first "if you cannot say it, do not merge it."** In pilot week 1 the younger engineer submits the email pull retry logic of the extraction pipeline, generated by an AI coding agent, all tests green. At review you ask only the first of the three questions. "Why three retries, and why does the interval double?" He pauses. "That is how the agent wrote it, and it runs." By clause three, sent back. Things were stiff afterward. The code is not wrong, so why not merge it? You slide the agreement across. The code passes, the explanation does not. When this code breaks at midnight, the person woken up is him, not the agent. Two days later it comes back, the retry ceiling, the backoff strategy (wait a little longer before each retry after a failure), which dead-letter queue (the queue that collects messages that keep failing) failed items land in, and he has had his hands on every one of them and can explain each. He said something himself that you still remember. "This time I know why it is written this way."
There is a second thing after that rejection. You hand over enforcement of the rule as well. From that week merges on the extraction pipeline are gatekept by the older engineer, and you are no longer this module's gatekeeper.
**The first rotating release.** Friday of pilot week 2, the first beat of the biweekly cadence. The people on stage are them, not you. The spot-check results of the extraction pipeline against two hundred real emails, the accuracy of missing-document detection, and a live walk-through of three extraction errors, what a wrong one looks like and how the spot check caught it. You sit in the audience and say nothing the whole time. Afterward Kevin keeps them back for another ten minutes. You notice one detail. For the first time he says both their names. For the six weeks before that, in Kevin's mouth they had been "those two I gave you."
Internal visibility starts compounding from that day. The rotating release runs one beat every two weeks, weekly after the team expands (Chapter 22), scheduled all the way to handoff, and the cadence sheet is in [Template 15](../appendices/template-15-cobuild.md). This cadence is also the prelude to Chapter 21's operating cadence.
There are three places where doing these three things inside a company is easier than outside. All three are written here as actions. First, the three explain-it questions do not have to go into this project's PR (pull request, the request to merge code) template alone. Open a PR against the company-wide PR template, add a field for the share generated by AI and the name of the person who explained it, change it once and every project benefits. Second, the rotating release needs no venue you build yourself. The company's internal tech talk series is a ready-made stage, put both their names on the sign-up sheet, and the audience gains a layer for free. Third, you can see the other side's performance sheet, so the place to win can be made real. Ask Kevin to write the rotating release demo into a line of their quarterly appraisal, and visibility walks out of the meeting room and into the appraisal form.
## Failure Modes
**1. Your team does it all, three fast months, dead at handoff.** You and the Digital Center's engineers write all the code, the progress looks good, and the co-build engineers' "participation" is down to a weekly meeting. Doing it all is rational in every single moment. You are fast, they are slow, and a deadline only looks at the moment. The payoff of co-build settles on handoff day, and nobody wants to pay today's cost for an account three months out. An AI coding agent digs this trap deeper. What you and an agent produce in one evening takes two weeks to teach someone else. The consequence is cashed in Chapter 22. On handoff day the code is complete and the capability is zero. One test. If the co-build engineers' share of commits has been zero two weeks running, you are already doing it all.
**2. Co-build engineers used as cheap labor.** There is a split of work in form, but what they get is writing documents, building test data and adjusting page styles, while the core modules are "too critical, we will take those first." Split by "who is faster now" and the core work necessarily goes to the experienced hands. That split is right every single time, and add them up and it is entirely wrong. Worse, it confirms their deepest fear. AI plus people from the department next door prove together that "you are optional." The identity problem gets worse, and silence ferments into resistance. The remedy is clause two. The test for staffing is "whoever owns it later," not "whoever is faster now."
**3. AI code nobody can maintain.** The repo swells at agent speed, all tests green, and not one piece of core logic can be explained face to face by anyone. The scissors gap between generation speed and comprehension speed, on top of a review whose default test is "does it run." CI can verify running. It cannot verify understanding. Leave "explanation" out of the merge bar and black boxes necessarily accumulate, because "tests passed, merge it" is faster every time. The consequence lands at the first change request after handoff. Nobody dares touch it, the system freezes at the state it was handed over in, and then it gets worked around and abandoned. There is only one defense, make understanding the bar for merging rather than a wish held after the merge.
**4. Co-build turns into custody.** All four clauses are signed, CODEOWNERS names the receiving side, and yet you make every design call, you give every PR its final review, and you rescue every release. Co-build in form, doing it all in cognition.
Custody reinforces itself. Your reviews are the best, so everything comes to you for review. Their judgment gets no practice, the gap widens instead of narrowing, and that further proves "it still has to be you."
The warning signal, three months in, you are still the merge gatekeeper on every module, and at the rotating release you answer Kevin's questions for them. The rule, gatekeeping rights (the power to decide a module's merges, [Template 15](../appendices/template-15-cobuild.md)) are handed over module by module (at Anchor & Helm, gatekeeping rights over the extraction pipeline moved the same week as that first rejection). At the rotating release you only add, you do not answer for them.
!!! note "Vendor View"
The vendor side's clause one is physical. The code goes into the client's repo, you commit with an external collaborator account, the account is revoked on your exit date, and ownership needs no ledger behind it. Inside one company, one GitLab, ownership comes down to three lines, CODEOWNERS, merge rights and oncall. There is no account to revoke, so headcount (no ops headcount reserved) and the ledger (real names on the transfer ledger) have to do what account revocation used to do.
## Next Monday
1. Open the project repo and look at ownership. Who is in the module's CODEOWNERS, who owns oncall, do the co-build engineers have merge rights. If all of it sits with you, change CODEOWNERS this week. The later, the more expensive.
2. Draw a backward staffing sheet. For each module write "who maintains it after handoff" first, then set it against "who leads the writing now." The modules where the two columns disagree are tomorrow's legacy, so reshuffle the lead writing this week.
3. Add one rule for merging. AI-generated code merges only after the maintaining side explains it face to face, using the three explain-it questions in [Template 15](../appendices/template-15-cobuild.md). Run it on your own next PR first.
4. Schedule one rotating release. Two weeks out, a co-build engineer demonstrates progress to their owner, and you sit in the audience.
**Want an agent to get you started?** In the repo you set up following [Start Here](../index.md), paste this to your coding agent:
```text
In the repo/ directory of the the-last-mile repository, help me with the Chapter 15 Next Monday actions. First open my project repo's CODEOWNERS
and list who owns each module now, and tell me if the file does not exist. Then build a backward staffing sheet following templates/cobuild/ownership-checklist.md.
I fill in "who maintains it after handoff," you count "who leads the writing now" from the commit history, and mark the modules where the two columns disagree. Merge the three explain-it
questions in pull-request-template.md into my project's PR template, show me collaborator-access.yaml as a sample only, and leave permission changes to me.
If any command errors, stop and show me the output.
```
---
## Chapter Kit
- **Judgment frameworks.** The four clauses of the co-build agreement (code ownership goes to the receiving side / backward staffing / if you cannot say it, do not merge it / the biweekly rotating release); the three explain-it questions (why it is written this way / where it can go wrong and how you would find out / where you start to change it)
- **Templates.** [Template 15](../appendices/template-15-cobuild.md), Co-build Agreement and Knowledge Transfer Cadence, fillable clause by clause, with the backward staffing sheet and the biweekly rotating release agenda
- **Key judgments**
- "What stops co-build is identity, not skill. Give the co-build engineers a place where they can win."
- "In the AI era, what co-build has to transfer is the ability to understand and evolve code. The code is only the carrier."
- "If you cannot say it, do not merge it."
- "The test for staffing is 'whoever owns it later,' not 'whoever is faster now.'"
---
# 16 ยท Production Engineering: Set the Running Budget Before You Launch
!!! info "Companion Templates"
๐ [Chapter Template](../appendices/template-16-production-readiness.md) ยท ๐ [Template Library](../appendices/template-library-index.md)
> **The Challenge.** Everything was fine in the pilot, so how does real traffic bring an exploding bill, latency over the line, and something acting up every few days?
>
> **What You Will Be Able to Do.** Stand up the four ledgers of the running budget for an AI system (cost / latency / error / degradation), build the minimum observability set, and turn "how much a month, how much per claim" into one page an owner can manage against.
---
## Monday of Week 12, One Bill
The pilot starts on Monday of project week 10 and runs eight weeks. The first two weeks are calm. The eight people on Linda Marsh's team use it every day, the eval spot checks turn up no unsafe, and Kevin Doyle's weekly report carries the phrase "chasing early" for the first time.
The pilot enters week 3, project week 12. On Monday Finance forwards the monthly bill to Kevin, Kevin forwards it to you with one line attached. "Is this number normal?" The LLM call spend is 4 times the estimate you gave him before the pilot started.
There is a second thing the same week. Tuesday at the morning peak, a queue refresh goes from 3 seconds to 40. On Wednesday Linda puts it to you politely. "The queue takes forty seconds to turn over. We cannot wait that long, so we have gone back to working out of the inbox. The inbox does not make you wait." The usage curve drops on cue. Two days of investigation turn up three causes.
1. **Full recomputation.** Every refresh by every person recomputes the whole team's several hundred claims from scratch, while fewer than two in ten claims actually have a new event on a given day.
2. **A retry storm.** An extraction call that fails is retried at once, with no cap on attempts, and a failed claim comes around again on the next full recomputation. Rate limiting, retries, more rate limiting, feeding each other.
3. **The whole operations manual stuffed into the prompt.** To lift the accuracy of the document checklist during the prototype, the entire manual was pasted into every call. Claim volume was small then, and nobody was watching tokens.
The sharpest part is that not one of these three decisions was wrong during the prototype. Full recomputation was the simplest, unlimited retries the least trouble, stuffing in the manual the fastest to show results. The prototype optimizes for learning speed, and all three moves were right for it. They just were not put on trial again when the stage changed (this is Chapter 14's changing stage without changing the rules, and this time the bill is what dragged it out).
## Why This Is Hard: An AI System's Running Characteristics Are Not the Same Shape
Traditional software's assumption of "launch it, then hand it to ops" rests on three premises. An AI system satisfies none of them.
Cost grows linearly with calls, and there is no marginal cost trending to zero. Double the users of traditional software and server cost barely moves. Every LLM call is close to full price. Even with the model provider's own caching, a stuffed context still pays a bill in latency and diluted attention, and the more irrelevant content it holds, the more easily the model misses the few lines that matter. Double the claim volume and the bill doubles, and wasted calls are billed to the last cent. Cost has become, for the first time, a running variable that occurs claim by claim, not a fixed investment bought once and never charged per transaction again.
The latency distribution has a long tail, and the tail lands exactly on the rhythm of the business. Average latency means nothing. P50 looks good and P95 kills you. Line up a period's requests by how long they took. The one standing in the middle is P50. The one at the 95% mark, with only a small share slower than it, is P95, and only 5% of requests are slower. Peak requests also cluster naturally into one window (reviewers all open the queue in the morning). Those 40 seconds were the long tail colliding with the morning peak, and an "average" cannot see it at all.
Quality drifts silently. The code has not changed and the system reports no error, but the input distribution has moved. Claim mix, fraud methods and the repair shop list are all in motion, and model behavior moves with them. No exception stack, no alert. The first to notice is usually a front-line hunch or a customer complaint.
The traditional trio of ops monitoring (uptime, error codes, resource usage) is blind to all three. So there is only one conclusion. The running budget has to be designed like a feature, holding a place in the architecture and the schedule and a vote at the launch gate, rather than written up as an ops manual after launch.
## Prior Art, and What AI Changed
The ledger our predecessors left is called the error budget. The core idea of the Google SRE school (paraphrased here). Reliability is a budget, not a wish. 100% availability is the wrong target. The right move is to set a reliability target, and the gap between it and 100% is the budget, which the team spends freely inside (releases, risks), freezing changes once the budget burns through. It turned "do not err" from a moral expectation into a manageable account. Its companion is observability discipline. A system has to be able to answer "what is it doing right now, and why."
What did the AI era change? The budget goes from one ledger to four. Beyond the error budget come the token cost ledger (cost became a running variable, claim by claim), the latency ledger (the long tail colliding with the business rhythm) and the quality ledger (the error ledger in the framework below). The per-category threshold table of Chapter 11 turns from an acceptance document into a running instrument, and the eval goes online (Chapter 18 puts it as the threshold table becoming the monitoring definitions as is, meaning which table monitoring reads from here on, the same thing).
And "degradation" gained a new meaning. Traditional degradation cuts features to protect the core. An AI system's degradation cuts intelligence to protect the process. The LLM goes down, gets slow, gets expensive, and the question is whether you can fall back to rules mode and keep the core process running.
Chapter 10's "start with the dumbest thing" pays out its second value here. Do not delete the dumbest thing that works. It is your degradation path.
## The Framework: The Four Ledgers and the Minimum Observability Set
**The four ledgers of the running budget.** One budget line, one action on breach and one named owner each. Read the first four columns first. The last column holds Anchor & Helm's actual numbers, looking back across the whole pilot, at a point later than this chapter's story, so take your time matching it up.
| Ledger | What It Records | How the Budget Line Is Set | Action on Breach | Anchor & Helm Magnitude |
|------|--------|--------------|----------|----------|
| **Cost** | Cap on cost per claim, plus the monthly total | Worked back from business value. What the labor hours one exception claim saves are worth, with system cost allowed only a fraction of it. Do not work back from last month's bill | Trigger a cost review, instead of finding out at month end | 4 times the estimate when it was out of control, cost per claim back to 1/6 after the four moves |
| **Latency** | The P95 cap (not the average), with the peak window on its own line | Worked back from the user's working rhythm. From opening the queue to being able to work, how many seconds can they wait | Degrade, or extend precomputation | Morning peak refresh back from 40 seconds to under 3 |
| **Error** | Per-category error rate caps | Carry over the Chapter 11 threshold table as is, build nothing new. unsafe zero tolerance, concern โค10%, useless โค20% with the trend not rising, measured by online sampling | Rework. A kill criteria trigger means pause (Chapter 14) | One suspected unsafe case across the whole pilot, caught by Linda in week 14 (Chapter 18), short of the two-in-one-week trigger line. concern under half the threshold, useless around one in ten, and it stopped surfacing after that batch of repair shop mislabels was fixed in week 13 |
| **Degradation** | A fallback for every AI dependency point | The budget line is "the drill passed." Really pull the dependency and the core process still runs | A dependency point with no fallback does not launch | The LLM dependency cut for half an hour, the queue fell back to rules mode and review carried on |
The cost ledger has a precondition inside a company. The account has to be separable in the first place. Your calls are most likely mixed into the Digital Center's shared gateway or the company's single cloud account, and what comes out at month end is one department total that shows nothing about which system spent what. Tagging this system's calls with FinOps (the practice of splitting cloud spend by system), or hanging them on a showback report (showback only shows the bill and deducts no budget), is the precondition for standing up a cost ledger, not a bonus. Without the split you do not even have the month-late reminder that is the month-end bill, and the first time anyone notices the cost will be at next year's budget meeting.
The owner column holds one more pitfall specific to the inside. The owners of the four ledgers are often not in the same department. Cost follows the budget account, and the action on breach needs the business side to make the call. Latency and errors live in your system, and the actions are done by your people. Degradation's fallback needs the business side to accept that half hour of no supply. So each ledger names exactly one first-line owner, and the test is whose budget or whose daily actions a breach hits directly. Everyone else related is written in as a countersignature (confirming together but not carrying the main responsibility). The first-line owner does not have to be the person who fixes it, but the breach alert can land in only one inbox. Two names is the same as no name.
**The minimum observability set.** Three pieces, and missing one means a ledger cannot be kept.
1. **End-to-end trace**, any single suggestion replayable. A trace is the complete record of one suggestion from input to display, the input snapshot, what was retrieved, the prompt and model version, which rules hit, the final display. It is not the same thing as Chapter 12's six decision trail fields. The six fields are an audit commitment, answering "who decided what," and they are for Victor Reyes and compliance. A trace is engineering replay, answering "why does this suggestion look the way it does," and it is for whoever is debugging. One chain, two readings. Do not conflate them, and do not build only one.
2. **Online eval sampling**, golden cases replayed on a cadence, plus a proportion of production output spot-checked by human review, booked by the Chapter 11 categories. The number of rows spot-checked decides how fine a ratio you can read, and the sample size follows Chapter 11's passage on how to read a ratio line. When people cannot keep up, let a model screen first, but the model's verdicts count only once they have been calibrated on a human-annotated sample, with the disagreement rate read separately by error category. On the unsafe class the model may only report, never release.
3. **Drift signals**, sudden-change alerts on the input distribution (claim mix, amount distribution, policy type composition) and on the override rate. The override rate is the share of system suggestions the reviewers push back, and when that share moves, usually the input or the model moved first.
The last two do not serve this chapter alone. They are ready-made components of Chapter 18's launch monitoring surface.
## At Anchor & Helm: One Week of Repairs, One Page
The repairs take you and the two claims-ops IT engineers (the headcount set in Chapter 15) one week. Four moves on cost.
**Full recomputation becomes incremental.** Only claims with a new event get recomputed, a new email, a status change, a timeout. Call volume drops by a large chunk right away.
**Cache the extraction results for stable fields.** The extraction result for one email does not change on its own, so extract once and store it, and anything that passes schema validation (all fields present, types matching) goes into the cache. Until then every round of full recomputation was paying again and again for the same email.
**The manual moves out of the prompt and into retrieval.** Each call now carries only the two or three passages tied to the current claim's document type, and token length collapses. Accuracy did not drop, confirmed by two weeks of spot checks (changing the call path also has to clear the eval. That is discipline, not ceremony).
**Long-tail claims routed to a smaller model.** The rules layer already absorbs most claims (Chapter 10's step two, the part where rules are the floor), so cut once more inside the long tail that needs the LLM. Claims whose signals are simple and merely uncovered by the rules go to a smaller model. Only the genuinely hard ones use the big one. The eval decides where the split goes, and if the smaller model clears the threshold on that subset, the subset is its.
With all four moves in, cost per claim falls to 1/6 of what it was, and the monthly total is back inside the estimate line.
The retry storm gets treated on its own. The source is the LLM call in the extraction step, while the email pull that was rewritten back then turns out to be fine, because that stretch has carried a cap and a backoff (wait a while before retrying after a failure, with the wait lengthening each time) all along. The fix for the extraction stretch is backoff plus a retry budget, at most two attempts per claim, and over the limit it goes to the "pending human" column. The email pull keeps the three attempts with doubling intervals set in Chapter 15, which is the standard for pulling. The new ledger for LLM calls tightens to two, because one call costs far more than one pull.
The person doing the work is exactly the engineer who rewrote that retry logic back under "If you cannot say it, do not merge it." (Chapter 15). Last time he learned to make one stretch of retries explainable. This time he opened a ledger for retries across the whole chain.
Latency gets two moves of its own. Precompute the queue, working the whole team's queue out before the morning peak. Make the refresh asynchronous, so opening it shows the most recent result at once (the interface marks the data's timestamp) while the background updates incrementally. The morning peak refresh comes back under 3 seconds, and usage climbs back to where it was a week later.
The degradation drill is set for Friday afternoon. You do something that surprises Linda's team, cutting the LLM dependency on purpose for half an hour. The queue falls back to rules mode, the five rules of thumb rank claims as usual, the reason column says as usual which rules matched, and the missing documents column grays out to "extraction paused, open the claim to read the original email." Review carried on through that half hour, and some people never noticed. The drill counts as passed on two conditions, that review carried on through that half hour with real users present, and that when asked afterward "did you notice," most say no. The cadence goes into the register, every dependency point drilled at least once more within 90 days, with a rerun mandatory after a model swap or a change to the call path ([Template 16.3](../appendices/template-16-production-readiness.md)). This is the second identity of the rules layer Chapter 10 left behind. In normal times it is the floor. When supply is cut it is the fallback.
You do not have to invent an occasion for the drill. The company most likely already has an annual disaster recovery drill or a business continuity drill on the books, so hang the pull-the-plug drill for AI dependency points onto that one. The schedule, the notification templates and the acceptance records are all there already, and it comes with a date somebody else keeps for you. All you have to do is add one line to that drill list, saying which dependency gets pulled, what it falls back to, and who is present to watch.
At the next weekly you bring no repair report. You bring one page, the four ledgers. Kevin looks at it and says, "When I signed off on that bill last month I felt hollow. I did not know whether that number counted as normal." Now he knows. This system has a cap on monthly running cost (back to the estimate line), a cap on cost per claim (1/6 of what it was before it ran away), a morning peak P95 of three seconds, error thresholds that are exactly the table he ruled on in Chapter 11, two degradation switches with their triggers, and one named owner per ledger.
"This page stays with me," he says. It is the first time Kevin can manage this system against numbers instead of against feeling. You add a line of small print in the footer while you are at it. At the handoff when the pilot ends, this page becomes a budget account, and running cost moves from project funds into claims operations' department budget (Chapter 22's setup starts on this page).
The hard part is not the accounting move. Moving this money into claims operations' budget account is half an hour of work for Finance. The hard part is whether claims operations will own that line in its own department budget next year. Owning it puts one more number on their cost sheet, and that number will get asked about at the quarterly business review. Which is why the money has to be separable before it can move. Kevin will not accept a total mixed into the Digital Center's shared gateway account, because he cannot read what makes it up and you cannot prove the money belongs to this one system. Split usage by system first, then talk about moving it.
The business side is a real owner only once it takes the budget. Whoever holds the budget holds the priorities (week 22 in Chapter 22). From the first month Kevin pays for it, what he asks when he schedules is whether his own people are worth it, not whether you can. If he does not take it, the money stays on the Digital Center's project funds, your cost center feeds it forever, and the team that feeds it ends up being its ops team. That is the financial source of permanent ops.
The four moves did not contribute equally. All the field left behind is a qualitative account. Full recomputation to incremental cut the number of calls, the manual moving into retrieval cut the length of each call, and those two are the largest in magnitude. The cache removed the part where the same email was paid for again and again, ranking behind them. How much the long-tail routing to a smaller model saved was not recorded at the time, and the calculation was done once before the handoff. About six in ten of the long-tail claims went to the smaller model. That move accounts for less than a tenth of what the four saved in total, the smallest of the four. This matches industry experience. Gateway-style routing usually saves only ten to twenty percent of inference cost, so do not expect it to carry the load. Take the moves in this order.
Here is something an outside team cannot do and you can, going to look at the account directly. The cloud platform console and the FinOps reports sit inside the company, so you do not have to wait for someone to forward you the monthly bill. You can pull it by day and by step yourself. The more valuable half is the denominator, the baseline number you can compare against. Go ask the department next door about the comparable system it launched a year ago, what its cost per claim is, and benchmark once sideways. An outside team can only get the summary sheet the other side is willing to give. You can get the detail, and you can get other people's numbers.
## Failure Modes
**1. No degradation path.** The LLM service twitches and the whole system is paralyzed, with reviewers staring at a spinning page. The cause has two layers. A degradation path is a "negative feature," invisible in normal times and unimpressive in a demo, so it always loses the schedule to new features. The other layer is the capability ceiling criterion coming back (the criterion named in Chapter 10, engineers look at what a model can do, enterprises look at what an error looks like when it is wrong), "we have an LLM now, what do we need rules for," and the rules layer gets deleted as transitional scaffolding. But the rules layer is an asset, not scaffolding (Chapter 10). The discipline. Every AI dependency point registers its fallback and really drills it before launch.
**2. Monitoring that looks only at uptime, that is, availability.** Availability is 100% month after month while suggestion quality has been drifting for three weeks, and the first to notice is Linda's hunch. The monitoring list is inherited from traditional software, uptime, error codes, resource usage. That system has no place for "quality drifting silently," because traditional software's behavior does not change while the code stays the same. When the quality ledger is missing, an all-green dashboard is precisely the most dangerous thing. A system 100% online does not mean it is still saying the right things. The discipline. Online eval sampling and drift alerts launch at the same rank as uptime monitoring, not in phase two.
**3. Cost noticed only after the fact.** The month-end bill is the only cost monitoring. A bill settles monthly, so the feedback is thirty days late by nature. The more structural layer is ownership. In an organization the bill belongs to Finance and the code belongs to engineering, so "cost per claim" is on no engineer's instrument panel, and a number nobody owns does not get managed. The discipline. The cost instrument produces numbers daily, split by step, and breaching the budget line triggers a review the same day. The month-end bill is a cost autopsy. It cannot serve as cost monitoring.
**4. A budget set that nobody reads.** The four ledgers are standing, the one page has been handed over, and three months later nobody has opened it. The budget lines are set on paper, nobody gets woken when one breaks, and the owner column holds a job title instead of a name. The launch gate is passed once, the running accounts need someone reading them every day, and the two get treated as one thing. Deeper down, reading the numbers has no moment. Anyone may read them, which is the same as nobody reading them. The discipline. Bind every ledger to a named owner and a fixed moment for reading its numbers, with a breach landing in his inbox automatically. The test. Pick any ledger and ask its owner "did it break a line last week." No answer means that ledger is not running.
## Next Monday
1. Work out cost per claim for your system once, last month's call bill divided by the number of claims handled. Take the number to the system owner and ask "do you know this number." If he does not, the cost ledger has no owner yet.
2. Open your largest prompt and find the static blocks inside it (the manual, whole rule texts, piles of examples). Ask one question. Can these move into retrieval or a cache?
3. Pull the LLM dependency in the test environment for ten minutes and see what is left of the system. If the core process is not left, you have no degradation path. Go back to Chapter 10 and get back the dumbest thing that works, the one you deleted.
4. Run the launch gate once with the checklist in [Template 16](../appendices/template-16-production-readiness.md), and fill in the owner column of the four-ledger table. Whichever ledger you cannot put a name to is the address of your next incident.
**Want an agent to get you started?** In the repo you set up following [Start Here](../index.md), paste this to your coding agent:
```text
In the repo/ directory of the the-last-mile repository, help me with the Chapter 16 Next Monday actions. First run python3 templates/production-readiness/cost_dashboard.py
with the built-in sample to show me the four-ledger instrument, then open drift-alerts.json and trace-schema.json and explain every field. Then I give you
last month's call bill and the number of claims handled, and you only work out cost per claim. Whom I take the number to is mine to decide. Build Template 16's checklist as a table, with the four ledgers' owner column
filled in by me with real names, leaving what I cannot fill blank and flagged red. You may find the static blocks in my largest prompt. Whether they move is mine to decide.
If any command errors, stop and show me the output.
```
---
## Chapter Kit
- **Judgment frameworks.** The four ledgers of the running budget (cost / latency / error / degradation, each with one budget line, one action on breach and one named owner); the minimum observability set (end-to-end trace / online eval sampling / drift signals)
- **Templates.** [Template 16](../appendices/template-16-production-readiness.md), Production Readiness Checklist and Running Budget Sheet, the item-by-item launch gate, the four-ledger template (with the Anchor & Helm example rows), the degradation path register
- **Key judgments**
- "The running budget has to be designed like a feature."
- "Do not delete the dumbest thing that works. It is your degradation path."
- "A system 100% online does not mean it is still saying the right things."
- "The month-end bill is a cost autopsy. It cannot serve as cost monitoring."
---
# 17 ยท From Dashboard to Action Queue: Put the Next Action in Front of the User
!!! info "Companion Templates"
๐ [Chapter Template](../appendices/template-17-action-queue.md) ยท ๐ [Template Library](../appendices/template-library-index.md)
> **The Challenge.** The business side loved the dashboard you delivered. Three months later nobody opens it. What is missing between "we can see it" and "someone acts on it"?
>
> **What You Will Be Able to Do.** Measure how deeply a deliverable is embedded in the workflow with the action integration ladder. Translate "I want a big screen" into queue columns. Design an action queue with a decision trail, so that override data becomes the fuel that makes the system smarter.
---
## Friday of Week 9, the Last Afternoon Before the Pilot Starts
The prep meeting is breaking up. Kevin Doyle, packing his things, says as if in passing, "It goes live next week. I still want a big screen, up on the wall in the department. Backlog, where things stand, all at a glance. And when Grant Whitmore asks, I have something to show him."
You know this sentence. In Chapter 0 it was Kevin's first ask, "Ideally with a dashboard..." In Chapter 8 the "not this phase" list kept it out. This is the big screen's third appearance, and this time there is no way around it. Next Monday Kevin is the pilot's owner. Leave the wish hanging and sooner or later he will build a big screen of his own, one that is always greener, competing with the queue for the right to explain. When two screens disagree, nobody knows which one to believe.
This time you did not block it. You asked three questions, and the conversation is replayed in the second half of this chapter. First, the reasoning.
## Why This Is Hard: Four Organizational Things Stand Between Seeing and Acting
For an insight (one conclusion read out of the data) to become an action, it has to clear four organizational gates.
- Owner. Who does this piece of information belong to?
- Next action. What is the specific thing to do right now?
- Reason. Why should he trust this judgment?
- Capture. Where is it recorded that it was done or not done, and where does the next person pick up?
All four are organizational matters, not information matters. And the dashboard's design choice is to leave all four to the viewer to solve on his own, so the viewer chooses not to solve them. Do not blame the viewer for laziness. This is structure. The people who look usually have no power to act, and the people who can act are usually not looking. The screen is visible to everyone, so everyone can assume "whoever should act will act." Information present, responsibility absent.
Chapter 0 gave the user version of this judgment, the workflow claim, whose next action the system will change. This chapter is its system design version.
> **A system's value is not in what it knows. It is in whose next action it changes.**
By that standard a dashboard delivers "knowing," and clears none of the four gates. It is not useless. Watching trends and giving managers a sense of presence (Kevin's real need, see below) are both real uses. It has only one way to die, being treated as the deliverable meant to change front-line action.
## Prior Art, and What AI Changed
**The lean tradition. A kanban is a next action, not an information display.** A Toyota kanban card says who, fetch what, how many, deliver where. The card arrives, the action happens. The software industry took the cards and the swimlanes, and what it most often dropped was exactly this half. The action queue brings it back. Every row is a real kanban card.
The metrics tradition. Watch controllable inputs, not outputs on display. The controllable input metrics idea in *Working Backwards* splits metrics in two. Output metrics (revenue, total cycle time) are results and can only be watched. What today's action can change is the input metric (how early a missing document is found, chase response time). A big screen that shows only output metrics gives the action no handle, by definition.
What did AI change? One thing removed an old excuse, and one thing created a new necessity.
First, the cost of translating data into a suggested action collapsed. Stopping at display used to have a respectable reason. Translating "backlog is rising" into "chase these seven claims first" meant hard-coding rules for every kind of situation, too expensive for anyone to write. LLMs made that layer of translation cheap (the pattern selection in Chapter 10 does exactly this job). Stopping at display has had no excuse since.
Second, the decision trail went from virtue to necessity. Traditional systems keep logs for audit, written for the day someone might look. An AI system's trail has one more recipient, and it shows up every day. Each person's accept or override of each suggestion, plus the reason, is the only continuous right-or-wrong signal produced outside eval, and it flows back into rule iteration and golden cases (Chapter 11). An AI system without a trail is pouring out its most expensive training data on the spot.
## Framework One: The Six-Level Action Integration Ladder
> **The action integration ladder measures how deeply a deliverable bites into the workflow, six levels from "can be seen" to "gets smarter."** Judge level by level, and stop at the first level you cannot produce evidence for.
Read only the first two columns to place yourself. The last two are for a closer look on review.
| Level | Gloss | Test for Standing Here | What the Next Level Costs |
|------|----------|------------|------------------|
| **visibility** | It can be seen | Data is aggregated and displayed, and someone looks at it | (Starting point) Just data and charts |
| **prioritization** | It is ranked | The ranking logic matches the users' real trade-offs (chase โ risk, Chapter 0) | One value judgment, what matters more and who gets to say |
| **recommendation** | It suggests | Every row gives a next action, and the reason can be challenged | Tacit knowledge, the suggestion logic eats front-line judgment (Chapters 6, 11) |
| **task creation** | It becomes a task | The suggestion lands on a named person, in the work entry point he already uses, with a status and a due date | Organizational ownership, owner negotiation + workflow embedding |
| **decision capture** | The decision is kept | Accept / override is recorded along with the reason | The front line's trust, willingness to put a real name on a decision (Chapter 12) |
| **closed loop** | The loop closes | Someone reviews the trail data on a cadence, and it has actually changed a rule or an eval | Operating cadence, a review mechanism that stays alive long term (Chapter 18 takes it over) |
The queue at Anchor & Helm stands at level five during the pilot. The owner column, the Human Call column, and reason codes have been in Linda Marsh's team's work entry point since day one of the pilot. Level six has been cashed in only once, the review in the second half of this chapter that changed one line of a condition and added golden cases. One catch is not a cadence. Only once the review grows into a fixed mechanism does the queue qualify for level six, and that last cell of the table is handed to the next chapter.
Two rules for reading the table. First, most dashboards die at level one, and most "AI assistants" die at level three, with suggestions, no owner, no trail. Second, the higher you go, the less engineering and the more organization. The lower half can be reached by writing code. Every level of the upper half has to ask the organization for something.
This makes three ladders in the book. The outcome ladder measures how far a project has gotten (Chapter 1). The data fitness ladder measures whether a data source deserves to drive an action (Chapter 9). The action integration ladder measures how deeply a deliverable is embedded in the workflow. Each measures its own dimension. They do not swap and do not convert.
What the upper half asks of the organization, the internal reader can get, with a shortcut outsiders do not have. Level four stalls on owner negotiation because the person writing code has no authority to claim responsibility on someone else's behalf. You are inside the company, and you can push directly to change the SOP. Write "high-priority claims in the queue are handled by the on-duty team lead the same day" into the claims operations work standard, and level four drops from an organizational negotiation to a document revision, through the existing standards revision process, signed off by the business side's supervisor, not you. Institutional anchoring is the biggest lever in the internal reader's hands (Chapter 21).
On the same stretch of ladder, being internal also adds a scheduling dependency. Level four requires the suggestion to appear in the work entry point he already uses, and inside a company that entry point is often the corporate portal or the core system, and the people who change those are the platform team, not you. You need two things from them, a slot on the portal, and one release window of the core system that carries your module. Both come on the platform team's own iteration schedule, usually counted in quarters, so ask one iteration cycle before your launch date. Do not wait until the pilot is running to remember.
The level six review cadence is actually easier to keep alive long term from the inside. An external deliverer is gone by the exit date, and the review cadence leaves with him. You are still here. You can hang the action on a regular meeting the business side already holds, making it a line on the agenda rather than an extra step someone has to remember. Once it hangs there, the review's owner is whoever chairs that meeting.
## Framework Two: Queue Design Patterns
Turn the upper half of the ladder into a table and you have the action queue, the productionized version of the Chapter 0 two-hour prototype ([Template 0.3](../appendices/template-00-field-mvp-pack.md)), with every column able to state its origin.
**Column structure = the projection of the four layers of the decision rights boundary (Chapter 8).**
| Decision Rights Layer | Columns in the Queue | Who Is Responsible |
|----------|-----------|--------|
| Sense | Claim ID, reason stuck, missing item, days waiting, risk signal, status-in-doubt flag | AI summarizes facts |
| Advise | Suggested priority, suggested next action, reason | AI stops here |
| Act | Owner, Human Call | People take over here. Without confirmation the system takes no action |
| Decide | No columns on the table for this layer | The payout red line. The table has no columns for this layer, and the blank itself is the red line (Chapter 8) |
**Three columns = the final form of the three oversight questions (Chapter 12).** The three questions. Is there time to look? The ability to judge? The authority to stop it? They land on the queue as three things. The Human Call column. If she says no, does it count? It counts. That is authority. The reason column. The information for judging right or wrong is in the same row. That is ability. Risk ranking plus a daily volume cap, with the number of rows entering human view derived backward from the review time budget. That is time, and it lives in the suggested priority column, which keeps the daily count within budget. Oversight ends up as three columns on a table. The boxes on the flowchart are only its shadow.
**Trail schema = the answer to that question from Chapter 9.** "Do fields written back by AI count as truth?" No, and they must never get the chance to. The schema (the field structure of the trail table) cuts cleanly, with four rules.
- Suggestions and facts are stored in separate tables. The fact table (the merged view and source system fields) records only the world and what people did.
- The suggestion table is append-only, never updated. Not one word of AI output is written back to the source.
- Every suggestion gets one trail row, matching the six decision trail fields of Chapter 12 item for item. The six are input snapshot and rule/model version, suggestion and reason, Human Call with the decider's real name, and timestamp.
- On override, one more field, a reason code, a few enumerated values plus an optional note, chosen in two seconds. Anchor & Helm's version has six. Risk judgment differs, priority judgment differs, information outdated, already handled offline, suggested action not feasible, other. The reason code is the easiest to skip and the one that must not be skipped. It is the entrance to level six.
## At Anchor & Helm: The Big Screen, the Fourteen Steps, and the Decision Trail's First Catch
### The Big Screen, Translated Live
Replay the three questions from Friday of week 9.
"Kevin, the big screen is up. You see the auto exceptions backlog rising. Then what?"
"Have the team lead handle it."
"How does it get to the team lead? Do you call her, or is she watching the screen too?"
"The system could alert her."
"The alert arrives. How does she know which one to handle first?"
Kevin laughed. "You are describing that queue of yours."
Three questions in, the big screen ask took itself apart. For a backlog number to become an action, there has to be an owner (the team lead), a next step (which claim to move first), and a basis (the reason). The queue has all three, and they already live in Linda's team's work entry point. What the big screen wanted to do, the queue's first four levels have done. The three questions are a field variant of Chapter 6's three-layer probing method. There it chases expert judgment, here it chases the action chain behind an ask, the same craft.
The conversation did not end there. Kevin's ask had one real core left. He wants to know the system is being managed. To have the numbers in his head, and an answer when Grant asks. That need is entirely legitimate, and meeting it takes no big screen. Three things, all much cheaper, are enough.
1. **A rollup view**, a page of numbers summed by team, line of business, and aging, where every number clicks through to the queue itself. It is only ever an entrance to the queue, never a parallel screen (not the same thing as the merged view in Chapter 9's data layer, one faces up, the other faces down).
2. **Exception escalation rules**. Queue aging past a threshold, an unsafe interception event, a lengthening status-in-doubt list. Three kinds of signal escalate to Kevin automatically, and the rest of the time nothing bothers him.
3. **A weekly report**, one page watching the North Star and the process metrics (the metric tree in Chapter 18 gives it a formal skeleton).
Kevin listened and said, "Fine, better than a big screen. A big screen I have to watch myself. This one comes to me." He may not have realized the weight of that sentence. "It comes to me" is the whole difference between visibility and task creation.
### The Fourteen Steps Revisited, Which the Queue Eats and Which It Leaves
The fourteen real steps counted from a folding stool in Chapter 6 are now the map of the queue's embedding points. Before the pilot starts, you and Linda rule on every step again. Absorb, transform, or keep?
| Original Step (Chapter 6's Fourteen) | Where It Goes |
|--------------------------|------|
| 1 scan new claims / 2 enter into Excel / 3 sort today's order by color | **Absorbed**. Claims enter the queue automatically, risk ranking replaces the color codes |
| 5 search the inbox for documents | **Absorbed**. The inbox extraction signal prompts "documents arrived" |
| 13 Tuesday and Thursday filter of claims over seven days, chase in the group chat | **Absorbed**. Queue aging escalates automatically. Two manual filters a week become one automatic watch every day |
| 4 check the document list | **Transformed**. The system pre-generates the missing list, a person reviews it (errors and omissions are concern class, Chapter 11) |
| 6 attach to the system / 8 send chase email | **Transformed**. The system gives the lead and drafts the text. Attaching and sending are done by a person, and the red line of no automatic outbound messages still stands (Chapter 0) |
| 9 amount scan / 11 flag risk red | **Transformed**. Five rules + the long tail suggest a red flag and give a reason, and the final call goes through the Human Call column. The word or two only the team understood, written beside the risk column in step 11 of Chapter 6, became the ancestor of the reason code |
| 12 change status | **Transformed**. The Human Call is captured in the trail automatically. Writing back to the core system is still a ten-minute daily operating discipline (Chapter 9, owner Kevin) |
| 7 / 10 call the surveyor | **Kept**. Judgment and relationships, the system does not touch them |
| 14 back up Excel before leaving | **Kept**. Linda still backs up. The day she stops on her own is the true measure of trust (Chapter 21) |
Five steps absorbed, six transformed, three kept. Everything absorbed is information hauling. Everything kept is judgment and relationships. That is the workflow projection of Chapter 0's "AI does the grunt work. People make the calls." Excel is not retired. In the early pilot it is still the source of truth for claim status (Chapter 9). Only after the week 13 retest passes does the merged view stop deferring to it, and the three columns she added to Excel herself (actual status, chase count, risk mark, Chapter 6) move into the queue as the trail. The warning that "the system will kill its own source of truth with its own hands" is guarded during the pilot by the status-in-doubt list.
### Week 13, the Decision Trail's First Catch
At the end of pilot week 3 (Monday of project week 13), with the cost scare just put out (Chapter 16), you review three weeks of decision trail for the first time. Overall override rate, somewhere in the teens. Split by suggestion category, one thorn jumps out. High-priority suggestions triggered by "prior claim linkage" were overridden by Linda's team 78% of the time. Ten times the system said "risky, escalate for review," eight times it was pushed back.
The reason code distribution saved you half a day. Of the 78%, the great majority picked "risk judgment differs," and the notes kept repeating the same phrase, repair shop filing for the customer. Two hours of tracing back. Annotation guide section 4.2 (Chapter 11) had two layers. A repeated phone number counts as a linkage signal, but a repair shop filing on behalf of customers is an exception scenario, and the same number reporting for different customers is not high risk. In implementation the first layer made it into code, and the exception line was missed. So every repair shop that used its own number to file for customers was treated as a fraud ring.
Why did eval not catch it? Of the fifty golden cases, only one had this shape, and it happened to overlap with other risk signals, so "high priority" counted as correct. Eval was not wrong. The ruler was not bent. This material was never put on it.
> **Override data is the last net for the errors eval cannot catch.**
The fix is one line of a condition. Claims with the wrong shape go into the golden cases, and "incidents go onto the ruler" becomes a fixed process in Chapter 18. The mechanism deserves the record more than the fix. This error set off no alarm, drew no customer complaint, and showed on no monitoring chart. It was fished out by Linda's team's two-second reason codes. Victor Reyes read the page of analysis and replied with one line, "The trail caught something live for the first time. That week on constraint engineering paid for itself."
## Failure Modes
**1. Insight without action, delivering at level one.** The project closes with "the management cockpit is live," and three months later nobody opens it. The dashboard is a conspiracy of supply and demand. The reporting chain wants "presentable progress," the engineering side wants "a deliverable nobody is accountable for" (a display cannot be wrong, a suggestion can), and its success metric can always be made positive.
**2. Recommendation without owner, suggestions left hanging.** The system generates dozens of "worth attention" items a day, every one reasonable, nobody claims them. Assigning an owner is an organizational negotiation, not an engineering task. The person writing code has no authority to claim responsibility on someone else's behalf, so the owner column defaults to blank, and a blank is painless in a demo (the presenter is the temporary owner). A suggestion with no responsible person is a decorated insight, not a suggestion. The owner column's legitimacy comes from the SOP and job responsibilities, not from the system and not from you.
The next four are pits you fall into after the queue is up.
**3. Five tools stitched into one workflow.** Look in system A, act in system B, record in spreadsheet C, ask for the reason in group chat D. Each existing system has its own owner and its own cost of change, and "open a new interface" is always cheaper than "embed in the old entry point." The cost shifts from the project to the user, and every system switch drops usage by half. One test. How many systems does completing one suggestion touch? More than two and you are bleeding users.
**4. Alert fatigue, everything important equals nothing important.** Escalation rules keep getting added, the manager receives 40 "important alerts" a day, and from week 3 on he swipes them all away. Same structure as the scope creep of Chapter 8. Everyone has the power to add an entry to "important," nobody has the power to remove one. Adding one is free for the proposer, and the recipient pays the whole cost. The fix is a budget. First set how many the recipient is willing to look at per day (the time question of the three), then derive the escalation thresholds backward.
**5. Not recording the override reason, the loop breaks at the last link.** The trail table has accept/override and no why. For the front line, writing the reason is pure expense, and the benefit goes to the system. Without design, the default is nobody writes. A free-text box equals forcing the front line to give up writing. Without reason codes, 78% tells you only that the system is not trusted, not which rule is wrong.
**6. Stopping at level five, a trail nobody reviews.** Reason codes get filled in for months, and the trail table is never opened once. The front line paid the cost as designed, and nobody is at the receiving end. Review needs an owner and a cadence, neither of which writing code can produce, so it is the easiest thing to postpone indefinitely after launch. Anchor & Helm's first catch came from the act of reviewing, not from the trail table's existence. The test. Ask on what date the trail last changed a rule or the golden cases. No date, and the ladder stops at level five.
!!! note "Frontier Sketch: The Ladder Is Being Climbed"
The act layer will not stay empty forever. A team building autonomous software engineering agents shared one test in 2026 (paraphrased). How long an agent can run unattended in a given step depends on the density of deterministic checks. The denser the loops that settle right or wrong on the spot, linters (code checkers), type checks, automated tests, the farther an agent goes without a human stepping in, and their internal processes with the clearest error shapes are already fully autonomous. This is the same principle as "put the LLM where its output can be checked" (Chapter 10). The upper half of the ladder is paved with verification loops, not with smarter models. The two forms of evidence for moving up are in Chapter 8, and Anchor & Helm's first move up is in Chapter 22.
## Next Monday
1. Put the last dashboard you delivered on the six-level ladder, and place it with one question. "Who changed an action last week because of it?" No name, and it is at level one.
2. Pick the metric you display most often and run the three-question translation. You see the change, then what? Who handles it? How does he know which to handle first? Write the answers as three items, owner, next action, reason. If you cannot fill all three, what you delivered is insight.
3. Check whether your system records override reasons. If not, add reason codes this week. No more than seven values, plus an optional note ([Template 17.2](../appendices/template-17-action-queue.md)).
4. If you already have a trail, split the override rate by suggestion category once, and pick the category with the highest override rate to trace back the rule implementation. Odds are an error eval never caught is waiting for you.
**Want an agent to get you started?** In the repo you set up following [Start Here](../index.md), paste this to your coding agent:
```text
In the repo/ directory of the the-last-mile repository, help me with the Chapter 17 Next Monday actions. First run python3 templates/action-queue/run_report.py
with the built-in sample to show me an override weekly report, then open schema.sql, reason-codes.json, and weekly-override-report.sql and explain each one.
Then I will name the metric displayed most often on the dashboard I delivered recently. You ask only the three questions, you see the change then what, who handles it,
how does he know which to handle first, and the three answers are mine to write. The reason code values are mine to define, no more than seven, and you check the count and mutual exclusivity.
If any command errors, stop and show me the output.
```
---
## Chapter Kit
- **Judgment frameworks.** The six-level action integration ladder (visibility โ prioritization โ recommendation โ task creation โ decision capture โ closed loop; the higher you go, the less engineering and the more organization); queue design patterns (column structure = the projection of the four layers of the decision rights boundary; three columns = the final form of the three oversight questions; the six decision trail fields + reason code, suggestions and facts stored separately)
- **Templates.** [Template 17](../appendices/template-17-action-queue.md), Action Queue Design Patterns, Decision Trail Schema, Before/After Workflow Map
- **Key judgments**
- "A system's value is not in what it knows. It is in whose next action it changes."
- "A dashboard leaves the four organizational things to the viewer to solve on his own, so the viewer chooses not to solve them."
- "Override data is the last net for the errors eval cannot catch."
- "Everything absorbed is information hauling. Everything kept is judgment and relationships."
---
# 18 ยท Launch and Measure: Prove the Gain, Not the Trend
!!! info "Companion Templates"
๐ [Chapter Template](../appendices/template-18-metric-tree.md) ยท ๐ [Template Library](../appendices/template-library-index.md)
> **The Challenge.** The system is live and the numbers are moving. How do you get everyone to believe the system deserves the credit? An incident happens. How do you keep one error from eating months of banked trust?
>
> **What You Will Be Able to Do.** Use the three layers of a metric tree to turn "the system is improving things" into defensible evidence. Use an eight-week baseline plus a run chart to tell fluctuation from change. Write the AI incident runbook before the incident happens, so the first incident is handled by process, not by mood.
---
## Pilot Week 5, Two Pieces of News in the Same Week
It is pilot week 5, project week 14. On Monday the weekly numbers come out. First-touch handling time for auto exceptions is down 18% against the pre-launch baseline. The monthly report to Grant Whitmore is booked for Friday, and that number was going to be the star.
The launch itself went smoothly. The pilot started on Monday of week 10. Before it started, the security review packet from Chapter 12 was updated to a launch version and served a second time, the same document used once more, with monitoring definitions and a rollback path (rehearsed once) added, signed by Victor Reyes, and filed as launch gate material.
On Wednesday at 10:40 in the morning, something else happened. The queue suggested one claim as routine. Linda Marsh glanced at it, overrode it in the Human Call column, picked the reason code "risk judgment differs," and wrote two words in the note. Repair shop. At the system level this is one override. At the organizational level it is a bomb. By eleven, the claims department was passing around "the AI waved a fraudulent claim through as routine." By lunch the version had evolved into "the AI got it wrong," and the rumor had dropped the half sentence "and Linda caught it."
So you are carrying two things into Friday's meeting room. A proof problem for the good news. On what grounds do you say the -18% is the system's doing? A trust problem for the bad news. How do you keep one error, one that was caught, from defining this system's reputation? These two things are the whole of the work in the launch period.
## Why This Is Hard: Numbers Fluctuate, and Memory Is Asymmetric
Proving is hard because operating numbers fluctuate by nature. The mix of claims shifts, people rotate, seasons move, and first-touch handling time rises and falls every week. Between "it went down after launch" and "it went down because of the launch" lies the whole of statistics. If you do not deal with that, sooner or later someone will deal with it for you. The week the number bounces back, someone will say, "See, it never had anything to do with the system."
Trust is hard because organizations have an asymmetric memory for AI incidents. An 18% improvement is remembered for three days, one incident for a year. The reason is structural. An incident is a story. It has a claim number, characters, and a "that was close" plot, and it can be retold in full over lunch. An improvement is a statistic, a curve drifting slowly downward, with no plot, and nobody retells it. People spread stories. They do not spread distributions.
So the real work of the launch period is two things. Turn the improvement into defensible evidence, and turn the incident into a recoverable process. The first relies on measurement discipline, the second on a runbook written in advance (the procedure manual for handling incidents, unfolded in framework two). The common enemy of both is improvisation.
## Prior Art, and What AI Changed
**The discipline of measuring improvement comes from healthcare quality improvement.** The tool first. A run chart is a chart of points plotted in time order with the baseline median as a reference line (you saw one in Chapter 14). *The Health Care Data Guide* (Provost and Murray, paraphrased here) faces a problem with exactly your shape. Clinical metrics fluctuate by nature, and "this improvement saved lives" has to be defensible. Their method uses two tools. The run chart is one. The other is a pair of distinctions, common cause variation (the random rise and fall built into the system) and special cause change (the signal of a structural change, for example six consecutive points on the same side of the median). The discipline is one sentence. Build the baseline first, and claim improvement only when a special cause signal appears. A line drawn between two points gives you no trend, only an illusion of slope.
**The discipline of incident handling comes from SRE (site reliability engineering, an operations discipline).** Tiered response, notice within a time limit, blameless retrospectives. An incident is a learning asset, and shame does not help, provided the incident is caught by a process rather than by emotions.
The AI era changed two things. First, monitoring gained a layer specific to AI, the monitoring surface. Most of it is what earlier chapters already settled, carried over as it is. The override rate and its sudden-change alerts. The error category distribution taken online (the threshold table of Chapter 11 becomes the monitoring definitions as is). Golden cases replayed on a cadence. Drift signals (the signs of the system's results quietly degrading over time, the minimum observability set of Chapter 16). The fuel all comes from the decision trail of Chapter 17.
Second, the incident retrospective gained a question that did not exist before. Was the error in the model, the data, the rules, or the interaction? The four layers have completely different repair actions. Attribute it to the wrong layer and the repair fixes the wrong place.
Together, these two are the same defense in a different state. AI uncertainty management, the one skill this method adds of its own (the single new skill Chapter 2 named), was delivered by Chapter 11 in its design state, the defense drawn on the blueprint and in the eval, with nobody yet watching it every day. This chapter delivers its operating state, the same defense standing watch in production through the monitoring surface and the runbook.
## Framework One, the Metric Tree Plus Baseline Discipline
A metric tree splits "is the system improving things" into three layers, with every metric carrying an owner and an action. The table below runs through the three layers, North Star, process, and balancing. The process layer has the most mechanisms and involves the most people, so that row is the longest.
| Layer | Question It Answers | Anchor & Helm Instance | Owner | Action on Deviation |
|----|-----------|----------|-------|----------|
| **North Star (outcome)** | Did the improvement happen | First-touch handling time (the charter North Star, Chapter 4), shown on a run chart | You + Kevin Doyle | Two consecutive points back above the baseline median โ go through the process layer for the cause |
| **Process (mechanism)** | Why the improvement happened, where the lever is | Queue aging (count of overdue claims not moved), missing-document detection lead time, override rate (split by suggestion category), status-in-doubt list length (Chapter 9's operating item, owner Kevin) | Linda Marsh (team lead), your team's engineer (renamed line by line at handoff), you, Kevin, in the same order as the left column | Over the line โ claimed and cleared the same day; lead time shrinks โ check the extraction pipeline; a sudden change in one category โ trace back that category's rule implementation and list version; continuous growth โ re-check the review team's ten-minutes-a-day status cleanup discipline |
| **Balancing (cost)** | Did the improvement shift the cost | Complaint rate, reviewer overtime hours | Kevin | Rising โ check chase scripts and frequency; rising โ check the queue's daily volume setting |
Three rules. **First, there is one North Star, and it comes from the charter.** No setting up a better-looking one after launch. **Second, every metric must have an owner and an action.** Who does what when the metric moves is written next to the metric. A metric with no owner and no action is only scenery. **Third, balancing metrics are shown on the same page as outcome metrics.** The balancing layer guards against pressing the problem down here and having it pop up over there. If handling time fell because every reviewer worked an extra hour a day, or because chasing drove customers up the wall, that is shifting the cost, and it does not count as improvement. The victims of a shifted cost are usually not in the reporting meeting. Only the metric speaks for them.
The second rule has a hurdle inside a company. Most process layer owners do not report to you. Queue aging belongs to Linda, the status-in-doubt list and both balancing metrics belong to Kevin, and their schedules are set by Claims Operations, not by you. So actions like "claimed the same day when over the line" cannot hang on your reminders. Get the whole metric tree onto the business department's own weekly meeting agenda, and get the claiming actions written into their team SOP. An action written into the SOP is their job at review time. An action sitting in your email is only your request.
The override rate in the process layer deserves its own mention. It is the last net for the errors eval cannot catch. Chapter 17 already showed it, when the 78% override rate on "prior claim linkage" suggestions exposed a condition the implementation had missed.
One more item hangs on the monitoring surface, the fairness spot check pre-planted in Chapter 12, the sentence written into the fairness row of the trust constraint matrix. The override distribution is spot-checked by customer segment every quarter, with the first check scheduled for the end of the first quarter after launch. It guards against the system systematically treating one class of customers worse with nobody noticing.
**Baseline discipline** in three sentences. The baseline is built before launch. Anchor & Helm used eight weeks of history, corrected claim by claim through reconciliation (Chapter 9's reconciliation paying off a second time, which is to say it gets used once more; without the correction, the baseline itself stands on distorted status fields). The run chart is kept up continuously, with the median as one reference line. Improvement is claimed only when a special cause signal appears.
In the launch period, the internal deliverer has three advantages to take, and they only count once written down as actions. The baseline does not have to wait for the business side to export data for you. Pull the eight weeks of history from the data warehouse yourself, and a median line is up the same day, so "nobody knows how slow it used to be" finds no excuse inside a company. You can hear the rumor first. On the day of the incident, standing in the claims department's work area for ten minutes gets you there half a day before it reaches your reporting line. The notice has to be drafted within two hours, and the calm of those two hours is bought with that half day. The run chart does not need a meeting of your own. Hang it on the claims line's existing business review as a fixed agenda item. That works better than chasing people to look at a chart every month.
## Framework Two, the AI Incident Runbook
The AI incident runbook is the handling process written before the incident happens. Levels, notice, retrospective, flow-back. The word "before" is the whole point, for the same reason as the kill criteria of Chapter 14. The process is written while calm. It cannot be written on the day of the incident.
**Levels.** P1 to P3 are response levels. They set how soon there must be action and who must be alerted, and they govern how fast someone takes charge. They are a separate numbering from unsafe / concern / useless, which governs how severe the error is. An unsafe-class error entering human view is P1, and it runs the full course of same-day notice plus a time-limited retrospective. Being caught does not downgrade it, because what caught it was already the last line of defense. Concern-class errors over threshold or appearing in batches are P2, retrospective the same week. A rising trend in useless-class errors is P3, folded into the monthly retrospective.
**The same-day notice** has three parts, and the order cannot change. Part one, the defense held. State first that the interception mechanism worked as designed, then the error itself. Part two, the scope of impact. How many claims, and whether there were real consequences. Part three, root cause under investigation, with a deadline for the retrospective. Drafted within two hours, sent the same day, with the distribution drawn by "how far the rumor can reach." Inside a company that line has to be drawn wider. You lack the buffer an outsider has. The person sending the notice is the person being talked about, and if you do not send it, someone sends it for you. Putting the defense first is the correct order of the facts, and has nothing to do with PR spin. This system's design premise is that AI will make mistakes and the Human Call column backstops them (Chapter 12's three oversight questions, Chapter 17's queue design). The incident is precisely the moment the design is validated.
Who issues the notice is a separate question inside a company. You hold the drafting right. The business-side owner holds the issuing right, and at Anchor & Helm this notice was signed by Kevin Doyle. The reason is the same as for the governance pledge of Chapter 20 (that chapter goes into detail). Only the person with the power to violate it can issue it. Your signature does not count. Send it out alone and the three parts read as the technical team defending itself. With Kevin's name at the bottom, the sentence in part one, that the defense worked as designed, becomes the claims line's own judgment. The price is that he has the right to edit the draft. The one thing you hold the line on is that the order cannot change. Move the scope of impact ahead of the defense, and this notice turns from fact into spin.
**The four retrospective questions.** The investigation runs in a fixed order. Was the model layer wrong (long-tail judgment falling short)? Was the rule layer wrong (a rule missing, or the implementation drifted)? Was the interaction layer wrong (a person saw it but had no time, no basis, or no authority to stop it)? Was the data layer wrong (input, list, or fields incorrect)? The first three questions are all elimination. Only when you have asked your way down and none of the three layers was wrong does the data layer come up, and by then the three eliminated layers are the evidence for the attribution. Wherever the fault is located, that is where the repair goes. Model layer, switch pattern or add review. Rule layer, change the rule and release a new version. Interaction layer, change the queue design. Data layer, fix the data and add operating discipline.
**Flow-back.** The incident case goes into the golden cases permanently within 24 hours (the historical incident category of Chapter 11, where the flow-back mechanism of the human review path left the interface ready long ago). Incidents go onto the ruler, meaning that once in the eval they become a lasting test standard, and from here on that is a fixed process. The kill criteria are checked by the book and recorded, whether or not they trigger.
Flow-back has to land on names. The easiest slip inside a company is for it to stall at "incident handling belongs to operations, eval belongs to development." The incident record closes on the claims line, and not one golden case gets added on your side. Anchor & Helm wrote this item with real names. Linda rules at the retrospective on which category the incident case goes into, you add it to the golden cases the same day and countersign in the incident record, and if either name is missing the loop is not closed. After you step out of the daily, this item is handed over under the three AI items going platform-level (Chapter 22 goes into how), and what receives it is the platform's eval capability, not "someone will always look after it."
## At Anchor & Helm: Forty-Eight Hours of One Incident, and One Honest Run Chart
**Wednesday, 12:10, the notice goes out**, less than two hours after Linda's override, to the whole claims department, Kevin, Grant Whitmore, and Victor Reyes.
> This morning the queue suggested a high-risk claim as routine. The review team lead caught it in the Human Call column and overrode it to high risk. That column exists to catch exactly this kind of error. The defense worked as designed, and the full trail is on record. Scope of impact. No real action was taken on the claim. Today's queue has been checked, and there is no error of the same shape. Root cause under investigation, retrospective conclusion within 48 hours.
In the afternoon the rumor changed versions. "The AI got it wrong" became "the AI got it wrong, but Linda caught it, and that is how the system was designed." Which version of an incident the organization remembers is set in the first two hours. Let the notice run a day behind the rumor and you spend a quarter correcting that version.
**Thursday, the retrospective.** You, Linda, Kevin, and the claims-ops IT engineer go through the four questions. The model layer? The long-tail judgment really was wrong, but the amount sat right at the top of the usual range, the claim was reported the same day, the photos were complete, and there was no prior claim linkage. The input held no usable signal, and a stronger model would still be guessing. The rule layer? The implementation of the five rules was checked line by line against the annotation guide. No drift. The interaction layer? Linda had the time, the reason column, and the authority. All three questions pass, and the successful catch is itself the evidence. The data layer? Hit. That repair shop registered only last month. Linda said at the meeting, "The address is right next door to the one I crossed off. The owner changed the name. How is my list supposed to keep up with that?"
In Chapter 6 she already said it, "it is not a fixed list. I crossed one off just last month." Chapter 11's annotation guide got its dated version because of that sentence. Back then it was a maintenance cost. Now it is the prophecy of an incident's root cause. The list drifts, a fact this system knew from day one and had not yet scheduled into operations.
Attribution decides the repair path. The retrospective locates the data layer, and the repair action follows. The list is promoted from Linda's rolling mental list to an operating asset with an owner. Her team updates it monthly, it goes into the annotation guide's dated version, and it becomes a config item the claims-ops IT engineer can change (the maintainability row of Chapter 15). The queue adds a weak-signal prompt for newly registered repair shops not on the list. Not high risk, only a prompt, "new entity, not on the list."
One more thing lands without waiting for the repair. The ADR memo to Grant Whitmore promised in black and white (Chapter 13) that if a single missed-risk case appears during the pilot, the whole category of suggestions is downgraded to human review. That week it is executed by the book. All risk-class suggestions go to human review until the repair passes a golden cases replay.
The incident case goes into the golden cases that day, kept permanently. Then, in front of everyone at the retrospective, you do something some people think unnecessary. You pull out the kill criteria written in Chapter 14. Two or more unsafe-class errors in a single week means pause and rework. This week, one. Not triggered, keep running, and the check goes into the incident record. "Why check if it did not trigger?" Check even when it does not trigger. The value of checking is in the act of checking every time. Each time the stop clause is executed in earnest is the only proof that it has teeth.
**Friday, the report.** The one page for Grant Whitmore is a run chart. An eight-week baseline, the median as one line, and all five pilot weekly points below the median, an average drop of 18%. On the chart the baseline points scatter above and below the reference line, and the pilot points are all below it, still one point short of a signal.
```
First-touch handling time, weekly median, all points
Baseline, eight weeks Pilot, five weeks
ยท ยท ยท ยท baseline points, above the median
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโ baseline median, the reference line
ยท ยท ยท ยท baseline points, below the median
ยท ยท ยท ยท ยท pilot weekly points, all five below the line
โ next week's sixth point, a signal only on the same side
```
You told him three things honestly. First, by the rules of improvement measurement, five consecutive points are one short of a shift signal. Only if next week's sixth point is still below the line does this 18% stand. Second, there is still a gap to the charter's -30%, and the next step on the gap is chasing missing documents earlier. The process layer's detection lead time still has visible room, and that is next month's main push. Third, one unsafe-class incident occurred this week. Caught, reviewed, added to the golden cases, kill criteria checked and not triggered.
Grant looked at the chart for a long time and said, "The last person who told me 'one point short does not count' was Audit. Keep running. Tell me the day the sixth point comes in." Then he added two operations staff on the chase side for Kevin. More trust, not blame. Those two people did not appear out of nowhere. What Grant could settle on the spot was only a move within Claims Operations, shifting people from another team to the chase side with total headcount unchanged. Adding two actual heads waits for the annual headcount review, a round of request and approval, slower than you would think. Resource commitments inside a company come in these two kinds. The one honored on the spot is moving people. The one that queues is adding headcount. When you hear a commitment, ask which kind it is first.
This is not luck. This is compound interest. The reliability term in the Trust Equation's numerator is built up one "said it, did it" at a time (Chapter 5). The harder backing comes from the charter. Nobody suspects you of dressing things up, because the clause that takes the resources back has been on the table all along, the charter's resource reassessment conditions (Chapter 4), and like the kill criteria checked by the book on Wednesday, it was gone through by the book this week and did not trigger. The thirty seconds of awkward silence when the exit conditions were written in Chapter 4 paid off all their interest at this moment.
This run chart and this retrospective record will not be used only once. They are ready-made raw material for the impact memo of Chapter 19 (that chapter goes into detail). The evidence you need at closeout is already banked this week.
## Failure Modes
Four, matching the four links of reporting, incident handling, balancing metrics, and attribution.
**1. Screenshot reporting.** Pick the week with the best-looking number, make it a slide, and leave out the weeks it bounced back. The reporting chain favors single-point good news at every level, and every retelling deletes some uncertainty. But trust is indivisible. Get caught once and every number loses credibility with it, including the true ones. The run chart is the antidote. Every point is on the chart, so you have no way to pick, and no way to be accused of picking.
**2. Incident silence.** Something goes wrong and it is digested internally first, "let's get to the bottom of it before we say anything," then a passive response three days later. The cost of the notice is immediate and concrete, and the risk of silence is delayed and probabilistic, the same time-discounting trap as Chapter 12's "our own tool" escape (quietly going live to dodge review). A small certain cost now outweighs a large probabilistic cost later. But rumor runs ten times faster than fact, and rumor always picks the worst version. Every day you stay silent is a day campaigning for the "AI is unreliable" narrative.
**3. Reporting outcomes only, ignoring balancing metrics.** Handling time fell and there is a big celebration, and the complaint rate and overtime hours are not in the monitoring at all. Balancing metrics measure the cost someone else pays for you, and the victims, the customers chased up the wall, the reviewers working overtime, are not in the reporting meeting. A shifted cost is invisible on every individual report. Three months later the complaints pile up into an event, and the improvement is given back with interest.
**4. Attributing to the wrong layer.** Something goes wrong and the model is suspected first, and a data layer problem gets repaired as a model layer problem. Attribute it to "the model is not good enough" and the next three weeks go to switching models, tuning prompts, and adding review gates, none of which fixes the point. The model layer is the easiest to think of and the easiest to touch, and one line of config lets you announce "we are fixing it." A data layer repair means finding an owner, setting an update cadence, and changing operating discipline. Slow, and it does not look like technical work. The four layers' repair actions do not transfer, and misjudge the layer and every hour is lost. The test. Write down the repair action from the last incident, see which layer it falls in, and check whether that is the same layer the retrospective record ruled on.
## Next Monday
1. Draw a metric tree for the system you have now. One North Star (from the charter), two to four process metrics, at least two balancing metrics, each with an owner and an action on deviation. If you cannot write a balancing metric, ask one question. Whom is this improvement most likely to shift its cost onto?
2. Check the baseline. If you are not live yet, start collecting baseline data today. If you are live with no baseline, start recording from today and say so honestly in the next report. An honest "no baseline" beats a fabricated comparison.
3. Use [Template 18](../appendices/template-18-metric-tree.md) to write a one-page AI incident runbook. What counts as P1, who the notice goes to, the four retrospective questions, how the incident case gets into the golden cases. Write it before the incident, like the kill criteria. It can be written while calm, and cannot be written on the day of the incident.
4. Run a drill. Take one wrong output from history, walk through the same-day notice template, and time it. The parts that run over are the parts to lock down in advance.
**Want an agent to get you started?** In the repo you set up following [Start Here](../index.md), paste this to your coding agent:
```text
In the repo/ directory of the the-last-mile repository, help me with the Chapter 18 Next Monday actions. First run python3 templates/metric-tree/run_chart.py
with the built-in sample to show me a run chart with the special cause marked. Then copy metric-tree.md to the working directory I name.
I will read you the North Star from the charter. The process and balancing metrics are mine to propose. You ask only for each one's owner and action on deviation, and mark what is missing [TBD].
If I cannot come up with a balancing metric, ask one question, whom is this improvement most likely to shift its cost onto. The P1 definition in the runbook is mine to write.
If any command errors, stop and show me the output.
```
---
## Chapter Kit
- **Judgment frameworks.** The three-layer metric tree (North Star outcome โ process mechanism โ balancing cost, every metric with an owner and an action); baseline discipline (eight-week baseline + run chart + claim improvement only on special cause); the AI incident runbook (unsafe is P1 / the three-part same-day notice / the four retrospective questions / the incident case goes into the golden cases permanently)
- **Templates.** [Template 18](../appendices/template-18-metric-tree.md), Metric Tree, AI Incident Runbook, Same-Day Notice
- **Key judgments**
- "Organizations have an asymmetric memory for AI incidents. An improvement is remembered for three days, an incident for a year."
- "The launch period has two jobs. Turn the improvement into evidence, and turn the incident into a process."
- "The first sentence of the incident story is that the defense held."
- "A metric with no owner and no action is only scenery."
- "Attribution decides the repair path. The retrospective asks which layer failed first, and how to fix it second."
---
# 19 ยท Talking to Executives: Three One-Page Memos
!!! info "Companion Templates"
๐ [Chapter Template](../appendices/template-19-memo-suite.md) ยท ๐ [Template Library](../appendices/template-library-index.md)
> **The Challenge.** An executive gives you ten minutes at a time. The project runs for months. How many times do you actually deal with him, and what do you say each time? Too early is an interruption, too late is losing control.
>
> **What You Will Be Able to Do.** Cover every contact between the deliverer and an executive with the three-memo system. Use a kickoff memo for authorization at the open, a decision memo for the call at the midpoint, an impact memo for renewed funding at the close, and let each one handle AI expectation management along the way.
---
> **Part V Navigation.** Chapters 19 to 22 run from the pilot's close to stepping out of the daily.
> The impact memo (Monday of week 18, this chapter) โ the annual budget and headcount review (Wednesday of week 18, this chapter and Chapter 22) โ the survey team lead pushes back (week 20, Chapter 20) โ three daily-active curves (week 21, Chapter 21) โ the gradual withdrawal and the five self-sufficiency tests (weeks 22 to 27, Chapter 22) โ the last retrospective (Friday of week 29, Chapter 22)
## Asking for Renewed Funding with a Number That Missed Target
The pilot closed out on Friday of pilot week 8, and the North Star stopped at -22%. The charter says -30%. Wednesday is the annual budget and headcount (staff positions) review, and whether the project expands, holds, or gets archived right there turns on the one page you send out on Monday. You have to trade a number that missed target for the next phase's investment.
Most people's instinct at this point is to explain. The people arrived late, the list fix landed late. But the harder you explain, the more it reads like a justification.
You spread everything you have sent Grant Whitmore so far across the table to get a feel for it, laid out in time order. The pre-mortem in week 1, the weekly report in week 6 that sank without a reply, the ADR memo in week 8, the run chart report in pilot week 5, and the one-line short notice the week after. Five pieces laid out, and the ones that actually did something all sit in the same structure, while the only one that drew no answer is the one written with the most effort and the most pages.
There is a pattern in it. Every contact between the deliverer and an executive comes down to three moments. Authorization at the open, a decision at the midpoint, renewed funding at the close. Three moments, three forms. This chapter grows Chapter 13's one page into a system that covers the whole project, then uses it to write Monday's memo.
## Why This Is Hard: Saying the Right Kind of Thing at the Wrong Moment
The common shape of failed executive communication is saying the right kind of thing at the wrong moment. Actually saying something wrong is rarer. At the open you talk implementation detail (he wants risk and commitments). At the midpoint you paint the vision (he wants the item to call). At the close you talk about how hard the work was (he wants the outcome and the next step). All of it is true, and all of it lands in the wrong moment. A mismatched form gets shelved and never even earns a rebuttal. An executive will not say "you wrote it wrong." He just does not reply.
You ran the controlled experiment yourself. Wednesday of week 6, with the reconciliation half done, you sent Grant the only "weekly progress report" email of the whole project. Five paragraphs of reconciliation progress, thorough and orderly, with one sentence that mattered buried at the end of paragraph three, "we propose taking the merged view as the source of truth for claim status, and will proceed on that basis unless there are objections." Zero replies. You told yourself silence meant consent.
At Friday's source of truth decision meeting, the merged view was settled on the spot with Kevin Doyle and Linda Marsh (Chapter 9). After the meeting you rewrote the request buried on Wednesday as half a page. One conclusion sentence (source of truth for claim status = the merged view), two lines of reasons, and one closing line, "This decides where the development effort goes next. Please confirm it for the record by Sunday." Kevin passed it up. The approval came back inside 48 hours, two words. "Proceed accordingly."
Same matter, same reader. The first time was five paragraphs of information and it sank without a reply. The second time was one decision request, closed out in 48 hours. The first email was not rejected. It was sorted. An executive's inbox sorts by "what do I have to do," and mail whose answer is "nothing" drops into white noise automatically. Executives do not reply to information. They reply to decision requests. This has nothing to do with arrogance. It is how the job is designed. His output is decisions, and what you hand him either joins that production line or joins no line at all.
## Prior Art, and What AI Changed
Chapter 13 already put up the structural skeleton. Open with SCQA (Situation, Complication, Question, Answer), then lead with the conclusion. That comes from the Minto pyramid, where every level is driven by the reader's question. Amazon's narrative memo tradition (*Working Backwards*) runs important meetings around one page of full sentences rather than slides. What this chapter adds is the system, promoting "one page" from a single tactic to a communication protocol that covers the whole project. When to send, which one to send, and what skeleton each one has.
The AI era changed two things. The first is in your favor. AI drafting takes "I cannot get it written" off the table. An SCQA draft, the evidence laid out, the charts tidied, all a matter of minutes. From here the one-page rule has one bottleneck left, thinking it through. Once the cost of writing has collapsed, sending a long and muddled email exposes the fact that you have not thought it through, and time stopped being an excuse a while ago.
The second is a new burden. Executive communication on an AI project carries an expectation management job of its own. Before your executive ever meets you, the media has already calibrated his AI expectations, calibrated them into a diode with two states, conducting or not conducting, and no scale in between. Either "AI can replace half a department" or "it is all a bubble, money burned for the noise." Both extremes are fatal. The first makes -22% look like failure. The second makes any added investment look like waste. So each of the three memos keeps one expectation management slot. The pre-mortem summary in the kickoff, the decision rights ladder in the decision memo, the honest gap in the impact memo (write the number that missed target exactly as it is, do not dress it up into "close to target"). Calibrate one memo at a time and acceptance day does not collapse.
## The Core Framework: The Three-Memo System
Monday of week 9, right after Grant Whitmore made the call (the pilot call in Chapter 13), he said the sentence this whole system starts from. "Everything you send me from now on gets written the way this page is written." Half of it is an order and half of it is a permission. From that day your one page carries default priority with him. Sorting it out later, you find that "the way this page is written" splits into three by moment.
> **The three-memo system. Every contact between the deliverer and an executive comes down to three moments, authorization at the open, a decision at the midpoint, renewed funding at the close. One form per moment, and all three share the same one-page rule.**
| Moment | Memo | His Question | Middle Section | The Anchor & Helm Piece |
|------|------|--------------|----------|----------|
| Authorization at the open | **kickoff memo** | What are you promising, what do you want me to commit, how will it die | The outcome promised + the investment needed + the pre-mortem summary | The week 19 expansion kickoff |
| A decision at the midpoint | **decision memo** | What am I calling, what does it cost, what happens when it goes wrong | Three pillars + the trade-off said out loud + the decision rights ladder + risk and backstop (Chapter 13's ADR as it stands) | The week 8 pilot ADR memo |
| Renewed funding at the close | **impact memo** | Was it worth it, how far short, which next step | Run chart evidence + the honest gap + the next-step options | The week 18 retrospective memo (the Monday after the pilot closed) |
Two new terms. **Kickoff memo**, the one page sent to the sponsor before a phase starts, putting on record the outcome promised, the investment needed, and the ways to die already rehearsed. Chapter 1's pre-mortem stops being a standalone document here and becomes the opening act of every phase. **Impact memo**, the one page sent to the person who calls it when a phase closes, presenting the result with running evidence, presenting the gap honestly, and giving next-step options with one of them recommended. The midpoint one is Chapter 13's ADR memo, and this chapter does not rewrite it.
The three share four rules.
- One page of body text, attachments unlimited.
- Lead with the conclusion, the pyramid apex test as before (Chapter 13). Who you want and what you want him to call, in one sentence. If you cannot write it, do not start writing.
- The ending always carries "what I need from you," even if it is as small as "hold ten minutes on your calendar."
- Send it early, delivered 48 hours before the meeting, so the meeting runs around the document.
The sameness and the difference are visible at a glance. The opening (SCQA) and the ending (the ask) are identical across all three, and the middles differ. The kickoff faces the future, the decision memo faces the present, and the impact memo faces the past plus one exit.
One more thing is easy to miss. The three memos link into a loop. The moment the impact memo's recommended option is approved is the Situation (S) of the next kickoff. However many phases the project has, that is how many times the loop turns.
Turning that loop inside a company needs recalibrating in four places. The first is the cost of reporting bad news. Your sponsor is most likely on your performance review chain too, or one layer away from it. Writing the gap section means admitting to the person who grades you that you did not make it, and the instinct is to write -22% as "close to target." Do not write that. An honest gap costs more inside a company, and it is worth more for exactly that reason. What it buys in one stroke is this person's default belief in every number you give him afterward, and that deposit is not for sale anywhere else.
The second is the nature of the ask. The "investment needed" line inside a company is usually not about making the business side honor what it already claimed. It is about asking the sponsor to require a peer department to send people. So the internal version splits it into two columns, who sends the people and who gives the word. Who sends the people takes real names and hours per week. Who gives the word takes whose mouth the sentence has to come out of before it counts, the sponsor or the other side's manager. The week 19 kickoff is split exactly that way. Split it and you see that every item landing in the "who gives the word" column is a political act, not a clause you can cite. You hold no authority over that department and can only rely on the sponsor to go persuade a peer manager, so before it goes into the memo, work out whose standing you are spending.
The third is timing. Your closeout calendar is not set by the project. It is set by the fiscal year. The impact memo's send date is worked backward from the day budget preparation starts, and a week late means your numbers can only make next year's pot. The kickoff is planted a quarter ahead, so that next year's headcount (that is, the staff positions) has a written origin before the headcount table freezes (once settled, nothing more can be added). And there is one more that only the internal version needs. A year after the system goes live, not one bill reminds anyone that it is still alive, and by then one page is your only tool for winning resources back on a regular basis. Written the way the impact memo is written, to the same recipient, except that this time nobody comes asking you for it.
The fourth is the copy line. All three memos go to the sponsor as the primary recipient, with three fixed copies, your manager Owen Hartley, the business-side owner Kevin Doyle, and the PMO. Owen's copy is not politeness. Your schedule and your review sit in his hands, and if he hears from someone else what you sent Grant, the next time your people get pulled away nobody will tell you first. Kevin's copy lets him see the ask before it goes out, since the people and the time you want mostly come through him. The PMO's copy keeps the project status matching the ledger. Going over a head has a real cost inside a company, and the copy line is the paperwork that turns going over a head into not going over one.
There is another half page only you have. The day before the memo goes out, hand it to Grant's assistant to read and ask what question has been chasing him this week. Rewrite the subject line on that answer, and reorder the first piece of evidence. Then pull two words out of the company strategy document into the conclusion paragraph, and find one comparable number from a similar system in another department for the evidence section. A consultant brought in from outside either does not have that assistant or cannot get those numbers. You have both.
## At Anchor & Helm: One Short Notice and Two Pages
**Monday of week 15, one line.** The weekly numbers come out, and the run chart's sixth weekly point falls below the baseline median. "Six points on one side," a special cause holds (Chapter 18), and -18% turns from "one point short" into a shift that stands up. That day you send Grant a one-line short notice with one chart attached.
> You said to tell you the day the sixth point comes in. It came in today, still below the median. Six points on one side, which by the rule is a structural shift, not fluctuation. Chart attached.
No request, no next step, nothing asked for. Ten minutes later the reply comes. "Noted." But it honored a promise made the previous Friday, and the reliability term in the Trust Equation's numerator is built up exactly this way (Chapter 5). Three weeks later you will see what that deposit is for.
**Monday of week 18, the impact memo.** The raw material is all ready-made, Chapter 18's run chart page, the incident retrospective record, and the kill criteria check records. Here it is in full.
---
> **To Grant Whitmore, from [you], Monday of week 18**
> **Subject. The exceptions queue pilot retrospective, for you to call the next step at Wednesday's annual budget and headcount review**
> *One page of body text. The full run chart, the incident retrospective record, and the metric tree detail are attached.*
>
> **Background.** The 8-week pilot closed out last Friday (auto exception claims, the eight people on Linda's team). **But**, the North Star finished at -22%, short of the charter's -30%. **So the question to answer**, was the pilot worth it, how does the shortfall get closed, and where does the next investment go.
>
> **Conclusion. The pilot validates, and I recommend approving an expansion to go deeper. Extend auto exception claims to two more review teams, and put the North Star on the -30% target line within eight weeks.**
>
> **Evidence (run chart attached)**
> 1. **The improvement is structural**, with all eight weekly points below the eight-week baseline median, and the day the sixth point landed it already formed the "six points on one side" signal. A shift, not fluctuation.
> 2. **The mechanism can be explained**, with most of the drop coming from chasing missing documents earlier and from clearing queue aging. The two chase-side people you added in week 14 and the repair shop list fix only entered the curve in the last three weeks.
> 3. **The defense has been tested in the field**, one unsafe-class incident during the pilot, caught in the Human Call column, same-day notice, root cause located within 48 hours (the data layer), and the incident case is in the golden cases. The kill criteria were checked by the book every week, zero triggers.
>
> **The gap.** *(this section is the expectation management slot)* -22% against -30%, 8 percentage points short. Two attributions, both verifiable. The chase-side people covered only half the stretch, and the weak-signal prompt for new entities not on the list went live only in week 15. Both levers are working and neither has run a full cycle yet, which is the basis for "-30% within eight weeks," and its whole basis. If the North Star stops falling four weeks into the expansion, this judgment is void and we go back to option three.
>
> **Three next-step options**
> 1. **Expand and go deeper (recommended)**, two review teams onboarded, the investment set out in next week's kickoff memo, using the validated mechanism to eat the remaining 8 percentage points.
> 2. **Start home property**, whose revival condition reads "the first extension after the pilot North Star hits target" (exclusive, the line written down in Chapter 8, meaning that once the target is hit it stands ahead of every other extension). The target line is -30% against -22% today, so the condition is not ripe and I do not recommend starting it now. The day the expansion hits target it goes first automatically, with no new project approval needed.
> 3. **Hold and watch**, no expansion and no withdrawal, four more weeks. Choose this one if you doubt that the last three weeks' drop will hold. The cost is stopping the scale-up just as both levers start working.
>
> **The two things I need from you**
> 1. **At Wednesday's annual budget and headcount review**, call one of the three options (we recommend 1).
> 2. **If the expansion is approved**, agree to start the responsibility transfer plan alongside it. Running and maintenance move to Kevin Doyle's team step by step through the expansion period, with the plan attached to next week's kickoff memo.
---
At Wednesday's annual budget and headcount review, two approvals come back. The expansion goes through as submitted, and the handoff plan is agreed to start. The meeting first goes through the graduation criteria (Chapter 14). Four of the five pass and only the North Star is missing, so what gets approved is the expansion plus an extension of the pilot rules, with the reassessment date set at week 26, not a conversion to formal project approval. Grant also pays a compliment at that meeting, and it pushes the handoff into the main subject (that sentence of his is picked up in Chapter 22). Record one detail only. He points at the "gap" section and says that had it been written as "expected to hit target soon," what he approved would have been option three. An honest gap, number against number, attributed to verifiable events, carrying its own void condition. It turned "missed target" from a justification into grounds for a decision, and the posture came along for free.
Of the three artifacts, the decision memo is not rewritten here, and its one-page structure and the full Anchor & Helm piece are in Chapter 13. One thing to add. Its expectation management slot is the decision rights ladder inside the trade-off section, and what an executive sees on that ladder is a boundary, not the AI in the media that can do anything.
**Monday of week 19, the expansion kickoff memo.** The loop turns back to the open, and this is exactly what goes out next once the expansion is approved.
---
> **To Grant Whitmore, from [you], Monday of week 19**
> **Subject. The expansion period starts, for the record. Promises, investment, and ways to die**
> *One page of body text. The handoff plan and the full pre-mortem are attached.*
>
> **Background.** The annual budget and headcount review has approved the expansion. **But**, the pilot's adoption was built on eight weeks of co-build with Linda's team, and the two teams coming in have neither Linda nor those eight weeks. **So the question to answer**, what do we promise in this phase, what do we need, and how is it most likely to die.
>
> **Conclusion. An eight-week expansion period, auto exception claims extended to two more review teams. We promise the North Star reaches the -30% target line, and one business win worth announcing inside the first month.**
>
> **Investment needed, who sends the people.** Kevin Doyle, schedule protection for the two new teams, 2 hours a week each on the precedent of Linda's team. Linda Marsh, taking part in transplanting her experience as the two teams come on, with the time agreed between her and Kevin. The two chase-side people continue to the end of the expansion period (already confirmed at the annual budget and headcount review, recorded here).
>
> **Investment needed, who gives the word.** The first two sit inside Kevin's own line, and his word is enough. The two on the chase side do not, so at the next monthly meeting please say a line to the chase side's manager about continuing them, with the review's own resolution as the wording.
>
> **Pre-mortem summary** (full text attached), the three most likely ways this dies.
> 1. The new teams have no Linda, daily actives start high and fall off, the system is "usable" and nobody uses it โ the defense is embedding the queue retrospective into both teams' existing morning standups, and announcing wins in business language.
> 2. The new teams accept everything as is, with none of Linda's kind of challenge, and the errors all land on the last line of defense โ the defense is that an override rate clearly below Linda's team's baseline in the first month triggers a spot check, because abnormally low is bad news too.
> 3. The larger claim volume overruns the monitoring surface, and alert fatigue drowns the real signal โ the defense is resetting the escalation rules' thresholds and recipients for the expanded volume.
>
> *(this section is the expectation management slot)*
>
> **The one thing I need from you.** This memo is for the record, no call needed. The only request is ten minutes at each monthly meeting, one look at the run chart in week 22 and one in week 26.
---
Put the three artifacts side by side, the week 19 kickoff, the week 8 ADR (full text in Chapter 13), and the week 18 impact memo. The same opening and the same ending, and three completely different middles. When you were writing the pre-mortem, writing the ADR, and reporting honestly, you did not know you were building a system. Grant's sentence strung them together, and later Kevin picked up the same way of writing (Chapter 22). That is how a system usually comes about. First a few pieces that did something, then someone names them, aligns them, and writes them down as rules. You do not have to wait for that step to happen on its own. Name the moment, pick the form, leave the ask, and in three moves the next one is written inside the system.
## Failure Modes
**1. Moment and form mismatched.** The opening memo is all architecture detail, and the closing memo is all war stories. The writer writes whatever he is most immersed in at the time. At the open your head is full of the plan, at the close you remember the overtime best. The form follows the writer's state of mind instead of the reader's decision moment. The fix is mechanical. Before writing, answer "which moment is this," and if you cannot answer, check the calendar, because phase boundaries are objective.
**2. Expectations left unmanaged.** The memo carries only numbers and requests and never calibrates "what AI can actually do." On acceptance day the executive measures your system against the AI in the media, and -22% gets read as "so much for AI." An executive's AI expectations are not a blank sheet. The media pre-calibrated them into a diode, and calibrating him is on nobody's task list. The fix is already sitting in the three expectation management slots, the pre-mortem summary, the decision rights ladder, and the honest gap. Calibrating once per memo beats one reconciliation on acceptance day.
**3. Showing up only when you want resources.** Zero contact outside the three memos, every appearance carrying a request, and the executive starts guarding the budget by reflex. Contact with no ask "produces nothing," so it is the first thing cut when you get busy. But honoring a promise mostly comes with no request attached. The short notice in week 15 asked for nothing. It only honored the line "tell me the day the sixth point comes in." Three weeks later the impact memo traded a number that missed target for renewed funding, and that short notice was the interest landing. One test. If every item in the communication record carries a request, you are the person who wants resources, not a partner.
## Next Monday
1. Spread out everything you have sent an executive on this project so far and label each one. Information, or a decision request? Compare the reply rates of the two kinds. That is your own controlled experiment.
2. Work out which of the three moments the project sits in right now (look at the calendar, not at how it feels), and check whether what you are writing matches that form.
3. Write the next memo with [Template 19](../appendices/template-19-memo-suite.md). Check one thing before sending, whether the ending carries "what I need from you" with a date.
4. Find a promise you made an executive and have not honored yet (the "tell you the day X shows up" kind) and set a reminder. The message that honors it carries no request at all.
**Want an agent to get you started?** In the repo you set up following [Start Here](../index.md), paste this to your coding agent:
```text
In the repo/ directory of the the-last-mile repository, help me with the Chapter 19 Next Monday actions. I will list everything I have sent
executives on this project so far. You only classify, information or decision request, and I fill in the reply rates. Then follow the instructions in
templates/memo-suite/prompt-memo-scaffold.md and draft the skeleton of the next memo for the moment I name (kickoff, decision, impact). The numbers
and the conclusion are mine to fill in, you invent none. Run python3 templates/memo-suite/remind_milestones.py and runchart_signal.py on the
built-in samples and tell me which moment each one is useful in. Before sending, check one thing only, whether the ending carries "what I need from you" with a date. If any command errors, stop and show me the output.
```
---
## Chapter Kit
- **Judgment frameworks.** The three-memo system (three moments ร three forms, kickoff at the open / decision at the midpoint / impact at the close, the loop closing back on itself); the four shared rules (one page, lead with the conclusion, an ask at the end, delivered 48 hours before the meeting); the three expectation management slots (pre-mortem summary / decision rights ladder / the honest gap)
- **Templates.** [Template 19](../appendices/template-19-memo-suite.md), Three-Memo Set (with ADR) (the kickoff and impact templates plus the Anchor & Helm examples; the decision memo is section 19.2 inside the same file)
- **Key judgments**
- "Executive communication has only three moments, authorization at the open, a decision at the midpoint, renewed funding at the close."
- "Executives do not reply to information. They reply to decision requests."
- "Three moments, three forms. Mix them and they stop working."
- "The evidence of communication is the other person's changed behavior, not your sent folder."
---
# 20 ยท Resistance Is a Signal: Decode It Before You Answer It
!!! info "Companion Templates"
๐ [Chapter Template](../appendices/template-20-resistance-decoder.md) ยท ๐ [Template Library](../appendices/template-library-index.md)
> **The Challenge.** Someone challenges you in front of the whole room. A team quietly stops cooperating. Are they being unreasonable, or did you miss something?
>
> **What You Will Be Able to Do.** When you are challenged in public, do not defend, name the emotion first. Use the resistance decoder to turn delay, sniping and silence back into interest, fear, pride or a design defect. Give "the system can be used to appraise people" a governance pledge written into policy.
---
## Week 20, Tuesday, the Claims Line Weekly Meeting
In week 17 the pilot closed out, first-touch handling time at -22%. The annual budget and headcount review in week 18 approved the expansion to go deeper. In week 19, auto review teams two and three came into the queue. It is now the second week after the expansion, that is, week 20, and Kevin Doyle is chairing the claims line weekly meeting. Item three on the agenda is the rollup view (Chapter 17), which gained a per-step dimension when the new teams came on, and this page of numbers summed by team and by step is in front of the whole claims line for the first time.
Among the bottleneck numbers broken out by step, the survey step is the sorest spot. Survey sits one step ahead of review, materials pass the surveyor before they reach the reviewer. The bulk of an exception claim's waiting time hangs on "waiting for documents," and "waiting for documents" is booked to the survey step. Kevin points at that page and says, "The claims on the survey side need tightening up."
The survey team lead puts his pen down. "Kevin, let me ask one thing first. This system of yours, is it here to appraise us? That BI system last time was yours too, and once it was built nobody looked after it."
The room goes quiet. Everyone looks at you.
You have three well-worn paths at this moment. Defend, lay out the facts, or go to Kevin after the meeting, or even to Grant Whitmore, and get the top to "align everyone's thinking." All three lead to a cliff, and this chapter's failure modes have one waiting for each. What they share is treating resistance as an obstacle to be cleared. This chapter is about the fourth path, treating it as a signal to be decoded.
## Why This Is Hard: Resistance Is Data, Not Noise
An engineer's default model of resistance is noise, irrational, unprofessional, something to overcome. This is one of the most expensive misreadings a deliverer can make. Resistance marks, precisely, where the system touched a real interest, a real fear, or the pride of someone who was never consulted.
Behind every piece of resistance sits a piece of information you never drew into the stakeholder map. Suppress the resistance and you delete the data. The structure of this one at Anchor & Helm is worth seeing clearly. The survey team lead is not on your Chapter 5 stakeholder map. Back then the actual user (the person who uses the system every day) was the reviewer, and the surveyor was only a step upstream in the workflow. But once the queue absorbed the drafting of chase notices and the Tuesday and Thursday manual filter of overdue claims (Chapter 17), chasing went from scattered emails and phone calls inside Linda's team to a systematized action, timestamped, and able to be aggregated. The rollup view then sums "how many days each claim sat with whom" onto one page. The surveyors' work got turned into data, and from beginning to end nobody asked them a single question. The system's radius of visibility outran its radius of consultation, and the ring outside is the resistance band. Put plainly, the system sees more people than you asked, and those extra people are where resistance comes out.
So the lead's challenge is not unreasonable. The constraint interview this project owed him (Chapter 12) came knocking on its own, in the ugliest way it could.
Inside a company there is one more layer. The lead's line, "that BI system last time was yours too," is not asking about this system. It is asking about your department's track record. The front line's opinion of the Digital Center was built up over several years, and the last system dropped and forgotten, the last promise made and not kept, all of it gets booked to this account. An outside deliverer carries no such ledger. You do. So decoding inside a company means decoding the history along with it. Take the old debt on first, then talk about what is different this time. You take it on not by apologizing but by giving a difference that can be checked. Which of the five capabilities (the test Chapter 3 mentioned, for whether the receiving side can run the system on its own) belongs to whom, whose annual goals name this system, which escalation tree gets walked when something breaks (Chapter 22). Refuse the old debt and every promise you make afterward is discounted automatically.
## Prior Art, and What AI Changed
The consulting tradition treats resistance as a physical phenomenon, and the way to handle it is naming, not rebuttal. Peter Block gives resistance a whole chapter in *Flawless Consulting*. Resistance is a normal physical phenomenon in consulting, an indirect expression of worry. Saying "I am afraid this thing is bad for me" outright is too dangerous, so it puts on a disguise, delay, a barrage of detail, "agreed in principle," a sniping remark. Reasoning with a disguise gets you nowhere. The reasoning answers the lines being spoken, not the worry underneath them.
Block's method has one move only. Say what you sense in neutral language. "I sense this plan worries you. Can you say more?" Then shut up. Naming gives the worry a way out of its disguise, and rebuttal only forces it into a costume harder to recognize.
**The change tradition, Kotter's warning.** The first cause of failed change is underestimating how hard it is to get others to work differently, and a wrong strategy ranks below it (quoted in Chapter 1). The corollary is just as direct. A launch plan that leaves no time for resistance is itself one of the causes of resistance.
The AI era added two sources of resistance that did not exist before. First, a sharp rise in visibility. An AI system turns work into data and makes it comparable by its nature. The decision trail, waiting time, the override rate (the share of system suggestions the front line pushes back), every one of these numbers, born to improve the process, can be picked up and used to appraise people. So the fear of being monitored is rational, not paranoid. Second, the fear of replacement. The system ate the front line's judgment (Chapters 6 and 11 did exactly that), and "it learned my judgment, and then what?"
The two share one thing. Neither can be answered with "the system does not mean it that way," because the system really does have that capability. Defending intent does not work on a fear of capability. The only thing that answers it is a governance pledge, writing "the system could do it and we pledge not to" into a written policy with an issuer and an appeal path. That is this chapter's second framework.
## Framework One: The Resistance Decoder
The resistance decoder, a table mapping the forms resistance takes onto their common real causes and the matching first response. Full version in [Template 20.1](../appendices/template-20-resistance-decoder.md), skeleton below.
| Form of Resistance | Common Real Cause | First Response |
|----------|----------|--------------|
| Delay, rescheduling, "too busy lately" | Interests not aligned, this is all cost to him | Go back and consult, work out what he fears and what he wins (Chapter 5) |
| A barrage of detail, endless technical challenges | Fear disguised as professionalism | Name it, ask the real worry out |
| Sniping at a meeting, a challenge in public | Pride, being turned into data, skipped over, compared | Do not defend, name the emotion first |
| Data withheld, a process forever "in progress" | An information gap plus risk, he does not know what you are up to | Go back and consult + a governance pledge |
| Agreed in principle, nothing moves | Interest or fear, politeness is the disguise | Narrow it to one concrete action, watch where it sticks |
| Silence, no objection and no use | The deepest kind, the conversation has been given up on | Go to them; check for a design defect |
The last row of the table did not play out in this chapter's field, but it is right in front of you at this moment. The only person who spoke up at that weekly meeting was the survey team lead, and teams two and three, just into the queue, said nothing at all. The trouble with silence is that it produces no event. It will never come to you if you do not go looking, so the first step in decoding silence is not guessing what they are thinking, it is finding a count that can be falsified, daily actives, the decision trail, override records. Any one of them at zero over a long stretch turns an impression into a diagnosis. The second step is going to them, and not asking "why are you not using it," which is a demand for an admission of fault. Ask instead, "walk me through one of your recent exception claims from the top," and watch which step he routes around the queue at. The third step is keeping the second half of the rightmost column, check for a design defect, because his silence may simply mean the thing has no use at all in his workflow.
The next chapter shows that usage did collapse in one of these two new teams, and the real cause is not on the front line. It is that team's own lead, who missed the weekly queue retrospective three weeks running.
Three rules of use. **First, form and real cause do not map one to one.** The table gives candidates, the diagnosis comes from an interview, and you ask anchored on a specific instance (Chapter 6's three-layer probing method reports for duty a third time). **Second, handle the emotion before the information.** Naming first, decoding second, redesign third. Reverse the order and not one step works. **Third, always keep one hypothesis alive, that he is right.** The rightmost column of the decoder has an action called "change the design." Some resistance does not point at an emotion, it points at a place where you really did get it wrong.
## Framework Two: The Three Visibility Pledges
The three visibility pledges, three policy answers to the fact that the system has the capability to appraise people, standard equipment for an AI-era deliverable (template at [Template 20.3](../appendices/template-20-resistance-decoder.md)).
1. **The data is used to improve the process, not to appraise individuals**, written into policy, not said out loud. A verbal pledge expires when the person who made it moves to another post, and a policy has an issuer and an appeal path.
2. **The team concerned sees its own data first.** Any number about a team goes to that team before it goes to a meeting. People's hostility to data that is "about me and the last thing I hear about" has nothing to do with whether the data is good or bad.
3. **Aggregate display takes priority over individual detail.** The entrance facing upward defaults to the step and the team, and individual detail lives only in the team's own view.
If the business owner (Kevin's role, not a risk owner like Victor) will not sign, do not treat it as a communication failure yet, because that is an answer in itself. Ask him which of the three he is stuck on. Stuck on the first usually means he is himself being appraised on these numbers by his own manager, so fall back to signing the second and third, land what can be landed, and tell the front line honestly which one you did not get. Do not make a promise on the owner's behalf that he never made. Signing is not the end either. Give the pledge a recheck action, so every later time someone asks for the numbers split by person, you go back to it and ask whether it still stands. A policy pledge with no recheck is only the written version of a verbal one.
A written pledge has one hole peculiar to the inside. Once the data is in the company warehouse, HR or management can pull it directly, going around you, and one query ranks waiting time by person. The policy you signed cannot govern a pull you never hear about. So inside a company the three pledges need technical guardrails, three actions, all of them implemented on the data platform side. Splitting a table by person is disabled at the warehouse layer. The individual dimension is de-identified in the shared layer. A pull that needs individual detail goes through approval and leaves a trail, who pulled what on which day and why, checkable afterward. Only you can do these three, they are the one hard means of honoring the three pledges inside a company, and they go into the pledge text alongside the issuer and the appeal path.
One more action belongs to the inside only. Issuers move to other posts. Kevin will not run the claims line forever, and a policy draws its force from the position, not from the person who signed. So add a line to the pledge text, an issuer change clause. When the business owner changes, the successor reissues within thirty days, and where there is no reissue, the pledge is flagged red on the PMO's AI system transfer ledger and walked through at the quarterly business review along with the system's other transfer items (Chapter 22). You will be at this company longer than any issuer. Nobody will remind you of this, so you have to raise it yourself.
Note what the three pledges are. They limit the right to use, and take nothing away from the system's capability. The trust constraint matrix of Chapter 12 answers the risk owner's "why should I trust you," and the three pledges answer the same question from the front line. Same craft, different recipient.
## At Anchor & Helm: Three Weeks of Decoding One Challenge
**On the spot, Tuesday of week 20.** You did not defend. You said, "That worry is fair. This system really can be used to appraise people, the waiting time is sitting right there, and I am not going to pretend it cannot. Let us talk about how we make sure it is not used that way." The air in the room changed. The lead was ready for a defense and not ready for an admission. Then you booked something. "I want to come sit half a day on the survey side this week and go through your claims from the top." Kevin nodded, the agenda moved on. You did two things and no more, naming, and moving the battlefield out of the meeting room and back to the field.
That sentence about coming to the field costs an internal reader far less to say. You need not go through a visitor process, and nobody has to approve your taking half a day in another department. You walk over. This right of access, being able to show up any time, is a lever only the inside gives you, and it has two uses. One, break the decoding interviews into several half days and follow the claims that actually happen that day, instead of saving it all up for one formal interview. What the front line says at its own desk is not the same batch of words it says in a meeting room. Two, use depth as evidence. You know what ranking people by name set off in this company last time, and which department's numbers have been untrustworthy ever since. You can look that history up and you remember it, and putting it into the draft pledge you hand the owner works better than any argument.
**That week, the decoding interview.** The surface cause is confirmed quickly, the fear of being monitored, and entirely rational. The queue really did make "how many days each claim sat with whom" visible across the board for the first time, and the rollup view really can rank people by name. But you dig further, following Chapter 6's discipline. You do not pick claims by sampling, you pick the few with the longest waiting time, because the mechanism hides in the extremes. The questions follow the three-layer probing method. First have him replay what he did first on that claim that day, then compare it against a claim from the same period that did not go overdue, then ask under what conditions he can afford to wait. There is a test for having dug to the bottom too. The answer has to land on a piece of design you can change, and stopping at "their attitude is the problem" means you have not got there.
You trace six claims "stuck at the survey step," following each one end to end, and dig out a second layer. In four of them the surveyor sent the document request to the repair shop the same day the chase arrived, and then waited. Waiting on the loss assessment list, waiting on repair photos, three or four days of waiting. Those three or four days are all booked by the system to the survey step. The real reason surveyors are slow to complete documents is that repair shops are slow to respond, and the system pinned the blame on the wrong step.
This layer is a debt you owe. Chapter 6's field archaeology put the folding stool beside Linda's desk. You did the archaeology on the fourteen steps inside the review step and never on a single step inside the survey step. Inside the lead's anger sits a real bug. Decode the resistance to the bottom and what comes out is a piece of workflow truth Chapter 6 missed. The repair shop has now crashed into your field of view a second time. Last time it was the list since promoted to an operating asset (Chapter 18), on the risk dimension. This time it is response time, on the efficiency dimension. The same blind spot, two alarms.
**Week 21, the fix goes live.** The queue adds a "waiting on external" status. While a claim waits on a repair shop, a customer or any other outside party, its waiting time is attributed separately and booked to no internal step. The rollup view's bottleneck attribution changes to the step, not the person. The numbers come out again, and the survey step's own waiting time drops by more than half. The bulk of it was "waiting on the repair shop" all along.
On Friday, Kevin issues the three visibility pledges, in writing, to the whole claims line, with one line at the bottom giving the appeal path. Anyone who believes the data is being used to appraise people can appeal directly to Kevin or to the union representative. You drafted it, but the issuer has to be Kevin. A governance pledge can only be issued by the person with the power to violate it. Your signature does not count. A reader doing this inside a company has all the more reason to accept that. You and the front line are colleagues, you cannot retreat to an outsider's position, and the policy still has to be signed by your business owner, drafted by you.
**Week 23, the retrospective.** Halfway through the weekly queue retrospective Linda chairs (it has existed since pilot week 2), the survey team lead walks in. Nobody invited him. He brings a page of numbers he broke down himself. "On the new standard, our own step's median waiting time is 1.8 days. Two new surveyors are dragging it, and I am already coaching them. Also, repair shop response time, should you not build a number for that too? The system watches us, it should watch them as well." From "is it here to appraise us" to "it should watch them as well," three weeks. The next stop on this line is Chapter 21. What a former opponent turns into deserves a chapter of its own.
Tally it up. Everything this resistance produced, a corrected attribution logic, a new status, a governance policy, a team lead who started checking his own numbers, none of it came from your original design. All of it came from decoding that one piece of sniping. Once resistance is decoded, the information in the opponent's hands becomes the system's improvement.
## Failure Modes
**1. Crushing emotion with logic.** You pull up the metric definition document on the spot, prove line by line that the numbers are right, and the other person has nothing left to say. An engineer is trained to rule on right and wrong, while an emotional appeal is not asking for evidence, it is asking for acknowledgment. Every point you win is deducted from the other person's pride. In the Trust Equation this zeroes out intimacy, and a person beaten in public will never tell you the truth again (Chapter 5). You win the argument and lose the system.
**2. Getting the sponsor (the executive who funds it and makes the call) to lean on people.** After the meeting you report to Grant that "the survey team is resisting the change" and ask him to weigh in. Power can change behavior, it cannot change willingness. Grant applying pressure wins this one round, but the whole organization learns the same lesson, that raising an objection to this system gets it taken to Grant. From then on nobody tells you the truth, and the resistance goes underground into the most expensive kind, silence. You also push the lead into permanent opposition along the way, and the truth in his hands, that repair shops are slow to respond, will now never reach you. Inside a company there is one more layer to this bill. You and the person leaned on are long-term colleagues. He is your resistance this time, and next time he is very likely the dependency of another one of your projects, the one who has to hand over data and put up people. Win by pressure this round and what you bought is a person who can quietly block you every time you need his cooperation from now on, in a way that appears on no ledger and that nobody will attribute for you. An outsider leans on people once and moves to the next client. You do not move.
**3. Reading silence as agreement.** Nobody objects at the meeting, nobody asks a question at the training, and the weekly report says "feedback from all teams is good." Resistance that speaks up at least still cares how this ends. The deepest resistance says nothing, it just does not use the thing. The cost structure of a meeting makes silence the cheapest form of objection there is, especially just after the organization has watched how someone who did raise an objection got treated. The daily actives curve after the expansion will put this on the table, and that is the opening of Chapter 21.
**4. Treating all resistance as misunderstanding.** The response to every challenge is to "explain it one more time," to run another training. "Misunderstanding" is the attribution that protects pride best, the plan is fine and they just did not understand it, so the cure is always more communication and the design does not move a line. But some resistance points at a real design defect. At Anchor & Helm this time, an attribution error was buried right under the "it is here to appraise us." One test. On every piece of resistance, force the question, "if he is right, which piece of the design is wrong?"
## Next Monday
1. Write down the last piece of resistance you ran into and run it through the decoder in [Template 20.1](../appendices/template-20-resistance-decoder.md). What form did it take? What are the candidate real causes? Was your response at the time naming, defending, or crushing?
2. Check your system. Whose work has it turned into data without ever consulting them? Radius of visibility minus radius of consultation, and the difference is your resistance band. Give the people in that band a constraint interview (Chapter 12).
3. If your system produces any number that can be used to appraise an individual, push the owner to issue a governance pledge this week ([Template 20.3](../appendices/template-20-resistance-decoder.md)). Written, with an issuer, with an appeal path. A verbal version does not count.
4. Next time you are challenged in public, do two things and no more. Name it ("that worry is fair"), and book the field ("I will come over and watch for half a day"). Scripts in [Template 20.2](../appendices/template-20-resistance-decoder.md). A script is a crutch, not a line to recite.
**Want an agent to get you started?** In the repo you set up following [Start Here](../index.md), paste this to your coding agent:
```text
In the the-last-mile repository, help me with the Chapter 20 Next Monday actions. This chapter has no script. Open docs/appendices/template-20-resistance-decoder.md,
build 20.1's decoder as an empty table. I will describe the last piece of resistance I ran into, and you record only the form and the candidate real causes in the table's columns.
Whether my response at the time was naming, defending or crushing is mine to judge. Then help me compute radius of visibility minus radius of consultation. The list of people
turned into data by the system but never consulted is mine to name, you only count. Draft the governance pledge from 20.3, leaving the issuer and the appeal path blank for me.
If any command errors, stop and show me the output.
```
---
## Chapter Kit
- **Judgment frameworks.** The resistance decoder (form โ real cause โ response action; name first, decode second, redesign third); the three visibility pledges (not for individual appraisal and written into policy / the team concerned sees it first / aggregate before detail)
- **Templates.** [Template 20](../appendices/template-20-resistance-decoder.md), Resistance Decoder, Hard Conversation Scripts, Visibility Governance Pledge
- **Key judgments**
- "Resistance is data. Suppress the resistance and you delete the data."
- "The fear of being monitored is rational, and the only thing that answers it is a governance pledge. 'The system does not mean it that way' answers nothing."
- "You win the argument and lose the system."
- "The deepest resistance says nothing, it just does not use the thing."
---
# 21 ยท Adoption Engineering: From Usable to Missed When It Is Gone
!!! info "Companion Templates"
๐ [Chapter Template](../appendices/template-21-adoption-plan.md) ยท ๐ [Template Library](../appendices/template-library-index.md)
> **The Challenge.** The training was held, the email went out, and the system really is good. The same few people are still the only ones using it. What does adoption actually run on?
>
> **What You Will Be Able to Do.** Use the four adoption mechanisms (super-user network / operating cadence / short-term win announcement / institutional anchoring) to turn "one team uses it" into "taking it away would hurt." Diagnose a falling adoption curve and find the root cause, which is almost always in management.
---
## Monday of Week 21, Three Daily-Active Curves
After the expansion was approved, auto review teams two and three entered the queue in week 19. Same system, same training, even the same instructor. Monday of week 21, you lay the three daily-active curves side by side. Daily actives, the share of reviewers who actually opened and worked the queue that day.
- Linda's team, 100%, not one day below since the pilot began.
- Team two, 60% in their first week on the queue, steady since.
- Team three, 60% in their first week, and already below 30% at the start of week 3.
Every variable is controlled and the curves still split three ways. Only one difference is left. Linda's team has Linda. And you cannot assign a Linda to every team.
That one sentence is the whole subject of adoption engineering. Turn the Linda effect into a mechanism you can reproduce.
## Why This Is Hard: Training Fixes Knowing How, Not Whether They Use It
The engineer's default model of adoption is the product model. Make it good and people will use it. The real model in an enterprise is the social model. The front line watches its peers, peers watch the person who leads, and that person watches whether the boss cares. Training solves a cognitive problem, whether they know how. Adoption is a habit and social problem, settled by three variables no feature list can touch. Who is using it (the peer environment), whether the people using it are doing well (visible winners), and whether not using it costs anything (institutions). The four mechanisms below grow out of these three variables. Only the operating cadence maps to none of them. It is the carrier that makes the other three happen again every week. Who is using it maps to the super-user network, doing well maps to the short-term win announcement, and whether it costs anything maps to institutional anchoring.
In one sentence, organizations do not adopt tools. Habits adopt tools. A tool enters daily work through ritual, visible winners and peer pressure, and not one of those three happens automatically because "the feature shipped." Between L3 on the outcome ladder (adopted, "taking it away would hurt") and L2 (production, the system is in production with an owner and monitoring) lies nothing but organizational engineering of this kind.
## Prior Art, and What AI Changed
**Kotter's eight steps for change, cut down to four for the field.** John Kotter's framework in *Leading Change* (paraphrased) is designed for organization-level change. The deliverer's battlefield is one size smaller, one department and one workflow, and eight steps cut to four are enough. This chapter's four mechanisms come from there, mapped as follows. Guiding coalition = the super-user network, short-term wins = early wins announced, communicating vision = said again in every weekly ritual, anchoring in culture = written into the institutions.
**Habit thinking.** The habit loop Charles Duhigg draws in *The Power of Habit* (cue-routine-reward, paraphrased), where new behavior forms through a loop of trigger, action and reward, not through one-time persuasion. Applied to adoption, the same hour every week, the same room of people, the same queue, a ritual that keeps recurring is the cue.
AI changed two things, the cold start and the fragility.
The cold start. This system gets better through use, and reason codes and override trails are the fuel for rule iteration and the golden cases (Chapter 17, the trail flowing back). But before it gets better, who uses a system that is not smart enough yet? Chicken and egg. The answer is that the super-user network carries the feeding period. A small group of seed users, given a formal identity, keep using the system at its dumbest and keep feeding it real judgment until it gets smart. The feeding period runs on design and on reward, not on goodwill you sit and wait for. And the reward has to be paid before the system gets smart. What Anchor & Helm did was put Linda's name on the five rules. The reason column of every suggestion carries the name of one of her rules, and the more the system is used the more it looks like her work. Seed users keep feeding it at its dumbest, and what they feed it is their own.
The fragility. Trust in AI is built one suggestion at a time, every suggestion is on trial, and organizations have an asymmetric memory for incidents (Chapter 18, an improvement is remembered for three days, an incident for a year). It follows that one unsafe the defense failed to catch can empty a team's daily actives. Traditional software that is hard to use only gets used slowly. An AI system gets one thing wrong and the story becomes "this thing is not reliable." The adoption curve is far more fragile than it is for traditional software, which is why adoption has to be engineered and cannot be left to grow on its own.
## The Core Framework: The Four Adoption Mechanisms
> **The four adoption mechanisms, the four that turn adoption from a wish into a structure. Network first, cadence fixed, wins amplified, institutions last.** (Template in [Template 21](../appendices/template-21-adoption-plan.md))
The four rows below correspond to those four steps in order.
| # | Mechanism | What It Is | What It Prevents |
|---|------|------|-----------|
| 1 | **Super-user network** | A super-user, a seed user with influence on the front line, given privileges (new features first, a direct channel for improvements) and identity (a formal name, credit in public). **Pick by influence, not by title** | Adoption resting on the project side's pitch alone; nobody to carry the feeding period |
| 2 | **Operating cadence** | A 30-minute queue retrospective every week, going over the reason code distribution, aging claims, and one improvement. **Embed it in a ritual that already exists, do not create a meeting** | A new ritual cannot win a calendar slot and dies in three weeks |
| 3 | **Short-term win announcement** | Manufacture one win worth telling inside the first month, in business language, announced by the business owner. **The winner is the business team** | The improvement goes unnoticed, or the credit goes to "AI" |
| 4 | **Institutional anchoring** | New-hire onboarding material covers queue operation; the exception claim SOP is updated to reference the system. **Lock the habit in with the cost of leaving it** | Held up by personal enthusiasm, and the curve goes to zero when the person leaves |
Two design points. The test for picking a super-user comes down to one question. A hard claim comes in, who does everyone get up and go ask? Inside a company you already have the answer to that question. You have been here a few years, and which team lead really has influence and which one only has the title is not something you need three months on site to see. Put that depth straight to work on the candidate list. The cost difference in the operating cadence is legitimacy. A new meeting has to win a calendar slot, win attention, and keep proving it deserves to exist, while a ritual that already exists comes with its own attendance. Embedding the system into an existing ritual is ten times cheaper than building a new one.
The agreement itself is one page ([Template 21.2](../appendices/template-21-adoption-plan.md)). Anchor & Helm gave Linda three privileges. New rules two weeks before everyone else, improvement suggestions going straight to the development board (the board your team schedules its own work on, not the review queue) with a reply guaranteed inside two weeks, and the right to chair the retrospective. What comes back is the feeding-period commitment. Keep using the system at its dumbest, fill in reason codes seriously, and the time invested goes into her workload. Exit is written up front. She may leave voluntarily at any time, four straight weeks of not using it or not attending voids the status automatically, and using data visibility to lean on colleagues means immediate removal.
Two clauses of this agreement go hollow most easily inside a company. One is workload. Writing it into workload does not happen by itself. It has to land on the super-user's own supervisor, either a written confirmation of how many hours a week or the item going into the super-user's OKRs for the quarter. A nod given in a meeting is eaten by the day job by week three. The other is the direct channel. Your team is split across several projects at once, and "a reply inside two weeks" is the first promise to be sacrificed, so reserve a fixed response capacity allowance in your own schedule (a fixed block each week for answering improvement suggestions) before you say that privilege out loud. A bad check can be written only once. Break faith once and the whole super-user network stops speaking for you.
### The Adoption Curve Diagnostic Order, What to Check First When a Curve Drops
- Check the unsafe trail first. Pull that team's scores and override records for the two weeks the curve turned, and look for an unsafe that landed and for concern crossing the pilot-period threshold. If there is one, the problem is the system and the incident story spreading from it, not adoption.
- Then check the manager's attendance. Pull the retrospective sign-in sheet and count how many of the last four weeks this team's manager showed up. Zero attendance and the root cause is right there. The fix sits with the owner, not with the system.
- Only then check features. When the first two come back clean, it is time to look at this team's reason code distribution and aging claims to find which feature it is actually missing.
## At Anchor & Helm: Four Weeks of One Network
**Week 19, the person standing at the front of the expansion training is Linda.** That is by design, not coincidence. A peer's testimony beats the project side's pitch, and the front line only believes the front line. It works the same way inside a company. The project side is your own team, and the front line will not count you as a peer just because you and they draw pay from the same company. Page one of the handout carries the five rules of thumb, amount, report delay, photo count, prior claim linkage, repair shop list, signed by Linda, with the system screenshots further back. That day Kevin gave her a formal identity in front of everyone, queue co-builder. Not "system administrator," not "training instructor." The name itself says that half the system is hers. One section of the training is devoted to the repair shop list, now an operating asset with an owner (Chapter 18), and the first thing a new team does on arrival is add the shops in its own district.
Standing in the audience you think back to Chapter 0. Her first feedback on this system was three unsafe marks. From skeptical scorer to the front of the room there was never one act of persuasion. Her judgments went into the system one at a time. The five rules (Chapter 6), the annotation rulings (Chapter 11), the implementation flaw the reason codes fished out (Chapter 17). You did not turn Linda into a believer in the system. You turned the system into Linda's work.
**The operating cadence was planted back in pilot week 2.** You did not create a "system weekly" then. You asked Kevin only for the last 30 minutes of his existing weekly, to go through this week's override reason code distribution and the top aging claims and settle on one improvement. The engineering side's prelude came earlier. The biweekly rotating release (Chapter 15) had already set the beat of "the claims-ops IT engineers demonstrate to the business owner," and the business side's weekly retrospective meshes with it into a complete cadence. From week 19 the retrospective is chaired by Linda, and you move back into the audience.
**After the week 21 retrospective, the 25% diagnosis.** The curve that dropped below 30% on Monday settles at 25% after this week's retrospective. Your first reaction is the lesson of Chapter 18. Was there an unsafe? You check the trail. Team three, three weeks, zero unsafe, and concern inside the threshold too. The system is not at fault. Then you check the retrospective sign-in sheet. Team two's lead is there every week, team three's lead has zero attendance across three weeks. This week is no exception. He leaves when the first half of the standing meeting ends, "the second half is your project's business." His team read that sentence instantly. This thing does not count in our team.
When an adoption curve drops, the first thing to suspect is not the system. It is the manager's attention. Whether team members use it depends first on whether their own boss looks at it, and how good the system is comes second. The fix sits with Kevin, not with the system. You added no feature and sent no email. You put the curve and the sign-in record side by side in front of Kevin. He read them and said one thing. "This is on me." In week 22, attendance at the retrospective goes into the team lead's job description. In week 23, team three's daily actives climb back above 80%. You did not change a line of code.
Inside a company this step needs one more calculation. Handing the sign-in record to Kevin means reporting on the attendance of a team lead who works for the head of a peer department, across a reporting line that is not yours, and it reads easily as tattling. The steadier order is to let the number walk into his view on its own. Per-team daily actives and retrospective attendance belong in claims operations' own monthly operating data anyway, and Kevin will see them when he turns the page of his own report. Or let Linda raise it at the retrospective as a super-user, because coming from the front line it is a peer's opinion and coming from you it is the department next door passing judgment. Only when neither road works do you lay the two sheets side by side, and even then you put down the data only, with no conclusion.
**Also week 21, something nobody announced.** Thursday evening you are on site and walk past Linda's desk. For six years the last thing she did before leaving was back up that Excel, step 14 of Chapter 6's fourteen steps, one of the three steps Chapter 17 ruled "Kept." Today she closes the laptop and goes. You catch up and ask her. She says, "It is all in the queue." She pauses, then adds, "I stopped the day before yesterday. You are only noticing now?"
You checked afterward. That six-year-old file was last opened on Monday of this week. No ritual, no email, nobody's approval. In Chapter 17 you wrote that the day she stops on her own is the true measure of trust. The measure arrived, on a Thursday evening nobody noticed. Trust never holds a launch event.
The source of truth did not fall with it. At the week 13 retest the core system's status field had already reached validated (Chapter 9), on the strength of the daily write-back discipline Kevin claimed at gate three (Chapter 14), and from that week the merged view stopped deferring to the Excel. Her stopping the backup is evidence that the discipline runs in her hands, not the case Chapter 9 warned about, where the system kills its own source of truth with its own hands.
**Week 23, two things land in the same week.** First, auto team two clears its exception claim backlog, the first short-term win worth telling. Kevin announces it at the monthly business review, one page, two minutes. On that page, one side is the date the backlog hit zero, the other is team two's first-touch handling time curve coming down. The subject from start to finish is "team two," and not once does the room hear "the AI launch was a success."
Second, the survey team lead comes to the retrospective with self-check numbers for the first time, the waiting time for his own step, which he pulled himself on the new standard (Chapter 20). The "waiting on external" status and the three visibility pledges (Chapter 20) let him see his own team's data first, and this time he brings data and comes looking for improvements. In week 24 Kevin names him the second super-user. The former opponent is the most convincing spokesman. One sentence from him, "this thing is not appraising us," lands harder than a hundred from you, and the whole department remembers the question he asked with his pen set down in front of everyone.
What landed this week is the first three mechanisms. The fourth is still owed. Institutional anchoring rides on two documents, the new-hire onboarding material (the set that teaches a new hire how the work is done) and the exception claim SOP (the standard handling steps for this class of claim). The person who changes them is not you. It is the process owner, and at Anchor & Helm that is Kevin.
That sentence is only half true inside a company. You can write the change as revision text and hand it straight to the process owner, and you can file the SOP revision request yourself in the process management system, an entrance an outside consultant does not have. The price is that once filed it goes through compliance or quality management review, one more gate, two more weeks. Of the four, institutional anchoring is the biggest lever in the internal reader's hands. The first three you have to push every week. The fourth, once changed, needs nobody to remember it, so the moment to start is earlier than you think. You do not have to guess the moment either. Attendance is already written into the team lead's job description, so these two documents are the next thing to move.
What you do is write the change as two sentences and hand it over. Add a section on queue operation to the onboarding material, and everywhere the exception claim SOP says "check the core system," change it to "check the queue." That onboarding section does not have to be assembled from scratch either. The company already runs a new-hire training program, and slotting a section into it lives longer than building a separate one of your own. At Anchor & Helm the two documents were settled in week 24 and issued by Kevin. The queue operation section of the onboarding was written by Linda, who by then was already an instructor in new-hire training ([Template 21.2](../appendices/template-21-adoption-plan.md)). The SOP was revised by Kevin's own operations team, and the revised version went back out to every team lead. Only after this step does opening the queue turn from a personal habit into a job duty.
The North Star stopped at -22% when the pilot ended, and it will reach -31% in week 26, but that is the next chapter's stretch of time. The more important measure right now is what the week 23 retrospective was made of. Linda chaired it, the data came from the survey team lead, and the attendance rule was signed by Kevin. The right to chair the operating cadence is already in the business side's hands. The handoff (Chapter 22) has already begun, and it does not wait to be announced.
## Failure Modes
Five. The first two treat adoption as training or as an order. The last three each get one thing wrong, in the super-user, in the win announcement, and in institutional anchoring.
**1. One training session treated as adoption.** Launch comes with one training session and a manual, and three months later daily actives are zero. Training is a cognitive intervention, adoption is habit engineering. A habit needs repeated cues and a peer environment, and a one-off event structurally cannot reach that far. Training is also easy to deliver, easy to tick off, and has a completion date, while adoption has no deadline, so only the first one survives in the performance review. The test, look only at the daily-active curve in weeks two and three after the training, not at the sign-in sheet.
**2. Pushed through by order.** The boss orders "you must use it," and daily actives hit target within a week. An order changes login behavior, not decision behavior. The front line will invent the cheapest compliance available, accept everything, and pick "other" for the reason code without looking. And an AI system eats the data it produces itself. Going-through-the-motions trails flow into rule iteration and the golden cases (Chapter 17), garbage in, the system gets dumber, and that in turn proves "it is not useful." Pushing traditional software through wastes licenses. Pushing an AI system through poisons its own fuel. The test, high daily actives with an override rate near zero and the share of "other" reason codes shooting up means nobody is actually judging. Do not read it as a perfect system.
**3. Picking champions by title.** Champion, what the change literature calls an internal advocate. Every team lead is appointed a "rollout ambassador." Titles are easy to find on the roster and appointing costs nothing to negotiate. Influence only shows up if you sit in the field, and "who does everyone ask when a hard claim comes in" and "who is the team lead" are often not the same person. Picking by title also insults the real opinion leader for free, the person who got skipped. Neither of Anchor & Helm's two super-users came by title. Linda came by twenty years of field authority, the survey team lead by the turnaround of "even he came around."
**4. The win announced as "AI succeeded."** At the monthly meeting, "our company's AI project has achieved phased results," with a system screenshot. The project side is under reporting pressure, and the "AI succeeded" story pays the project side immediately. But it rewrites the business team from winner into backdrop. Team two did three weeks of work and the credit went to the system. The super-users' account stops balancing at once. They staked their standing to vouch for the system and the return was taken away. Steal the credit once and the network falls apart. A side cost, raised expectations (Chapter 19) make the next incident a longer fall. The rule, the subject of the announcement must be the business team, the measure must be a business number, and the system's name goes in the last paragraph or nowhere. Inside a company, the business team means the department that uses the system, not the platform or data team you sit in, and drawing pay from the same company does not change that rule. The credit has to go on their side of the ledger.
**5. No institutional anchor, held up by people.** The network is built, the cadence is running, the win has been announced, and the one thing nobody went back to do is change the onboarding material and the exception claim SOP. For the months the super-users are around, the curve looks fine. But people move. The super-user transfers, team leads rotate, and a new hire still gets the old process that says "check the core system," so the old process is what he learns, and the system slides from "should be used" back to "may be used." The first three anchor to people, the fourth anchors to the job, and only the fourth needs nobody to remember it. The test, find someone who joined recently, ask how a claim of this kind should be handled, and see whether what he recites is the queue or the old SOP.
## Next Monday
1. Pull per-team daily actives and depth of use (override rate, reason code fill rate) once, and find your "Linda's team" and your "team three." For a team with an abnormal curve, check the trail for an unsafe first, then that team's manager's attendance at the adoption ritual over the last four weeks, and only then the features.
2. Write down three names, the people everyone asks first when a hard claim comes in. That is your super-user candidate list. Check it against the org chart, and if all three are team leads, pick again.
3. Find a standing meeting that already exists and negotiate its last 30 minutes for the queue retrospective. Do not create a meeting. Lay out the cadence with [Template 21.3](../appendices/template-21-adoption-plan.md), and use 21.2 to write down the privileges and the identity with your first super-user.
4. Check the most recent announcement that went out. Is the subject the business team or the system? If it is the latter, rewrite the next one, and confirm the business owner is willing to say it in his own voice.
5. Dig out the new-hire onboarding material and this workflow's SOP, list every paragraph still carrying the old system's name into one list of changes to make, and hand it to the process owner along with the wording. If the documents do not change, the first four all walk out with the people.
**Want an agent to get you started?** In the repo you set up following [Start Here](../index.md), paste this to your coding agent:
```text
In the repo/ directory of the the-last-mile repository, help me with the Chapter 21 Next Monday actions. First run python3 templates/adoption/retro_onepager.py
on the built-in sample to show me a one-page queue retrospective, then open weekly-usage.sql and explain how per-team daily actives and depth of use are computed. I will give you the decision trail export,
you only run the queries and draw the tables, and which team is "Linda's team" and which is "team three" is my call. The three super-user names are mine to write, you only check them against
the org chart and flag whether all three are team leads. Build the operating cadence sheet per Template 21.3, and the name of the standing meeting and how to negotiate its last 30 minutes are mine to handle.
If any command errors, stop and show me the output.
```
---
## Chapter Kit
- **Judgment frameworks.** The four adoption mechanisms (super-user network picked by influence / operating cadence embedded in an existing ritual / short-term win announced in business language / institutional anchoring into onboarding and the SOP); the adoption curve diagnostic order (check the unsafe trail first, then the manager's attendance, then the features)
- **Templates.** [Template 21](../appendices/template-21-adoption-plan.md), Adoption Plan, Super-user Agreement, Operating Cadence Sheet
- **Key judgments**
- "Adoption is a habit and social problem. Training only solves the cognitive one."
- "Embedding the system into an existing ritual is ten times cheaper than building a new one."
- "An adoption problem is mostly a problem of the manager's attention. The fix sits with the owner, not with the system."
- "The former opponent is the most convincing spokesman."
- "The business winner must be the business team. Steal the credit once and the network falls apart."
---
# 22 ยท Make the Business Side Self-Sufficient: Stepping Out of the Daily Transfers Capability, Not Files
!!! info "Companion Templates"
๐ [Chapter Template](../appendices/template-22-handoff.md) ยท ๐ [Template Library](../appendices/template-library-index.md)
> **The Challenge.** The project is winding down. How do you step out of the daily so that the system does not die three months after you leave?
>
> **What You Will Be Able to Do.** Use the five self-sufficiency tests to judge whether the receiving side can hold the system. Withdraw gradually over the four stages of the handoff cadence, instead of moving documents in the last week. Land each of the three AI-specific handoff items on a named owner.
---
## Wednesday of Week 18, a Compliment That Should Alarm You
Wednesday of week 18, the annual budget and headcount review, two days after the impact memo went out (Chapter 19). This meeting is not a celebration. Finance and the PMO spend the first twenty minutes on the Group's new policy. Starting next year, support for every live system must not be tied to specific people, and the ops headcount the Digital Center has sitting on each subsidiary's systems will be counted one by one. Owen Hartley's response is to widen the scope and keep headcount as is. Each side argues its case, then both look at the same page, the run chart (the week-by-week line of the North Star) stopped at -22% where the pilot ended, and then both look at Grant Whitmore.
Grant finishes the page and says the sentence that makes most team leads in the room feel their headcount is safe.
> "You are not going anywhere. This system does not run without you."
Everyone in the room laughs, Owen included. What he hears is that the team is needed and next year's headcount will be easy to get. You do not laugh. Praise is also a diagnosis. It announces that your success so far hides a failure. You have become the system's single point of failure. It also turns what Finance just asked for into a paradox. The system cannot do without you, and people must be replaceable. Those two sentences cannot both hold, and the key that unlocks them is not in the headcount table.
Take stock of what this system lives on every day. Who guards the unsafe line (the grade among the four where acting on the suggestion causes harm)? Who watches for list drift? When cost breaks the line, who decides what to cut? Who has the say on switching model versions? The answer to every one of them right now is the same word. You.
The signals of adoption are there, and taking it away would hurt (Chapter 21), but the L2 account is not settled yet, the North Star is still eight points short. And the ruler Chapter 1 set up has one more rung. L4, self-sufficient, the business side can run, maintain, and improve it on its own. The deliverer's achievement settles only here.
You say to Grant, "Right now it really cannot run without us, and that is the last defect this system has left. Give me a quarter to make it run without us, then decide what to keep funding. Finance wants 'replaceable people.' We will give a more thorough version. The capability grows in claims operations' own people." Grant approves two things, expanding the team to go deeper, and a handoff plan. The review's resolution gains a line. Next year, the Digital Center's ops headcount on this system is cut according to the results of the five self-sufficiency tests. Pass all five, headcount goes to zero. Whichever fails, its headcount stays, and every one that stays is booked on the Digital Center's headcount. What it hangs on is the five-test table below, not a date (Chapter 14's rule, this time Finance wrote it into the resolution on its own). Kevin Doyle picks up the second half. "Grant, I will lead the handoff on our end."
## Why This Is Hard: Documents Are the Shadow of Capability, Not the Capability
The usual picture of a handoff is "write documents and run training." But the system lives on five daily capabilities growing in the business side's people, and documents cannot hold it up. Run, so someone can respond when it fails. Configure, so someone dares to touch a rule that needs adjusting. Exceptions, so someone can judge a situation never seen before. Evolve, so someone keeps feeding the eval. Teach, so someone can train the next newcomer. Documents record "why we decided this back then." What the system needs every day is "how to judge this now." Judgment cannot be written into a document. It can only be practiced into people. This is the mirror image of Chapter 6. Back then you dug the judgment out of Linda Marsh's head and put it into the system. Now you have to put the judgment that keeps the system alive back into the business side's people.
The other layer is your own team's economics. Staying on is comfortable. The business side has no worries, Owen's headcount is easy to ask for, you are needed. Dependency is an asset in the service business and a liability for the deliverer. Inside a company it has a more precise name, capacity debt. Every extra day you stand duty on this system is a day of capacity the Digital Center cannot give another department next year, and the headcount won at the review gets eaten by this debt one system at a time. It also locks out pattern reuse (whether the same approach can be carried to the next business unit, Chapter 23 expands on it), and dependency itself is fragile. The week you are on vacation is the system's risk exposure.
This is the paradox you face. You cannot leave, and you must be able to leave. You are in this company, three months from now still in the same building, and no day comes when someone revokes your access. So your stepping out of the daily is an event of responsibility, not a physical event. The five capabilities (run, configure, exceptions, evolve, teach) transfer from your team to the business side and ops, and you turn to the next project. The other half of the paradox is more dangerous. The day you could leave never arrives on its own, the handoff can always be put off a little longer, and so your team turns, project by project, into permanent ops for every AI system in the company, with capacity for new projects at zero. That is the organizational version of failure mode 3 (your team becomes permanent ops). No date forces anyone to start a handoff. The force to start has to be built by hand, and the three mandatory mechanisms below exist for that.
## Prior Art, and What AI Changed
**The consulting tradition, the ending is written into the beginning.** Peter Block's philosophy of landing the work, in *Flawless Consulting* (paraphrased). The ultimate goal of consulting is that the client no longer needs you. Measure an engagement by how much stronger the client's capability is after you leave than before you came. How pretty the deliverable is comes after.
**The teaching tradition, scaffolding withdrawn gradually.** The gradual release of responsibility in teaching theory (Pearson & Gallagher, paraphrased). Responsibility moves over step by step. I do, you watch. We do it together. You do, I watch. You do. The whole value of scaffolding is in taking it down. Scaffolding put up and never taken down ends up part of the building, bearing load in place of the wall.
What did AI change? The handoff checklist gains three items, and the vocabulary of a traditional IT handoff does not have them.
1. **The capability to keep the eval updated.** Who adds golden cases, who approves them? Who guards the thresholds? Who manages annotation guide versions? The eval is a living operating asset (Chapter 11). An eval frozen on handoff day means drift (the silent shift in input distribution or suggestion quality) goes undetected from then on.
2. **The capability to re-verify on model and dependency changes.** When the model provider upgrades a model version or retires an endpoint, who runs the golden cases replay, who signs "safe to switch"? Miss this one and the first model retirement notice is the system's death sentence.
3. **The capability to review decision trail data.** Who chairs the override review? Who analyzes the reason code distribution? The closed loop, the last level (Chapter 17), is fed by it.
A traditional handoff checklist asks "are the documents complete, were the accounts handed over." It does not ask "who tests when the model changes." These three are the easiest to miss and the most fatal.
Inside a company these three have one more way out, and teams delivering externally do not have it. They rebuild these three from scratch at every company, leave them there, and build them again at the next. You build once. The three AI items go platform-level. Continuous eval updating, re-verification on model and dependency changes, and decision trail review become company-wide shared capabilities, one eval platform, one change re-verification process, one trail review template, and permanent ops drops from one copy per system to one copy per company. The day Anchor & Helm's golden cases replay process becomes a company-level one, the home queue and the next subsidiary's system no longer each keep their own. This is an action, not a consolation. The first step is to pull the replay process out of Anchor & Helm's repo and hand it to the Group platform team to maintain, while your team keeps only the right to set the questions, that is, to keep deciding what the golden cases should test.
## The Framework: The Five Self-Sufficiency Tests and the Four-Stage Cadence
The five self-sufficiency tests. Can the receiving side hold the system's five daily capabilities? The test for each is Show Me, not a signature. (Template at [Template 22.2](../appendices/template-22-handoff.md))
| Capability | What It Covers | Show Me (Anchor & Helm Version) |
|------|------|--------------------|
| **Run** | Daily operation and first-line incident response | Cut the LLM dependency without warning, and the receiving side degrades, notifies, and recovers on its own (Chapter 16's degradation drill, run by the receiving side this time) |
| **Configure** | Adjusting rules and thresholds | The monthly repair shop list update + one rule parameter through the full change, test, release cycle |
| **Exceptions** | Handling and escalating situations never seen before | Inject a new class of situation and watch whether the first reaction is the judgment path or a call for help |
| **Evolve** | Eval updates + small change development + change re-verification | A new request from proposal to rotating release with you nowhere in it; a minor model version upgrade passes the golden cases replay |
| **Teach** | Training the next newcomer | A new reviewer is onboarded by the business side's own mentor, and you only look at the result |
Three rules. Show Me means you are in the room with your hands behind your back. One failure sends it back for rework, and what gets added is a drill, not a document. Only when all five pass does stage four begin.
The four-stage handoff cadence, the engineering of scaffolding withdrawn gradually. Each stage sets responsibility with three questions, who chairs the retrospective, who touches production, who answers outside questions, plus one gate for entering the next stage. (Template at [Template 22.3](../appendices/template-22-handoff.md))
| Stage | Who Chairs the Retrospective | Who Touches Production | Who Answers Outside Questions | Condition for the Next Stage |
|------|--------------|----------|--------------|--------------------|
| 1 Your team leads (business side observes) | Your team | Your team | Your team | Co-build agreement in force, gatekeeping rights of the first module handed over (Chapter 15) |
| 2 Shared lead | Business side chairs, your team adds | Both, gatekeeping rights handed over module by module | Business side answers, your team backstops | Gatekeeping rights of every module on the business line side, four consecutive retrospectives chaired by the business side |
| 3 Business side leads (your team advises) | Business side | Business side (your team no longer commits) | Business side | All five self-sufficiency tests passed ([Template 22.2](../appendices/template-22-handoff.md)) |
| 4 Stepping out of the daily | Business side | Business side | Business side | Response window closed with written confirmation, write access reduced to read-only, no open rework items |
Laying out this table in week 20, you discover something. Stage one finished long ago, and stage two is already more than half done. Linda has chaired the retrospective since week 19 (Chapter 21), and gatekeeping rights over the extraction pipeline were handed over in week 1 of the pilot (Chapter 15). That is by design, not coincidence. The handoff began on day one of co-build. The repo belongs to the business line, backward staffing, the rotating release, "If you cannot say it, do not merge it." The four clauses of the co-build agreement (building with the business line's engineers) are development management and the first stage of handoff at once. All that really needs scheduling is the three transfers of stage three, and the date of stage four.
The date of stage four is the softest cell in an internal handoff. No outside date forces you to fill it, and once filled, nobody forces you to keep it. So the force that starts the handoff has to be built by hand. Three mechanisms, and missing any one of them slides you back into permanent ops. One, next year's staffing budget reserves no ops headcount for this system. Two, completing the handoff goes into your own quarterly goals, and the date all five tests pass is a line in your OKRs. Three, the PMO keeps an AI system transfer ledger, one line per live system, which of the five capabilities belongs to whom, blanks flagged red, walked through at the quarterly business review. The three draw their force from different places. The first is Finance's budget line, the second is Owen's review of you, the third is the PMO's standing meeting, and none of them relies on your willpower. They are the same shape as Chapter 14's resource gates. Those three built the pilot's expiry date, these three build the handoff's. In a company with no PMO, the third hangs on the standing agenda of the sponsor weekly, with the same effect. At Anchor & Helm this time, the first was written into the resolution at that review. The other two you have to go write, and go ask for, yourself.
The internal receiving side is not one party but three, and the five capabilities land separately. The business side takes configure, exceptions, and teach. Whether a rule changes, how a never-seen claim is judged, who mentors the newcomer, these are judgments, and they can only grow in claims operations' own people. Ops takes the first-line incident response and degradation notices inside run. That is what they already do, one more system is all. Evolve, together with the three AI items, lands directly on the two claims-ops IT engineers Anchor & Helm has.
Most business units cannot take this one, because they have no engineers of their own. Then the three AI items fall through the word "ops" into the traditional ops ticketing system, and the traditional ops checklist has no line for golden cases replay. When the model retirement notice arrives they run it through change management as a dependency upgrade, testing whether the endpoint responds, not whether the suggestions are right. There is one backstop. The three AI items get their own lines on the transfer ledger, with real names for the receivers, and where no real name can be written, the platform team takes execution and the business side's eval guardian signs. Writing "ops" and moving on is not allowed.
## At Anchor & Helm: Eleven Weeks of Gradual Withdrawal, as It Happened
**Week 22, where the budget sits.** With the new budget cycle, running cost moves from project funds into claims operations' department budget. That is exactly what the small print in the footer of Chapter 16's four-ledger page recorded, and this week it is honored. At the budget meeting Finance asks Kevin "why is this number this much," and he answers from the four ledgers. From this week on you no longer sit in on cost reviews. The system owner is now in place. Whoever holds the budget holds the priorities. This step is the financial source of whether an internal team slides into permanent ops. As long as the money hangs on the Digital Center's project funds, the system is still yours, and every "can you change this" from the business side comes to you with a clear conscience, because they are not the ones paying. Once the money is in claims operations' budget line, Kevin asks his own people "is it worth it" when he schedules, instead of asking you "can you." The first of the three mandatory mechanisms above has its budget line from this week.
**Week 23, the eval guardian.** Linda takes on three things, approving new golden cases, the weekly review, and annotation guide version control. She asks a good question. "When I approve or reject, what is the standard?" Your answer is the same standard as claim 2093 back then (the disagreement that blew up at the annotation session in Chapter 11). Write the disagreement down, find someone with decision rights to rule, put it in writing (Chapter 11). The first version number she signs onto the annotation guide is the moment the eval turns from a project artifact into her operating asset.
**Week 24, the last module.** After the team expanded, the release rhythm went from biweekly to weekly. At this rotating release, lead-writer rights for the merged view are handed over, and the line on the staffing sheet that says "lead-writer rights handed over before handoff" (Chapter 15, who leads the writing of the merged view's code moves from your team to the business line) is honored. The older engineer presents, and not one of the three explain-it questions is skipped. With that, gatekeeping rights for every module sit on the business line side, and the backward staffing account is settled. In every row of "whoever owns it later writes it now," the two columns are finally equal.
**Week 25, the "exceptions" test fails.** Run and configure pass without trouble. Then a class of claim nobody had seen arrives. A rainstorm, and dozens of interrelated auto damage claims pour into the queue at once, the missing documents identical, the risk signals tangled together, and the queue ranks them as independent claims, getting messier with every pass. The claims operations team's first reaction is to call you. You take the call, help work through it, and then write on the test sheet, failed. The problem is not that they cannot judge. It is that "call you" is their only escalation path. Two weeks of rework. Together you draw the escalation tree, what to check first (the decision trail and drift signals), who judges (Linda), who it escalates to beyond scope (Kevin, with that class of claims paused and routed to manual handling if needed), then two injection drills. Your phone number is deleted from the flowchart. Retest in week 27, another constructed class of new situation injected. The younger engineer checks the trail, Linda rules to route to manual, the incident case goes into the golden cases the same day. The phone does not ring.
**Week 26, the North Star is reached.** Evolve and teach pass in the same week. One small request goes from Kevin's scheduling to the rotating release without a single commit from you. A reviewer who joined in the week 19 expansion is onboarded by a veteran from Linda's team. The same week, the run chart's weekly median reaches -31%. -18% (pilot week 5) โ -22% (end of pilot) โ -31% (now), six consecutive points below the pilot-period median. Chapter 18's test, and this time it is not you saying it, it is Kevin writing it into his own impact memo. The charter's North Star of -30% is met. The last item of the graduation criteria is filled in the same week, the reassessment date set in week 18 comes due and settles, the system runs under production rules from this week, and L2 is settled. The revival condition for home property, "the first extension after the pilot North Star hits target" (Chapter 8), formally holds and goes on the agenda. That is the next chapter's story.
Which week each of the five passed, laid out as a table. All five pass in week 27, and by the last of the three rules, stage four counts from that week, not from week 26 when the North Star was reached.
| Capability | First Test | Result | Passed |
|------|------|------|------|
| Run | Week 25 | Passed | Week 25 |
| Configure | Week 25 | Passed | Week 25 |
| Exceptions | Week 25 | Failed, two weeks of rework | Retest passed in week 27 |
| Evolve | Week 26 | Passed | Week 26 |
| Teach | Week 26 | Passed | Week 26 |
**Week 29, Tuesday, the ladder moves up a rung.** The older engineer brings a proposal. Claims that hit a rule and carry no long-tail signal (signals the rules cannot cover, which need the LLM's judgment) get assigned by the queue straight into a reviewer's personal list, with no wait for manual confirmation. Chapter 8 said "evidence" comes in two forms, and he has both ready. The human acceptance record, the assignment suggestions for this class of claim have run above 99% acceptance for eight consecutive weeks, with overrides near zero. The machine verification loop, this path runs on the rules layer, a hit is reproducible, the errors are enumerable, a wrong assignment is reassigned in one click, and the trail covers the whole way.
The review follows that ladder. The three oversight questions are run through again, the scope decision log gains a line, Kevin signs. The unsafe row of the kill criteria is rewritten the same day, one claim of this class and it stops. The sentence written in week 9 in Chapter 14 is honored today.
You ask only one question the whole way. What about the long-tail claims? The answer is right. Claims the LLM backstops all stay on the advise layer. No automated outbound messages, no payout decisions, both red lines untouched. The ladder on the whiteboard climbs a rung for real for the first time, and the proposer is not you.
**Week 29, Friday, the last retrospective.** Linda chairs. On the agenda, a ruling on one override disagreement, two new golden cases, the monthly list update, one small request to schedule, and Kevin sets the priority. You sit in the seat nearest the door, the seat the claims-ops IT engineer sat in four months ago, and say nothing the whole time. There is no farewell ceremony. The meeting ends, the system keeps running, and that is the acceptance.
After the meeting you hand Kevin a one-page responsibility transfer agreement. One page. Your write access drops to read-only, you are removed from the system's oncall (on-duty) rotation and from the standing attendee list of the retrospective, and a three-month response window is kept, taking only P1 incidents and exceptions the escalation tree ran to the end and did not catch. Closing the window needs written confirmation from the business owner and the ops owner. No confirmation means extension by default, and extension by default means permanent ops. P1 is the top level of the incident scale, and unsafe is P1. Outcome ladder L4 reached. Chapter 1 said the deliverer's achievement is settled only at L3/L4, and this project's achievement is settled only now. And "able to step out of the daily" opens two doors, distilling the Anchor & Helm approach into a reusable playbook (an operating manual you can follow step by step, Chapter 23), and your own next, bigger field (Chapter 26).
This matters to you too. Inside a company, stepping out of the daily is also the pass to the next rung of a career. Sitting on one system as the irreplaceable person, you never get the next level of complexity.
## Failure Modes
**1. Handoff = sending documents.** Handoff week produces a fifty-page ops manual, two training sessions, photos for the record. Documents are visible, acceptable, reportable. Capability is none of the three. A project acceptance checklist can always list "documents" and never "judgment," and with no external acceptance as a check inside a company, the checklist tilts even further toward documents, and writing documents is ten times cheaper than training people, so under closeout pressure the cheap one wins. But documents are the shadow of capability, not the capability. However complete the shadow's shape, it cannot hold the system up. The test, look at the ratio of nouns (documents, accounts, code) to verbs (can respond, can judge, can evolve) on your handoff checklist.
**2. Turning the capability tests into sign-offs.** Every box on the checklist ticked, both sides signed, ceremony complete. Three weeks later, at the first real failure, the business side's first reaction is still to call. A real test is expensive for both sides. The business side has to put in people and time and risk the embarrassment of "flopping the demo." A signature is cheap for both. The business side saves the effort, you get scheduled onto the next project a day earlier, Owen frees a person a day earlier. The two comforts stack, and "Show Me" decays into "we all understand it now." Anchor & Helm's failure in week 25 was precisely the test doing its job. The value of a test is not in passing. It is in exposing. The test, whether every item on the test sheet has a drill date and an injected scenario after it. An item with only a signature and no scenario is a sign-off.
**3. Your team becomes permanent ops.** Two years after the project "ended," you are still handling its alerts. The dependency trap is comfortable for three parties. The business side has no worries (someone handles it when things break), Owen feels the team is needed (headcount is easy to get at the review), you are needed (who does not like being needed). Nobody has any reason to break the equilibrium, except the system itself. It has an organizational version too. Project stacks on project, the Digital Center becomes the ops department for every AI system in the company, and capacity for new projects is zero. Knowledge stops being passed on, the owner is permanently missing, and it dies the day you resign. The test, check the calendar and the transfer ledger. If the date for stepping out of the daily has never appeared in any plan, or this system's line still has blanks, you are already in the trap.
**4. Starting the handoff too late.** "Knowledge transfer" starts the month before you are pulled onto the next project, there is only time to send documents, and it decays into failure mode 1. The payoff of a handoff settles at the end of a project, and every day in the middle has something more urgent. It is important and not urgent, so it always gives way, until only the last four weeks are left. The right answer was written in Chapter 15. Co-build is the handoff that starts on day one. A handoff that starts in the "handoff period" can only ever hand over a legacy. The test, check in which week the first gatekeeping rights were handed over. If that line is blank, or it is scheduled for the week you get pulled onto the next project, you are already late.
**5. The clean break.** Zero response after the handoff ceremony, "go to ops from now on." Three weeks later a small failure that ten minutes could have fixed is not caught by the new team, confidence collapses, the front line drifts back to the old process, and the system is abandoned. The internal root cause is the opposite of this one. In a clean break nobody takes it, inside a company everyone assumes you will take it, and both lead to the same result. Response has no owner, not because you refuse, but because you are still in the company, everyone assumes you will take it, so nobody formally takes it over. Ops schedules no oncall, the business side draws no escalation tree, your phone number stays in everyone's contacts. Default takeover and the clean break are two faces of one thing. Neither writes response down as a responsibility with an owner, a duration, and a scope. And the new team's confidence is most fragile in the first three months. The outcome of the first failure they face alone decides whether they dare touch the system or route around it from then on. The response window is part of the handoff design, not a favor done in passing. The last stretch of gradual withdrawal is still gradual, not a jump off a cliff. The test, whether the responsibility transfer agreement states the length of the response window, which two kinds of things it takes, and who confirms its closing in writing. Missing any one, and it is the clean break, or its internal version, the break that never comes.
!!! note "Vendor View"
The vendor side's exit is a physical event. The response window expires, external collaborator accounts are revoked, access goes to zero, people leave, and the date is set by the contract, so nobody can extend it by default. Your exit is an event of responsibility. You are always on the premises, so the responsibility transfer agreement has to do for you what the contract does there.
## Next Monday
1. Make an "only I can do this" list. What in this project can nobody but you do? Every line is a single point of failure, and a handoff debt.
2. Draw the owner placement sheet ([Template 22.1](../appendices/template-22-handoff.md)). System owner (budget and priorities, usually the business owner), maintenance owner (the person who responds in daily operation), eval guardian, AI dependency owner (who evaluates and signs when the model or a dependency changes), one real name in each of the four rows. A row you cannot fill is a cause-of-death candidate for the system. While you are at it, ask the PMO whether the company has an AI system transfer ledger. If not, use this sheet to start its first line, and leave the blanks flagged red.
3. Check whether your handoff plan has the three AI items. Who adds golden cases, who re-verifies when the model changes, who chairs the review. If none is there, you are still using a traditional IT handoff template.
4. Pick the one of the five capabilities you are least sure of, and put its Show Me on the calendar this week, with the scenario to inject and who is in the room settled. While it runs, keep your hands behind your back. If it fails, congratulations. You found it during the test period, not after stepping out of the daily.
**Want an agent to get you started?** In the repo you set up following [Start Here](../index.md), paste this to your coding agent:
```text
In the repo/ directory of the the-last-mile repository, help me with the Chapter 22 Next Monday actions. First run python3 templates/handoff/check_handoff.py
with the built-in sample to show me the five self-sufficiency tests check output, then open checklist.yaml and explain each item. Then build two sheets. The "only I can do this" list
I dictate and you record. The four rows of the owner placement sheet (system owner, maintenance owner, eval guardian, AI dependency owner) get real names from me. Leave the ones I cannot fill blank and flagged red,
do not guess names for me. Ask me about the three AI items one by one (who adds golden cases, who re-verifies when the model changes, who chairs the review). Which one gets the Show Me booking is my pick.
If any command errors, stop and show me the output.
```
---
## Chapter Kit
- **Judgment frameworks.** The five self-sufficiency tests (run / configure / exceptions / evolve / teach, test method = Show Me); the four-stage handoff cadence (your team leads โ shared lead โ business side leads โ stepping out of the daily, each stage sets who chairs the retrospective / who touches production / who answers outside questions); the three AI handoff items (continuous eval updating / re-verification on model and dependency changes / decision trail review)
- **Templates.** [Template 22](../appendices/template-22-handoff.md), Handoff Plan, Five Self-Sufficiency Tests, Handoff Cadence Sheet
- **Key judgments**
- "Dependency is an asset in the service business and a liability for the deliverer."
- "Documents are the shadow of capability, not the capability."
- "The handoff began on day one of co-build."
- "The best exit. The meeting ends, the system keeps running, and you said nothing the whole time."
---
# 23 ยท The Pattern Library: Turn Hero Stories into Assets
!!! info "Companion Templates"
๐ [Chapter Template](../appendices/template-23-pattern-extraction.md) ยท ๐ [Template Library](../appendices/template-library-index.md)
> **The Challenge.** Every project feels like starting from zero, and the team gets through on individual heroics. How do you make the second project cost half of what the first one did?
>
> **What You Will Be Able to Do.** Use the field learning flywheel to turn a project retrospective into reusable assets. Build a pattern library that will not rot, using the four asset classes and the three admission criteria. Before the next project starts, tell apart what can be carried over as is and what has to be dug out from zero.
---
> **Part VI Navigation.** Chapters 23 and 24 are two angles on the same asset inventory, Monday of week 30. Chapter 25 goes back to the quarterly business review in week 28, and Chapter 26 is the weekend of week 30.
> The quarterly business review, where you say no (week 28, Chapter 25) โ the last retrospective (Friday of week 29, Chapter 22) โ the asset inventory (Monday of week 30, this chapter and Chapter 24) โ Swiftway's whiteboard (week 30, this chapter) โ two ledgers (the weekend of week 30, Chapter 26) โ week 1 of the home queue (week 34, this chapter)
## Week 30, Another Whiteboard
The week after the last retrospective at Anchor & Helm, you are standing in a meeting room at Swiftway Logistics. Swiftway is another subsidiary of the Group, a mid-sized logistics business, and the Group has routed its ask to your team. The ask, in their own words: "Build an AI dispatch assistant."
The VP of Operations spends twenty minutes on the current state. Exception waybills pile up at the dispatch desk, drivers chase for assignments in the group chat, a waybill marked "delivered" in the system is not delivered and "in transit" is not in transit, and the dispatch team lead keeps a scheduling spreadsheet that the whole team treats as the only real one.
You stand at the whiteboard for ten minutes without writing a word. In the tenth minute you understand why nothing comes. Ideas are not missing. The ideas are suspiciously familiar. Exception waybills piling up is exception claims piling up. Drivers chasing for assignments is customers chasing claims. A status field you cannot trust is the core system's "in progress." The dispatch team lead's private Excel is the sheet Linda kept for six years.
You have seen this project before. It is Anchor & Helm with the nouns swapped.
The problem is that this "seen it before" lives, right now, in the back of one head, yours. When the Group routed the ask over, the schedule was set as starting from zero, and it eats into your team's headcount for the year. If the person standing at this whiteboard today were any colleague of yours who never worked Anchor & Helm, he would start from week 1 interviews and step into every hole you stepped into over those twenty-nine weeks, one for one. The organization does not know it already knows how to run this class of project. It only knows that one person pulled one off. That is the difference between a hero story and an organizational asset. This chapter is about turning what sits in the back of your head into what sits on the shelf.
## Why This Is Hard: Custom on the Surface, Reusable Underneath
A deliverer's work looks custom every time. Different industry, different systems, different vocabulary, different people across the table. So the organization assumes none of it transfers, and your team degrades into supplying a body to each department in turn. Headcount is fixed, so one more department means overtime, or a refusal, and the price of refusing is in Chapter 25. Scale produces no leverage, and capacity is locked to headcount. That is the ceiling of an internal delivery team. What separates this role from that ceiling is a bet that a reusable layer exists across projects, and that it thickens the more you accumulate (Palantir, mentioned in Chapter 2, is the company that turned that bet into a method, and the name of this role is one it made popular).
Where is the reusable layer? Not in the code. Move the exception queue's code to Swiftway and it dies on the field names on day one. What is reusable is the judgment structure. The exceptions queue and the waybill queue share one skeleton. Find the private source of truth โ three-way reconciliation (the system field, the manual sheet, and the person handling it, cross-checked) โ action queue โ decision trail โ adoption mechanisms. It took you twenty-nine weeks at Anchor & Helm to see this structure whole. At Swiftway's whiteboard it took ten minutes.
Two structural facts make it hard. The first, a judgment structure hides in people by nature, not in things. The things a project leaves behind, the repo, the documents, the configs, go to the business line under clause one of the co-build agreement (Chapter 15), and they all sit on the Group's own GitLab, and the organization forgets anyway, for three reasons. A reorganization comes, teams split and merge, and nobody claims the assets. Repos scatter across subsidiaries and outside contractors, and nobody can say how many copies exist. The judgment structure was never written down, the code is there and "why it was designed this way" is not. The day someone changes posts, all three fire at once. The second structural fact, nobody owns admission. The retrospective ends and people are scheduled onto the next department's ask. Accumulating assets is important and not urgent, the same lesion as starting the handoff too late (Chapter 22's failure mode, where the important and not urgent always gives way). Without an explicit mechanism, every generation of deliverers reinvents the same wheel, and the organization pays full price for each reinvention.
## Prior Art, and What AI Changed
**Maister's practice economics.** David Maister's argument in *Managing the Professional Service Firm* (paraphrased), how much output a professional services organization can move is decided not by how smart it is but by its leverage structure, how much delivery by non-senior people a senior person's judgment can move. Leverage comes from two things, reusable IP (methods, templates, tools) and the master-apprentice ladder (judgment passing down the ladder). The argument does not care who employs you. A team with neither, however good, is a set of expensive sole traders sharing a department name.
**The SRE runbook tradition.** Google's SRE practice writes down the knowledge of firefighting at midnight as a runbook. The engineer on duty facing an alert opens what a predecessor wrote, check this first, then that, escalate at this point. The core insight is that turning the hero moment into a checklist raises the organization's floor from the worst person on the team to the checklist. A pattern library is the same thing at project scale.
**What did AI change? A double-edged blade.**
The good side, the carrier of an asset moved up. A runbook can only be read by a person. A pattern library can also be fed to a coding agent. An agent carrying the queue skeleton, the decision trail schema, and the reconciliation checklist in its context starts the next project at 80% of a skeleton instead of an empty folder. For the first time an asset can become productive capacity directly, without passing through a person reading it and retelling it.
The bad side, the cost of pollution collapsed to zero. One prompt gets AI to produce a professional-looking playbook, complete in structure, standard in vocabulary, with no field provenance at all. On the shelf it looks exactly like a real asset, right up to the day the next project starts from it and it blows up. So the first principle of building a library in the AI era is that the admission standard matters more than the library. A pattern with no field provenance is a negative asset, worse than an empty shelf, because it occupies trust in the shape of an asset.
## The Core Framework: The Flywheel, the Four Classes, the Three Criteria
Two definitions first. **Pattern**, a reusable practice that still holds once the field context is stripped off a specific project. **The pattern library**, an organization-level store that accumulates patterns under one admission standard. A full set of patterns for one class of project, from discovery (finding out how things actually are) through to handoff, is together called a **playbook**.
**The field learning flywheel** makes accumulation a loop rather than a one-time action.
```
deploy (deliver the project)
โ capture (the closeout retrospective produces candidate patterns)
โ generalize (strip the field context, write where it applies and where it does not)
โ reuse (field-test it on the next project)
โ feed back and revise (the test results go back into the asset) โ back to deploy
```
The flywheel breaks most easily at two points. Capture has no owner (the retrospective ends and everyone leaves), and reuse has no feedback (it works badly, nobody fixes it, and the asset stays at 1.0 forever). So every step needs a named owner, the same as every metric on the metric tree (Chapter 18).
Inside a company there is one more break. The moment of capture never arrives at all. There is no closeout event. The first project has not wrapped up and the second department's ask is already on the schedule. You have to manufacture the moment, and there are two ways to hang it. Hang it on the first week after the responsibility transfer agreement (Chapter 22) is signed, or hang it on a retrospective a fixed number of weeks after launch, written into the project approval resolution so the meeting happens automatically when the date comes due, without waiting for anyone to remember. Anchor & Helm used the first, Monday of week 30.
**The four asset classes.** Different assets are reused in different ways. Mix them together and the library becomes unusable.
| Class | What It Holds | Anchor & Helm Example | Reuse Profile |
|----|------|---------|---------|
| Template | Documents and process | The charter template, the three-layer probing script, the four adoption mechanisms | Easiest, usable across industries as is |
| Component | Code and schema | The queue skeleton, the decision trail schema, the extraction pipeline interface | Medium, needs adapting to the data source |
| Judgment rule | Transferable judgments | "Chase priority โ risk priority", "If you cannot say it, do not merge it" | Carried away in one sentence, with the boundary attached |
| Metric model | The structure of a metric definition | The three tiers of the metric tree, the error severity vocabulary | Swap the North Star, keep the structure |
**The three admission criteria.** All three pass or it does not go on the shelf.
1. **At least one field validation.** Running in a demo does not count. Surviving on the business side's floor counts.
2. **A written boundary.** Where it does not apply is worth more than where it does.
3. **A named owner.** An asset with no owner rots in six months. The vocabulary goes stale, the links break, it drifts from reality, and then it starts misleading people. What is frightening about rot is not that the asset stops working. It is that it keeps being cited after it stops working.
Two rules that exist only inside a company. One, the owner column often holds a real name from another team, the platform team or the engineering productivity team, and you have no authority to hand another team an owner role. There are only two routes to a claim. A shared manager settles it in one meeting, or you push the library into a carrier that already exists, the company knowledge base, the platform team's component repo, the architecture review checklist, so an existing owner picks it up along the way. Two, the library has to declare its force level, reference, recommended, or mandatory, one of the three, written on the library's front page. Declare nothing and it is taken as mandatory by default, and every time a business department finds it awkward to use, that is on you.
Turn the flywheel one more notch and the carrier of an asset can move up another level, straight into an agent. Teams doing deployment engineering are already on this road (from several teams' talks in 2026, paraphrased. Some give every delivery engineer a delivery-cycle assistant agent, some let an agent take over the first half of the way from gathering requirements to writing the spec, and some make automating themselves out of the job a team goal). Put inside this framework, it is not a fifth asset class. It is the executable form of the four classes. The reconciliation checklist grows into a reconciliation agent, the three-layer probing script grows into an interview preparation agent.
Two rules do not move. The three admission criteria still apply, and field provenance is waived for nothing. What an agent produces is always a draft, and the standard and the boundary are ruled on by a person. A deliverer who uses AI to amplify the business side's front line should also use agents to amplify himself. In Maister's leverage structure this is the third lever, after reusable IP (methods, templates, tools) and the master-apprentice ladder (judgment passing down the ladder).
## From Anchor & Helm to Swiftway: A Reuse Field Test
**Monday of week 30, the team's own asset inventory.** It is a different thing from the week 28 account of the launch given to the business side (Chapter 24 writes this meeting from another angle). That one had Grant Whitmore in the room, it was the account for the business side, and it came first. This one has only your team, hung on the first Monday after the responsibility transfer agreement was signed. The queue's weekly retrospective carries on as usual, and the last one you attended (Friday of week 29) came after that account, with this inventory last of all. The meeting has one rule. Do not ask how the project went. Ask only what here the next project can still use. That is capture.
First on the admission list, in the component class, is the extraction pipeline interface. In Chapter 8 you said one sentence to the claims-ops IT engineer, "Whoever writes it keeps the interfaces clean. That is the first brick of the future platform." At the time it was a line for catching enthusiasm. Now it cashes out literally. Emails in, schema out, validation, the spot-check tool. The interface has nothing to do with the line of insurance. Strip off Anchor & Helm's fields and it takes waybills. First brick, admitted. The rest go in by the four classes.
- Components, the queue skeleton and the decision trail schema.
- Templates, the three-way reconciliation checklist, the four adoption mechanisms, the three-layer probing script.
- Judgment rules, "chase priority โ risk priority," with a boundary of any ranking scenario where the people being served can apply pressure.
- Metric models, the metric tree structure and the error severity vocabulary.
The same meeting went through the three lines on Chapter 8's revival list. Home property, whose revival condition "the first extension after the pilot North Star hits target" held in week 26, is on Anchor & Helm's own agenda. Full automation still moves up the decision rights ladder one layer at a time, each layer decided on its own, and is not opened here. The general platform's revival condition reads "after the second slice lands, distill what is common and revisit." Nobody knew then who the second slice would be. Now it is known, Swiftway. And distilling what is common is itself the pattern library. The way that revival condition is honored is by becoming this chapter.
**Weeks 31 to 32, Swiftway discovery, two weeks.** Anchor & Helm took six weeks from taking it on to settling the reconciliation. Swiftway took two. These carried over as is. The queue skeleton (waybills replace exception claims, the column structure unchanged, suggested priority, suggested next action, reason, owner, Human Call, not one column missing). The three-way reconciliation method (the system field, the dispatch team lead's Excel, asking the handler, and in week 1 it located the largest single class of waybill status mismatch). The four adoption mechanisms (pick the super-user by influence, and the dispatch team lead who maintains the scheduling spreadsheet is Swiftway's Linda candidate). Plus the three-layer probing method for mining tacit knowledge, anchor on an instance, compare, boundary counterexample, works as is.
The grunt work of reconciliation was not written from scratch either. The admitted reconciliation checklist, together with its boundary, was fed to a coding agent, which generated the sampling and comparison scripts. People ruled only on the standard and on who owns the source of truth.
What has to be dug from zero is just as clear. The golden cases are all new. A waybill is not a claim, and not one carries over. The rules of thumb reset to zero. The dispatch team lead's "one look and I know this one is going bad" runs on a combined feel for shipper, route, and driver, and does not overlap the repair shop list by a single word. You mined for two afternoons with the three-layer probing method, and every rule that came out was new. Not one of them came from Anchor & Helm.
| Class | Carried Over As Is | Has to Reset to Zero |
|----|---------|---------|
| Template | The three-way reconciliation method, the four adoption mechanisms, the three-layer probing method | What goes into the templates. The super-user is picked again for the new field |
| Component | The queue skeleton, waybills replacing exception claims, not one column missing | Fields and data sources. Move them over and they die on the field names on day one |
| Judgment rule | Sentences that hold across scenarios, like "chase priority โ risk priority," with the boundary attached | The rules of thumb. The dispatch team lead's combined feel does not overlap the repair shop list by a single word |
| Metric model | The metric tree structure and the error severity vocabulary | The golden cases. A waybill is not a claim, and not one carries over |
> **Tacit knowledge does not transfer. The method for mining it transfers.** Chapter 6's line, "'Seeing that the Excel exists' is in no dataset," gets its second confirmation here.
There is one asset an outside team can never accumulate, and it belongs to the metric model class. The auto and home property queues share one decision trail schema, so the override reason codes can be laid side by side, and the overlapping part is a company-level answer to which kinds of judgment AI is not trusted on. The waybill queue will be the third once it launches. An outside team is bound by confidentiality and cannot aggregate across companies. You do not have to wait for anyone's approval. All three queues sit in the Group's own warehouse, and one query gets you the table. Register it under the metric model class, owner column, yourself.
The four weeks saved are the pattern library's price. They land on no invoice. They land on capacity. Your headcount is fixed, so four weeks cash out as a few more departments covered this year, and the half the second department is cheaper by is capacity, not money. Only you keep this account. The business side sees a system go live and does not see the four weeks you did not spend. Owen Hartley's inventory sheet carries headcount, not weeks. So the bookkeeping has to go into the asset register. Log the weeks saved on every reuse, and at the annual review that column is a number you can write down, with no need for vague phrasing like "efficiency gains." The part where the second project is cheaper than the first is the interest the organizational asset pays.
These two weeks held one more deliberate arrangement. The person running the reconciliation changed. The engineer on your team who had worked Anchor & Helm took the reconciliation checklist and led the writing, and you only gatekept. The three explain-it questions moved up a level, from the claims-ops IT engineer explaining it to you (Chapter 15) to your apprentice explaining it to you. This is where Maister's leverage structure lands, reusable IP ร the master-apprentice ladder = capacity. The IP saves the apprentice from working it out from nothing, and the ladder saves the IP from having to be executed by your own hands. As for Swiftway's dual Group-and-subsidiary structure, that is another kind of training ground, and the last chapter takes it up.
**Week 34, an email.** From Kevin Doyle, subject line "week 1 of the home queue." The body says the home property exceptions queue went live and the run chart put down its first point. The whole thing reused that one set of patterns from the auto queue. Linda led the annotation build, supplying the method and not the judgments, with the disagreement records, the ruling process, and the guide format all following the rules set back then, and the judgments themselves coming from home property's own reviewers. The claims-ops IT engineer changed the configuration and shipped it, with little new code written. The golden cases were gathered from zero, and they knew on their own that they had to be, with nobody reminding them.
The escalation tree caught two new exceptions, and one of them was not free. Priorities were ranked wrong for half a day, and the escalation path caught it and routed it to manual only in the afternoon. The balancing metrics raised a small bump that day and came back down the next, and the incident case went into the golden cases that evening. The last line of the email. "Did not have to bother you this time."
You checked. The exit date you set yourself in the responsibility transfer agreement (Chapter 22) has two months left. P1s received, zero. Nobody is coming to revoke your access. You wrote that date yourself, and you are the only one watching the countdown.
This email is proof of three things. First, the end of Chapter 22 said the home property revival condition formally held and went on the agenda, and that it was the next chapter's story. This is the story. Second, L4 was defined in Chapter 1, the business side can run, maintain, and improve it on its own. The highest form of improving it is opening a second front with the same method, and fixing bugs is only the start. Third, the playbook's first reuse was done by the business side, with no hand from you. The best acceptance a pattern library gets is that it runs when you are not in the room. And home property's rules of thumb were dug from zero all over again, the third confirmation that tacit knowledge does not transfer. This time, not even the digging was yours.
## Failure Modes
**1. Keeping the code, not the judgment.** The library is all repo links and starter kits, with not one sentence of "why." Code is visible, countable, and can be signed off. A judgment structure is none of the three. What gets kept is decided by what is easy to appraise, not by what is worth something. But code becomes scrap in another industry (the field names kill it on day one), and a judgment structure does not. The test, if the library has fewer words on when this does not apply than on code comments, what you left behind is a specimen, not an asset.
**2. Assets with no owner, and the library becomes a document graveyard.** The last team's "best practices" lie on the wiki, and nobody knows which of them are still alive. Building a library is a project. Keeping one is operations. Organizations will approve and decorate projects, and will not reserve headcount for operations. Leave the owner column blank and the updates sit forever behind everything urgent. Six months later nobody dares use anything in the library. Someone burned once by a stale asset never comes back. The test, pull three assets from the library at random and check whether the owner column holds a real name and which project the last revision came back with. The ones you cannot answer for are already in the graveyard.
**3. Treating custom work as reuse, forcing the answer to fit.** Taking Anchor & Helm's answers to Swiftway and deploying straight past discovery. The pleasure of reuse eats the sense of boundary. Taste the sweetness of saving four weeks once and you want to save six. And a pattern's boundary is written in the least conspicuous column of the register, which the people racing a deadline never read. The result is the five gaps stepped into again on the spot. The golden cases are auto's, the rules are Linda's, and the adoption mechanisms have no super-user of Swiftway's own. The test, if nothing on your "reuse" list is ruled "has to be redone," what you are doing is forcing a fit, not reusing.
**4. Retrospectives that write down only successes.** The library is all winning moves and not one way to die. A retrospective is a social occasion, and writing down a failure names the person responsible, so failures get smoothed into three lines of "lessons learned" and the information about where the real holes are evaporates at the meeting room door. But failure modes are the most valuable entries in the library. Every project's successful path looks different. The holes are the same batch of holes. The book you are reading has a "Failure Modes" section in every chapter. That is not a writing preference. That is the same discipline applied to a book. The test, count the output of your last retrospective, how many winning moves and how many ways to die. Ways to die at zero, and those notes are a victory record, not an asset.
**5. Filling the library with AI-generated entries.** The problem has two layers, that the count rises fast, and that what comes up has been validated by nobody. The library swells from 12 entries to 200 in one quarter, every one of them impeccably tidy. Once building the library carries an entry-count KPI, AI is the perfect tool for running up that KPI, and professional-looking is exactly what a large model is best at producing. Trust in a library is indivisible. Someone burned by one fake pattern doubts all two hundred. The defense is not deleting the library. The defense is at the door. Library admission has to be able to answer which project, which weeks, who was in the room, where the validation record is ([Template 23.3](../appendices/template-23-pattern-extraction.md)). AI can draft, structure, and rewrite. The one thing it cannot supply is field provenance. The test, look at the library's growth rate. The stretch where the entry count grew faster than the project count is the stretch with no field provenance.
!!! note "Vendor View"
The vendor side's pattern library carries a price tag. The four weeks saved land on the quote, four weeks of day rate not billed, and the pattern library is what gives a firm the nerve to price on outcome instead of day rate. Your four weeks land on capacity, an account only you keep, so the bookkeeping has to go into the register.
## Next Monday
1. Reopen the retrospective of your most recently finished project for 30 minutes and ask one question only. "What here can the next project still use?" Force out at least one entry in each of the four classes, and strip the field context with [Template 23.1](../appendices/template-23-pattern-extraction.md).
2. Write a "does not apply when" line for every candidate asset. The one you cannot write is not admitted. If you cannot write the boundary, you do not understand it yet.
3. Find a "best practice" on the team wiki that nobody has touched in over six months, and check its owner and its last field validation. Both come back missing, run it through 23.3. Re-claim it, or delete it.
4. Pick one validated component-class asset, feed it together with its boundary to a coding agent, and have it generate the next project's skeleton in a sandbox. See how far that is from ready to start. The gap is the part of the asset you have not written down clearly yet.
**Want an agent to get you started?** In the repo you set up following [Start Here](../index.md), paste this to your coding agent:
```text
In the repo/ directory of the the-last-mile repository, help me with the Chapter 23 Next Monday actions. First run python3 templates/pattern-library/downgrade_stale.py
with the built-in sample to show the demotion output for stale assets, then open asset-register.md and pattern-template.md. The candidate assets I name during the retrospective,
you strip of field context in the 23.1 format and write up as drafts. The "does not apply when" line for each is mine to write, and the one I cannot write you mark "not admitted."
Last, read through packing-example/ and tell me which fields are needed to feed an asset together with its boundary to a coding agent.
Which asset gets fed is my pick. If any command errors, stop and show me the output.
```
---
## Chapter Kit
- **Judgment frameworks.** The field learning flywheel (deploy โ capture โ generalize โ reuse โ feed back and revise); the four asset classes (template / component / judgment rule / metric model); the three admission criteria (one field validation / a written boundary / a named owner)
- **Templates.** [Template 23](../appendices/template-23-pattern-extraction.md), Pattern Extraction Sheet, Asset Register, Library Admission Checklist
- **Key judgments**
- "What is reusable is the judgment structure, not the code."
- "Tacit knowledge does not transfer. The method for mining it transfers."
- "An asset with no owner rots in six months."
- "The four weeks the second project saves are the pattern library's price."
- "The admission standard matters more than the library. A pattern with no field provenance is pollution."
---
# 24 ยท From the Field to the Platform: Field-to-Product
!!! info "Companion Templates"
๐ [Chapter Template](../appendices/template-24-f2p-memo.md) ยท ๐ [Template Library](../appendices/template-library-index.md)
> **The Challenge.** The project has stepped out of the daily, and the custom code and dozen-odd platform judgments the field produced are rotting in the business line's repo and inside your head. The platform team knows none of it, and the next project writes it all again from zero.
>
> **What You Will Be Able to Do.** Run the asset recovery checklist at the asset inventory and pick out what should cross the river. Use the field-to-product memo's four sections plus the three filters to turn a field signal into a proposal the platform team can consume. Give n=1 signals a candidate register instead of forwarding them or throwing them away.
---
## Monday of Week 30, the Internal Closeout Retrospective
Friday of week 29, the last retrospective at Anchor & Helm, and you said nothing the whole time (Chapter 22). Monday morning of week 30, before you leave for Swiftway, your team holds its own asset inventory. Nobody asked for this meeting. The account of the launch given to the business side happened in week 28, and that one was Claims Operations' account to Anchor & Helm Insurance. This one has no business side in the room, and one item on the agenda. What did this project leave us?
Start with the "delivered assets" column (this meeting is the same asset inventory as Chapter 23's, written there from another angle). You put the module list from the Anchor & Helm repo on the screen and ask the same question of every item. This piece here, will the next similar project write it again? When the count was done the room went quiet for a few seconds. About 40% of the code is the kind any similar project will need. The decision trail schema (suggestion and fact tables kept separate, the six fields, reason codes, Chapters 12 and 17), the eval harness (the management and replay scripts for the golden cases), the queue skeleton (column structure, the rollup view, escalation rules as configuration). All of it sits in the repo of that one line, Anchor & Helm's claims operations.
The repo belonging to the business line is not the mistake. Clause one of the co-build agreement should be signed exactly that way (Chapter 15). The mistake is losing contact inside the same git. The same Group, the same GitLab, code anyone can click open, so everyone assumes recovery has already happened. It has not. That 40% of structure was never pulled out and never declared reusable, which is the same as it not existing. This hides better than a gap between two companies, because the code is plainly right there.
The more expensive item sits outside the code. Across twenty-nine weeks you banked a dozen or so judgments about what the platform's next version should have. The front line wants a queue, not a dialog box. The reason code is the cheapest ground-truth signal there is. The decision trail is a hard requirement and logs are only its by-product. The Group platform team knows none of it. Nobody failed at his job. Between the field and the platform runs a river nobody owns.
## Why This Is Hard: Symmetrical Complaints, a Missing Mechanism
One bank of this river is your delivery team, the other is the Group platform team. The complaints on the two banks are strikingly symmetrical. You say the platform does not understand the field and builds its roadmap behind closed doors. The platform says field requests are all noise, that every department calls its own request the most general one, and that every deliverer shouts loudest for his own department. Both sides are right and neither has a way out, because what is missing is a mechanism for filtering and translation. Goodwill is short on neither bank. Which signals are worth crossing, in what format they cross, who rules. The internal river has one more drop in it. The platform team's KPI is adoption rate, yours is business outcome, and the two appraisal sheets do not share a single line. They want you to use off-the-shelf components, the capability you want is not on their roadmap, and neither side has any obligation to give way first.
With no mechanism, a signal has two default exits. It stays locked in the deliverer's head (the organization pays the whole bill on the day he moves posts or leaves), or it gets forwarded to the platform as is (the platform, once flooded, learns to listen to none of it, and even the real signals stop crossing). Done well, your team is the platform's most expensive radar. You stand behind real users and see what no piece of research can see. Done badly, your team is a source of noise on the platform's roadmap. The radar and the noise source are the same people with the same observations. The whole difference is the mechanism.
With no shared KPI across the two banks, you have to build the channel yourself, and the way to build it is to put the f2p memo (field-to-product memo, defined in the next section) into a fixed line of the quarterly report-out. How many memos went out this quarter, what n each carried, how many the platform took, one row of numbers, read by Owen Hartley, copied to the platform lead. What gets written into the report-out gets done. Everything else is a favor done in passing.
## Prior Art, and What AI Changed
**The Amazon tradition, write backward from customer value.** The generalization case in an f2p memo has one hard rule. You may not start from what you built at Anchor & Helm. You must write backward from which other departments will hit this. Cannot name a second one, cannot send the memo. The rule comes from the PR/FAQ discipline recorded in *Working Backwards* (Bryar and Carr), where a project starts by writing a future press release and the FAQ forces answers to the hardest questions, whose behavior changes because of this and on what grounds. Its spirit is the direction of the argument. Write backward from customer value, not forward from what we made.
**The lean tradition, genchi genbutsu (go and see the actual thing in the actual place).** Toyota's discipline. Whoever makes the decision has to go to the field and see with his own eyes, and a secondhand report is no substitute. Landed on a platform role it becomes one hard rule. The platform's decision-makers must go into the field on a cycle. You can book it for him. You cannot look for him.
**What did the AI era change? Generality can be measured for the first time.** "This request is very general" used to be a piece of rhetoric, and whoever was loudest was the most general. AI delivery changed three things.
1. **Decision trail data aggregated across departments is platform-level eval insight.** Every project's override reason codes describe which class of judgment people do not trust AI on (Chapter 17). Put two departments' reason code distributions side by side, and the overlap is the platform's problem to solve, not the project's.
2. **A cross-department diff of prompts and rules is the platform candidate list.** A diff means comparing two departments' versions line by line to see what changed and what did not. Diff the extraction prompts of Anchor & Helm and Swiftway. What changes is the industry vocabulary and the field names. What does not change (output structure, validation method, the backstop on exceptions) is the platform. The line between the custom layer and the platform layer can be drawn with a diff for the first time.
3. **n turns from a claim into a number.** "How general is it" can now be argued quantitatively. That is an opportunity and a discipline at once. In an era where you can count, not counting is laziness.
All three are stronger inside a company. Anchor & Helm's and Swiftway's decision trails both land in the Group's warehouse, so putting two reason code distributions side by side is one query away, with no cross-company data export approval to clear, and a prompt diff inside the same GitLab is close to free. An outside delivery team can never do either of those two things, because every one of its projects sits behind a company wall. This is a pure advantage, and cashing it has one precondition. The decision trail schemas of the two projects have to match, which is exactly the problem the first memo in the Anchor & Helm section below sets out to solve.
## The Framework: The Field-to-Product Memo and the Three Filters
> **The field-to-product memo (f2p memo) is the one-page proposal that flows a field signal back to the platform team. Four sections, phenomenon โ the generalization case โ the product recommendation โ the cost of not doing it.**
| Section | The Question It Answers | Rules |
|----|------------|------|
| **Phenomenon** | Which field, at what frequency, on what evidence | Quote the actual words and the decision trail data, not impressions |
| **The generalization case** | Which other departments or scenes will hit this, and what n is | Write backward from the side with the need. n must be verifiable |
| **The product recommendation** | What **capability** you are asking for | A capability, not a feature. "Ship a decision trail schema component" is a capability. "Add an export button for Anchor & Helm" is a feature. Features belong to projects, capabilities belong to the platform |
| **The cost of not doing it** | What every project pays over again, what is lost structurally | Quantify it as engineering effort. Say the most expensive one out loud |
The format discipline is not newly invented. The f2p memo obeys the one-page rule of Chapters 13 and 19, one page, lead with the conclusion, an ask at the end, delivered 48 hours before the meeting. The full statement of those four sits in Chapter 19, and this chapter only uses them, it does not re-teach them. The only difference is the reader, an executive on the business side becomes the Group platform lead. What is new is three thresholds for sending, called the three filters here. Fail any one and it does not go out.
1. **Report only at nโฅ2.** n=1 is custom work. nโฅ2 is a signal. n counts departments. The same need appearing in two business departments is what makes it the platform's problem to solve. However reasonable one department's request is, it goes into the candidate register first. The candidate register is where n=1 signals are kept, each with its trigger condition written down, promoted to a memo the second time it appears. The register has five fields, signal, source (department / proposer / date), n, status, and the trigger or revival condition.
The live example is at Anchor & Helm. In week 23 the survey team lead proposed at the retrospective, "repair shop response time, should you not build a number for that too" (Chapter 20). Real need, real proposer, real instinct for data, but it has appeared at Anchor & Helm exactly once. It goes into the register with one line of trigger condition, "any second project that raises an 'outside party response time' need promotes this to a memo." The register is an incubator for n=1, not a wastebasket.
2. **Describe the judgment structure, not the interface.** "The business side wants a Gantt chart" does not cross the river, because a Gantt chart is this department's interface habit. "The business side needs to see waiting attributed across steps" is the transferable judgment structure. An interface description turns the platform team into an outsourcing shop. A structural description is what gives them design room.
3. **Attach what you are willing to give up for it.** A recommendation with no trade-off is a wish list item, and the platform team gets dozens a day. Which of your own requests will you give up priority on for this one? No answer means you do not much believe in it yourself.
The third one is the hardest inside a company. An outside delivery team bargains with its own schedule priority. Your hand is different, and it holds four cards. Your department goes first as the pilot. Your department retires its own ready-made custom implementation and switches to the platform version. Your department supplies people to co-build. The hardest card, file a PR straight into the platform repo and turn the proposal into a fact on the ground, which routes around the roadmap negotiation entirely, and which an outside team cannot do. Cannot produce a single one of the four, and what you want is the platform doing your work, not the platform providing a capability.
## At Anchor & Helm: One On-Site Day, Three Memos
**Tuesday of week 24, an on-site day.** Halfway through the handoff period, you cleared it with Kevin Doyle and pulled the Group platform team's lead for this capability over to Anchor & Helm to sit for a day. An outside team that wants to bring its own product manager into the field first has to clear a visitor confidentiality filing. You do not. The platform lead walks down two floors. In the morning he sat behind Linda Marsh's team and watched the queue. Eight people, and not one of them "asked" the system anything. Scan the row, confirm, override, two seconds to pick a reason code, and pick up the phone when a hard claim needs a judgment.
In the afternoon he sat in on the retrospective Linda chairs. One override disagreement, two new golden cases, one revision to the annotation guide, all through in five minutes. Three days after he got back he cut a feature already in development, a "conversational query entry point" that would let users ask a claim's status in natural language. He had sat in the field for one morning and seen with his own eyes that the user this feature assumed does not exist. The front line does not want to ask the system questions. The front line wants to be told the next action.
The chatbot that died at the start of this book (Chapter 1) nearly came back to life inside a capability the Group platform was already building. What saved it was one day in the field, which no requirements document could have done. A negative signal from the field is worth as much as a positive one, and not building this feature saved a full quarter of development.
**Week 32, memo #1.** Work starts on Swiftway's queue skeleton, and you catch yourself hand-writing the same decision trail schema a second time. The nโฅ2 filter fires on the spot. The memo went out that evening, four sections plus the give-up line, not one of them skipped.
> **F2P Memo #1, the decision trail component.** To the platform lead, week 32
> **Phenomenon.** Projects at two subsidiaries, Anchor & Helm (insurance claims) and Swiftway (logistics dispatch), hand-wrote the same thing one after the other. Suggestions and facts stored in separate tables, the suggestion table append-only, six-field decision trails, override reason codes. About one week of engineering each time, and the structures are near identical.
> **The generalization case.** The decision trail is a structural requirement of any "AI suggests, a person decides" system, and it has nothing to do with the industry. As long as decision rights stop at the advise layer, the system has to answer "who made the call and on what grounds." n=2, the two subsidiaries Anchor & Helm and Swiftway, and the second implementation needed no industry adaptation at all.
> **The product recommendation (a capability).** A decision trail component. The trail schema plus reason code configuration plus a review query that splits overrides by category, available to any project out of the box.
> **The cost of not doing it.** A week of rewriting on every new project is the small bill. The big one is that project schemas diverge and cross-department override aggregation analysis cannot be built at all. Both subsidiaries' decision trails sit in the Group warehouse. An outside team can never do this, we could have, and it is the only source of platform-level eval insight.
> **Willing to give up for it.** If the schedule conflicts, this team sends one person into the platform repo to co-build, and withdraws its request this quarter for custom columns in the queue interface.
The two below give only the key points. The full form is in memo #1.
**Memo #2, productizing the eval harness.** Golden cases are managed in spreadsheets at both subsidiaries today, added and removed by hand, versioned by filename, replayed by script. Every department that runs an eval will hit this, two subsidiaries already have, n=2 clears the line, and pushed further by judgment structure, nobody who manages golden cases in a spreadsheet escapes it. The capability recommended, a golden case management interface with versions, the annotation disagreement log, admission of incident cases, and replay reports.
**Memo #3, the queue scaffold.** Column structure as configuration, the rollup view, escalation rules. Refused by the platform team. It conflicts with the existing roadmap and will not be scheduled this quarter. You did not argue. You made a record. The memo goes into the register together with the reason for refusal, plus one line of revival condition (events, not dates, the same form as Chapter 8's scope decision log). Half a year later the platform roadmap shifted and somebody dug it back out of the register, but that comes later. Only one sentence needs keeping right now. A memo that was refused but recorded is still an asset. One that was refused and never recorded is the one written for nothing.
The interface at the inventory end needs a line too. From here on the internal closeout retrospective carries one more fixed segment, closeout asset recovery, four classes, code, documents, judgment, metrics, which is Chapter 23's four asset classes seen from the counting side (code = component, documents = template, judgment = judgment rule, metrics = metric model).
Give every item a destination. There are only three exits. Admit it to the library, send a memo, or leave it in the business line repo. The first two are not mutually exclusive. Anything at platform capability level usually goes into the library and gets reported as well, and a refused report still stays in the library. Only "leave it in the business line repo" is an exclusive exit. The blank sheet is [Template 24.2](../appendices/template-24-f2p-memo.md).
| Class | What Goes Through, Item by Item | Test Question | Destination |
|----|--------------|----------|------|
| **Code** | Decision trail schema / eval harness / queue skeleton / extraction pipeline structure / configuration patterns | Does it still hold in another industry | Admit and send a memo. Business-line-specific fields and integrations stay in the business line repo |
| **Documents** | charter / review packet / the three memos / runbook / annotation guide framework | The framework transfers, the content does not | The framework is admitted. The content stays in the business line repo |
| **Judgment** | Key judgments / failure modes and incident retrospectives / annotation disagreement rulings as precedent | Tacit knowledge does not transfer. The method for mining it transfers | The method is admitted. The rules of thumb themselves stay in the business line |
| **Metrics** | Metric definition tree / thresholds and acceptance lines / reason code enumeration / balancing metrics | The structure is general, the vocabulary is business-line-specific | The structure is admitted or rides along with a memo. The vocabulary stays in the business line |
The "admit it" exit runs through Chapter 23's admission review and the [Template 23](../appendices/template-23-pattern-extraction.md) register, and this chapter does not repeat that standard. Every item has to have a destination. "Leave it for now" is not allowed.
What gets recovered is structure. The code files belonged to the business line all along. And structure, as Chapter 23 put it, belongs to no repo at all.
There is one more bill after a memo is taken up, and the recovery sheet does not carry it. On the day the platform version launches, Anchor & Helm's custom implementation will not vanish by itself. The dual-track period where both run side by side is a quarter at the short end and a year at the long end, and three things go unmanaged if nobody asks. Who maintains the old one, when it gets retired, who makes the call. The way to write it is three more cells on this system's row in the PMO's AI system transfer ledger (Chapter 22), the maintainer of the old implementation, the retirement condition (the day the platform version clears Anchor & Helm's golden cases), and who makes the call (the system owner, Kevin Doyle). Leave any one of the three blank and you are feeding two systems at once.
One note in passing. Anyone who gets good at bridging this river gains one more career route. The starting point of a platform product owner is where field-to-product ends (Chapter 26).
## Failure Modes
**1. Every department reports its requests straight to the platform.** You forward the business side's actual words to the platform as they were said, and the more diligent you are the more responsible you look. Forwarding is free, filtering is expensive, and filtering means carrying the "I suppressed the business side's request" responsibility. On top of that you live with the business side day in and day out, and empathy naturally amplifies "my department is the most general one." Inside a company the temptation is larger still. You and the platform team work for the same company, raising a request is one message away, forwarding is far cheaper than writing a memo, and the channel burns out that much faster. The platform team's rational response once flooded is to listen to none of it, the channel is burned, and even the real signals stop crossing. This is exactly what the three filters are for. Filtering is the sender's responsibility. Fail to filter and the recipient will filter for you by not listening. The test, count the field signals you forwarded to the platform last quarter and see how many carried an n and a give-up line. None at all, and what you were doing was forwarding, not reporting.
**2. The platform team never goes into the field.** All the platform's knowledge of the field comes from your account of it. A trip to the field is permanently "important, not urgent" on the platform's schedule. The better-hidden part is that any account distorts. You will translate what you saw into requirement language without noticing, and the negative signals are the first thing lost in translation. "Nobody asks the system anything" does not form a requirement, so it never gets passed on, and yet it is precisely what cut a quarter of wrong development. The fix is to write the on-site day into the rhythm of the platform role. However good you get at retelling, it will not cover this. Genchi genbutsu is an institution, not a nice story. The test, go through your platform counterpart's calendar. If a whole year holds no record of one full day sitting behind real users, your shared knowledge of the field is entirely secondhand.
**3. Nobody recovers the custom code.** Roughly 40% of general structure rots in the business line repo and the next project rewrites it. An outside team at least has a closeout date, with appraisal and celebration both settling on acceptance day, and the recovery window opens in the unattended stretch where the project is dead and the next one has started, narrow, but it does open. Inside a company there is not even that window. There is no closeout event, people slide from one system to the next, and recovery never fires. The fix is to hang three triggers on recovery, the fixed retrospective three months after launch, the moment before any team member moves posts, and the quarterly asset inventory. Hang it on any one of them and recovery has a date. This is the same move as Chapter 14's resource gates and Chapter 22's three mandatory handoff mechanisms, manufacturing a moment that comes due automatically for something with no outside date. The one at Anchor & Helm hung on the third, and it happened to open on the morning you left for Swiftway, two triggers stacked on the same day, which is why it opened at all. The test, go through the retrospective notes of the last project you stepped out of the daily on. If "asset recovery" is not an item on the agenda, recovery never happened.
## Next Monday
1. Open the repo of the last project you stepped out of the daily on and go through the module list, asking of every item, "will the next project rewrite this?" The number you get is your "rot-in-the-field rate." Use [Template 24.2](../appendices/template-24-f2p-memo.md) to give each item a destination, admit / send a memo / leave it in the business line repo. "Leave it for now" is not allowed.
2. Find the one field discovery you are most sure the platform should build, and run it through the three filters. What is n? Judgment structure or interface description? What are you willing to give up for it? All three clear, and you write it up as a one-page memo with [Template 24.1](../appendices/template-24-f2p-memo.md) and send it.
3. Anything that cleared only two of them goes into a candidate register. One line of phenomenon, one line of trigger condition. Refused memos go into the table too, with the reason for refusal and the revival condition attached.
4. Book an on-site day for your platform counterpart. Not a demo day, a morning sitting behind real users. You are responsible for booking it, not for retelling it.
**Want an agent to get you started?** In the repo you set up following [Start Here](../index.md), paste this to your coding agent:
```text
In the repo/ directory of the the-last-mile repository, help me with the Chapter 24 Next Monday actions. Open the repo of the last project I stepped out of the daily on,
list the modules, "will the next project rewrite this" is mine to answer item by item, and you compute the rot-in-the-field rate. Then copy templates/asset-recovery/recovery-checklist.md
over to me, the destination of each item (admit, send a memo, leave it in the business line repo) is mine to set, and no "leave it for now" allowed. For the field discovery I am most sure about,
you only ask the three filters, what is n, judgment structure or interface description, what am I willing to give up, and only after all three clear do you draft from f2p-memo-template.md,
with candidates registered in candidates.csv. If any command errors, stop and show me the output.
```
---
## Chapter Kit
- **Judgment frameworks.** The field-to-product memo's four sections (phenomenon โ the generalization case โ the product recommendation (a capability, not a feature) โ the cost of not doing it) + the three filters (report only at nโฅ2 / by judgment structure, not by interface / attach what you will give up); closeout asset recovery (four classes, code / documents / judgment / metrics, ร three exits)
- **Templates.** [Template 24](../appendices/template-24-f2p-memo.md), Field-to-Product Memo and Closeout Asset Recovery Checklist
- **Key judgments**
- "Done well, your team is the platform's most expensive radar. Done badly, it is a source of noise on the roadmap."
- "n=1 is custom work. nโฅ2 is a signal."
- "A negative signal from the field is worth as much as a positive one."
- "A memo that was refused but recorded is still an asset."
---
# 25 ยท When to Say No: Intake Rules and Red Lines
!!! info "Companion Templates"
๐ [Chapter Template](../appendices/template-25-intake-redlines.md) ยท ๐ [Template Library](../appendices/template-library-index.md)
> **The Challenge.** A project you know is a trap gets pushed at you, and the person proposing it is the one who trusts you most. How do you refuse without burning the relationship? One layer down, how do you get the organization to take fewer traps, instead of betting every time on your being in the room?
>
> **What You Will Be Able to Do.** Use an organization-level intake rubric (five dimensions of scoring plus three red lines) to rule on whether a project should be taken, before any demo. Use the three steps to no to translate refusal from a relationship event into a judgment event. Build a kill register, so the judgment behind saying no evolves the way an eval does.
---
## Week 28, the Light in His Eyes
Week 28, Anchor & Helm's quarterly business review, which is also the meeting where Claims Operations accounts to Anchor & Helm for the launch (Chapters 23 and 24). The first one since the auto queue hit target. The run chart is on the wall. Week 26, first-touch handling time for auto exceptions, -31%, North Star met. Grant Whitmore, Kevin Doyle, and Victor Reyes are all in the room. Halfway through the agenda, Grant closes the minutes.
"Next step, full automation of the whole claims process. AI does loss assessment and payout approval directly. I have had Finance set the budget aside, and I will handle the board."
There is light in his eyes. And this is not an outsider's enthusiasm. The framework he cites is the one you taught him. "You said decision rights move up on evidence. The evidence is hanging on the wall now."
This is the hardest kind of situation to say no in. The budget is there, so you cannot hide behind resources. The results are real, so you cannot say the timing is not right. The proposer is your most important sponsor, and every "I will tell you when the time comes" over the past half year, he remembers. And the five-question framework in your head (Chapter 7) is already sounding the alarm. Risk, 1. The payout decision is a regulatory red line, the error is irreversible, and this company's human oversight capacity cannot carry a move of decision rights like that.
He asked about the same direction once in Chapter 8. That one was easy. The project had just started and one ladder drawing was enough. This time he comes with a budget and results in hand, ready to call it. The hardest no is the one you say to the person who trusts you most. This chapter is about how to get it said, and how to keep your organization from needing you personally to say it every time.
## Why This Is Hard: Refusal Has Only Two Default Meanings
Refusal inside an organization has only two default meanings. "I do not want to help you," or "I cannot." The first damages the relationship, the second damages your professional standing. So most people facing a project pushed at them take a third road that looks harmless, take it now and sort it out later. Three months on the project dies on the five gaps, and the relationship and the standing go down together.
The deliverer has to build a third meaning. The evidence says not yet. You have used it twice already. In Chapter 7 the elimination table went on the table, you pointed at the scores and read the evidence column, and the service chatbot forwarded by the board was eliminated. That was its first rehearsal. The hold and watch option in the Chapter 19 impact memo is its midpoint form, turning "no expansion yet" into an option with its own expiry conditions, handed to the person who decides. The mechanism is the same sentence both times. Translate refusal from a relationship event into a judgment event. A relationship event is settled by position and rank. A judgment event is settled by evidence, and nobody in the room has evidence harder than yours.
But an individual who can say no solves only half the problem. The other half is at organization scale. The demo inflation of Chapter 1 and the idea inflation of Chapter 7 reappear unchanged at the organization's intake layer, and because your service carries no price, the cost of proposing an idea is zero, so idea inflation floods toward you at triple speed. After this win at Anchor & Helm, your team becomes the mounting point for every AI wish in the company, new projects arrive every week, each with a budget and a light in someone's eyes. A delivery team with no intake discipline has a settled fate. Packed with impossible projects, record L0 output, L3/L4 achievement at zero. Personal scripting cannot hold back structural traffic. Only a mechanism can. This chapter raises "the evidence says not yet" from a personal script to an organizational mechanism.
## Prior Art, and What AI Changed
**Solution Selling's qualification discipline, upgraded this time into an organizational process.** Chapter 7 quoted the iron rule of the Bosworth line (paraphrased). Chasing the wrong deal costs ten times what losing one costs, so the qualification standard goes up front. The real core of that work is not the script. It is moving the qualification standard out of personal craft and into a process. Resources go in only after a review, the standard is written on paper, and nobody gets an exemption on enthusiasm. What the delivery team copies is exactly this. Make the intake standard explicit, rule on it at a meeting, leave a trail for the retrospective.
And inside a company you need not build a new gate for it. The company is already full of ready-made carriers, the project approval review, the architecture review, the investment committee, the quarterly ask review, each with its own agenda, its own minutes, and a group of people who already have to sign. Hanging the three red lines and the five dimensions on one of them takes a tenth of the effort of opening your own meeting, and it is harder to route around than a new meeting would be.
**Maister's economics of practice, which projects not to take.** *Managing the Professional Service Firm* (its ideas paraphrased) carries an account most service organizations ignore. The wrong project destroys more than profit. It burns the team, because your best people spend themselves on a project that was going to fail, and then they leave. It burns the reputation, because other departments and the executive layer define who you are by the projects you have done. So the intake standard is not an operating detail. It is the team's position. Are we a delivery team or a demo crew. What you take is what you are.
**The AI era changed two things.** First, the moment for saying no was forced earlier. In traditional software, "can be demoed" and "can be delivered" were not far apart, and the demo was itself the filter. AI tore a crack between the two. The five gaps (Chapter 1) are all invisible at the demo stage, so the infeasible projects are often the ones that make the most stunning demos. Wait for the demo and then say no? Too late. The budget is approved, the executive has reported it up, the applause has already sounded, and sunk cost has taken everyone hostage, you included. The gate has to sit before the demo, and that is exactly where intake sits.
Second, red lines acquired an external source of hardness. In traditional projects, "not taking it" was mostly an economic judgment. For AI projects, part of "not taking it" comes from regulation, ethics, and irreversible harm. Those boundaries are not settled by weighing things up inside your organization, so they have no business appearing in any scoring sheet. A red line is a boundary, not a preference. Inside a company, use that hardness to the full. State red lines through regulation, compliance, audit, and the Group's AI governance policy wherever you can, so that "this is not ours to vote on" is literally true, and you do not have to carry an executive's will on your personal judgment.
## The Framework: An Organization-Level Intake Rubric, Five Dimensions of Scoring Plus Three Red Lines
> **Intake rubric, the organization's ruling structure at the door of a project. Clear the three red lines first (any one of them vetoes), then score the five dimensions, and rule on a candidate project, take it, do not take it, or take it with conditions, before any demo.** (Intake, the review that decides which asks the team takes on.)
The three verdicts turn on two things, whether the red lines were cleared and whether any dimension is a weak link. All three red lines clear and no weak dimension, take it. Any red line touched, do not take it, and however high the other dimensions score, nobody looks. Red lines all clear with exactly one dimension at the floor, that is take it with conditions. Taking it with conditions means writing three things down on the spot, which dimension is short, who closes it by when, and who re-scores it once it is closed. If those three cannot be written, record it as not taken then and there. Do not leave an ownerless candidate sitting in the sheet.
Its relation to the five-question framework of Chapter 7 fits in one line. The five questions assess "is this use case any good," the deliverer's personal tool in the field. The intake rubric assesses "can this organization carry it, is it worth carrying," the gate at the team's door. These five dimensions are intake's five, and what they grade is a project. They are not the same set as the five-axis capability radar that grades people in Chapters 3 and 26. You can see the five questions' shadow in the five dimensions, but there are two upgrades, taken up after the table.
| Dimension | Test Question |
|------|----------|
| **Strategic value** | Win it, and what does the organization get? One department's thanks, or agenda power and standard-setting power over a class of problem? (the organizational version of Pain + ROI) |
| **Data readiness** | How far do the key data exist, how reachable are they, how far can they be trusted? Have they been reconciled? (inherited straight from the Data question) |
| **Owner in place** | Is there someone on the business side who signs for the outcome? Who backstops the errors, who maintains it after launch? (the ownership gap at the door) |
| **Production path** | From demo to production, do review, oversight, and change discipline get through? Or can only the demo live? |
| **Reuse potential** | What can go into the pattern library (Chapter 23)? How much cheaper is the second delivery? (Maister's leverage) |
Scoring follows the weakest-link logic of Chapter 7. 1 to 5, any dimension โค2 does not enter the schedule, and a re-evaluation condition gets written. No weighting, no averaging. The same rule bites harder in Chapter 7, where a question at the floor eliminates the use case. Here it only keeps the project out of the schedule. That notch is deliberately looser. Assessing whether an organization can carry something runs one notch wider than assessing a single use case.
Both verdicts show up in the meeting that follows. Loss assessment and payout approval touches a red line, not taken. "Move decision rights up one more layer" gets its condition and its re-scoring moment nailed down, and that one goes through as take it with conditions.
The first upgrade is the two dimensions of organizational economics added in (strategic value, reuse potential). An individual assesses the use case. An organization also has to assess what this fight does to itself. The second upgrade matters more. The five questions' Risk column has disappeared here. It was promoted into a red line and no longer gets scored. Three red lines, each with one test question.
1. **Automated decisions at an irreversible-harm step.** Can the harm from the single worst output be taken back? Anchor & Helm's "never touch payout decisions" (Chapter 8) is this one's instance.
2. **A move of decision rights with no human oversight capacity behind it.** After the move, is there still someone with the time to look, the ability to judge, and the authority to stop it (the three oversight questions, Chapter 12)?
3. **Outward-facing output in a regulatory grey zone.** When this output causes a dispute, what do you answer the regulator with?
Each of the three has a typical misjudgment. On the first, reading "a person signs at the end" as not automated. If the signer has no time to actually look, nobody is reviewing, and that is automated. On the second, reading a timetable of "assist first, automate later" as oversight capacity already built. On the third, reading "the regulator has not said no" as "the regulator said yes."
The red lines carry exactly one rule of discipline. Any one of them vetoes, and none of them enters the scoring. The meaning of a scoring sheet is trade-off, and everything that enters the sheet is by default compensable by a high score elsewhere. A red line gets its force from outside. A regulator does not fine you one unit less because strategic value scored 5, and irreversible harm does not turn reversible because reuse potential is high. Pulling Risk out of the scoring columns and making it an off-sheet veto is not a wording change. It admits a fact. Some risks are there to be managed. Some risks are there to be avoided.
One step remains after the ruling, saying the no out loud. Three steps.
1. **Affirm the goal.** Declare that what you refuse is the path, not the intention. "The direction is right, and we get there sooner or later."
2. **State the evidence.** Put the rubric on the table and go through it item by item. Put Risk in the language the other side cares about. Talk the regulatory account and the brand account, not model probabilities.
3. **Offer a path.** A step plus a condition. "First get to [measurable threshold]. When the data is there, decision rights move up one layer, each layer decided on its own."
The third step is where it succeeds or fails, and the rule is one sentence. A no must come with the conditions for a yes. An unconditional "no" is a refusal. A conditional "no" is a roadmap.
Saying no inside a company carries a notch more political cost than saying it outside. Refusing a project is not losing a piece of business. Refusing an executive's use case can be read as picking a side, and next quarter's budget for you sits at the desk beside his. This does not change the structure of the three steps. It raises the weight of the last two. The evidence in step two and the alternative path in step three cannot be skipped inside. A no carrying only an attitude gets you routed around, and a team that gets routed around cannot even keep its gate.
## At Anchor & Helm: The Page Goes on the Table
Back to the meeting room in week 28. You knew on the day the target was hit in week 26 that this proposal would come back. That line of revival condition in the Chapter 8 scope decision log (the revival list), "once the advise layer's acceptance data hits target, move up one layer at a time, each layer decided on its own," is owed a settlement. So there is a page in your bag.
"Half a sentence before the conclusion. You have not got the direction wrong. Efficiency across the whole process is what this project pointed at on day one, and if claims cost is to come down another notch, that really is the only way to go." Catch the goal first. Only then does a refusal earn the right to begin.
Then you put the page on the table and go through the five dimensions one row at a time. Strategic value, high, no argument. Data readiness, low. What the queue reconciled is process status data. Loss assessment and payout approval feed on a different set, repair hours, parts prices, the reading of survey images, and not one of them has been reconciled. Owner, "AI sets the loss. Who signs the payout approval?" Nobody answers. Then your finger stops on the three rows boxed out separately, outside the scoring area.
"This one does not get a score," you say. "Get one loss assessment wrong, the customer takes the payout approval to the regulator, and the nature of it becomes a payout decision the company made and got wrong. Calling it a system fault will not cover it. The regulatory consequence and the brand consequence, you know better than I do. I do not need another line on this page."
You did not say the model will get things wrong. Model probability is your ledger. Regulation and brand are his. Speak in his ledger and he had it in three seconds. Victor Reyes added half a sentence from the side. "Until the regulator's line on AI payout approval is out, this review does not pass with me."
"So what I recommend is rebuilding the steps, not dropping it." You turn the page over. "Next step is estimate assist. AI produces the loss estimate, the payout approver reviews it, the whole thing leaves a trail, the same skeleton as the exceptions queue. The step condition is nailed down. When estimate assist's override rate drops below 10%, when the approver changes fewer than one suggestion in ten, the data will tell us which classes of claim can have their decision rights moved up a layer. One layer at a time, each decided on its own. That is the concrete form of that revival condition from back then. When the data is there we go, instead of we will see. As for the payout decision itself, until the regulator speaks it is a red line, and it is not ours to vote on."
Grant studies the page for a while. His first sentence is, "The last ladder you drew me, I took to the board. Give me this page too." Then the second one.
"Then we go with your steps."
After the meeting he holds you back at the door for one more line. "We go with the steps you laid out. But the window with the board, I can only hold it two quarters. When the time comes, the estimate assist numbers had better be presentable."
There is no loser in the room, but the steps now carry a clock. This no could be said because of money saved up over half a year, and the script is only the wrapping. The honest report in Chapter 18, "one point short does not count," is why he does not doubt your evidence for a second. Self-orientation in the denominator of the Trust Equation (Chapter 5), how visibly you are working for yourself, is as low as it can go at the moment you push away a fully funded project.
The relationship took no damage, because the way you refused proved you were carrying the judgment on his behalf. What he wants is someone who can hold off a wrong decision for him. Executors who agree to everything, he has plenty of. The moment you say no is both a withdrawal against the trust balance and the next deposit.
## From One No to a Gate
From week 30 on, that page stops being your personal tool. You harden it, together with the retrospective on that meeting, into the team's intake process. A new candidate clears the three red lines first, then the five dimensions, then gets ruled on at a meeting. A killed candidate is not allowed to disappear. It goes into the kill register. Project, proposer, kill reason, revival condition, retrospective conclusion a year later, five columns (template at [Template 25.4](../appendices/template-25-intake-redlines.md)). The first two rows are ready-made. Full automation (loss assessment and payout approval), revival condition as above. Service chatbot (Chapter 7), re-evaluation condition carried over unchanged.
The sheet has an owner and a rhythm. You hold it from week 30, and after you leave it goes to whoever takes over intake. The fifth column gets a pass once a year, two questions per row, did it get built in the end, did the revival condition come true.
Whether this gate stands inside the company turns less on how good the sheet is than on where its force comes from. Three routes. The first is the cheapest, build no new review. Hang the three red lines and the five dimensions on the company's existing project approval review, architecture review, or quarterly ask review, and the gate comes with an agenda, minutes, and signatories. Second, have the CIO or the responsible VP endorse the standard once at a meeting, and after that "this is the company's intake standard" is not something you have to explain each time. The third is the hardest, write the three red lines into the company's AI governance policy. From then on the red lines are not your team's to vote on and not the proposer's to bid on.
Stronger than a gate is putting a price on the service. Where the company has showback (shows the bill only) or chargeback (actually deducts budget), the compute and labor bill is booked straight to the department using it (Chapter 14). Where it does not, use a non-monetary price. Every project supplies one full-time counterpart from the business side, named, written into the project approval resolution. The price need not be money. As long as the person with the idea has to put something up, inflation comes down by itself.
Owner in place is the hardest dimension to get inside a company. No written agreement can force the business side to produce a signer, and most people genuinely believe that "once it launches someone will look after it." Forcing out a named owner uses the same trade. You want our people and our time, so give us the person who signs for the outcome after launch, name into the project approval resolution. No name, and this dimension is a 1, and by weakest-link logic it does not enter the schedule. This is not obstruction. It turns owner from a courtesy into a claim, and what the claim buys is that after launch this system's priority and iteration schedule are set by him (Chapter 4).
There is one more account that holds only inside a company. The projects you turn down do not disappear. The business department can have it built by an outside firm, or buy a SaaS product, and when it goes badly the ops work still routes back to you, because you are the only people in the company who understand this stuff. You are not moving to a different company and starting over, so inside, "not taking it" does not equal "not my problem." This account pushes the internal no toward the take-it-with-conditions notch, keeping the project on a path you can see, which is cheaper than letting it take a worse road and picking it up afterward. That is also why, among the rubric's three verdicts, the middle one gets the most use inside.
The fifth column is empty right now, and it is the soul of the whole sheet. A year later you read back row by row. The kills you got right, someone else built it later and it died of exactly the cause you predicted. The kills you got wrong, the revival condition had long since come true, nobody re-scored it, and the opportunity went to someone else. Both conclusions feed back into the rubric. Whichever dimension keeps misjudging is the dimension whose test question you fix. This is the spirit of Chapter 11 reappearing at the intake layer. The judgment that judges intake also has to be judged. Without this column, saying no is the one decision in the organization that never gets checked, and it will stop at the level you are at today and never evolve again. Inside a company the sheet carries one more benefit. People come and go and it is still there, and you get to see what actually became of the projects you rejected, so the fifth column's evidence is easier to come by than it would be outside. It turns "why we did not do it back then" into organizational memory, and what it defends against is every change of executives stepping into the same hole again.
## Failure Modes
**1. Taking everything for next year's budget.** Half the team's schedule reads "let us build a demo and see." Inside a company there is no sales team, and nobody pushes you to hit a number. The pressure only changed source. A project an executive named is hard to push back. Next year's headcount and compute get requested on the strength of this year's presentable AI highlights, and your department's OKRs have a few "landed scenarios" written on them. What is truly fatal is that the shape of the account changed. Taking the work and delivering it used to be two groups of people, with benefit and cost on two separate ledgers. Inside, those two ledgers merged into one, and the taker and the deliverer are the same person. You take a project at the start of the year for the budget, knowing it is a trap, and a few months later you are the one who pays.
This is no longer misaligned incentives. It is a bet you placed on yourself, and the payoff on the day you bet is certain while the cost on the day you settle is still invisible, which makes it harder to quit than misalignment. With no gate, taking it is the rational choice in the moment. The consequences settle globally. The team turns into a demo factory, record L0 output, while the deliverer's unit of value settles only at L3/L4 (Chapter 1). Achievement goes to zero and the best people leave first. The test. Go down the running projects on the schedule one by one and ask how many can state the release criteria written down at intake. The ones that cannot are the ones that got in on enthusiasm and a budget narrative.
**2. Using delay instead of refusal.** "Great idea, the schedule is full, let us look next quarter." Delay defers the social cost of refusal and looks free in the moment. But it also cancels the judgment. The other side got no evidence and no step, only a brush-off, and a brush-off reads no differently from contempt. Next quarter the project comes back unchanged and your credibility is thinner than last time. Relationship and judgment both lose. Delay rolls the interest on the no into the principal. The test. Go through the projects you said "let us look next quarter" about in the past six months and count how many were actually re-scored the next quarter. If none were, what you used was delay, not refusal.
**3. Negotiable red lines.** The scoring sheet shows "compliance risk 2," and then strategic value's 5 averages it away. The meaning of the sheet is trade-off, and everything that enters it is by default compensable, while a red line's force comes precisely from not taking part in the trade-off. Make the first exception for a "specially important project" and the red line turns from a boundary into a price. After that every proposer comes to bid. Once a red line is in the scoring sheet, it is no longer a red line. The test. Look at which column your red lines live in right now. If they carry a score and can be averaged away by another dimension, they already are not red lines.
**4. Killing without a retrospective.** A rejected project is never mentioned again, and the register either does not exist or has a permanently empty fifth column. Organizations only review what they did. What was done has data, an owner, and meetings. What was not done produces nothing, and nobody schedules an agenda for it. So refusal becomes the one judgment in the organization exempt from inspection, nobody ever learns which kills were right and which were wrong, and intake's accuracy stops where it was on day one. The test. Pick any candidate rejected last year at random. If nobody can state its revival condition back then and nobody can say whether the condition holds today, this column is empty.
**5. Saying no with no alternative path.** The meeting ends, the other side leaves politely, and next time the project approval routes around you. A no with no conditions carries nothing but attitude. The other side cannot tell "this road is blocked" from "he does not want to do it," and by ordinary human nature only the second is left. From then on he stops bringing you his ideas to look over. The idea skips intake and goes straight to a budget, dies three months later, and the bill still gets charged to "AI does not work." What you lost is more than this project. It is the gate itself. The test. Read back the words you sent the last time you refused. If there is not one sentence in there about what evidence would reopen it, attitude is all the other side received.
## Next Monday
1. Take the three projects your team took on most recently and run them back through the five dimensions plus three red lines after the fact. Which one should not have been taken? Where is it stuck on the outcome ladder (Chapter 1) now? That is your evidence that your organization needs intake.
2. Write down your three red lines and check where they live now. The ones living in the scoring sheet, move them out, list them separately, and mark them "any one vetoes."
3. Build the kill register ([Template 25.4](../appendices/template-25-intake-redlines.md)). Add the candidates you said "no" or "let us look again" to in the past six months, fill in a revival condition for each, and send it back to the proposer.
4. Think back to the last time you refused. If you gave no condition, add a line now and send it. "What evidence, if it appears, makes this worth reopening."
**Want an agent to get you started?** In the repo you set up following [Start Here](../index.md), paste this to your coding agent:
```text
In the repo/ directory of the the-last-mile repository, help me with the Chapter 25 Next Monday actions. First run python3 templates/intake/revival_reminder.py
with the built-in sample to show the reminder for a revival condition coming due, then copy intake-scorecard.md into the working directory I name. The three projects the team
took on most recently I score myself on the five dimensions plus three red lines. You only build the sheet and keep the record, and which one should not have been taken is
mine to judge. The three red lines are mine to write, and you check whether they are listed separately and marked "any one vetoes." Build the kill register per 25.4. The
revival condition on each row is mine to fill in, and who it goes back to is mine to decide. If any command errors, stop and show me the output.
```
---
## Chapter Kit
- **Judgment frameworks.** The organization-level intake rubric (five dimensions, strategic value / data readiness / owner in place / production path / reuse potential; three red lines, automated decisions at an irreversible-harm step / a move of decision rights with no human oversight capacity behind it / outward-facing output in a regulatory grey zone, any one vetoes, none of them scored); the three steps to no (affirm the goal โ state the evidence โ offer a path)
- **Templates.** [Template 25](../appendices/template-25-intake-redlines.md), Intake Rubric, Red Line List, Saying No Scripts, Kill Register
- **Key judgments**
- "Translate refusal from a relationship event into a judgment event."
- "Wait for the demo to say no and sunk cost has already taken everyone hostage."
- "A no must come with the conditions for a yes. An unconditional 'no' is a refusal. A conditional 'no' is a roadmap."
- "Once a red line is in the scoring sheet, it is no longer a red line."
---
# 26 ยท The Deliverer's Career Path: F1 to F5
!!! info "Companion Templates"
๐ [Chapter Template](../appendices/template-03-capability.md) (F1 to F5 anchors and the annual growth agreement, 3.7 and 3.8) ยท ๐ [Template Library](../appendices/template-library-index.md)
> **The Challenge.** Three years in this work, then five, and what are you? Will you just be a firefighter who costs more every year?
>
> **What You Will Be Able to Do.** Rate yourself with the F1 to F5 behavior anchors. Run an axis-by-axis self-assessment on the project you just closed out and find "the part you cannot carry." Write a one-page annual growth agreement that turns your next project into a step in career design.
---
## The Weekend of Week 30, Two Ledgers
Friday of week 29, at the last retrospective at Anchor & Helm, you said nothing the whole time (Chapter 22). By week 30 you are already standing in a meeting room at Swiftway Logistics. Swiftway is another subsidiary of the Group, and the Group routed its ask to your team. The ask is "build an AI dispatch assistant," and what happened on that whiteboard is Chapter 23's story.
This weekend you spread out the Anchor & Helm project file and work a second ledger. The project's ledger is settled, -31%, L4, an impact memo Kevin Doyle wrote himself. This ledger is about you. In that project, which stretches did you run through cleanly, and which did you paper over with luck and overtime?
The honest list is not pretty.
- Cleanly. The Field MVP (Chapter 0's 2-hour minimal prototype), the charter, co-building the eval, the three memos. You could run every one of these again tomorrow at any company.
- On luck. Victor Reyes going from blocker to ally started with a pre-mortem that happened to draw him into asking for a meeting (Chapter 1). What if he had not come?
- On overtime. The week the cost ran away (Chapter 16). The four fixes were clean, but the bill had gone to four times the estimate before anyone noticed, and that in itself is a failure to anticipate.
- Taking the hit. The moment the survey team lead put his pen down (Chapter 20). You only learned to price visibility after the fact.
Then the colder question. If tomorrow's project is three times as complex, five business lines, three executives opposed to each other, data ten times worse, can you carry it?
You cannot. And that is exactly the good news. The part you cannot carry is your next level. This chapter gives you a ruler for measuring it.
## Why This Is Hard: An Invisible Ladder
Growth in this role has a paradox. There is no ready-made job ladder for it. It does not fit the seniorโstaff line of code depth, because your code keeps getting zeroed out by AI and by handoffs. It does not fit the manager line either, because half of your "team" belongs to the business side. Ask "how does this role get promoted" and most organizations have no answer, and the Digital Center's job ladder has no row for it either.
People who cannot see the ladder move sideways. Five years, ten projects, case studies ready on demand, and all ten projects at the same difficulty, one sponsor, one business line, data bad in roughly the same way. That is experience 1ร10, not 10ร1. Project count is not experience. Complexity jumps are.
So the ladder has to be drawn first. The deliverer grows along one dimension only, the ceiling on the complexity of judgment, the ceiling on the ambiguity, the intensity of conflicting interests, and the organizational complexity you can handle. Anchor & Helm had exactly one Grant Whitmore from start to finish. Two Grant Whitmores pulling against each other is a problem of another order.
One distinction has to be set first. This book already has a ladder, the outcome ladder (L0 to L4, Chapter 1), and that one is the system's ladder. F1 to F5 in this chapter is the person's ladder. A system's outcome is counted in L. A person's capability is counted in F. Anchor & Helm climbed to L4, and what you graduated is F2. The two ladders do not convert into each other. An F4 on the wrong project can still watch the system die at L1. An F2 on the right project can take a system to L4, which is exactly what you just did.
## Prior Art, and What AI Changed
**Maister, growth in professional services is an apprenticeship.** The three-layer structure in *Managing the Professional Service Firm*, finder (winning the work), minder (managing the relationship), grinder (doing the work), served Chapter 3 as a lens on time leverage (which layer a senior person's time should go to). Here it is a career path. A professional grows by moving up those three layers, and moving up runs on apprenticeship, which no classroom can teach. You follow a senior person on a real project and watch how he makes trade-offs where there is no right answer.
**The Dreyfus model of skill acquisition, from rules to situations.** The Dreyfus brothers' account of skill acquisition (novice โ advanced beginner โ competent โ proficient โ expert, five levels from beginner to expert) turns on one change in kind. A novice runs on rules, sees X and does Y. An expert runs on situations, stops retrieving rules, and simply "sees" what this situation is missing. This book gave you dozens of checklists and templates, and they are the novice's handrail. The book's goal is precisely that one day you stop checking them line by line. That is internalization, not forgetting.
**What did the AI era change?** AI is compressing, precisely, the value of technical execution inside F1 and F2 (the next section's table gives the levels). The part a coding agent can do devalues fastest, and the rate of devaluation is the rate at which models iterate. Two consequences.
First, the bottom step of the apprenticeship is disappearing. A newcomer used to trade grunt work for the right to be in the room. The grunt work now belongs to AI, and F1's ticket has changed from "can do the work" to "can review the work AI did."
Second, the whole growth curve tilts toward judgment and relationships. The earlier you move your assets out of execution and into judgment, the safer you are.
The other side of the same fact is that the deliverer is one of the few engineering roles in the AI era that gets more valuable the longer you do it. Chapter 6 said it. "Seeing that the Excel exists" is in no dataset. Scaled up to a career, it still holds. Field judgment is the hardest thing for a model to learn. It grows in specific rooms, the Thursday Linda Marsh stopped backing up the Excel, the weekly meeting where the survey team lead pushed back. Models cannot read those moments. You were there.
## The Core Framework: F1 to F5, Five Levels ร Five Axes
F1 to F5 is the deliverer's capability rating (F = field, meaning the level of a person's capability). It invents no new coordinate system. It reuses Chapter 3's five-axis capability self-assessment radar, engineering depth, AI engineering, business grasp, narrative, field judgment. The five axes here are capability axes. Chapter 25's intake rubric has five of its own, the five dimensions that score whether to take a project, and those are a different subject. Every level is given as behavior anchors, with adjectives set aside. "Senior" and "can stand on his own" cannot fix a level. "Has done this thing or has not" can. The full five-level ร five-axis anchor table is in [Template 3](../appendices/template-03-capability.md), and here is the main line.
| Level | In One Sentence | Graduation Behavior Anchors |
|----|--------|--------------|
| **F1 Can Execute** | Delivers inside a given charter | Work inside the boundary done to quality and on time; can spot exceptions and escalate them; can review what AI produced |
| **F2 Can Deliver Alone** | Goes from a vague ask to L3 or above alone | One sponsor, one business line, from "build an AI assistant" to a system in daily use and ready to hand off |
| **F3 Can Handle a Complex Business Side** | Delivers inside a conflicting interest structure | Several business lines, sponsors pulling against each other, strong opposition; a charter that opposing parties will both sign |
| **F4 Can Lead and Productize** | Amplifies judgment with multi-project leverage | Accountable for several projects at once; assets in the pattern library with you as owner; field insight carried across into product capability (the work of Chapters 23 and 24) |
| **F5 Can Design the Practice** | Designs the mechanisms that grow other people | The intake mechanism, the economic model, and the talent pipeline came from you; because of you the organization takes on fewer bad projects and produces more F3s (Chapter 25's organizational mechanisms are a work sample of F5) |
Take one axis as a sample. The rating follows the lowest axis, and field judgment is the shortest plank in the self-assessment in the "At Anchor & Helm" section. Its five anchors are below, and the other four axes are in the same table.
| Level | Behavior Anchor |
|----|----------|
| F1 | Spots anomalies and escalates the same day; can recite the red line list and holds it |
| F2 | Ranks by irreversibility under time pressure; writes real causes of death in a pre-mortem; dares to write a readout (the conclusions report) that says "stop" |
| F3 | Judgment still holds under conflicting interests. When two decision-makers pull against each other, handles it by written principle instead of picking a side in the moment |
| F4 | Trade-offs at the portfolio level. Has judged, across several projects, which to save, which to kill, whom to send |
| F5 | Has said no to an organization-level opportunity and offered a way out; the kill register and the retrospective discipline were built by you (Chapter 25) |
Three rules for using it.
1. **The rating follows the lowest axis, not the highest**, the same as Chapter 3's rules for reading the chart. The five axes multiply, they do not add. Narrative at F4 and field judgment at F2 makes you F2.
2. **The evidence for moving up is having carried it. Having taken part does not count.** Serving as the F1 executor on an F3 project still grows F1 experience. The verbs in the anchors, signed, ruled, claimed, designed, must take you as the subject.
3. **Each level absorbs the one before it. It does not replace it.** An F4 still has to be able to deliver alone, only no longer by hand on everything. Whichever axis caves in, the person slides back to the level of that axis.
### Three Paths Out, What Comes After the F Levels
Above F3 the road forks. Three common paths out, each with its own capability transfer map.
**Internal incubation or a spinout.** Take the proven pattern with you and start a new business or spin out a company. This role is one of the best founder trainings there is, because it covers the whole distance. What you did in these two years is spread across several chapters, and put together it is exactly a founder's capability list. Find the real need (the five questions and discovery, Chapters 7 and 6), deliver an outcome someone is willing to pay for (L3/L4, Chapter 1), win resources and renewed funding (the Trust Equation and the three memos, Chapters 5 and 19). Every core move of a startup's first two years you have already run for real. What you have to add is a shift in mindset. From serving one business line to selling one product, from "what does this business line want" to "which thousand companies want the same thing" (Chapter 24's three filters are the exercise).
**Platform product owner.** Take Chapter 24's field-to-product translator all the way and you are the platform product owner, distilling what the business lines have in common into an internal platform. You take the seat carrying assets nobody else has, the field radar, the nโฅ2 discipline, and real usage evidence in the decision trail data, and every business line's decision trail data is yours to read on one sheet. What you have to add is a switch in the scale of trade-offs. From one business line's outcome to a platform roadmap, learning to say no to the field the way you once said no to the platform for the field. You used to say no to the platform on behalf of the front-line business. Now you sit on the platform side and have to learn to say no to a business line on behalf of the platform.
**Head of AI for the Group.** Stay in the field and take F3 to F5 all the way. On one side, technical judgment the executives of every business line trust. On the other, the designer of the whole company's intake, eval, and delivery standards, and Chapter 25's mechanisms are the work sample for this seat. This path is the least glamorous and the scarcest. It demands fluency in both technology and organization, and both sides are short of people who have it.
Each of the three paths has an original outside the company, founder, product leader, field CTO. The capability transfer maps are the same, and only the market where you cash them in differs. All three share one hole card, the judgment banked by walking the whole distance. All three are ways of cashing in F-level assets, and none of them counts as escaping this role.
There is one more direction, not a path out but still a variable career design has to account for. Titles are drifting. Across the industry the boundary between field delivery roles and product engineering is thinning, and some companies have already merged the two hiring bars and reporting lines (from what two AI companies' deployment teams shared in 2026, paraphrased). Product engineers increasingly face users directly, and field delivery people go increasingly deep into product. Job names inside a company drift the same way. The day your title becomes product engineer or platform engineer, or the reverse, is not a change of career. Chapter 2 said it. The label drifts, the core does not move. The five axes travel with you, and the complexity you have carried is never zeroed out.
## At Anchor & Helm: The F2 Graduation Exam and the Next Map
**Axis by axis.** Take the file you spread out over the weekend and run it through the five axes, accepting behavioral evidence only. The verdict comes in three grades. Solid, every anchor at that level has evidence with you as the subject, and the next level has none. Close to the next level, that level is solid and the next level's anchors have scattered evidence but not the full set, which does not change the rating and only tells you what the growth agreement should say. Barely, every anchor at that level has evidence, but some of it rests on luck or on after-the-fact repair, and in another field you might not carry it again. Put simply, solid is no holes, close to the next level is showing but not yet counted, and barely is just qualified but not certain to hold up somewhere else.
| Axis | Behavioral Evidence | Verdict |
|----|----------|------|
| Engineering depth | On a ten-year-old system under strict security constraints, took the merged view from design through to handing over lead-writer rights | F2 solid |
| AI engineering | Wrote the five-part eval spec (Chapter 11) alone, designed the monitoring surface alone, closed the loop on incidents and drift (Chapter 16) | F2 solid, close to F3 |
| Business grasp | Dug out the fourteen-step real workflow; -31% converted into money and taken to the board | F2 solid |
| Narrative | The three-memo system; bad news volunteered within 24 hours; argued "the evidence says not yet" until it was accepted | F2 solid, close to F3 |
| Field judgment | The red lines held, but the team lead's pushback was not anticipated, the cost blowout was handled by repair, and the Victor Reyes thread had a component of luck | F2 barely |
The conclusion, F2 graduated, with a complete evidence chain. Going from two vague sentences to L4 on your own is the question paper of the F2 graduation exam.
Chapter 25's intake mechanism does not rewrite that conclusion. Having drafted an F5 exercise once is not rating evidence. A rating follows complexity carried again and again, and one draft is only another form of "having taken part."
The hardest line on the diploma is the week 26 impact memo, the one Kevin Doyle wrote himself. Your "able to step out of the daily" (Chapter 22) is not only a successful handoff. It is also the pass to the next rung of the F ladder. Sit on the last field as the irreplaceable person and you never free your hands for the next level of complexity.
**The F3 gap.** The weak planks on all five axes point at the same thing. Anchor & Helm's political environment was easy mode. One Grant Whitmore, an ally from day one. Even the heaviest no you said (week 28, when he proposed full automation, Chapter 25) was said to someone who trusted you. You have never handled two decision-makers pulling against each other. Who signs the charter, who hears the bad news first, whose North Star wins when the two conflict. On these questions you have not even given a wrong answer yet.
**Swiftway is not a coincidence.** During the handoff period there were two candidates for the next project, and the one you went to Owen Hartley to ask for was Swiftway Logistics, precisely because it is awkward. Dual Group and subsidiary sponsors, two decision-makers pulling against each other, exactly the complexity you have not carried. Whether its business looks like Anchor & Helm's is Chapter 23's question. Whether its politics look like Anchor & Helm's is this chapter's.
Inside a company, complexity jumps come from switching subsidiaries, switching divisions, and taking on more contested sponsor structures, from building a tool for one department to delivering one system between two competing divisions. And you have one thing people outside do not. Every ask queued in the company sits on the PMO register and the intake list (Chapter 25), so you can see which kinds of sponsor structure next year holds, and all you have to do is ask. Project selection is career design. Projects that come to you keep you moving sideways. Projects you choose are the ones that move you up.
**The annual growth agreement.** Land this self-assessment on one page. What complexity you will carry next year, which piece is missing, which project fills it (template at [Template 3.8](../appendices/template-03-capability.md)). Your worked example is below.
> What you carried this year, one sponsor, one business line, from a vague ask to L4 (graduating F2).
> What you will carry next year, deliver an L3 or above once under the conflicting interests of dual sponsors (F3's first question).
> The gap, the F3 anchors on field judgment and narrative, a charter co-signed by opposing parties, and information discipline under conflict (rules of the "who hears the bad news first" kind).
> Which project fills it, Swiftway Logistics (a dual Group ร subsidiary structure).
> Six-month checkpoint, draw the dual-headed stakeholder map with both sides agreeing to it; the first time instructions from the two sides conflict, handle it by the principle written down in advance rather than picking a side in the moment.
This page is worth something not on the day you fill it in but at the reconciliation six months later. Like a pre-mortem, it turns growth from "something that happens to you" into "something you arranged to happen."
## Failure Modes
**1. Moving sideways.** Ten projects in five years, all at the same difficulty, a resume that keeps getting longer and a level that will not move. Project assignment is usually not yours to control, and the assignment logic is "give it to whoever has done something like it." The organization wants certainty and you are attached to the comfort zone, and the two conspire to pin you at one level. Moving sideways is better hidden inside a company, because one company's projects naturally converge in difficulty. Five years on you may be the person in the company who understands one business line's AI best, and nothing more. And the F ladder is invisible, so repetition sets off no alarm. The fix is this chapter. Make the ladder explicit, negotiate project selection as career design, and go ask for the awkward business lines. The test, look back at your last three projects. If the sponsor structure and the hard problems are all of one kind, you are moving sideways.
**2. The permanent firefighter.** Wherever something breaks, there you are, the business side asks for you by name, and you are quietly proud of it. This is a two-way addiction. The organization needs you to fight fires (you are its cheapest reliability), and you are addicted to being needed (firefighting feedback is a strong hit on an hourly cycle, Chapter 3's firefighter mode trap inside the operator identity, recurring at career scale). Every fire postpones your F3. Firefighting spends nothing but the judgment you already hold, and it grows none. And "being needed" is the exact opposite of "able to step out of the daily," so the pass upward is one you burn with your own hands. The test, think back to the judgment you used in your last firefight. Was any of it grown in that field? If it was all judgment you already had, you are being consumed.
**3. Growing technically without growing in judgment.** Every new model and framework tried first, the left half of the radar (engineering depth, AI engineering) rising year after year, and the right half (business grasp, narrative, field judgment) not moving at all. Technology has courses, documentation, and instant feedback. Judgment has no textbook. It grows only out of real projects with retrospectives, and a retrospective you skip is a saving you keep forever. This leg is the one AI is eating, precisely. You are racing a coding agent's price curve, and it does not sleep. The test, open your last archived radar. Did the behavioral evidence on the three axes in the right half change? Scores up with the evidence unchanged means nothing went up.
**4. Leaving the field too early.** Two years in, one full cycle done, and you move into investing, consulting, or evangelism, more and more stage and less and less field. The title premium on F2 is at a high right now ("has shipped AI in production" sells well in the market), and the temptation to cash out is real. But judgment has not taken root. One cycle is only enough to run the book's frameworks through once, not enough to know what to do when a framework fails. The test, all the stories you tell are other people's, or you have been telling the same story of your own for two years.
!!! note "Vendor View"
On the vendor side, complexity jumps come from changing clients. The next project is set by the sales pipeline, and the difficulty is not yours to pick. The three paths out come in their original form, founder, product leader, field CTO. The capability transfer maps are the same, and only the market where you cash them in sits outside the company.
## Next Monday
1. Run an F-level self-assessment on the project you just closed out (or the one in hand). Find behavioral evidence axis by axis, and rate by the lowest axis. Write down the "papered over with luck and overtime" list. That is the syllabus for your next level.
2. Write a one-page annual growth agreement with [Template 3.8](../appendices/template-03-capability.md). The "six-month checkpoint" has to be filled with verifiable behavior, not an adjective.
3. Look at the next project waiting for you. Does it hold complexity you have not carried? If not, go through the asks queued on the PMO register, pick the one with the most awkward sponsor structure, and go talk to Owen Hartley this week with your growth agreement in hand.
4. Archive today's five-axis radar ([Template 3](../appendices/template-03-capability.md)). What you compare a year from now is not the scores. It is whether the behavioral evidence behind each axis changed.
**Want an agent to get you started?** In the repo you set up following [Start Here](../index.md), paste this to your coding agent:
```text
In the repo/ directory of the the-last-mile repository, help me with the Chapter 26 Next Monday actions. First run python3 templates/f-levels/radar-to-flevel.py
with the built-in sample to show how historical radars map to F levels, then copy questionnaire.md into the working directory I name. The behavioral evidence for the five axes
I will dictate axis by axis, and you only record it. The rating follows the lowest axis, and the level is mine to set, not yours to compute. Build the skeleton of the annual growth
agreement from Template 3.8, and if I write an adjective in the "six-month checkpoint" hand it back and make me replace it with verifiable behavior. Archive today's radar CSV by date,
side by side with my Chapter 3 one. If any command errors, stop and show me the output.
```
---
## Chapter Kit
- **Judgment frameworks.** The F1 to F5 capability rating (five levels ร five axes of behavior anchors; kept strictly apart from the outcome ladder, systems counted in L, people counted in F); the capability transfer maps for the three paths out (internal incubation or a spinout / platform product owner / head of AI for the Group); the annual growth agreement (what you carried this year โ what complexity you will carry next year โ the gap โ which project fills it โ six-month checkpoint)
- **Templates.** [Template 3](../appendices/template-03-capability.md), *Capability Self-Assessment and F1โF5 Rating*, the F-level half (3.7 and 3.8), behavior anchor table plus the annual growth agreement template
- **Key judgments**
- "Project count is not experience. Complexity jumps are."
- "Field judgment is the hardest thing for a model to learn. It is in no dataset."
- "Project selection is career design."
- "A system's outcome is counted in L. A person's capability is counted in F."
---
## From That Monday Morning to This One
The Monday morning of Chapter 0, nine o'clock, fifteen minutes, and all the information in your hands was two sentences, "we want an AI assistant" and "I want to see something in three months."
Thirty weeks later, take stock of what you carry away. Not the code, which stays in Anchor & Helm's repo, and even the lead-writer rights are not yours. The -31% counts for only a small share. What you really carry away is a body of judgment built by walking the whole distance.
- When to build a 2-hour MVP instead of a six-week study.
- Which signal means it is time to switch identity.
- When data can be trusted.
- How to co-build an eval.
- How to deliver bad news.
- When to say no.
- How to step out of the daily.
And a map. You know where you stand between F2 and F3, you know what the next level's questions look like, and you know why the project in hand is this one.
Next Monday morning, in Swiftway's meeting room, someone will say something vague to you again. This time you will not be nervous, and you should not be casual either. Everything Anchor & Helm taught you will be examined again there, plus one new question you have never worked.
The road this book can walk with you ends here. The rest of the road, like judgment, is in no dataset. It is in your next field.
---
# Template 0 ยท Field MVP Pack
> Companion chapter(s): Chapter 0. The five templates (0.1โ0.5) follow the order of the two-hour process and can be copied and modified directly. 0.6, the red line declaration, is an input set before the clock starts and is not counted among the five.
> License: Every template in this book may be modified freely and used in your work, no attribution needed.
---
## 0.1 Workflow Claim
One sentence, two blanks:
> This system will change **[who: role and name]**'s **[which next action]**.
**Self-Check**:
- "Who" is a specific person you can book for 40 minutes, not a department.
- "Action" is something he already had to do today, not something new added to him.
- Cannot write it โ go back and find the "who". Do not go further.
**Wrong โ Right** (the first right example is from this book's Anchor & Helm case):
| Wrong (not a claim) | Right |
|---|---|
| Use AI to improve claims efficiency | Let claims reviewer Linda Marsh, on opening an exception claim, see directly what material is missing and whom to chase first |
| Build a knowledge base assistant | Let a newly hired customer service rep, on taking a refund dispute call, get a citable handling position within 30 seconds |
---
## 0.2 Case Table
10 representative cases. Adjust the fields to the scenario. The skeleton:
| # | Case ID | Current Status | Reason Stuck | Time Waiting | Key Context Fields (add or drop per scenario) | Source (real, de-identified / synthetic) |
|---|---------|----------|----------|----------|------------------------------|--------------------------|
| 1 | | | | | | |
| โฆ | | | | | | |
**Case Selection Rules**:
- 6โ7 typical claims + 2โ3 edge cases + 1 case "even the front line finds hard."
- Mark the source of every case. **All synthetic data = the data feasibility hypothesis is entirely unvalidated**, and that must go into the Readout.
- Red line: no real, identifiable customer information. If de-identified data cannot be had, use AI to generate realistic synthetic cases.
---
## 0.3 Action Queue Prototype
One table is enough (Excel / any spreadsheet tool / a one-screen web page, do not spend more than 25 minutes on the interface):
| Case ID | Reason Stuck | Missing Item | Suggested Priority | Suggested Next Action | Owner | Reason | **Human Call** |
|---------|----------|--------|------------|------------|--------|------|--------------|
**Design Rules**:
- The "Reason" column is mandatory. Every AI suggestion must give a reason that can be challenged. That is the target of the scoring step.
- The "Human Call" column is the soul. It declares the system's position, **advising people, not deciding for them**.
- Every row must have an owner. A suggestion with no owner is a dashboard, not an action queue.
---
## 0.4 Scoring Sheet (Real User Marks Each Line)
| Case ID | Grade | Reason (User's Own Words) | Follow-up |
|---------|------|------------------|------|
| | pass / concern / unsafe / useless | | |
**The four grades defined** (read to the user before scoring):
- **pass.** Follow this suggestion and nothing goes wrong.
- **concern.** Right direction, but a detail makes me hesitate (write down what).
- **unsafe.** Following it causes harm (top-priority signal, ask "what would happen").
- **useless.** Not wrong, but no use to me (value signal, ask "what do you actually need"; this grade is reserved for real users taking a position. When engineers read traces to judge evidence, the fourth grade is unclear. Do not mix them).
**Facilitation Rules**:
- The scorer must be someone "whose daily work changes after the system launches," not that person's boss.
- No 1โ5 scores. Numbers let people politely give a 3. Four grades force a position.
- Record the user's own words. They are the most expensive raw material of the discovery stage.
---
## 0.5 MVP Readout Memo
One page, five sections, answer first:
```
Conclusion: [continue / narrow / redirect / get more data / stop] (one sentence + one numeric piece of evidence)
1. What was validated: whether the workflow claim holds; the score distribution (x pass / y concern / z unsafe / w useless)
2. Key finding: the 1โ2 sentences of judgment worth showing an executive (e.g., "chase priority โ risk priority")
3. Risks exposed: data risk / trust risk / boundary risk, one line each
4. Unvalidated hypotheses: what this MVP did not cover (synthetic data โ data feasibility unvalidated, mandatory)
5. Recommended next step: scope, people needed, data access needed, time box; inside a company add three more, whose schedule gives way, how many people at how many hours a week, and by when it gets reassessed (the Chapter 4 charter locks them in)
```
**Rules**:
- "Stop" is a legitimate conclusion, and a high-return one. You spent two hours saving the company three months. Inside a company, stopping cannot live only in the readout. It goes into the kill register with revival conditions attached (Chapter 25). You will see the person who made the ask tomorrow.
- Every conclusion comes with evidence (score numbers, the user's own words). Never write "overall feedback was positive."
---
## 0.6 Red Line Declaration (Set Before 0:00, Not at 1:59)
Confirm in writing with the business side and the risk owner before the clock starts, and record it in the memo's notes. Silence is not consent:
- [ ] No real production data / no identifiable customer information
- [ ] No messages sent outside automatically
- [ ] Do not touch [this scenario's core decision red line, e.g., the payout amount decision]
- [ ] This MVP promises no launch date. Its deliverables are a basis for judgment, not version 0.1
---
## 0.7 Counterexample: A Tidy-Looking Wrong Answer
Every cell in the Pack excerpt below is filled in, and it would not be sent back if submitted. That is exactly what makes it dangerous.
```
Workflow Claim: Let the claims department use AI to improve exception handling efficiency.
Case Table: 10 cases, all AI-synthesized, all typical "missing material" claims.
Scoring Sheet: scorer Kevin Doyle (Director of Claims Operations); result 9 pass / 1 concern.
Readout: Overall feedback positive, direction validated, recommend continuing.
```
Line by line:
1. **The claim has no "who" and no "action".** "The claims department" cannot be booked for 40 minutes, and "improve efficiency" is nobody's next action today. Rewrite against the 0.1 right example: let Linda, on opening an exception claim, see directly what material is missing and whom to chase first.
2. **All synthetic + all typical.** The data feasibility hypothesis is entirely unvalidated (0.2 requires this to go into the readout), and there are no edge cases. The scoring only sat the ten easiest questions.
3. **The scorer is the boss.** Kevin judges "whether it looks respectable in a report," not "what happens when it is really used tomorrow" (Chapter 0, failure mode 2). Nine pass from a director's hand, and not a single unsafe could be drawn out. The person who should be at the table is Linda.
4. **The readout has no numbers and no quotes.** "Overall feedback positive" is exactly the phrase 0.5 bans by name. Treating "continue" as the default conclusion is Chapter 0's failure mode 4.
The overall test. Take this Pack and ask "what will Linda do differently next Monday because of it," and there is no answer. On the four-grade scale, this deliverable is itself useless.
---
## Code Hooks
The companion repo provides (this repository's `repo/` directory):
- [`templates/field-mvp/`](https://github.com/hallieren/the-last-mile/tree/main/repo/templates/field-mvp/): the table scaffolding and AI prompt scripts for this Pack
---
# Template 2 ยท Deliverer Role Charter
> Companion chapter(s): Chapter 2. Write it inside week 1 and walk your sponsor through it line by line in 15 to 30 minutes. It comes before the Chapter 4 deployment charter. The charter defines the boundary of the **project**, this one defines the boundary of **your role**. Say who you are first, then talk about what the project does.
> License: Every template in this book may be modified freely and used in your work, no attribution needed.
---
## 2.1 What I Am Accountable For (and to Which Level)
```
The outcome I am accountable for on this project:
- Take [the vague ask / project name] to a verifiable production outcome,
accountable through L3 (real adoption) on the outcome ladder; L4 (business side self-sufficient) is reached through handoff
(outcome ladder: the five-level outcome scale from L0 demo to L4 self-sufficient, defined in Chapter 1)
- Full coverage: discovery, design trade-offs, eval, prototype to production, adoption and handoff
What I do not promise:
- Not to settle the outcome on "the demo went well" (a demo is a means, not a milestone)
- Not to promise a schedule that bypasses security or compliance review
- Not to force a launch when the data or the eval does not support it
- Not to promise [this project's red line, e.g. touching automatic payout decisions]
```
**Self-Check**: If you cannot write "to which level" in the first section, you have not talked to your sponsor about the definition of the outcome. Talk about that first, then fill in the table.
## 2.2 Whom I Do Not Replace
Read it out loud to the sponsor, row by row. This table does not guard against other people taking work. It guards against **you being pressed into an old slot** (the four molds in Chapter 2).
| I Am Not | Where the Difference Is | What You Still Need Them For |
|--------|----------|------------------------|
| The POC demo squad | My settlement point is a production outcome, not a demo; a demo's target audience is real users, not visiting executives | Product showcases for executive visits, still arranged by the business side itself |
| Outsourced development (a team) | I have an obligation to question every ask (it must answer a workflow claim); I refuse "shut up and code" | High-volume feature work with clear boundaries can still be outsourced or moved to the platform team |
| Internal consulting (an advisor) | My advice is delivered and validated as a running system, not wrapped up in a report | Consulting at the organizational and strategic level still needs dedicated advisors |
| Ops ticket support | I am accountable for the outcome through L3 on the outcome ladder, not for answering and closing tickets; launch is not my finish line | Day-to-day ops tickets and on-duty response still belong to the ops team |
| The platform team | I am accountable for this one business side's field outcome; field findings flow back, but I do not set the platform roadmap | The platform roadmap and trade-offs on cross-department shared features still belong to the platform team |
## 2.3 The Two-Sided Expectations Table
| What You Can Expect From Me | What I Need From You |
|----------------|----------------|
| A one-page sync every week, answer first, bad news the moment it appears | One sponsor who can make the call, 15 minutes a week |
| Every AI suggestion in the system carries a reason, can be challenged, and can be overridden by a person | Real users' time (scoring, shadowing, co-building the eval) |
| I am present for production incidents: response, rollback, retrospective | Access to de-identified data and a clear data boundary |
| I teach your people to run and improve the system rather than leaving a black box | Security and compliance reviewers on the project team from week 2 |
| I say stop when it is time. When data or value does not support it, I recommend narrowing or terminating | One consistent way of introducing me (see 2.4), and do not call me "the AI expert" |
| Committed input written into the charter, with the fulfillment rate reported monthly | Named business-side commitments and schedule protection (people, hours per week, written into their calendars) |
| The handoff arrangement stated well before launch, with no vacuum period | Ops ownership and on-duty arrangements after launch, by name |
## 2.4 Alignment Process and Rules
- **Week 1**: book 15 minutes with the sponsor and go through it line by line. Record any item you disagree on right there. Leave nothing to "assumed agreement."
- **Introduction line**: give the sponsor one repeatable sentence, for example, "They are accountable for taking this project from an idea to launched and used, working alongside our people the whole way." Correct the name on the spot whenever it is wrong. The name decides how you get used.
- **Position**: this charter is attached ahead of the deployment charter (Chapter 4), and the charter negotiation takes it as a premise.
- **Citation**: each time you are used according to an old mold ("come give us a demo," "just build this ask"), cite the matching item and correct it in one sentence. Do not argue, do not escalate it into a principle.
- **Review it three times**: when the charter is signed, before launch, and at the ops handoff after launch. The role boundary moves with the project's stage (the builder's weight gives way to the teacher's, see Chapter 3), and the move has to be confirmed explicitly, not left to unspoken understanding.
---
## Code Hooks
The companion repo provides (this repository's `repo/` directory):
- [`templates/role-charter/`](https://github.com/hallieren/the-last-mile/tree/main/repo/templates/role-charter/): the fill-in guidance prompt and the one-page layout template for this charter
---
# Template 3 ยท Capability Self-Assessment and F1โF5 Rating
> Companion chapter(s): Chapter 3 (capability self-assessment), Chapter 26 (F1โF5 rating and the growth agreement). The first half (3.0โ3.6) is the five-axis radar self-assessment: one paragraph of definition per axis, 1โ5 behavior anchors, three self-check questions, and shoring-up advice by score. The second half (3.7โ3.8) lays out two tools in the order you use them: rate yourself with the anchor table first (3.7), then turn the level gap into a plan with the growth agreement (3.8). Chapter 26's F1โF5 capability rating reuses the same coordinate system, so archive your scores. The last section (3.9) is for engineers who want to enter this kind of work. It translates the experience you already have into evidence on this coordinate system.
> License: Every template in this book may be modified freely and used in your work, no attribution needed.
---
## 3.0 How to Use This
- **Score on behavioral evidence, not on "I should be able to."** Read the 1โ5 anchors for each axis first, and take the row that most resembles you on your most recent real project.
- **The self-check questions are for calibration.** Three per axis. Fewer than two yes answers, and you drop that axis by one point.
- **Get a coworker to score you blind and compare.** On any axis where the two of you differ by โฅ2 points, take theirs. The systematic bias of self-assessment is overrating.
- **You may let AI prefill the evidence from your project records, but the score has to be yours.** The value of a self-assessment is honesty, and AI cannot be honest for you.
- **Re-score and archive every quarter (or after every project milestone)**. The F-level rating in 3.7 (Chapter 26) draws on your historical radars.
---
## 3.1 Engineering Depth
Without AI, can you still stand a system up. Data, APIs, debugging, deployment. This axis guards the floor under "AI does it, you review it." For someone who cannot review it, moving up the leverage is not a conversation.
**Anchors**:
| Score | Behavior |
|----|----------|
| 1 | Can modify code someone else built; building from scratch alone is hard |
| 2 | Completes modules alone inside a familiar stack; needs support to cross stacks or go to production |
| 3 | Can take a prototype to a working pilot alone, data integration, deployment, logging, basic monitoring |
| 4 | Can enter a business-side environment on an unfamiliar stack and locate and fix integration problems within two weeks |
| 5 | Can design and land a maintainable integration in a constrained environment (legacy systems, no documentation, strict security requirements) |
**Self-check questions**:
1. Handed an error log from an unfamiliar system, can you locate which layer the problem is in without calling anyone?
2. Was the last time you deployed a service from scratch to reachable, with your own hands, within the past six months?
3. Was the last time a review of yours caught a subtle error in AI-generated code (boundary conditions, concurrency, silent failure) within the past month?
**Shoring-up advice**:
- **1โ2**. Do not fly solo on this kind of project yet. Do two complete small projects from scratch to deployment. AI assistance throughout is fine, but force yourself to understand every line and to handle every deployment failure by hand.
- **3**. Fill in the "unfamiliar environment" experience. Volunteer for one job integrating with a legacy system or working under constraints. Depth grows fastest where you are uncomfortable.
- **4โ5**. Your risk is not inability, it is unwillingness to let go (Chapter 3's "coder mode"). Explicitly cede the builder time you save to the axes on the right.
---
## 3.2 AI Engineering
The engineering capability to manage uncertainty. Eval, error taxonomy, human oversight, drift. Traditional engineering delivers certainty. AI engineering delivers **managed uncertainty**. This axis is the addition that separates this role from its five predecessors.
**Anchors**:
| Score | Behavior |
|----|----------|
| 1 | Calls model APIs and writes prompts; judges quality by feel |
| 2 | Can build basic retrieval augmentation and tool calling; can read failure cases but has no method for classifying them |
| 3 | Can write an eval for one feature, golden cases, acceptance thresholds, error categories (the method in Chapter 11) |
| 4 | Can design the human-machine division of labor, what goes to the model, what stays with people, how human overrides are recorded and flow back |
| 5 | Can manage uncertainty across a system's whole lifecycle, drift monitoring, rollback strategy, an eval that evolves with the business |
**Self-check questions**:
1. For your most recent LLM feature, can you write "an acceptance threshold plus three error categories," or only "it works pretty well"?
2. In your system, at which step do people override the AI's output? Did you design that on purpose, or did it just grow that way?
3. After launch, when the model's behavior changes (an upgrade, data drift), within how many days will you know? By what means?
**Shoring-up advice**:
- **1โ2**. Start from the eval, the fastest-returning move on this axis. Write 20 golden cases and one error taxonomy table for any AI feature on your desk (Chapter 11 plus Template 11).
- **3**. Move toward the operating period. Design one complete loop for an existing system, human override into the decision trail โ weekly review โ rules flowing back.
- **4โ5**. Settle the method into a teachable asset (Chapter 23's playbook). The scarcest output from an expert on this axis is getting other people to 3.
---
## 3.3 Business Grasp
Read the business side's process, metrics, and money. What the real workflow looks like, which business number your system moved, who watches that number.
**Anchors**:
| Score | Behavior |
|----|----------|
| 1 | Can follow a description of the request; does not know how the business side makes money or where it bleeds |
| 2 | Knows the official version of the target workflow (the SOP level) |
| 3 | Knows the real workflow, workarounds, exception handling, how the front line gets around the system (the output of Chapter 6's archaeology) |
| 4 | Can convert a system change into business metrics, handling time, leakage rate, cost |
| 5 | Can anticipate the organizational reaction, which department will resist, which metric will trigger gaming, who ends up uncomfortable |
**Self-check questions**:
1. Can you draw the real version of the target workflow (workarounds included) and have front-line users agree "that is exactly it"?
2. Can you name which of the business side's numbers your system moved, and who looks at that number in which monthly meeting?
3. On the business side, who is made uncomfortable by your system doing well? Do you have a name?
**Shoring-up advice**:
- **1โ2**. Go to the field. Use Chapter 6's method to shadow one real user for half a day. That half day moves this axis further than ten industry reports.
- **3**. Learn the money math. Convert your current project's North Star metric into an annualized dollar figure, and take it to the sponsor to be corrected in person. The correcting is the classroom for this axis.
- **4โ5**. Business intuition is your scarcest asset. Spend it where the leverage is highest, opportunity elimination and intake judgment (Chapters 7 and 25).
---
## 3.4 Narrative
Get every level of the organization the judgment it needs, from the front line's own words to an executive memo, from good news to bad.
**Anchors**:
| Score | Behavior |
|----|----------|
| 1 | Can write clearly what you did (an activity-list report) |
| 2 | Can report by layer, detail for engineers, progress for managers |
| 3 | Answer first. The memo leads with the conclusion and the decision you want, then the evidence (Chapter 13's pyramid structure) |
| 4 | Can deliver bad news, risks, delays, "I recommend stopping." Trust goes up afterward, not down |
| 5 | Can lead the narrative, so the project is told correctly inside the business side (short-term wins, retrospectives, public postings) |
**Self-check questions**:
1. Could the reader of your latest weekly report make a decision without asking a single follow-up question?
2. When did you last volunteer bad news within 24 hours? Did the relationship get better or worse afterward?
3. For an executive, a director, and the front line, do you tell the same project in three versions with different content?
**Shoring-up advice**:
- **1โ2**. Force every report into one page, "conclusion plus the decision you want from them plus three pieces of evidence." Do ten in a row and the habit grows.
- **3**. Practice bad news. The next time a risk appears, volunteer it within 24 hours and attach your defense. Bad news plus a defense is a deposit. Bad news discovered by someone else is a withdrawal.
- **4โ5**. Lend the narrative capability to the business side. Help the champion and the owner tell this project well inside their own organization (Chapter 21).
---
## 3.5 Field Judgment
Make the right trade-off on incomplete information. Red lines, priorities, the timing of an identity switch, when to say no. This axis is the hardest to build in a hurry, and it is where a deliverer at F4 and above separates from a senior engineer.
**Anchors**:
| Score | Behavior |
|----|----------|
| 1 | Executes to plan; on an exception, goes and asks |
| 2 | Can spot an anomaly and escalate, knowing what has to be said today |
| 3 | Can make trade-offs under time pressure, ranking by irreversibility (Chapter 3), with a steady sense of the red lines |
| 4 | Can anticipate risk and move it forward, writing real causes of death in the pre-mortem and turning defenses into work items at kickoff |
| 5 | Can judge whether it should be done at all, deciding before taking it on whether it is worth taking (Chapter 25's intake), and daring to say no with evidence |
**Self-check questions**:
1. Looking back at your last project, can you point to three specific moments where you should have switched identity or said no?
2. In a pre-mortem you wrote, was there an entry that later actually happened, with the defense doing its job?
3. In the past year, did you kill a direction you were part of with your own hands, and argue the reason until the other side accepted it?
**Shoring-up advice**:
- **1โ2**. Judgment grows out of retrospectives. Spend 15 minutes every Friday writing "the three trade-offs I made this week and what they rested on," and after a quarter look back at which were wrong and why.
- **3**. Write a pre-mortem at the start of every new project (Template 4) and check the answers when it ends. The pattern in your missed predictions is your judgment blind spot.
- **4โ5**. Teach the judgment out. Judgment is the hardest asset to hand off, and it is the core proposition of F5 in Chapter 26.
---
## 3.6 Reading the Chart and Acting
**Draw the radar**. Connect the five scores into a five-point radar chart (by hand is fine; the code hook is at the end of this page).
**Three rules for reading the chart** (the same as Chapter 3):
1. **The lowest axis decides how large a project you can own alone**, and the highest axis decides nothing. The five axes multiply, they do not add.
2. **The shape predicts the failure mode**:
| Radar Shape | High-Risk Failure Mode (Chapter 3) |
|----------|--------------------------|
| Engineering depth + AI engineering high, narrative / judgment low | Coder mode |
| Narrative + business grasp high, engineering depth thin | Consultant mode (advice you cannot verify with your own hands) |
| Field judgment low, the rest passable | Firefighter mode (cannot tell "mine to fix" from "mine to teach") |
| All five around 3, no strong suit | Busyness in place of judgment (nothing only you can do) |
3. **The 30-day shoring-up rule**. Work one axis at a time, and pick the lowest. The target has to be verifiable behavior ("the next three weekly reports draw zero follow-up questions"), not an adjective ("improve communication").
**Archive**. Record the date, the scores, and one sentence of evidence for every self-assessment. The behavior anchor table in 3.7 (Chapter 26) uses the same coordinate system to define the F1โF5 capability levels. Your historical radars are your growth curve.
---
The next two sections and the self-assessment above are three instruments on the same five-axis coordinate system. The self-assessment radar fixes position (how strong each axis is at a given moment). The behavior anchors fix level (what the next level requires on each axis). The growth agreement fixes the path (turning the level gap into a year's plan). Chapter 3 makes the first assessment with the self-assessment sheet, Chapter 26 re-assesses with the anchor table and the growth agreement, and your quarterly archive is the evidence base for the rating. That is how the two ends close the loop.
## 3.7 The F1โF5 Behavior Anchor Table
**Rules**:
1. The cells hold behavior anchors ("have you done it or not"), not adjectives. An axis sits at a level only when every anchor at that level has behavioral evidence with **you as the subject** (having taken part does not count, having carried it does).
2. **Overall level = the lowest of the five axes.** The five axes multiply, they do not add (Chapter 3's rule for reading the chart).
3. Each level absorbs the one before it. Rating at F3 assumes you can still produce the F2 anchors.
### Engineering Depth
| Level | Behavior Anchor |
|----|----------|
| F1 | Completes modules inside an existing codebase and a familiar stack; needs support for deployment and integration; catches the obvious errors in AI output on review |
| F2 | Takes a prototype to production alone, data integration, deployment, monitoring, rollback, the whole loop by one person |
| F3 | Still lands a maintainable integration on an unfamiliar stack and in constrained environments (legacy systems, no documentation, strict security requirements) |
| F4 | Sets the technical baseline for several parallel projects; locates the layer of an architecture problem in someone else's project quickly |
| F5 | The organization's principles for technology choice and its standard for reusable components came from you, and were proven on several real projects |
### AI Engineering
| Level | Behavior Anchor |
|----|----------|
| F1 | Calls models, writes prompts, runs an existing eval; can describe failure cases but has no method for classifying them |
| F2 | Writes the five-part eval spec alone and uses it to rule on a plan; has designed a human oversight and decision trail loop |
| F3 | Maintains evals for several business lines at once; has ruled between conflicting error tolerances with all parties accepting it |
| F4 | The eval method settled into a pattern library asset that other projects reuse; decision trail data aggregated across projects into product-level insight (Chapter 24) |
| F5 | The organization's AI quality standard (error taxonomy, error severity vocabulary, acceptance discipline) was defined by you and stayed in use |
### Business Grasp
| Level | Behavior Anchor |
|----|----------|
| F1 | Follows the request; can repeat back the official process |
| F2 | Dug out the real workflow and the source of truth by hand; converted the North Star metric into money and had the sponsor accept it |
| F3 | Reads the interest structure of several departments at once; has anticipated who would resist and which metric would trigger gaming, and changed the design in advance |
| F4 | Recognizes isomorphic problems across industries (reusing the judgment structure, Chapter 23); can write the evidence for an nโฅ2 generalization argument |
| F5 | Has judged whether a class of business scenario is worth the organization's investment; the strategic value axis of the intake rubric was calibrated by you |
### Narrative
| Level | Behavior Anchor |
|----|----------|
| F1 | Reports what was done, accurately, leaving out no bad news |
| F2 | Runs the three-memo system alone; volunteers bad news within 24 hours; has traded a one-page memo for a decision |
| F3 | Holds trust on both sides between opposed sponsors; tells the same fact to conflicting parties in versions each can accept, without distorting it |
| F4 | Has designed the narrative for team members and for the business side's champion; the format that translates between the field and the product came from you |
| F5 | Speaks publicly for this method (industry exchanges, open talks); the way the organization tells its own methodology carries your hand |
### Field Judgment
| Level | Behavior Anchor |
|----|----------|
| F1 | Spots anomalies and escalates the same day; can recite the red line list and holds it |
| F2 | Ranks by irreversibility under time pressure; writes real causes of death in a pre-mortem; dares to write a readout that says "stop" |
| F3 | Judgment still holds under conflicting interests. When two decision-makers pull against each other, handles it by written principle instead of picking a side in the moment |
| F4 | Trade-offs at the portfolio level. Has judged, across several projects, which to save, which to kill, whom to send |
| F5 | Has said no to an organization-level opportunity and offered a way out; the kill register and the retrospective discipline were built by you (Chapter 25) |
**A worked rating** (mapped from Chapter 26's Anchor & Helm self-assessment). The evidence on the five axes reads F2 / F2 (close to F3) / F2 / F2 (close to F3) / F2 (barely) โ overall level F2. "Close to F3" does not change the rating. It only tells you what the growth agreement should say.
---
## 3.8 The Annual Growth Agreement Template
**How to use this**. Write it at the end of a closeout season or at year end, one page or less. Once written, align it with one senior person you trust (a mentor, your manager, or a long-standing business-side lead). It must be reconciled six months later. The agreement is how Chapter 26's sentence lands. Project count is not experience. Complexity jumps are.
| Field | What to Fill In |
|----|----------|
| **What you carried this year** | A description of complexity plus behavioral evidence, not a project list ("did three projects" fails) |
| **What complexity you will carry next year** | One sentence, pointing at a specific F-level gap |
| **The gap** | Which axis, and which anchors still have no evidence with you as the subject |
| **Which project fills it** | Whether one is in the existing pipeline; if not, whom you ask for what kind of project |
| **Six-month checkpoint** | Verifiable behavior, not an adjective ("increase influence" fails) |
| **Reconciliation record** (fill in after six months) | What happened; where the agreement was wrong, and why |
**Three rules**:
1. **The checkpoint has to be a behavior**. Did it happen or not, judged at a glance.
2. **If the projects do not match, go negotiate**. Take this page to whoever assigns the work. It turns "I want good projects" into "I need evidence at this level of complexity."
3. **The reconciliation is worth more than the filling in**. The fields you got wrong are your judgment blind spots, the same way checking a pre-mortem's answers works.
**A worked example (Chapter 26, "your" Swiftway version)**:
> **What you carried this year**: one sponsor, one business line, from a vague ask all the way to L4 on your own (the top of the outcome ladder, the business side self-sufficient, L0 demo โ L4, defined in Chapter 1; Anchor & Helm, graduating F2; evidence, primary responsibility from charter through handoff, with the last impact memo written by the business-side owner himself).
> **What you will carry next year**: deliver an L3 (real adoption) or above once under the conflicting interests of dual sponsors (F3's first question).
> **The gap**: the F3 anchors on field judgment and narrative. A charter co-signed by opposing parties, and information discipline under conflict, neither has evidence.
> **Which project fills it**: Swiftway Logistics (dual Group and subsidiary sponsors, two decision-makers pulling against each other).
> **Six-month checkpoint**: draw the dual-headed stakeholder map with both sides agreeing to it; the first time instructions from the two sides conflict, handle it by the principle written down in advance rather than picking a side in the moment.
> **Reconciliation record**: (fill in after six months.)
---
## 3.9 The Bridge for Newcomers: The Last Mile from Engineer to This Role
This section is for engineers who have not done this job and want to enter this kind of work (it applies to solutions engineer and similar titles). Score a radar with 3.1โ3.6 first. The typical newcomer shape is high on the left and low on the right. Engineering depth and AI engineering have evidence, the three axes on the right do not. This section does three things: it translates the experience you already have into this language (3.9.1), turns the book's frameworks into interview ammunition (3.9.2), and turns internal cross-department delivery into F1 evidence (3.9.3).
### 3.9.1 How to Word Your Resume
There is one rule, taken from Chapter 0's workflow claim structure. Rewrite each item as "whose action it changed, and which business number it moved," with you as the subject of the verb (the same source as 3.7's rating rule).
| Background | Radar Strength | Weakest Axes | Rewrite Example (Original โ the Deliverer's Language) |
|----------|----------|----------|------|
| Recommendation / search ML engineer | AI engineering | Business grasp, narrative | "Ranking model AUC +2%" โ "Co-built the evaluation definition with the business side, converted the model improvement into $X per order of conversion revenue, and drove operations to adjust the strategy on that basis and ship it" |
| Backend engineer | Engineering depth | AI engineering, business grasp | "Refactored the order service, tripled QPS" โ "Rebuilt the order path under a no-downtime constraint on a legacy system, aligned the cutover plan with the ops and risk teams, and ran two weeks of gradual rollout with zero incidents" |
| Data engineer | Engineering depth (data side) | Narrative, field judgment | "Built the ETL, X hundred million rows a day" โ "Aligned three data sources with conflicting definitions into one source of truth the business side accepted, surfaced N status fields that did not match actual business, and drove the fixes" |
| Pre-sales SA | Business grasp, narrative | Engineering depth, AI engineering | "Led X POCs, win rate Y%" โ "Did the data integration and deployment in the POCs with my own hands, and N of them reached production and daily use on the business side (L3 and above)" |
### 3.9.2 Common Interview Questions โ Chapter Ammunition Index
Six common questions, each with the chapter numbers and a skeleton answer. The skeleton is the order of trade-offs. Go back to the chapter for the detail.
**1. "A business-side executive does not trust the model's output. What do you do?"** (Chapters 5, 11, 12) Trust is the first deliverable, accumulated through small verifiable promises, not by explaining how the model works. Bring the business experts the executive trusts into co-building the eval. They set the golden cases and the acceptance thresholds. Then show him the human oversight design explicitly, which outputs must be confirmed by a person, and how errors get found and rolled back.
**2. "The demo went well and the project will not move. How do you diagnose it?"** (Chapters 1, 17, 20) Use the outcome ladder first to locate which rung it is stuck on. A successful demo is only L0, and being stuck is usually an adoption problem, not a technical one. Then read resistance as a diagnostic signal. Whose workflow changed, who is made uncomfortable, get the names before discussing a plan. Last, check the delivery form. A dashboard changes nobody's next action, and it usually has to be rebuilt into an action queue embedded in the workflow.
**3. "The business side demands 99% accuracy and there is no ground truth. How do you take it?"** (Chapters 4, 11) Do not take the number, take the definition first. Eval as spec. Co-build the golden cases and the error categories with the business side's own experts, and "99%" decomposes, in front of concrete cases, into different tolerances for different errors. Negotiate thresholds by severity. Which class of error gets zero tolerance, which can be backstopped by a person. The thresholds you settle go into the charter's success metrics and exit conditions, never into a verbal promise.
**4. "What would you do in the first week?"** (Chapters 0, 4, 5) Run a two-hour Field MVP within 48 hours. Write the workflow claim, get 10 cases (if the data cannot leave the database, copy the structure by hand at the data owner's screen), and have a real user score them on four grades (pass / concern / unsafe / useless). Take the scoring results into the deployment charter negotiation, goal, users, data boundary, success metrics, exit conditions. Draw the stakeholder map at the same time, and confirm the actual user, the owner, the data side, the risk side, and the maintainer are all identified.
**5. "The data is bad. Do you still do the project?"** (Chapters 9, 25) Use the data fitness ladder first to locate which rung it is bad at (exists, accessible, interpretable, timely, traceable, validated, actionable), because the fix is completely different at each. Then narrow. Use a thin slice to find the one narrow business line with the best data and get that running, and write the data problems into the readout as evidence. Hit a red line (no source of truth can be built) and give the conclusion "stop or redirect" with the evidence attached.
**6. "How do you prove your project produced business value?"** (Chapters 18, 19) The metric is set in the charter before work starts, not looked for before the report. The North Star metric is converted into money and accepted by the sponsor in advance. After launch, prove the improvement instead of only showing a trend, with a baseline comparison, the human override rate, and decision trail data. Report with a pyramid memo, answer first. The hardest evidence is an impact memo the business-side owner wrote himself.
### 3.9.3 Internal Cross-Department Delivery Is the Main Stage
Internal cross-department delivery is not a second-best substitute. It is another room for the same exam. Check it against 3.7's F1 anchors. The ticket is not an external site, it is "complete a delivery inside a given boundary, review AI output competently, escalate exceptions." You can sit that exam inside your company right now, and it is easier to sit than an external project, because the requester and the data are in the same building and you are not waiting on someone else to approve access. Three steps.
**Step one, pick a cross-department delivery project.** The users are not on your team (operations, sales, finance, another department), the ask is vague (at the "help us be more efficient" level), and there is real data and there are real users. Internal tools, process automation, and report rebuilds all count. The test. The requester and the user are not the same person, and if nobody uses it when it is done, you can observe that. Meet both and it is a mock exam room for F1.
**Step two, produce the artifacts the way the book does.** A one-page charter before work starts (Template 4). A two-hour Field MVP with real users scoring on four grades (Template 0). Twenty golden cases plus acceptance thresholds before launch (Template 11). A one-page impact memo at closeout, with the business numbers confirmed by the using side. The four artifacts make up your F1 evidence base, and you are the subject of every one.
**Step three, write it into your resume by 3.9.1's rule.** Not "developed an internal tool," but behavioral evidence. "Converged department X's vague ask into a workflow claim, and narrowed the scope after real users scored it line by line. After launch, metric Y improved Z%, confirmed by the using side. The system stayed in daily use after I stepped out." That last half sentence (the system stayed in daily use after you left) is the scarcest evidence on a newcomer's resume.
When an interviewer says "you have never worked an external site," answer with the artifacts from these three steps. The requester, the user, and the data owner in a cross-department project are a scaled-down interest structure, and the four artifacts prove you carried it (3.7's rating rule counts only having carried it).
---
## Code Hooks
The companion repo provides (this repository's `repo/` directory):
- [`templates/self-assessment/`](https://github.com/hallieren/the-last-mile/tree/main/repo/templates/self-assessment/): the radar chart generation script and team roll-up template for the capability self-assessment sheet (3.0โ3.6)
- [`templates/f-levels/`](https://github.com/hallieren/the-last-mile/tree/main/repo/templates/f-levels/): the F-level rating questionnaire and the "historical radar โ F level" comparison script, sharing one data format with the radar chart script in [`templates/self-assessment/`](https://github.com/hallieren/the-last-mile/tree/main/repo/templates/self-assessment/), so quarterly self-assessment archives feed straight into the annual growth agreement
---
# Template 4 ยท Deployment Charter and Pre-mortem Memo
> Companion chapter(s): Chapter 4 (the pre-mortem sections, Chapter 1). The fillable one-page charter with its seven elements, the 60-minute contracting meeting agenda, and the pre-mortem memo, ready to copy and modify.
> License: Every template in this book may be modified freely and used in your work, no attribution needed.
---
## 4.1 The Deployment Charter One-Pager (Seven Elements)
> Rules: **one page maximum; you draft it, it gets changed together in the meeting, and both sides sign**. Every cell must be changed by the business side at least once during the contracting meeting. A charter with no business-side fingerprints on it is not an agreement. The charter and the project approval form are two documents. The approval form governs budget and process compliance, the charter governs the work and the expectations.
```
Deployment Charter ยท [project name]
Version: v[ ] Date: [ ] Drafted by: [ ] Signed by: [business-side decision-maker] [business-side owner] [you] [your manager]
```
### Element 1, North Star Outcome Metric (One, Measurable)
| Item | Content |
|---|---|
| Metric definition | [the business outcome metric, precise down to how it is measured, e.g., first-touch handling time for auto exceptions (defined as calendar days from claim report to first handling completed)] |
| Baseline | [current value plus data source; if the baseline data is still unverified, write "as set by discovery's reconciliation"] |
| Target change | [e.g., -30%; state how this number was derived (cite the Field MVP / the data breakdown analysis)] |
| Settlement point | [when it gets measured, over what window, and who produces the number] |
**Self-Check**:
- Only one North Star. A second "important metric" is demoted to an observation metric and listed separately.
- An outcome (handling time, leakage rate), not an activity ("development complete," "launch training done"). The test, could this sentence fail to come true? A goal that cannot fail is not a goal.
- When the metric gets worse, does it hurt anyone? A metric nobody hurts over is not the North Star (Chapter 1).
### Element 2, actual user and owner (by Name)
| Role | Name / Title | Notes |
|---|---|---|
| actual user | [e.g., Linda Marsh and her review team (eight people)] | The people whose daily work the system changes after launch |
| owner | [e.g., Kevin Doyle, Director of Claims Operations] | The person who owns and operates this system after launch (not the deliverer) |
| executive sponsor | [e.g., Grant Whitmore, COO] | The person who settles resources and metrics |
**Self-Check**: Write names, not departments. If the owner cell holds you or your own team, the handoff (Chapter 22) is already destined to fail.
### Element 3, Scope and Red Lines
| Category | Content |
|---|---|
| This phase does | [e.g., the auto exceptions action queue, covering missing documents / abnormal amount / disputed liability] |
| This phase does not | [write out explicitly what was discussed and excluded, e.g., the management dashboard, non-auto lines of business, routine claims] |
| Red lines (not touched at all) | [e.g., no automated payout decisions; no automated outbound messages; no handling of identifiable customer information] |
**Self-Check**: The "not this phase" list is the main defense against scope creep. A request that was turned down has to leave its name on paper, or it revives at every meeting.
### Element 4, Data Boundary and Access Commitments
| Item | Content |
|---|---|
| List of data needed | [dataset ร form (read-only / export / de-identified sample) ร purpose] |
| De-identification and data-exit rules | [what data may leave the business side's environment (usually none); who sets the de-identification standard] |
| Approvers and deadlines | [the named approver for each dataset plus the committed delivery date] |
| Security review timing | [review start and finish written into the schedule, with a named owner] |
**Self-Check**: Every row of data needs "who approves it, when it arrives." An access commitment with no date is not a commitment, it is that week 9 email nobody answers.
### Element 5, Business-Side Commitment (People ร Time, by Name)
| Person / Team | Commitment | Purpose | Term | Constraints |
|---|---|---|---|---|
| [e.g., Linda's team] | [2 hours a week] | [case annotation, rule confirmation] | [from discovery through one month after launch] | [booked in advance, not during peak hours] |
| [e.g., business-line engineers ร2] | [x days a week] | [co-build, take over maintenance] | [ ] | [ ] |
| [the owner in person] | [the weekly meeting plus a decision response time] | [ ] | [ ] | [ ] |
| **Your team's commitment and protection conditions** | [you and your team members, x days a week, by name] | [discovery, co-build and handoff] | [from discovery through the completion of handoff] | [countersigned by your manager, stating whose weekly must clear it before anyone is pulled] |
**Self-Check**: This cell left blank = a one-sided love letter. "What you want" (element 1) and "what you are willing to put in" (this element) must be settled in the same room, and if you cannot agree you do not sign. The hours written down have to enter the other side's calendar system, not stay in their wishes. Your team's commitment row cannot be blank either. Your schedule is a resource too, and it only counts once your manager countersigns.
### Element 6, Launch Release Conditions (Tied to the Eval)
| Item | Content |
|---|---|
| Release mechanism | The eval spec both sides co-build is the authority, [n] golden cases plus agreed thresholds, and a score over the line releases the launch |
| Releasers | [the business-side owner plus the risk owner, not you] |
| Eval spec status | [usually not built yet at signing; state who co-builds it, by when, and that it becomes an attachment to this charter once written] |
| Release methods explicitly excluded | Subjective methods, the demo passed, the boss is happy, trial feedback was good, are not grounds for release |
**Self-Check**: Sign the mechanism now, fill in the numbers later. "Release criteria to be determined" is unacceptable. "Release mechanism settled, thresholds per the attachment" is acceptable. How to build the eval spec is in Chapter 11.
### Element 7, Exit and Resource Reassessment Conditions (Three Parties)
| Direction | Trigger | Action |
|---|---|---|
| The delivery side may call a resource reassessment | [e.g., the business side's promised data access / staffing goes unmet two weeks running] | [three tiers: escalate to the sponsor weekly / the project turns to "awaiting inputs" on the PMO register / your team's people are released and it is recorded openly, see the three tiers of resource reassessment (Chapter 4)] |
| The business side may call a resource reassessment | [e.g., a key risk event; a change in compliance requirements] | [escalate to the sponsor weekly, where the sponsor decides between cutting scope, going manual, or adding input] |
| The sponsor may call a re-prioritization | [e.g., a higher-priority project appears; the budget cycle shifts] | [re-place this project in the schedule, and both sides reconfirm the delivery cadence against the new schedule] |
**Self-Check**: All three parties can raise it. A charter written for the business side only is a disclaimer. One written for the delivery side only is a one-sided contract. Inside a company there is no pause card (Chapter 4), salaries get paid either way, so all three rows are about reassessment and re-prioritization, and none of them says who may pause. The function of exit and resource reassessment conditions is to make sure the person who calls a stop does not have to be the villain. The full judgment on kill criteria is in Chapter 14.
---
## 4.2 The 60-Minute Contracting Meeting Agenda
**Before the meeting** (missing any one item, reschedule):
- [ ] The decision-maker has confirmed attendance (someone who can settle metrics and commitments on the spot)
- [ ] The role charter (Template 2) is aligned with the sponsor (say who you are first, then talk about what the project does)
- [ ] The Field MVP evidence pack, the original scoring results, the readout memo, the key data breakdown analysis
- [ ] The charter draft (all seven elements pre-filled, with the cells you expect to be contested marked)
- [ ] Extracts of the acceptance and scope wording from any existing project approval form or ticket (if it conflicts with the charter, say so in the meeting and agree how to handle it)
**Agenda**:
| Time | Segment | Key Points | Facilitation Rule |
|------|------|------|----------|
| 0:00โ0:10 | Open with evidence | Retell the MVP scoring results and the readout conclusion, no vision deck | Anchor every negotiation to evidence; when someone returns to the vision story, pull them back to the scoring sheet |
| 0:10โ0:25 | North Star negotiation | Metric definition โ baseline (with the data risk stated) โ how the target was derived | Answer a disagreement over the target with a breakdown of the data, never with "I think"; if the baseline is not clean, write "as set by the reconciliation" |
| 0:25โ0:35 | Names and commitments | actual user and owner by name; the business-side commitment settled row by row | Names, not departments; get the person or their manager to confirm the hours in the room |
| 0:35โ0:45 | Scope, red lines and the data boundary | The three lists of "do / do not / do not touch"; each dataset gets an approver and a date | Excluded requests go on the "not this phase" list by name; the security review date goes into the schedule |
| 0:45โ0:55 | Acceptance and exit | Sign the acceptance mechanism first (eval plus thresholds to follow), then ride it into the exit conditions | Warn them the silence is coming; explain that exit conditions mean the person who calls a stop does not have to be the villain, two-way and equal |
| 0:55โ1:00 | Read back and commit | Read the seven elements back out loud; agree on writing it up within 48 hours and signing within a week | The facilitator does the read-back, and the business side's corrections are the last round of alignment |
**Within 48 hours of the meeting**:
- [ ] The charter is written up and sent, with every "changed in the meeting" spot marked (so the business side sees its own fingerprints)
- [ ] Collect the edits (every edit is a cheap expectation gap blowing up, so welcome it)
- [ ] Complete the signing within a week; where it conflicts with existing project approval wording (the old ticket's acceptance clause, say), file it through a project approval change review
**Change rule**: the charter is a living document, but a change = a renegotiation. Any side changing any element must notify all signatories and get their confirmation, and the version number goes up by one. A charter changed quietly is worse than no charter.
---
The pre-mortem memo below sits in the same template unit as the charter, and not by accident. The pre-mortem is written around the time the charter is signed. Writing down the causes of death the night before signing is a stress test of the seven elements. The causes of death and defenses it lists are the best raw material for element 7's exit conditions (and for the self-checks in every cell). Draft the two documents together and put them in front of the sponsor together, and only then is contracting complete.
> Companion chapter(s): Chapter 1. Write it at project kickoff (around when the charter is signed), 15โ30 minutes. The method comes from Gary Klein's pre-mortem (a post-mortem done in advance). This template adapts it for the deliverer's setting.
## 4.3 How to Write It
The premise. **It is six months from now and this project is dead.** Work backward to how it died.
Each cause of death must: (1) map to one of the five gaps, (2) come with a defense you can start this week, (3) name an owner for that defense.
## 4.4 Template
```
To: [sponsor]
Subject: The five most likely ways [project name] dies
Assume this project has failed six months from now. Here are the most likely causes of death and their defenses.
Cause of death 1 (data gap):
How it happens: [e.g., the core system's status field cannot be trusted, and the suggestion is built on the wrong status]
Defense: [e.g., reconcile the data during discovery] Owner: [ ] Start by: [ ]
Cause of death 2 (workflow gap):
How it happens: [e.g., the system never enters the screen the user opens every day, and no one logs in after two weeks]
Defense: [ ] Owner: [ ] Start by: [ ]
Cause of death 3 (trust gap):
How it happens: [e.g., one wrong suggestion causes an incident, and the front line stops trusting any suggestion from then on]
Defense: [ ] Owner: [ ] Start by: [ ]
Cause of death 4 (ownership gap):
How it happens: [e.g., the security review does not start until week 10, and it blocks launch]
Defense: [ ] Owner: [ ] Start by: [ ]
Cause of death 5 (value gap):
How it happens: [e.g., three months in, no one can say what was saved, and the project gets "archived as a success"]
Defense: [ ] Owner: [ ] Start by: [ ]
```
## 4.5 Rules
- Make each cause of death concrete enough that you can picture the meeting on that day. Do not write vague phrases like "low user adoption."
- Cover at least four of the five gaps. If every cause of death lands in one gap, you have not yet seen the others.
- Send the finished memo to the sponsor. The pre-mortem's second function is as a stakeholder detector. Whoever it draws out to come talk to you is often a key role you had not identified yet.
- Review it monthly. The cause-of-death list is a living document. Add new ways to die as you find them during the pilot.
## 4.6 Reference Library of Common Ways to Die (by Gap, to Check Against While Drafting)
- **Data**: the source of truth is one team's private Excel; the status field lags; the join keys are unstable; the data owner will not grant production access
- **Workflow**: the user is asked to open an N+1th system; suggestions have no owner; exceptions have no exit; the time saved gets eaten by new review work
- **Trust**: one unsafe incident in the first month decides everything; suggestions cannot be questioned because they carry no reason; the front line feels watched instead of helped
- **Ownership**: the security or compliance review starts too late; no owner after launch; responsibility never fully transfers, so the system goes dark when you change roles or take leave; the business-line engineers never co-built it; you never step out of the daily and the team becomes permanent ops
- **Value**: the metric is an activity, not an outcome; nobody claims the North Star; the success criteria get discussed the week before the demo; ROI can only be expressed as accuracy
---
## 4.7 Counterexample: A Tidy-Looking Wrong Answer
Every one of the seven elements has words in it, and both sides will sign, because it binds nobody. Excerpts:
```
Element 1 North Star: complete the exceptions action queue launch and train everyone; also improve reviewer satisfaction.
Element 2: actual user: the relevant colleagues in Claims; owner: the delivery team (holding it during the transition).
Element 3 not this phase: (blank, "no requests currently need excluding")
Element 4 data: IT will give strong support and open the relevant access as soon as possible.
Element 5 business-side commitment: the business department strongly supports this and agreed verbally at its weekly.
Element 6 release: the specific release criteria will be agreed between the two sides after launch.
Element 7 exit: if the business side is not satisfied, it may terminate the cooperation at any time.
```
Item by item:
1. "Complete the launch and the training" is an activity, not an outcome. It cannot fail to come true, so it fails element 1's test. And "also improve satisfaction" stuffs in a second North Star. The correct form is one measurable outcome metric plus a baseline plus a derivation of the target.
2. "The relevant colleagues" is not a name. Filling owner with the delivery team itself writes the Chapter 22 handoff failure into the charter back in Chapter 4 (see element 2's self-check).
3. "Not this phase" left blank means every request that was ever turned down can revive at the next meeting.
4. "Strong support" has no named approver and no date. This is that week 9 email nobody answers.
5. The commitment cell, "the business department strongly supports this and agreed verbally at its weekly," has no name, no hours, and has entered nobody's calendar system. That is a one-sided love letter.
6. "Criteria to be determined" is exactly the wording element 6 declares unacceptable. What is acceptable is "the release mechanism is signed, thresholds to follow in the eval spec attachment."
7. The exit clause is one-way with no trigger condition. That is a disclaimer, not a mechanism for making sure the person who calls a stop does not have to be the villain. Only when all three parties are equal and each trigger condition is written down explicitly do you have workable resource reassessment conditions.
---
## Code Hooks
The companion repo provides (this repository's `repo/` directory):
- [`templates/deployment-charter/`](https://github.com/hallieren/the-last-mile/tree/main/repo/templates/deployment-charter/): the document scaffolding for this template and the AI prompt script that turns meeting minutes into a charter
- [`templates/premortem/`](https://github.com/hallieren/the-last-mile/tree/main/repo/templates/premortem/): the pre-mortem facilitation prompt
---
# Template 5 ยท Stakeholder Map
> Companion chapter(s): Chapter 5. One map plus one self-check. Draw it in week 1 of the project, then update it every two weeks.
> License: Every template in this book may be modified freely and used in your work, no attribution needed.
---
## 5.1 The Six-Role Map (Main Table)
| Role | Real Name | What They Fear | What They Win | Trust Balance (- / 0 / +) | Contact Plan (see 5.3) |
|------|------|----------|----------|------------------------|---------------------|
| Decision-maker | | | | | |
| Actual user | | | | | |
| Data owner | | | | | |
| Risk owner | | | | | |
| Maintainer | | | | | |
| Blocker | | | | | |
**Rules**
- **Real names.** A cell with a department name counts as blank. Departments do not fear and do not win. Only people do.
- **One person can fill several cells.** Mark anyone in several cells with an asterisk (a high-leverage node, cultivate first).
- **Blank = project risk.** Blank maintainer โ cause of death at handoff. Blank blocker โ not that he does not exist, but that you have not found him yet.
- **The opening trust balance is not zero, it is inherited.** Working relationships, the reputation the last system left in this department, the standing impression of "system builders" as a type, all of it enters the books before you do. The inherited value is occasionally positive, usually negative. Fill in this column's starting point first, then add your own deposits on top.
- Two mandatory questions once the map is drawn. **Which name has never been in the meeting room? Which cell's content did you guess?** The first is a blind spot, the second goes for verification. A wrong guess at "What They Fear" is more dangerous than a blank.
- AI project reminder (the double trust deficit, see Chapter 5). If the actual user's "What They Fear" does not mention "being replaced" or "taking the blame for AI," most likely you did not ask. Not that he is unafraid.
---
## 5.2 Identifying Questions per Role and the "Fear / Win" Prompt Library
The "common fears / wins" below are a prompt library, not answers. Every entry must be replaced by field evidence.
### Decision-maker
- **Test.** Who can decide this project's next phase of resources, or kill it, in one sentence? And whom does he answer to?
- **Common fears.** The project becomes his failure; he cannot give the board or his superiors a number; he gets cornered by the "all the competitors are doing it" narrative.
- **Common wins.** A number and a story he can tell; the "mastered AI" label.
- **Contact notes.** A one-page memo on a fixed cadence (pyramid structure, Chapter 13); bring only judgments and decision requests, no process detail.
### Actual user
- **Test.** Whose daily actions change after launch? Who gets interrupted by the system's suggestions every day?
- **Common fears.** Being replaced; taking the blame for AI's mistakes; experience ignored; being monitored.
- **Common wins.** Less repetitive work; his judgment written into the system with his name on it; standing in the department.
- **Contact notes.** Go to the desk, not the meeting room; use his jargon; a fixed time slot written into the charter (say "2 hours a week"); say the fear of replacement out loud, to his face (Chapter 5).
### Data owner
- **Test.** Who owns the data you need **on the business side**? (Not who holds the database key)
- **Common fears.** Definition problems landing on him after the data is misused; poor data quality exposed.
- **Common wins.** Someone fixes a data definition for him; his own reporting gets easier.
- **Contact notes.** Solve one of his own data pains first, then talk about access; when the request stalls, let him speak in his own language of power.
### Risk owner
- **Test.** When something goes wrong, whose name is on the accountability email? Who signs the security / compliance / legal review?
- **Common fears.** Being bypassed, last to know; taking the fall; an unauditable black box.
- **Common wins.** A project that writes risk on page one; a case of prudence he can show upward.
- **Contact notes.** Invite him into the project team in week 2, do not send it for review passively in week 10; hand-deliver the pre-mortem memo (Template 4).
### Maintainer
- **Test.** A year after launch, whose annual goals list this system?
- **Common fears.** Inheriting a black box; an undocumented mess that becomes his later.
- **Common wins.** New skills; ownership of the code; a co-build credit.
- **Contact notes.** Co-build from the first line of code (Chapter 15), not a handoff document in the final week.
### Blocker
- **Test.** Without whose nod does everything stop, even though he never attends? (PMO, finance, information security, a peer digital counterpart, the architecture committee)
- **Common fears.** Every blocker is different, but behind the blocking there is always something he is protecting. Find it.
- **Common wins.** Once what he protects is written explicitly into the plan and protected, he loses his reason to block you (resistance is a diagnostic signal, Chapter 20).
- **Contact notes.** One on one, early, in person; never ambush a blocker in a big meeting.
---
## 5.3 Contact Plan
| Person | Frequency | Format | What to Bring Next Time (**what he wants, not what you want**) | Last Contact Date |
|----|------|------|------------------------------------------|--------------|
| | | | | |
**Rules**
- Bring something useful to **him** every time. A one-page risk list, a fixed spreadsheet, a data point he cares about. Contact empty-handed spends trust, it does not save it.
- A key cell with no contact for two weeks โ mark it red.
- Contact โ meeting. A hallway, a desk, a three-line email all count.
---
## 5.4 Trust Equation Self-Check
Source. The Trust Equation is paraphrased from Maister / Green / Galford, *The Trusted Advisor*: trust = (credibility ร reliability ร intimacy) / self-orientation. Run it on yourself every two weeks. Any question you cannot answer with a concrete example is the next behavior you should deliberately create.
**Credibility**
- [ ] In the past two weeks, I said to someone's face "I do not know, I will have an answer by X," and delivered on time.
- [ ] Most of my last three judgments carried evidence from this business unit's floor (numbers, cases, users' own words).
- [ ] The business side quoted me in a meeting I was not in.
**Reliability**
- [ ] Of my last five "you will have it Friday" commitments, at least four landed on time.
- [ ] Meeting action items go out in writing within 24 hours.
- [ ] At least once, I shrank a commitment I could not keep, early and unprompted, instead of explaining on the deadline.
**Intimacy**
- [ ] Someone has told me something at the level of "I am only telling you this."
- [ ] I have said the tension in the room out loud, to people's faces, at least once.
- [ ] The front line uses their own jargon with me, not the official vocabulary.
**Self-orientation (lower is better, a check on any of these four is bad news)**
- [ ] "Our system / our plan" shows up in my speech more than "your claims / your metrics."
- [ ] When the business side suggests shrinking the project scope, my first reaction is defense.
- [ ] I have never advised the business side "let's not do this yet."
- [ ] What my reporting says is "the system we built," not "the Claims queue."
---
## Code Hooks
The companion repo provides (this repository's `repo/` directory):
- [`templates/stakeholder-map/`](https://github.com/hallieren/the-last-mile/tree/main/repo/templates/stakeholder-map/): the table scaffold for this template and the periodic update reminder script
---
# Template 6 ยท Field Archaeology Kit
> Companion chapter(s): Chapter 6. The three tools are in order of use. Use the half-day guide to book and complete one shadowing session, keep the friction log throughout, and switch on the three-layer probing script when you hear "I can tell at a glance."
> License: Every template in this book may be modified freely and used in your work, no attribution needed.
---
## 6.1 Half-Day Shadowing Guide
### Booking It
- **Who to book**: the actual user, the person whose daily moves the system will change after launch, not their supervisor. If the supervisor arranges it, name the person explicitly.
- **Opening script** (three elements: apprentice posture + capped duration + zero-interruption promise):
> "I would like to sit with you for half a day and watch how you normally handle these claims. I am here to learn. I am not here to evaluate you, and I am not here to sell a system. Work as you normally do, no need to explain anything to me. I will sit to the side and take notes, save my questions, and ask when you have a free moment. I will take at most 30 minutes of your wrap-up time."
- **When you get rescheduled**: wait, book again, do not escalate to a superior to apply pressure. The reschedule itself is information, and front-line time is the most expensive. The test is the reason, not the count. If the reason changes every time, you are being brushed off and escalation is warranted. If the reason is backlog every time, the place you want to look is exactly where it hurts most, so wait. Consider escalating only after three or more in a row, and when you do, talk about schedule protection, not attitude, and land it in the business-side commitment line of the charter (Chapter 4).
### Prep Checklist for the Day Before
- [ ] Print one copy of the official flowchart / SOP (your dig map, for marking differences)
- [ ] Print a blank friction log (see 6.2)
- [ ] Confirm with the person: may I look at your screen? May I take notes? Where sensitive data is involved, the default is **no recording, no screen photos** (notes are enough, trust is worth more; having permission does not mean you should use it)
- [ ] Think through your red lines and keep them: no promising features, no judging how they work, no suggestions on the spot
- [ ] Prepare 1โ2 candidate claims to trace: after following the person, pick one claim and follow it end to end
### Half-Day Schedule (4 Hours as an Example)
| Slot | Action | Discipline |
|------|------|------|
| 0:00โ0:10 | Opening (repeat the script above) | Make the role clear: apprentice, not audit |
| 0:10โ2:00 | Pure observation | No interruptions; write every question down and save it |
| 2:00โ2:15 | Tea-break question window | Ask only about actions you observed, no more than 3 questions at a time |
| 2:15โ3:30 | Continue observing + pick a claim to trace | Count steps, count tool switches, note off-screen actions (phone calls, calling over a colleague) |
| 3:30โ4:00 | Wrap-up questions | The three-layer probing method (6.3); the three closing questions (below) |
### What to Count While Observing
- **Step count**: real steps vs. official process steps (the difference is the width of the workflow gap)
- **Tool switches**: core system / inbox / spreadsheet / phone / chat tool, how many switches to each
- **Off-screen actions**: phone calls, calling over a colleague, leafing through paper files (the easiest to miss, and often the most valuable)
- **Workaround tools**: anything open on the screen that is not on the IT asset register (private spreadsheets, sticky notes, personal folders, small group chats)
- **Ten-second judgments**: the moments the person pauses and then decides outright (mark the time, save for the three-layer probe)
### Code of Conduct (Break Any One and the Half Day Is Wasted)
1. No suggesting ("actually, you could...", you are teaching the master their craft)
2. No correcting (see a "noncompliant" workaround, note it, do not call it out)
3. No selling the system (not even one "the system will be able to do this for you later")
4. No promising features (a promise makes everyone you observe from then on start performing)
5. Noting without calling out is not permanent secrecy (not correcting on the spot governs this half day; whether to report is a separate question; if you really must report, tell the person first, face to face)
### The Three Closing Questions
1. "Was today a typical day? What was not typical about it?"
2. "If you took a week off, who would do this work? Which step would they get stuck on?"
3. "Of everything in this half day, which thing did you feel was least worth your time?"
---
## 6.2 Friction Log Template
> friction log (defined in Chapter 0): a running list of friction, "the places where reality and paper do not match." During shadowing it is the main recording tool. After discovery it is the upstream of the data reconciliation checklist and the eval material.
| # | Time | Paper version (what the SOP / system / report says) | Field version (what actually happened) | Type | Follow-up |
|---|--------|--------------------------------------|----------------------------|------|----------|
| 1 | | | | data / tool / process / judgment | reconcile / probe / into eval / report risk |
**The four-way type split**:
- **Data**: a field does not match reality (status "in progress," actually "waiting two weeks for documents") โ follow-up is usually reconciliation
- **Tool**: a workaround tool carries the function of the formal system (Excel is the source of truth) โ follow-up is usually bringing it into the data source inventory
- **Process**: real steps added to, removed from, or changed against the SOP (official 5 steps, actual 14) โ goes into the real flowchart
- **Judgment**: a human judgment that cannot be derived from rules (the ten-second red flag) โ switch on the three-layer probe
**Recording rules**:
1. One line per friction. Write keywords on the spot, complete within 24 hours. Details recalled the next day are invented.
2. Record facts, not conclusions ("status field lags on 3 claims" is fine, "their system is terrible" is not).
3. Every line must have a follow-up entry, or the log becomes a gripe list.
4. AI use: dictating or photographing your notes and feeding them to AI for structuring and clustering is fine, but **the raw observation must be what you wrote down while present**. AI organizes archaeology notes. It does not produce archaeology facts.
---
## 6.3 Tacit Knowledge Probing Script (the Three-Layer Probing Method for "I Can Tell at a Glance")
### Trigger Phrase List (Switch On When Heard, Mark the Time on the Spot, Probe in the Wrap-up Window)
- "I can tell at a glance" / "obviously fake" / "gut feeling" / "from experience"
- "Hard to say" / "I cannot explain it, it just feels off"
- "Cases like this are always like that" / "Do it long enough and you get it"
- And any judgment **made within ten seconds whose reason you cannot derive yourself**
### Layer One: Anchor on an Instance, Replay the Actions, Do Not Ask Why
> "That claim just now, you looked at [action 1] first, then went through [action 2], and then you [judgment]. Right?"
- Principle: ask "why" straight out and you get an on-the-spot rationalization or "I can tell at a glance." Replay **the concrete action that just happened** and the person cannot answer with a stock phrase. A stock phrase does not match actions.
- Key points: it must be a real instance observed in this session; hypothetical questions ("if one came in that was..., how would you judge it?") are banned throughout the script.
### Layer Two: Compare, Find the Difference Between Two Similar Claims
> "This one and the one just now are both [same type], and the [surface indicator] is about the same. Why did you [flag] this one and not that one?"
- Principle: a single instance gets you a general description; a pair of minimally different instances forces out the **boundary condition**. The real shape of the rule is in the difference.
- Key points: best to pick the comparison claim on the spot from the ones they handled today; once the difference comes out, follow with "any other difference?", the second answer is often worth more than the first.
### Layer Three: Boundary Counterexamples, When Does It Not Hold
> "When would you not [judge it this way], no matter how [extreme the indicator]?"
> "Have you ever [judged it this way] and then found you were wrong? What did you change after that?"
- Principle: the boundary of a rule and its failure cases are the part experts themselves are rarely asked about; "what did you change after that" also digs out the **rule's update mechanism**. That is the key evidence against hardcoding a snapshot.
- Key points: when asking for counterexamples the tone is curiosity, not challenge; if the person answers "never been wrong," switch to "when you train new people, what mistake do they make most often on this kind of claim?"
### Wrap-up: Restate as Rule Sentences, Have the Person Correct Them
Write every judgment you dig out in one uniform sentence form:
> **When [observable signal] appears, I [action], because the risk of not doing so is [consequence].** (Boundary: [when it does not apply]; source: [name + date + instance claim number]; stability: stable / changes with [X])
Three disciplines:
1. **Verify back**: take the write-up back within 48 hours for the person to circle the mistakes. Ask "where is it written wrong," not "is it right."
2. **Mark stability**: rules that drift (list types, threshold types) must have their drift source marked; this entry directly decides how the rule enters the system. Stable rules can be made explicit as suggestion logic. Drifting rules must keep the Human Call and a decision trail for overrides (Chapter 6 failure mode 4, Chapter 17 mechanism).
3. **Mark destination**: every rule notes its downstream use, into the real flowchart / into an eval error category (Chapter 11) / into the risk list. A rule with no destination dies in the notes.
---
## Code Hooks
The companion repo provides (this repository's `repo/` directory):
- [`templates/field-archaeology/`](https://github.com/hallieren/the-last-mile/tree/main/repo/templates/field-archaeology/): structured friction log template and AI cleanup prompt scripts (transcribe โ structure โ cluster)
---
# Template 7 ยท Five-Question Opportunity Rubric
> Companion chapter(s): Chapter 7. The three pieces are ordered by use. Score question by question first (7.1), then summarize and eliminate (7.2), then report (7.3).
> License: Every template in this book may be modified freely and used in your work, no attribution needed.
---
## 7.1 Five-Question Scorecard
One card per candidate use case. Score each question 1โ5. **Any score with a blank evidence column is treated as a 2.** Numbers, users' own words, and data samples count. Adjectives do not.
**Blank scorecard** (the card header carries the candidate use case name / scorer / date):
| Question | Score (1โ5) | Evidence (numbers / users' own words / data samples) |
|---|---|---|
| Pain | | |
| Data | | |
| Decision | | |
| Risk | | |
| ROI | | |
### Question 1 ยท Pain
**Test question.** Whose action is bleeding? Is the bleed rate (loss per unit of time, hours, complaints, fines, churn) measurable? The implication follow-up. If it goes unfixed for three months, what happens? Who suffers?
| Score | Anchor |
|----|------|
| 5 | A specific person + a specific action + a measured bleed rate; leave it and someone keeps taking blame or paying out |
| 3 | The pain is real, but the bleed rate has not been measured. Measure it first, then come back and score |
| 1 | Cannot name the person or the action; or after the follow-up "what happens if it is not fixed," the answer is "not much" |
**What the evidence column requires.** Users' own words / ticket volume / waiting time / complaint records. "Everyone feels that way" is not allowed.
**Note.** A candidate whose pain vanishes under the follow-up (for example, "nobody reads the monthly report closely") is not a regrettable low score. It is this card's most valuable output.
### Question 2 ยท Data
**Test question.** What data does this action need? Does it exist? Can you get it? Can you trust it? (A pre-check here. The full grading method is the data fitness ladder in Chapter 9.)
| Score | Anchor |
|----|------|
| 5 | Data exists, the business owner is named, samples are in hand, spot checks hold up |
| 3 | It exists and is reachable, but trustworthiness is unverified. A conditional pass, with the reconciliation or validation action and its deadline written down |
| 1 | The key data does not exist, or there is no realistic path at all to getting it |
**What the evidence column requires.** Data samples / field spot-check results / the name of the data's business owner. "The business unit says it is all there" does not count as evidence, and neither does "we assumed it was all there" (Chapter 9 will tell you what that sentence is worth).
### Question 3 ยท Decision
**Test question.** Which decision point in the real workflow does AI enter? How many times may it be wrong at that step? Who backstops it, and how fast does the error surface?
| Score | Anchor |
|----|------|
| 5 | It enters at the advise layer; the backstop is named; errors are visible on the spot and can be overridden |
| 3 | There is a backstop, but discovery of the error is delayed, or the backstop has not yet confirmed |
| 1 | Real-time and outward-facing with nobody backstopping; or a decision with near-zero error tolerance (money, compliance, safety) executed directly by AI |
**What the evidence column requires.** Where the decision point sits in the real workflow (the output of the field archaeology in Chapter 6, not its position on the official flowchart) + the backstop's name.
### Question 4 ยท Risk
**Test question.** What does the worst single output look like? Does it reach only inside, or customers, regulators, money? Can the loss be closed out (found, corrected, made good)?
| Score | Anchor |
|----|------|
| 5 | The worst output's loss stays inside and can be closed out, and it sits outside the agreed red lines |
| 3 | There is a path outward, but a gate sits on it (human review, delayed sending, an amount cap) |
| 1 | One bad output can reach a customer, a regulator, or money, and it cannot be taken back |
**What the evidence column requires.** Write out the imagined text of that "worst output" word for word (for example, "Your loss assessment payment is expected within three business days"). If you cannot write it, you have not thought about it seriously.
### Question 5 ยท ROI
**Test question.** What is the North Star outcome metric? What is the arithmetic that turns the improvement into money? Who owns that bill?
| Score | Anchor |
|----|------|
| 5 | An outcome metric + the conversion arithmetic + a person willing to carry that number in his own reporting |
| 3 | The metric is measurable, but the conversion is rough or the person owning the bill has not confirmed |
| 1 | Only activity metrics ("it launched," "92% accuracy"). Chapter 1 covered why they do not survive the boardroom |
**What the evidence column requires.** The arithmetic itself, and the name of the person who owns the bill.
---
## 7.2 Elimination Table
All candidates summarized on one table:
| Candidate | Proposed By | Proposer's Rank / Does He Control Your Budget | Pain | Data | Decision | Risk | ROI | Lowest Score (Which Question) | Conclusion | Referral Route (Fill In When the Conclusion Is Elimination) | Condition / Re-evaluation Condition (Who Rechecks It, at Which Standing Meeting) |
|------|--------|--------------------------------|------|------|----------|------|-----|------------------|------|------------------------------|----------------------------------------|
| | | | | | | | | | Winner (conditional) / Eliminated / Needs more evidence | Refer to a standard tool / Refer to self-serve / Refer to an outside purchase (with ops ownership) | |
**Seven hard rules:**
1. **Any question scoring โค2 is eliminated outright.** No weighting, no averaging, no "considering it in the round." The five questions stand in a multiplying relationship. One factor close to zero makes the product zero whatever it multiplies. Wherever you see a "weighted total" row, delete it.
2. **A 3 = a conditional pass.** The condition (what evidence to add, what validation to run) goes into the table with a deadline. Not met by the deadline, it drops to 2.
3. **Elimination โ never.** Every eliminated candidate gets one line of re-evaluation condition, what evidence, if it appears, makes it worth running the five questions again. The re-evaluation condition also needs a recheck owner and a place to be rechecked. Writing it down is not the same as being remembered, and a re-evaluation condition nobody rechecks is a politely worded permanent refusal. The default is to hang the elimination table on the team's quarterly ask review and read it out each quarter. "Not now" needs an exit before elimination can be enforced at all.
4. **The evidence column is all that counts.** In 7.1, any score with a blank evidence column is treated as a 2. Imagined scores are this table's most common forgery.
5. **Eliminating everything is a legitimate result.** Write the conclusion as "back to discovery to find an opportunity," not "loosen the standard and screen again."
6. **Do not mix the scales.** This table's 1โ5 screens opportunities (you are the scorer, the logic is finding the weakest link). The Field MVP's pass / concern / unsafe / useless is for real users testing a single output. The two scales are not interchangeable.
7. **Elimination needs an exit.** Every row concluding in elimination also gets a referral route, refer it to a standard tool, refer it to self-serve, or refer it to an outside purchase with ops ownership written down. Leave it blank and the one who feeds it from then on is you.
**Worked example** (the three Anchor & Helm candidates, the full working is in Chapter 7):
| Candidate | Proposed By | Proposer's Rank / Does He Control Your Budget | Pain | Data | Decision | Risk | ROI | Lowest Score (Which Question) | Conclusion | Referral Route | Condition / Re-evaluation Condition |
|------|--------|--------------------------------|------|------|----------|------|-----|------------------|------|----------|------------------|
| Service chatbot | The board (forwarded by Grant Whitmore) | Top level, does not directly control your team's schedule | 3 | 1 | 2 | 1 | 3 | 1 (Data / Risk) | Eliminated | / (none of the three ready-made routes applies, gather evidence first) | Re-measure the call mix after exception handling is cured; the standard-answer library gets built (recheck owner, you; where, the team's quarterly ask review) |
| Automated monthly reports | Kevin Doyle | Business unit head, does not control your team's budget | 1 | 3 | 4 | 4 | 2 | 1 (Pain) | Eliminated | Refer to self-serve (turned into a favor done in passing, forty minutes' worth) | No project approval (recheck owner, you; where, the team's quarterly ask review) |
| Exceptions action queue | The Field MVP readout | Not applicable (produced inside the team, not an individual's proposal) | 5 | 3 | 4 | 4 | 4 | 3 (Data, conditional) | Winner | / | Finish the core system ร Excel data reconciliation during discovery (recheck owner, you; where, a precondition for pilot launch) |
---
## 7.3 Reporting One-Page Format
For reporting the screening result to the sponsor. One page, answer first:
```
Conclusion. [winning candidate] moves into charter negotiation / discovery.
One-sentence reason + one number (example, "about six tenths of the waiting time is 'waiting on documents with nobody chasing,' and the queue can eat half of it").
Candidates and results. The full elimination table (with scores and which question the lowest score sits on).
For each eliminated candidate, three lines.
Cause of death. Question X, N points.
Evidence. The single hardest piece of evidence (a number / a quote / a sample), not an opinion.
Re-evaluation condition. What evidence, if it appears, triggers a re-evaluation.
Next step. The winning candidate's time box, the data access and staffing it needs
(this column is the input to the charter negotiation, see Chapter 4).
Opportunity cost, the other things this team is not doing because of it (the candidates ranked behind, which row each is stuck on, which quarter each waits for, written for the sponsor and for the department that ranked behind).
```
**Facilitation rules:**
- **Language.** Do not say "I do not recommend it." Say "the evidence says not yet." A position gets haggled over, and evidence can only be rebutted by better evidence.
- **When eliminating an executive's proposal.** Let the table speak and read only the evidence column, and give him one sentence he can repeat upward (example, "treat the source of the calls first, and the foundation for service automation gets laid along the way"). What he needs is not to be persuaded. It is a reason he can pass on.
- **Say it in person. Do not fire the table off by email.** Eliminating someone's idea is a trust event, not an information transfer (Chapter 5).
- **One page, no more**, and the most an attachment may be is the 7.1 scorecard itself.
- File the elimination table after the report. An internal elimination table has to live for years, not three months, so when the subject comes back up, what you open is the evidence and the re-evaluation condition from the time, not your memory.
---
## Code Hooks
The companion repo provides (this repository's `repo/` directory):
- [`templates/five-questions/`](https://github.com/hallieren/the-last-mile/tree/main/repo/templates/five-questions/): the scoring tables for this rubric and the script that generates the reporting one-pager
---
# Template 8 ยท Thin Slice Definition Sheet and Scope Decision Log
> Companion chapter(s): Chapter 8. The four tools work together. Fill in the definition sheet once at project approval and recheck it at every milestone. Draw the boundary map at the same time as the definition sheet. Keep the scope decision log continuously from project approval on. Run the no slice self-check after every narrowing.
> License: Every template in this book may be modified freely and used in your work, no attribution needed.
---
## 8.1 Thin Slice Definition Sheet (the Five Ones)
> Thin slice: the smallest working slice that cuts vertically through all five layers, "real user, real decision, real data, real risk control, measurable outcome." Thin is the width. The depth must go all the way down.
**Blank template**:
| The Five Ones | What to write | Self-check question |
|--------|----------|----------|
| **One user group** | Specific to the team and the names, not a department | Who in this group will say "this solves my problem"? Cannot name them โ back to Chapter 5 to find the actual user |
| **One decision** | One decision point the user **already** makes, not a newly invented action | If the system's suggestion is wrong, can the user see it on the spot and override it? |
| **One data path** | One verifiable path from data source to user interface, listing every source it passes through | Has the fitness of every source on the path been checked (Template 9)? Has "after reconciliation" actually been delivered? |
| **One risk boundary** | The decision rights boundary (which layer AI stops at) + the red lines (what it never touches) | Besides the documents, do the red lines also appear in the review materials and the product interface? |
| **One measurable outcome** | Cite the charter North Star directly (Template 4), no separate metric | When this number gets worse, does anyone hurt? |
**Anchor & Helm example** (the slice from Chapter 8):
| The Five Ones | Anchor & Helm's entry |
|--------|----------|
| One user group | Linda Marsh's auto claims review team (not "the claims department") |
| One decision | The reviewer's next step on each exception claim, which documents are missing and whom to chase first (an existing decision point, the suggestion can be overridden on the spot) |
| One data path | Core system status fields + the review team's Excel + the survey-review mailbox โ merged read-only view โ the queue (fitness of each source, see the Template 9 example rows) |
| One risk boundary | AI stops at the advise layer (8.2); red lines, no payout decisions, no automated outbound messages |
| One measurable outcome | First-touch handling time for auto exceptions (charter North Star, Template 4 element 1) |
**Rules**:
- Any cell containing "and / as well as / etc. / two kinds" โ that cell is already creeping. Negotiate the narrowing first, then fill in the sheet.
- The five cells must lock together into one readable sentence: **"This user group, making this decision, along this data path, inside this risk boundary, improves this metric."** Wherever it does not read smoothly is where the cut is not clean.
- The definition sheet is not a one-time document. Recheck it at every milestone (prototype โ pilot โ launch). Have any of the Five Ones quietly become plural?
---
## 8.2 Decision Rights Boundary Map
The four-layer skeleton (bottom up). Rewrite each layer's content for your workflow:
```
Decide [this workflow's final decision, e.g. pay or not, how much] <- red line: never touched
Act [actions with external effect, e.g. send chase notices, change status] <- humans take over here (name who)
Advise [priority, next step, each with a reason] <- AI stops here (this phase)
Sense [gather, extract, reconcile, flag signals] <- AI does this
```
Three entries per layer:
| Layer | Owner this phase (AI / human, name the human) | Evidence needed to move up | Move-up approver |
|----|------------------------------|--------------|-----------|
| Decide | | | |
| Act | | | |
| Advise | | | |
| Sense | | | |
**Rules**:
- Moving up is a **milestone decision**, not routine tuning. "Let it just send this one automatically" is where the boundary starts to fall.
- "Evidence needed to move up" takes two forms, a human acceptance record (e.g. the acceptance rate for one class of suggestion โฅ threshold for N consecutive weeks), or a deterministic verification loop on that layer's output (errors automatically detectable, enumerable, reversible, with proof of verification coverage). Both are written as measurable facts, never as "once results stabilize."
- Every move up reruns the three oversight questions (Chapter 12). Does the overseer have time to look? The ability to judge? The authority to stop it?
- This map appears in at least three places, the scope decision log's attachment, the security review packet (Template 12), and the memo to decision makers (Template 19).
---
## 8.3 Scope Decision Log
| Date | Expansion | Proposer | Proposer's relation to you | Reason for refusal (the cost, one sentence) | Revival condition (a checkable event) | Status |
|------|----------|--------|--------------------|--------------------------|------------------------|------|
| | | | | | | shelved / revived / abandoned / transferred |
| | | | | | | |
Anchor & Helm example (as recorded in Chapter 8):
| Date | Expansion | Proposer | Proposer's relation to you | Reason for refusal (the cost) | Revival condition | Status |
|------|----------|--------|--------------------|------------------|----------|------|
| Proposed the day of signing, logged the next day | Home property exceptions | Kevin Doyle | pilot owner | User group, annotation system, and metric baseline all double, diluting the pilot evidence | The **first** extension after the pilot North Star hits target (exclusive) | shelved |
| Day of signing | Full automation | Grant Whitmore | sponsor | The decision rights boundary stops at the advise layer this phase, and the act layer has no acceptance data | Once the advise layer's acceptance data hits target, move up one layer at a time, each layer decided on its own | shelved |
| Day after signing | General-purpose workflow platform | The older engineer | co-build engineer (claims-ops IT, same project team) | Not one of the five gaps narrows with a platform, and two widen | After the second slice lands, distill what is common | shelved |
**Rules**:
- **Log it on the spot, send it back to the proposer within 48 hours.** The log itself is the proposer's answer. Most proposals want not a "yes" but an answer they can repeat. For a proposal from another department, cc the sponsor when you send it back. The cc is not tattling. It sends a lateral request into the hands of the person with the authority to rank it.
- Write the reason for refusal as a **cost**, not an attitude ("not for now" and "not a priority" are attitudes; "eval doubles, pilot slips six weeks" is a cost).
- Write revival conditions as **events**, not dates. "Look again in Q3" and nothing happens when Q3 comes. "After the pilot hits target" and when the event happens the proposer comes to open this page himself. Once the system enters the operating period, pilot events like "hits target" no longer occur, and revival conditions switch to operating-period anchors. One kind hangs on the periodic resource review, written as "ranked with the others at next quarter's resource review." The other hangs on an operating metric, written as "revisit after such-and-such operating metric holds within threshold for N consecutive weeks." An entry that cannot be given either anchor is really a new project already and should go through project approval.
- The charter's "not this phase" list (Template 4) is this log's first batch of entries. After signing, add the cost and revival condition to each.
- **The half-day rule**: any "favor" over half a day (scattered support that goes through no project approval and no charter) also goes in this log. The proposer writes their name and their relation to you, and the refusal reason column states which item of this phase it crowds out.
- **Reread the revival list** at every milestone retrospective. Entries whose revival condition is met either start or get their condition explicitly rewritten. Renege once and this log loses its credibility. On the day the pilot hits target, this log is the ready-made phase-two roadmap (Chapter 23).
- Platform proposals all get the same revival condition, "after the second slice lands." A good abstraction grows out of the second case. It is not guessed from the first.
---
## 8.4 No Slice Self-Check (Over-Narrowing Check)
Run this after every narrowing. Any item marked โ means you cut too far, and the thin slice has become a no slice:
- [ ] At least one real user in the "one user group," after hearing the slice described, will say "this solves my problem" (verify with their own words, do not answer for them).
- [ ] The slice runs end to end, data โ sense โ advise โ human decision โ outcome measurement, one whole path, with no gap where "this stretch is handled by hand for now."
- [ ] The North Star's improvement is still attributable to this slice (if what is left cannot move the metric, the metric has become an empty promise).
- [ ] The change users make to use it is smaller than the effort it saves them (otherwise the resistance of Chapter 20 arrives early).
**Judgment**: the "minimum" in thin slice is the minimum that is complete in value, not the minimum that is easy to engineer. A slice that avoids every risk also avoids the risk of success.
---
## Code Hooks
The companion repo provides (this repository's `repo/` directory):
- [`templates/thin-slice/`](https://github.com/hallieren/the-last-mile/tree/main/repo/templates/thin-slice/): table scaffolds for this template set and a lightweight maintenance script for the scope decision log
---
# Template 9 ยท Data Source Inventory, Fitness Scorecard, Reconciliation Checklist, Source of Truth Decision Log
> Companion chapter(s): Chapter 9. The four tools are in order of use. Inventory first (9.1), then score (9.2), verify the most expensive rung by reconciliation (9.3), and finally settle the conclusion into a source of truth decision (9.4).
> License: Every template in this book may be modified freely and used in your work, no attribution needed.
---
## 9.1 Data Source Inventory
Inventory first, score second. Anything that carries a business fact counts as a data source, including the ones you are embarrassed to call data sources, private Excel files, mailboxes, chat groups, paper ledgers. In the LLM era they are legitimate candidate sources, and often the only ones telling the truth.
| Data Source | Type | Business Owner | Key Holder (Permission Path) | Coverage | Update Method and Frequency | Known Problems | Target Action Supported |
|--------|------|-----------|------------------------|----------|----------------|----------|----------------|
| e.g. core claims system, claim status | Core system field | Director of Claims Operations | IT data team (request process + business owner sign-off) | All claims | Triggered at process nodes, lags by the week | "In progress" carries inconsistent meanings | Exception queue ordering |
| e.g. review team tracking Excel | Private spreadsheet | Review team lead | The team lead herself | This team's auto exception claims only | Manual, same day | Maintained by one team only, manual daily backup only | Candidate source of truth for claim status |
| e.g. survey-to-review mailbox traffic | Unstructured | / | Mailbox administrator + compliance | Document request correspondence | Real time | Needs LLM extraction, carries an error rate | "Waiting on documents" signal |
**Rules**:
- **Fill in business owner and key holder separately**. The one holding the key (IT) approves the permission. Only the data's business owner (the business lead) can make the approval move. The way out of a permission standoff lies with the latter (Chapter 5).
- **Coverage is required**. The most common trap with a private source of truth is that it covers one team only, or one line of business only. Using a local source of truth as a global one is another way of lying.
- **"Target action supported" is required**. Fitness is relative to an action, and with no target action there is nothing to score.
- Always ask one question in an inventory interview: **"Outside the system, where else do you record this yourselves?"** A private source of truth never appears on any architecture diagram of its own accord.
---
## 9.2 Fitness Scorecard
Score one column per "data source ร target action" pair. **Judge upward from exists, stopping at the first rung that cannot produce evidence.** That is where it really stands. No skipping, and no verbal assurance in place of evidence.
| Rung | Test Question | Qualifying Evidence (Example) | Verdict (pass / fail / unknown) | Evidence Record |
|------|----------|------------------|---------------------------|----------|
| **exists** | Has this information been recorded? In which system, in which field? | You can point to a specific field or file location | | |
| **accessible** | Can it be read in a compliant, repeatable way (including machine access for the future system)? | Permission is in hand, and it is not a one-off export | | |
| **interpretable** | Does the same value mean one thing across departments and across periods? | A value dictionary confirmed by two or more consumers | | |
| **timely** | Does the update frequency keep up with the target action? | Measured update lag โค the latency the action allows | | |
| **traceable** | Where did the value come from, who changed it? Can the same entity be matched across systems? | Lineage can be stated, and the join key spot check passes | | |
| **validated** | Has it been reconciled against reality? | Three-way reconciliation inconsistency rate + pattern (9.3) | | |
| **actionable** | Holding it, can the "next action" be carried out by a specific person, and can an error be caught? | A trial run of the target action passes, errors detectable and traceable | | |
**Rules**:
- "Unknown" is not a middle state. It is the polite way of writing fail. A launch decision treats it as fail.
- exists and accessible are the two cheapest rungs, and an organizational assurance or a successful demo proves only this far.
- validated has no shortcut. It goes only through the reconciliation in 9.3.
- Date the scorecard. **Fitness is a state, not a property, and a scoring conclusion has a shelf life** (retest triggers in 9.4).
---
## 9.3 Three-Way Reconciliation Checklist
### Before the Reconciliation
- [ ] **Define the standard first**: write down in black and white what counts as inconsistent (suggested standard = cannot drive the target action at face value; report the total rate and the largest single class together). Define it before you touch the data.
- [ ] Choose the sample: 50โ200 records, covering typical cases + edge cases + aged claims (the longest-waiting records, where field discipline is worst).
- [ ] Assemble all three sources: the system field export, the private source of truth (spreadsheet or email threads), the list of handlers and the interview times.
- [ ] Give the data's business owner a heads-up: the reconciliation is for assigning a source of truth, not for assigning blame, and sensitive fields and handler names in the report are all de-identified, with nobody named. Say this in advance, or you will be handed a "tidied up" second version of the scene.
### During the Reconciliation (Record Row by Row)
| # | Record ID | System Field Version | Private Source of Truth Version | Handler's Version | Verdict (consistent / inconsistent = cannot drive the target action at face value) | Pattern (lagging / forgotten / semantic divergence / broken join) | Notes (the handler's own words) |
|---|---------|--------------|----------------|------------|--------------------------------------------|-------------------------------------------|---------------------|
- [ ] Two ways can only find "they differ." Every disputed item has to reach the third way (asking the handler) before "who is right" can be ruled on.
- [ ] Spot-check the cross-system join key while you are at it: match 20 random records across systems and record the match failure rate.
- [ ] Write down the handler's own words. A line like "we never look at that field" is worth more than the inconsistency rate.
### After the Reconciliation
- [ ] Produce the **inconsistency rate**, total rate and largest single class reported together ("44% cannot be used at face value, and the largest class, the lagging pattern, is about three in ten" is more honest than a single number).
- [ ] Produce the distribution of **inconsistency patterns**, and triage by pattern:
| Pattern | Signature | Prescription | Belongs To |
|------|------|------|------|
| Lagging | The value is true, only slow | A faster source / extract the upstream signal | Engineering |
| Forgotten | An update that no process forces and that nobody is affected by gets skipped | Operating discipline + the system produces a "status in doubt" list to help | Mostly operating |
| Semantic divergence | One word, two departments, two meanings | Hold a semantic alignment meeting first, talk about syncing second | Organizational |
| Broken join | The same entity cannot be matched across systems | Fix the join key, and where it cannot be fixed, demote the signal | Engineering |
- [ ] Write the result up as a source of truth decision (9.4) and set the retest date.
---
## 9.4 Source of Truth Decision Log
One class of business fact, one source of truth. Every row is a defensible decision, formatted after the scope decision log (Template 8):
| Business Fact | Source of Truth | Fallback Source | Reason for the Decision (Cite Reconciliation Evidence) | Is the Derived View Read-Only | Retest Cadence | Owner (Name + Date Claimed) |
|----------|--------|--------|--------------------------|------------------|----------|----------------------------|
| e.g. claim status | Review team Excel | Core system | Reconciliation, 61/200 system statuses lagging; Excel updated same day | Yes, writes back to no source | Monthly during the pilot | You (during the pilot) โ claims-ops IT after handoff (note the date claimed, Chapter 15) |
| e.g. payout amount | Core system | / | Runs through finance reconciliation, covered by audit discipline | / | Quarterly spot check after launch | Director of Claims Operations (note the date claimed) |
**Rules**:
- [ ] **Derived views are read-only without exception**, writing back to no source system. AI output exists as a decision trail (suggestion + reason + Human Call + timestamp), physically separated from the source fields (Chapter 17 works it out).
- [ ] Operating-side repair items (such as "status write-back") name the business-side owner and the action. That half is not something your system can fix.
- [ ] **Retest triggers** (any one of them hit reruns 9.3, without waiting for the retest date):
- Rotation in a key role (the maintainer of the source of truth changes)
- Team staffing change (the engineer on your team who owns this line changes, an internal-only trigger, counted separately from the business-side rotation above)
- A source system upgrade or a process revision
- Your system has been live for a month (it changes the front line's motivation to maintain the private source of truth, and your system will kill its own source of truth with its own hands)
- The override rate or the length of the "status in doubt" list rises abnormally (Chapter 18's monitoring signal)
---
## 9.5 Counterexample: A Tidy-Looking Wrong Answer
Scoring target: the core system's "claim status" ร exception queue ordering. The table is full, not one cell blank:
| Rung | Verdict | Evidence Record |
|---|---|---|
| exists | pass | The core system has this field, and the data platform says it is all synced over |
| accessible | pass | IT confirmed verbally that read-only access can be opened |
| interpretable | under confirmation | Emails already sent to the departments asking about the value meanings |
| timely | under confirmation | Waiting for IT to reply on update frequency |
| traceable | pass | Kevin says every change goes through approval |
| validated | pass | The demo ran on last month's export with no errors |
| actionable | pass | Kevin thinks the approach is workable |
Item by item:
1. accessible: a verbal confirmation is an organizational assurance, not evidence. The bar is "permission is in hand, and it is not a one-off export."
2. The two "under confirmation" entries: unknown is not a middle state, it is the polite way of writing fail, and a launch decision treats it as fail. And the rule is to stop at the first rung that cannot produce evidence. interpretable did not pass, so the four rows below it should carry no verdict at all.
3. traceable and actionable substitute a title's endorsement ("Kevin says," "Kevin thinks") for evidence. The qualifying evidence for actionable is a trial run of the target action, not anyone's opinion.
4. validated is scored pass on the strength of a demo running. A successful demo proves only as far as accessible, and validated has no shortcut, only the three-way reconciliation in 9.3. Chapter 9's lesson is exactly this. After the reconciliation, nearly half of this field cannot be used at face value.
5. No scoring date anywhere on the table. Fitness is a state, not a property, and a conclusion with no date has no shelf life and no retest to speak of.
6. The exists evidence, "the data platform says it is all synced over," is equally an organizational assurance, not field-level evidence. A data platform solves the hauling. Permission approvals, definition dictionaries, update latency and cross-system keys are equally unsolved (Chapter 9).
---
## Code Hooks
The companion repo provides (this repository's `repo/` directory):
- [`templates/data-fitness/`](https://github.com/hallieren/the-last-mile/tree/main/repo/templates/data-fitness/): this appendix's scorecard scaffolding and three-way reconciliation record script
---
# Template 10 ยท Pattern Decision Table and Anti-pattern List
> Companion chapter(s): Chapter 10. The three tools are in the order you use them. Learn the spectrum first (10.1), then take every step where AI intervenes through the criteria card by card (10.2/10.3), and last, give the design a physical with the anti-pattern list (10.4).
> License: Every template in this book may be modified freely and used in your work, no attribution needed.
---
## 10.1 The Pattern Spectrum at a Glance
The six patterns are ordered from dumbest to smartest by how clear the error shape is. The further down, the higher the capability ceiling and the harder it is to draw a shape for the error.
| Pattern | In One Line | The Most Typical Misuse |
|------|-----------|--------------|
| **Rules** | Deterministic logic. Same input, necessarily the same output | Hardcoding a tacit judgment that will drift (Chapter 6, failure mode 4) |
| **Structured extraction + rules** | The LLM does only the unstructured โ fixed schema conversion, and everything downstream is deterministic logic | No schema validation, and the extraction output goes into the store as fact |
| **RAG** | Retrieval over a trustworthy corpus + generation of an answer with citations | Nobody maintains the corpus itself, or it is used to answer exact-field questions (query the store when a query is what is called for) |
| **Single-step LLM** | One call completes one judgment or conversion, and the output goes into a human review step | Output with no format limit and no reason attached, leaving the reviewer nothing to review |
| **Agentic workflow** | Multi-step autonomous planning and execution, path decided at run time | Put into a step where the errors go outside and nobody backstops them |
| **Fine-tuning** | Changing model weights with annotated data | Using it on a problem a prompt could solve |
---
## 10.2 Pattern Decision Cards (Six Patterns ร Five Criteria, Full Version, Ready to Copy)
**Usage**. This section re-lays Chapter 10's decision table as cards (one card per pattern, the five criteria as labels, judgments identical to the chapter's decision table, so the full content fits). Run it once per AI intervention step, not once per project. Read the judgments card by card in the 10.1 spectrum order and hold them against the reality of your step. Where a judgment plainly does not hold (the step's tolerance is zero, say, while the pattern's ways of failing are open-ended), that pattern is out for this step.
### Card One ยท Rules
- **Error tolerance**. Ways it fails are enumerable and foreseeable. When it is wrong it stays wrong, but one fix ends it. The default choice for zero-tolerance steps.
- **Verifiability**. Fully reproducible, explainable rule by rule, exhaustively testable.
- **Data requirements**. Fields only need to be interpretable.
- **Latency and cost**. Milliseconds, near zero cost.
- **Maintainability by the receiving side**. The business-side IT or Group ops can read it, change it, and test it themselves, and the handoff cost is the lowest there is.
### Card Two ยท Structured Extraction + Rules
- **Error tolerance**. The error is fenced tight inside the extraction step. Schema validation intercepts format errors, and spot checks watch for semantic errors. Usable in near-zero-tolerance steps, provided the validation layer really exists.
- **Verifiability**. The extraction output can be validated against the schema and spot-checked by sample. The rules part is the same as the "Rules" card.
- **Data requirements**. The unstructured source has to be accessible. Keep a separate sample set for manual spot checks. The extraction output is derived data, and it is trustworthy only once it climbs to validated (Chapter 9).
- **Latency and cost**. One call per claim, batchable and cacheable.
- **Maintainability by the receiving side**. The rules go to the receiving side. The schema and the prompts need handoff training. The spot-check process needs an owner on the receiving side, and for as long as the spot-check owner is you, this row is your own permanent liability.
### Card Three ยท RAG
- **Error tolerance**. Missed retrieval and wrong retrieval both happen, and the ways it fails are open-ended. Only for steps where a person checks the citations.
- **Verifiability**. Verified by "checking the citations." Whether a citation is real can be checked, but "should have been retrieved and was not" is hard to verify. A rotten corpus rots everything.
- **Data requirements**. It needs a corpus that is genuinely trustworthy and has someone responsible for maintaining it. The corpus's fitness decides everything.
- **Latency and cost**. Two hops, retrieval plus generation. Maintaining the corpus is a long-term hidden cost.
- **Maintainability by the receiving side**. What the receiving side maintains is really the corpus. The corpus decays slowly and out of sight, so it needs a named maintenance owner and a cadence.
### Card Four ยท Single-Step LLM
- **Error tolerance**. Every call can be wrong and the ways it fails are open-ended. Only for steps where a person reviews row by row.
- **Verifiability**. The output cannot be verified exhaustively. Pinning the output format + requiring a reason brings verification down to "a person checks the reason."
- **Data requirements**. No training data needed. But no eval set means no threshold, which means no acceptance.
- **Latency and cost**. One hop, latency and spend both controllable and budgetable.
- **Maintainability by the receiving side**. Prompt drift (model updates, quiet edits to the wording) going unnoticed is the biggest hazard. Prompts go into version control.
### Card Five ยท Agentic Workflow
- **Error tolerance**. Errors compound and amplify across steps (95% on one step leaves 77% over five). If any step sends an action outside, the whole chain is reviewed as zero tolerance.
- **Verifiability**. Many intermediate states and the path varies. Failures are hard to locate and hard to reproduce, and the same input may take a different path.
- **Data requirements**. The data each step consumes has to pass fitness on its own, each step's output needs its own eval, and the cost multiplies by the number of steps.
- **Latency and cost**. Multiplied by hops, and the number of steps is settled only at run time, so latency and spend are both unpredictable (Chapter 16 does this arithmetic).
- **Maintainability by the receiving side**. The hardest to hand off. Debugging requires understanding every agent's intent and interactions, and "only the builder can fix it" means, inside a company, that you are locked in, because the maintainer is often your own team.
### Card Six ยท Fine-Tuning
- **Error tolerance**. Errors set into the weights. A newly found way of failing means retraining, and it cannot be patched.
- **Verifiability**. Behavior changes globally and cannot be explained locally. Every retrain requires a full regression eval.
- **Data requirements**. Thousands of high-quality annotations and up. Most sites cannot accumulate them, and once accumulated they drift out of date.
- **Latency and cost**. Training cost comes up front and recurs with every update. Inference may be cheaper, but ask first whether that training bill is worth it.
- **Maintainability by the receiving side**. The receiving side can hardly take it over at all. A retraining pipeline, an annotation team, a regression eval, and not one of them is optional.
### Three Rules for Using It (Same as the Chapter)
1. **Pick per step, not per project.** Inside one system, tolerance differs from step to step, and the pattern follows the step.
2. **Read top down and stop at the first pattern that passes all five criteria.** "Works" = all five judgments pass, not the most impressive result.
3. **An upgrade has exactly one legitimate channel, eval data showing that the dumber pattern cannot clear the threshold (Chapter 11).** The criteria cards hold in reverse too. If the evidence says upgrade, "never use agents" is just as much a dogma.
---
## 10.3 Filling-In Process Checklist
- [ ] **Get the sponsor's endorsement of "ruling by the table" as a procedure first**, not his endorsement of one design. With the procedure endorsed, every step from then on goes row by row through the same table, all proposals are treated alike, including the ones you make yourself, and what is argued at the review is a proposal against a table, not you against the proposer.
- [ ] List the AI intervention points first. Which steps in this system does AI intervene in? One line per step (an intervention point = one decision or one conversion, not one feature module).
- [ ] For every step, write down three things before you look at the cards:
- The step's **error tolerance** (how many times may it be wrong? who sees it when it is, how fast, and can it be recalled?), carrying on from Risk, one of Chapter 7's five questions.
- The **real position** on the fitness ladder of the data the step consumes (cite the conclusion of the Template 9.2 scorecard directly).
- Which layer of the **decision rights boundary** the step sits in (sense / advise / act / decide, Template 8.2). For steps at the act layer and above, tolerance is preset to zero tolerance.
- [ ] Work down from the "Rules" card, taking each pattern through the five criteria, and stop at the first pattern that passes all five.
- [ ] When somebody proposes jumping to a pattern further down, **have the proposer fill in the table himself** (what Anchor & Helm did in Chapter 10). A judgment can only be rebutted. It cannot be overruled by rank.
- [ ] A rejected pattern proposal goes into the scope decision log (Template 8), and the revival condition is written as an event. "The eval shows the current pattern cannot clear the threshold" is the standard form.
- [ ] Mark the selection conclusion with its date and its premises. Fill it in again when a premise changes (data fitness improves, the eval goes live, the decision rights boundary moves up). Pattern selection, like fitness, is a state, not a property.
---
## 10.4 Anti-pattern List
Matched to Chapter 10's failure modes and extended. Each entry gives a warning signal (the actual words you can hear at a review) and a corrective action.
| # | Anti-pattern | Warning Signal (Actual Words) | Structural Root Cause | Corrective Action |
|---|--------|------------------|-----------|----------|
| 1 | **Resume-driven architecture** | "This project is a good chance to get fluent in the agent framework" | How new the stack is calibrates an engineer's market value. The project account settles in three months, the resume account settles the same day. The department's annual report also needs a line that says "we shipped agents," and that pressure runs top down, onto the line that runs your schedule and your review, which is harder to push back on than a resume | Proposals go through the decision table, and the criteria recognize neither learning value nor the wording of a report. Schedule the learning separately, off the critical path |
| 2 | **The demo's pattern straight into production** | "The demo already runs, do not tear it down" | A demo selects on how impressive it is, production on whether it can be backstopped. "It already runs" manufactures the illusion of a sunk cost | Go through the table again before production. Demo assets can be reused piece by piece (prompts, schema). The architecture is not inherited |
| 3 | **Uncertainty into a step that cannot be wrong** | "Just let the model work out this amount while it is in there" | Probabilistic behavior is invisible on an architecture diagram. Tolerance is a property of the step, and if it is not drawn nobody aligns on it | Mark the tolerance and the backstop at every AI intervention point on the architecture diagram. Before a generated number enters a deterministic step there has to be a layer of validation or a person in between |
| 4 | **Reverse dogma** | "We had an incident last time, so no LLM goes to production" | An incident sets into a ban, and a ban is cheaper than judging case by case | Replace the ban with the decision table. If eval evidence says upgrade, upgrade, one layer at a time |
| 5 | **Generalizing into a platform, agents ahead of time** | "Build it general-purpose and it will run any process later" | The legitimate raw material for abstraction is repetition (Chapter 8). The generality of LLMs made a platform demo cheap, and production cost did not move | Go back to the thin slice. Platform proposals go into the scope decision log, with the revival condition "after the second slice lands" |
| 6 | **Fine-tuning a prompt problem** | "The results are unstable, a bit of fine-tuning will fix it" | Fine-tuning sounds "more engineering" than editing a prompt, and it justifies a budget | Exhaust prompts and schema constraints first, comparing with the eval. Fine-tuning has to answer "where the annotated data comes from, and who retrains when it drifts" first |
| 7 | **Using RAG as a database** | "Ask it how much this claim paid out and it can answer" | The generative interface hides the fact that an exact field calls for an exact query. A retrieval hit is not the same as a correct fact | Exact fields go through queries and rules. RAG answers only questions whose answer lives in a document, and always with citations |
| 8 | **Chained LLMs in place of deterministic logic that works** | "Have an agent judge whether these two records are the same claim" | Doing deterministic work with a probabilistic component feels "smarter." The right answer to a broken linkage is repairing the join key (Chapter 9) | What can be joined is not generated, what can be looked up is not inferred. Leave the LLM only the conversions deterministic logic genuinely cannot do |
| 9 | **Talking about upgrades with no eval** | "Users say it is not smart enough, let us go multi-step" | "Feelings" are cheap and cannot be rebutted. With no eval in place, the loudest feeling wins | Freeze the upgrade question and build the eval first (Chapter 11). An upgrade proposal has to cite which class of cases the current pattern fails the threshold on |
| 10 | **A backstop in name only** | "There is human review anyway" | Review is treated as a disclaimer rather than a design object. Review with no time, no ability, and no authority is review in name only | Accept it against Chapter 12's three oversight questions. Is there time to look? The ability to judge? The authority to stop it? The reason column and the daily volume are design parameters, not footnotes |
**How to use it**. Before a design review, send this table to the whole team and have each person tick anonymously "which of these we are committing right now." Any entry with two or more votes goes on the review agenda. The value of an anti-pattern is not in naming it afterwards. It is in making it unsayable beforehand.
---
## Code Hooks
The companion repo provides (this repository's `repo/` directory):
- [`templates/pattern-selection/decision-table/`](https://github.com/hallieren/the-last-mile/tree/main/repo/templates/pattern-selection/decision-table/): the decision table in electronic form and the script that generates the intervention point list
- [`templates/pattern-selection/schema-check/`](https://github.com/hallieren/the-last-mile/tree/main/repo/templates/pattern-selection/schema-check/): schema validation scaffolding for extraction-type steps (field type / enum value / required field validation + a rejection log)
- [`templates/pattern-selection/spot-check/`](https://github.com/hallieren/the-last-mile/tree/main/repo/templates/pattern-selection/spot-check/): the spot-check log (sampling, ruling, semantic error classification)
---
# Template 11 ยท Eval Spec, Golden Case Collection List, Annotation Disagreement Log
> Companion chapter(s): Chapter 11. The three tools are in order of use. Set up the spec skeleton first (11.1), then collect cases by the list to fill it (11.2), and record every disagreement during annotation in 11.3, rule on it, and settle it into a rule.
> License: Every template in this book may be modified freely and used in your work, no attribution needed.
---
## 11.1 Eval Spec Template (Five Parts, Fill Section by Section)
One eval spec evaluates one task. Fill the document header with: project name / version and date / co-builders (actual users by name) / trigger condition for the next retest.
### 11.1.1 Task Definition
| Item | Fill In |
|----|------|
| One-sentence task (workflow claim level) | e.g., for every auto exception claim entering the queue, give the next action, the owner, and a risk flag, for the reviewer to follow or override with a stated reason |
| Input | Data sources and fields (cite the source of truth decision log, Template 9.4) |
| Output | Field list (suggestion + reason + Human Call column are mandatory) |
| Users | Named down to the team |
| Red line (what this task does not do) | e.g., no suggested payout amounts (cite the scope decision log, Template 8) |
### 11.1.2 Golden Cases List
Size 50โ200 cases. Register each one:
| Case ID | Source Type (typical / edge / historical incident / annotation disagreement) | Input Summary | Correct Answer | Basis for Answer (reconciliation evidence / confirmed by the person involved) | Annotator | Date |
|---------|------|----------|----------|----------|--------|------|
**Rules**:
- [ ] Every fact a "correct answer" cites must stand at validated or above on the data fitness ladder. Candidate claims that do not reconcile go through the three-way reconciliation first (Template 9.3), and those that cannot be ruled on are dropped.
- [ ] The annotation basis cites an entry in the annotation guide (e.g., "rule of thumb 4: prior claim linkage"), and the guide carries a dated version.
- [ ] Check the makeup of the set against the mix lines in 11.2.
### 11.1.3 Error Taxonomy
| Error Category | Definition | Business Consequence | Error Severity Vocabulary (unsafe / concern / useless) | Tolerance |
|----------|------|----------|-----------------------------------|--------|
| e.g., missed risk | A high-risk claim judged routine | Grows into a major case, regulatory complaint | unsafe | Zero tolerance |
**Rules**: the unsafe class must come from asking the actual user "which kind of error can you not accept even once." Each category states a "consequence," not a "technical cause." The taxonomy is layered by cost, not by bug type.
### 11.1.4 Per-Category Thresholds and Acceptance Lines
| Error Category | Threshold (on golden cases) | Acceptance Line | Retest Cadence |
|----------|------------------------|--------|----------|
| e.g., missed risk | 0 cases | 0 cases | Rerun on every update to the set |
| e.g., material error | โค10% | โค10% and the trend does not rise over three consecutive replays | Monthly |
**Rules**: **no overall score**. Any "overall accuracy X%" is the shortest path to burying the unsafe class. Thresholds and launch release conditions cite each other (Chapter 4).
### 11.1.5 Human Review Path
| Item | Fill In |
|----|------|
| Trigger condition | e.g., hits a risk rule with low confidence / suspected unsafe-class output |
| Destination and owner | e.g., the queue's "pending human" column, the review team lead |
| Time limit | Review must happen within how long |
| Flow-back mechanism | Review conclusions are added to the golden cases monthly; suspected unsafe-class errors are reviewed weekly |
| Flow-back owner | Who is responsible for adding review conclusions to the golden cases on cadence and bumping the guide version with them, named to a person |
| Reminder mechanism | Hang it on a meeting that owner already holds (e.g., the monthly QA review), no new reminder, do not rely on anyone remembering |
---
## 11.2 Golden Case Collection List
### Four Source Types and Where to Find Them
| Type | Definition | Where to Find | Mix Line |
|------|------|----------|--------|
| **Typical cases** | High frequency, answer undisputed | Random sample from the source of truth | **Under half** |
| **Edge cases** | Cases right at a rule's threshold | Aged claims, claims at the amount threshold, claims where rules conflict | โฅ30% |
| **Historical incident cases** | Cases that once went wrong or were misjudged | Complaint log, retrospective notes, internal audit reports, regulatory inspections, veterans' "the one I got wrong" stories | โฅ10% |
| **Annotation disagreement cases** | Cases the two annotators marked differently (entered after ruling in 11.3) | The annotation session itself | Added continuously |
Edge cases and historical incident cases are the soul of the eval. That is where the system's value and its risk both live.
### Collection Steps
- [ ] 1. Finish the 11.1.1 task definition first. Without a task, there is no way to judge whether a case is representative.
- [ ] 2. Draw candidates from the source of truth (not the easiest database to reach), preparing 1.5 times the target count.
- [ ] 3. **Clear the data check case by case**: where the three sources agree, in; where they disagree, rule by three-way reconciliation; where no ruling is possible, drop (Chapter 9's baseline: skip this step and nearly half the answers may be wrong).
- [ ] 4. AI pre-fills fields and drafts answers (AI does the grunt work); the actual user rules on the final answer (people make the calls).
- [ ] 5. Double independent annotation, disagreements go through 11.3.
- [ ] 6. Enter with version and date; bump the annotation guide version in step.
- [ ] 7. Set the rolling update cadence: cases flowing back from review are added monthly; drift in business rules (e.g., a change to a list-type rule) triggers an immediate update.
---
## 11.3 Annotation Disagreement Log
An annotation disagreement is not noise. It is a business rule that was never aligned, surfacing. One row per disagreement, and the loop closes only when it is "settled into a rule":
| Disagreement ID | Case ID | Annotator A Mark + Reason | Annotator B Mark + Reason | Root Cause (different standards / different information / different reading of the task) | Type (within-department / cross-department / compliance-related) | Ruled By | Ruling | Settled into a Rule (which entry of the annotation guide) | Date |
|---------|---------|----------------|----------------|--------------------------------|--------------------------------|--------|----------|----------------------------------|------|
| e.g., D-01 | 37 | High risk: phone number appearing for the third time in six months | Routine: shop-filed claims are common, count only plate + filer | Different standards, whether "prior claim linkage" counts the phone number | Within-department | Kevin Doyle (Director of Claims Operations) | Phone number counts toward the signal; on its own it only triggers human review | New exception scenario added to guide section 4.2 | / |
**Rules**:
- [ ] Trace the root cause before talking about a ruling: **different standards** need a business decision; **different information** means fill in the information and annotate again; **different reading** means go back and fix the 11.1.1 task definition.
- [ ] Route disagreements one of three ways by type, and never let one hang. **Within-department disagreement**, find the owner of this workflow, rule on the spot, and write the ruling into the annotation guide. **Cross-department disagreement**, pause annotation, write both definitions and their business consequences on half a page, and send it to the two sides' common superior or put it on the sponsor weekly as an agenda item. **Compliance-related disagreement**, cross-department or not, goes to the risk owner (Chapter 11).
- [ ] The person ruling must hold business decision rights (the workflow owner or above), not a vote or a compromise between the two annotators.
- [ ] Every ruling must be settled into one executable line of text in the annotation guide, or the same disagreement comes back on another claim.
- [ ] Disagreement cases themselves go into the golden cases (the fourth type in 11.2). They are the cases with the highest information density in the set.
- [ ] Report the disagreement rate and the number ruled on. The disagreement rate is not poor annotation quality, it is the number of missing pages in the requirements document. Every ruling is one more business rule the organization has aligned.
---
## 11.4 Counterexample: A Tidy-Looking Wrong Answer
An excerpt of the five parts. Every section has content, and it looks like a spec:
```
Task definition: Comprehensively evaluate the quality of AI output in the claims scenario.
Golden cases: 50, randomly exported from the test database; answers filled in by Victor Reyes (security reviewer) against the system fields.
Error taxonomy: one category, "output error"; tolerance: as few as possible.
Threshold: overall accuracy โฅ 90% passes acceptance.
Human review: handled by human intervention when necessary.
```
Line by line:
1. The task definition is not at workflow claim level. "Comprehensively evaluate output quality" has no action, no user, and no red line, so every section after it has no target.
2. The golden cases break three rules: drawn from the easiest database to reach, not the source of truth; answers taken from the system fields, which by Chapter 9's baseline means nearly half may be wrong without reconciliation, and the rule requires standing at validated or above; Victor annotating alone, so with no double annotation there are no 11.3 disagreements to record, and disagreements are where business rules surface, and he is not the actual user.
3. The taxonomy has one category and no business consequence. The taxonomy is layered by cost; the unsafe class must come from asking Linda "which kind of error can you not accept even once."
4. "Overall accuracy 90%" is exactly the overall score 11.1.4 forbids: 2 missed risks out of 50 still scores 96%, and unsafe is buried in the average.
5. "Human intervention when necessary" is missing four things: no trigger condition, no owner, no time limit, no flow-back. Review conclusions cannot get back to the golden cases, so this eval is one-shot.
---
## Code Hooks
The companion repo provides (this repository's `repo/` directory):
- [`templates/eval-spec/`](https://github.com/hallieren/the-last-mile/tree/main/repo/templates/eval-spec/): this appendix's eval run scaffolding and per-category threshold report script
---
# Template 12 ยท Trust Constraint Matrix and Security Review Packet
> Companion chapter(s): Chapter 12. The two tools are used in time order. The matrix (12.1) starts being filled after the week 2 constraint interview and updates all the way through the design, a living document of the design period. The packet (12.2) is assembled from the matrix and the design documents in the week before the review, a snapshot of the matrix at review time. Chapter 18's launch gate reuses the packet (updated to a launch version).
> License: Every template in this book may be modified freely and used in your work, no attribution needed.
---
## 12.1 Trust Constraint Matrix
### 12.1.0 Constraint Interview (Do This Before Filling the Matrix)
Who, the risk owner (the name in that cell of the stakeholder map, Template 5). When, week 2 of the opening, before the design takes shape. Everything can still be changed at that point. Go with questions, not with a design. Three must-asks:
- [ ] **What has gone wrong before?** Real incidents this company has had with data or systems. The answer marks out his actual minefield.
- [ ] **What are you afraid of?** How responsibility lands on him when something goes wrong. The answer fills the "What They Fear / What They Win" columns (Template 5), and it also explains every reaction he has later.
- [ ] **What evidence do you want to see at the review?** The review checklist in his head. The answer generates the table of contents of the 12.2 packet directly.
Interview output, a first draft of the matrix's first column plus an open questions list. Rows you cannot fill stay as open questions, not deleted, not hidden.
### 12.1.1 Blank Template (Six Constraints ร Three Columns)
| Constraint | The Specific Requirement for This Project | How the Design Satisfies It (Write the Cost on the Last Line) | Who Signs Off |
|------|------------------|------------------------------|-----------|
| Privacy | | | |
| Security | | | |
| Audit | | | |
| Human oversight | | | |
| Fairness | | | |
| Maintainability | | | |
**Rules for filling the three columns**:
- **Column one**. Specific to this project, taken from the constraint interview. "Complies with the company data standard" does not count as filled in. "Customer-identifiable information may not leave the claims domain" does.
- **Column two**. A mechanism, not a promise. "We will be careful" does not count. Being able to write "at which layer, by what means, and how an error gets caught" does. **One line at the end of every cell, "Cost,"** who spends how much extra time under this way of satisfying it, and what information or capability is given up. Costs that never reach the table let constraints accumulate in one direction only (Chapter 12, failure mode 4).
- **Column three**. Real names. The day a name goes on, the reviewer turns from judge into co-author. Blank = this row has no owner yet, handle it as an open question.
**Row-by-row filling guide** (the mechanism material is long, moved into numbered notes below the table):
| Constraint | Ask Yourself for Column One | Common Qualifying Mechanisms | Anti-formality Check |
|------|-----------|--------------|----------------|
| **Privacy** | Whose information, and how far may it flow before the boundary stops it? | Note 1 | Spot-check whether de-identified data can re-identify a person; whether the right side of the line really has no plaintext |
| **Security** | Who can access what, and change what? What attack surface is new? | Note 2 | Whether the new system's permissions exceed what the same person has in the source system |
| **Audit** | When something goes wrong, can you reconstruct who did what, when, and on what basis? | Note 3 | Take one historical suggestion and actually rehearse a trace |
| **Human oversight** | At which step is a person present, and in what way is that presence real? | Note 4 | **Three questions. Is there time to look? The ability to judge? The authority to stop it?** Daily volume ร review time per row, do the division in front of the reviewer |
| **Fairness** | Will the system systematically treat one class of subject worse? | Note 5 | It counts as a mechanism only if you can say "what data would reveal unfair treatment" |
| **Maintainability** | After you are reassigned, leave, or the system enters its third year with nobody watching, who can safely change it? | Note 6 | Have the future maintainer (the co-build counterpart of Chapter 15) try one change |
**Common qualifying mechanisms (material for column two, numbered by constraint)**:
1. **Privacy**. De-identification and placeholder substitution at the extraction layer; the de-identification boundary drawn on the data flow diagram; the minimum-fields principle.
2. **Security**. Permissions inherit existing system roles (no separate account system); externally writable content (emails, attachments) is handled as data, never as instructions.
3. **Audit**. A trail the whole way, including input snapshot, rule and model version, suggestion, reason, Human Call, timestamp; derived views purely read-only, AI output not written back to the source (Chapter 9's trail form).
4. **Human oversight**. The Human Call column carries a real name; every suggestion carries its reason; risk ranking with a daily volume cap.
5. **Fairness**. Ranking and suggestion logic use no identity attribute; the override distribution is spot-checked by customer segment (Chapter 18 monitoring).
6. **Maintainability**. Rules, de-identification config and thresholds become configuration items the claims-ops IT engineers can change; changes go through the existing release process; a model upgrade runs the regression eval (Chapter 11).
### 12.1.2 Anchor & Helm Example (Excerpt from the Week 8 Review Version)
| Constraint | The Specific Requirement for This Project | How the Design Satisfies It | Who Signs Off |
|------|------------------|--------------|-----------|
| Privacy | Customer-identifiable information may not enter the queue or any model call | Name / ID number / plate number / phone are replaced with placeholders at the extraction layer, no plaintext to the right of the line. Cost, same-name customers across claims are checked by hand, about 1 extra minute per claim in review | Victor Reyes |
| Security | The queue may not enlarge anyone's existing permissions | Permissions inherit core system roles, no separate accounts; email content is handled as data, not parsed as instructions. Cost, new members depend on the core system process for access, up to 3 days | Victor Reyes + the IT data team |
| Audit | Every suggestion is fully traceable | The six end-to-end decision trail fields; the merged view is purely derived, purely read-only, no write-back. Cost, log storage and the hours for the quarterly spot check | Victor Reyes |
| Human oversight | No action without human confirmation | Risk ranking with a daily volume cap (time); suggestions carry their reason (ability); the Human Call column names a person (authority). Cost, a cap on daily throughput, with the backlog going through the old manual process | Kevin Doyle |
| Fairness | Ranking may not systematically treat a customer segment worse | Priority logic carries no identity attribute; the override distribution is spot-checked by customer segment each quarter. Cost, half a day per quarter for the spot-check report | The head of compliance |
| Maintainability | After handoff, claims-ops IT can make changes on its own | Rules and de-identification config are configuration items claims-ops IT can change; a model upgrade passes the regression eval. Cost, about one extra week of development to make it configurable | Kevin Doyle + claims-ops IT |
---
## 12.2 Security Review Packet Template
**Organizing principle. Organize by the reviewer's point of view. Whatever answer he wants, write it in advance.** Open every section with one line naming which of the reviewer's questions this section answers. Send it a week ahead. Confirm section by section at the meeting, and debate only the open questions (ยง8).
### ยง0 Cover
- The system in one sentence (precise to workflow claim level) + the review scope (which components are under review, which are explicitly not)
- **Byline, two lines. Solution design (the deliverer) / constraint design (the risk owner and the other signers)**, and everyone named in column three appears here. The review conclusion memo goes out from these two jointly, recipients see two senders, and it is not downgraded to one approver's signature
- **Rule for the maintainability row's signer**. The signer may not come from this project's delivery team. Either it is signed by the business line's IT or ops, or it is left blank and handled as an open question. Self-signing does not count, that is signing yourself an ops commitment with no end date
- Nature of the review, a confirmation meeting (every row of the matrix is signed) / a ruling meeting (open questions exist, list them)
### ยง1 Answers-First Page (One Page)
The reviewer's top questions ร a one-line answer ร the section with the detail. The third question of the constraint interview decides what goes on this page. For example:
| The Reviewer's Question | One-Line Answer | See |
|--------------|----------|------|
| Does customer data leave the boundary? | No. Identifiable information is de-identified at the extraction layer, boundary shown as the red line on the data flow diagram | ยง2/ยง5 |
| Who audits the model's output? | Every suggestion carries a trail the whole way, schema and query permissions in the trail notes | ยง4 |
| Who is responsible when it is wrong? | No action without human confirmation; the Human Call carries a real name | ยง3/ยง6 |
### ยง2 Data Flow Diagram (Figure Slot)
[Figure slot: source systems โ extraction and de-identification layer โ merged view โ queue โ Human Call]. Must be labeled. The **de-identification boundary** (one solid line, plaintext on the left, none on the right); the **read-only / writable** attribute of every arrow; the **external call points** (model calls, outbound notices); and where data does not flow (draw "does not leave the boundary" explicitly).
### ยง3 Permission Model
| Role | What They Can Do in the Queue | Inherits From (Source System Role) | Exceptions and Why |
|------|----------------|----------------------|------------|
One line of principle. The new system enlarges nobody's existing permissions, and every exception must be listed item by item and signed.
### ยง4 Trail Notes
- Audit log schema (sample in Code Hooks at the end), suggestion ID / input snapshot / rule and model version / suggestion content / reason / Human Call + the decider / timestamp
- One line of declaration. AI output is written back to no source field and exists in trail form (Chapter 9)
- Retention period, query permissions, spot-check cadence
### ยง5 De-identification Boundary
| Field | Treatment (De-identify / Substitute / Keep) | Where It Happens | Reason for Keeping (If Kept) |
|------|---------------------------|----------|---------------------|
Attach the verification method. Take N de-identified records, run a re-identification test, and record the result.
### ยง6 Human Oversight Design
One line each for question ร mechanism ร evidence (time, the arithmetic of daily volume divided by the review time budget; ability, a sample of the reason column; authority, a demonstration of the stop path).
### ยง7 Rollback Path
Trigger conditions (which signal sends it back) / the rollback action and how long it takes / which old process runs after a rollback / the owner / the rehearsal record (Chapter 18 requires one real rehearsal before launch).
### ยง8 Open Questions
Unsettled item ร suggested handling ร what this meeting has to rule on. **Open questions are not hidden.** It is the blank row you hid that blows up at the review, and an open question you list yourself is a deposit into credibility (Chapter 5).
### ยง9 Attachments
The full trust constraint matrix (12.1) + the signature page.
**Assembly discipline**. The packet writes nothing new. Every section should be a rearrangement of material already in the matrix and the design documents. A section you cannot write while assembling means the design has a hole. Fix it at the design layer, do not paper over it at the document layer.
---
## Code Hooks
The companion repo provides (this repository's `repo/` directory):
- [`templates/trust-constraints/`](https://github.com/hallieren/the-last-mile/tree/main/repo/templates/trust-constraints/): audit log schema sample and redaction rule config sample
---
# Template 14 ยท Stage Gate Checklist and Pilot Kill Criteria
> Companion chapter(s): Chapter 14. The three tools are in order of use. Use 14.1 first to judge which game you are in and which rules you have to keep, rule item by item with 14.2 at the escalation decision meeting, and in the same meeting where the resolution passes, use 14.3 to write the stop standard hard.
> License: Every template in this book may be modified freely and used in your work, no attribution needed.
---
## 14.1 Stage Rules Table (Printable)
| | Prototype | Pilot | Production |
|---|---|---|---|
| **Optimizing for** | Learning speed, changing fast beats not erring | Evidence quality, measurable and attributable under real conditions | Operating reliability, not erring beats changing fast |
| **Data** | De-identified samples / synthetic claims, read-only | Real data, controlled scope (one team, one claim type) | Full real data, decision trail, closed-loop write-back |
| **Users** | A few volunteers, using it and cursing it to your face | A named real user group that depends on it in daily work | All target users, the most unwilling group included |
| **Change discipline** | Change anytime, ship the same day, redo it if wrong | Changes released in batches, announced ahead, rollback available | Changes follow the process, clear the eval regression before shipping |
| **Typical way to die** | Over-polished into a deluxe demo | The zombie pilot, no definition of graduation or death | Demo code shipped carrying demo discipline |
**Rules**:
- Judge the stage by the "change discipline" row, not by what the system calls itself. Change anytime = prototype, batched with rollback = pilot, regression first = production, whichever environment it runs in.
- Mapped onto the outcome ladder (the five-rung outcome ladder from L0 demo to L4 business-side self-sufficiency, defined in Chapter 1), prototype sits between L0 and L1, pilot = L1, production = L2.
- The danger is not inside a cell. It is in changing column without changing the rules. Changing stage = changing the whole column of discipline, checked row by row.
---
## 14.2 Three Escalation Gates Checklist (prototype โ pilot)
Standing agenda for the escalation decision meeting. Go through the three gates item by item โ rule โ write 14.3 in the same meeting. Fill the evidence into each column before the meeting, and do only the ruling in it.
### Gate One: Eval Over the Line
- [ ] The five-part eval spec is complete (Template 11), and the per-category threshold table is signed off by both sides
- [ ] Most recent golden case replay, unsafe class **0 cases**
- [ ] Every concern-class item inside its threshold
- [ ] The useless class inside its threshold, and the trend has not risen across three consecutive replays
- [ ] The version replayed = the version that will enter the pilot (no "we tested the previous build")
### Gate Two: Willingness to Use It Daily
- [ ] The target user group used it on its own for โฅ2 consecutive weeks (no reminders), with usage records as proof
- [ ] At least one behavioral signal that taking it away would hurt (someone asks when the system is missing, someone complains)
- [ ] The evidence comes from behavior records, not from a satisfaction survey
- [ ] At least one user can say "which next action of mine it changed"
### Gate Three: Owner in Place
- [ ] The pilot's operating owner (the business line's head) claims it by name and is present in person (not assigned in absentia)
- [ ] The owner's three items go into the minutes, schedule protection for the users' input / ownership of the operating discipline / first signatory on "stop or not"
- [ ] Consistent with the charter's owner field, or the charter is updated on the spot
### Ruling Rules
- All three gates **clear** โ set the pilot launch day (within a week is a good default), and write 14.3 in the same meeting.
- Any one not cleared โ write down which evidence is missing, the action to fill it and who owns it, and set a date to reconsider.
- **No "basically passed, start it first"**, and no using "more polish / more watching" in place of a ruling. Ask first, which gate is the polishing for?
### Reusing It for pilot โ production
Same structure, three gates with different evidence sources. Eval over the line becomes per-category targets met on live data across the whole pilot (kill criteria zero triggers, or triggered and closed out). Willingness to use it daily becomes willingness to roll out beyond the controlled group with balancing metrics not worsening (metric tree, Template 18). Owner becomes the production operating owner and confirmed budget ownership (the budget line lands in the business side's annual budget, Chapter 16's four ledgers, Chapter 22's handoff).
---
## 14.3 Pilot Kill Criteria Template
**Timing rule**. Written in the same meeting as the escalation resolution, signed the same day. Kill criteria are written during the excited period. They cannot be written during the disappointed one. After entering the pilot, adding items is allowed, loosening existing ones is not.
**Header**:
| Item | Fill In |
|---|---|
| Pilot name / start, end and duration | e.g. the auto exceptions queue pilot, 8 weeks |
| Owner (first signatory) | By name |
| Signatures | The delivery side's lead + the business owner, filed with the sponsor, dated |
| Relationship to the charter's resource reassessment conditions | This sheet is a pilot-level operating trigger, and a trigger is executed by the action in its own row. Project-level shutdown is still ruled by the charter's resource reassessment conditions (Template 4, element 7), and the power to shut down runs one way, to the sponsor / the project approval committee. Neither replaces the other |
| Resource renewal point | [The approval conditions for the next round of people and compute budget, written as the three gates and the graduation criteria, no dates] |
| Reassessment date | [Set together with the pilot launch day; if it comes due with no escalation resolution, settle on the graduation criteria and convert to formal project approval or shut down] |
**Item table** (every row must carry a hard number, no "assess as the situation warrants"):
| # | Trigger (the data appears, it stops) | Data Source | Action on Trigger | Recovery Condition | Recheck Cadence |
|---|---|---|---|---|---|
| 1 (Anchor & Helm example, unsafe) | Unsafe-class errors โฅ2 in a single week | Weekly retrospective record of suspected unsafe claims | Pilot pauses for rework, that whole class of suggestion is degraded to human review, root cause retrospective | Golden case replay passes after the fix, owner signs the restart | Every Friday |
| 2 (Anchor & Helm example, North Star) | Weekly median of first-touch handling time above the eight-week pre-pilot baseline, two consecutive weeks | Weekly points on the run chart (Template 18) | Pause the expansion, check the data source (fitness retest) before the system | Two consecutive weeks back below the baseline | Every Friday |
| 3 (Anchor & Helm example, cost) | System cost per claim above half the value of the labor hours that claim saves, two consecutive weeks | The cost page of the four ledgers (Template 16.2), not waiting for the month-end bill | Pause the expansion, cost review | Cost per claim back under half the value line and held two weeks | Every Friday |
| 4 (candidate, rewrite as needed) | The target user group's weekly usage rate below the agreed line, two consecutive weeks | Usage logs | Pause the expansion, go back to users to locate the reason | The reason is closed out and usage recovers for a week | Weekly |
| 5 (required, business-side input) | The business side's committed data access or people go two consecutive weeks unmet | Scheduling records | Triggers the charter's resource reassessment conditions (Template 4, element 7), handled at the matching tier | Input resumes | Weekly |
**Rules**:
- [ ] A trigger means executing the action in its row, with no meeting reopening the standard itself. The time to argue the standard has passed.
- [ ] A trigger is not a failure, it is the defense holding. The first sentence of the outward notice says the mechanism held, then the root cause investigation.
- [ ] Record every recheck (triggered / not triggered), and archive the whole sheet when the pilot ends. This sheet gets checked line by line again in the incident retrospective after launch (a fixed action of Chapter 18's AI incident runbook).
- [ ] A case that has triggered goes into the golden cases permanently (Template 11.2, third category, historical incident cases).
---
## 14.4 Counterexample: A Tidy-Looking Wrong Answer
An excerpt from the minutes of an escalation decision meeting. All three gates "cleared," and the pilot has a launch day:
```
Gate one: eval basically over the line. Unsafe class 1 case, an edge case, the team judged it acceptable;
the replay used the build from two weeks back (this week's build did not change much).
Gate two: satisfaction survey 4.6/5, users generally said "very helpful."
Gate three: owner, Kevin Doyle's team will arrange a dedicated person later.
Ruling: basically meets the bar, start the pilot first, fill the open items while running.
Kill criteria: assess flexibly as the pilot runs.
```
Item by item:
1. The gate one unsafe threshold is 0 cases. There is no "edge case" discount, and the definition of unsafe is that not one is acceptable (Template 11.1.3). "The team judged it acceptable" steals the ruling power back from the threshold table and hands it to people in the excited period.
2. The replayed version โ the version that will enter the pilot, landing exactly on gate one's prohibition, "we tested the previous build."
3. The satisfaction survey is the evidence gate two names and excludes. What it wants is behavior records, two consecutive weeks of use with no reminders, someone asking when the system is missing.
4. "The team will arrange a dedicated person later" = owner not in place. Gate three requires a claim by name with the person present. There is no name in the minutes, so during the pilot "stop or not" has no first signatory.
5. "Basically passed, start it first" and "assess flexibly" are the exact phrases 14.2 and 14.3 forbid. Kill criteria must be written with hard numbers in the same meeting as the escalation resolution, and they can only be written during the excited period, because they cannot be written during the disappointed one.
---
## Code Hooks
The companion repo provides (this repository's `repo/` directory):
- [`templates/stage-gates/`](https://github.com/hallieren/the-last-mile/tree/main/repo/templates/stage-gates/): the eval-over-the-line automatic check script (reads the eval replay report, compares thresholds category by category, and outputs the checkbox state for 14.2's gate one directly)
- [`templates/stage-gates/`](https://github.com/hallieren/the-last-mile/tree/main/repo/templates/stage-gates/): the kill criteria weekly report template
---
# Template 15 ยท Co-build Agreement and Knowledge Transfer Cadence
> Companion chapter(s): Chapter 15. The two tools are used in time order. The agreement (15.1) is signed after the pilot is called and before it starts, one page, signed by both sides, the "constitution" of the co-build period. The cadence sheet (15.2) is the vehicle for clause four, running from pilot week 2 to the handoff (Chapter 22), one beat every two weeks (weekly at Anchor & Helm after the team expanded).
> License: Every template in this book may be modified freely and used in your work, no attribution needed.
---
## 15.1 Co-build Agreement Template
### 15.1.0 Agreement Header
| Item | Fill In |
|----|------|
| Project | |
| Receiving-side owner (real name) | |
| Receiving-side engineers (real names, at least 1) | |
| Digital Center (real names) | |
| Effective date / expected handoff date | |
| Signatures (everyone above) | |
Signing timing, after the pilot is called and before it starts. Later than that, the do-it-all habit has already formed. Earlier than that, the staffing sheet has no factual basis yet. The owner must sign. He is clause four's audience, and clause two's labor cost is booked to his account (the co-build engineers' project time is a commitment clause, of the same rank as the user commitment, stating weekly hours: ____, and the accounting method (booked to the project / chargeback / written into their quarterly objectives, pick one) ____).
### 15.1.1 Clause One: Code Ownership Goes to the Receiving Side
| Item | Fill In | Check |
|----|------|------|
| Module CODEOWNERS | | Names the receiving side, not your team |
| Merge rights | | With the receiving side (gatekeeping rights are allocated in 15.1.2) |
| Oncall ownership | | The rotation spells out that these modules belong to the receiving side |
| CI / release process ownership | | Runs on the receiving side's existing release process, no second one built |
| Account form for your team's members | | Collaborator status, exiting after handoff |
| Exception list | | Which code does not go into the receiving side's repo (your team's internal tooling, for instance), listed item by item with reasons |
Red line: the day the AI team holds gatekeeping rights for the long run is the day permanent ops begins. There is no option of "develop on our side first and migrate at handoff."
### 15.1.2 Clause Two: Backward Staffing Sheet
The filling order is a rule. **Fill the second column first (who maintains it after handoff), then work backward to the third (who leads the writing now)**. Any row where the two disagree must have a convergence plan written out in the last column.
| Module | Maintainer After Handoff | Where the Maintainer Is Appraised | Lead Writer | Pairing / Review Side | Current Gatekeeping Rights and Transfer Condition |
|------|-------------|------------------|------|----------------|----------------------|
| | | | | | |
| | | | | | |
| | | | | | |
- **Leading the writing is not writing alone**: the lead writer is accountable for the code, can explain it, and decides on merges; the other side pairs.
- **Where the maintainer is appraised**: this column answers "when it breaks at midnight, whose annual objectives say to fix it." In most cases it should match the maintainer after handoff. A row where it does not is a hidden accountability vacuum, and the convergence plan has to say how that gets filled.
- **Gatekeeping rights** (the power to decide merges) get a column of their own. Your team may hold them early, but every row states the transfer condition (for example, "handed over after two consecutive reviews of that module with no rejection"). Before handoff day, gatekeeping rights on every row should sit with the receiving side.
- A module whose complexity is beyond the receiving side for now (the Anchor & Helm example, the queue core) still gets "maintainer, the receiving side (longer term)." The pairing on that row is not a courtesy, it is the convergence path.
### 15.1.3 Clause Three: The AI Code Review Rule
Scope, **all code generated or rewritten by an AI coding agent, on both sides alike** (your team's commits are checked too).
The rule in one sentence, **if you cannot say it, do not merge it.** The module's maintaining side reviews, and the submitter goes through the three explain-it questions face to face (or over video):
1. **Why is it written this way?** Every non-obvious choice (retry count, timeout value, algorithm, edge handling) has a reason he can state. "That is how the agent wrote it" is not a reason.
2. **Where can it go wrong, and how would you find out?** The failure path, who catches it, what the log or the alert looks like.
3. **If it has to change, where do you start?** The next foreseeable change request, where the change lands and how far it reaches.
Any question unanswered โ sent back. A rejection does not mean the code failed, it means the explanation failed. Before resubmitting, the submitter must actually have had his hands on it (rewriting, simplifying or adding tests all count). "Memorize it and say it again" is not accepted.
Suggested addition to the PR template, one field for the share of this commit generated by AI and the signature of the person who explained it (a PR template carrying that field is in Code Hooks at the end).
### 15.1.4 Clause Four: Rotating Release Cadence
| Item | Fill In | Default |
|----|------|--------|
| Frequency | | Every two weeks, 30โ45 minutes |
| Presenter | | The receiving side's engineers in rotation. Your team does not take the stage, only adds, never answers for them |
| Standing audience | | The receiving-side owner (required) + all co-build members |
| Date of the first one | | Within pilot week 2 |
The rotating release is not a training session (there is no one-way teaching) and not a status report (nothing is reported to your team). It is the receiving side's engineers showing their own organization's owner the progress that stands under their own names. Agenda in 15.2.1.
### 15.1.5 Anchor & Helm Example (Extract from the Version Signed in Week 9)
- Clause one: code goes into modules under claims-ops IT's name, CODEOWNERS and merge rights with claims-ops IT, oncall ownership spelled out. Digital Center engineers (you included) commit as collaborators, and cross-subsidiary code repository access was approved by Victor Reyes.
- Clause two: extraction pipeline, maintainer claims-ops IT / lead writer claims-ops IT / pairing Digital Center. Queue core, maintainer claims-ops IT (longer term) / lead writer Digital Center / pairing claims-ops IT. Merged view, lead writer Digital Center, lead-writer rights handed over before handoff. Rules and de-identification as configuration, lead writer claims-ops IT (the trust constraint matrix's maintainability row).
- Clause three: first enforced in pilot week 1, when the email pull retry logic was sent back for failing the first of the three questions. After the rewrite was merged, gatekeeping rights for that module moved to claims-ops IT.
- Clause four: Fridays every two weeks, from pilot week 2. Kevin Doyle as owner is required.
---
## 15.2 Knowledge Transfer Cadence Sheet
### 15.2.1 Biweekly Rotating Release Agenda Template (30โ45 Minutes)
| Time | Segment | Content and Rules |
|------|------|------------|
| 0:00โ0:05 | Review of the last beat | Reconcile last time's commitments item by item, with reasons for anything unfinished |
| 0:05โ0:20 | Demonstration | The receiving side's engineers demonstrate the effect on real data. Show "what it can do" and also "what a wrong one looks like and how it gets caught." A release that shows only successes does not pass |
| 0:20โ0:30 | Owner Q&A | The owner asks, the presenter answers. Your team adds only when named, and never answers for them |
| 0:30โ0:40 | Commitment for the next beat | The presenter, not your team, states the goal for the next two weeks |
| 0:40โ0:45 | Cadence sheet update | Update the milestones against 15.2.3. Were gatekeeping rights handed over? Were lead-writer rights rotated? |
### 15.2.2 The Three Explain-It Checks (Use Before Merging)
| Question | What a Passing Answer Looks Like | Failing Signal |
|----|--------------------|------------|
| Why is it written this way? | He can state the alternatives and the reason for not picking them | "That is how it came out" / "It runs" |
| Where can it go wrong, and how would you find out? | He can point to the specific failure path and the matching log, alert or backstop | "It should not go wrong" / "The tests all passed" |
| If it has to change, where do you start? | He can state the change point and how far it reaches | A long scroll through the code hunting for the entry point |
Usage, this is the enforcement tool for clause three and a self-check tool as well. Before you merge AI-generated code yourself, put the questions to yourself first.
### 15.2.3 Cadence Sheet: Biweekly Milestones from Pilot to Handoff
Fill it in by release beat. Besides the demonstration, every beat marks one "transfer action." The progress of knowledge transfer is not read from how many training sessions were held, it is read from how many items of power and responsibility were handed over:
| Beat | Time | What Was Demonstrated | Transfer Action This Beat (Examples) |
|------|------|----------|----------------------|
| 1 | Pilot week 2 | | Gatekeeping rights of the first module handed over |
| 2 | Pilot week 4 | | The receiving side's engineer handles a production issue alone for the first time (your team watches) |
| 3 | Pilot week 6 | | Lead-writer rights of another module rotate to the receiving side |
| 4 | Pilot week 8 | | The receiving side's engineer leads the pilot technical retrospective |
| โฆ | After launch | | Gatekeeping rights of the remaining modules handed over beat by beat; your team's share of commits declines |
| Last beat | Handoff (Chapter 22) | The receiving side's engineers demonstrate the whole system to the owner | Gatekeeping and release rights on every module sit with the receiving side, and the handoff checklist (Template 22) should have no new items |
Health self-check (three numbers every beat), the receiving side's engineers' share of commits (should rise), the number of modules where your team holds gatekeeping rights (should fall), the number of questions at the rotating release that your team answers for them (should trend to zero). If any one of the three runs the wrong way for two beats running, go back to Chapter 15's failure modes and find your match.
---
## Code Hooks
The companion repo provides (this repository's `repo/` directory):
- [`templates/cobuild/`](https://github.com/hallieren/the-last-mile/tree/main/repo/templates/cobuild/): the repo structure and CI ownership checklist, the external collaborator permission config sample, and the PR template (with the three explain-it questions field)
---
# Template 16 ยท Production Readiness Checklist and Running Budget Sheet
> Companion chapter(s): Chapter 16. The three tools are in order of use. Before launch (prototype up to pilot, pilot to production) run the gate item by item with 16.1. Fill in 16.2's four-ledger running budget sheet before launch and reconcile it weekly after. Register the fallback for every AI dependency point in 16.3, and only a drill that passes counts as "having a degradation path."
> License: Every template in this book may be modified freely and used in your work, no attribution needed.
---
## 16.1 Production Readiness Checklist
Tick item by item. Any "no" is either fixed, or accepted out loud at the escalation decision meeting with the accepting person written down. Passing in silence is not passing.
### Cost Ledger
- [ ] Is the cap on cost per claim set? Worked back from business value (what the labor hours one claim saves are worth, with system cost allowed only a fraction of it), not worked back from last month's bill.
- [ ] Does the cost instrument produce numbers daily, split by step (extraction / priority / other call points)? The month-end bill should only ever be a confirmation, never news.
- [ ] Do retries have a budget and a backoff (a per-claim cap on attempts, and where it goes over the limit, such as the "pending human" column)?
- [ ] Is recomputation incremental, or does every refresh run the whole thing again?
- [ ] Have the static blocks in the prompt (the manual, whole rule texts, piles of examples) been checked? Has everything that can move into retrieval or a cache been moved?
- [ ] Are the results for stable inputs cached (the same email is not paid for twice)?
- [ ] Are calls tiered? Are claims that do not need the big model using the big model? Is there eval evidence for where the split goes?
### Latency Ledger
- [ ] Is the P95 cap set? Worked back from the user's working rhythm. Averages do not count.
- [ ] Has the peak window (the morning peak, say) been measured on its own (with real concurrency, not one person clicking once)?
- [ ] Is everything that can be precomputed computed before the user arrives?
- [ ] Is the refresh asynchronous (show the most recent result with its timestamp marked first, update in the background)?
### Error Ledger
- [ ] Do the per-category error rate caps carry the eval spec threshold table over as is (Template 11.1.4)? Build nothing new, set no overall score.
- [ ] Is the online sampling cadence set (the golden case replay period, the spot-check proportion of production output, who reviews)?
- [ ] Are spot-check results booked by category, with the action on breach written down (rework / kill criteria, Template 14)?
### Degradation Ledger
- [ ] Does every AI dependency point have its fallback registered in 16.3?
- [ ] Has the degradation path been drilled (the kind where the dependency is really cut and real users are present)?
- [ ] Is the degradation trigger automatic (a failure rate or latency breach switches it over), or does it count on somebody switching by hand in the middle of the night?
- [ ] Is the rules layer still alive (an owner, tests, updated as business rules change)? Before deleting it, answer this. On the day supply is cut, what do you rely on?
### Observability
- [ ] Can any single suggestion be replayed from the trace (input snapshot, what was retrieved, prompt and model version, which rules hit, final display)?
- [ ] Are the trace and the audit decision trail kept apart? The trail answers "who decided what" (an audit commitment, Chapter 12), the trace answers "why does this suggestion look the way it does" (engineering replay). You need both.
- [ ] Are drift alerts configured (a sudden change in the input distribution, a sudden change in the override rate)?
- [ ] Have the alerts been verified? An alert that has never fired needs one firing manufactured for it. An alarm that does not sound is more dangerous than no alarm.
### Operating Ownership
- [ ] Does each of the four ledgers have a named owner (is the last column of 16.2 filled in)? A ledger you cannot put a name to will not be managed.
- [ ] Is the budget account for running cost settled? Has the business owner accepted that account in writing? After the handoff (Chapter 22), whose budget does this money come from?
- [ ] Is the cost booking standard settled (chargeback / showback / not booked)? Without a standard, calls only ever mix into one department total, and what this system spent cannot be split out.
- [ ] Is the cross-department first-line owner settled for each of the four ledgers? The test is whose budget or whose daily actions a breach hits directly. Everyone else related is written in as a countersignature.
- [ ] Is the reconciliation cadence set (who, which day of the week, which page)?
---
## 16.2 Running Budget Sheet (Four-Ledger Template)
One page, posted where the owner can see it. Set the alert line ahead of the budget line (at 80%, say), and write the action on breach hard in advance, rather than inventing it in front of the bill.
| Ledger | Metric | Budget Line | Alert Line | Current Value | Action on Breach | Ledger Owner |
|------|------|--------|--------|--------|----------|-----------|
| Cost | Cost per claim | | 80% of the budget line | | Cost review the same day | |
| Cost | Monthly call total | | 80% of the budget line | | Cost review the same day | |
| Latency | P95 on the core operation | | | | Degrade / extend precomputation | |
| Latency | P95 in the peak window | | | | Same as above, peak on its own line | |
| Error | unsafe class | 0 | Highest level on the first one | | kill criteria (Template 14) | |
| Error | concern class | Carry over the Template 11.1.4 threshold table (e.g., โค10%) | | | Rework retrospective | |
| Error | useless class | Carry over the Template 11.1.4 threshold table (e.g., โค20% with the trend not rising) | | | Rules and prompt retrospective | |
| Degradation | Drill status of each dependency point | Passed, and no more than 90 days ago | | | No change allowed until the drill is made up | |
**Anchor & Helm example rows (the pilot week 4 version, numbers stated relatively, your sheet needs absolute numbers)**:
| Ledger | Metric | Budget Line | Current Value | Ledger Owner |
|------|------|--------|--------|-----------|
| Cost | Cost per claim | โค the post-repair level ร 1.5 (post-repair = 1/6 of what it was before it ran away) | Inside the budget line | The claims-ops IT engineer (the instrument) โ follows the budget account to Kevin Doyle after the handoff |
| Cost | Monthly call total | The estimate line from the pilot approval | Back inside the line | Same as above |
| Latency | Morning peak queue refresh P95 | โค3 seconds | Met | The claims-ops IT engineer |
| Error | The three category thresholds | The Chapter 11 threshold table as is | Spot checks met | Linda Marsh's team (spot checks) + you (booking) โ the maintaining side after the handoff |
| Degradation | Two points, extraction and long-tail priority | Drilled (Friday of pilot week 3) | Passed | The claims-ops IT engineer |
**Rules**:
- [ ] Not one row of the error ledger is newly invented. It and Template 11.1.4 reference each other, and changing that side means changing this side.
- [ ] "Current value" is updated at each weekly reconciliation. Two consecutive weeks blank means the minimum observability set has a hole.
- [ ] The owner column takes a person's name, not a department's. At the handoff (Chapter 22) rename it row by row. This sheet is the working ledger of the operating handoff.
---
## 16.3 Degradation Path Register
One row per AI dependency point. A fallback has to be able to run the core process. "System under maintenance, please try again later" is not a fallback.
| AI Dependency Point | Failure Shape (Down / Slow / Expensive) | Fallback | What the User Sees | Trigger and How | Date Last Drilled | Owner |
|-----------|--------------------------|------|--------------|----------------|--------------|-------|
| e.g., missing document extraction | Down / slow | Pause extraction, rules rank as usual | The missing documents column grays out to "extraction paused, open the claim to read the original email" | Failure rate or latency breaches the line, switched automatically | | |
| e.g., long-tail priority suggestion | Down / slow / expensive | Everything goes through the rules layer, the reason column still says which rules matched | Long-tail claims marked "ranked by rules only" | Same as above. On a cost breach, route to the smaller model first, then degrade fully | | |
**Rules**:
- [ ] "Expensive" is a failure shape too. Write the degradation action for a cost breach (route to a smaller model, lower the sampling frequency, narrow the scope) in advance.
- [ ] The three elements of a drill, really cut the dependency, production data, real users present. Ask users afterward "did you notice." The best degradation is the one users did not notice.
- [ ] Every dependency point is drilled at least once every 90 days, and a change to a dependency point (a model swap, a change to the call path) makes a rerun mandatory.
- [ ] Chapter 10's line is honored on this sheet. Do not delete the dumbest thing that works. It is your degradation path. The rules layer needs an owner, tests, and updates as business rules change, or the fallback itself has already collapsed.
---
## Code Hooks
The companion repo provides (this repository's `repo/` directory):
- [`templates/production-readiness/`](https://github.com/hallieren/the-last-mile/tree/main/repo/templates/production-readiness/): the cost per claim instrument script, the drift alert rule sample, and the trace field scaffold
---
# Template 17 ยท Action Queue Design Patterns, Decision Trail Schema, Before/After Workflow Map
> Companion chapter(s): Chapter 17. The three tools are in design order. Set the column structure first (17.1), then the trail schema (17.2), and last, check the embedding points with the workflow map template (17.3).
> License: Every template in this book may be modified freely and used in your work, no attribution needed.
## 17.1 Action Queue Column Structure Checklist
**Design rule**: The column structure is the projection of the four layers of the decision rights boundary (Template 8.2). Layers AI owns get columns, layers people own get columns, and the red line layer **gets no column**. Go through column by column, and for each answer "which layer, who writes, who consumes." Longer design points move to numbered notes below the table, and the table keeps a phrase.
| Layer | Column | Who Writes | Who Consumes | Design Point |
|----|----|--------|--------|----------|
| Sense | Claim ID / business primary key | System | Everyone | Matches the source system's primary key, can be traced back |
| Sense | Reason stuck / missing item | AI extraction | Handler | Output can be checked (note 1) |
| Sense | Days waiting / aging flag | System | Handler, escalation rules | Counted in business days, not calendar days |
| Sense | Risk signal | Rules + AI | Handler | Signal traceable to its source (note 2) |
| Sense | Status-in-doubt flag | Reconciliation logic | Handler, owner | Lights up on mismatch (note 3) |
| Advise | Suggested priority | Rules + AI | Handler | Chase โ risk (note 4) |
| Advise | Suggested next action | Rules + AI | Handler | Starts with a verb, executable (note 5) |
| Advise | Reason | Rules + AI | Handler | The "ability" of the three questions (note 6) |
| Act | Owner | Pick up / assign | Everyone | Real name, never blank (note 7) |
| Act | Human Call | Handler | Trail, downstream actions | The "authority" of the three questions (note 8) |
| Decide | **(no column)** | / | / | The physical form of the red line (note 9) |
**Design point notes**:
1. Output can be checked. Put the LLM where its output can be checked (Chapter 10).
2. Every signal traces to a source rule or model version.
3. Lights up when the source system and the source of truth disagree (the in-row form of Chapter 9's status-in-doubt list).
4. The ranking logic eats the front line's real trade-offs. Chase โ risk.
5. Starts with a verb and can be executed directly. Outbound sending actions produce a draft only.
6. The "ability" of the three oversight questions. The information for judging right or wrong is in the same row.
7. Must be a named person, with the basis for assignment written down, one of SOP clause / job responsibility / supervisor assignment. Blank equals failure mode 2 (suggestions left hanging, a suggestion with no responsible person is not a suggestion).
8. The "authority" of the three oversight questions. Without this column the system takes no action.
9. The physical form of the red line. The payout decision has no place on the table.
**Two companion pieces** (Chapter 17, At Anchor & Helm):
- **Rollup view**: a page of numbers summed by team / category / aging, where every number clicks through to the queue itself. It is the entrance to the queue, not a parallel big screen.
- **Exception escalation rules**: set how many rows the recipient is willing to look at per day (the "time" of the three oversight questions), then derive the escalation thresholds backward. Keep signal categories to three or fewer, against alert fatigue.
## 17.2 Decision Trail Schema
**Guiding principle**: **Suggestions and facts are stored separately**. The fact table records only the world and what people did. The suggestion table is append-only, never updated, and not one word of AI output is written back to the source (Chapter 9's trail form, turned into a schema). Fields align item for item with Chapter 12's "six end-to-end decision trail fields."
| Field | Content | Aligns With |
|------|------|------|
| Input snapshot | Fingerprint / snapshot reference of the input data at the moment the suggestion was generated | Based on what |
| Rule / model version | Version number of the rule hit, model and prompt version | Based on what |
| Suggestion content | Priority + next action | What it said |
| Reason | The reason sentence shown to the handler, stored as is | What it said |
| Human Call | accept / override / hold | Who made the call |
| Decider | Real name, inherited from the source system identity | Who made the call |
| Timestamp | Suggestion time + decision time, both | When |
| **Reason code** | Required on override: enumerated value + optional note | Loop entrance |
**Sample reason code values** (Anchor & Helm's version, rewrite in the front line's own language for each project; no more than seven values):
1. Risk judgment differs (the word that keeps showing up in the notes is the name of the next rule)
2. Priority judgment differs
3. Information outdated (data lag, feeds into the fitness retest)
4. Already handled offline (workflow escape signal, the action happened outside the system)
5. Suggested action not feasible
6. Other (note required; "other" above thirty percent = the values need a revision)
**Review cadence** (the operating side of level six, closed loop, taken over by Chapter 18): split the override rate by suggestion category, look once a week; trace back the rule implementation for the category with the highest override rate; add claims with the wrong shape to the golden cases (Chapter 11).
## 17.3 Before/After Workflow Map Template
**Usage**: Before designing the embedding, rule on where every real step (field archaeology output, Template 6) goes. Three destinations, one line of test each:
| Where It Goes | Test | Example (Chapter 17's Fourteen Steps) |
|------|------|---------------------|
| **Absorbed** | Pure information hauling: find, copy, move, watch | Scan new claims, enter into the sheet, the Tuesday and Thursday manual filter of overdue claims |
| **Transformed** | Information hauling + human confirmation: the system does the first half, the person does the second | Pre-generated missing list reviewed by a person, drafted chase sent by a person |
| **Kept** | Judgment and relationships: persuade, coordinate, decide | Calling the surveyor, backing up Excel before leaving |
Template table:
| # | Real Step (Before) | Where It Goes | After Form | Notes (Red Line / Dependency) |
|---|--------------------|------|-----------|---------------------|
| 1 | | Absorbed / Transformed / Kept | | |
**Three checks**:
- [ ] Is everything absorbed pure information hauling? A judgment absorbed = the decision rights boundary is drawn wrong, go back to Chapter 8.
- [ ] Does the "person does the second half" of each transformed step land on a queue column (Human Call / review), not in another system? (Failure mode 3, five tools stitched into one workflow)
- [ ] Are the kept steps written down explicitly? "What the system does not touch" matters as much as "what the system does." It is the boundary promise to the front line.
## Code Hooks
The companion repo provides (this repository's `repo/` directory):
- [`templates/action-queue/`](https://github.com/hallieren/the-last-mile/tree/main/repo/templates/action-queue/): trail table DDL (fact table / suggestion table split) + sample reason code enumeration config
- [`templates/action-queue/`](https://github.com/hallieren/the-last-mile/tree/main/repo/templates/action-queue/): sample weekly report query splitting the override rate by category
- [`templates/action-queue/`](https://github.com/hallieren/the-last-mile/tree/main/repo/templates/action-queue/): sample config for the rollup view and escalation rules (thresholds, recipients, signal categories)
---
# Template 18 ยท Metric Tree, AI Incident Runbook, Same-Day Notice
> Companion chapter(s): Chapter 18. The three tools are in launch order. Before launch, build the metric tree and baseline (18.1), write the incident runbook (18.2), and have the same-day notice template ready (18.3). All three must be done before the first incident happens.
> License: Every template in this book may be modified freely and used in your work, no attribution needed.
---
## 18.1 Metric Tree Template
### 18.1.1 Three-Layer Structure (Fillable)
One row per metric. **The owner and "action on deviation" columns may not be left blank.** Delete any metric you cannot write an action for. It is only scenery.
| Layer | Metric | Definition and Measurement | Data Source | Baseline | Owner | Owner's Reporting Line | Action on Deviation (Who Does What) |
|----|------|-----------|----------|------|-------|-----------------|----------------------|
| North Star (outcome) | | From the charter, the only one | | โฅ8 weeks before launch | | | |
| Process (mechanism) | | Can be changed directly by an action | | | | | |
| Process (mechanism) | | | | | | | |
| Balancing (cost) | | The cost the improvement may shift onto someone | | | | | |
| Balancing (cost) | | | | | | | |
### 18.1.2 Anchor & Helm Example
| Layer | Metric | Definition and Measurement | Owner | Owner's Reporting Line | Action on Deviation |
|----|------|-----------|-------|-----------------|----------|
| North Star | First-touch handling time | From an auto exception claim entering exception status to first handling complete, weekly median, shown on a run chart | You + Kevin Doyle | Dual owner, inside your team + business side | Two consecutive points back above the baseline median โ go through the process layer for the cause |
| Process | Queue aging | Count of overdue claims in the queue not moved | Linda Marsh (team lead) | Business side, does not report to you | Over the line โ claimed and cleared the same day |
| Process | Missing-document detection lead time | Time from a claim entering exception status to the missing document being identified and a chase sent, against baseline | Your team's engineer (renamed line by line at handoff) | Inside your team | Lead time shrinks โ check the extraction pipeline |
| Process | Override rate | Split by suggestion category (per-category measurement, no overall average) | You | Inside your team | A sudden change in one category โ trace back that category's rule implementation and list version |
| Process | Status-in-doubt list length | Count of claims the queue flags as status in doubt (Chapter 9's operating item) | Kevin Doyle | Business side, does not report to you | Continuous growth โ re-check the review team's ten-minutes-a-day status cleanup discipline |
| Balancing | Complaint rate | Complaints related to exception claims, shown on the same page as the outcome | Kevin Doyle | Business side, does not report to you | Rising โ check chase scripts and frequency |
| Balancing | Reviewer overtime hours | Weekly overtime for Linda's team | Kevin Doyle | Business side, does not report to you | Rising โ check the queue's daily volume setting |
| Monitoring surface (fairness) | Override distribution spot check by segment | Override distribution by customer segment, once a quarter (Chapter 12's fairness row) | You + Victor Reyes | Inside your team + risk owner | Significant skew โ trigger a fairness review |
### 18.1.3 Checklist
- [ ] One North Star, word for word the charter North Star; no new metric set up after launch.
- [ ] **Every metric has an owner and an action on deviation**; the owner is a named person, not a department.
- [ ] For metrics whose owner does not report to you, the action on deviation is written into their team SOP or weekly meeting agenda, not left hanging on your reminders; an action written into the SOP is their job at review time.
- [ ] Balancing metrics โฅ2, shown on the same page as the outcome metrics; a balancing metric on a separate page is always "next time."
- [ ] Baseline: โฅ8 weeks of data before launch, cleared through the data check claim by claim (whatever does not reconcile is corrected first, Chapter 9).
- [ ] The run chart shows every point, no picking weeks; improvement is claimed only when a special cause signal appears (such as "six points on one side": six consecutive points on the same side of the baseline median, paraphrased from *The Health Care Data Guide*).
- [ ] The monitoring surface (override alert thresholds, error category distribution, golden cases replay cadence) is maintained in the same table as the metric tree, with definitions citing the eval spec (Template 11).
---
## 18.2 AI Incident Runbook
### 18.2.1 Level Table
| Level | Definition | Response | Anchor & Helm Example |
|------|------|------|----------|
| **P1** | An unsafe-class error enters human view; **being caught by a person does not downgrade it** (what caught it was already the last line of defense) | Draft the same-day notice within two hours; retrospective within 48 hours; incident case into the golden cases within 24 hours | A risk claim suggested as routine |
| **P2** | Concern-class errors over threshold or appearing in batches | Retrospective the same week; trace back rules and data sources | Missing-document lists wrong in batches |
| **P3** | A rising trend in useless-class errors | Folded into the monthly retrospective | Share of empty suggestions rising two weeks running |
### 18.2.2 The Four Retrospective Questions
Go layer by layer. Wherever the fault is located, that is where the repair goes. **Attribution decides the repair path. Attribute it to the wrong layer and the repair fixes the wrong place.**
| Layer | Test Question | Typical Repair Path |
|----|----------|--------------|
| **Model layer** | The input signal was sufficient and the model still judged wrong? | Switch pattern / add a human review gate / lower the decision rights layer |
| **Data layer** | The input itself was wrong (list not updated, fields distorted, source drift)? | Fix the data + add operating discipline (give the data asset an owner and an update cadence) |
| **Rule layer** | A rule missing, or the implementation inconsistent with the annotation guide? | Change the rule, release a new dated version of the annotation guide |
| **Interaction layer** | A person saw it but had no time / no basis / no authority to stop it? | Change the queue design (fix the three oversight questions item by item, Chapter 12) |
Note: the table is in definition order of the four layers. For the order of investigation see Chapter 18's four retrospective questions (model โ rule โ interaction โ data).
The attribution walkthrough for Anchor & Helm's misclassification incident: model layer (no usable signal in the input, not a model capability problem); rule layer (implementation consistent with the guide); interaction layer (all three questions pass, the catch succeeded); **data layer hit** (the repair shop list did not cover a newly registered entity). Repair: the list promoted to an operating asset with an owner (monthly update, dated version, config item changeable), not a model swap.
### 18.2.3 Incident Case Flow-Back into Eval
- [ ] 1. The incident case goes into the golden cases within 24 hours (historical incident category, Template 11.2), kept permanently.
- [ ] 2. The kill criteria are checked by the book and recorded, **whether or not they trigger**; the check record is attached inside the incident record (Chapter 14).
- [ ] 3. The retrospective conclusion is settled in writing. Rule layer changes go into a dated version of the annotation guide; data layer changes state the new owner and update cadence.
- [ ] 4. After the repair goes live, the golden cases (including the new incident case) are retested once, and the result is attached to close the incident record.
- [ ] 5. If an error of the same shape recurs, respond at the next level up. The runbook itself gets a retrospective too.
---
## 18.3 Same-Day Notice Template
Three parts, and the order cannot change: the defense first, then the scope, and root cause last. **Drafted within two hours, sent the same day**; the distribution is drawn by "how far the rumor can reach," and no individual is named as responsible.
> **[Notice] On today's [error category] by [system name]**
>
> Issued by [business-side owner; you are the drafter]
>
> โ [The defense held] Today at [time], the system suggested a claim that should have been [X] as [Y]. [Review role] caught it in the Human Call column and overrode it. **That step exists to catch exactly this kind of error. The defense worked as designed, and the full trail is on record.**
> โก [Scope of impact] No real action was taken on the claim; all suggestions from [today / the same period] have been checked, [no error of the same shape / N more claims handled together].
> โข [Root cause under investigation] Preliminary direction [one-sentence direction], retrospective conclusion and improvement measures within [deadline].
Anchor & Helm example (pilot week 5):
> Issued by Kevin Doyle. This morning the queue suggested a high-risk claim as routine. The review team lead caught it in the Human Call column and overrode it to high risk. That column exists to catch exactly this kind of error. The defense worked as designed, and the full trail is on record. Scope of impact: no real action was taken on the claim; today's queue has been checked, and there is no error of the same shape. Root cause under investigation, preliminary direction the repair shop list not covering a newly registered entity, retrospective conclusion within 48 hours.
**Rules**:
- [ ] Part one always states the interception mechanism first, then the error. This is the correct order of the facts, not spin. The system's design premise is that AI will make mistakes and people backstop them.
- [ ] You hold the drafting right, the business-side owner holds the issuing right; only when the person with the power to violate this defense issues it does the notice avoid being read as the technical team defending itself.
- [ ] The scope of impact states only verified facts; what is not verified is written as "being checked," never a guess.
- [ ] The promised retrospective deadline must be met. The notice itself is a reliability deposit (Chapter 5).
- [ ] Before any incident, drill once with a historical wrong output and time it; lock down in advance any step that runs past two hours.
---
## Code Hooks
The companion repo provides (this repository's `repo/` directory):
- [`templates/metric-tree/`](https://github.com/hallieren/the-last-mile/tree/main/repo/templates/metric-tree/): run chart generation script (baseline median and special cause signal marked automatically) and sample override alert threshold config
---
# Template 19 ยท Three-Memo Set (with ADR)
> Companion chapter(s): Chapters 13 and 19. The four tools are in project order. Fill the kickoff template (19.1) before a phase starts, the decision memo, which is the ADR memo (19.2), at any midpoint where an executive has to call something, and the impact template (19.3) before a phase closes. Run the shared rules list (19.4) before every send.
> License: Every template in this book may be modified freely and used in your work, no attribution needed.
---
## 19.1 Kickoff Memo Template (Before a Phase Starts)
> Kickoff memo. The one page sent to the sponsor before a phase starts, putting on record the outcome promised, the investment needed, and the ways to die already rehearsed. It usually asks for no call (the authorization already came from the previous impact memo or from project approval). It translates that authorization into commitments that can be checked.
```
To: [sponsor] From: [you] Date: [Monday of the week the phase starts]
Subject: [phase name] starts, for the record. Promises, investment, and ways to die
One page of body text; the full pre-mortem and the plan attachments are there for reference.
```
- Three fixed copies, your manager, the business-side owner, and the PMO (Chapter 19)
| Section | What Goes In It | Rules |
|----|--------|------|
| โ SCQA | S = the previous impact memo's approval outcome / C = the new variable in this phase / Q = what is promised, what is needed, how it will die | For C, write "where this phase differs from the last one." If there is no new variable, no new phase is needed |
| โก The outcome promised | The North Star target to move toward + one early visible win, each with a date | Keep the charter's metric, do not set a new one; give the early win a date on the order of "inside the first month" |
| โข Investment needed, who sends the people | People and time from other departments, named to the person, with hours per week spelled out | Cite a precedent where one exists ("on the precedent of team X"); investment left out of this section will never be recognized later |
| โข Investment needed, who gives the word | Whose mouth this commitment has to come out of before it counts, the sponsor himself or the other side's manager | Every item landing in "who gives the word" is a political act. Work out whose standing you are spending before you write it |
| โฃ Pre-mortem summary | The three most likely ways to die + the defense for each | Take the top three from Template 4's output; this is the memo's expectation management slot |
| โค What I need from you | Can be as small as "hold ten minutes on your calendar" | It has to be there, even if it is only "one look at the run chart in week N" |
**Anchor & Helm worked example** (the week 19 expansion kickoff, full text in Chapter 19):
- โก Promise: the North Star reaches the charter's -30% target line (eight weeks); one business win worth announcing inside the first month.
- โข Investment, who sends the people: Kevin Doyle, schedule protection for the two new teams, 2 hours a week each on the precedent of Linda's team; the two chase-side people continue to the end of the expansion period (already confirmed at the annual budget and headcount review, recorded here).
- โข Investment, who gives the word: the first two sit inside Kevin's own line and his word is enough; the two on the chase side do not, so Grant Whitmore says a line to the chase side's manager at the next monthly meeting about continuing them.
- โฃ One way to die, as an example: the new teams accept everything as is, with none of Linda's kind of challenge โ the defense, an override rate clearly below Linda's team's baseline in the first month triggers a spot check.
- โค The ask: no call needed; ten minutes at each monthly meeting, one look at the run chart in week 22 and one in week 26.
## 19.2 Decision Memo / ADR Memo (the Midpoint)
The decision memo is Chapter 13's ADR memo, the midpoint piece of the three-memo system (kickoff at the open / decision at the midpoint / impact at the close). Three tools work together. Fill the template (19.2.1) once at every point where an executive has to call something. Run the self-check list (19.2.2) before sending. Use the anti-pattern table (19.2.3) to diagnose when a memo comes back or gets read with nothing following.
**When to use it**. Any midpoint where an executive has to call something (a plan settled, added investment, a scope change, stay or go after kill criteria trigger). The test is the pyramid apex test. If you can say "who I want and what I want him to call" in one sentence, it is a decision moment.
**How it joins the other two**. When reality tests a kickoff's pre-mortem defense and the decision has to change, go to a decision memo. What a decision memo gets approved enters the evidence section of the next impact memo. Once an impact memo's option is called and a new decision point comes up during execution, you come back to a decision memo again. The split in one line. **The kickoff puts commitments on record, the decision memo asks for a call, the impact memo delivers the evidence.**
### 19.2.1 ADR Memo Template (One Page of Body Text)
> ADR memo. One page presenting one pending decision to the person with the authority to call it (conclusion, evidence, the cost said out loud, risk and backstop, and what he has to do). Its purpose is not to make him understand the plan. It is to let him call it responsibly.
**The bar before you start writing (the pyramid apex test)**. Who you want and what you want him to call, in one sentence. If you cannot write that sentence, do not open the template.
```
To: [the person who calls it, one name] From: [you] Date: [ ]
Subject: [decision name], for you to call [N] things by [date]
One page of body text; [the technical plan / the reconciliation report / the review packet, and so on] are attached for reference.
```
- Three fixed copies, your manager, the business-side owner, and the PMO (Chapter 19)
#### โ SCQA Opening (Three Lines)
| Line | What Goes In It | Rules |
|----|--------|------|
| S (Situation) | A fact both sides already agree on | Write only what he already knows and accepts. New information in S and the opening loses its anchor of agreement |
| C (Complication) | The change or risk that threatens that agreement | Use the single hardest piece of evidence (one scoring run, one distortion rate), do not pile them up |
| Q (Question) | The question that rises in the reader's mind as a result | It has to be **his** question ("should we invest now, and how much"), not yours ("which architecture do we pick") |
#### โก The Conclusion in One Sentence
The call to be made, one sentence, in bold. It carries scope, duration, and the validating metric. Give the source of the key number in one parenthesis (cite the charter or the reconciliation report), and do not unpack it.
#### โข Three Pillars (One Line of Evidence Each)
- Pillar 1: [reason], [number / signature / scoring result]
- Pillar 2: [reason], [evidence]
- Pillar 3: [reason], [evidence]
Rules. MECE, three that do not overlap and together are enough to hold up the conclusion. Evidence is nouns ("200 records reconciled, 44% distorted," "co-signed by the security reviewer"), not adjectives ("a solid data foundation"). One line each. If it will not fit on one line, the evidence is not hard enough yet, or you are stuffing the reasoning process into a pillar.
#### โฃ The Trade-off Said Out Loud (the Soul of the Memo)
- Given up 1: [what], revival condition: [a checkable event]
- Given up 2: [what], [required for an AI system, the decision rights boundary map, which layer AI stops at and on what evidence it moves up (paste in Template 8.2's four-layer ladder as it stands)]
- What it buys: [one sentence, usually "a result that can be validated inside N weeks"]
Rules. Give every item given up a **place** (a revival condition), turning "no" into "not now" (Chapter 8). Transcribe the material straight from the scope decision log (Template 8.3), do not invent it here. Self-check. This section must contain at least one sentence that hurt to write. Not one, and you are selling, not reporting. **Rule. Once a hidden cost is exposed, you permanently lose the right to lead with the conclusion.**
#### โค Risk and Backstop
Write four things. What errors will occur (**error categories**, not a vague "there may be some error") โ the threshold and tolerance for each category (zero-tolerance classes listed separately) โ the backstop path (at which step human review catches it, and whether overrides leave a trail) โ the resource reassessment conditions (what signal triggers a resource reassessment, citing the three tiers of resource reassessment, Chapter 4).
Rules. Put it the way a non-technical decision maker hears it. Promise no "it will not make mistakes," and say instead "when it is wrong, it can be found and it can be caught." Panic does not come from "it will make mistakes," it comes from "nobody handles it when it goes wrong." This section is a commitment, not an apology. Every sentence needs a subject and a mechanism, with no "we will try" and no "in principle."
#### โฅ The Three Things I Need From You (Each With a Date)
1. [date / occasion]: [action, starting with a verb]
2. [date]: [action]
3. [date]: [action]
Rules. Three at most. Each one is **his** action (call it, approve it, confirm it), not your plan. Dates down to the day. Without this section the memo is a report, not a request.
- **Anchor & Helm example** (the ADR memo in Chapter 13): 1. Call the pilot's scope and start at next Monday's weekly; 2. By next Wednesday, approve the development effort for the merged view, two people from the Digital Center and two from claims-ops IT, six weeks; 3. By next Friday, confirm with Kevin Doyle that Linda's team's 2 hours a week keeps running through the pilot, with the business-side owner confirmed down to a name, and the landing action and the person accountable written out.
### 19.2.2 SCQA Writing Self-Check List (Run It Before Sending)
Ask yourself item by item. Any "no" and you go back and fix it, you do not send:
- [ ] **The pyramid apex test**. Can you say in one sentence what you want him to call? (If you cannot โ you have not thought it through, and AI cannot think it through for you)
- [ ] Is S a fact he already accepts? Does it smuggle in new information that needs arguing?
- [ ] Does C use the single hardest piece of evidence? Will Q rise in him on its own after reading it?
- [ ] Is Q his question, or your question with the subject swapped?
- [ ] Is the conclusion on the first screen? Delete the rest of the body, keep only SCQA plus the conclusion, and does he know what you want?
- [ ] Do the three pillars cover his three "can I answer for it" questions (why believe it / what does it cost / what happens when it goes wrong)? Are any two of them overlapping?
- [ ] Is every piece of evidence a verifiable noun? Do "significant," "solid," or "basically" appear?
- [ ] Does the trade-off section have the sentence that hurt to write? Does every item given up have a revival condition?
- [ ] AI system item. Is the decision rights boundary drawn in? Are errors put as categories and thresholds, or as one vague accuracy number?
- [ ] Does the backstop path have specific people and steps? Do the resource reassessment conditions cite the three tiers of resource reassessment (Chapter 4)?
- [ ] Do all three things carry dates? Are they all his actions?
- [ ] Is the body inside one page? (attachments unlimited)
- [ ] The four-minute test. Give it to someone who does not know the project for 4 minutes. Can he repeat back "what is the decision, what is the cost, what happens when it goes wrong"?
### 19.2.3 Anti-pattern Table
| Anti-pattern | Symptom | Root Cause | Fix |
|--------|------|------|------|
| **Chain-of-reasoning narrative** | It starts from "project background" and the conclusion sits on the last page | It copies the order of the work, not the order of the reader | Once it is written, move the last paragraph to the front; at every paragraph ask "what will he ask when he gets here" |
| **Benefits only** | Upside throughout, with the cost left waiting to be asked about | Treating the memo as sales copy; fear that bad news scares the decision away | Make the trade-off section mandatory; remember that one exposure takes credibility to zero |
| **Asking for understanding, not a decision** | The logic is perfect, the reader nods, and then nothing happens | Treating "explained it clearly" as the finish line of delivery; a vague request dodges the risk of refusal | Add "the three things I need from you + dates"; the nod is for you, not for the project |
| **A body over one page** | "It is all important, none of it can go," and the body runs ten pages | The cost of the trade-off is passed to the reader; the apex has not been found yet | Treat going over as a diagnostic signal. Back to the pyramid apex test, think it through and then write; all detail goes into the attachments |
## 19.3 Impact Memo Template (Before a Phase Closes)
> Impact memo. The one page sent to the person who calls it when a phase closes, presenting the result with running evidence, presenting the gap honestly, and giving next-step options with one of them recommended. Its purpose is not to sum up the past. It is to trade the past for the next step.
```
To: [the person who calls it] From: [you] Date: [48 hours before the retrospective or the annual budget and headcount review]
Subject: [phase name] retrospective, for you to call the next step at [occasion]
One page of body text; the full run chart, the incident retrospective record, and the metric tree detail are attached.
```
- Three fixed copies, your manager, the business-side owner, and the PMO (Chapter 19)
| Section | What Goes In It | Rules |
|----|--------|------|
| โ SCQA | S = the phase and scope in one sentence / C = the hardest result number (including a missed target) / Q = was it worth it, what next | Write the number straight into C, target hit or missed. A gap you hide ends up found for you by the reader |
| โก The conclusion in one sentence | Whether the validation holds + which option is recommended | In bold; carries scope and duration |
| โข Evidence | The run chart cited + the mechanism explained + the defense's field record | See "run chart citation rules" below |
| โฃ The honest gap | Target number against actual number; attributed to verifiable events; levers started but not run through; a void condition of its own | See "the three parts of an honest gap" below |
| โค Next-step options | 2โ4 of them, each with content / investment / when to choose it | See "option rules" below |
| โฅ What I need from you | Call one of the options, with a date and an occasion | Mark the recommendation as "we recommend N" |
**Run chart citation rules**. Every point goes on the chart, the baseline median as one reference line, and the name of the special cause rule used is stated (such as "six points on one side"). No cropped stretches, and no trend drawn between two points. Write the defense record whether it looks good or bad. Incident count, what was caught, which layer the retrospective attributed it to, flow-back into the golden cases, and the kill criteria check result. **Zero triggers still gets written as "zero triggers,"** because it is the evidence that the clause for stopping has teeth.
**The three parts of an honest gap**. โ Number against number ("-22% against -30%, 8 percentage points short," never "close to target"). โก Attributed to verifiable events (when the people arrived, when the fix took effect), never to "we need more time." โข Carrying its own void condition ("if the metric stops improving inside X weeks, this judgment is void and we go back to option N"). If you cannot write โข, then โ and โก are a justification, not an attribution.
**Option rules**. Options other than the recommendation have to be real options, each stating "when to choose it." Two standing options. **Revival list items** are stated by their revival condition (is the condition ripe, how far short), not by mood. **Hold and watch** always sits last. It is the exit for "the evidence says not yet" (the closeout form of Chapter 7's re-evaluation condition). State the ceiling on the watching period and its cost, and when it comes due with no new evidence, "another month of watching" has to go through project approval again (Chapter 14).
**Anchor & Helm worked example** (the week 18 pilot retrospective, full text in Chapter 19):
- โ C: the North Star finished at -22%, short of the charter's -30%.
- โข One piece of evidence, as an example: all eight weekly points below the eight-week baseline median, and "six points on one side" holds. A shift, not fluctuation.
- โฃ The gap: the chase-side people covered only half the stretch and the list's weak-signal prompt went live only in week 15, so neither lever has run a full cycle; void condition = if it stops falling four weeks into the expansion, we go back to option three.
- โค Options: expand and go deeper (recommended) / start home property (revival condition = "the first extension after the pilot North Star hits target" ยท exclusive; the -30% target line against -22% today, condition not ripe) / hold and watch (four more weeks, cost = stopping the scale-up just as the levers start working).
## 19.4 The Shared Rules List for All Three (Run It Before Sending)
- [ ] One page of body text, attachments unlimited (over a page = not thought through)
- [ ] Lead with the conclusion, a three-line SCQA opening (self-check with 19.2.2)
- [ ] The ending carries "what I need from you," a concrete action with a date. A kickoff's ask can be as small as "hold ten minutes," but it has to be there
- [ ] The expectation management slot is in place. The kickoff has the pre-mortem summary / the decision memo has the decision rights ladder / the impact memo has the honest gap
- [ ] Numbers match the charter and the run chart, with no orphan numbers (every number traces back to an attachment)
- [ ] The "I will tell you the moment X shows up" promised in the last memo, has it been honored? (the message that honors it = one line plus one chart, carrying no request)
- [ ] Is the copy line complete (your manager / the business-side owner / the PMO)? If the investment needed means going over a peer department to ask for people, does that need cross-department administrative coordination, and is it marked?
**Send timing table**:
| Memo | When to Send | Meeting Shape |
|------|----------|----------|
| kickoff | Monday of the week the phase starts; inside one week of the approval meeting | Usually no meeting needed. For the record plus one calendar request |
| decision | Delivered 48โ72 hours before the meeting where it gets called | A 10-minute meeting. Silent reading first, then the pillars |
| impact | Delivered 48 hours before the retrospective or the annual budget and headcount review | The meeting works the options around the document, and the approval comes as it breaks up |
The shared rule on timing. **Document before meeting.** The executive reads the one page before the meeting, and only then is the meeting the place where it gets called rather than the place where it gets reported for the first time. Within 24 hours of the meeting, write the approval outcome back into the project record (that is the S of the next memo).
## Code Hooks
The companion repo provides (this repository's `repo/` directory):
- [`templates/memo-suite/`](https://github.com/hallieren/the-last-mile/tree/main/repo/templates/memo-suite/): outline scaffolding for the three memos (ADR memo included), with sample prompts that take phase data (charter metrics, run chart, pre-mortem items) and generate draft SCQA and evidence sections; the conclusion sentence is forced to stay empty, because only a person places the apex
- [`templates/memo-suite/`](https://github.com/hallieren/the-last-mile/tree/main/repo/templates/memo-suite/): the run chart plotting script plus the generator for the one-line "point N" short notice (it takes the metric tree data source of Template 18)
- [`templates/memo-suite/`](https://github.com/hallieren/the-last-mile/tree/main/repo/templates/memo-suite/): a send timing reminder sample, generating the three memos' deadline reminders from the phase calendar (worked back from 48 hours before the meeting)
---
# Template 20 ยท Resistance Decoder, Hard Conversation Scripts, Visibility Governance Pledge
> Companion chapter(s): Chapter 20. The three tools are in order of use. Decode first (20.1), then have the conversation (20.2), and when it hits the fear of visibility, give the governance pledge (20.3).
> License: Every template in this book may be modified freely and used in your work, no attribution needed.
## 20.1 Resistance Decoder
**Rules.** The table gives candidate real causes, not a diagnosis. The diagnosis comes from an interview anchored on a specific instance (the three-layer probing method, Template 6.3). Order discipline. Name the emotion first, decode the real cause second, change the design last. Reverse the order and not one step works.
| Form of Resistance | Common Real Cause | Response Action |
|----------|----------|----------|
| Delay, rescheduling, "too busy lately" | **Interest**. This is all cost to him, with nothing he fears or wins at stake | Go back and consult. Find the fight he is already in, and help him win once first (Chapter 5) |
| A barrage of detail, endless technical challenges | **Fear** disguised as professionalism. Occasionally he really does know | Name it. "I sense there is another worry behind the detail. Can you say more?" And go through the detail seriously at the same time, because he may be right |
| Sniping at a meeting, a challenge in public | **Pride**. Being turned into data, skipped over, compared | Do not defend, name the emotion first, book the conversation into the field (one on one) |
| Data withheld, a process forever "in progress" | **Information gap + risk**. He does not know what you are up to, or who carries it when something goes wrong | Go back and consult + a governance pledge. Open again in the other person's language of power (Chapter 5) |
| Agreed in principle, nothing moves | Interest or fear, politeness is the disguise | Narrow it to one concrete action and one concrete date, watch which step it sticks at |
| Silence, no objection and no use | The deepest kind. The conversation has been given up on, or he was never invited into it | Go to them and shadow. Force a check of the design defect hypothesis |
**Shorthand for the four real causes.** Interest (what does he lose) / fear (how will the data be used, what will replace him) / pride (is he the last to know) / information gap (does he know what you are up to). One piece of resistance often comes in layers, one real cause on the surface and another underneath.
**The Anchor & Helm surveyor row (the full decoding of the Chapter 20 instance)**:
| Item | Content |
|----|------|
| Form | Sniping at a meeting. "This system of yours, is it here to appraise us?" |
| Surface cause | Fear (monitoring), and entirely rational. The queue really did make "how many days each claim sat with whom" visible |
| Deeper cause | A design defect. The real reason surveyors are slow to complete documents is that repair shops are slow to respond, and the system attributed it to the wrong step |
| Response action | Name it on the spot โ the decoding interview in the field that week โ change the design ("waiting on external" status, bottleneck attribution to the step) โ the governance pledge (the three pledges, issued by Kevin Doyle) |
| Verification | Three weeks later the person concerned comes to the retrospective on his own, carrying his own numbers |
## 20.2 Hard Conversation Scripts
> **A script is a crutch, not a line to recite.** Use your own words, but keep the function of every sentence. A script read out loud is worse than a defense.
**Opening without defending** (when challenged in public):
- "That worry is fair. This system really can [be used to appraise people / see everyone's data], and I am not going to pretend it cannot. Let us talk about how we make sure it is not used that way." (Function, admit the capability, do not defend the intent)
- "I sense this plan worries you. Can you say more?" (Block's original form. Function, name it, then shut up)
- "You know this step better than I do. I want to hear from you first, where it is wrong." (Function, hand back the pride)
**Probing to confirm the real cause** (one on one or in the field):
- "Was there a specific claim last week that made this number feel unfair to you? Let us start from that one." (Function, anchor on an instance, get out of the abstract argument)
- "If these numbers were pulled tomorrow, would your worry go away? What would be left?" (Function, separate the fear of visibility from the other real causes)
- "In the worst case, who do you worry will use this data, and how?" (Function, turn the fear into a concrete scenario, because only a scenario can be answered by policy)
- "Under what conditions would this system feel like it is helping you rather than watching you?" (Function, get the other person to state the acceptance condition)
**Closing with a commitment and a follow-up**:
- "I will come sit half a day with you this week and go through the claims from the top. You will hear back from me by Friday. If we recorded it in the wrong place, we change the system, not the wording." (Function, give a date and give a way out. "Change the system, not the wording" is the key pledge)
- "I will draft these three pledges, [owner's name] issues them, and they go to the whole line. Which one do you think is missing?" (Function, turn the other person into a co-author of the governance pledge)
**The banned list** (every one of them defends intent and dodges capability). "The system does not mean it that way." / "The data does not lie." / "This is a company-level decision." / "Do not get emotional."
## 20.3 Visibility Governance Pledge Template
**Generic template**, *Governance Pledge on the Use of [system name] Data*:
| Pledge | Landing Mechanism | What a Violation Looks Like (for Self-Check) |
|------|----------|------------------------|
| One. Data the system produces is used to improve the process, not to appraise individuals | The performance appraisal metric list may not cite any individual number the system produces. This clause goes into the policy document, and no verbal version is accepted | Someone's waiting time or override records show up in a performance review conversation |
| Two. The team concerned sees its own data first | Any material containing a team's numbers goes to that team's lead [N] working days before the meeting | A team sees numbers about itself for the first time in a meeting room |
| Three. Aggregate display takes priority over individual detail | The upward-facing view goes no finer than the step / the team. Individual detail is visible only inside the team's own view | A table ranking individuals circulates outside the team |
| Four. Technical guardrails (the hard means behind the three pledges) | Splitting a table by person is disabled at the warehouse layer. The individual dimension is de-identified in the shared layer. A pull that needs individual detail goes through approval and leaves a trail (who, on which day, why, and what was pulled) | Someone goes around approval and one query ranks the data by person |
**Issuing elements** (missing any one voids it):
- **Issuer.** It has to be the person with the power to violate the pledge (the business owner), not the project team. A pledge issued by the project team binds nobody.
- **How it takes effect.** Issued in writing, sent to every team the system covers, entered into the policy document. When the system extends to a new team, it is reissued along with the extension.
- **Issuer change.** When the business owner changes, the successor reissues within thirty days. Where there is no reissue, the pledge is flagged red on the PMO's AI system transfer ledger and walked through at the quarterly business review along with the other transfer items (Chapter 22).
- **Appeal path.** Anyone who believes a pledge has been violated raises it with [the issuer] or [an independent third party, the union representative / HR], with a written reply within [N] working days. Appeals do not go through the project team.
**Anchor & Helm version** (Friday of week 21, issued by Kevin Doyle, sent to the whole claims line):
> One. All data produced by the exceptions queue and the rollup view (including waiting time, bottleneck attribution and override records) is used to improve the process and is not a basis for any individual performance appraisal. The claims line's performance metric list may not cite the above data.
> Two. Any page of data summed by team goes to that team's lead two working days before the meeting.
> Three. Upward display goes no finer than the step and the team. Individual detail is visible only inside the team concerned's own view.
> Issuer: Kevin Doyle (Director of Claims Operations). Takes effect: from the date of issue. Appeals: anyone who believes the above pledges have been violated may raise it directly with the issuer or the union representative, with a written reply within five working days.
## Code Hooks
This appendix has no code companion. (The queue config sample for the "waiting on external" status and bottleneck attribution belongs to the code hooks of Template 17.)
---
# Template 21 ยท Adoption Plan, Super-user Agreement, Operating Cadence Sheet
> Companion chapter(s): Chapter 21. The three tools are in order of use. Do the plan first (21.1), then sign the agreement with your first super-user (21.2), and last lay the cadence into the calendar (21.3).
> License: Every template in this book may be modified freely and used in your work, no attribution needed.
## 21.1 Adoption Plan Template (One Row Per Mechanism)
**Usage**: fill it in before launch, and none of the four rows may be blank. Every mechanism needs a named owner and a moment in time. A mechanism with no owner is not a mechanism, it is a wish. If you cannot fill in the "goal" column, go back and reread that mechanism in Chapter 21.
| Mechanism | Goal (Verifiable) | Action | Owner | When |
|------|---------------|------|-------|------|
| Super-user network | | | | |
| Operating cadence | | | | |
| Short-term win announcement | | | | |
| Institutional anchoring | | | | |
**Anchor & Helm worked example**:
| Mechanism | Goal (Verifiable) | Action | Owner | When |
|------|---------------|------|-------|------|
| Super-user network | โฅ1 named super-user before the expansion; โฅ1 contact point per team after it | Linda Marsh named "queue co-builder", credited by name on the five published rules, lead instructor for the expansion training; the survey team lead in an observation period | Kevin Doyle (naming rights sit with the business, not with the project side) | Week 19, the day the expansion starts |
| Operating cadence | 100% attendance at the weekly retrospective (team lead level) | Embedded in the last 30 minutes of Kevin's weekly, reason code distribution / top aging claims / status-in-doubt list / one improvement; chairing handed to Linda | Linda Marsh (chairing), Kevin Doyle (attendance discipline) | From pilot week 2; chairing handed over in week 19 |
| Short-term win announcement | One business win worth telling inside the first month | Team two's backlog cleared, announced at the monthly business review, the date it hit zero plus the drop in first-touch handling time, with team two as the subject | Kevin Doyle (the announcer must be the business owner) | The week 23 monthly meeting |
| Institutional anchoring | New reviewer onboarding material covers queue operation; the exception claim SOP references the system | Onboarding gets a one-page queue operation section plus the five rules; the SOP's exception claim section rewritten to "claim it from the queue" | Kevin Doyle (SOP), Linda Marsh (content of the operation page) | Weeks 23 to 24, settled as the expansion goes |
**Three checks**: how many times does your own team appear among the four owners? More than once and adoption is still growing on you, and the handoff will go wrong. Which of the four already has a hook in the business side's existing institutions? If not one of them does, do not launch yet. What do you check first when daily actives drop? The order is fixed, unsafe trail โ manager's attendance โ features.
## 21.2 Super-user Agreement
**Usage**: one page, walked through face to face with every super-user, with the business owner in the room. This is not an employment document. It makes identity and reward explicit, and it guards against the hidden exploitation of "take the feeding, give nothing back."
### Selection Criteria (All Must Hold to Be a Candidate)
- [ ] **Influence test**: a hard claim comes in and he is the one colleagues get up to go ask. Pick by that question, not by rank
- [ ] **Actual user**: the person whose daily work the system changes, not his supervisor
- [ ] **Dares to say unsafe**: has a record of rejecting the system's output to its face (Anchor & Helm: Linda Marsh's three unsafe marks in Chapter 0; the survey team lead's challenge in week 20). Someone who only says nice things cannot feed a good system
- [ ] **Voluntary**: an appointed champion is not a champion
### Privileges (What You Give)
| Privilege | What It Is |
|------|------|
| First use | New features and new rule versions two weeks before everyone else, and his view can veto the release cadence |
| Direct channel | Improvement suggestions go straight to the development board, skipping the ticket process; every one gets a reply inside two weeks; reserve a fixed response capacity allowance in your own schedule before you say this privilege out loud |
| Chairing | The queue retrospective is chaired by the super-user, not by the project side |
### Identity (What Is Made Public)
- A formal name (Anchor & Helm: "queue co-builder"), granted by the business owner in front of everyone. Naming rights sit with the business side, not with the project side
- Personal contributions credited by name in public (Anchor & Helm: the five rules signed by Linda Marsh). The self-esteem loop has to close. Promoting the system = promoting your own work
- The super-user takes the training instructor role. A peer's testimony beats the project side's pitch
### Obligations (What You Get Back)
- The feeding-period commitment. Keep using the system at its dumbest, fill in reason codes seriously. Trail quality is how the system gets paid back
- The time spent chairing the retrospective and answering new teams' questions goes into his workload, on the precedent of the business-side commitment line in the project approval memo or the charter (Anchor & Helm: "2 hours a week"); his supervisor confirms in writing how many hours a week, or it goes into the super-user's OKRs for the quarter. No free rides
### Exit Conditions (Written Down in Advance)
- Voluntary exit, at any time, no reason required. The identity is an honor, not a shackle. Ask once for him to stay, and apply no pressure
- Four straight weeks of not using the system or not attending the retrospective voids the status automatically (against a title running empty)
- Red line. Using data visibility to appraise or pressure colleagues means immediate removal. The three visibility pledges (Template 20.3) apply to super-users too
## 21.3 Operating Cadence Sheet
**Usage**: one row per tier of ritual, laid into the calendar. The "Embedded in an Existing Meeting" column is the soul of this sheet. Embed wherever you can, and creating a new ritual takes a written explanation of why no existing ritual can hold it.
| Frequency | Ritual | Agenda (Time-boxed) | Chair | Embedded in an Existing Meeting |
|------|------|-------------|------|-------------|
| Weekly | Queue retrospective (30 minutes) | Reason code distribution / top aging claims / status-in-doubt list / settle one improvement | Super-user | โ Embedded in the tail of the business owner's existing weekly (Anchor & Helm: Kevin's weekly) |
| Biweekly | Rotating release (Chapter 15) | The claims-ops IT engineers demonstrate progress to the owner, the three explain-it questions | Co-build engineers | โ Aligned to the same week as the retrospective, the heavier of the two alternating week by week (Anchor & Helm went weekly after the expansion, Chapter 22) |
| Monthly | Metrics and win announcement | Run chart of the North Star plus balancing metrics (Chapter 18) / short-term win announced (business language) / per-team review of the adoption curves | Business owner | โ Embedded in the one-page slot of the monthly business review |
**Two red lines**: the retrospective crowded out twice in a row by "something more important" = an adoption warning, handled the same way as a drop in daily actives; the subject of the monthly announcement must be the business team. The project side has no name on this page.
**Handoff marker** (used in Chapter 22): the day all three tiers of ritual are chaired by the business side is the day the operating cadence handoff is complete.
## Code Hooks
The companion repo provides (this repository's `repo/` directory):
- [`templates/adoption/`](https://github.com/hallieren/the-last-mile/tree/main/repo/templates/adoption/): sample weekly query for per-team daily actives and depth of use (daily actives = the share of users who took a Human Call action that day; depth = override rate, reason code fill rate, share of "other" codes, the last two being the detector for failure mode 2's going-through-the-motions compliance)
- [`templates/adoption/`](https://github.com/hallieren/the-last-mile/tree/main/repo/templates/adoption/): sample script that generates the retrospective one-pager automatically (reason code distribution plus top aging claims, taken from the trail table in Template 17.2)
---
# Template 22 ยท Handoff Plan, Five Self-Sufficiency Tests, Handoff Cadence Sheet
> Companion chapter(s): Chapter 22. The three tools are in order of use. Fill in the plan template (22.1) when the handoff starts, because until the owners are named, the two sheets after it have no subject. Run the test sheet (22.2) item by item in stage three. Lay out the cadence sheet (22.3) from handoff start to the closing of the response window. It picks up from the "handoff checklist" at the last beat of Template 15's cadence sheet.
> License: Every template in this book may be modified freely and used in your work, no attribution needed.
---
## 22.1 Handoff Plan Template
### 22.1.0 Plan Header
| Item | Fill In |
|----|------|
| Project / system | |
| Handoff start date / target exit date | |
| Response window (length and scope, terms in 22.3) | |
| Signatures (four owners + the delivering team's lead) | |
| Transfer ledger entry number (PMO AI system transfer ledger) | |
### 22.1.1 Owner Placement Sheet
Rules: one real name per row. One person may hold several rows, but each row holds only one name. "The team is responsible" means nobody is responsible.
**Blank template**:
| Role | Responsibility | Real Name |
|------|------|------|
| System owner | Budget and priorities, the four-ledger running budget (Chapter 16), request scheduling, first signature on whether to stop | |
| Maintenance owner | Code and production, gatekeeping rights over every module, releases, first-line incident response | |
| Eval guardian | Approving new golden cases, the weekly review, annotation guide version control | |
| AI dependency owner | Model and platform changes, evaluating and signing off on provider model version upgrades, dependency library changes, and platform migrations | |
**Anchor & Helm example**:
| Role | Anchor & Helm Placement |
|------|----------|
| System owner | Kevin Doyle (from the week 22 budget cycle) |
| Maintenance owner | The two claims-ops IT engineers hold gatekeeping rights by module (one person per module; every module on the business line side from week 24) |
| Eval guardian | Linda Marsh (from week 23) |
| AI dependency owner | The younger engineer runs the replay, Linda Marsh (quality) and Kevin Doyle (cost) sign |
### 22.1.2 Three AI Handoff Items Owner Sheet
The vocabulary of a traditional IT handoff does not have these three. They are the easiest to miss and the most fatal. Name an owner for each, no blank rows allowed:
**Blank template**:
| Item | Question to Answer | Receiver Type (Business / Ops / Platform) | Real Name |
|----|--------------|------------------------------|------|
| Continuous eval updating | Who adds golden cases, who approves them? Who guards the per-category thresholds? | [business / ops / platform] | |
| Re-verification on model and dependency changes | On a model version switch or dependency upgrade, who runs the golden cases replay? Who signs "safe to switch"? | [business / ops / platform] | |
| Decision trail review | Who chairs the override review? Who analyzes the reason code distribution? Who runs the fairness spot check? | [business / ops / platform] | |
**Anchor & Helm example**:
| Item | Anchor & Helm Owner | Receiver Type |
|----|----------|------------|
| Continuous eval updating | Adding: review flow-back + the review team; approving: Linda Marsh | Business |
| Re-verification on model and dependency changes | Runs the replay: the younger engineer; signs: Linda Marsh (quality) + Kevin Doyle (cost) | Business |
| Decision trail review | Chairs: Linda Marsh; analyzes: the older engineer; the quarterly fairness spot check likewise (Chapters 12/18) | Business |
---
## 22.2 Five Self-Sufficiency Tests Sheet
Rules: the test method is always **Show Me**, you are in the room and do not step in; inject the scenario into the real system wherever possible; any failure sends the item back for rework, and what gets added is a drill, not a document; only when all five pass does stage four of the cadence sheet begin.
| Capability | Show Me Scenario | Pass Standard | Rework Action on Failure |
|------|----------------|----------|--------------|
| Run | Cut one AI dependency point without warning | The business side degrades, notifies by the template, and recovers on its own, with no instruction from the deliverer the whole way | Rerun the degradation drill (Template 16.3), add an incident notice drill |
| Configure | One real configuration change (list / threshold / rule parameter) | The full change, test, release cycle runs on the business side's gatekeeping rights, and the change passes the eval replay | Walk it again as a pair, check configuration item permissions and documents |
| Exceptions | Inject a class of situation the system has never seen | The first reaction is the escalation path, not the deliverer; the ruling gets settled somewhere (into the golden cases or a rule) | Draw the escalation tree together + two injection drills, then retest |
| Evolve | One small request end to end + one model version upgrade | From scheduling to rotating release with no commit from the deliverer; the upgrade passes the golden cases replay and gets both signatures | Back to co-build pairing (Chapter 15), shrink the request and retest |
| Teach | The business side onboards one newcomer | The newcomer handles the queue independently within two weeks; the training material is maintained by the business side itself | Have the newcomer recount where they got stuck. Fix the mentoring path, not the manual |
**Anchor & Helm example (the "exceptions" failure, as it happened)**: Week 25, the rainstorm batch (dozens of interrelated auto damage claims entering the queue at once, identical missing documents, tangled risk signals), and the claims operations team's first reaction is to call you โ recorded as **failed**. Two weeks of rework: escalation tree (check the decision trail and drift signals first โ Linda Marsh judges โ beyond scope, escalate to Kevin Doyle, with that class of claims paused and routed to manual handling if needed) + two injection drills. Retest in week 27 passes. The ruling routes to manual, the incident case goes into the golden cases the same day, the phone does not ring. Note that the failure is not an incident. It is the test doing its job. The value of a test is not in passing. It is in exposing.
---
## 22.3 Handoff Cadence Sheet
| Stage | Who Chairs the Retrospective | Who Touches Production | Who Answers Outside Questions | Condition for the Next Stage |
|------|--------------|----------|--------------|--------------------|
| 1 Your team leads (business side observes) | Your team | Your team | Your team | Co-build agreement in force, gatekeeping rights of the first module handed over (Chapter 15) |
| 2 Shared lead | Business side chairs, your team adds | Both, gatekeeping rights handed over module by module | Business side answers, your team backstops | Gatekeeping rights of every module on the business line side, four consecutive retrospectives chaired by the business side |
| 3 Business side leads (your team advises) | Business side | Business side (your team no longer commits) | Business side | All five self-sufficiency tests passed (22.2) |
| 4 Stepping out of the daily | Business side | Business side | Business side | Response window closed with written confirmation, write access reduced to read-only, no open rework items |
**Response window terms** (written into the responsibility transfer agreement):
- Length: 3 months recommended, the period when the new team's confidence is most fragile.
- Scope: only two kinds are taken, P1 incidents (an unsafe reaching human eyes, scale in Template 18.2.1) and exceptions the escalation tree ran to the end and did not catch. Daily ops and routine requests are outside the window. Taking one means falling back to stage three.
- Record: log every request for help inside the window; review once before the window closes, which capability each request pointed to, and whether to add one targeted drill.
- On expiry: write access drops to read-only, removal from the oncall rotation and the retrospective's standing attendee list (the design of co-build clause one is honored here, Chapter 15).
- Closing the window: needs written confirmation from the business owner and the ops owner. No confirmation means extension by default, and extension by default means permanent ops.
---
## 22.4 Counterexample: A Tidy-Looking Wrong Answer
An excerpt from a handoff completion report, and it looks fine at the retrospective:
```
Handoff deliverables: ops manual (52 pages), architecture diagram, code repository and account transfer form,
2 training sessions (sign-in photos on file), 8 screen recordings of operations.
Owner placement: system owner: Kevin Doyle; maintenance owner: the business line engineering team;
eval guardian: to be decided by the business side after handoff.
Five self-sufficiency tests: covered by documentation, training, and Q&A, deemed passed.
Response window: we are right here in the company, come find us anytime.
```
Line by line:
1. The deliverables list is all nouns (manual, accounts, code, recordings) and not one verb (can respond, can judge, can evolve). Chapter 22's test is exactly this ratio. Documents are the shadow of capability, and a shadow cannot hold the system up.
2. "The business line engineering team" breaks the placement sheet's first rule. Each row holds only one name, and team responsible equals nobody responsible.
3. Eval guardian "to be decided after handoff" = a blank row on the three AI handoff items owner sheet. The easiest to miss and most fatal item really was missed. Nobody approves golden cases, and the eval starts rotting from handoff day.
4. "Training deemed passed" swaps out 22.2's test method. All five capabilities are Show Me, cut a dependency, inject an exception, run a real change end to end. Having heard the lesson and being able to catch it are two different things, and Anchor & Helm's rainstorm batch in week 25 is the evidence.
5. "We are right here in the company, come find us anytime" looks considerate and is in fact never exiting. The internal team was always in the company, and nobody comes to revoke your access. The response window must state a length (3 months recommended), a scope (only P1 and exceptions the escalation tree did not catch), and written confirmation of closing. Missing any one is extension by default, and extension by default is permanent ops.
---
## Code Hooks
The companion repo provides (this repository's `repo/` directory):
- [`templates/handoff/checklist.yaml`](https://github.com/hallieren/the-last-mile/tree/main/repo/templates/handoff/checklist.yaml): sample machine-readable handoff checklist, validates the named-owner fields for the four owners and the three AI items (a blank name raises an alert)
- [`templates/handoff/`](https://github.com/hallieren/the-last-mile/tree/main/repo/templates/handoff/): five self-sufficiency tests state machine (pending / failed / passed, failed must carry a rework action and a retest date)
- [`templates/handoff/`](https://github.com/hallieren/the-last-mile/tree/main/repo/templates/handoff/): response window countdown and help request log template
---
# Template 23 ยท Pattern Extraction Sheet, Asset Register, Library Admission Checklist
> Companion chapter(s): Chapter 23. The three tools are arranged in flywheel order. At the closeout retrospective, use the extraction sheet to turn hero stories into candidate patterns (23.1), register them once they pass review (23.2), and let the review checklist keep the door (23.3).
> License: Every template in this book may be modified freely and used in your work, no attribution needed.
## 23.1 Pattern Extraction Sheet
One page per candidate pattern, six fields.
**Blank template**:
| Field | What to Fill In |
|----|---------|
| **Phenomenon** | Which project, which scene, at what frequency. Record it in the business side's own words as they were said, do not abstract yet |
| **Strip the field context** | Swap the business side's nouns for structural nouns, and state the transferable judgment in one sentence |
| **Applies when** | What structural features have to be present for it to be usable |
| **Does not apply when** | Where using it breaks. Required. Cannot write it, cannot admit it |
| **Field provenance** | Project, date, link to the validation record, who was there |
| **owner** | Real name plus date claimed. The owner is responsible for revisions and the quarterly recheck |
**Anchor & Helm example** (using "chase โ risk"):
| Field | Anchor & Helm's Entry |
|----|----------|
| **Phenomenon** | The Anchor & Helm Field MVP. All 3 unsafe items Linda marked came from the system treating "chased many times" as a high-priority signal |
| **Strip the field context** | In any ranking scenario where the people being served can apply pressure, pressure intensity and business risk are two independent variables, and they must be modeled separately and shown separately |
| **Applies when** | Queue ranking scenarios with an outside party applying pressure (chasing a claim, chasing an assignment, chasing an order) |
| **Does not apply when** | Scenarios where the party applying the pressure is the risk (safety incident reporting, for example, where the most urgent caller is often the most dangerous case) |
| **Field provenance** | The week 1 Anchor & Helm MVP scoring sheet (Linda's original annotations). The week 31 re-verification at Swiftway (drivers chasing assignments โ waybill risk) |
| **owner** | [name] |
**The three stripping steps** (the operational definition of the generalize step):
1. **Swap the nouns**: exception claim โ queue item, reviewer โ handler, repair shop โ outside party;
2. **Drop the numbers**: -31% is Anchor & Helm's result, not a property of the pattern. Numbers go into "field provenance," not into the judgment;
3. **Keep the structure**: after stripping, read it out to a colleague who never worked that project. If he can repeat back when to use it and when not to, the stripping is done.
## 23.2 Asset Register
Nine columns: class / name / one-line description / source project / validation count / owner / last updated / status (active, pending-reverify, retired) / force level (reference, recommended, mandatory, and with nothing declared it is taken as mandatory). The upstream entrance to this sheet is Template 24.2, the closeout asset recovery checklist, and the field mapping between the two sheets is in the last table of 24.2.
**Anchor & Helm's first batch admitted** (first admitted at the internal closeout retrospective in week 30. The validation counts below are a snapshot updated in week 34, covering the home property business side's own build and the items already field-tested at Swiftway. Swiftway's queue has not launched, so the queue-class components' reuse at Swiftway is not counted yet and gets added after launch. The sample table drops the "last updated" and "force level" columns):
| Class | Name | In One Line | Source | Validations | owner | Status |
|------|------|-----------|------|------|-------|------|
| Component | Queue skeleton | Column structure = the projection of the four layers of the decision rights boundary, the decide layer gets no column | Anchor & Helm | 2* | [name] | active |
| Component | Decision trail schema | Suggestion and fact tables kept separate, append-only, reason codes enumerated | Anchor & Helm | 2* | [name] | active |
| Component | Extraction pipeline interface | Unstructured input in, schema out, validation, spot-check tool. The interface is independent of the business domain | Anchor & Helm | 2* | [name] | active |
| Template | Three-way reconciliation checklist | System field / private source of truth / asking the handler, producing the inconsistency rate and pattern | Anchor & Helm | 2 | [name] | active |
| Template | Four adoption mechanisms plan | super-user / operating cadence / announcing wins / institutional anchoring | Anchor & Helm | 2* | [name] | active |
| Template | Three-layer probing script | anchor on an instance โ compare โ boundary counterexample | Anchor & Helm | 3* | [name] | active |
| Judgment rule | "Chase priority โ risk priority" | Pressure intensity and business risk modeled separately | Anchor & Helm | 2 | [name] | active |
| Judgment rule | "If you cannot say it, do not merge it." | AI-written code merges only after the maintaining side passes the explanation | Anchor & Helm | 2 | [name] | active |
| Metric model | Metric tree three-tier structure | North Star (outcome) / Process (mechanism) / Balancing (cost), every metric carrying an owner and an action | Anchor & Helm | 2* | [name] | active |
| Metric model | Error severity vocabulary | pass/concern/unsafe/useless plus the per-category threshold structure | Anchor & Helm | 2* | [name] | active |
\* Starred entries have one validation carried out by the business side (Anchor & Helm's home property team built it themselves). **The reuser does not have to be you. A third party's reuse counts as validation too, and it is the highest-grade kind.**
**Maintenance rules**:
- Entries with a validation count of 1 are marked "awaiting a second field test" and are down-weighted when a coding agent packs them;
- Every owner sweeps the assets in his name once a quarter. Anything not updated in over six months is automatically demoted to "pending-reverify" (the mechanized form of "an asset with no owner rots in six months");
- Retirement is not deletion. A retired entry keeps its original text and states why it was retired (the mirror of failure mode 4, a dead pattern is teaching material too).
## 23.3 Library Admission Checklist
When to review, in the same session as the closeout retrospective, no separate meeting (capture comes with its own review). Reviewers, the library owner plus one peer deliverer **who did not work on that project**. A pattern an outsider cannot read is a pattern whose stripping is not finished.
**The three criteria, item by item**:
- [ ] **Field validation**: at least one record of real use on the business side's floor. A demo, an internal team rehearsal, or a sandbox test does not count
- [ ] The validation record is linkable and checkable (scoring sheet, run chart, retrospective notes, decision trail data)
- [ ] **Boundary**: both the applies-when and the does-not-apply-when fields are filled in
- [ ] At least one does-not-apply condition comes from a real "we used it and it did not work" or from serious reasoning, not from boilerplate
- [ ] **owner**: claimed by name, with the owner knowing and agreeing (assignment without consent is not a claim)
- [ ] The owner can state the next expected reuse scenario for this asset
**AI pollution check** (any one unanswered and it is refused):
- [ ] Can you say which project, which weeks, who was in the room?
- [ ] Is there someone who can verify it? When the proposer was the only person present, it needs an endorsement from the business side or a teammate at the time
- [ ] Is there a failure record or a does-not-apply condition? A pattern with only upsides is a promotional piece, not an asset
- [ ] Does the content carry any "industry best practice" statement that cannot be traced back to a field? If so, delete it line by line or add the provenance
**Where AI's share stops**: AI may draft, structure, rewrite the language, and fill in formatting to the template. AI does the grunt work. People make the calls (Chapter 0's division of labor, recurring at library scale). AI may not supply "field provenance" or "validation record." Those two fields can only be filled in by someone who was on site.
## Code Hooks
The companion repo provides (this repository's `repo/` directory):
- [`templates/pattern-library/`](https://github.com/hallieren/the-last-mile/tree/main/repo/templates/pattern-library/): asset register structure (table and machine-readable versions), sample script that automatically demotes anything not updated in six months to "pending-reverify"
- [`templates/pattern-library/`](https://github.com/hallieren/the-last-mile/tree/main/repo/templates/pattern-library/): fillable Pattern Extraction Sheet (includes the three-step stripping check)
- [`templates/pattern-library/`](https://github.com/hallieren/the-last-mile/tree/main/repo/templates/pattern-library/): packing sample for feeding a component-class asset to a coding agent, the asset itself plus its boundary plus its validation record traveling with the package
---
# Template 24 ยท Field-to-Product Memo and Closeout Asset Recovery Checklist
> Companion chapter(s): Chapter 24. The two tools are arranged in the direction the signal flows. At the asset inventory, run the recovery checklist (24.2) first to count the assets, then take the platform-level candidates it turns up across the river with the memo template (24.1). A signal that trips the three filters mid-project can go straight to 24.1.
> License: Every template in this book may be modified freely and used in your work, no attribution needed.
## 24.1 Field-to-Product Memo Template
**Self-check before sending (the three filters)**, fail any row and it does not go out:
- [ ] **nโฅ2**: at least two business departments or projects have hit it, and n is verifiable. n=1 โ into the candidate register (24.1.3), do not send.
- [ ] **Judgment structure**: what you describe is a transferable judgment structure, not one department's interface habit ("waiting attributed across steps" passes, "a Gantt chart" does not).
- [ ] **The give-up line**: name what you are willing to give up or yield for it. Cannot write it โ you do not much believe in it yourself.
**Format rules**: obey the four one-page rules of Chapter 19, one page of body text, lead with the conclusion, an ask at the end, delivered 48 hours before the meeting. The reader changes from an executive on the business side to the Group platform lead. The rules do not.
### 24.1.1 The Four Sections
> **F2P Memo #[number]: [the capability name in one line].** To: [the platform lead], From: [you], [date]
>
> **Phenomenon**: which field, at what frequency, on what evidence. Quote the actual words, the decision trail data, the engineering effort records, not impressions.
>
> **The generalization case**: which other departments or scenes will hit this, and on what grounds. Write backward from the side with the need (the *Working Backwards* discipline), not forward from "what I built." State **what n is**, the name behind each n, and whether the nth implementation needed any industry adaptation.
>
> **The product recommendation (a capability, not a feature)**: what **capability** you are asking the platform to provide. Self-check, rewrite it as "add a [button / page] for [Claims Operations]" and see whether it still holds. If it does, it is a feature. Rewrite it.
>
> **The cost of not doing it**: the small bill (the engineering effort every project repeats) + the big bill (what is lost structurally, consistency, aggregation analysis, response speed).
>
> **Willing to give up for it**: one sentence.
> **What I need from you**: an ask with a date on it (a scheduling review / 30 minutes in person / a written reply).
### 24.1.2 A Worked Example (Anchor & Helm, the Decision Trail Component)
> **F2P Memo #1: the decision trail component.** To: the platform lead, From: [you], week 32
>
> **Phenomenon**: projects at two subsidiaries, Anchor & Helm (insurance claims) and Swiftway (logistics dispatch), hand-wrote the same trail layer one after the other. Suggestions and facts stored in separate tables, the suggestion table append-only, the six end-to-end decision trail fields, override reason codes (an enumeration plus an optional note). About one week of engineering each time, and the structures are near identical.
>
> **The generalization case**: the decision trail is a structural requirement of any "AI suggests, a person decides" system, and it has nothing to do with the industry. As long as decision rights stop at the advise layer, the system has to answer "who made the call and on what grounds." n=2 (the two subsidiaries Anchor & Helm and Swiftway), and the second implementation needed no industry adaptation at all.
>
> **The product recommendation (a capability)**: a decision trail component (the trail schema + reason code configuration + a review query that splits overrides by category), available to any project out of the box. (Counterexample self-check, "add a decision trail export button for Claims Operations" is a feature, not this proposal.)
>
> **The cost of not doing it**: the small bill, about a week of rewriting on every new project. The big bill, project schemas diverge and cross-department override aggregation analysis cannot be built at all. Both subsidiaries' decision trails sit in the Group warehouse. An outside team can never do this, we could have, and it is the only source of platform-level eval insight.
>
> **Willing to give up for it**: if the schedule conflicts, this team sends one person into the platform repo to co-build, and withdraws its request this quarter for custom columns in the queue interface.
> **What I need from you**: a written reply before the next platform scheduling meeting ([date]), on the roadmap / refused (with the reason, which this team records in the register).
### 24.1.3 The Candidate Register (Where n=1 and Refused Memos Are Kept)
| Signal | Source (Department / Proposer / Date) | n | Status | Trigger / Revival Condition |
|------|--------------------------|---|------|----------------|
| Repair shop response time as a metric | Anchor & Helm / the survey team lead / week 23 | 1 | candidate | Any second project that raises an "outside party response time" need promotes this to a memo |
| The queue scaffold | Anchor & Helm + Swiftway / you / after week 32 (memo #3) | 2 | refused, recorded (conflicts with the existing roadmap) | Raise it again when the platform roadmap touches a rework of queue shape |
Three rules. n=1 is registered, not reported. A refused memo goes into the table together with the reason for refusal, and its revival condition is written as events, not dates (the same form as Chapter 8's scope decision log). Go through the table once a quarter. **A memo that was refused but recorded is still an asset**.
## 24.2 Closeout Asset Recovery Checklist
**Three triggers**, the fixed retrospective N weeks after launch / before anyone moves posts / the quarterly asset inventory. Hang it on any one of them and recovery has a date, and past that window it does not happen by default. **Guiding principle**, the code may sit in the same git, but structure that was never declared is the same as structure that does not exist. Code files belong to the business line repo (Chapter 15). Schema design, the error taxonomy, and the key judgments belong to no repo at all. The four classes correspond to Chapter 23's four asset classes (component / template / judgment rule / metric model), ordered here the way an asset inventory counts them.
**Three questions per item**. Is it general (does it still hold in another industry)? What is n (how many projects have hit it)? Destination, **admit** (through Chapter 23's admission review, registered into the Template 23 asset register, must have an owner) / **send a memo** (nโฅ2 and at platform capability level, through 24.1) / **leave it in the business line repo** (business-line-specific, with one line on why it is not recovered). **The first two exits are not mutually exclusive.** A component at platform capability level usually goes into the library and gets reported as well (a refused report still stays in the library), and only "leave it in the business line repo" is an exclusive exit.
| Class | What Goes Through, Item by Item | General? | n= | Destination | owner |
|----|--------------|--------|----|------|-------|
| **Code** | Decision trail schema / eval harness / queue skeleton / extraction pipeline structure / configuration patterns | | | | |
| **Documents** | charter / review packet / the three memos / runbook / annotation guide **framework** (the content does not transfer) | | | | |
| **Judgment** | Key judgments / failure modes and incident retrospectives / annotation disagreement rulings as precedent | | | | |
| **Metrics** | Metric definition tree / thresholds and acceptance lines / reason code enumeration / balancing metrics | | | | |
**Anchor & Helm's sample rows** (what actually happened in Chapter 24, including the note added after the asset inventory, memo #1 went out in week 32):
| Asset | Class | General? | n= | Destination |
|------|----|--------|----|------|
| Decision trail schema | Code | Yes (a structural requirement) | 2 | Admit + send a memo (#1) |
| eval harness | Code | Yes | 2 | Send a memo (#2) |
| Queue skeleton | Code | Yes | 2 | Send a memo (#3, refused โ the register). Admitted at the same time (Template 23) |
| The five rules of thumb | Judgment | No. Tacit knowledge does not transfer. The method for mining it transfers (Chapter 23) | 1 | Stays with the business side. The three-layer probing method is admitted |
| "Chase priority โ risk priority" | Judgment | Yes (a key judgment transfers) | 1 | Admitted (Template 23. Validation count 1 at the asset inventory, marked "awaiting a second field test." Raised to 2 after the week 31 re-verification at Swiftway) |
| Reason code enumeration (the Anchor & Helm version) | Metrics | The structure is general, the vocabulary is business-line-specific | 2 | The structure rides along with memo #1. The vocabulary stays with the business side |
**Three closing checks**:
- [ ] Every item has a destination, and "leave it for now" is not allowed. A count with no exits is the same as no count.
- [ ] Every admitted item has an owner (Chapter 23, an asset with no owner rots in six months).
- [ ] Every n=1 candidate has its trigger condition written down (24.1.3).
**Field mapping, the 24.2 recovery checklist โ the 23.2 asset register** (the bridge across the river for the "admit" exit. The counting finishes in this sheet and the registration lands in Template 23.2, which keeps two registration systems from growing apart):
| 24.2 Recovery Checklist Field | 23.2 Asset Register Field | How It Maps |
|------------------|--------------------|----------|
| Class (code / documents / judgment / metrics) | Class (component / template / judgment rule / metric model) | One to one (Chapter 23's four asset classes), code โ component, documents โ template, judgment โ judgment rule, metrics โ metric model |
| What goes through, item by item | Name + in one line | The one-line description is filled in at the admission review (23.3) |
| General? | / (not in the register) | A test field. Answer "no" and it takes the "leave it in the business line repo" exit, producing no register row |
| n= | Validation count | The same count. After registration it keeps accumulating with reuse. n=1 is marked "awaiting a second field test" |
| Destination | / (not in the register) | "Admit" generates one register row. "Send a memo" goes through 24.1 separately, and the two exits are not mutually exclusive |
| owner | owner | The same person. Claimed at the recovery count, confirmed at the admission review as knowing and agreeing (assignment in absentia does not count) |
| / (the recovery checklist has no such three columns) | Source project / Last updated / Status | Filled in at admission. Source = the project the asset inventory sits in. Last updated = the admission date. Status defaults to active |
## Code Hooks
The companion repo provides (this repository's `repo/` directory):
- [`templates/f2p-memo/`](https://github.com/hallieren/the-last-mile/tree/main/repo/templates/f2p-memo/): the f2p memo Markdown template plus the candidate register CSV structure
- [`templates/asset-recovery/`](https://github.com/hallieren/the-last-mile/tree/main/repo/templates/asset-recovery/): the closeout asset recovery checklist tables plus the field alignment notes for the Template 23 register
- [`templates/f2p-memo/override-compare.sql`](https://github.com/hallieren/the-last-mile/blob/main/repo/templates/f2p-memo/override-compare.sql): sample query comparing override reason code distributions across departments (the minimal implementation of platform-level eval insight)
---
# Template 25 ยท Intake Rubric, Red Line List, Saying No Scripts, Kill Register
> Companion chapter(s): Chapter 25. The four tools are in order of use. A candidate project clears the red lines first (25.2, any one vetoes), then gets scored (25.1). When the ruling is no, use the conversation script (25.3). Once it is said, it lands in the register (25.4) for a retrospective a year later.
> License: Every template in this book may be modified freely and used in your work, no attribution needed.
## 25.1 Organization-Level Intake Rubric (Five-Dimension Scoring Table)
**Relation to the Chapter 7 five-question sheet.** The five questions assess "is this use case any good" (the deliverer's personal tool in the field), and the intake rubric assesses "can this organization carry it, is it worth carrying" (the gate at the team's door). The five questions' Risk column is promoted here into a red line (25.2) and no longer gets scored.
**Scoring rules.** 1โ5, weakest-link logic (same as the Chapter 7 elimination table). Any dimension โค2 does not enter the schedule, and a re-evaluation condition gets written. A 3 is a conditional pass, with the condition and the deadline written into the sheet. No weighting, no averaging. The evidence column takes only numbers, users' own words, and samples.
**Where the gate's force comes from** (fill this in first, or the rubric is only talking to itself; pick one of the three and name the carrier it hangs on)
- Hang it on the company's existing project approval review / architecture review / quarterly ask review (name the meeting)
- Endorsed once by the CIO or the responsible head
- Written into the company's AI governance policy
| Dimension | Test Question | What a 1 Looks Like | What a 5 Looks Like |
|------|----------|-----------|-----------|
| Strategic value | Win it, and what does the organization get? One department's thanks, or agenda power and standard-setting power over a class of problem? | A single department's goodwill, with no organization-level benefit you can state | It opens a repeatable class of problem, and the ROI arithmetic has a business owner who stands behind it |
| Data readiness | How far do the key data exist, how reachable are they, how far can they be trusted? Have they been reconciled? | The key data do not exist, or there is no realistic path to getting them | Samples in hand, a reconciliation conclusion, a named data owner |
| Owner in place | Is there someone on the business side who signs for the outcome? Who backstops the errors, who maintains it after launch? | "Once it launches someone will look after it" | A named owner has been through the meeting and has claimed actions |
| Production path | Do review, oversight, and change discipline get through? Or can only the demo live? | No path through security or compliance review, no oversight headcount in existence | The review path is agreed, and the oversight role and its hours are committed |
| Reuse potential | What gets banked into the pattern library (Chapter 23)? How much cheaper is the second delivery? | Pure custom work, the judgment structure does not transfer | The skeleton transfers, and halving the cost of the second delivery has a basis |
**Note on owner in place.** No written agreement can force the business side to produce a signer, so forcing out a named owner runs on a resource trade. You want our people and our time, so first give us the person who signs for the outcome after launch, name into the project approval resolution. No name, and this dimension scores 1, and by weakest-link logic it does not enter the schedule.
**Scorecard** (one per candidate; a 3 row must fill the "Condition / Deadline" column, same convention as Template 7.2):
| Dimension | Score (1โ5) | Evidence (numbers / users' own words / samples) | Condition / Deadline (a 3 is a conditional pass, required) |
|------|-----------|---------------------------|----------------------------------------|
| Strategic value | | | |
| Data readiness | | | |
| Owner in place | | | |
| Production path | | | |
| Reuse potential | | | |
**Anchor & Helm example, full automation (loss assessment and payout approval), week 28**
| Dimension | Score | Evidence | Condition / Deadline |
|------|-----|------|-------------|
| Strategic value | 5 | The only path to another notch off claims cost; a board agenda item | / |
| Data readiness | 2 | Loss assessment data (repair hours / parts prices / image reading) not reconciled; what the queue reconciled is only process status data | / (โค2 does not enter the schedule, write a re-evaluation condition) |
| Owner in place | 2 | "Who signs for a loss the AI set" went unclaimed | / (same) |
| Production path | 1 | The regulator's line is not out, and Victor Reyes (Head of IT Security and Architecture) is clear that the review does not pass | / (same) |
| Reuse potential | 4 | The estimate assist component reuses across lines of business | / |
> Note. This example never really gets as far as scoring. Red lines 1 and 3 are already triggered (see 25.2), and the veto comes first. Scoring it anyway has exactly one use, seeing which dimensions the alternative path has to rebuild (data readiness โ reconcile the loss assessment data; owner โ land a payout signer; production path โ wait for the regulator's line, build the oversight capacity). The steps get built up from the low-scoring dimensions.
## 25.2 Red Line List
| # | Red Line | Test Question (One Line) | Anchor & Helm Instance |
|---|------|------------------|----------|
| 1 | Automated decisions at an irreversible-harm step | Can the harm from the single worst output be taken back? | "Never touch payout decisions" (Chapter 8) is this one's instance. A payout decision is irreversible the moment it takes effect |
| 2 | A move of decision rights with no human oversight capacity behind it | After the move, is there still someone with the time to look, the ability to judge, and the authority to stop it? (the three oversight questions, Chapter 12) | With the payout approval team having no spare capacity to review every AI loss estimate, raising "assist" to "automatic" crosses the line |
| 3 | Outward-facing output in a regulatory grey zone | When that output causes a dispute, what do you answer the regulator with? | An insurer's outward statements are regulated. Nothing goes outward before the line on AI payout approval is clear |
**The three rules of "red lines do not enter the scoring sheet"**:
1. Any one vetoes. Not scored, not weighted, not compensated by any high score. The meaning of a scoring sheet is trade-off. The meaning of a red line is boundary.
2. Adding or removing a red line is not the project team's vote. It is re-argued only as the external source changes (the regulator gets clear, the oversight capacity gets built, a reversal mechanism appears), and the re-argument leaves a trail.
3. Keep the red lines few (three or so). Every extra fake red line that could have been scored dilutes the veto force of the real ones.
## 25.3 Saying No Scripts
> **A script is a crutch, not a line to recite.** Use your own words, but keep the function of every sentence (same convention as Template 20.2).
**Step one, affirm the goal** (declare that what you refuse is the path, not the intention)
- "The half sentence before the conclusion. You have not got the direction wrong, and this is the road we have been laying all along." (Function, catch the goal, so the refusal earns the right to begin)
- "We get to this goal sooner or later. What we are settling today is which road and when." (Function, switch the question from whether to do it to how to go)
**Step two, state the evidence** (the rubric goes on the table, read the evidence item by item)
- "This is the sheet we use for intake. Going through it, the evidence on [dimension] is [numbers / users' own words / samples]." (Function, let the sheet talk, you only read it out)
- How to put Risk. "The single worst output is [concrete scenario]. That account, [regulatory / brand / customer, whichever the other side cares about most], you know better than I do." (Function, put risk in the other side's ledger, not in model probabilities)
- The banned lines. "The model will make mistakes." "AI is not mature enough." Both demote a judgment question into a technical one, which invites the other side to rebut you with "then get a stronger model."
**Step three, offer a path** (a step plus conditions; an unconditional "no" is a refusal, a conditional "no" is a roadmap)
- "What I recommend is not dropping it, it is rebuilding the steps. First do [the assist form one layer down], with the condition nailed down. [Measurable threshold] met, [which class of decision] moves up one layer, each layer decided on its own." (Function, turn the "no" into a "yes" with milestones)
- "I am putting this into the register, and the revival condition is [an event, not a date] (how to write it, Template 8). When the day comes you will not have to raise it, I will." (Function, turn the promise into a trail anyone can check)
**Banned throughout**. "The schedule is full" (delay is not refusal) / "This cannot be done" (an unconditional "no") / a "policy does not allow it" that cannot name which policy (it outsources the judgment to the institution, and the other side goes looking for whoever can change the policy). Borrowing external hardness for a red line is fine. If you can name the regulation number or the policy clause, name it. A "policy does not allow it" that cannot say which clause is a shield, not a red line.
## 25.4 Kill Register
| Project | Proposer | Kill Reason | Revival Condition (events, not dates) | Retrospective Conclusion a Year Later |
|------|--------|-----------|---------------------------|----------------|
| Full automation (loss assessment and payout approval) | Grant Whitmore | Red lines 1 and 3 triggered; data readiness 2 | Estimate assist override rate <10% and the regulator's line clear, re-argued layer by layer | (to fill) |
| Service chatbot | The board (forwarded by Grant Whitmore) | Eliminated in Chapter 7, Data 1 / Risk 1 | Re-measure the call mix after exception handling is cured; the standard-answer library gets built | (to fill) |
**Retrospective discipline** (once a year, two questions per row):
1. Did it get built later, here or somewhere else? How did it go?
2. Did the revival condition come true? Once it did, did anyone re-score it?
**The fifth column's three conclusions**. The kill was right (built elsewhere, and it died of the cause predicted back then) / the kill was wrong (the condition came true long ago with nobody re-scoring, or the predicted cause of death never happened) / the evidence has changed (back through the rubric).
**The rubric's own eval** (the spirit of Chapter 11, the judgment that judges intake also has to be judged). Whichever dimension the wrong kills cluster on is the dimension whose test question you fix. A register that stays empty for a long time is also a signal. Not that every judgment was right, but that the gate is not working.
## Code Hooks
The companion repo provides (this repository's `repo/` directory):
- [`templates/intake/`](https://github.com/hallieren/the-last-mile/tree/main/repo/templates/intake/): table samples for the five-dimension scorecard and the red line checklist; kill register sample
- [`templates/intake/`](https://github.com/hallieren/the-last-mile/tree/main/repo/templates/intake/): lightweight script sample for revival condition due reminders (register scan plus event-triggered prompts)
---
# Field Template Library
> The master entry point to the book's appendices. The library holds 24 templates, numbered to match the chapter each one serves ("Template N goes with Chapter N"), grouped by the six Parts of the main text. Three merges: the Chapter 1 pre-mortem folds into Template 4 (4.3โ4.6), the Chapter 13 ADR memo folds into Template 19 (19.2), and the Chapter 26 F1โF5 anchors and growth agreement fold into Template 3 (3.7โ3.8). In each case the two halves were already one tool. Every template may be copied, modified, and used in your own projects. The companion code scaffolding lives in this repository's `repo/` directory (see `repo/README.md` for the mapping).
## Chapter to Template Map
| Chapter | Template | Contents | The Sheet Most Often Copied on Its Own |
|----|------|------|-------------------|
| 0 | Template 0 | Field MVP Pack | Scoring sheet (pass/concern/unsafe/useless) |
| 1 | โ Template 4 (4.3โ4.6) | Pre-mortem Memo | Reference library of common ways to die |
| 2 | Template 2 | Deliverer Role Charter | Role charter one-pager |
| 3 | Template 3 | Capability Self-Assessment (3.0โ3.6) + F1โF5 Rating and Growth Agreement (3.7โ3.8) | Five-axis self-assessment sheet |
| 4 | Template 4 | Deployment Charter (4.1โ4.2) + Pre-mortem (4.3โ4.6) | Charter seven-element one-pager |
| 5 | Template 5 | Stakeholder Map | Six-role map with names |
| 6 | Template 6 | Field Archaeology Kit | Three-layer probing script |
| 7 | Template 7 | Five-Question Opportunity Rubric | Five-question scorecard |
| 8 | Template 8 | Thin Slice Definition Sheet + Scope Decision Log | Scope decision log |
| 9 | Template 9 | Data Source Inventory + Fitness Scorecard + Reconciliation Checklist | Fitness scorecard |
| 10 | Template 10 | Pattern Decision Table + Anti-pattern List | The six pattern cards |
| 11 | Template 11 | Eval Spec + Golden Case Collection + Annotation Disagreement Log | Annotation disagreement log |
| 12 | Template 12 | Trust Constraint Matrix + Review Packet | Six-constraint matrix |
| 13 | โ Template 19 (19.2) | ADR Memo + SCQA Self-Check | ADR one-pager |
| 14 | Template 14 | Stage Gate Checklist + Kill Criteria | Kill criteria card |
| 15 | Template 15 | Co-build Agreement + Knowledge Transfer Cadence Sheet | The four co-build clauses |
| 16 | Template 16 | Production Readiness Checklist + Running Budget Sheet | Four-ledger budget sheet |
| 17 | Template 17 | Action Queue Design Patterns + Decision Trail Schema | Queue column design sheet |
| 18 | Template 18 | Metric Tree + Incident Runbook + Same-Day Notice | Three-part same-day notice |
| 19 | Template 19 | Three-Memo Set (kickoff / decision=ADR / impact) | Impact memo template |
| 20 | Template 20 | Resistance Decoder + Hard Conversation Scripts + Visibility Governance Pledge | Resistance decoder |
| 21 | Template 21 | Adoption Plan + Super-user Agreement + Operating Cadence Sheet | Four adoption mechanisms checklist |
| 22 | Template 22 | Handoff Plan + Five Self-Sufficiency Tests + Handoff Cadence Sheet | Five self-sufficiency tests sheet |
| 23 | Template 23 | Pattern Extraction + Asset Register + Library Admission Review | Asset register |
| 24 | Template 24 | F2P Memo + Closeout Asset Recovery Checklist | Closeout recovery checklist |
| 25 | Template 25 | Intake Rubric + Red Lines + Saying No Scripts + Kill Register | Three red lines card |
| 26 | โ Template 3 (3.7โ3.8) | F1โF5 Anchors + Annual Growth Agreement | Annual growth agreement |
## Grouping by Part, and the Main Chain
- **Part I ยท Position and Mindset** (Templates 0, 2, 3). For the opening week. Template 0 is where the book starts, the five deliverables due within two hours.
- **Part II ยท Discovery** (Templates 4, 5, 6, 7). The entry to the main chain. The charter in Template 4 is the authority everything after it rests on.
- **Part III ยท Design** (Templates 8, 9, 10, 11, 12). The core of the main chain. Order of output = 8 (boundary) โ 9 (data) โ 10 (pattern) โ 11 (eval) โ 12 (constraints).
- **Part IV ยท Build and Run** (Templates 14, 15, 16, 17, 18). The eval from Template 11 is the gate input for 14. The decision trail sheet in 17 is raw material for monitoring in 18 and for assets in 23.
- **Part V ยท Adoption and Transfer** (Templates 19, 20, 21, 22). The three memos in Template 19 run through the whole cycle. The rest of this group follows the order "resistance โ adoption โ handoff".
- **Part VI ยท Reuse and Growth** (Templates 23, 24, 25, plus the second half of Template 3). For closeout. The asset register in 23 and the recovery checklist in 24 share fields (mapping in Template 24).
**The template main chain** (the minimum set for one full delivery cycle): Template 0 โ 4 โ 8 โ 9 โ 11 โ 14 โ 17 โ 19 โ 22. Take the rest as needed.
---
# Vendor Crosswalk
> **Who this is for.** You are the vendor-side FDE or delivery engineer, on site at a client under contract, pushing an AI system from demo to production. The text keeps saying "the business side", "your manager", "project approval", "renewed funding", and you wonder what they mean for you.
---
## One Judgment First
The last mile is defined by crossing one organizational boundary and keeping an AI system alive in the real workflow beyond it. Whether that boundary is a company wall or a department wall changes how visible the mechanisms are, not the mechanisms themselves. The text is set inside a company because that scene hides them deepest. Contract, price tag, payment milestones, exit date, none exist there, so the book rebuilds their functions one by one. Your scene is the explicit version of the same mechanisms. Trust starts at zero, commitment is contractual, exit has written clauses, every mechanism is on the table.
Read it this way. Learn the mechanism from the text, then use this table to land it in the document you hold. Usually what the text labors to rebuild is one line in your contract, and the reminder is that a line is not a working mechanism.
## Concept Crosswalk
| Concept in the Text | Vendor-Side Equivalent | Chapter |
|----------|------------|--------|
| Business side / requesting side | The client | Whole book |
| Your team, the Digital Center | Your vendor firm, delivery team | Whole book |
| Your manager (Owen Hartley) | Your commercial and delivery leads | Chapters 2, 4, 22 |
| Project approval form, old ticket's wording | Procurement contract, acceptance clauses | Chapter 4 |
| Four charter signatures (business line head, business owner, you, your manager) | Three signatures, client decision-maker, client owner, you; commercial contract signed separately | Chapter 4 |
| Business-side commitment (charter element 5) | Client commitment, same names, hours, calendar slots | Chapters 4, 15 |
| Your team's commitment and protection conditions | Staffing clauses in contract | Chapter 4 |
| Three tiers of resource reassessment | Exit clauses, "if the client's committed input goes unmet two weeks running, we may propose a pause" | Chapters 4, 13, 18 |
| Launch release conditions | Acceptance clauses tied to the eval | Chapters 4, 11 |
| Three resource gates (named commitments claimed / renewed funding tied to the gates / fixed reassessment date) | Three money gates (paid pilot / payments tied to thresholds, not dates / expiry exit in contract) | Chapter 14 |
| The two claims-ops IT engineers, co-build engineers | Client engineers | Chapters 8, 10, 15 |
| Module CODEOWNERS, merge rights, oncall ownership | Code in client's repo, external collaborator accounts | Chapter 15 |
| Annual budget and headcount review, renewed funding | Renewal meeting, keeping the engagement alive | Chapters 19, 22 |
| Three mandatory handoff mechanisms (no ops headcount / into your OKRs / transfer ledger) | Contract end date | Chapter 22 |
| Responsibility transfer agreement, stepping out of the daily | Exit agreement, account revocation, physical exit | Chapter 22 |
| Capacity ledger | Day-rate pricing and margin | Chapter 23 |
| Group platform team | Your own product team | Chapter 24 |
| Intake gate and its non-monetary price | Screening requests by quote | Chapter 25 |
| Switching subsidiaries, dual Group and subsidiary sponsors | Switching clients | Chapter 26 |
| Stakeholder map, field archaeology, data fitness, eval as spec, action queue, adoption engineering | Apply as is, change nothing | Chapters 5โ21 |
## Five Real Divergences
The crosswalk translates. These five are structural. Six chapters carry a "Vendor View" sidebar (4, 14, 15, 22, 23, 26). Here is the summary.
**1. You have an exit (Chapter 22).** The text spends three paragraphs rebuilding mandatory handoff mechanisms because the internal deliverer cannot leave. You have a contract end date, so the deadline is ready-made. The cost is the flip side, a clean break. The response window is not a given. It goes into the exit agreement, with the account revocation date. The five self-sufficiency tests and four stages apply as is, and for you the fourth stage is a physical event.
**2. You can name a price (Chapter 25).** The intake gate and its non-monetary price filter for teams that cannot quote. Your quote alone screens out half the ideas. The cost is sales pressure. The gain from taking a job is booked at once to the taker, the cost of delivery late, to you. The vendor-side failure mode 1 is still taking everything. The three red lines and the three steps to no hold as is.
**3. You are the hired expert (Chapter 5).** The internal deliverer is no prophet at home, low start, low ceiling. You are the reverse, negative opening balance, high ceiling. "Hired" carries authority, and the expectation of "use and dismiss". You have no depth here. The last failed project's real cause of death takes you six weeks to unearth, and what an internal reader says in a sentence, you probe for with a pre-mortem. Helping Kevin Doyle for 40 minutes and no more is a strong signal. You are an outsider, so not overstepping is itself evidence.
**4. Money has a shape (Chapters 4, 14).** The text replaces the price tag with named commitments, funding milestones, and a reassessment date. You have a price tag. "Let's wait and see" costs the client an invoice, so a zombie pilot dies faster than inside. The cost is contract language dragging the charter into legalese, the vendor-side failure mode 1. Charter and contract must be two documents.
**5. Complexity jumps by changing clients (Chapter 26).** The internal reader must ask for the awkward business lines, because one company's projects converge in difficulty. Changing clients changes complexity, and the risk of moving sideways is more visible. Ten clients in five years at one difficulty, and your rate will not move. The three paths out are the original for you, founder, product leader, field CTO.
## Two Natural Assets, Two Traps
**Asset one, explicit mechanisms.** Contract, price tag, and exit date force out expectations, filtering, decisions, and handoff, so the three chapters that rebuild them (4, 14, 22) read faster. **Asset two, the neutral seat.** You are an outsider. "That is not mine to decide" is a buffer, and the client owner issuing the governance pledge comes naturally.
**Trap one, dependency addiction.** Stable renewals, comfortable revenue, being needed. Three comfortable parties, none with a reason to break it, the vendor-side failure mode 3 of Chapter 22. **Trap two, resetting to zero.** Field archaeology, trust balance, constraint baseline, dug up anew with every client. You lack the internal reader's depth, so the pattern library (Chapter 23) is priced higher.
## Reading Advice
Start at Chapter 0 as written. The opening 48 hours hold, only the clock starts the day you arrive. Read straight through, and at project approval, renewed funding, or your manager's signature, come back here to translate. Chapters 4, 14, 15, 22, 23, 26 each end with a "Vendor View" on where that chapter's boundary truly differs.