Skip to content

6 · Field Archaeology: Dig Out the Real Workflow and Its Tacit Knowledge

Companion Templates

📋 Chapter Template · 🗂 Template Library

The Challenge. Interviews do not get you the truth, and the flowchart the business side gave you does not match reality. How the work actually happens, you do not know, and you are starting to suspect that nobody in this company knows the whole of it.

What You Will Be Able to Do. Reconstruct the real version of a workflow with the six steps of the workflow truth loop. Know the three places where tacit knowledge (the knowledge people hold but cannot state) hides. Use the three-layer probing method to turn "I can tell at a glance" into rules that can be written down and verified.


Week 3, Wednesday, 8:20 a.m., a Folding Stool

Over tea on Thursday of week 1, you and Linda Marsh agreed on "come sit with me for half a day next week." That appointment got pushed twice, both times for the same reason, backlog. By the time you actually set a folding stool behind and to the side of her desk, it is Wednesday of week 3, the charter meeting is done, and the signatures are still a few days off. Two reschedules are data in themselves. Front-line time is the most expensive resource in this company, more expensive than Grant Whitmore's. Grant's calendar has open slots. Linda's queue does not.

At the second reschedule, a road opens to you that an outsider does not have. Go to Kevin Doyle, have him say a word to Linda's supervisor, and the half day gets scheduled on the spot. That road really works, and the price is really high. A half day bought with a supervisor's authority is a half day Linda will treat as an evaluation. Every move you see from behind her shoulder will be a performed move, that Excel file will stay closed all day, and that spreadsheet is the one thing you must see today. The test is the reason, not the count. If the reason changes every time, you are being brushed off and escalation is warranted. If the reason is backlog every time, the place you want to look is exactly where it hurts most, so wait. Consider escalating only after three or more in a row, and when you do, talk about schedule protection, not attitude, and land it in the business-side commitment line of the charter (Chapter 4).

It is not that you have not seen her these two weeks. Since week 2 her team has spent 20 minutes a day annotating cases for you. But the Linda in the annotation room is answering your questions, and the Linda at her desk is doing her job. Those are two different people, and today you are here to watch the second one.

In your hand is a printout, the Exception Claim Handling Process Kevin gave you last week.

  1. System assigns the exception claim → 2. Review document completeness → 3. Notify for missing documents → 4. Second review and approval → 5. Write back to the core system, close

Five steps, straight arrows. Your plan is to use it as a checklist and see which step Linda gets to.

At 8:31 Linda logs in. The first move is on the paper. She opens the core system and scans the exceptions assigned overnight. The second move is already off it. She opens an Excel file on the shared drive, enters the new claims one by one, then sorts today's order by her own color codes in the sheet, not by the system's assignment order. You swallow the question, write a line in the friction log (the Chapter 0 list of "where reality and paper do not match"), and keep watching.

By 9:40 your count stands at step fourteen. Below is what you wrote down on the spot. Steps in bold happen only inside that Excel file, steps marked (phone) get done by a phone call, and the rest are moves that exist in the official five.

  1. Scan newly assigned claims in the core system;
  2. Enter the new claims into the Excel tracker;
  3. Sort today's order by the Excel color codes;
  4. Check the document list claim by claim;
  5. For missing documents, search the inbox by claim number and surveyor name. The documents often arrived long ago, nobody attached them to the system;
  6. If the inbox has them, download and attach to the system, flip the Excel mark from yellow to green;
  7. If the inbox does not have them and the surveyor is someone she knows, call and ask (phone);
  8. If not, send a chase email, copy the supervisor, and add 1 to the Excel "Chases" column;
  9. For claims with doubtful amounts, scan the claim report and the photos, ten seconds, flag or do not flag;
  10. When unsure, call the surveyor to verify details from the scene (phone);
  11. For claims judged abnormal, mark the Excel "Risk" column red and write a word or two only the team understands;
  12. When done, go back to the core system and change the status, but that waits for second review, so Excel has a separate "Actual Status" column;
  13. Filter out claims older than seven days and post a few in the department group chat to chase them;
  14. Before leaving, save a dated copy of the Excel file, "the shared drive lost it once."

Fourteen steps. Six of them happen only in an Excel file that does not exist on the official flowchart, and two are phone calls to a surveyor she knows. Not one of the official five steps is wrong. They are all there, like a riverbed. But the shape of the water is nowhere on the paper.

At the 10:15 tea break you ask the question you have held for two hours. "That spreadsheet, is it in the system?"

Linda's answer, which you later quoted verbatim in your memo to Grant: "The status in the system is for the people upstairs. That sheet is our own." The sheet has columns the core system does not have, actual status, chase count, risk marks, notes in team shorthand. This team is the only one in the company that maintains it, and has for six years. This sheet is the source of truth, the data everyone actually trusts and acts on. You write down the most expensive line of the half day in the friction log. The source of truth for exception claims is not in the core system. It is in an Excel file on the shared drive. The data reconciliation in Chapter 9 starts from this line.

Why This Is Hard: Organizations Are Systematically Wrong About Their Own Workflows

Nobody lied to you. The flowchart Kevin gave you was sincere, and it is also "correct" in the sense that audit and training use the word. The problem is that every organization holds three versions of a workflow at the same time.

  • The SOP version describes what should happen. The SOP's job is compliance, audit, and onboarding. By design it is not responsible for describing reality.
  • The reporting version is what management sees. Every step up the reporting chain smooths away a bit more of the workarounds (the homegrown ways of getting around the system). The front line does not report "that sheet," because reporting a workaround tool means confessing a violation, or inviting a round of "systematization," and both are losses. A rational front line stays silent forever.
  • The actual version lives only at the desk. It lives in exception handling, private spreadsheets, and the feel of veteran staff, and it is alive. Surveyors change, fraud methods change, and it quietly updates every month without telling anyone.

The three versions coexist indefinitely and never correct each other. This is not an illness of this company. It is the normal state of every organization. The trouble is that an AI system must embed in the actual version, which is what the workflow gap among the five gaps in Chapter 1 means. Build the system from the official flowchart and it will wait at step 2 for a "document completeness result," while in reality that result is scattered across the inbox, phone calls, and three columns of an Excel file. The official flowchart is not wrong. It belongs to a parallel universe. Build the system from it and you are writing software for a parallel universe.

This is also why interviews fail. Ask a manager and you get the SOP version plus the reporting version. Ask the front line "what is your process" and you still get the SOP version. When people are asked about "the process," they recite the one they were taught, not the one in their hands. She was not hiding anything. She herself never counted the fourteen steps as "a process." That was just "doing the work." The real workflow is in nobody's words, only in their actions. There is one way to get it. Go and watch. For you this sentence carries one more layer. You have been at this company for years and thought you already knew how the claims department works, and the version you knew is precisely the one that traveled up the reporting chain to you.

The internal reader's biggest risk is skipping this chapter. The reason sounds solid. I am not an outsider, why would I need to carry a folding stool and sit for half a day. But look back at where this dig started, and the honest answer is not flattering. Why did you first sit beside Linda's desk only in week 3? Because you have been at this company three years and thought you knew how the claims department works. In those three years you saw their monthly reports, heard their fifteen minutes at the business review, took the tickets they filed, all of it reporting version and SOP version. You never saw that Excel file, because by design it does not flow in your direction. The two reasons the front line does not report it are only stronger in front of you. You are exactly the person who would systematize it out of existence.

Familiarity is the only enemy in this chapter. The outsider knows he knows nothing, so he goes to watch. You know a lot, so you go to meetings. And sitting for half a day is far cheaper for you than for an outsider. Two floors down, no appointment, no background briefing, no week spent explaining why you are sitting there. If the cost is that low and you still do not go, that is not a judgment. That is familiarity deciding for you.

Presence is an asset, but it only counts once it is cashed into action, and this chapter has three actions you can take now that an outsider cannot. First, you need not do the half day in one go. You can go any time, so split it into five one-hour visits on different days. Monday's backlog and Friday's wrap-up are not the same workflow. An outsider can only buy four consecutive hours. You can buy five different cross-sections. Second, go through the review archives. In the three-layer probing method below, the third layer asks Linda whether she has ever flagged a claim and misjudged it. From memory she gives you one claim from the year before last, while the same batch of claims has a complete record in the QA reviews and the customer complaint log, which you have permission to pull. Lying there are the counterexamples she cannot fully remember. Third, ask for the first version of that Excel file. It has been maintained for six years. The person who built it may still be at the company, and the email proposing it may still exist. Why a workaround tool got built is page one of a requirements document written over six years, and you can get it.

Prior Art, and What AI Changed

Going to watch has two mature traditions. You need not remember the names, only one thing. Watch first, ask second.

The consulting tradition, fact-driven, firsthand first. The first rule of McKinsey's problem-solving discipline (as retold in Ethan Rasiel's The McKinsey Way) is fact-based, and facts have ranks. Firsthand observation in the field outranks the business side's reports, which outrank industry press. Facts should be taken from where they are produced.

The anthropological tradition, observe before you ask. From Malinowski's fieldwork to the contextual inquiry that the design world engineered out of it (Beyer & Holtzblatt), the core is the same. Enter the field as an apprentice, with the user as the master. Watch the master work first, anchor questions on the concrete action that just happened, never on hypothetical situations. "Why did you do that" is always asked after seeing it. Reverse the order, ask before you watch, and all you get is a rationalization invented for you on the spot.

What did AI change? The cost of writing up the dig collapsed. The digging itself has not moved an inch. Half a day of shadowing (following a person and watching them work. In this book, shadow means only a person following a user to watch real work. It has nothing to do with the engineering practice of running a system in parallel on live traffic with its output disabled, "shadow mode," and this book does not use the word for that) used to leave scrawled notes that took an evening to clean up, and clustering exception types out of three thousand tickets took a week. Now AI does it in half an hour, transcribing, structuring the friction log, clustering the reasons claims get stuck. All the time saved should go back into the field.

But AI cannot shadow. "Seeing that the Excel exists" is in no dataset. That sheet has no API, no documentation, no entry on the IT asset register. The only evidence of its existence is Linda's second move at 8:31 in the morning. The cheaper the information AI can organize, the more expensive the tacit knowledge that only a person on site can get. It is the raw material for the eval in Chapter 11, and the real moat of this system. Models are a public good. Linda's five rules are not.

The Framework: The Workflow Truth Loop

Turning "go and watch" into a repeatable method takes a loop of six actions.

Step Action Key Discipline
observe Be present and watch them work, do not interrupt Write questions down, do not ask on the spot; keep the friction log as you go
shadow Follow a person for half a day Follow the person, not the process. A process does not open Excel, a person does
trace Follow one claim end to end Across people, systems, and tools, count how many hands and how many systems it passes through
ask Ask about exceptions "When would you not do it this way?" anchored on the concrete action you just saw
reconstruct Draw the real process Including workaround tools and verbal coordination, not one step prettified
validate Have the person correct it Ask "where is it drawn wrong," not "is it right." The former gets corrections, the latter gets politeness

Three rules of use. First, the order cannot be skipped. The most common cheat is to skip the first three steps and go straight to ask, and then what you get is the recited version again. Second, it is a loop, not a pipeline. Validate exposes new exceptions, and new exceptions deserve a new round of observe. Run it two or three times, and the rate of convergence tells you when it is enough. Third, trace speaks through a claim. Shadowing looks at one person's day, trace looks at one claim's whole life, and only the two cross-sections together make a workflow.

The Anchor & Helm instance. You pick the claim ending in 4471 and follow it end to end. The claim report is in the core system, the survey photos are email attachments, the customer's chase calls are in the call center ticketing system (no join key to the claim, found by searching the phone number), and the actual progress is in Excel. One claim, four homes. This one trace is worth ten pages of interview notes.

The loop digs for the workflow, but the richest vein hides elsewhere. The three hiding places of tacit knowledge are these.

  1. Exception handling. The SOP covers the main path. All the judgment is on the branches. The probe is "when would you not do it this way?"
  2. Workaround tools. Private Excel files, sticky notes on the desk corner, personal email folders, small group chats. Like shadows, these tools appear on no architecture diagram. The test is anything open on the screen that is not on the IT asset register. Each one is a requirements document, written over six years.
  3. "I can tell at a glance" judgments. The expert has compressed the rules into intuition and cannot state them directly. The signal is three trigger phrases, "I can tell at a glance," "from experience," "hard to say." A trigger phrase is an outcrop of the vein. Mine it with the three-layer probing method below.

The workaround tools cell holds one more question for you than for an outsider. Linda's sheet sits on the shared drive, with actual status, chase counts, risk marks, and notes only the team understands, and very likely fields that can be matched to the person who filed the claim. It has never been through a data review, and it is not on the IT asset register. An outsider sees it, writes it down, puts it in the report, and that is the end of it. You see it, and you are an employee of this company whose job description says risks found must be reported, and half an hour ago you told her, I am not here to evaluate you.

The rule still holds. See a noncompliant workaround, note it, do not call it out (Template 6, code of conduct rule 2). But it governs that half day, not the half year. You do not correct on the spot because your role while present is apprentice, and correcting would make everyone you observe from then on start performing. Whether to report it afterward is a separate question. Noting without calling out is not permanent secrecy. Merge the two and you either wreck this dig on the spot, or become the person who knew and said nothing six months later.

Draw the line before you go into the field. What must be reported is real compliance risk. Data that can be matched to an individual sitting on a personal computer or in a personal mailbox, outbound sending that bypasses approval, operating under someone else's account, a step that requires a record being skipped. These four are not efficiency problems, they are institutional problems, and if you know and stay silent, your silence will count as consent when something goes wrong. The great majority of the rest is not in this class. Sorting in Excel, color-coding priority, calling a surveyor you know, these are not violations. They are scars left by underdesigned process, and their proper destination is the real flowchart and the data reconciliation of Chapter 9, not a compliance report. When you cannot tell, the test is whether this would become a problem if audit saw it, not whether it looks proper.

If you really must report, one more discipline. Tell the person first, face to face. It can be short. Your sheet has fields that can be matched to individuals, and this one I have to report, and at the same time I will write down why you need this sheet and what breaks if it is shut down before a replacement exists. That is not an easy thing to say. But you and Linda will still be in meetings together next year. Report around her once, and the second time nobody opens any spreadsheet in front of you. The reverse also holds. One report that was announced first, and that really spoke for her in the write-up, tells the whole claims department that opening a sheet in front of you is safe. An outsider cannot build that credit. You can, and you get exactly one chance to build it.

At Anchor & Helm: A Ten-Second Judgment, Five Rules

At 11:20 Linda pauses on a claim for under ten seconds and flags it red. Half an hour earlier Sam, the team's new reviewer, came over with a laptop to ask about a similar claim, and Linda glanced at it. "Reported three days late, and it is that repair shop again. Flag it." Sam asked how she could tell. She said, "You see enough of them and you know." That is how this team's judgment gets passed down. Verbal, at random, ten seconds at a time, never on paper.

At 11:40, the last half hour, you start probing. Asking "why did you flag it" straight out is useless. You have already heard her first response, "I can tell at a glance." The three-layer probing method starts here.

Layer one, anchor on an instance, replay the actions, do not ask why. "That claim just now, 4471, you went to the report time first, then opened the photos and flipped through two, and then you flagged it. Right?" She paused. "...Right. Rear-end collisions rarely go unreported the same day, and this one waited three days. Only five photos, a surveyor at the scene normally takes a dozen or more." Two rules surface. Where is the leverage? You replayed the actions she had just taken herself, and a stock phrase does not match actions.

Layer two, compare, find the difference between two similar claims. You pull up another rear-end collision from the morning with a similar amount. "This one is two thousand higher than 4471. Why did you not flag it?" "That one was at a major intersection, with a police report, and the repair shop is a dealer's service center." The amount is judged against the usual range for this line of coverage and this vehicle model, not in absolute terms. And "the repair shop" is a list she keeps in her head.

Layer three, ask for boundary counterexamples. "When would you not flag it no matter how high the amount?" "Have you ever flagged one and been wrong?" She thought about it. "I have flagged wrong. The year before last there was a claim that turned out to be the owner paying for the repair out of pocket first... so now I also check whether the same plate has come through in the last six months."

Three layers in, "I can tell at a glance" resolves into five rules of thumb.

  1. Amount deviation. The repair amount is well above the usual range for the same coverage and vehicle model (judged relative, no hard threshold);
  2. Report delay. The gap between the incident and the report is beyond the ordinary (a rear-end collision not reported the same day, reported three days later);
  3. Survey photo count. Clearly fewer photos than a normal claim, or key angles missing;
  4. Prior claim linkage. The same plate, filer, or phone number appears again within six months;
  5. Repair shop list. The shop is on her rolling mental list of shops that "come up a lot."

On Friday you print the fourteen-step real flowchart and the five rules, put them on Linda's desk, and ask one question. "Where is it drawn wrong?" She circles two places. Step 13 (filter claims over seven days and post them in the group chat) has the wrong frequency. Tuesdays and Thursdays, once each. She does not do it daily. And the repair shop list, "it is not a fixed list. I crossed one off just last month." The weight of the second one, failure mode 4 will explain.

Where these five rules go, you can be told now. When you co-build the eval spec with Linda in Chapter 11, every one of them becomes an error category for the golden cases (sample cases with the correct answer settled in advance), and "missed risk" gets split along these five into five testable ways to fail. That is the step where tacit knowledge turns from interview color into a system asset. The half day's accounts are easy to settle. Four hours on a folding stool produced three things. The existence of the Excel file, the starting point of the data reconciliation in Chapter 9. The fourteen-step real process, the map of where the Action Queue embeds in Chapter 17. The five rules of thumb, the raw material of the eval in Chapter 11. Every key deliverable of this project later traces back to these four hours.

One more thing to arrange right now. The conclusions of these four hours will expire, and most likely you will be the one who expires them. Chapter 9 goes a layer harder. Your system will kill its own source of truth with its own hands. Of the fourteen steps, six happen only in Excel. After the queue goes live the system takes some of them over, the rest grow new workarounds around it, and nobody will come to tell you about the new workarounds either. The real flowchart you drew today starts going stale the day the system launches.

An outsider does not handle this. Next time he comes, he digs from scratch. You will not. You only do increments, so the cadence of increments is yours to set. Put it on the calendar. Go sit once in week 4 and once in week 12 after launch, one hour each, doing two things only. Count the steps again, and ask whether any new spreadsheet appeared this month. After that, once every six months, as the queue's periodic review (Chapter 17 expands on it). The real flowchart and the five rules both get a date and a next-dig date. A flowchart with no date is the same thing as an SOP three months later.

Failure Modes

1. Interviewing only managers. Discovery (the stage of finding out how things actually are) ran three weeks, met eight directors and team leads, produced a thick requirements document, and never sat at a desk once. There are two layers of cause. Managers are the people you can book. They have calendars, meeting rooms, and a desire to talk, while the front line is "in the backlog." Worse, a manager's information itself comes from the reporting chain, so interviewing him is looking at the field through two layers of distortion, and everything you get is the should-be universe. One discipline. In discovery, the actual user must account for more than half of interview time. If you cannot book them, wait. Two reschedules, still wait.

2. Treating the flowchart as the truth. You get the SOP, treasure it, and design the system state machine straight from it. The SOP's institutional functions (compliance, audit, training) mean it describes what should be. The organization has motives to maintain the should-be version and no mechanism to update the actual one. It was never a map, so it cannot be out of date. The right use. The SOP is the starting point of the dig, not the end. Print it, bring it to the field, and hunt specifically for where it differs from reality. Every difference is a line in the friction log.

3. Asking "what features do you want." The requirements interview becomes an ordering session, and you bring back a wishlist. Users can only imagine the future in the language of their current tools. Ask Linda what she wants and she will say "add an automatic reminder to that sheet," because that sheet is the whole world she has seen. Users can order from the menu, but users are not the chef. What she reports is a solution projected onto the old tool, and the problem itself is still buried. And "what do you want" is a hypothetical question, and the anthropological discipline is precisely not to ask hypotheticals. The right path. Observe the problem (that inbox search you watched this morning, hunting every day for documents that arrived long ago), and let the solution grow between you and her.

4. Hardcoding tacit judgment. You get five rules and that night write them as five if-else branches and ship. Rules of thumb are a snapshot of living rules. The repair shop list lost an entry just last month, fraud methods evolve, Linda's rules move with them, and if-else does not. Three months later the rules have drifted, the system still flags from the old snapshot, and because "the rules came from Linda," nobody questions it. Hardcoding also quietly moves the decision from the person to the system, and when it errs, there is nowhere to put the responsibility. That is exactly why the Chapter 0 MVP queue keeps a "Human Call" column. The right path. Make the rules explicit, put them in the suggestion and reason columns, keep the decision with the person, and record overrides in the decision trail. The decision trail is the drift detector for rules (Chapter 17 expands on this mechanism).

Next Monday

  1. Find your project's actual user and book half a day of shadowing (Template 6 has the opening script and a half-day schedule). Do not be discouraged by a reschedule. The reschedule itself is information. Where the most expensive time is, that is usually where the most painful workflow is.
  2. Print the official flowchart and take it to the field as your dig map. Count the real steps. Look for tools open on the screen that are not on the IT asset register. Find one and you have found a requirements document the user wrote over several years.
  3. Next time you hear "I can tell at a glance," use the three-layer probing method (anchor on an instance → compare → boundary counterexample) to recover the rule, write it down, and take it back for the person to circle the mistakes.
  4. Go through your existing requirements document and tag the source of every line. Manager interview, SOP, or firsthand observation? If the first two make up more than ninety percent, you are writing software for a parallel universe.

Want an agent to get you started? In the repo you set up following Start Here, paste this to your coding agent:

In the repo/ directory of the the-last-mile repository, help me with the Chapter 6 Next Monday actions. Copy
templates/field-archaeology/friction-log.md into the working directory I name, and leave it blank for now. When I am back
from shadowing I will paste my dictation or notes to you. Organize them into a structured friction log as described in
prompt-organize.md, keeping the original wording in every entry, and do not add steps you did not hear. While organizing,
mark every "I can tell at a glance" style judgment and probe me on each one with the three layers (anchor on an instance,
compare, boundary counterexample). The recovered rules are written by me, and the last step is taking them back for the
person to circle the mistakes, not handing them to you.
If any command errors, stop and show me the output.

Chapter Kit

  • Judgment frameworks. The workflow truth loop (observe → shadow → trace → ask → reconstruct → validate); the three hiding places of tacit knowledge (exception handling / workaround tools / "I can tell at a glance" judgments)
  • Templates. Template 6, the Field Archaeology Kit, half-day shadowing guide + friction log template + three-layer probing script
  • Key judgments
  • "The official flowchart is not wrong. It belongs to a parallel universe."
  • "Users can order from the menu, but users are not the chef."
  • "AI cannot shadow. 'Seeing that the Excel exists' is in no dataset."
  • "The SOP is the starting point of the dig, not the end."