Skip to content

Chapter 11 · Failure Modes Unique to Research

Chapter companion

📋 Chapter 11 templates · 🗂 Template index · 💻 code/persona-panel

This chapter's ladder. This chapter cuts across all seven steps. Every failure mode grows on a step of its own, so the chapter cannot be filed under any single step. Recognizing failure modes is itself work AI currently does at assistant level. Formal errors, fabricated citations, numbers that do not add up, you can let it scan for those. You will also meet a class of error it structurally cannot catch, and in that class your accomplice is you. The conservative note on the "read and catch errors" row of the snapshot in Chapter 3, section 3.5 is about exactly this.

Spine update. Part III opens no new level. It settles a debt. The pit I fell into in Chapter 4 goes on the dissection table here, as the most expensive specimen in the chapter.

This chapter delivers. Three things. Three error amplifiers, an "AI attribute × research step" failure-mode table, and a thirty-minute failure-mode census, with an AI self-check prompt set attached.


11.1 The ninth hour before the meeting

Eleven on a Tuesday night, the architecture review nine hours away. You are on the last pass. On the desk is a technology selection report AI was deeply involved in, twenty-six pages, clean structure, plenty of citations, your group's main output of the past two weeks, concluding that the retrieval layer should be swapped for a new engine.

The core input to the migration benefit model is one performance judgment, that the new engine cuts "p99 latency by 31%" under comparable load. The report gives a source, the 2025 Vector Retrieval Benchmark Review, with a page number. You want to see the test workloads in the original, so you search. Nothing. Three different keyword sets, still nothing. You go back to the chat window and ask the AI. It apologizes, then "corrects" the source to the name of another evaluation report. That one does not turn up either.

It is half past midnight. The report is not garbage. Most of the other citations you spot-checked are real, and the framework is genuinely useful. The trouble is that the number holding up the entire migration benefit now hangs on a citation nobody can find, in exactly the same tone as every true sentence in the report. The next thought is colder. What if you had not run that one extra search tonight? The 31% goes into the benefit model, the benefit model goes into the architecture decision record, the decision record goes into quarterly planning. Six months later another team runs a selection and cites your decision record. The lawyer in the Avianca case from Chapter 1, who filed AI-invented precedents with the court and was fined after the judge checked, is the man who failed to ask one more question at exactly this spot.

This chapter does one thing. It builds a case file for moments like this. First the format of the file. It is not organized as a bug list for AI. You will see later that reading these errors as a bug list is precisely why they do not get stopped. The questions it answers are structural. Why are these errors more lethal in research than elsewhere? Where do they grow from? Handed a polished output, how do you know where to poke?

11.2 Old errors, a new production line

Start by taking apart a convenient illusion, "these are all new AI defects, and one more model generation will fix them."

Turn back to Chapter 2 and not one of these errors is new. Spurious significance means a conclusion that looks statistically sound but is a false positive squeezed out by adjusting the data and the analysis method over and over, p-hacking in the jargon. The Simmons paper demonstrated it live in 2011. That was eleven years before ChatGPT. Bad data riding into an authoritative conclusion, the Reinhart-Rogoff Excel formula that skipped five rows, moved fiscal policy debates in several countries.

Citation rot is an older disease still, one that bibliography has always had. Dead links, drifting paraphrase, qualifiers shedding layer by layer through secondhand citation, all of it dates to the print-journal era. A meta-analysis of 28 studies of medical journals, meaning the results of many studies computed together (Jergas and Baethge, 2015, PeerJ), measured a total citation error rate of 25.4%, and about one in ten of those citations simply did not support the claim they were cited for. The comparable audit in ecology journals (Todd and Yeo, 2007) landed around 24%. Long before ChatGPT, academic citation carried a standing one to two in ten duds.

AI added not one new kind of error. What it changed is the production function of error. Every coefficient moved.

Volume changed. Manufacturing one fabricated citation good enough to pass used to take effort, so only people who set out to fake did it. Now it is a free byproduct of a language model, a dozen out of one generation, and nobody has to set out to do anything. This is row 6 of the transfer map, AI turns "looking rigorous" into one generation.

Fluency changed too. Pre-AI errors often carried a tell, a sudden shift in style, data that did not hang together. Errors AI produces have no tell. The wrong paragraph and the right paragraph come off the same text engine, and their surface properties are identical. That is row 7 of the transfer map, hallucination and fabricated citations hiding behind polished prose.

Marginal cost went to zero. In a hand-written review, every citation passed through the author's hands at least once. In an AI-generated review, the length of the citation list and the probability that any one entry passed through anyone's hands are fully decoupled. Pre-AI error detection ran mostly on that one gate, human attention.

So read this chapter the right way. An old set of opponents on the research field just got industrial equipment. Your defense has to industrialize with them, and that is Chapter 12's business. This chapter settles "recognize it" first.

11.3 The three amplifiers

Why is the same error more lethal in research than elsewhere? Research fits errors with three amplifiers. These three are the foundation of the chapter's taxonomy, and of every failure mode below you can ask which of them it feeds on.

Amplifier one, invisibility. Research errors raise no error. Write code wrong and the compiler hands you a red answer in seconds. Get a research conclusion wrong and nothing happens at all. The wrong conclusion sits quietly in the document, looking exactly like a right one. Chapter 2 calls this "research has no compiler" (failure condition one). A research error can only be caught by a step in the process, and steps can be skipped, especially when you are short on time.

Amplifier two, compounding. The definition of a research output is "knowledge downstream will use." A wrong conclusion gets written into a report, the report gets cited into a review, the review gets fed into the training data of the next generation of models, and all the while it is supporting real decisions. Chapter 2's contrast (failure condition three) applies here. Bad code can be rolled back. A bad conclusion has no rollback button. If the 31% in section 11.1 had survived that night, its blast radius would widen month by month.

Amplifier three, motivated collusion. The nastiest of the three. When an error happens to grow into the shape you wanted, supporting your hypothesis, filling the gap in your argument, making the project look like a contribution, your motivation to check it drops to its lowest point. This amplifier differs from the other two in one place. Those two are properties of the environment. This one grows on you. The inspector and the error become accomplices, usually without noticing. AI's role here is an amplifier of the amplifier, not a deceiver. It is trained to talk in the direction you are facing, and whatever you want, it wraps to look more like fact. Section 11.5 has a firsthand specimen, and the victim is me.

The three amplifiers multiply, they do not add. An error that is invisible, that compounds, and that happens to be what you wanted, that is the recipe the most expensive accidents in research are produced from.

11.4 Attribute times step, building the case file

Now build the taxonomy. First the chapter's core claim.

Failure modes are the interaction product of "AI attribute × research step," and no bug list written about AI alone will produce them.

Take it apart. On the AI side, three attributes offend again and again, and all three are neutral in themselves. They are even why you bought it. Fluency, the output is always coherent and confident regardless of whether the content is true, which is a language model's default property. Sycophancy, it is trained to produce answers people are satisfied with, and human raters' preferences are etched into its gradients. Corpus prior, its "knowledge" is a statistical compression of past text, not an observation of your situation. On the research step side, it is Chapter 3's seven steps. The same attribute landing on different steps grows into different errors, the way one pathogen causes different diseases in different organs.

The summary table first, then a dissection of each class.

High-incidence step Failure mode Chief attribute One-line signature
Master the field, deliver Hallucination and fabricated citations Fluency × corpus prior The most on-point citation is the most suspect
Test plan, interpretation Spurious significance and the criteria backdoor Sycophancy × fluency The criterion appears after the result
Execution Data leakage and contamination Corpus prior The score is too good to be true
Interpretation, red team Sycophancy drift Sycophancy The conclusion flips with how you ask

The four rows are not everything. One more class stays off the table because it does not pick a step, and section 11.5 dissects it on its own.

Hallucination and fabricated citations. These grow on the master-the-field and deliver steps. Fluency writes the invented content so it cannot be faulted, and corpus prior makes it look like a typical sample of the field, typical title, common authors, plausible year. There are two signatures. The first is counterintuitive. Fabricated citations are built to order for your needs, and their distribution is not random. The model is completing "a paper supporting this claim belongs here," so the one that happens to fill the gap in your argument, the one whose title hits your question dead center, are the entries in the whole list to check first. The odds are not small either. Walters and Wilder (2023, Scientific Reports) examined citations generated by GPT and found GPT-3.5 fabricated 55% of them, GPT-4 still 18%. The second signature is paraphrase drift. The citation is real, and AI's retelling carries fewer qualifiers than the original. "Improves on arithmetic tasks" becomes "improves", and Chapter 4, section 4.7 covered the mechanism. The scenarios most likely to catch you are first drafts of reviews, the citation list of a report, and anywhere "AI added a source while it was at it."

Spurious significance and the criteria backdoor. Move to test plan and interpretation and this one runs the house. Say to AI "have a look at whether there is anything in this data" and it will almost certainly give you something. The sycophancy attribute guarantees a finding, the fluency attribute guarantees the finding looks well founded. This is the AI edition of p-hacking, more dangerous than the manual edition. When a person picks the data, that person knows roughly what they did. When AI picks for you, what you receive is an analysis that looks objective and neutral. The criteria backdoor is its twin. The decision standard gets written down only after the results are in, or gets quietly revised, and the conclusion then always "just happens" to clear the line. Its signature is one question. Was this criterion locked in before the data was seen, or after? If you cannot answer, or the answer is "after," treat that significance as zero. The scenarios most likely to catch you, one is exploratory analysis delivered as a confirmatory conclusion, the other is the boss asking "could you look at this from another angle." Chapter 6 locks criteria first and signs off on them to guard against exactly this class. For auditing someone else's output after the fact, see Chapter 12.

Data leakage and contamination. On the execution step this one owns the ground. Corpus prior takes its most concrete shape here. The questions and answers of a public benchmark are, with high probability, already in the model's training corpus. It saw the exam paper before the exam, and you cannot know how much of it it saw. Another variant is analysis code AI writes that lets test-set information leak into the training side, feature engineering using statistics computed over the full data, normalization performed before the split, which is to say things that should have been computed on training data only get quietly computed on everything. This class of error is just as common in code people write, and the AI version is fluent enough that you want to look closely even less. There is only one signature. When the score is too good to be true, treat it as leakage first, especially when the score collapses on a fresh set of questions from the same distribution. The scenario most likely to catch you is evaluating model capability on a public benchmark, which is this book's spine case as daily routine. Chapter 5's "choose a contamination-resistant variant set" when picking task families and the data contamination check in Chapter 6's criteria list are both sentries the spine case posted for this class of error at design time. Whether the sentries hold, Chapters 7 and 8 give the verdict.

Sycophancy drift. It shows up in interpretation and red team. The sycophancy attribute takes its purest form here. You ask with a lean, it answers along the lean. "Does this result show our method works?" and "Could this result be nothing but noise?", one batch of data, two ways of asking, and AI can hand you two fluent and opposite readings. Five frontier assistants consistently showed sycophancy (the sycophancy attribute described above, talking along with the asker) across many task types, and human preference data itself rewards the behavior (Sharma et al., 2023, arXiv:2310.13548). The signature here is operational. Ask the same question twice with opposite leans. If the conclusion flips with the lean, what you measured is your own phrasing, not the world. The scenarios most likely to catch you share a structure, letting the AI that produced the conclusion serve as its own judge, including asking it to "check itself for problems." That directly violates the orthogonality principle stolen in Chapter 2, section 2.4. For how to build an independent channel, see Chapter 12. For how to use AI as a critic instead of a yes-man, Chapter 10 already ran the demonstration.

Not one of the four classes needs the model to "get dumber" before it happens. All four are byproducts of the model working normally. That is why "wait for a stronger model" cannot serve as a defense strategy. The attributes will not disappear. They will only get more fluent.

11.5 Spine update · the "finding" that died on the check

One class is still missing from the table, amplifier three, motivated collusion. It offends on whichever step carries your wish. Textbook specimens are hard to find, because the people it hits usually do not know they were hit. I can offer a firsthand one, a pit this book fell into itself. Chapter 4 wrote the process down honestly. Here it gets dissected from another angle.

The summary is one sentence. I took "a head-to-head comparison on a cost-matched basis" for an unclaimed patch of open ground, that judgment survived a full round of scanning, and it died on the forward-citation check. The comparisons exist, and they cluster on the skeptics' side.

The value of the dissection is in the contrast. In the same project, a paper title I wrote from memory had one word wrong, "Key" in Wang's ACL 2024 paper, which I had remembered as "Answer." The verification channel picked it up while doing the forward check, a mechanical step, dull work. The "gap" judgment lived a whole round. Why? Look at the three attributes. Fluency? I wrote that judgment myself. Corpus prior? The core error was not in AI's retelling. Sycophancy? AI did not contradict me, true, but it counts only as an accessory. The principal offender is what the "gap" was worth to me. A gap meant my project had a contribution, the two-hour demo was not wasted, and this book's spine case held. A title one word wrong died on a mechanical step that does not care where anything came from. The gap judgment went unchecked far longer, because it was my wish. The line at the end of Chapter 4 bears saying again here. The more dangerous one is your own hallucination when you want a gap to exist.

From this specimen you can lift the general signature of motivated collusion. The conclusion you are most excited about is usually the conclusion you checked least. That is the default output of the motive structure, and it has nothing to do with character. "Be careful next time" is not a countermeasure. What saved me in Chapter 4 was the coverage check as a step in the process, not my own vigilance. For what that process looks like, see Chapter 12. First an ugly thing said in full, because it governs how you use this chapter's templates. AI self-check cannot catch this class of error. Ask AI to check your conclusion and it checks the form, whether the citation exists, whether the numbers add up, whether the reasoning chain breaks. Your motive it not only leaves unchecked, it goes along with. That warning is written out again in bold in the appendix prompt set.

11.6 The perfect respondent

Now point the whole taxonomy at this book's subplot case and run a full workup. Readers who care only about the spine case can read just the names of the five ways of passing for real and the closing paragraph, but keep in mind that synthetic data in your own eval set fakes in the same five ways. Persona research means using a persona-bearing model ("a 25-year-old mom with one child who works at X") in place of real people for interviews and survey rehearsals. It is the perfect specimen of the plausible trap. Its output is almost impossible to fault one answer at a time, and it reads more like "a real person" than plenty of real interviews do. Chapter 5 ground that rough question into a testable shape, the answer distribution agrees with the real population and no systematic stereotype drift appears. This chapter answers a different question. An output that resembles a real person that closely, in exactly which ways can it be fake?

There are five, and each has a seat on the table in section 11.4.

One, the fluency illusion. The answers are coherent, detailed, in a voice that fits the persona, and they trigger your "like a real person" intuition. Fluency has zero correlation to distributional fidelity. This one is a mechanism fact and needs no experiment. Fluency is the very target the model is optimized for, and it holds for any persona and any question. This is the standard case of "fluency × interpretation." The answer itself may be fine. What goes wrong is the reasoning step where you take fluency for evidence.

Two, the average face problem. A persona's answers are a weighted average of the corpus's stereotype of that group. Every single answer is reasonable, and the aggregate distribution is narrower and more homogeneous than the real population, variance collapse. It is like stacking a hundred faces into one "average face," well proportioned, and no such face exists in the world. Use it for all hundred and you will conclude that "humans all look alike." Corpus prior offends on the data generation step. Nominally you are sampling a population. In operation you are sampling the same statistical compression over and over.

This one has peer-reviewed measurement behind it. Early work leaned optimistic. Argyle et al. (2023) reported that "silicon samples" fit real surveys reasonably well at the level of the overall distribution. Then Bisbee et al. (2024, Political Analysis) measured the standard deviation of synthetic answers as significantly smaller than that of real respondents. Shrunk by how much? Run the same statistical power calculation, which is the calculation of how many respondents are enough, twice. On the variance of the synthetic data, 33 partisan respondents are "enough." On the real variance from the American National Election Studies (ANES), you need about 300. The center of the distribution may be right, the spread within the group has been squeezed out, and inference built on synthetic data is therefore systematically overconfident. Verified in the political survey setting, and the effect size in other fields is still exploring.

Three, stereotype drift. In answers written for the "25-year-old mom," the concentration of parenting anxiety runs systematically high. This is the average face one step further. Answers for a subgroup are more homogeneous and also systematically more extreme. The distribution narrows the same way, and its center lands near the stereotype, off the real mean, and the stereotype itself carries an offset. This one has measurement too. The Marked Personas study by Cheng et al. (2023, ACL) found that in character portraits generated by GPT-3.5 and GPT-4, stereotype vocabulary appeared at a higher rate than in a human-written control group, and leaned toward "seemingly positive" essentializing narratives (writing a group as born that way), exoticizing, "strong Black woman" archetypes. Train a simple classifier on portrait wording alone to guess which demographic group was being written about, and GPT-4's portraits get guessed right 96% of the time (GPT-3.5, 92%). Verified at the phenomenon level. For business decisions this is the most dangerous of the five. What you hear is the corpus's caricature of this kind of user, not the user.

Four, sycophancy drift. When the interviewer's phrasing carries a lean, a persona follows the question further than a real person does, the survey edition of sycophancy. Real respondents wander off topic, push back, and say "that question is put wrong." Personas almost never do. So every implicit assumption in the interview guide gets gently confirmed once more. This is the fourth class from section 11.4 replayed on the subplot, same signature. The grade of the evidence, though, is not like the previous class, and it splits into two layers. Sycophancy in the general setting has solid evidence, Sharma et al. (2023) cited in section 11.4. Whether role-play makes it worse has so far only one 2026 preprint giving a preliminary signal (Shah et al., arXiv:2604.10733, in 9 of 13 open-source models, the higher a persona's agreeableness, a personality-test score for being easygoing and pleasant to deal with, the higher the sycophancy rate), publication status unconfirmed, so it can only be used as preliminary evidence. Still exploring.

Five, time dislocation. What you interview in 2026 is a "25-year-old mom" compressed out of text from before 2024. A persona reflects a time slice of the training corpus, not the population of today, and every shift in attitude after the corpus cutoff is missing. The symptom has measurement. A 2026 study ran 9 LLM configurations against the real polls of the 2024 US election and found support for Harris systematically overestimated by 10% to 40% (the paper's own words are overestimated by 10% to 40% relative to the polls, without defining whether that is proportional or percentage points, arXiv:2602.06302), and hooking up web retrieval does not fix it. That it cannot be fixed is worth writing down. It says "training cutoff date" is an explanation with partial evidence behind it, not the confirmed sole mechanism. Symptom confirmed, cause still exploring. For fast-moving fields (consumer preference, attitudes to technology, public opinion topics), this dislocation is a systematic error of direction and cannot be treated as noise.

The five ways of passing for real collect into one table, and the two columns "which layer it lives on" and "evidence grade" are enough.

Way of passing for real Lives in the single answer or in the distribution Evidence grade
One, the fluency illusion Single answer, it fools your intuition Mechanism fact, no experiment needed
Two, the average face problem Distribution, variance collapse Verified in political surveys, still exploring in other fields
Three, stereotype drift Distribution, center pulled toward the stereotype Verified at the phenomenon level
Four, sycophancy drift Distribution, it follows how you ask Still exploring, one 2026 preprint's preliminary signal and nothing else
Five, time dislocation Distribution, the whole thing stopped at the corpus cutoff Symptom confirmed, cause still exploring

The five share one property. Not one of them is visible at the level of a single answer. The fluency illusion fools your intuition, and the other four all live at the distribution level. Every single answer is unfaultable, and a thousand answers together are wrong. This is the top form of the plausible trap. Every item passes its check and the conclusion is still wrong. Against it, reading the interview transcripts more carefully does nothing. You have to compare against the distribution of a real population as ground truth, dual-source, three criteria, a preregistered falsification shape. For how to expose it, see Chapter 12, which is also where the subplot's suspense is settled.

11.7 Swap in your project, run a failure-mode census

Your project should have walked a few steps by now, at least a controversy map and a sharpened question. Run a failure-mode census on it, budget thirty minutes.

  1. List the steps. Lay your project out along the seven steps and mark the three where AI is most deeply involved. The census covers those first;
  2. Ask two questions of every step. Which AI attribute am I consuming at this step (fluency / sycophancy / corpus prior)? At this step, which class of error from the table in section 11.4 is that attribute most likely to grow into? Put the answers into the census sheet, a step × error type matrix, and the templates have a fillable version;
  3. Write one signature for every high-risk cell. No copying the book's own wording. Translate it into a concrete signal in your project. What score exactly counts as "the score is too good to be true" here? Which entries exactly are "the most on-point citation" here? A cell where you cannot write a concrete signal means you have not yet worked out what the error there looks like;
  4. Circle two cells. One is the cell where being wrong costs most, and that is your operating room. The other takes honesty, the cell holding the conclusion you are currently most excited about. By the law in section 11.5, that is the high-incidence zone for motivated collusion, the place where your checking should double and where you are in fact most likely to cut corners;
  5. The closing self-test. With the filled sheet in hand, answer one sentence. If this project blows up three months from now, which cell is it most likely in, and what does the crime scene look like? Answer that and the census passes.

The census sheet and the "have AI self-check for failure modes" prompt set, complete and fillable, are in this chapter's templates. The use warning in there is a direct corollary of section 11.5. Do not skip it as boilerplate.

Want an agent to run it with you? Paste this to your AI assistant or coding agent:

Help me run the Chapter 11 failure-mode census, budget thirty minutes. Following Template 1 in docs/appendices/ch11-templates.md,
build an empty step × error type matrix, then walk the seven steps asking me two questions each, which AI attribute I am consuming at this step and which class of error it is most likely to grow into. I fill the answers in.
The signature for every high-risk cell is written by me as a concrete signal in my project. You do one thing only, send it back if I copied the book's wording.
The cell holding the conclusion I am most excited about is circled by me. You may not circle it for me, and you may not say that conclusion looks fine. When the census is done, run the
"have AI self-check for failure modes" prompt set from Template 2, and before running it read out the use warning in the templates. It checks form. It cannot check my motive.
If any command errors, stop and show me the output.

11.8 Sober reminders

  • This classification list will itself expire. Row 7 of the transfer map, new capability manufactures new error types in bulk, is a regularity still running, not a summary of history. The next generation of model capability will manufacture errors that are not on the table in section 11.4, the way the error type "the formula references one cell off" did not exist before the spreadsheet was invented. So what this chapter gives you is classification instinct, not a complete list. When a new error surfaces, you can ask "which attribute hit which step" and add a row to the table yourself. People who can add their own rows do not fear an expiring list.
  • Recognizing is not preventing. This chapter has done diagnosis and nothing else. The full defense is in Chapter 12, layered verification, the independent channel, and check dispatch all live there. The red-team routine for having AI attack your conclusions systematically was delivered in Chapter 10. Do not hold this chapter's list of signatures and feel safe. A signature only tells you where to look.
  • Nobody has measured the relative incidence of the four classes. The order in this chapter follows the steps, not a ranking by frequency or by damage. There are hard measurements in places, the Walters and Wilder fabrication rate cited in section 11.4 is one, and that is a number for one class of error on one step. Which class is most common in real research workflows and which causes the most loss still has no unified measurement covering the whole workflow. This line is still exploring, and it is a row that will appear on the honest map in Chapter 13.
  • Motivated collusion has no tool solution, still exploring. Fluency is handled by a checking step, sycophancy by an independent channel, corpus prior by fresh data. Motivated collusion cannot handle itself, and "itself" is exactly the problem. It has only a process solution. Let a channel that does not share your motive, another person, or a verification agent that does not know the answer you expect, touch the conclusion you are most excited about. That is how this book got saved that one time (section 11.5). Turning it into your standard equipment is Chapter 12's job.

11.9 The unfair advantage you now hold

Handed any polished research output, your own included, you can say within ten minutes which class it is most likely broken in and where to poke. Check the most on-point citation, ask whether the criterion was written before the result or after, look at whether the score deserves suspicion, or ask the question again from the opposite side. Most people facing the same output have two settings available, "something feels off" and "looks fine to me."