Chapter 5 · Questions and Hypotheses
Chapter companion
📋 Chapter 5 templates · 🗂 Template index · 💻 code/persona-panel
This chapter's ladder. At this step AI is stable at assistant level. Batch-generating candidate questions, candidate hypotheses, and counterexample angles can all be handed off. But this step has no half-open door. "Which question is worth answering" has no reliable automated path, still exploring. Any "autonomous" claim you see at this step is mostly marketing.
Spine update. Last chapter I found a brawl, and watched the "gap" I had found for myself get narrowed a notch by a citation check. This chapter I grind the surviving question into a sentence I dare sign, and the word "tie" gets forced into a number.
This chapter delivers. Three things. The five steps for sharpening a question, one hypothesis with a falsification condition, and a question-sharpening card.
5.1 Two weeks of reading, still not started
Your boss drops a line in the weekly meeting, "look into whether we can rebuild the support workflow with agents." The question is as big as the weather. You cannot take it and you cannot refuse it. You open a new document titled "Agent Research." A week goes by and the body is still empty. You have not been idle. You just do not know which first stroke would not be the wrong one.
The academic version is the same play. A PhD student says at group meeting, "I want to study interpretability in LLMs." The advisor nods, good direction, go pin down the question. Two weeks later his progress is a reading list thirty papers longer. The advisor asks what the question is, and he says he needs to read a bit more.
Both scenes have the same root illness. What you hold is a direction, not a question. Direction and question differ by whether there is a finished state, not by size. A direction has no finished state. Under the word "interpretability" you can always "read a bit more." A question has a finished state. It gets answered, or it gets falsified, and the thing is over. People stuck in a direction are not necessarily lazy. Squeezing out a question is a hard step with no feedback. Think in silence for three days and you cannot tell whether what you squeezed out is gold or waste paper, so the safest move is always to go back to reading.
Chapter 3 defined this step, grind a blur of curiosity into a question worth answering, one whose answer might embarrass you. In Chapter 4 you got the controversy map. The map tells you where the battlefield is. It will not write the declaration of war for you. This chapter teaches the step from map to declaration.
5.2 From squeezing out questions to filtering them
The way AI rewrites this step is direct. The cost of producing candidate questions has collapsed. Growing a question out of a direction used to take talent plus soaking, staying in a field until one day a question surfaced on its own. Now you hand AI the direction, the controversy map, and your constraints, and in a few minutes you get twenty candidates back. Much of it is silt, but for the first time the candidate space is fully on the table. Your work shifts from making something out of nothing to trimming and choosing. That is a change of trade, not a demotion.
Filtering needs a sieve. A good question passes three tests, and none can be skipped.
- Testable, you can say what evidence would kill it. A question that cannot be killed is not a question, only a position.
- Worth answering, the answer has to change action. If you do the same thing whether the answer is yes or no, the question is decoration.
- Affordable, the road from question to evidence is walkable with your time, money, data, and skill. A good question out of reach belongs to someone else.
The three tests frame the human/machine line for this step. Generating candidates can be handed off. AI enumerates faster than you and brings more angles. Judging which candidate deserves the next few weeks of your life cannot be delegated. The reason is that the three variables in Chapter 3, section 3.4 all press onto the human side at this step. "Is it worth answering" has no cheap ground truth, picking the wrong question raises no error, and picking again is priced in months. That is where this chapter's ladder line comes from.
The half you hand off has a hidden pit too, named the mirror risk. AI's candidate questions come from the corpus, and questions in the corpus are by definition questions already asked. Chapter 2's transfer failure condition four said the same thing. So the candidate space AI lays out has a systematic shape. "Already asked" is covered densely, "nobody has asked" thinly. The practical corollary follows. Use AI's candidate list as a floor, not a ceiling. The genuinely novel one or two, you will mostly have to add yourself.
5.3 Lessons stolen from writing issues
You have dispatched work to a coding agent many times. An issue, a spec, that is the question you feed the agent. The tuition you have paid on this step comes to three lessons. Below, only where each carries over to research and where it breaks.
Lesson one, the clearer the question, the less rework. Throw "this page feels slow, optimize it" at an agent and it optimizes the wrong place. You know this. The research side is identical. Feed "I want to study X" to AI and back comes fluent, hollow skimming. Feed it a question with criteria and back comes evidence. Chapter 4's lesson four said the unit of "can ask it something" is the question, not the document, which is the front of this lesson. This chapter is its back. The quality of the question itself is the first load-bearing wall of the whole workflow. Where it breaks. Write a bad issue and the agent's code fails the tests, and you know the same day. Sharpen a question badly and it takes weeks before you notice you are precisely answering a question nobody cares about.
Lesson two, fire a tracer bullet first, sharpen the smallest testable version first. Do not expect to grind a "perfect question" at your desk in one pass. Grind the smallest version you can start testing today, and let the evidence flow back and correct the question itself. This also explains again what Chapter 3 said, "the seven steps are a loop." Your first version of the question is almost certain to be sent back for regrinding by a later step. That is part of the process, not a failure. Where it breaks. A tracer bullet in code flows back every few hours. In research one round takes days to weeks, so the first round has to be smaller.
Lesson three, where candidates get cheap, judgment gets expensive. Row 8 of the transfer map wrote it. Search engines devalued memory and raised the value of "what to look up and what to trust." Judgment did not die. Judgment went up in price. The same pattern replays at this step. Candidate questions go from scarce to surplus, and the scarce resource moves, from "being able to think of a question" to "being able to see which question is worth answering." Read that row in full, only generation was devalued. Hand judgment over to AI along with it and you have opted out of this round of appreciation.
5.4 The five steps for sharpening a question
Below is a workflow you can copy as is. The input is a direction plus the controversy map you built in Chapter 4. The output is one hypothesis with a falsification condition, or an honest "it will not sharpen." Budget one hour.
Step 1, diverge candidates. Have AI put the candidate space on the table.
My direction: [one sentence].
Background: here is my controversy map: [paste or summarize the table from Chapter 4].
My constraints: [time budget / available data and hardware / the edge of my skills].
Generate 20 candidate research questions. Requirements:
1. Cover different grain sizes, from "worth a paper" to "worth an afternoon";
2. Cover different positions, at least 5 of them questions the opposition or a skeptic would ask first;
3. Attach to each question one line, "what evidence could refute it." If you cannot write that line, do not list it.
At the end, mark separately which questions have most likely already been asked in the literature, and what the clue is.
That last requirement is a detector fitted for the mirror risk. It makes AI confess which items on the list are recitation. Candidates marked "already asked" still have a use, they amount to a free novelty check. Your novelty space sits near the lines that were not marked.
Step 2, score the three tests. You do it, AI sits as a juror. Run each candidate through the three tests and score 0, 1, or 2 on each. The value of scoring is that it forces you to state a reason. The score itself is secondary. Any candidate you dare give a 0 on "testable" is out on the spot, and the top three by total go to the next step. You can have AI take the other side ("attack the testability of this question"), but the pen stays in your hand.
Step 3, rewrite it falsifiable. Rewrite the survivors into a falsifiable shape. Practice this sentence pattern.
Under [conditions/basis], the [measurable metric] of [subject], compared with [control], is [direction and threshold].
If [specific observed outcome] is observed, the hypothesis is falsified.
If you cannot write both lines, it is still a direction. Go back to step 2. The most common disease here is the wish sentence, "explore the possibility of X," "validate the effectiveness of Y," and what they share is that no observation could make them wrong. The cure is to pour nouns into the sentence. Which metric, compared with whom, how big a gap counts.
Step 4, rehearse "does the answer change action." Assume the answer is yes, then assume it is no, and write down what you, or your reader, or your boss, would do differently in each case. Same action either way? The question is decoration. Go back to step 2 and take the next one.
Step 5, reality-check the resources. Take the falsifiable sentence and ask one last time. Is the road to that "specific observed outcome" walkable for you? Can you get the data? Are the budget and the compute enough? Are the skills you need inside the range of you plus AI? If it is not walkable, shrink it. Shrink the task scope, shrink the metric, and a control arm can be cut too, until it is walkable, or until you honestly admit that this question does not belong to you right now.
The output of the five steps lands on a question-sharpening card, with fields for direction, candidates, three-test scores, the falsifiable sentence, the action rehearsal, the resource check, and the sign-off date. The full fillable version of this chapter's template is in the appendix.
5.5 Spine update · forcing "tie" into a number
The sharpening below is a synthetic narrative reconstructed after the fact, with the sequence rearranged for teaching. The hypothesis H at the end, the sentence this chapter grinds out, together with its criteria and the reasons for the values, is a recorded fact. Now watch the five steps run. My starting point is the rewritten controversy map at the end of Chapter 4. The empty ground it left was the decision-grade comparison nobody had done in full, and my private odds were that narrow tasks have a shot and a general tie is doubtful. But odds are intuition, and intuition cannot be preregistered. Preregistration means locking in the criteria and filing them before the run, with no changes after. To go further, "can small models team up" had to become a sentence the data could condemn to death.
At the diverge step AI taught me a lesson first. I fed in the controversy map and the constraints and asked for twenty candidates, and more than half of what came back was recitation of the literature brawl. "The effect of debate rounds on accuracy," "the scaling curve for the number of agents," all battlefields Chen 2024 and Smit 2024 had fought over long ago, the two papers from the skeptics' side on last chapter's controversy map. The mirror risk needs no argument. It is the default shape of a candidate list. What was worth money in this round was the "already asked" column. It pushed me back to the empty ground left after Chapter 4's narrowing.
The direction did not change. Three words had to be sharpened, task family, cost alignment, tie.
Task family knocked out a tempting option first. The candidate I liked most in the first version was "can the army, meaning the small-model team system, tie GPT-4o on the AlpacaEval benchmark." Chapter 4 checked it. Mixture-of-Agents scored above GPT-4o exactly there, a ready-made comparison target.
It died on "testable" in the three-test scoring. AlpacaEval uses an LLM as judge. The judge's taste for long answers has a paper devoted to correcting it, and the correction itself was criticized as incomplete (arXiv:2404.04475), a limitation that hangs on the Mixture-of-Agents paper card from Chapter 4. Test the army on a task judged by an LLM and what am I measuring, "the army is stronger" or "the army is better at pleasing the judge"? There is no telling them apart. Judge preference would contaminate the measurement into a second research question.
One question with two unknowns in it equals zero answerable questions. So I set an exclusion rule. No open-ended generation judged by an LLM. Scoring has to be programmatic, cheap, and uncontested.
Out of the space that remained, I chose three families. One family alone will not do, since my question is exactly "which tasks have a shot." All of them will not do either, affordability goes bankrupt on the spot. I took three that sit far apart in type, ① math and logic reasoning (a contamination-resistant variant set) ② knowledge Q&A (an MMLU-Pro subset, a knowledge quiz in multiple-choice form) ③ code (a HumanEval+ or LiveCodeBench slice, two programming test sets that score themselves). The three families stand for reasoning, knowledge, and automatically verifiable generation, and every score goes through a program. The first family says "variant set" on purpose, not the original problems. The originals may long since have entered the training corpus, and Chapter 8 opens that contamination account. Team topology is the independent variable, taking the three plays the two camps in the literature have fought over, voting (sample and take the majority), debate (multi-round mutual review), division of labor (role decomposition).
Then the hardest word, "tie." Right now it is a marketing word anyone can claim. Forced into a number, it means a gap ≤ ε counts as a tie. How big ε should be is the first thing I signed in this case. Set it to 0 and you demand scores match to the decimal, when the jitter across repeated runs alone is bigger than that. Set it to 5 points and "clearly worse" counts as a tie, and the skeptics laugh first. I signed ε = 2 percentage points, with the reason written down for the record. Below that gap, a team really choosing a technology mostly will not change its decision. Above it, the word "tie" does not deserve its name. This is a judgment, not a theorem. Together with all the criteria it gets its final sign-off at the Chapter 6 preregistration.
The action rehearsal in step 4 gave me an unexpected bonus. Rehearsing "yes" and "no" both went smoothly. H holds, local small-model teams become a serious candidate in enterprise technology selection, and my memo would recommend a pilot. H is falsified, everyone saves the trouble and the paid API continues. But halfway through the rehearsal I ran into a middle outcome. The army neither ties nor loses badly, it catches up at three times the cost. My first version of the hypothesis was completely silent on that outcome. It wrote a gap threshold and no cost boundary, so "tie" could be bought with unlimited money. A tie bought at more than double the money is not a tie, it is burning cash. So the falsification condition gained a cost clause. Lesson noted in passing, the action rehearsal rehearses not only the answer but the falsification condition itself.
A job budgeted at one hour actually took close to two, recorded honestly. The final version follows, where pp is short for percentage points.
Hypothesis H, under cost alignment, the accuracy gap between an open-source small-model team system and a single frontier model on the selected task families is ≤ ε (ε = 2 percentage points), which counts as a "tie."
Falsification condition, if the army trails by > 5pp on all three task families, or catches up only at > 2× the cost, H is falsified.
Note the shape of the falsification condition. Losing one task family does not kill H, since "which tasks have a shot" was part of the question to begin with. Total defeat, or catching up only by burning cash, is what earns a death sentence. Reading is done family by family with no aggregation across families, because a total averaged over three families represents nobody's decision. There is also a gray band. A single family trailing by 2 to 5 percentage points is neither a tie nor a falsification, and it sits between the tie line (≤2pp) and the falsification line (>5pp). Its verdict gets locked in when Chapter 6 writes the criteria. One word has still not been cashed in. How exactly is cost alignment accounted, in dollars or in compute, and how is amortization defined? That is the first pillar of the test design, and Chapter 6 opens it.
The three readings together make the table below. Read it family by family, and it goes into the file as is when Chapter 6 writes the criteria.
| Result on a single task family | Verdict |
|---|---|
| Trails by ≤ 2 percentage points | Tie |
| Trails by 2 to 5 percentage points | Undecided, neither a tie nor a falsification, report it as is, no picking a side |
| Trails by > 5 percentage points | This family lost. All three families lost, or catching up only at cost > 2×, and H is falsified |
5.6 The same rasp on another field, persona research
The book's second case line enters here, and it comes back in Chapters 6, 8, 11, and 12. It is one concrete shape of the question, can synthetic data stand in for real data? Product teams use persona-bearing models for user interviews and survey rehearsals, the "a 25-year-old mom with one child who works in manufacturing" kind, and save a round of recruiting real people, in time and in money. The answers read a lot like a real person. So the question arrives, can AI personas replace interviews with real people? That proposal on your desk to "run it with synthetic users before launch" is the same question.
Hold it up to the three tests and it is a direction, not yet a question. "Replace" cannot be killed. Supporters demo ten interviews indistinguishable from the real thing, opponents point out ten distortions, and both sides can talk past each other forever. Run the same five steps on it.
The questionnaire itself does not matter in this case. Watch the process. Rewriting it falsifiable, first swap "replace" for something measurable, the distribution of persona answers against the distribution of answers from real people. That needs a ready-made control. Some large public survey question bank, where how real people answered is on the record, let the personas answer the same set, and compare the distributions. The first falsifiable sentence takes shape. Does the agreement between a persona model's answer distribution on a public survey question bank and the answer distribution of the corresponding real-population subgroup reach a preset threshold?
The "questions the opposition would ask" from the diverge step earned its keep here. The opposition's hardest question, the averages match, is that enough? If every "25-year-old mom" gives a textbook-consistent answer, matching averages become the most dangerous illusion of all. The decision maker holds a stereotype recitation machine and takes it for a population. So the question gained a second clause, and no systematic stereotype drift appears (answers for a subgroup more extreme and more homogeneous than real people's).
The action rehearsal passed cleanly. Both clauses met, persona rehearsal can enter the formal research process as a coarse screen. Either one missed, it is only a toy for brainstorming, and its output is banned from decision documents. The resource check walks through too, public survey data is free to get. Which question bank to use as ground truth, the real answers used for checking, and where to set the threshold, that is test design work, and Chapter 6 picks up the hook.
What this demonstration is really about lives in the contrast between the two lines. The spine compares models, the subplot compares data, one substitute is a small model and the other a synthetic person, and the same five steps carried both through. What this process sharpens is the shape of the question. The metric is measurable, the control is ready-made, the falsification condition can be written, and domain knowledge counts for little of it.
5.7 Swap in your project
Last chapter you built a controversy map. Now give it a declaration of war. Budget one hour, and run the clock.
- Write down your direction, however vague. This is raw material, not output;
- Run the diverge prompt (the template in 5.4), get 20 candidates, look first at the "already asked" marks, your novelty space sits near the lines that were not marked;
- Score the three tests, pick the top three, and remember "testable" is a veto;
- Write a falsifiable sentence for each, and whichever you cannot write is out on the spot;
- Run the action rehearsal and the resource check on the survivors, put the last one standing into the question-sharpening card, and sign the date.
The five steps can also end in total defeat, with no candidate clearing all three gates. The hour was still not wasted. You just confirmed that this direction will not yield a question of your own right now, and what you saved is the weeks you would have spun in place. Go back to Chapter 4 for a different battlefield, or loosen the resource constraint, and run another round.
Want an agent to run it with you? Paste this to your AI assistant or coding agent:
Help me run the Swap in your project of Chapter 5, clock running, one hour. I give you a vague direction, and you use the diverge prompt of
step 1 in docs/appendices/ch05-templates.md to produce 20 candidate questions, marking each "already asked" or "not seen," with a source for the mark, and "unsure" when you cannot give one.
The three-test scoring is mine, you only write my scores into the table, testable is a veto, and if I forget to veto you remind me. The falsifiable sentence is mine to write, and you judge
one thing only, whether the sentence says "what counts as losing." Run the action rehearsal and the resource check with me using the prompts of steps four and five,
and the last one standing I put into the question-sharpening card and sign the date. If any command errors, stop and show me the output.
5.8 Sober reminders
- Mistaking a direction for a question is the most common crash at this step, and the person rarely feels it, since a direction can keep you very busy. There is only one test. Does it have a finished state? If you cannot say "under what conditions this thing is over," what you hold is still a direction.
- Wish-list hypotheses come second. "Validate the effectiveness of X," "explore the potential of Y," grammatically research, logically a wish, with no observation that could make them fall through. The falsifiable sentence pattern (5.4, step 3) is the targeted antidote. A statement that will not fit that pattern, do not call it a hypothesis.
- Letting AI pick your topic turns the knob past the safety line. This step has no cheap ground truth. When AI says "I recommend number 3," hold it up to the four criterion questions from Chapter 3. Who answers when it's wrong? Nobody. Stack the mirror risk on top and what it recommends is usually the mode of the corpus, while what you are looking for sits by definition at the corpus edge. Generation handed off, the ruling taken back, a line this chapter has now drawn three times.
- Over-sharpening is the pit in the other direction. Sharpening a question can become an advanced form of procrastination, the question always one round of polish short, which conveniently means no work has to start. The five steps are budgeted at one hour. Run well over and what you lack is the first tracer bullet (lesson two), not a better question. Walk into the next step with a "good enough" question and let the evidence keep sharpening it for you.
-
Still exploring deserves its own entry. Research systems that generate hypotheses automatically, some claiming end-to-end "propose and verify a finding," are iterating fast. Look closely at the loudest results and every one carries qualifiers. Sakana's AI Scientist-v2 passed a workshop at ICLR 2025, meaning a session hanging under the main conference with a lower bar, and the review was only semi-informed, reviewers knew AI papers were mixed into the batch but not which ones. Zochi claims acceptance at the ACL 2025 main conference, the highest venue, meaning place of publication, among this batch of systems, but the result is vendor self-reported with no independent audit. The overall signal is still weak, and the case-by-case interrogation is in Chapter 13.
You also have to see clearly how this class of system actually wins. For the places they take in public competitions, the ideas come almost entirely from human published papers and community discussion, and their most striking skill is picking up ideas others abandoned as too hard to implement, combining them, landing them. That is an extension of execution, strong and valuable, but it is not asking questions. Packaging combinatorial search as "autonomously proposing research directions" is the most common line on this front. As of this writing they are good at mass-producing candidates in the neighborhood of solved problems, which is assistant level doing its job. On the judgment of "which one is worth answering," no system's performance makes me willing to write a name here. The book's online case library tracks this front.
5.9 The unfair advantage you now hold
Give you any vague direction and one hour, and you can grind out a question with a falsification condition, a rehearsed action, and a walkable resource path, or honestly find that it will not grind. Both outcomes are worth more than "I need to read a bit more."