Working notes, not prose. This page is a research packet for drafting the chapter: the brief, source material, questions, examples, reader perspectives, and the earlier draft.
Brief
- Must: Walk through observe, explain and test with enough care that the reader sees why one result rarely settles a cause and why retrieval keeps knowledge usable.
- Serves: It is the mode that keeps our maps answerable to the territory.
Quote options
-
“No isolated experiment, however significant in itself, can suffice for the experimental demonstration of any natural phenomenon… a phenomenon is experimentally demonstrable when we know how to conduct an experiment which will rarely fail to give us a statistically significant result.” — R. A. Fisher, The Design of Experiments (1935; 2nd ed. 1937), pp. 13–14 (p. 16 in 2nd ed.). [verified: https://things-people-say.blogspot.com/2014/03/fisher-in-1937-p-16-of-design-of.html ; scan: https://gwern.net/doc/statistics/decision/1937-fisher-thedesignofexperiments.pdf] Why: States the chapter’s core claim outright: one result rarely settles a cause; a reliable procedure does.
-
“non-reproducible single occurrences are of no significance to science” — Karl Popper, The Logic of Scientific Discovery (1934; Eng. 1959), §8, p. 66 (Routledge 1992 ed.). [verified: https://en.wikipedia.org/wiki/Reproducibility ; https://pmc.ncbi.nlm.nih.gov/articles/PMC2981311/] Why: Short, epigraph-ready version of the same point from the book’s own Popper reference. (It is a sentence fragment; the original opens “we may say that…”)
-
“if you can spray them, then they are real.” — Ian Hacking, Representing and Intervening (1983), ch. 1 (Hacking’s fuller wording is “So far as I’m concerned, if you can spray them then they are real”; check the page against the book). [verified: https://en.wikipedia.org/wiki/Entity_realism] Why: Makes “manipulate” the test of reality: intervention, not only representation, keeps the map answerable to the territory.
-
“Essentially, all models are wrong, but some are useful.” — George E. P. Box & Norman R. Draper, Empirical Model-Building and Response Surfaces (1987), p. 424. The idea first appears as a section title in Box, “Robustness in the Strategy of Scientific Model Building” (1979). [verified: https://en.wikipedia.org/wiki/All_models_are_wrong ; https://gwern.net/doc/statistics/decision/1979-box.pdf] Why: Covers the “model” step: explanation is always a useful simplification that stays open to revision.
-
“Testing is a powerful means of improving learning, not just assessing it.” — Henry L. Roediger III & Jeffrey D. Karpicke, “Test-Enhanced Learning,” Psychological Science 17(3) (2006), p. 249 (abstract). [verified: https://learninglab.psych.purdue.edu/downloads/2006/2006_Roediger_Karpicke_PsychSci.pdf] Why: Backs up the retrieval claim: you keep knowledge usable by pulling it back out, not by reading it again.
Research
Core argument
- Start from a failure, not a textbook. Learning begins when a prediction breaks (the loaf didn’t rise, the child’s tower fell, the medicine didn’t work). Peirce’s point: doubt is an irritation, and inquiry is the struggle to get rid of it. The temptation is to end the irritation fast with the first satisfying story (“the new yeast is bad”).
- Measure: deciding what counts as a difference. Measuring is a cut (Act II): it picks one quality, gives it a scale, and ignores everything else. That choice happens before any data exists. Scales differ in what they allow (Stevens: naming, ranking, intervals, ratios), and it’s easy to do arithmetic a scale can’t support (“twice as happy”). A dated note and an honest comparison count as measurement; a thermometer is optional.
- Model: a proposed relationship that sticks its neck out. A model says “if this, then that” and so predicts things not yet seen. It earns trust by risky predictions, not by fitting what already happened. Several models always fit the same data (bad yeast, cold kitchen, rushed kneading). All models leave things out; the question is whether they leave out the right things for the job.
- Correlation is where the model starts, not where it stops. Observation alone leaves confounders (Pearl’s first rung, association). The barometer predicts storms but doesn’t cause them. Snow’s pump, Semmelweis’s clinics and Hawthorne’s lights each show how an observed pattern can have more than one cause.
- Manipulate: change one thing and watch. Intervening separates “goes with” from “makes happen” (Woodward: causes are handles for manipulating effects). Controls, blinding and randomization exist because the experimenter is also part of the system (Clever Hans; Feynman’s rats). A good manipulation cuts every route except the one being tested.
- Why one result rarely settles a cause. Any single result could come from noise, a hidden variable, the experimenter’s own cues, a measurement artifact or luck. Fisher’s standard is a procedure that reliably gives the result, not one striking result. The replication crisis puts a number on this: in 2015 only about 36% of 100 psychology replications came out significant. Hill’s point: no single criterion proves causation. What makes the case is several kinds of evidence that fail in different ways and still agree.
- Some causes can’t be tested by manipulation, and learning still goes on. You can’t move a star, rerun history or randomly assign smoking. Astronomers, epidemiologists and historians triangulate with natural experiments (Snow’s two water companies), dose-response evidence and the order of events. Intervention is the strongest route to causal knowledge, but it isn’t the definition of learning.
- Knowledge you can’t retrieve does no work. A map in a drawer doesn’t steer. Retrieval is a manipulation too: pulling knowledge out changes how well it stays (Roediger and Karpicke 2006). The feeling of familiarity misleads here too. Students who reread felt more confident and remembered less a week later. So the loop has to measure itself as well.
- The loop closes and turns. Results change what you look at next time. Sometimes you fix the procedure, and sometimes the result makes you rethink the goal (Argyris’s double loop, which Ch. 18 covers). Knowing when evidence is good enough for the stakes, and stopping to act, hands off to Creating (Ch. 15).
- Why this serves the book: this is how you avoid being cheated. When someone says X causes Y, ask how they measured it, what model they used, what they changed, and whether anyone has reproduced it.
Key questions
- When you say you “know” why something happened, how many other stories would also explain it?
- What did your measurement leave out, and could the cause be hiding there?
- Is a guess that never turns out wrong a good guess, or just a vague one?
- Why is a surprising prediction that comes true worth more than an old fact the model “explains”?
- If two things always happen together, how could you ever tell which one is doing the pushing, or whether a third thing pushes both?
- What does a control group actually control? (Child version: why does the scientist need a plant she doesn’t water differently?)
- Can you learn causes without touching anything? What do astronomers and historians do instead?
- Why did the horse seem to count? And how could the questioner be the thing being measured?
- If a famous study didn’t replicate, was the original a lie, a fluke, or a true result under conditions nobody wrote down?
- How many times should you see something before you believe it, and does the answer depend on what’s at stake?
- Is “more data” always better, or can it bury the one difference that matters?
- Why do we feel we know something right after rereading it, then find we can’t say it back?
- Is testing a way of checking knowledge, or a way of making it?
- (Cynic) Isn’t “the science” just whatever got funded, published and not yet disproven? How is that different from authority?
- (Cynic) If most single studies are shaky, why believe any of them? And how is that different from believing none?
- (Child) Why do grown-ups say “it’s because of X” when they didn’t check?
- When does going on investigating become a way of hiding from a decision?
- Can a measure change the thing it measures (Hawthorne, being watched, grades)? What does that do to learning?
- What would it take to change your mind about something you believe strongly? If nothing could, is it still knowledge?
- How is a cell or an animal doing measure-model-manipulate without words? And what does language add?
Examples
- Bread that won’t rise (kitchen scale). Keep it from the earlier draft as the chapter’s running example. Four changes at once (yeast, water temperature, kneading, open window) show confounding; two small batches with one change show manipulation; a second rerun shows why one result isn’t enough.
- Bacterial chemotaxis (cell scale). E. coli compares chemical concentration now with a moment ago (a measurement over time), “models” that the gradient is improving, and changes tumbling frequency (a manipulation of its own path). It’s learning without a mind, and a good bridge from Act I’s wanting cells.
- A toddler dropping food off a high chair (child scale). Developmental psychologists call it hypothesis testing. Gopnik’s “scientist in the crib” work shows infants run informal interventions. It’s the book’s “purely curious” perspective in action. (Check the specific studies before citing: Gopnik, Meltzoff & Kuhl, The Scientist in the Crib, 1999.)
- Clever Hans (animal plus observer). Berlin, c. 1904–1907. The horse “did arithmetic” by tapping. Oskar Pfungst (1907) varied whether Hans could see the questioner and whether the questioner knew the answer: accuracy was about 89% when Hans could see a questioner and about 6% when he couldn’t. The cue was an involuntary head movement by the humans. Lesson: the observer is part of the system, which is why blinding exists. https://en.wikipedia.org/wiki/Clever_Hans
- Feynman’s “Mr. Young” and the rats (lab scale). Young’s rats kept finding the “right” door by cues nobody had controlled for (smell, light, and finally the sound of the floor), until he laid the corridor in sand. Feynman called it “an A-Number-1 experiment” that others ignored. Lesson: a manipulation is only as clean as the routes you’ve cut off. https://calteches.library.caltech.edu/51/2/CargoCult.htm
- James Lind’s scurvy trial, HMS Salisbury, 20 May 1747. Twelve sick sailors in six pairs: cider, elixir of vitriol, vinegar, seawater, oranges and lemons, or a spicy paste. The citrus pair recovered fastest. It was a real controlled comparison, yet the Navy didn’t adopt lemon juice until 1795, and Lind himself didn’t fully believe his result. With two sailors per arm, one result settled nothing socially. https://en.wikipedia.org/wiki/James_Lind
- Semmelweis, Vienna General Hospital, 1847. First Clinic (doctors and students) maternal mortality averaged about 10%, the midwives’ Second Clinic under 4%. He tested and ruled out several hypotheses (crowding, the priest’s bell, birthing position) before Kolletschka’s scalpel death pointed to “cadaveric particles.” With chlorinated-lime handwashing from mid-May 1847, mortality went from 18.3% in April to 2.2% in June. He was still rejected, partly because he had no mechanism and didn’t publish well. Lesson: a working manipulation without a model people accept doesn’t spread. https://en.wikipedia.org/wiki/Ignaz_Semmelweis (Hempel’s Philosophy of Natural Science, 1966, ch. 2 uses this as the textbook case of hypothesis testing.)
- John Snow and cholera, London 1854. Broad Street outbreak: began 31 August, with 127 deaths in the first three days (confirm the final toll; sources differ, roughly 500–616). The pump handle came off on 8 September. Snow himself admitted the outbreak was already waning, so the handle removal proved nothing on its own. The stronger evidence was his “Grand Experiment”: households on the same streets were supplied by Southwark & Vauxhall (sewage-laden intake) or Lambeth (upriver), and the first had many times the cholera mortality. A natural experiment stood in for one he couldn’t run. https://en.wikipedia.org/wiki/1854_Broad_Street_cholera_outbreak
- The barometer and the storm (thought experiment). The needle predicts storms perfectly well. Force the needle down by hand and the storm still comes. That shows the difference between predicting something and controlling it. Pearl’s ladder: association, then intervention, then counterfactual. https://plato.stanford.edu/entries/causation-mani/
- Smoking and lung cancer, 1950–1965. Doll & Hill (1950) and later cohort studies. Nobody could randomize smoking, and Fisher himself argued for a confounding genotype. Hill’s 1965 “viewpoints” (strength, consistency, specificity, temporality, dose-response and others) show how several imperfect lines of evidence add up. Nice irony: the father of randomized experiments was on the wrong side of an observational question. https://en.wikipedia.org/wiki/Bradford_Hill_criteria
- The Reproducibility Project: Psychology (2015). 100 studies from 2008 journals: 97 originals significant, about 36% of replications significant, and effect sizes about half as large. Cancer biology (2021): about 26% could be replicated, and effects were about 85% smaller. Scale: a whole field learning that its single results weren’t settled. https://en.wikipedia.org/wiki/Reproducibility_Project:_Psychology
- Hawthorne Works, 1924–1932. Lighting went up and output rose; lighting went down and output rose again. Being measured changed the thing measured. Levitt and List’s 2011 reanalysis of the original data found only weak effects, so even the famous result about results needed replicating. https://en.wikipedia.org/wiki/Hawthorne_effect
- Stevens’s scales, and “twice as happy.” A 1–10 pain or happiness score is ordinal, so averaging it, or saying 8 is “twice” 4, assumes more than the scale gives you. Everyday measurement mistake. https://en.wikipedia.org/wiki/Level_of_measurement
- Roediger & Karpicke (2006), retrieval. Experiment 2: after 5 minutes, the students who only reread (SSSS) recalled most (83% vs 71% for STTT). After a week the order flipped: STTT 61%, SSST 56%, SSSS 40%. The rereaders were the most confident they’d remember. It’s a measure, model, manipulate experiment about memory, and the chapter’s retrieval payoff. https://learninglab.psych.purdue.edu/downloads/2006/2006_Roediger_Karpicke_PsychSci.pdf
- Ebbinghaus, 1880–1885 (one person as the whole lab). He memorized nonsense syllables (WID, ZOF), tested himself at intervals, and measured “savings” on relearning. It gave us the forgetting curve. Civilizational version: libraries, scriptures learned by heart, and the oral recitation traditions of the Vedas, which keep knowledge retrievable across generations (links to Ch. 19). https://en.wikipedia.org/wiki/Forgetting_curve
Source material
-
“The irritation of doubt causes a struggle to attain a state of belief. I shall term this struggle inquiry, though it must be admitted that this is sometimes not a very apt designation.” Charles S. Peirce, “The Fixation of Belief,” Popular Science Monthly 12 (November 1877), §IV. [verified: https://en.wikisource.org/wiki/Popular_Science_Monthly/Volume_12/November_1877/Illustrations_of_the_Logic_of_Science_I] Use: Opening. Learning starts from felt doubt, not curiosity in the abstract, and the risk is settling doubt too fast.
-
“the assignment of numerals to objects and events according to rules” S. S. Stevens, “On the Theory of Scales of Measurement,” Science 103(2684), 7 June 1946, p. 677. [verified: https://en.wikipedia.org/wiki/Level_of_measurement] Use: Measure section. The rules are a choice, which ties measurement back to the Act II blade. (Check against the Science original for exact punctuation.)
-
“causal relationships are relationships that are potentially exploitable for purposes of manipulation and control” James Woodward, “Causation and Manipulability,” Stanford Encyclopedia of Philosophy (first pub. 2001; rev.), §1. The same idea runs through Making Things Happen (2003). [verified: https://plato.stanford.edu/entries/causation-mani/] Use: Defines the Manipulate step. The entry also opens with causes “as handles or devices for manipulating effects,” which fits the book’s “handles” language from the Preface. Confirm the entry’s author credit and revision date on the page.
-
“The first principle is that you must not fool yourself—and you are the easiest person to fool.” Richard P. Feynman, “Cargo Cult Science,” Caltech commencement address, 1974; Engineering and Science 37(7), June 1974. [verified: https://calteches.library.caltech.edu/51/2/CargoCult.htm] Use: Why controls and replication exist. Pair it with the Mr. Young rats story from the same speech (“that is an A‑Number‑1 experiment”).
-
“the attacks had so far diminished before the use of the water was stopped, that it is impossible to decide whether the well still contained the cholera poison in an active state.” John Snow, On the Mode of Communication of Cholera, 2nd ed. (London: Churchill, 1855), Broad Street section (page to be confirmed). [verified: https://en.wikipedia.org/wiki/1854_Broad_Street_cholera_outbreak] (Wording is from Wikipedia’s quotation; check against the 1855 text at the UCLA John Snow site before printing.) Use: The iconic “one intervention” did not settle the cause, and Snow said so himself. That makes it the chapter’s best case for “one result rarely settles a cause.”
-
“none of my nine viewpoints can bring indisputable evidence for or against the cause-and-effect hypothesis and none can be required as a sine qua non.” Austin Bradford Hill, “The Environment and Disease: Association or Causation?” Proceedings of the Royal Society of Medicine 58(5), 1965, pp. 295–300 (the passage is near p. 299). [verified: https://en.wikipedia.org/wiki/Bradford_Hill_criteria] (Secondary source; the PMC scan https://pmc.ncbi.nlm.nih.gov/articles/PMC1898525/ is image-only, so confirm the page there.) Use: How a case for a cause gets built when you can’t intervene: separate lines of evidence, none sufficient on its own.
-
“However, on the delayed tests, prior testing produced substantially greater retention than studying, even though repeated studying increased students’ confidence in their ability to remember the material.” Henry L. Roediger III & Jeffrey D. Karpicke, “Test-Enhanced Learning,” Psychological Science 17(3), 2006, p. 249 (abstract). [verified: https://learninglab.psych.purdue.edu/downloads/2006/2006_Roediger_Karpicke_PsychSci.pdf] Use: Retrieval section. It complements the quote-options line already chosen: feeling that you know something isn’t evidence that you do.
-
“a sudden slight upward jerk of the head” Oskar Pfungst, Clever Hans (The Horse of Mr. von Osten), trans. C. L. Rahn (1911), as quoted in Wikipedia. [verified: https://en.wikipedia.org/wiki/Clever_Hans] (Secondary; confirm page in the Project Gutenberg text of Pfungst.) Use: The tiny, unintended cue that produced a whole “result.” Put it at the center of the blinding and controls passage.
-
“Experimentation has a life of its own.” Ian Hacking, Representing and Intervening (Cambridge UP, 1983), Part B, “Intervening” (commonly cited p. 150). [UNVERIFIED] Use: Manipulation isn’t just theory-testing; experimental craft builds knowledge in its own right. Check wording and page against the book.
Counterarguments and limits
- “Observe, explain, test” is a tidy myth of scientific method. Real inquiry is messy: theory often comes first and decides what gets measured (Kuhn, Hanson’s “theory-laden observation”), and discovery is often abduction or accident. Present MMM as a loop you can enter anywhere, not a recipe with steps in order.
- Popperian falsification is too simple. The Duhem–Quine problem: a failed prediction can always be blamed on an auxiliary assumption (the thermometer, the yeast storage). Lakatos: research programmes survive anomalies for good reasons. So “one result rarely settles a cause” cuts both ways, and one refutation rarely settles it either.
- Interventionism has limits. Many causes can’t be manipulated (race, sex, history, the Big Bang), and “ideal interventions” are idealizations. Don’t imply that anything you can’t manipulate isn’t knowledge. Astronomy, geology, evolutionary biology and history are real knowledge.
- The replication crisis can be overplayed. Low replication rates partly reflect underpowered replications, differences in context and publication bias. It isn’t proof that science fails. A cynical reader may slide from “single studies are shaky” to “nothing is known.” The chapter should block that move by showing how converging evidence (smoking, cholera) became settled.
- Measurement pessimism. Saying every measure “selects” can sound like every measure is arbitrary. Some measures are very stable across instruments and cultures (temperature, mass), and that sameness is itself evidence.
- The retrieval claim is narrower than it sounds. The testing effect is robust across many studies, but the original results used prose passages and free recall. It supports retention, not understanding or transfer, and effects differ with feedback, complexity and delay. Don’t make it a theory of wisdom.
- Anthropomorphizing cells. Calling chemotaxis “learning” or “modeling” stretches the words. Either label it as analogy or define learning broadly enough (a change in behavior from past input) and say so.
- Values and stakes. How much evidence is “enough” depends on what’s at risk (Heather Douglas, inductive risk). That isn’t a flaw in science, but it means the chapter can’t give a single threshold.
- Tacit and embodied knowledge (a baker’s hands, Polanyi) is learned without explicit models. MMM may overintellectualize craft; Ch. 15 and Ch. 21 should get that ground.
Connections
- Prologue: its method (observe, separate story from event, mark what is unknown) is the seed of this chapter; say so explicitly.
- Ch. 4: subjectivity and measured self-doubt. Clever Hans and Feynman are Ch. 4’s warnings turned into procedures. Also Ch. 4’s use of the Open Science Collaboration reference; don’t double up, cross-reference.
- Ch. 6: representation looping back on reality. Manipulation is the loop run on purpose to test the map.
- Ch. 9: measures replacing purposes (Campbell, Goodhart and Strathern). Hawthorne and teaching to the test are the flip side of retrieval practice.
- Ch. 10: measurement scales and diagnoses as thingification.
- Ch. 11: epistemology overview. This chapter is epistemology in practice.
- Ch. 12: mathematics proves; experiments don’t. Contrast certainty with reliability.
- Ch. 13: the dinner that introduces MMM. Reuse its scene or refer back to it.
- Ch. 15: stopping rules. “When is evidence enough for the stakes?” hands off to Creating.
- Ch. 16: Marvel as the receptive attention that precedes good measurement (noticing before measuring).
- Ch. 18: single versus double loop (Argyris), Kolb’s cycle, OODA. How learning loops can stall or run away.
- Ch. 19: retrieval at civilizational scale, meaning how knowledge stays alive.
- Ch. 20: program theory is Measure-Model-Manipulate for institutions.
- Ch. 22: review dates, and the retrieval exercise as a weekly practice.
Exercise ideas
-
One-change trial. Instruction: Pick something small at home that sometimes works and sometimes doesn’t (your sleep, coffee strength, a plant, a bread rise, how long your phone battery lasts). Write down your current explanation in one sentence. Then write down two other explanations that would fit what you’ve seen. Change exactly one thing and hold the others as steady as you practically can. Do it at least three times, and record the result each time on a dated line. Notice: How many rival explanations came easily. Whether your first result agreed with the second and third. Which variables you couldn’t hold still. Why: It runs the whole loop at the scale of a kitchen and shows directly why one result rarely settles a cause.
-
Close the book and say it back. Instruction: Right after finishing this chapter, close it. On a blank sheet, write the three verbs and, for each, one sentence on what it does and one example from the chapter. Don’t look back. Then check against the text and mark what you missed or got wrong. Repeat the next day, and again a week later, without rereading in between. Notice: The difference between how familiar the chapter felt and what you could actually produce, and whether the second and third attempts come more easily. Why: It lets the reader test the retrieval claim on this chapter. It also measures the reader’s own knowledge, which is the loop turned inward.
-
Interrogate a headline. Instruction: Find one news story that says “X linked to Y” or “X causes Y.” Write four short answers: What was measured, and how? What model links X to Y? Did anyone change X, or only observe it? Has it been found more than once? Then name one third factor that could produce both X and Y. Notice: How often the answer to “did anyone change X?” is no, and how the headline’s wording (linked, boosts, causes) shifts with that. Why: It turns the chapter into the book’s practical defense against being cheated. A reader who can ask these four questions is harder to fool.
Open questions for the author
- Is bread still the running example, or will you reuse Ch. 13’s dinner so the three MMM chapters share one scene?
- How far into philosophy of science do you want to go: Popper, Duhem–Quine, Kuhn and Lakatos named, or kept in the background for Ch. 11?
- Should retrieval be a full section, or a coda? It sits oddly beside causal inference unless it’s framed as “measuring your own map.”
- Do cells and animals “learn” in this chapter’s sense? Decide whether MMM applies below language, and say so. It matters for the Act III close (“unnatural in our loops”).
- Which historical case is the anchor: Snow (natural experiment, honest uncertainty), Semmelweis (right answer, rejected), or Lind (controlled trial, ignored for 48 years)? Using all three risks a catalog.
- How should the chapter handle the cynic’s move from “single studies fail” to “trust nothing”? Is that argued here or in Ch. 4?
- Should the Fisher–smoking irony go in? It’s a great story, but it’s about a real person being wrong and needs careful sourcing.
- Hacking’s “if you can spray them, they are real” makes manipulation a test of reality, not just of causation. Do you want that stronger metaphysical claim, or keep Manipulate epistemic?
- The earlier draft’s “When the evidence is adequate for the stakes…” handoff: keep it as the ending, or move stopping rules entirely into Ch. 15?
Reader perspectives
Curious young child
First reactions:
- The flat bread “like wet cement” is great. They’d want to poke it.
- They like that the first step is not saying “the yeast is bad.” That’s the moment a sibling gets blamed for something without proof.
- Making two small batches is the most exciting part, because it’s basically a science fair.
- They’d lose interest at astronomers and epidemiologists; those are words, not bread.
- The “close this chapter and remember the three verbs” bit is a quiz, and they’d groan. Then they’d be pleased when they got it right.
Questions they’d ask:
- “What is yeast? Is it alive? Is it eating the bread?” (It is alive. That’s a weird fact that would hook them.)
- “Why can’t you just change everything at once and see if it works?”
- “How do you know which thing made it go wrong if you changed four things?”
- “If it works once, doesn’t that mean it works?”
- “What’s a model? Like a toy car model?”
- “If you can’t move a star, how do you learn about stars?”
- “Why do I forget my spelling words even though I read them ten times?”
- “Why does a quiz help you remember? Quizzes are for checking, not learning!”
- “Can you measure being happy?”
- “What if you measure the wrong thing? Like measuring how tall my plant is when it’s actually sad because it’s thirsty?”
- “Can a result lie?”
Where they’d get lost, bored, offended or unconvinced:
- “A successful prediction is rarely a coronation” uses a word they don’t know to make a point they’d get instantly with a better picture: winning one game of rock-paper-scissors doesn’t mean you have a secret trick.
- “Intervention is one route to causal knowledge, not the definition of learning” is a mouthful. The star example underneath it is good; lead with the star.
- The Roediger and Karpicke study is the part most relevant to their actual life (school), yet it’s told in careful adult hedges. They’d want to know what the students did: some read the story again and again, some tried to remember it, and later the rememberers did better. Say that plainly (the hedges can stay).
- “Double-loop” question at the end (“identical loaf or a cook who can respond?”) is lovely, but they’d need it phrased as “do you want to follow the recipe or become a person who can cook without one?”
- They’d be unconvinced that one result “rarely settles a cause” unless they see it fail: the classic “I wore my lucky socks and we won” story.
Examples they’d bring:
- Lucky socks: “I wore these and we won the match!” Then they lose in them. One result didn’t settle the cause. Perfect for Fisher’s point.
- Paper airplanes: Fold two planes the same except one thing (the wing tip), throw them from the same spot. That’s Manipulate with one variable, done in a hallway.
- The bean in a cup: Every school does it. One by the window, one in the cupboard. That’s the two-batch bread experiment they’ve already done.
- Which snack makes the dog come fastest: Measuring with a stopwatch, guessing why (smell? crunch?), and testing. Also funny, and it has the problem that the dog gets full.
- Forgetting spelling words: A direct way into retrieval practice: covering the word and trying to write it beats staring at it.
- Mentos in cola: Every kid has seen the fountain video. Is it the Mentos, the fizz, or the shaking? Great for “several explanations fit the same result.”
What would win them over:
- Present the chapter as detective work. The bread is the crime scene; the four changes are suspects; you can’t arrest the yeast on a hunch.
- A real “try this at home” that works: two cups, two spoons of sugar, warm vs. cold water, same yeast, and watch which one foams. It’s safe, visual and about the chapter’s own example.
- Tell them the testing effect is a “brain trick” that makes studying shorter. They’d use it tomorrow.
- Admit that sometimes you can’t do the experiment (stars, dinosaurs), and show how people still find out.
Cynical adult
First reactions:
- “Observe, hypothesize, test. That’s the scientific method as taught to twelve-year-olds. The bread is cute, but I learned this in school.”
- The bread example is actually good: several changes at once, and the obvious culprit (the yeast) might not be it. That’s a real, familiar trap. It’s the one strong thing here.
- “One result rarely settles a cause” is true and important, and it’s precisely what the news, my boss and my own gut ignore daily. The chapter should hit that harder and in higher-stakes places than a kitchen.
- The retrieval-practice bit feels bolted on. Why is flashcard research in a chapter about causation?
- “Close this chapter and retrieve its three verbs” is the book testing me. I’ll do it once, maybe.
Questions they’d ask:
- “What’s here that isn’t in a high school science textbook?”
- “How many results do I need before I believe something? Two? Ten? Give me a number or a rule, not ‘rarely.’”
- “The news says ‘new study shows coffee causes X.’ What exactly should I do with that sentence using this chapter?”
- “If astronomers and historians can’t intervene, and they still count as knowledge, why is ‘Manipulate’ one of the three steps at all?”
- “Why is retrieval practice in here? Isn’t that a study tip, not epistemology?”
- “Isn’t ‘one result rarely settles a cause’ the exact line tobacco companies used for decades to cast doubt? How do I tell honest caution from paid-for doubt?”
- “What’s the difference between a model and a story I tell myself?”
- “Where’s the line between learning and over-analysis? The chapter says ‘eventually you have to bake.’ When?”
- “My doctor changes one thing at a time and it takes months. Is that good practice or just slow?”
- “What if the thing you’re measuring is the wrong thing? The chapter mentions it for a thermometer but doesn’t take it anywhere.”
- “Does Hacking’s ‘if you can spray them’ really apply to my bread, or is it just a clever quote?”
Where they’d get lost, bored, offended or unconvinced:
- The bread is too small to make “why one result rarely settles a cause” feel urgent. Nobody gets hurt by a flat loaf. The chapter needs at least one case where believing a single result cost something real.
- The retrieval section is under-motivated. The brief asks for it, but the draft transitions with “another quiet problem,” which reads like a to-do list item being ticked off. The link (knowledge you can’t recall can’t be tested or used) should be stated explicitly.
- “Manipulate means intervention, not domination”: the word still has a manipulative flavor. The cynic hears it and thinks of salespeople.
- The ending question “Are you trying to produce an identical loaf every week, or become a cook who can respond to a changing kitchen?” sounds like a motivational poster.
- Hedges like “under their experimental conditions” and “It does not become infallible” are correct but accumulate; too many and the reader feels the author won’t commit to anything.
Examples they’d bring:
- Andrew Wakefield’s 1998 Lancet paper linking MMR vaccine to autism: twelve children, later retracted (2010) and found fraudulent, yet it drove vaccine hesitancy for decades. The textbook case of one result being treated as settling a cause.
- Hormone replacement therapy: observational studies suggested it protected women’s hearts; the Women’s Health Initiative randomized trial (results published 2002) found increased risk of heart disease and stroke in the tested regimen. Observation vs. intervention, with real stakes.
- Amy Cuddy’s “power posing” (Carney, Cuddy & Yap, 2010) became a hugely popular TED talk; later replications (e.g., Ranehill et al., 2015) failed to find the hormonal effects. A lived example of a result becoming a brand before it was checked.
- A/B testing at work: a manager sees one test with a lift and ships it company-wide. Next quarter it disappears. Regression to the mean in an office.
- Personal: “I cut out gluten and felt better.” You also started sleeping more and cooking at home. That’s the bread problem in the reader’s own body.
What would win them over:
- Keep the bread as the teaching example but put a costly real-world case right next to it (Wakefield or HRT), so the stakes are clear.
- Give a short, portable checklist for single-result claims: Was it repeated? Was anything else changed? Who funded it? What would have counted as failure?
- Tie retrieval directly to not being cheated: if you can’t state what you know without the source in front of you, you can’t check the next person who contradicts it.
- Address “merchants of doubt” head-on (Oreskes & Conway, Merchants of Doubt, 2010) so caution isn’t mistaken for a weapon.
Believer / spiritual reader
First reactions:
- A believer who bakes will like the bread example, and may notice that leaven is one of scripture’s favourite images (Matthew 13:33; the unleavened bread of Passover). The loaf could carry more meaning than the draft gives it.
- The retrieval section will feel familiar. Religious traditions are the oldest retrieval-practice systems on earth: the Shema recited morning and evening, Qur’an memorisation, Vedic chanting, catechisms in question-and-answer form. The draft cites Roediger and Karpicke (2006) as if the idea were new.
- They’d be relieved by “intervention is one route… not the definition of learning” and by the point that learning “need not yield control.” That leaves room for knowledge you can’t get by experimenting on it.
- They’d be wary of where the chapter is heading. If manipulation is the gold standard, then God, grace and the sacred become the untestable things the book will later discount.
Questions they’d ask:
- Scripture tells me both “Prove all things; hold fast that which is good” (1 Thess. 5:21) and “You shall not put the LORD your God to the test” (Deut. 6:16). Your method has no second clause. Are there things one shouldn’t manipulate to learn about, such as people, relationships or the holy?
- If I can’t run a controlled trial on my marriage, is what I know about it second-class knowledge?
- You say “several explanations can predict the same loaf.” Believers and naturalists often predict the same observations. How does your method choose between them without adding a philosophical assumption?
- Hacking says “if you can spray them, they are real.” What about things that act on you but that you can’t act on? Doesn’t that make reality depend on our power over it?
- Is memorizing scripture learning, or only storing? Recitation in my tradition is meant to change the reciter, not just keep the text. Does your retrieval research measure that?
- When prayer studies find no effect, what does that show: that prayer doesn’t work, or that prayer isn’t a mechanism?
- Who does the measuring, and do they count the cook’s patience (which the draft admits the thermometer ignores) as real?
- Is a sign, or an answered prayer, “one result that rarely settles a cause”? Isn’t that what wise spiritual directors already say about discernment?
- When is further analysis “shelter from action”? Traditions call endless deliberation a failure of trust. Is that the same thing?
- What does learning do to the learner, rather than to the model?
Where they’d get lost, bored, offended or unconvinced:
- Offended: if the Popper line “non-reproducible single occurrences are of no significance to science” is used without the qualifier “to science.” For a believer, one-time events (a conversion, a death, a revelation at Sinai) can be the most significant things there are. Keep “to science” visible and say plainly that it’s a limit of the method, not a verdict on those events.
- Unconvinced: by retrieval practice presented as a finding from 2006 with no history. It will look like the book doesn’t know that liturgy and catechism have run on this principle for millennia. That undercuts Act III’s claim that practices come before patterns.
- Lost: at the jump from bread to “close this chapter and retrieve its three verbs.” It works, but a reader would welcome a traditional example of retrieval here to make the practice feel human rather than clinical.
- Bored: by bread as the only case. Add one example where the thing studied can’t be manipulated without harm, such as grief, a child or a sacred site, to show what the method does when intervention is off-limits.
Examples they’d bring:
- Daniel 1:12–15. Daniel proposes a ten-day trial of vegetables and water for himself and his companions, compared against the young men eating the king’s food, and the outcome is judged by appearance. It’s often cited as an early comparison trial, and it’s a scriptural example of Measure, Model, Manipulate.
- 1 Kings 18 (Elijah on Mount Carmel). Two altars, two bulls, the same conditions, and water poured on one side to make the test harder. It’s a deliberately controlled public test, and it shows scripture isn’t simply against testing.
- Deuteronomy 18:21–22. A prophet is judged false if what he predicts doesn’t happen: a built-in prediction test for religious authority, which is useful for the book’s anti-cheating theme.
- The STEP study (Benson et al., “Study of the Therapeutic Effects of Intercessory Prayer,” American Heart Journal 151(4), 2006). Intercessory prayer showed no benefit for cardiac-bypass patients, and patients certain they were being prayed for had more complications. Many theologians answered that prayer isn’t a dose-response mechanism. It’s a clean case for asking what a manipulation can and can’t show.
- Vedic recitation. Texts were preserved orally for centuries using pada, krama, jata and ghana patterns, which recite the same text in reordered and doubled word sequences as error-checking. UNESCO recognised Vedic chanting as intangible cultural heritage in 2003. It’s retrieval practice engineered for accuracy.
- Al-Ghazali, Deliverance from Error (al-Munqidh min al-Dalal). An 11th–12th century account of methodically doubting sense perception and reasoning before rebuilding knowledge. It’s a religious thinker running a Measure-and-Model crisis on his own mind.
What would win them over:
- Saying explicitly that some things are known well without being manipulated, and that some things shouldn’t be manipulated in order to be known. Deuteronomy 6:16 could sit next to Hacking.
- Crediting religious recitation traditions as the long history behind retrieval research.
- Presenting prayer studies (or similar cases) fairly: stating what they measured and what they couldn’t.
- Keeping the closing turn from “identical loaf” to “becoming a cook.” That’s the change-of-heart question believers care about.
Skeptical scientist
First reactions:
- This is my home turf, and the bread example is well chosen: several variables changed at once (yeast, water temperature, kneading, draught), with one outcome. That is confounding in a kitchen. But the draft never names it “confounding,” and the reader would benefit from the word.
- The two-batch experiment is a good start but leaves out the two things that make a comparison informative: replication (one batch each is n=1) and knowing how much loaves vary when nothing changes. Without a baseline you can’t tell a difference from noise.
- “Intervention is one route to causal knowledge, not the definition of learning” is correct and important. Astronomy and epidemiology deserve a concrete case rather than a name-check.
- The retrieval section is fair and properly hedged (“under their experimental conditions”). But the brief says retrieval keeps knowledge usable, and the Roediger and Karpicke result is about retention. Usability (transfer to new problems) is a weaker, more mixed literature.
- The Fisher quote in the list is the best thing about this chapter’s plan. Use it as the spine.
Questions they’d ask:
- How much does your bread vary when you change nothing? Without that number, how would you know the yeast mattered?
- Why change the yeast and not the temperature first? Which variable does the reader test when several are suspects, and how do they choose?
- Why is one result rarely enough? The draft says so but doesn’t give the three reasons: noise, confounding, and flexible analysis (trying things until something works).
- When you can’t intervene, as with smoking and lung cancer, how did we become confident anyway? Can the chapter show observational causal reasoning actually working?
- “Measure” assumes the instrument is trustworthy. How was the thermometer itself validated? Where does measurement bottom out?
- What is a “model” here: a verbal story, an equation, a mechanism, a statistical fit? Scientists use the word for all four.
- Does the chapter distinguish prediction from explanation? A model can predict well with the wrong mechanism (epicycles).
- Does the testing effect hold for understanding and application, or mainly for recall of studied material?
- Is “Measure, Model, Manipulate” really a sequence? In practice the model decides what you measure. Should the loop be drawn starting from Model?
- How does a reader avoid fooling themselves: blinding, writing the prediction down beforehand, deciding in advance what would count as failure?
- Do non-human organisms do this loop? Bacteria “measure” gradients. Where does the human version differ in kind rather than degree?
Where they’d get lost, bored, offended or unconvinced:
- Unconvinced: “If one rises and the other does not, the yeast explanation gains support.” With one batch each, it gains very little. That is the chapter’s own thesis, so the example should show it rather than wave at it: run it three times, or run a same-yeast pair as a control.
- Lost: “a successful prediction is rarely a coronation” is a lovely line, but the idea behind it (several models predict the same data) needs a concrete two-model example with a test that tells them apart.
- Bored: the retrieval section is a detour as written. Tie it to the loop: retrieval is the Measure step applied to your own knowledge. The draft’s “close this chapter and recall the three verbs” nearly does this; make that link explicit.
- Offended: “When the evidence is adequate for the stakes, continued analysis can become shelter from action” is fair, but a scientist will want to hear the opposite failure too (acting on a single striking result) with equal force.
- Loose terms: “evidence,” “support,” “model,” “cause,” “retention” vs “usable.”
Examples they’d bring:
- Measuring the measure. Hasok Chang, Inventing Temperature (2004), shows how thermometers were validated with no prior reliable thermometer to check them against, a bootstrapping problem. It gives “every measure selects” historical depth.
- Causal inference without intervention. Doll and Hill’s British doctors’ studies, beginning in the 1950s, and Bradford Hill’s 1965 lecture “The Environment and Disease: Association or Causation?”, which set out considerations for inferring causation from observational data. Snow’s 1854 Broad Street cholera map is the classic natural experiment.
- Why one result isn’t enough. Simmons, Nelson and Simonsohn, “False-Positive Psychology,” Psychological Science (2011), showed that flexible analysis choices can make almost anything come out “significant.” Begley and Ellis, Nature (2012), reported that Amgen scientists could confirm only 6 of 53 landmark preclinical cancer findings.
- The retrieval literature, and how far it reaches. Rowland, “The effect of testing versus restudy on retention,” Psychological Bulletin (2014), is a meta-analysis supporting the testing effect. Pan and Rickard, “Transfer of test-enhanced learning,” Psychological Bulletin (2018), found that transfer to new material is real but more limited and depends on conditions. Dunlosky et al., Psychological Science in the Public Interest (2013), rated practice testing and spaced practice as high-utility study techniques.
- The learning loop in biology. Berg and Brown, Nature (1972), tracked E. coli swimming: it compares concentrations over time and tumbles less when things improve, a loop with no model in it at all. Schultz, Dayan and Montague, Science (1997), linked dopamine signals to reward-prediction error, the brain’s version of “the model was wrong by this much.” Together they show how far down the loop goes and what humans add (explicit, shareable models).
- Two models, one prediction. Ptolemaic epicycles predicted planetary positions well. Separating geocentric from heliocentric models took new observations, notably Galileo’s report in 1610 that Venus shows a full set of phases. A clean case of a successful prediction that was no coronation.
What would win them over:
- Rework the bread experiment to show the thesis: a baseline of how much loaves vary, a controlled pair, a repeat, and one result that surprises and forces a model change.
- Name the three reasons one result rarely settles a cause (noise, confounding, flexible analysis), each with one real case.
- An honest split in the retrieval claim: strong evidence for retention, more conditional evidence for transfer. Then keep “usable” only in the narrower sense.
- A short passage on a non-human learning loop, so the Act III “unnatural in our loops” claim has a baseline.
Earlier draft
From “Learning: Measure, Model, Manipulate”
Imagine that last week’s bread rose beautifully and today’s loaf sits in the bowl like wet cement. You remember changing the yeast. You also used colder water, hurried the kneading, and left the dough beside an open window.
Learning begins by resisting the satisfying sentence the new yeast is bad.
Measure means becoming more careful about differences. Which yeast? How warm was the water? How long did the dough rise? Measurement need not involve a digital instrument; a dated note and an honest comparison may be enough. But every measure selects. A thermometer records temperature and ignores the cook’s impatience. What we choose to notice shapes what can later be explained.
Model means proposing relationships. Perhaps the yeast was inactive. Perhaps the dough was cold. A model earns trust when it helps us anticipate what we have not yet seen, but a successful prediction is rarely a coronation. Several explanations can predict the same loaf.
Manipulate means changing something to learn what difference it makes. Make two small batches, keeping the water, flour, kneading, and location as similar as practical while changing the yeast. If one rises and the other does not, the yeast explanation gains support. It does not become infallible: the packets may have been stored differently, and a kitchen is not a sealed laboratory.
Intervention is one route to causal knowledge, not the definition of learning. Astronomers cannot move a star to see what happens. Epidemiologists and historians often reason from observations they did not arrange. Even a controlled experiment depends on assumptions about what was held steady and how the result was measured. Learning can yield description, interpretation, prediction, or better uncertainty. It need not yield control.
There is another quiet problem: knowledge that cannot be recalled when needed does little work. Research by Henry Roediger and Jeffrey Karpicke found that, under their experimental conditions, retrieving material improved later retention compared with repeatedly studying it. This does not mean every quiz transforms understanding. It suggests a modest practice.
Close this chapter for a moment and retrieve its three ordinary verbs. Then ask what each contributes. Look back. The gap between what seemed familiar and what you could produce is useful evidence.
Learning loops because results alter the next observation. The loaf may teach you about yeast, or reveal that temperature matters more. Sometimes the deeper revision concerns the aim. Are you trying to produce an identical loaf every week, or become a cook who can respond to a changing kitchen? Correcting a procedure and reconsidering its goal are different achievements.
When the evidence is adequate for the stakes, continued analysis can become shelter from action. You eventually have to bake. That is where Creating begins.