Appearance
Text size

Chapter 14: Learning: Measure, Model, Manipulate

Observe, explain, test

Working notes, not prose. This page is a research packet for drafting the chapter: the brief, source material, questions, examples, reader perspectives, and the earlier draft.

Brief

Quote options

  1. “No isolated experiment, however significant in itself, can suffice for the experimental demonstration of any natural phenomenon… a phenomenon is experimentally demonstrable when we know how to conduct an experiment which will rarely fail to give us a statistically significant result.” — R. A. Fisher, The Design of Experiments (1935; 2nd ed. 1937), pp. 13–14 (p. 16 in 2nd ed.). [verified: https://things-people-say.blogspot.com/2014/03/fisher-in-1937-p-16-of-design-of.html ; scan: https://gwern.net/doc/statistics/decision/1937-fisher-thedesignofexperiments.pdf] Why: States the chapter’s core claim outright: one result rarely settles a cause; a reliable procedure does.

  2. “non-reproducible single occurrences are of no significance to science” — Karl Popper, The Logic of Scientific Discovery (1934; Eng. 1959), §8, p. 66 (Routledge 1992 ed.). [verified: https://en.wikipedia.org/wiki/Reproducibility ; https://pmc.ncbi.nlm.nih.gov/articles/PMC2981311/] Why: Short, epigraph-ready version of the same point from the book’s own Popper reference. (It is a sentence fragment; the original opens “we may say that…”)

  3. “if you can spray them, then they are real.” — Ian Hacking, Representing and Intervening (1983), ch. 1 (Hacking’s fuller wording is “So far as I’m concerned, if you can spray them then they are real”; check the page against the book). [verified: https://en.wikipedia.org/wiki/Entity_realism] Why: Makes “manipulate” the test of reality: intervention, not only representation, keeps the map answerable to the territory.

  4. “Essentially, all models are wrong, but some are useful.” — George E. P. Box & Norman R. Draper, Empirical Model-Building and Response Surfaces (1987), p. 424. The idea first appears as a section title in Box, “Robustness in the Strategy of Scientific Model Building” (1979). [verified: https://en.wikipedia.org/wiki/All_models_are_wrong ; https://gwern.net/doc/statistics/decision/1979-box.pdf] Why: Covers the “model” step: explanation is always a useful simplification that stays open to revision.

  5. “Testing is a powerful means of improving learning, not just assessing it.” — Henry L. Roediger III & Jeffrey D. Karpicke, “Test-Enhanced Learning,” Psychological Science 17(3) (2006), p. 249 (abstract). [verified: https://learninglab.psych.purdue.edu/downloads/2006/2006_Roediger_Karpicke_PsychSci.pdf] Why: Backs up the retrieval claim: you keep knowledge usable by pulling it back out, not by reading it again.

Research

Core argument

  1. Start from a failure, not a textbook. Learning begins when a prediction breaks (the loaf didn’t rise, the child’s tower fell, the medicine didn’t work). Peirce’s point: doubt is an irritation, and inquiry is the struggle to get rid of it. The temptation is to end the irritation fast with the first satisfying story (“the new yeast is bad”).
  2. Measure: deciding what counts as a difference. Measuring is a cut (Act II): it picks one quality, gives it a scale, and ignores everything else. That choice happens before any data exists. Scales differ in what they allow (Stevens: naming, ranking, intervals, ratios), and it’s easy to do arithmetic a scale can’t support (“twice as happy”). A dated note and an honest comparison count as measurement; a thermometer is optional.
  3. Model: a proposed relationship that sticks its neck out. A model says “if this, then that” and so predicts things not yet seen. It earns trust by risky predictions, not by fitting what already happened. Several models always fit the same data (bad yeast, cold kitchen, rushed kneading). All models leave things out; the question is whether they leave out the right things for the job.
  4. Correlation is where the model starts, not where it stops. Observation alone leaves confounders (Pearl’s first rung, association). The barometer predicts storms but doesn’t cause them. Snow’s pump, Semmelweis’s clinics and Hawthorne’s lights each show how an observed pattern can have more than one cause.
  5. Manipulate: change one thing and watch. Intervening separates “goes with” from “makes happen” (Woodward: causes are handles for manipulating effects). Controls, blinding and randomization exist because the experimenter is also part of the system (Clever Hans; Feynman’s rats). A good manipulation cuts every route except the one being tested.
  6. Why one result rarely settles a cause. Any single result could come from noise, a hidden variable, the experimenter’s own cues, a measurement artifact or luck. Fisher’s standard is a procedure that reliably gives the result, not one striking result. The replication crisis puts a number on this: in 2015 only about 36% of 100 psychology replications came out significant. Hill’s point: no single criterion proves causation. What makes the case is several kinds of evidence that fail in different ways and still agree.
  7. Some causes can’t be tested by manipulation, and learning still goes on. You can’t move a star, rerun history or randomly assign smoking. Astronomers, epidemiologists and historians triangulate with natural experiments (Snow’s two water companies), dose-response evidence and the order of events. Intervention is the strongest route to causal knowledge, but it isn’t the definition of learning.
  8. Knowledge you can’t retrieve does no work. A map in a drawer doesn’t steer. Retrieval is a manipulation too: pulling knowledge out changes how well it stays (Roediger and Karpicke 2006). The feeling of familiarity misleads here too. Students who reread felt more confident and remembered less a week later. So the loop has to measure itself as well.
  9. The loop closes and turns. Results change what you look at next time. Sometimes you fix the procedure, and sometimes the result makes you rethink the goal (Argyris’s double loop, which Ch. 18 covers). Knowing when evidence is good enough for the stakes, and stopping to act, hands off to Creating (Ch. 15).
  10. Why this serves the book: this is how you avoid being cheated. When someone says X causes Y, ask how they measured it, what model they used, what they changed, and whether anyone has reproduced it.

Key questions

  1. When you say you “know” why something happened, how many other stories would also explain it?
  2. What did your measurement leave out, and could the cause be hiding there?
  3. Is a guess that never turns out wrong a good guess, or just a vague one?
  4. Why is a surprising prediction that comes true worth more than an old fact the model “explains”?
  5. If two things always happen together, how could you ever tell which one is doing the pushing, or whether a third thing pushes both?
  6. What does a control group actually control? (Child version: why does the scientist need a plant she doesn’t water differently?)
  7. Can you learn causes without touching anything? What do astronomers and historians do instead?
  8. Why did the horse seem to count? And how could the questioner be the thing being measured?
  9. If a famous study didn’t replicate, was the original a lie, a fluke, or a true result under conditions nobody wrote down?
  10. How many times should you see something before you believe it, and does the answer depend on what’s at stake?
  11. Is “more data” always better, or can it bury the one difference that matters?
  12. Why do we feel we know something right after rereading it, then find we can’t say it back?
  13. Is testing a way of checking knowledge, or a way of making it?
  14. (Cynic) Isn’t “the science” just whatever got funded, published and not yet disproven? How is that different from authority?
  15. (Cynic) If most single studies are shaky, why believe any of them? And how is that different from believing none?
  16. (Child) Why do grown-ups say “it’s because of X” when they didn’t check?
  17. When does going on investigating become a way of hiding from a decision?
  18. Can a measure change the thing it measures (Hawthorne, being watched, grades)? What does that do to learning?
  19. What would it take to change your mind about something you believe strongly? If nothing could, is it still knowledge?
  20. How is a cell or an animal doing measure-model-manipulate without words? And what does language add?

Examples

  1. Bread that won’t rise (kitchen scale). Keep it from the earlier draft as the chapter’s running example. Four changes at once (yeast, water temperature, kneading, open window) show confounding; two small batches with one change show manipulation; a second rerun shows why one result isn’t enough.
  2. Bacterial chemotaxis (cell scale). E. coli compares chemical concentration now with a moment ago (a measurement over time), “models” that the gradient is improving, and changes tumbling frequency (a manipulation of its own path). It’s learning without a mind, and a good bridge from Act I’s wanting cells.
  3. A toddler dropping food off a high chair (child scale). Developmental psychologists call it hypothesis testing. Gopnik’s “scientist in the crib” work shows infants run informal interventions. It’s the book’s “purely curious” perspective in action. (Check the specific studies before citing: Gopnik, Meltzoff & Kuhl, The Scientist in the Crib, 1999.)
  4. Clever Hans (animal plus observer). Berlin, c. 1904–1907. The horse “did arithmetic” by tapping. Oskar Pfungst (1907) varied whether Hans could see the questioner and whether the questioner knew the answer: accuracy was about 89% when Hans could see a questioner and about 6% when he couldn’t. The cue was an involuntary head movement by the humans. Lesson: the observer is part of the system, which is why blinding exists. https://en.wikipedia.org/wiki/Clever_Hans
  5. Feynman’s “Mr. Young” and the rats (lab scale). Young’s rats kept finding the “right” door by cues nobody had controlled for (smell, light, and finally the sound of the floor), until he laid the corridor in sand. Feynman called it “an A-Number-1 experiment” that others ignored. Lesson: a manipulation is only as clean as the routes you’ve cut off. https://calteches.library.caltech.edu/51/2/CargoCult.htm
  6. James Lind’s scurvy trial, HMS Salisbury, 20 May 1747. Twelve sick sailors in six pairs: cider, elixir of vitriol, vinegar, seawater, oranges and lemons, or a spicy paste. The citrus pair recovered fastest. It was a real controlled comparison, yet the Navy didn’t adopt lemon juice until 1795, and Lind himself didn’t fully believe his result. With two sailors per arm, one result settled nothing socially. https://en.wikipedia.org/wiki/James_Lind
  7. Semmelweis, Vienna General Hospital, 1847. First Clinic (doctors and students) maternal mortality averaged about 10%, the midwives’ Second Clinic under 4%. He tested and ruled out several hypotheses (crowding, the priest’s bell, birthing position) before Kolletschka’s scalpel death pointed to “cadaveric particles.” With chlorinated-lime handwashing from mid-May 1847, mortality went from 18.3% in April to 2.2% in June. He was still rejected, partly because he had no mechanism and didn’t publish well. Lesson: a working manipulation without a model people accept doesn’t spread. https://en.wikipedia.org/wiki/Ignaz_Semmelweis (Hempel’s Philosophy of Natural Science, 1966, ch. 2 uses this as the textbook case of hypothesis testing.)
  8. John Snow and cholera, London 1854. Broad Street outbreak: began 31 August, with 127 deaths in the first three days (confirm the final toll; sources differ, roughly 500–616). The pump handle came off on 8 September. Snow himself admitted the outbreak was already waning, so the handle removal proved nothing on its own. The stronger evidence was his “Grand Experiment”: households on the same streets were supplied by Southwark & Vauxhall (sewage-laden intake) or Lambeth (upriver), and the first had many times the cholera mortality. A natural experiment stood in for one he couldn’t run. https://en.wikipedia.org/wiki/1854_Broad_Street_cholera_outbreak
  9. The barometer and the storm (thought experiment). The needle predicts storms perfectly well. Force the needle down by hand and the storm still comes. That shows the difference between predicting something and controlling it. Pearl’s ladder: association, then intervention, then counterfactual. https://plato.stanford.edu/entries/causation-mani/
  10. Smoking and lung cancer, 1950–1965. Doll & Hill (1950) and later cohort studies. Nobody could randomize smoking, and Fisher himself argued for a confounding genotype. Hill’s 1965 “viewpoints” (strength, consistency, specificity, temporality, dose-response and others) show how several imperfect lines of evidence add up. Nice irony: the father of randomized experiments was on the wrong side of an observational question. https://en.wikipedia.org/wiki/Bradford_Hill_criteria
  11. The Reproducibility Project: Psychology (2015). 100 studies from 2008 journals: 97 originals significant, about 36% of replications significant, and effect sizes about half as large. Cancer biology (2021): about 26% could be replicated, and effects were about 85% smaller. Scale: a whole field learning that its single results weren’t settled. https://en.wikipedia.org/wiki/Reproducibility_Project:_Psychology
  12. Hawthorne Works, 1924–1932. Lighting went up and output rose; lighting went down and output rose again. Being measured changed the thing measured. Levitt and List’s 2011 reanalysis of the original data found only weak effects, so even the famous result about results needed replicating. https://en.wikipedia.org/wiki/Hawthorne_effect
  13. Stevens’s scales, and “twice as happy.” A 1–10 pain or happiness score is ordinal, so averaging it, or saying 8 is “twice” 4, assumes more than the scale gives you. Everyday measurement mistake. https://en.wikipedia.org/wiki/Level_of_measurement
  14. Roediger & Karpicke (2006), retrieval. Experiment 2: after 5 minutes, the students who only reread (SSSS) recalled most (83% vs 71% for STTT). After a week the order flipped: STTT 61%, SSST 56%, SSSS 40%. The rereaders were the most confident they’d remember. It’s a measure, model, manipulate experiment about memory, and the chapter’s retrieval payoff. https://learninglab.psych.purdue.edu/downloads/2006/2006_Roediger_Karpicke_PsychSci.pdf
  15. Ebbinghaus, 1880–1885 (one person as the whole lab). He memorized nonsense syllables (WID, ZOF), tested himself at intervals, and measured “savings” on relearning. It gave us the forgetting curve. Civilizational version: libraries, scriptures learned by heart, and the oral recitation traditions of the Vedas, which keep knowledge retrievable across generations (links to Ch. 19). https://en.wikipedia.org/wiki/Forgetting_curve

Source material

  1. “The irritation of doubt causes a struggle to attain a state of belief. I shall term this struggle inquiry, though it must be admitted that this is sometimes not a very apt designation.” Charles S. Peirce, “The Fixation of Belief,” Popular Science Monthly 12 (November 1877), §IV. [verified: https://en.wikisource.org/wiki/Popular_Science_Monthly/Volume_12/November_1877/Illustrations_of_the_Logic_of_Science_I] Use: Opening. Learning starts from felt doubt, not curiosity in the abstract, and the risk is settling doubt too fast.

  2. “the assignment of numerals to objects and events according to rules” S. S. Stevens, “On the Theory of Scales of Measurement,” Science 103(2684), 7 June 1946, p. 677. [verified: https://en.wikipedia.org/wiki/Level_of_measurement] Use: Measure section. The rules are a choice, which ties measurement back to the Act II blade. (Check against the Science original for exact punctuation.)

  3. “causal relationships are relationships that are potentially exploitable for purposes of manipulation and control” James Woodward, “Causation and Manipulability,” Stanford Encyclopedia of Philosophy (first pub. 2001; rev.), §1. The same idea runs through Making Things Happen (2003). [verified: https://plato.stanford.edu/entries/causation-mani/] Use: Defines the Manipulate step. The entry also opens with causes “as handles or devices for manipulating effects,” which fits the book’s “handles” language from the Preface. Confirm the entry’s author credit and revision date on the page.

  4. “The first principle is that you must not fool yourself—and you are the easiest person to fool.” Richard P. Feynman, “Cargo Cult Science,” Caltech commencement address, 1974; Engineering and Science 37(7), June 1974. [verified: https://calteches.library.caltech.edu/51/2/CargoCult.htm] Use: Why controls and replication exist. Pair it with the Mr. Young rats story from the same speech (“that is an A‑Number‑1 experiment”).

  5. “the attacks had so far diminished before the use of the water was stopped, that it is impossible to decide whether the well still contained the cholera poison in an active state.” John Snow, On the Mode of Communication of Cholera, 2nd ed. (London: Churchill, 1855), Broad Street section (page to be confirmed). [verified: https://en.wikipedia.org/wiki/1854_Broad_Street_cholera_outbreak] (Wording is from Wikipedia’s quotation; check against the 1855 text at the UCLA John Snow site before printing.) Use: The iconic “one intervention” did not settle the cause, and Snow said so himself. That makes it the chapter’s best case for “one result rarely settles a cause.”

  6. “none of my nine viewpoints can bring indisputable evidence for or against the cause-and-effect hypothesis and none can be required as a sine qua non.” Austin Bradford Hill, “The Environment and Disease: Association or Causation?” Proceedings of the Royal Society of Medicine 58(5), 1965, pp. 295–300 (the passage is near p. 299). [verified: https://en.wikipedia.org/wiki/Bradford_Hill_criteria] (Secondary source; the PMC scan https://pmc.ncbi.nlm.nih.gov/articles/PMC1898525/ is image-only, so confirm the page there.) Use: How a case for a cause gets built when you can’t intervene: separate lines of evidence, none sufficient on its own.

  7. “However, on the delayed tests, prior testing produced substantially greater retention than studying, even though repeated studying increased students’ confidence in their ability to remember the material.” Henry L. Roediger III & Jeffrey D. Karpicke, “Test-Enhanced Learning,” Psychological Science 17(3), 2006, p. 249 (abstract). [verified: https://learninglab.psych.purdue.edu/downloads/2006/2006_Roediger_Karpicke_PsychSci.pdf] Use: Retrieval section. It complements the quote-options line already chosen: feeling that you know something isn’t evidence that you do.

  8. “a sudden slight upward jerk of the head” Oskar Pfungst, Clever Hans (The Horse of Mr. von Osten), trans. C. L. Rahn (1911), as quoted in Wikipedia. [verified: https://en.wikipedia.org/wiki/Clever_Hans] (Secondary; confirm page in the Project Gutenberg text of Pfungst.) Use: The tiny, unintended cue that produced a whole “result.” Put it at the center of the blinding and controls passage.

  9. “Experimentation has a life of its own.” Ian Hacking, Representing and Intervening (Cambridge UP, 1983), Part B, “Intervening” (commonly cited p. 150). [UNVERIFIED] Use: Manipulation isn’t just theory-testing; experimental craft builds knowledge in its own right. Check wording and page against the book.

Counterarguments and limits

Connections

Exercise ideas

  1. One-change trial. Instruction: Pick something small at home that sometimes works and sometimes doesn’t (your sleep, coffee strength, a plant, a bread rise, how long your phone battery lasts). Write down your current explanation in one sentence. Then write down two other explanations that would fit what you’ve seen. Change exactly one thing and hold the others as steady as you practically can. Do it at least three times, and record the result each time on a dated line. Notice: How many rival explanations came easily. Whether your first result agreed with the second and third. Which variables you couldn’t hold still. Why: It runs the whole loop at the scale of a kitchen and shows directly why one result rarely settles a cause.

  2. Close the book and say it back. Instruction: Right after finishing this chapter, close it. On a blank sheet, write the three verbs and, for each, one sentence on what it does and one example from the chapter. Don’t look back. Then check against the text and mark what you missed or got wrong. Repeat the next day, and again a week later, without rereading in between. Notice: The difference between how familiar the chapter felt and what you could actually produce, and whether the second and third attempts come more easily. Why: It lets the reader test the retrieval claim on this chapter. It also measures the reader’s own knowledge, which is the loop turned inward.

  3. Interrogate a headline. Instruction: Find one news story that says “X linked to Y” or “X causes Y.” Write four short answers: What was measured, and how? What model links X to Y? Did anyone change X, or only observe it? Has it been found more than once? Then name one third factor that could produce both X and Y. Notice: How often the answer to “did anyone change X?” is no, and how the headline’s wording (linked, boosts, causes) shifts with that. Why: It turns the chapter into the book’s practical defense against being cheated. A reader who can ask these four questions is harder to fool.

Open questions for the author

  1. Is bread still the running example, or will you reuse Ch. 13’s dinner so the three MMM chapters share one scene?
  2. How far into philosophy of science do you want to go: Popper, Duhem–Quine, Kuhn and Lakatos named, or kept in the background for Ch. 11?
  3. Should retrieval be a full section, or a coda? It sits oddly beside causal inference unless it’s framed as “measuring your own map.”
  4. Do cells and animals “learn” in this chapter’s sense? Decide whether MMM applies below language, and say so. It matters for the Act III close (“unnatural in our loops”).
  5. Which historical case is the anchor: Snow (natural experiment, honest uncertainty), Semmelweis (right answer, rejected), or Lind (controlled trial, ignored for 48 years)? Using all three risks a catalog.
  6. How should the chapter handle the cynic’s move from “single studies fail” to “trust nothing”? Is that argued here or in Ch. 4?
  7. Should the Fisher–smoking irony go in? It’s a great story, but it’s about a real person being wrong and needs careful sourcing.
  8. Hacking’s “if you can spray them, they are real” makes manipulation a test of reality, not just of causation. Do you want that stronger metaphysical claim, or keep Manipulate epistemic?
  9. The earlier draft’s “When the evidence is adequate for the stakes…” handoff: keep it as the ending, or move stopping rules entirely into Ch. 15?

Reader perspectives

Curious young child

First reactions:

Questions they’d ask:

  1. “What is yeast? Is it alive? Is it eating the bread?” (It is alive. That’s a weird fact that would hook them.)
  2. “Why can’t you just change everything at once and see if it works?”
  3. “How do you know which thing made it go wrong if you changed four things?”
  4. “If it works once, doesn’t that mean it works?”
  5. “What’s a model? Like a toy car model?”
  6. “If you can’t move a star, how do you learn about stars?”
  7. “Why do I forget my spelling words even though I read them ten times?”
  8. “Why does a quiz help you remember? Quizzes are for checking, not learning!”
  9. “Can you measure being happy?”
  10. “What if you measure the wrong thing? Like measuring how tall my plant is when it’s actually sad because it’s thirsty?”
  11. “Can a result lie?”

Where they’d get lost, bored, offended or unconvinced:

Examples they’d bring:

What would win them over:

Cynical adult

First reactions:

Questions they’d ask:

  1. “What’s here that isn’t in a high school science textbook?”
  2. “How many results do I need before I believe something? Two? Ten? Give me a number or a rule, not ‘rarely.’”
  3. “The news says ‘new study shows coffee causes X.’ What exactly should I do with that sentence using this chapter?”
  4. “If astronomers and historians can’t intervene, and they still count as knowledge, why is ‘Manipulate’ one of the three steps at all?”
  5. “Why is retrieval practice in here? Isn’t that a study tip, not epistemology?”
  6. “Isn’t ‘one result rarely settles a cause’ the exact line tobacco companies used for decades to cast doubt? How do I tell honest caution from paid-for doubt?”
  7. “What’s the difference between a model and a story I tell myself?”
  8. “Where’s the line between learning and over-analysis? The chapter says ‘eventually you have to bake.’ When?”
  9. “My doctor changes one thing at a time and it takes months. Is that good practice or just slow?”
  10. “What if the thing you’re measuring is the wrong thing? The chapter mentions it for a thermometer but doesn’t take it anywhere.”
  11. “Does Hacking’s ‘if you can spray them’ really apply to my bread, or is it just a clever quote?”

Where they’d get lost, bored, offended or unconvinced:

Examples they’d bring:

What would win them over:

Believer / spiritual reader

First reactions:

Questions they’d ask:

  1. Scripture tells me both “Prove all things; hold fast that which is good” (1 Thess. 5:21) and “You shall not put the LORD your God to the test” (Deut. 6:16). Your method has no second clause. Are there things one shouldn’t manipulate to learn about, such as people, relationships or the holy?
  2. If I can’t run a controlled trial on my marriage, is what I know about it second-class knowledge?
  3. You say “several explanations can predict the same loaf.” Believers and naturalists often predict the same observations. How does your method choose between them without adding a philosophical assumption?
  4. Hacking says “if you can spray them, they are real.” What about things that act on you but that you can’t act on? Doesn’t that make reality depend on our power over it?
  5. Is memorizing scripture learning, or only storing? Recitation in my tradition is meant to change the reciter, not just keep the text. Does your retrieval research measure that?
  6. When prayer studies find no effect, what does that show: that prayer doesn’t work, or that prayer isn’t a mechanism?
  7. Who does the measuring, and do they count the cook’s patience (which the draft admits the thermometer ignores) as real?
  8. Is a sign, or an answered prayer, “one result that rarely settles a cause”? Isn’t that what wise spiritual directors already say about discernment?
  9. When is further analysis “shelter from action”? Traditions call endless deliberation a failure of trust. Is that the same thing?
  10. What does learning do to the learner, rather than to the model?

Where they’d get lost, bored, offended or unconvinced:

Examples they’d bring:

What would win them over:

Skeptical scientist

First reactions:

Questions they’d ask:

  1. How much does your bread vary when you change nothing? Without that number, how would you know the yeast mattered?
  2. Why change the yeast and not the temperature first? Which variable does the reader test when several are suspects, and how do they choose?
  3. Why is one result rarely enough? The draft says so but doesn’t give the three reasons: noise, confounding, and flexible analysis (trying things until something works).
  4. When you can’t intervene, as with smoking and lung cancer, how did we become confident anyway? Can the chapter show observational causal reasoning actually working?
  5. “Measure” assumes the instrument is trustworthy. How was the thermometer itself validated? Where does measurement bottom out?
  6. What is a “model” here: a verbal story, an equation, a mechanism, a statistical fit? Scientists use the word for all four.
  7. Does the chapter distinguish prediction from explanation? A model can predict well with the wrong mechanism (epicycles).
  8. Does the testing effect hold for understanding and application, or mainly for recall of studied material?
  9. Is “Measure, Model, Manipulate” really a sequence? In practice the model decides what you measure. Should the loop be drawn starting from Model?
  10. How does a reader avoid fooling themselves: blinding, writing the prediction down beforehand, deciding in advance what would count as failure?
  11. Do non-human organisms do this loop? Bacteria “measure” gradients. Where does the human version differ in kind rather than degree?

Where they’d get lost, bored, offended or unconvinced:

Examples they’d bring:

What would win them over:

Earlier draft

From “Learning: Measure, Model, Manipulate”

Imagine that last week’s bread rose beautifully and today’s loaf sits in the bowl like wet cement. You remember changing the yeast. You also used colder water, hurried the kneading, and left the dough beside an open window.

Learning begins by resisting the satisfying sentence the new yeast is bad.

Measure means becoming more careful about differences. Which yeast? How warm was the water? How long did the dough rise? Measurement need not involve a digital instrument; a dated note and an honest comparison may be enough. But every measure selects. A thermometer records temperature and ignores the cook’s impatience. What we choose to notice shapes what can later be explained.

Model means proposing relationships. Perhaps the yeast was inactive. Perhaps the dough was cold. A model earns trust when it helps us anticipate what we have not yet seen, but a successful prediction is rarely a coronation. Several explanations can predict the same loaf.

Manipulate means changing something to learn what difference it makes. Make two small batches, keeping the water, flour, kneading, and location as similar as practical while changing the yeast. If one rises and the other does not, the yeast explanation gains support. It does not become infallible: the packets may have been stored differently, and a kitchen is not a sealed laboratory.

Intervention is one route to causal knowledge, not the definition of learning. Astronomers cannot move a star to see what happens. Epidemiologists and historians often reason from observations they did not arrange. Even a controlled experiment depends on assumptions about what was held steady and how the result was measured. Learning can yield description, interpretation, prediction, or better uncertainty. It need not yield control.

There is another quiet problem: knowledge that cannot be recalled when needed does little work. Research by Henry Roediger and Jeffrey Karpicke found that, under their experimental conditions, retrieving material improved later retention compared with repeatedly studying it. This does not mean every quiz transforms understanding. It suggests a modest practice.

Close this chapter for a moment and retrieve its three ordinary verbs. Then ask what each contributes. Look back. The gap between what seemed familiar and what you could produce is useful evidence.

Learning loops because results alter the next observation. The loaf may teach you about yeast, or reveal that temperature matters more. Sometimes the deeper revision concerns the aim. Are you trying to produce an identical loaf every week, or become a cook who can respond to a changing kitchen? Correcting a procedure and reconsidering its goal are different achievements.

When the evidence is adequate for the stakes, continued analysis can become shelter from action. You eventually have to bake. That is where Creating begins.