Working notes, not prose. This page is a research packet for drafting the chapter: the brief, source material, questions, examples, reader perspectives, and the earlier draft.
Brief
- Must: Show how we cut ourselves with the blade: failing to recognize the other side of a distinction makes one side feel more real than reality warrants. Measures replacing purposes and labels becoming destiny are the working examples.
- Serves: It turns the act’s descriptions into responsibility, which the epilogue’s “use the blade carefully” depends on.
Quote options
-
“There is an error; but it is merely the accidental error of mistaking the abstract for the concrete. It is an example of what I will call the ‘Fallacy of Misplaced Concreteness.’” — Alfred North Whitehead, Science and the Modern World (1925), ch. III, p. 72 (Macmillan 1925 ed.). [verified: https://archive.org/details/sciencemodernwor00whit] Why: Names the core harm: one side of a cut feeling more real than reality warrants.
-
“When a measure becomes a target, it ceases to be a good measure.” — Marilyn Strathern, “‘Improving ratings’: audit in the British University system,” European Review 5(3) (1997), p. 308. [verified: https://gwern.net/doc/statistics/decision/1997-strathern.pdf] Why: The canonical form of “measures replacing purposes”; it is often misattributed to Goodhart, so cite Strathern.
-
“The more any quantitative social indicator is used for social decision-making, the more subject it will be to corruption pressures and the more apt it will be to distort and corrupt the social processes it is intended to monitor.” — Donald T. Campbell, “Assessing the Impact of Planned Social Change,” Evaluation and Program Planning 2(1) (1979). [verified: https://en.wikipedia.org/wiki/Campbell’s_law] Why: The longer, mechanistic version: the proxy loops back and corrupts what it was meant to track.
-
“People spontaneously come to fit their categories.” — Ian Hacking, “Making Up People,” in Reconstructing Individualism (1986), opening section (p. 161, approx.). [verified: https://serendipstudio.org/oneworld/system/files/Hacking_making-up-people.pdf] Why: “Labels becoming destiny” in six words.
-
“You cannot think without abstractions; accordingly, it is of the utmost importance to be vigilant in critically revising your modes of abstraction.” — Alfred North Whitehead, Science and the Modern World (1925), ch. IV, pp. 82–83 (Macmillan 1925 ed.). [verified: https://archive.org/details/sciencemodernwor00whit] Why: Turns the cost into a responsibility (use the blade, but carefully), which sets up the epilogue. Alt: “But in order to do so, evidently it must first cut itself up into at least one state that sees, and at least one other state that is seen. In this severed and mutilated condition, whatever it sees is only partially itself.” — G. Spencer-Brown, Laws of Form (1969), pp. 104–105 per Wikiquote. [verified: https://en.wikiquote.org/wiki/G._Spencer-Brown and https://fmdolan.com/remembering-george-spencer-brown/; not checked against the book]
Research
Core argument
- Start from the gain. Every cut buys something: a measure lets a group see change, a label lets a stranger act quickly, a diagnosis gathers scattered suffering into something treatable. The chapter should grant this fully first, or it reads as anti-measurement.
- Name the mechanism of the cost. A distinction has two sides: what it marks and what it leaves out (the unmarked side, the residue, the “not-this”). The cost begins when we stop seeing the second side. The marked side then feels more real, more solid and more complete than reality warrants. This is Whitehead’s misplaced concreteness in the book’s own vocabulary.
- Show why forgetting is the default rather than a failure of character. Cuts are cheap to reuse and expensive to reopen. Once a cut is shared, written into a form or tied to money, remembering the unmarked side takes deliberate effort. Chapter 8’s counterexamples show the other side is always there. This chapter shows how it drops out of view.
- Working example 1: measures replacing purposes. A proxy stands in for a purpose (reading minutes for reading, body counts for winning, accounts per customer for service). Put pressure on the proxy and people optimize the proxy. The purpose sits on the unmarked side and quietly degrades. Campbell, Goodhart and Strathern are three angles on the same move. The Yankelovich steps give the progression: first ignore what can’t be measured, then decide it doesn’t exist.
- Working example 2: labels becoming destiny. A label picks out one real feature and makes it the whole explanation. Because people respond to how they are classified, the label loops back and reshapes the person (Hacking’s looping, Beauvoir’s “becomes”). The unmarked side of a person, everything the label doesn’t say, stops being asked about, and sometimes stops being allowed.
- Generalize the pattern. The same structure appears in states (Scott’s legibility and the scientific forest), in markets (Lukács’s reification), in ideologies (systems that can no longer be surprised) and inside one mind (self-labels such as “I’m not a math person”).
- Turn description into responsibility. Once the reader sees that the cut is ours, forgetting its other side is also ours. We cannot stop cutting (there is no uncut perception, and triage is necessary). What we can do is keep the purpose visible, date our measures, and listen to the people a category fits badly.
- A short diagnostic. Three questions to ask of any cut: What does it let us do? What does it leave out? Who would notice first if it were failing? Add one sign of trouble: when a counterexample can no longer count, the cut has hardened.
- Hand-off. Chapter 10 (Thingification) shows what happens when cuts harden into shared, followable objects, so this chapter should end on “the cut feels like a thing” without yet explaining the whole social machinery.
Key questions
- If every distinction leaves something out, how do I know when leaving it out has become a problem?
- Why does the marked side of a distinction feel more real than the unmarked side? Is this a property of attention, of language, or of institutions?
- Is a measure ever just a measure, or does being watched always change what is measured?
- When a teacher “teaches to the test,” who is at fault: the teacher, the test, or the people who tied money to the test?
- What was the number for in the first place? Could anyone in the organization still say?
- Can you name a label that described you once and that you then started to live up to (or down to)?
- When is a label a gift (a diagnosis that finally explains things, a name for an identity) and when is it a cage? Can it be both at once?
- If a child is called “the smart one” and a sibling “the kind one,” what happens to each of them?
- If a category makes people legible to the state, who benefits from being legible, and who benefits from staying illegible?
- Is reification a mistake we could avoid, or the price of being able to think at all (Whitehead: “You cannot think without abstractions”)?
- How can you tell an ideology from a strong, well-founded commitment? Is “can it still be surprised?” a fair test?
- What does it cost to reopen a cut: money, time, status, identity? Who pays?
- Are some measures safe from Goodhart’s law? For example, measures that nobody’s pay or status depends on, or measures that are the purpose (the temperature of a baby’s bath).
- A young child’s question: “If I stop being called shy, will I stop being shy?”
- A cynic’s question: “Isn’t all this just an excuse for people who don’t want to be held accountable?”
- If we can’t avoid cutting, is “use the blade carefully” anything more than a platitude? What would careful look like on a Tuesday?
- Who is best placed to notice that a category is failing: the people who made it, the people using it, or the people it is applied to?
- Is a person’s self-image a measure that has become a target?
- When the unmarked side is ignored for long enough, does it disappear in fact (the understory of the scientific forest), or does it return with interest (the dying forest)?
Examples
- Reading log (the earlier draft’s scene). A school counts minutes read each night. Children log minutes beside unopened books. This is the domestic opening, small and familiar.
- Hanoi rat bounty, 1902. The French colonial government paid a bounty per rat tail, reportedly 1 cent. Catchers cut off the tails and released the rats to breed. The proxy (a tail) separated from the purpose (fewer rats). https://en.wikipedia.org/wiki/Cobra_effect
- The “cobra effect” as a cautionary tale about stories. The famous Delhi cobra-bounty story may be apocryphal: the Wikipedia article reports a 2025 investigation that found no contemporary documentation. It makes a useful meta-example: a tidy story is a cut too, and it can feel more real than the evidence. https://en.wikipedia.org/wiki/Cobra_effect
- Wells Fargo, 2011–2016. The “Going for Gr-Eight” cross-selling push aimed for eight products per customer. Later estimates put fraudulent accounts at about 3.5 million. About 5,300 employees were fired, and about $3 billion was paid in the 2020 DOJ/SEC settlement. The measure (products per customer) was meant to indicate relationship depth. https://en.wikipedia.org/wiki/Wells_Fargo_cross-selling_scandal
- Atlanta Public Schools, 2009. Cheating on the CRCT involved 44 of 56 schools and 178 educators. Eleven were convicted of racketeering in April 2015. Teachers cited “inordinate pressure” to meet targets. This is Campbell’s law in a courtroom. https://en.wikipedia.org/wiki/Atlanta_Public_Schools_cheating_scandal
- Mid Staffordshire (Stafford Hospital), 2005–2008. A board chasing financial and national targets cut staff. Press estimates of excess deaths ran from 400 to 1,200, though the Healthcare Commission warned against tying the failures to a specific number. In Jeremy Hunt’s summary of the inquiry’s findings, “concentrating on national targets led to managers deprioritising the safety and well-being of patients.” https://en.wikipedia.org/wiki/Stafford_Hospital_scandal
- Vietnam body counts (McNamara fallacy). Enemy dead became the metric of progress in a war whose outcome turned on things nobody counted. Pair it with the four Yankelovich steps. https://en.wikipedia.org/wiki/McNamara_fallacy
- Hospital length of stay and the h-index. Targets for shorter stays were followed by premature discharges and more readmissions. The h-index’s link to scientific awards weakened once it became widely used. https://en.wikipedia.org/wiki/Goodhart%27s_law
- Scientific forestry (18th–19th-century Prussia and Saxony, via Scott). Foresters simplified the forest into standard timber yield and replanted it as a monoculture. It was easy to count and manage, but it was less resilient, and the next generations of trees suffered. This is the civilizational-scale image of the unmarked side (understory, fungi, deadwood, soil life) returning as damage. Scott, Seeing Like a State, ch. 1. https://en.wikipedia.org/wiki/Seeing_Like_a_State
- Pygmalion in the classroom (Rosenthal and Jacobson, 1968), told with its caveats. Teachers were told that randomly chosen pupils were “bloomers,” and the study reported IQ gains. Thorndike attacked the methods, and later meta-analyses find expectancy effects that are real but “usually small and temporary.” Use it to show “labels become destiny” honestly: the effect exists, and it is weaker than the legend. https://en.wikipedia.org/wiki/Pygmalion_effect
- Stereotype threat (Steele and Aronson, 1995). Black students did worse on hard GRE verbal items when the test was framed as “diagnostic of intellectual ability.” The 2015 Flore and Wicherts meta-analysis suggests the effect is small and inflated by publication bias. It is another example where the book should model care with claims (Chapter 8). https://en.wikipedia.org/wiki/Stereotype_threat
- Hacking’s “making up people.” Multiple personality disorder went from a rare diagnosis to an epidemic in the 1970s and 80s, and the “kind” and the people fitting it co-evolved. This is the looping effect: classifications change the classified, which changes the classification. (Hacking, “Making Up People,” 1986; Rewriting the Soul, 1995.)
- Beauvoir’s “becomes.” A role imposed from outside is lived as nature, and the unmarked possibilities (“One is not born a genius… the feminine situation has… rendered this becoming practically impossible”) are closed off in practice. https://en.wikiquote.org/wiki/Simone_de_Beauvoir
- Scale down to cells and animals. Evolution also cuts by proxy. Sweetness stood in for calories, and in an environment of refined sugar the proxy outruns the purpose. Supernormal stimuli do the same: Tinbergen’s birds preferred oversized dummy eggs to their own, and male jewel beetles (Julodimorpha bakewelli) tried to mate with brown beer bottles in Australia (Gwynne and Rentz, 1983). These show that the cost of the blade is older than humans: any creature that acts on a sign can be captured by the sign. [Verify beetle citation before use; from memory.]
- Self-labels. “I’m not a math person” or “I’m bad with names.” One remembered failure becomes an identity, and the identity steers away from practice, which confirms the label. This is a thought experiment the reader can check against their own life.
Source material
-
“Any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes.” — Charles Goodhart, “Problems of Monetary Management: The UK Experience” (1975). [verified: https://en.wikipedia.org/wiki/Goodhart%27s_law] (secondary source; check against the paper in Papers in Monetary Economics*, Reserve Bank of Australia, 1975, if possible)* Use: the original, drier economist’s wording. It shows the law was first about statistical regularities (the cut) collapsing under use, not about bad people.
-
“Achievement tests may well be valuable indicators of general school achievement under conditions of normal teaching aimed at general competence. But when test scores become the goal of the teaching process, they both lose their value as indicators of educational status and distort the educational process in undesirable ways.” — Donald T. Campbell, “Assessing the Impact of Planned Social Change,” Evaluation and Program Planning 2(1) (1979), pp. 67–90 (written 1976). [verified: https://en.wikipedia.org/wiki/Campbell%27s_law] Use: supports the reading-log scene directly. The test works as an indicator until it becomes the goal. (It is a different passage from the Campbell line already in Quote options.)
-
“The first step is to measure whatever can be easily measured. This is okay as far as it goes. The second step is to disregard that which can’t be easily measured or give it an arbitrary quantitative value. This is artificial and misleading. The third step is to presume that what can’t be measured easily really isn’t very important. This is blindness.” — Daniel Yankelovich, “The New Odds,” speech to the Marketing Strategy Conference, Sales Executives Club of New York, 15 October 1971 (as cited in Wikipedia, “McNamara fallacy”; confirm against a printed source). The next sentences run: “The fourth step is to say that what can’t be easily measured really doesn’t exist. This is suicide.” [verified: https://en.wikipedia.org/wiki/McNamara_fallacy] Use: a step-by-step account of how the unmarked side disappears, which fits the chapter’s mechanism exactly. The quote runs slightly over 60 words with the fourth step, so quote three steps and paraphrase the fourth, or quote the fourth alone.
-
“Certain forms of knowledge and control require a narrowing of vision.” — James C. Scott, Seeing Like a State (Yale, 1998), ch. 1, “Nature and Space”. [verified: https://en.wikiquote.org/wiki/James_C._Scott] Use: the gain and the cost in one sentence. Narrowing enables control, and that is the bargain.
-
“The forest as a habitat disappears and is replaced by the forest as an economic resource.” — James C. Scott, Seeing Like a State (1998), ch. 1. [verified: https://en.wikiquote.org/wiki/James_C._Scott] Use: the scientific-forestry example compressed. The marked side (resource) replaces the whole.
-
“A map is not the territory it represents, but, if correct, it has a similar structure to the territory, which accounts for its usefulness.” — Alfred Korzybski, Science and Sanity (1933), p. 58. [verified: https://en.wikipedia.org/wiki/Map%E2%80%93territory_relation] Use: the full sentence, not just the slogan. It keeps the balance (maps are useful because they share structure) and sets up “the cost is forgetting the ‘not’.”
-
“Since all models are wrong the scientist cannot obtain a ‘correct’ one by excessive elaboration.” — George E. P. Box, “Science and Statistics,” Journal of the American Statistical Association 71 (1976), p. 792. [verified: https://en.wikiquote.org/wiki/George_E._P._Box] Use: against the tempting fix of “just add more measures.” You can’t get back the unmarked side by making the marked side more elaborate.
-
“Its basis is that a relation between people takes on the character of a thing and thus acquires a ‘phantom objectivity’, an autonomy that seems so strictly rational and all-embracing as to conceal every trace of its fundamental nature: the relation between people.” — Georg Lukács, “Reification and the Consciousness of the Proletariat,” in History and Class Consciousness (1923), section I.1, trans. Rodney Livingstone (1967). [verified: https://www.marxists.org/archive/lukacs/works/history/hcc05.htm] Use: the origin of “reification” and a bridge to Chapter 10. It is a relation that feels like a thing and hides that it was made.
-
“One is not born a genius, one becomes a genius; and the feminine situation has up to the present rendered this becoming practically impossible.” — Simone de Beauvoir, The Second Sex (1949), Book I, Part 2, ch. 8 (p. 133 in the edition Wikiquote cites). [verified: https://en.wikiquote.org/wiki/Simone_de_Beauvoir] Use: a less-quoted companion to “One is not born, but rather becomes, a woman.” It shows a label working as destiny by closing off options, not by describing a nature.
-
“Where there is dirt there is system.” — Mary Douglas, Purity and Danger (1966), ch. 2. [UNVERIFIED] Use: the residue of every cut. What doesn’t fit gets labeled dirt, danger or deviance, which is the unmarked side taking its revenge.
Counterarguments and limits
- “Measurement is how we escaped anecdote.” Numbers exposed hospital infection rates, lead in water and pay gaps. Without measures, the powerful’s intuitions rule. The chapter must say plainly that the alternative to a bad measure is usually a better measure plus judgment, not no measure (see Muller, The Tyranny of Metrics, 2018, who argues for judgment alongside metrics rather than against metrics).
- Goodhart is overused. Many measures under pressure stay informative, for example mortality after surgery or GDP roughly tracking material output. The law describes a tendency whose strength depends on how easy the proxy is to game compared with the purpose. Overstating it makes the chapter unfalsifiable.
- Label effects are smaller than the legends. Pygmalion and stereotype-threat effects have shrunk under replication. If “labels become destiny” leans on these studies, a skeptical reader will catch it. Lean instead on structural cases (who is allowed into which school or job), where the label’s effect runs through institutions rather than subtle psychology.
- Labels also liberate. Diagnoses (ADHD, autism, long COVID) and identity terms often come as relief and open access to help and community. A chapter that treats labels only as cages misreads many readers’ experience. Hacking’s looping works in both directions.
- “Accountability evasion” (the cynic). “Every measure is incomplete” can shield incompetence. The chapter needs a criterion that doesn’t let anyone off the hook, for example: the person criticizing the measure owes a better account of the purpose.
- The “unmarked side” metaphor can overreach. Not every distinction has a single, symmetric other side (Chapter 8’s hunt for non-dualities). The chapter should talk about “what the cut leaves out” rather than insist on a neat obverse.
- Is it really “we cut ourselves”? Many of these harms are done to people by institutions with power (colonial bounties, states, employers). Calling it a shared human error risks flattening responsibility. Distinguish the universal cognitive tendency from who profits from the forgetting.
- Historicity of tidy stories. The cobra story, the Soviet nail factory and similar parables are often apocryphal. Use documented cases, or label parables as parables.
Connections
- Ch. 4 (From Subjectivity to Something That Will Have to Do): knowing the limits of your own view is the first defense. This chapter shows what happens socially when those limits are forgotten.
- Ch. 5 (The First Blade: Me and World): treating yourself as an isolated system is useful and false in the same way a measure is. You can “cut yourself” by forgetting the environment.
- Ch. 6 (The Second Blade: Reality and Representation): the proxy is a representation that loops back into reality. Goodhart is that loop gone wrong.
- Ch. 7 (The Third Blade: Good, Bad): metrics and labels survive for reasons unrelated to accuracy (they are cheap, legible and defensible), which is Chapter 7’s question of why a belief survives.
- Ch. 8 (What Is Not a Cut?): supplies the counterexamples and the other sides this chapter says we forget. Also note the risk that the chapter overclaims symmetrical dualities.
- Ch. 10 (Thingification): the direct sequel. Forgotten cuts harden into things (roles, money, diagnoses, institutions). Lukács and Hacking bridge the two.
- Ch. 11 (Knowledge): the repair is knowledge as correctable practice, the earlier draft’s closing hand-off.
- Ch. 14 (Learning: Measure, Model, Manipulate): measurement done well, where the model stays answerable to the territory. It is the positive twin of this chapter.
- Ch. 18 (Loops That Learn and Loops That Don’t): Campbell’s corruption and Hacking’s looping are runaway feedback loops.
- Ch. 20 (Raising the Floor): program theory makes a collective effort’s assumptions explicit so a proxy can’t quietly replace the purpose.
- Ch. 22 and Epilogue: “use the blade carefully.” The diagnostic questions here become Chapter 22’s review-date practice.
Exercise ideas
-
Find one number that has replaced its purpose. Instruction: Pick one number you or your workplace tracks (steps, inbox zero, sales calls, grades, followers, weight, screen time). Write the number at the top of a page. Below it, write in one sentence what the number was for. Then list three ways you could raise the number without serving the purpose, and check whether you or anyone around you already does any of them. Notice: how easy the list is to write, and whether the purpose sentence was hard to recall. Why: it makes Goodhart and Campbell personal and shows the unmarked side (the purpose) fading out of view.
-
Label autopsy. Instruction: Write down one label that has been applied to you for years (“the responsible one,” “not creative,” “difficult,” a diagnosis, a job title). Draw two columns. In the first, list what the label got right. In the second, list three times you acted against it and what happened to those memories: were they forgotten, explained away, or treated as “not really you”? Last, write down one person who would describe you without that label. Notice: whether the counterexamples were harder to recall than the confirming ones, and whether you have been choosing situations that fit the label. Why: it shows “labels become destiny” through looping, from the inside, without claiming the label is false.
-
Three questions to one rule. Instruction: This week, when a rule, score or category decides something about you or someone else (a form, a rubric, a policy), ask three questions and write down the answers. What does it let them do? What can it no longer see? Who would notice first if it were failing, and can that person reach anyone who could change it? Notice: whether the third question has an answer. Often nobody’s job is to notice. Why: it turns the chapter’s diagnosis into the “responsibility” the epilogue depends on, as a repeatable habit.
Open questions for the author
- Anchor scene. Keep the reading-log scene, or open with something more vivid and documented (Hanoi rats, Wells Fargo)? The reading log is relatable, and the documented cases carry stakes.
- Vocabulary. Will you use “the unmarked side” (Spencer-Brown), “the other side of the cut,” or “the residue” (Douglas)? Pick one and keep it consistent with Chapter 8.
- Reification and ideology. How much to name “reification” and “ideology” here versus leaving the vocabulary to Chapter 10. The outline keeps “things” for Chapter 10, and the earlier draft’s reification-to-ideology paragraph may belong there.
- Contested studies. Include Pygmalion and stereotype threat with explicit caveats, or leave them out? Including them with caveats would model the Chapter 8 humility contract.
- Who is “we”? How strongly to separate the universal cognitive tendency from the power that benefits from forgetting (colonial bounties, employers, states), and whether “we cut ourselves” stays literal or becomes “we cut ourselves and each other.”
- Non-human cases. Include the supernormal-stimulus and sweetness examples to keep the book’s cells-to-civilizations range, or keep this chapter human?
- Positive role for labels. How much space to give labels that liberate (diagnoses, identity terms) so the chapter isn’t read as anti-category.
- Link targets. The earlier draft links to /practice/labels and /practice/maps. Keep them, or replace them with the new exercises above?
Reader perspectives
Curious young child
First reactions:
- The reading-log opening is personal: “That’s MY reading log!” Many kids have faked minutes on one. They’d feel caught, and then feel understood. This is the strongest possible opener for them.
- “Cutting yourself with the blade” is a scary phrase. Taken literally, it sounds like self-harm. The metaphor needs softening or explaining for younger readers, maybe “the knife slips.”
- Labels like “lazy” matter to them. They’ve been called things: “the loud one,” “the baby,” “the smart one.” This chapter could feel very personal, in a good or a hurtful way.
- “Reification” and “ideology” are words they’d skip over completely.
Questions they’d ask:
- “If I read a really good book for 10 minutes and my friend reads a boring one for 30, who read more?”
- “Why does the teacher want the number and not just to know if I liked the book?”
- “If I get a gold star, does that mean I’m good? What if I got it by accident?”
- “If everyone calls me ‘the shy one,’ do I have to stay shy forever?”
- “Why can I go on the big roller coaster only if I’m tall enough? My friend is braver than me but shorter.”
- “If I get a bad grade on the spelling test, am I bad at spelling, or did I just have a bad day?”
- “Why do grown-ups count steps on their watch? Do the steps count if you shake your arm?”
- “If my brother is ‘the naughty one,’ does he get blamed even when it was me?”
- “Why do we have rules that don’t make sense anymore, like no running when there’s nobody there?”
- “Can you stop being a label? How?”
- “If you can’t see anything without cutting, how do you know what you cut off?”
Where they’d get lost, bored, offended or unconvinced:
- The brief’s core idea is that forgetting the other side of a cut makes one side feel more real than it is. That’s abstract, and the draft never states it in a kid-sized way. Try: “When you only look at the number, you forget the book.”
- The draft says “This is often called Goodhart’s law” right after Campbell. The quote packet notes the Strathern phrasing is misattributed to Goodhart. A kid wouldn’t care, but the “reference exactly” standard means the draft’s sentence needs fixing.
- The list of hidden reasons behind “lazy” (grief, chronic illness, impossible childcare, a bad manager) is a grown-up list. The kid version is tired, sad, bored, confused, or hungry.
- Beauvoir and “social destinies” won’t reach them.
- Possible offense: a child labeled with a diagnosis (ADHD, autism, dyslexia) reading “labels becoming destiny” might hear “your diagnosis is bad.” The chapter must say, as the Chapter 10 draft does, that labels can also help: they get you the right glasses and extra time.
- Unconvinced moment: “But you need numbers! How else would the teacher know?” The draft agrees (triage, water rules), but that agreement should come earlier, so the kid doesn’t think the book is anti-measuring.
Examples they’d bring:
- The faked reading log: the lived version of Strathern’s “When a measure becomes a target, it ceases to be a good measure.”
- Sticker charts and star-of-the-week: kids learn to do the sticker thing rather than the good thing, like tidying only when someone is watching.
- Height sticks at theme parks: a measure standing in for “safe enough to ride,” which is fair most of the time and wrong for some kids. A good case of a useful proxy with a known cost.
- Screen-time minutes: “I have 20 minutes left” turns into racing, not enjoying.
- Family nicknames: “the messy one” or “the sporty twin,” where Hacking’s “People spontaneously come to fit their categories” plays out at home. Kids know the feeling of acting the part because everyone expects it.
- Sorting a class into reading groups by animal names: kids know which group is the “slow” one no matter what it’s called.
What would win them over:
- Opening with the reading log and never talking down about it: “You weren’t cheating reading. The log was cheating reading.”
- Three simple questions from the exercise, made kid-sized: What’s this number for? What can’t it see? Who could tell us it’s wrong?
- A clear message that labels can be taken off or changed, and that you’re allowed to surprise people.
- Keeping it fair: numbers and labels help too, and the knife is useful. It just slips sometimes.
Cynical adult
First reactions:
- This is the chapter I came for. Goodhart and Campbell, measures eating purposes, labels becoming destiny: that’s my working life, and it’s real.
- It’s also well-trodden ground: Jerry Z. Muller’s The Tyranny of Metrics (2018) covers the measures half in detail. What does this chapter add?
- The reading-minutes example is gentle. Real metric corruption involves fraud, deaths and prison sentences. Soft examples make the warning feel optional.
- The draft says “this is often called Goodhart’s law” right after quoting Campbell, while the quote notes say to cite Strathern and not misattribute. Clean up the attribution before a pedant does.
- “Reification… ideology… surprise has become impossible” is sharp. It’s also exactly the charge a reader could level at a book with a tidy three-act, three-blade system.
Questions they’d ask:
- If every measure gets gamed, what do I use instead? “Keep the purpose visible” isn’t a management system.
- Is the claim that all targets corrupt, or only high-stakes ones? Campbell says “the more… used.” Where’s the threshold?
- “Labels becoming destiny”: for whom? Some labels (a diagnosis, a disability status) get people help. When is the label the problem and when is it the lack of one?
- How do I tell a useful label from a cage in the moment, not in hindsight?
- The draft lists possible causes behind “lazy.” Isn’t that just the charitable interpretation? Sometimes people are lazy. Does the chapter allow that?
- Hacking says people “spontaneously come to fit their categories.” Always? Is there evidence, or is it a philosopher’s observation?
- “Ideology is a system that already knows what every new event means.” Does that include the book’s framework? How would I know if it had turned into one?
- Who wrote the metric in the first place, and who profits from it? The chapter talks about forgetting, but a lot of this is deliberate.
- What does “the other side of a distinction” mean in the measure case? What’s the other side of “minutes read”?
- The Whitehead quotes are from 1925. Is there anything empirical here, or is it all philosophy plus anecdote?
- Isn’t “invite reports from people the category fits badly” just a feedback form? Those get ignored.
Where they’d get lost, bored, offended or unconvinced:
- The repair paragraph (“keep the purpose visible, revisit the measure, invite reports…”) is consultant boilerplate. Every failed reorg I’ve sat through promised this.
- “They do not abolish a blade. They keep it from becoming a prison.” Another poster line.
- The brief’s framing (“failing to recognize the other side of a distinction makes one side feel more real than reality warrants”) is abstract. The draft never connects that sentence to the examples, so I don’t see what “the other side” is in each case.
- Treating metric gaming as innocent forgetting will offend anyone who’s been on the receiving end. Often someone decided the number mattered more than you. Name incentives and power, not just cognition.
- The practice questions are good but have no teeth. What happens after I answer them?
Examples they’d bring:
- Wells Fargo (exposed 2016). Aggressive sales targets led staff to open millions of unauthorized customer accounts; the bank paid large settlements, including a $3 billion resolution with the US Department of Justice and SEC in 2020. Measures replacing purposes with a legal record attached.
- The US Veterans Affairs wait-time scandal (2014). At the Phoenix VA, staff kept off-the-books waiting lists so reported wait times met targets, while veterans waited far longer. The number looked fine and patients were harmed.
- The Atlanta Public Schools cheating scandal (investigation 2011, convictions 2015). Educators altered standardized tests under pressure to hit targets. A much stronger school example than reading minutes.
- Body count in the Vietnam War. Robert McNamara’s emphasis on enemy-dead counts as a measure of progress is the textbook case of a metric crowding out judgment; Muller discusses it.
- The “cobra effect.” A bounty on cobras supposedly led people to breed them. It’s widely told, but the historical sourcing is thin; if used, flag it as a parable, not a fact. That flag is itself a demonstration of the book’s method.
- Performance reviews and “stack ranking” (forced ranking, famously associated with GE under Jack Welch). A label (bottom 10%) that becomes destiny by policy.
What would win them over:
- High-stakes, documented examples instead of the reading log, with the incentive and who benefited named in each.
- A practical rule of thumb for measures: pair every metric with a check that would catch it being gamed, and say who is allowed to overrule the number.
- A case where the label helped, set beside one where it caged, to show the chapter isn’t anti-label, only anti-forgetting.
- Turning the “ideology” test on the book itself and passing it in view.
Believer / spiritual reader
First reactions:
- “One side feels more real than reality warrants” is, to a believer, very nearly the definition of idolatry: mistaking a made thing for the ultimate. Psalm 115 says it outright, and the reader will be surprised the chapter doesn’t.
- “Measures replacing purposes” will make them think at once of religion’s own worst habit: ritual performed for the count, like tallying rosary decades, or a Sunday attendance figure standing in for a parish’s life. They’ll respect a chapter that says this, and resent one that only points it at schools and dashboards while saving religion for the cheating column.
- “Labels becoming destiny” will make Hindu and Dalit readers think of caste. An honest chapter can’t skip it, and a believing Hindu reader will want it handled with knowledge of the reform traditions inside Hinduism, not only as an indictment from outside.
- They’ll like the ending’s “the repair is… keep the purpose visible.” That is what prophets did.
- The Whitehead quote, “Fallacy of Misplaced Concreteness,” will please process theologians. Whitehead’s metaphysics founded process theology, and a Christian reader who knows that will feel at home.
Questions they’d ask:
- Isn’t this chapter a secular version of the prophets’ critique of empty ritual? “I desire mercy, and not sacrifice” (Hosea 6:6). Will you credit it?
- Jesus: “The sabbath was made for man, and not man for the sabbath” (Mark 2:27). And the rabbis had their own version: “The Sabbath is given to you, and you are not given to the Sabbath” (Mekhilta de-Rabbi Ishmael, Ki Tisa). Isn’t that Strathern’s law, applied to the holiest measure a tradition has?
- If labels become destiny, what about positive renaming? Abram became Abraham, Jacob became Israel, Simon became Peter. Traditions use naming to call people into a destiny. Is that the same cost, or its redemption?
- Is scrupulosity, the religious anxiety of never having done a ritual exactly right, a case study for you? Ignatius of Loyola wrote rules for it in the Spiritual Exercises.
- Can a tradition that has lasted three thousand years keep a purpose visible better than a school district can? Or worse?
- Who gets to say what the “purpose” behind a measure is? In a religious community that’s God, scripture or tradition. In your book, who?
- Is sin, in the sense of hamartia, “missing the mark,” a kind of measure-replacing-purpose? You aim at a proxy good and miss the true one.
- You say “there is no clean escape into uncut perception.” Contemplatives say there is, at least briefly. Are you ruling that out?
- Does grace have a place here? My tradition’s answer to being trapped by labels is not better measurement but being seen whole by God. Is there a secular analogue?
- If ideology is “a system that already knows what every new event means,” isn’t that sometimes true of confident secularism as well? Will the chapter admit it?
Where they’d get lost, bored, offended or unconvinced:
- Offended if one-sided: If the working examples are all school metrics and workplace labels while religion appears only elsewhere as the cheat, the reader will see a double standard. Religion’s own self-critique is the strongest material here. Use it.
- Unconvinced: “The sign [of ideology] is that surprise has become impossible.” Many faithful people would say their faith is the source of constant surprise, through grace, conversion and doubt. The claim needs care, or it will read as “anyone with firm faith is ideological.”
- Bored: The reading-minutes example is apt, but it is small next to what’s at stake: caste, psychiatric labels, “heretic,” “unclean.”
- Lost: The chapter moves from Campbell to Hacking to Beauvoir in quick succession. The reader loses the idea of the other side of the distinction (the brief’s key phrase). Say plainly what the forgotten other side is in each example.
- Stung: “Labels becoming destiny” without any mention of how religious communities have lifted labels as well: lepers touched in the Gospels, the Buddha admitting outcastes such as Upali, the barber, to the Sangha.
Examples they’d bring:
- Psalm 115:4–8: “They have mouths, but they speak not: eyes have they, but they see not… They that make them are like unto them” (KJV). Three thousand years ago this said that a made thing treated as ultimate reshapes its makers into its own image. That is labels-as-destiny and the looping effect in one verse.
- Matthew 23:23: “ye pay tithe of mint and anise and cummin, and have omitted the weightier matters of the law, judgment, mercy, and faith.” A precise case of the measure (the tithe) crowding out the purpose (justice). Pair it with the rabbinic Sabbath saying so it doesn’t become an anti-Jewish trope.
- B. R. Ambedkar, Annihilation of Caste (1936): the most forceful critique of a religious label becoming birth-destiny, by someone who later led a mass conversion to Buddhism (Nagpur, 1956). It shows a label escaped through religion, not only away from it.
- Leviticus 13 and Mark 1:40–42: the priestly category “unclean,” which exiled the sick, and Jesus touching the leper. Within one scriptural canon you get both the cost of a category and a deliberate act of crossing it.
- Renaming in the Bible (Genesis 17:5, 32:28; Matthew 16:18): labels used deliberately to change destiny. That is Hacking’s looping effect aimed on purpose, and it raises the question of when that is good.
What would win them over:
- Credit the prophetic and contemplative traditions as the oldest critics of measures replacing purposes, including Hosea, Mark 2:27, the Mekhilta and Matthew 23. The chapter becomes stronger and fairer at the same time.
- Turn the cost-of-the-blade lens on religion honestly (caste, “unclean,” scrupulosity) and show religion’s own repairs. The reader can accept criticism when it is paired with recognition.
- Make clear that firm commitment is not the same as ideology. The draft already says “the sign is not that a person has strong commitments.” Keep that line prominent.
- Consider idolatry as a named frame: an ancient word for exactly the harm this chapter describes.
Skeptical scientist
First reactions:
- The draft says Campbell’s warning “is often called Goodhart’s law.” These are separate formulations: Goodhart (1975, on monetary targets), Campbell (1976/1979, on social indicators), and Strathern’s 1997 paraphrase that became the popular wording. The quote notes get this right and the prose blurs it. For a book about being careful with attribution, fix it.
- The reading-minutes vignette is plausible, but it’s invented. There are well-documented real cases of measures replacing purposes, including in medicine and in science itself. They’d be more convincing and more humbling.
- “Labels becoming destiny” is where pop psychology often overreaches. Pygmalion effects, stereotype threat and “gifted” labels all have contested or shrunken effect sizes. The chapter should use the evidence that holds up and say openly which famous findings didn’t.
- The “reification → ideology” step (“surprise has become impossible”) is the right idea. Scientists will know it from inside their own fields. Give them that mirror.
Questions they’d ask:
- Do you distinguish Goodhart, Campbell and Strathern, and the Lucas critique in economics? They are related but not the same.
- What’s the strongest case where optimising a proxy killed people? (There is one: see CAST below.)
- Isn’t science itself vulnerable? What happens to research when publication counts and p < 0.05 become targets?
- The Pygmalion study (Rosenthal & Jacobson 1968): how big are teacher-expectation effects after decades of follow-up? Jussim & Harber (2005, Personality and Social Psychology Review, “Teacher expectations and self-fulfilling prophecies: knowns and unknowns”) conclude that they’re real but usually small, and larger for some groups.
- Are you citing Rosenhan’s “On being sane in insane places” (1973, Science)? Its methods and data have been seriously questioned (Susannah Cahalan, The Great Pretender, 2019).
- The cobra-bounty story is widely told and poorly documented. Do you have a primary source, or should you use the documented Hanoi rat-tail bounty instead (Michael Vann, 2003, “Of rats, rice, and race: The Great Hanoi Rat Massacre,” French Colonial History)?
- How do you tell a label that shapes a person from a label that correctly describes one? What evidence would separate the two?
- When a measure has a blind spot built into its instrument, whose side of the cut is “more real”?
- What does “repair” look like with numbers? Rotating metrics, using several indicators, auditing by people who don’t benefit from the score?
- Is there a case where dropping a category did measurable good?
Where they’d get lost, bored, offended or unconvinced:
- “‘Lazy’ can hide grief, a chronic illness…” is humane, but unsupported. As written it reads as a moral appeal, not an analysis of cost.
- “A stack of categories can become ideology” jumps from measurement to politics in one sentence without a mechanism. The sceptic will want the intermediate steps: incentives, feedback loops, selection on what gets reported.
- The draft treats reification as the failure mode but never says how to detect it early. “Invite reports from people the category fits badly” is good advice. Show it working somewhere.
- If the chapter uses stereotype threat, a well-read reader will think of the replication problems (e.g. Flore & Wicherts 2015, Journal of School Psychology, meta-analysis finding signs of publication bias), and trust will drop.
Examples they’d bring:
- The Cardiac Arrhythmia Suppression Trial (CAST Investigators 1989, New England Journal of Medicine). Encainide and flecainide suppressed arrhythmias on ECG, the surrogate measure, but increased deaths compared with placebo. It’s the most sobering case of a proxy standing in for its purpose.
- English A&E four-hour targets (Bevan & Hood 2006, Public Administration, “What’s measured is what matters: targets and gaming in the English public health care system”). Documented gaming: patients held in ambulances, corridors relabelled.
- “The natural selection of bad science” (Smaldino & McElreath 2016, Royal Society Open Science). A model showing that rewarding publication count selects for low-power methods even when no one cheats. Measures replacing purposes inside the institution meant to guard against it.
- Pulse oximetry and skin colour (Sjoding et al. 2020, NEJM, “Racial bias in pulse oximetry measurement”). Hidden low blood-oxygen levels were missed about three times as often in Black patients. The measure’s blind spot became a clinical blind spot.
- Race-adjusted eGFR: the kidney-function equation’s race coefficient was removed after the NKF–ASN task force’s 2021 recommendation, because a category built into a formula delayed transplant eligibility. A label hardened into arithmetic, then deliberately undone.
- ADHD and birth month (Morrow et al. 2012, CMAJ; Layton et al. 2018, NEJM). The youngest children in a school year are more likely to be diagnosed and medicated. A cut made by the calendar ends up treated as a trait of the child.
What would win them over:
- Real, cited cases with outcomes, especially from medicine and science, and not only from education and management.
- Honest calibration: saying which famous “labelling” findings shrank or failed to replicate, and which held up.
- Accurate attribution of Goodhart, Campbell and Strathern.
- Concrete repair mechanisms that have actually been tried (multiple indicators, preregistration, clinical-endpoint trials, removing a coefficient), not only a stance.
Earlier draft
From “The Cost of the Blade”
The school wants to improve reading. It picks one number: minutes read each night. Soon children are logging time beside books they have not opened, parents are negotiating with clocks, and a teacher who reads aloud to a restless class has no number for the work. The measure was not foolish. Reading matters. The trouble began when the number became the task.
Every blade has this risk. A distinction makes action possible by omitting detail. A category makes a crowd legible. A diagnosis gathers a confusing pattern. A target helps a group see whether it is improving. These are gains. The cost appears when the omitted material is precisely what a situation needs us to notice.
Campbell’s warning is useful here: a social indicator can become corrupted when too much pressure is put on it as a decision tool. This is often called Goodhart’s law. The point is not that measurement is bad. It is that a proxy is a stand-in. If we forget the difference between the stand-in and the purpose, we begin serving the dashboard rather than the child, patient, customer, river, or neighborhood it was meant to help.
The same forgetting happens with people. “Lazy” can hide grief, a chronic illness, impossible childcare, boredom, a bad manager, or ordinary avoidance. “High-performing” can hide exhaustion, fear, and someone else’s invisible care work. A label may catch something real. It becomes cruel when it is treated as the whole explanation or as a reason to stop looking.
This is reification: a working cut hardens into an object that seems to explain itself. From there, a stack of categories can become ideology, a system that already knows what every new event means. The sign is not that a person has strong commitments. The sign is that surprise has become impossible: every counterexample is dismissed before it can count.
There is no clean escape into uncut perception. A hospital cannot operate without triage; a garden cannot share water without some rule; a person cannot keep every promise without deciding which obligations come first. The repair is more demanding and more ordinary: keep the purpose visible, revisit the measure, invite reports from people the category fits badly, and make changes before the system has to defend itself.
Use Labels and their edges or Map and territory. Ask of a rule, score, or identity: What does it let us do? What can it no longer see? Who could tell us that it is failing? Those questions do not abolish a blade. They keep it from becoming a prison.
After all this cutting, the book needs to ask what we build with the pieces: not just information as a pattern, but knowledge as a practice that can be corrected.