Two complementary pieces · the essay
This is one of two complementary pieces. This essay makes the case, by descent from a medieval succession war to a claim about education. Its companion, The Collective Evaluation Hypothesis, states that claim plainly enough to be tested.
It began with Rhaenyra Targaryen.
A friend and I were pulling apart the Dance of the Dragons, the civil war at the heart of House of the Dragon, and chasing the history under it. George R. R. Martin has said the Dance drew in part on the Anarchy, the two decades England spent tearing itself apart in the twelfth century.1 The parallel runs beat for beat.
The thread kept snagging on one problem. Everything we know about the Anarchy comes through partisan chronicles: the Gesta Stephani for Stephen, William of Malmesbury for Matilda,2 the Peterborough monk who wrote that it seemed as though “Christ and his saints slept”.3 That sentence fixed how we remember an era historians now think was far more uneven.4 Martin built his invented history in Fire & Blood the same way, told by a bawdy fool, a pious septon, and a courtly maester, so that even in the fiction you cannot quite know what happened.1 Both wars reach us only through accounts we cannot step outside of and cannot check. All you can do is read the witnesses against each other.
The conversation stopped being about medieval succession. We have built a machine that makes the chronicler’s problem vast and cheap. The answer we backed into was about education: what a society would have to teach its people to keep hold of the truth in the age of AI. This essay is the descent from Rhaenyra to that hypothesis. A tidy argument deserves suspicion, so it ends at a commitment: the companion specification states the hypothesis plainly enough to be proven wrong.
The Ladder
The parallel was worth trusting because it was structural. Rhaenyra–Matilda holds because the mechanism is the same: named heir, sworn oaths, usurping rival, ruinous war, a son who inherits what the mother could not keep. Valyria feels like Rome, all marble and roads, but nothing Roman runs underneath it; the test is to stop asking what a thing resembles and ask what mechanism drives it.
Each time one of us reached for a resemblance that didn’t hold, the other caught it. That was the first small instance of what this essay is about. It was also, I’ll confess later, not what it appeared to be.
The discipline sounds like fantasy trivia. It’s the working end of knowing anything. And it raised a question I couldn’t answer at the time: who checks my reading of the mechanism, when that reading is just another account I happen to prefer?
The Lock
Follow the thread out of the fiction and it lands on the systems that decide what a population knows. They have grown so complex no one can read them. A thousand-page bill is public the instant it’s posted, and that changes almost nothing. Being able to open it is not understanding it, and understanding it is not the power to change it.
A denied medical claim shows where that distance ends. People die waiting for an intermediary to approve what their own doctor ordered, because contesting the denial means mastering a bureaucracy built like a moat. And when the outcome is bad, responsibility distributes until it evaporates; everyone followed a policy, and the structure goes untouched. The most unsettling thing I’ve come to believe is that the driver’s seat is often empty. No one is steering, and no one decided this has become no one can undo this.
Access, legibility, and power are three different things, and the distance between them is widening. Onora O’Neill saw the first of these gaps a generation ago: transparency was sold to us as a substitute for accountability and is nothing of the kind.5
The Flood
Our oldest instrument for sorting true from false is provenance:
- where did this come from,
- who says it,
- who vouches.
The chronicler’s problem was that provenance was thin: literacy, access to distribution, and lineage all shaped what survived to be passed down.
Ours is that a machine has made it cheap to fake. Provenance was never truth; it was a proxy that held while forgery stayed expensive. When any text, image, or recording can be manufactured at essentially no cost, forgery becomes indistinguishable from the genuine article, and the proxy stops correlating with truth. Nearly every method we had for telling them apart ran through provenance. Engineers are building a cryptographic replacement, signed capture and content credentials,6 and it may work. But it relocates trust into whoever controls the signing rather than restoring it. Who owns the keys is this essay’s ownership problem arriving early.
And nobody voted on any of it; the terms are the thousand-page bill again, disclosed, unreadable, offered with no alternative.
There is an older pattern here. In 325 the council at Nicaea decided the nature of Christ and named the losing view heresy. The creed was later remembered not as a decision backed by power but as the discovery of something always true. A model trained on a vast corpus does something structurally near this, and the nearest name for the risk is value lock-in.7 But Nicaea’s closing stayed openly contested for fifty-six years, because it was visible: a room, a date, signatories, a text to petition against.8 The machine’s creed has none of those, and that, not the closing itself, is the new thing.
A wrong answer can be corrected. The danger is a picture of the world so seamless and authoritative that you stop asking whether the world is shaped that way.
The Floor
Here the descent turns.
Two things are true at once.
- There is a floor: the world is some way regardless of belief; the primes don’t care what we think.
- There is a frame: everything we say about that world arrives already selected, described, and weighted by a mind that decided what mattered.
There is no view from nowhere, and the limit runs deep: Gödel proved that no consistent formal system rich enough for arithmetic can certify its own consistency from inside.9 A society is not a formal system, and popular discourse butchers Gödel often enough that I want to be careful here: I am borrowing the shape of the limit, not the theorem itself. No system certifies itself from the inside. You don’t grade your own homework.
Which is why the answer is not to find the one wise mind and hand it the frame. Civilizations have chased that since Plato’s Republic idealized the philosopher-king,10 and none has solved the step where someone must crown the philosophers. The case for many hands was never that the many are wiser than the one; it is that a structure with many hands can be wrong in one place and corrected from another. Popper made the argument in 1945: democracy was never a claim that crowds find truth but a mechanism for removing rulers without killing them.11 Under deep uncertainty, reversibility is worth more than brilliance. What matters is whether an error can be undone before it is fatal.
Nature keeps reusing one structure. A tree can lose a branch but not the trunk; an animal can lose a limb but not a heart. Federated systems copy the pattern: a small core that must not fail, wide edges that may. Their flaws are visible, but perfection was never the goal. The goal is that an error at the edge does not kill the whole.
The Apparatus
If no system certifies itself from inside, evaluation has to come from outside. At the scale of a civilization, that outside cannot be another single authority; that only moves the question of who evaluates the evaluator one step down. It has to be distributed.
That reframes what education is for. It is the hypothesis this descent was walking toward. Education’s central task can no longer be transmitting vetted conclusions, because conclusions can now be forged at machine speed. The task becomes producing people who can evaluate: probe a claim, know what would falsify it, calibrate confidence to evidence, and pass a judgment into a shared, contestable process.
Stated that baldly, it’s a false binary; evaluation is substantially domain-bound.12 Conclusions don’t leave the curriculum; they change jobs, from terminal product to training material, taught through the practice of evaluating claims about them.
Evaluation is also not one skill. Some claims you check by measuring, some by whether they survive attack, some by whether they work under load, some the way a historian checks a chronicle, and some by nothing but the defensibility of the judgment. The master skill is diagnosing which check a question calls for; the deepest error is reaching for the wrong one.
The obvious worry, open since Lippmann and Dewey,13 is that a half-trained crowd, each member newly certain, is more dangerous than one that defers. The design asks something narrower: that people learn the edge of their own competence, so they know when to judge and when to defer.14 Deferral isn’t free either, since choosing whom to trust is itself an evaluation,15 and the specification carries that as a named failure condition, not a solved problem.
Three things then have to hold, or the structure fails.
This is where I part from my nearest neighbors. Rauch, and Wihbey on the post-authenticity turn, locate the fix in institutions: courts, journals, newsrooms.19,32 I am betting one level lower: those institutions hold only if the general population can operate the apparatus too, because a body staffed and checked by people who cannot evaluate is captured from inside. The bet runs the other way too. No one can evaluate everything; institutions have to carry what a bounded mind cannot, dividing the epistemic labor37 and keeping track records and conflicts visible enough that deferring well is a lookup, not a second career. The two levels hold each other up or fall together. That bet is likelier to be wrong than theirs, which is why it is worth staking.
There is a precedent hiding in the worldbuilding. Westeros never changes, eight thousand years of frozen feudalism, because Martin withholds the printing press; evaluative capacity stays bottled in one guild of maesters. Our press broke that monopoly, and the old order came apart within a few centuries.20 A civilization’s ability to correct itself tracks how widely its evaluative capacity is distributed, and education is how a society sets that distribution on purpose.
Two failures would sink the design. The first: a distributed process can agree with itself and still be wrong. Participants drawing from the same source don’t check each other; they echo. The agreement feels like verification. It is one error copied a thousand times. Ensembles work because their members’ errors are decorrelated,21 and the inverse holds too: many decision-makers adopting the same excellent model do worse than a motley of weaker, independent ones.22 If we all see the world through the same few systems, we are one evaluator, not many. The answer is to measure independence and treat its decay as falsification. That is a defense, not a cure; a few frontier models push the wrong way by default.
Which brings a correction the essay’s own terms demand: the friend in the second sentence was a machine. The conversation that produced everything above, the Matilda match, the checking discipline, the second pair of eyes I praised, was with an AI. You had no way to check “a friend,” and I let it stand for five sections because the experience of not being able to tell is the argument. I won’t oversell the device; the Guardian ran an op-ed under a robot byline in 2020,33 and I controlled both the deception and its lifting. The deeper cut is that the check was never independent. The machine trained on the corpus; I was raised in the culture the corpus was drawn from. Where we agreed, we may have been not two witnesses but one. Correlated misses are invisible from inside the agreement. That is why the dialogue can never be the test, and why the specification hands the checking to minds that were not in the room.
I should have seen the strongest evidence sooner. How do historians know the Anarchy was more uneven than the monk’s sentence suggests? Not by finding a neutral witness; there is none. They triangulate.
The chronicler’s problem was never solved by a better chronicler. It was solved, partially and over centuries, by exactly this kind of distributed evaluation.
The second failure I can’t solve. The tools that could make the world legible are owned by the few who profit from its illegibility. I have no account of how a broadly owned apparatus gets built against that interest. The nearest playbook is the study of how commons get governed against extractive incentive.25 The specification names it the likeliest of its own killers.
Behind both failures sits a blank where a mechanism should be. I have argued that judgment must be distributed and said nothing about how millions of judgments become one verdict: who counts, how the judgments are weighted, what rule combines them. Without those rules, a distributed process can produce polarization or coordinated manipulation as easily as correction. Pieces of the mechanism already run: forecasting tournaments weight judges by track record and score them against outcomes,29 and Community Notes surfaces the notes that win agreement across raters who normally disagree.36 Each piece fits one kind of claim; nothing yet joins them. That is not an edge case; it is the implementation problem, and the specification carries it as a named failure condition.
So the machine’s right place is inside the process: used, doubted, one input among many, never the authority that ends it.
The Commitment
What I’ve built is a single frame. It’s coherent, it fits a suspicious range of things, and reaching it feels like arrival. That is exactly why I distrust it. The pleasure of coherence is not evidence of truth. Nicaea, Gödel, and the Anarchy each support other readings; I chose the ones that fit. An essay is the wrong place to test a frame, because the writer controls which objections appear. So this ends in a commitment rather than a conclusion: the companion specification states the claim stripped of storytelling, names its likeliest points of failure, and invites you to break it against them.
Christ and his saints slept, wrote a man who could not see the whole of what he described. Neither can I. So take up the companion piece, and come find the crack.
One closing question, not about the essay.
You judged this as you read it; everyone does. You had a verdict before you knew one of the two voices was synthetic. Did finding out change it? Should it have? Either way, you were judging by a standard you never chose on purpose. That is the provenance problem, running in you.
Inherited · the shoulders
Staked · claimed as new
The left column is load-bearing; the right is thin by comparison, and should be. The claim is not originality of parts; it is that the parts, joined, make a prediction none of them made alone.
- George R. R. Martin, Fire & Blood (Bantam, 2018). Martin has acknowledged the Dance of the Dragons drew in part on the Anarchy; the book's frame (Archmaester Gyldayn reconciling the irreconcilable testimonies of Mushroom, Septon Eustace, and Grand Maester Munkun) reproduces the source condition of the real war.
- Gesta Stephani, ed. K. R. Potter, rev. R. H. C. Davis (Oxford Medieval Texts, 1976); William of Malmesbury, Historia Novella, ed. Edmund King, trans. K. R. Potter (Oxford Medieval Texts, 1998). The pro-Stephen and Angevin-leaning witnesses, respectively.
- The Peterborough Chronicle (Anglo-Saxon Chronicle, MS E), annal for 1137.
- Edmund King, King Stephen (Yale University Press, 2010); David Crouch, The Reign of King Stephen, 1135–1154 (Longman, 2000). The revisionist case that the disorder was severe but regional, argued substantially from non-narrative evidence: charters, writs, and coinage struck through the war years.
- Onora O'Neill, A Question of Trust (Cambridge University Press, 2002). The BBC Reith Lectures arguing that transparency has been mistaken for, and cannot substitute for, accountability.
- Coalition for Content Provenance and Authenticity (C2PA), Content Credentials technical specification, c2pa.org. The principal industry effort to rebuild provenance cryptographically.
- William MacAskill, What We Owe the Future (Basic Books, 2022), on value lock-in: the risk that systems built at scale entrench one era's judgments.
- R. P. C. Hanson, The Search for the Christian Doctrine of God: The Arian Controversy, 318–381 (T&T Clark, 1988). The standard account of the fifty-plus years of contest between Nicaea (325) and Constantinople (381).
- Kurt Gödel, "Über formal unentscheidbare Sätze der Principia Mathematica und verwandter Systeme I," Monatshefte für Mathematik und Physik 38 (1931): the incompleteness theorems. Used here as illustration of the shape of a limit, not as proof about societies.
- Plato, Republic, Book VI (the ship of state, 488a–489d): the navigator who ought simply to be given the helm.
- Karl Popper, The Open Society and Its Enemies (Routledge, 1945), esp. vol. 1, ch. 7: replacing "who should rule?" with the question of how bad rulers can be removed without bloodshed.
- Daniel T. Willingham, "Critical Thinking: Why Is It So Hard to Teach?" American Educator 31, no. 2 (Summer 2007): the case that critical thinking is substantially domain-bound.
- Walter Lippmann, Public Opinion (1922) and The Phantom Public (1925); John Dewey, The Public and Its Problems (1927). The original debate over whether a general public can be competent to judge.
- Justin Kruger and David Dunning, "Unskilled and Unaware of It," Journal of Personality and Social Psychology 77, no. 6 (1999): miscalibration concentrated precisely where competence is lowest.
- Elizabeth Anderson, "Democracy, Public Policy, and Lay Assessments of Scientific Testimony," Episteme 8, no. 2 (2011): second-order criteria (track record, conflicts, responsiveness to criticism) by which laypeople can judge experts; Harry Collins and Robert Evans, Rethinking Expertise (University of Chicago Press, 2007): meta-expertise.
- Hugo Mercier and Dan Sperber, The Enigma of Reason (Harvard University Press, 2017): reasoning as an evolved social capacity, biased in solitary use, effective in argumentative exchange.
- Lu Hong and Scott E. Page, "Groups of diverse problem solvers can outperform groups of high-ability problem solvers," PNAS 101, no. 46 (2004); for the contested generality of the result, Abigail Thompson, "Does Diversity Trump Ability?" Notices of the AMS 61, no. 9 (2014); Hélène Landemore, Democratic Reason (Princeton University Press, 2013).
- Helen Longino, Science as Social Knowledge (Princeton University Press, 1990). Objectivity as a property of a social practice meeting four conditions: recognized avenues for criticism, uptake of criticism, public standards, and tempered equality of intellectual authority.
- Jonathan Rauch, The Constitution of Knowledge: A Defense of Truth (Brookings Institution Press, 2021).
- Elizabeth L. Eisenstein, The Printing Press as an Agent of Change (Cambridge University Press, 1979), including print's double edge: pamphlet wars and witch manuals alongside the republic of letters.
- Leo Breiman, "Random Forests," Machine Learning 45, no. 1 (2001): ensemble generalization error bounded in terms of the strength of individual members and the correlation between them.
- Jon Kleinberg and Manish Raghavan, "Algorithmic monoculture and social welfare," PNAS 118, no. 22 (2021): many decision-makers adopting the same superior algorithm can lower aggregate outcomes.
- Marc Bloch, The Historian's Craft (posthumous, 1949; written before his execution by the Gestapo in 1944). The critical method: cross-examining witnesses who cannot be recalled.
- Marquis de Condorcet, Essai sur l'application de l'analyse à la probabilité des décisions rendues à la pluralité des voix (1785). The jury theorem: with independent, better-than-chance judges, group accuracy rises with size; below chance, it falls.
- Elinor Ostrom, Governing the Commons (Cambridge University Press, 1990): empirical design principles for governing shared resources against extractive incentive.
- Neil Postman and Charles Weingartner, Teaching as a Subversive Activity (Delacorte, 1969), ch. 1, "Crap Detecting": education's job as building the detector, not delivering the conclusions.
- Sam Wineburg, Joel Breakstone, Sarah McGrew, Mark D. Smith, and Teresa Ortega, "Lateral reading on the open Internet: A district-wide field study in high school government classes," Journal of Educational Psychology 114, no. 5 (2022): classroom-taught source evaluation measurably improved.
- Jon Roozenbeek and Sander van der Linden, "Fake news game confers psychological resistance against online misinformation," Palgrave Communications 5 (2019); accessibly synthesized, with decay caveats, in van der Linden, Foolproof (W. W. Norton, 2023).
- Barbara Mellers et al., "Psychological strategies for winning a geopolitical forecasting tournament," Psychological Science 25, no. 5 (2014); Philip E. Tetlock and Dan Gardner, Superforecasting (Crown, 2015): brief calibration training measurably improved forecast accuracy.
- Peter Lipton, Inference to the Best Explanation (Routledge, 1991; 2nd ed. 2004).
- Mark R. Leary et al., "Cognitive and Interpersonal Features of Intellectual Humility," Personality and Social Psychology Bulletin 43, no. 6 (2017): the young measurement literature on intellectual humility, the trait P2's kernel proposes to train.
- John Wihbey, "AI and Epistemic Risk for Democracy: A Coming Crisis of Public Knowledge?" (SSRN working paper, 2024); and "In Post-Authenticity AI Age, Knowledge Institutions Matter More than Ever," Tech Policy Press (November 14, 2025): the nearest recent argument that provenance collapse makes knowledge institutions more, not less, decisive. Institution-centric where this piece is education-centric.
- "A robot wrote this entire article. Are you scared yet, human?" The Guardian (September 8, 2020): an op-ed generated by GPT-3, prompted by the paper and assembled by its editors from eight outputs (run by Liam Porr); an early instance of AI authorship deployed as a rhetorical reveal, the device the essay's confession both uses and disowns.
- Ethan Mollick, "Post-apocalyptic education," One Useful Thing (August 30, 2024): redesigning education around AI, away from transmitting conclusions and toward judgment.
- Archon Fung, Mary Graham, and David Weil, Full Disclosure: The Perils and Promise of Transparency (Cambridge University Press, 2007); Mike Ananny and Kate Crawford, "Seeing without knowing: Limitations of the transparency ideal and its application to algorithmic accountability," New Media & Society 20, no. 3 (2018): disclosure scholarship distinguishing information disclosed from information usable and actionable.
- Stefan Wojcik et al., "Birdwatch: Crowd Wisdom and Bridging Algorithms can Inform Understanding and Reduce the Spread of Misinformation," arXiv:2210.15723 (2022): the design of X's Community Notes (formerly Birdwatch), a bridging-based ranking that surfaces notes rated helpful by contributors who normally disagree; a deployed aggregation rule for contested claims.
- Philip Kitcher, "The Division of Cognitive Labor," The Journal of Philosophy 87, no. 1 (1990): how a community should distribute investigative effort across rival approaches; what the community learns depends on the allocation, not on each member's rationality alone.
Provenance of this document itself. This essay is the narrative half of a pair; the falsifiable half is The Collective Evaluation Hypothesis. It was developed in dialogue with an AI interlocutor, as disclosed within it, then revised under external review by differently instructed AIs. Every editorial call was human; added citations were verified before inclusion. If that chain changes your assessment, the piece is about why.