Two complementary pieces · the spec
§1 Claim
When AI systems can produce plausible claims and indistinguishable forgeries faster than any individual can verify them, the decisive skill becomes evaluation, not access to information or its production. Evaluation cannot be centralized without recreating the problem it solves: no evaluator can fully verify its own output. It must therefore be distributed across many independent evaluators. The primary function of education should shift from teaching vetted conclusions to training people to evaluate.
Scope The claim is conditioned on claim-type (§6). Distribution is predicted to outperform lone expertise where independent errors can cancel (forecast-like empirical claims, contested judgments, historical reconstruction) and to underperform it on formal proofs and rare-skill determinate tasks; one competent radiologist beats a crowd reading an MRI.24
Sub-claims:
- Evaluation is several distinct methods, each with its own standard of validity and blind spot (§6); selecting the right one is a prerequisite skill.
- Distribution beats a single evaluator only when the evaluators are genuinely independent and individually better than chance; correlated evaluators produce agreement that proves nothing.21,24
- AI functions as an input to the process (retrieval, drafting, first-pass filtering), never the final authority; an evaluator that is not itself evaluated moves the problem up one level.
- Distributed evaluation reduces error relative to a single evaluator; it does not eliminate it, and it cannot verify its own foundational assumptions.
§2 Definitions
- Floor / frame. Floor: facts that hold independently of belief. Frame: the selection, description, and weighting applied to any account of those facts. Two failure modes: treating the frame as the floor, and denying the floor.
- Access / legibility / power. Three independent properties of information; having one does not confer the others. Defined in the figure below.
- Provenance. The source and endorsement of a claim, historically a proxy for truth. At near-zero forgery cost, testimonial provenance stops correlating with truth. Cryptographic provenance (signed capture, content credentials6) relocates rather than restores trust: control of signing authorities is an instance of F3.
- Calibration. Matching stated confidence to evidential warrant.
- Node. Any individual or system that submits a judgment to a shared evaluation process. Apparatus: the set of nodes and the process connecting them.
- Independence operationalized Low pairwise error correlation between nodes, measured on a common claim set with known answers. The measure is direct only where answers are known. On contested and historical claims the floor arrives late, incomplete, or inseparable from the frame; independence there is a proxy carried over from the checkable case, and that transfer is an assumption of the framework, not a result. Agreement between correlated nodes carries no evidential weight beyond one node.21
- Capture. The condition in which the apparatus’s standards, membership, or inputs come under control of a small set of actors, measurable as source concentration across nodes and error-correlation drift over time.
§3 Problem
Four conditions motivate the claim:
- Provenance collapse. Forgery at near-zero cost breaks source-based verification; producing false claims scales, verifying them does not.
- Illegibility. Systems too complex for non-specialists to evaluate; accountability diffuses until no single actor is responsible.5
- Consent deficit. The systems that mediate public information were adopted without informed consent (unread terms, no alternative).
- Scaled canon formation. Model training fixes defaults for what is treated as true, set by no accountable or contestable process, a diffuse form of value lock-in.7
§4 Why single-authority solutions fail
- Central arbiter. Raises the question of who evaluates the arbiter, with no stopping point.
- Expert rule. No neutral method selects the qualified experts without that method being controlled by someone. Also removes reversibility.11
- Mandated method (“teach correct reasoning”). Whoever defines correct reasoning fixes the permitted conclusions in advance.
§5 Framework
Six operating principles, stated in the figure below. They substantially restate, for a general population, Longino’s conditions for objectivity in science18 (see §10). Competence boundaries carry their own failure mode (F7); continuous operation is what keeps the apparatus from hardening into another uncontestable authority.
§6 Evaluation methods
How a claim is judged, split by the kind of correctness available: whether a determinate answer exists to check against, or not. Choosing the right branch is the prerequisite skill; the common error is applying the wrong one.
Formal
Empirical
Pragmatic
Historical
Aesthetic
Ethical
Adversarial PROCEDURE
Calibration META-CHECK
The historical method was missing from the first version of this spec. Claims about the Anarchy, the war the companion essay descends from, are determinate (the war happened some way) but cannot be replicated, derived, or load-tested; the check historians developed is inference to the best explanation30 from convergent independent witnesses, disciplined by source criticism.23 It is the framework’s oldest running instance.
§7 Implications for education
The trained outcome is a person who can:
- identify which evaluation method a claim requires;
- apply several methods and state each one’s blind spot;
- calibrate confidence, including recognizing the limit of their own competence;
- distinguish floor from frame;
- submit and revise judgments within a shared process.
This does not replace conclusions with content-free “skills.” Evaluation is substantially domain-bound12; vetted conclusions change role rather than disappear, taught through evaluating claims within each domain. The untested kernel, after prior evidence is subtracted (§10), is whether competence-boundary recognition is trainable at population scale.14,31 Training sets the boundary; institutions make it usable by routing what lies past one person’s edge to evaluators inside theirs (F2).
§8 Failure conditions
- F1. Correlated evaluators. Nodes sharing sources, biases, or incentives produce correlated errors; their agreement is not evidence of correctness.21 Widespread use of the same few AI systems raises correlation toward monoculture.22 Tracked as pairwise error correlation and source concentration across nodes. The framework’s most fragile requirement.
- F2. Competence limits. A population trained enough to feel competent but not to be competent may lower the reliability of the process below one that defers,14 the standing objection since Lippmann.13 Partial patches: division of epistemic labor, so that no node evaluates everything and each question routes to evaluators working inside their competence;37 and institutional scaffolds (credential registries, disclosure of conflicts, published track records) that turn the second-order checks of F7 into a lookup rather than a computation no bounded mind can run.15 Both patches presuppose institutions not yet captured (F3). The objection stands.
- F3. Ownership and capture. The systems capable of making information legible are controlled by parties that benefit from illegibility; no path around that incentive is specified here. The nearest playbook is the empirical study of commons governance.25 This is the likeliest failure.
- F4. Inaction. Emphasis on uncertainty can suppress action while less cautious actors proceed. Whether calibrated action is sustainable under competitive pressure is unresolved.
- F5. Residual limit. The apparatus is itself a bounded system and cannot verify its own foundational assumptions.9 It reduces error; it does not reach certainty.
- F6. Overfit framework. One structure explaining a wide range of cases may be a real pattern or a frame fitted after the fact. It must be tested against cases it predicts should fail; P1’s two-sided form supplies those.
- F7. Deference regress. Selecting whom to defer to is itself an evaluation, often made at or beyond the deferrer’s competence; the regress escaped at the top (§4) re-enters at the bottom. Partial patches: second-order criteria for judging experts (track record, conflicts of interest, responsiveness to criticism) and meta-expertise.15 The framework carries this problem; it does not solve it.
- F8. Aggregation failure. The framework requires many independent judgments and never specifies how they compose into a collective verdict: how judgments are weighted, what rule aggregates them, how dissent is resolved. Absent that mechanism, distribution can produce polarization, duplicated effort, status competition, or coordinated manipulation instead of error-correction. Partial patches: aggregation rules matched to claim type (§6). Forecast-like empirical claims have proper scoring rules and track-record-weighted pooling, tested in forecasting tournaments.29 Contested claims have bridging-based ranking, which surfaces judgments endorsed by evaluators who normally disagree; Community Notes runs that rule at platform scale.36 Historical reconstruction has convergence of independent witnesses disciplined by source criticism.23,24 Setting those rules is itself a higher-order evaluation, made before the apparatus exists to check it (the regress of §4 and F7, arriving at the design stage). This is the central implementation problem, not a secondary exception.
§9 Falsification
The hypothesis makes the following predictions; each is checkable, and it fails if one proves false.
- P1 two-sided On claim sets matched to method type, a distributed process whose nodes are (i) measured-independent (§2) and (ii) individually better than chance outperforms a single qualified expert on forecast-like empirical claims, contested claims, and historical reconstruction, and fails to outperform the expert on formal and rare-skill determinate claims. The winning half is three tests, not one. Forecast-like, contested, and historical claims are judged by different methods (§6); each runs as a separate leg on its own claim set. Any leg failing, on either side, falsifies (F1, F6).
- P2 narrowed Competence-boundary recognition can be taught to a general population and measurably improves aggregate evaluation accuracy. The target variable is the Kruger–Dunning over-placement gap.14 The bet is that a brief, teachable intervention narrows it by a pre-registered margin and that the narrowing persists past the training window rather than decaying as inoculation effects do.28 If the gap does not move, moves only for those already competent, or teaching it suppresses participation wholesale, the education claim fails (F2, F4).
- P3 design criterion An apparatus built to §5, under real adoption, maintains measured node independence above a pre-registered threshold for a pre-registered horizon, with capture metrics (source concentration, correlation drift) published continuously. Independence decaying below threshold within the horizon, across attempts, falsifies feasibility (F3, F1).
- forecast-like empirical claims
- contested judgments
- historical reconstruction
- formal proofs
- one radiologist reading an MRI
Status These predictions are checkable in principle, not yet in procedure. A real test requires named datasets, sampling procedures, comparison conditions, aggregation rules, numerical independence thresholds, minimum effect sizes, and time horizons; none are fixed here, and P3 says “pre-registered” where the numbers should stand. The forecast leg has a ready instrument in resolved forecasting-tournament question sets;29 the other legs do not. This section is falsifiability-conscious, not a pre-registration. Stating the debt is cheaper than pretending it is paid.
§10 Prior art & prior evidence
A specification claiming testability must say which of its predictions others already tested. What remains novel is the composite, the taxonomy’s scoping, and P2’s kernel.
| Component | Nearest prior work | Status |
|---|---|---|
| Distributed error-correction over authority | Popper (1945); Longino (1990); Rauch (2021)11,18,19 | inherited frame |
| Independence condition & when crowds beat experts | Condorcet (1785); Breiman (2001); Kleinberg & Raghavan (2021)24,21,22 | formalized elsewhere; adopted |
| Benefit of cognitive diversity | Hong & Page (2004; generality contested, Thompson 2014); Landemore (2013); Mercier & Sperber (2017)17,16 | suggestive, contested |
| Teachability of evaluation components | lateral reading (Wineburg et al. 2022); inoculation (Roozenbeek & van der Linden 2019); calibration training (Mellers et al. 2014)27,28,29 | partially confirmed pre-spec |
| Domain-boundedness constraint | Willingham (2007)12 | constraint adopted |
| Deference criteria (F7 patches) | Anderson (2011); Collins & Evans (2007)15 | partial patches only |
| Building commons against capture (F3) | Ostrom (1990)25 | nearest playbook; unproven at this scale |
| Historical method as running instance | source criticism (Bloch 1949); IBE (Lipton 1991)23,30 | oldest confirmed instance |
| Method-selection as master skill; boundary-recognition kernel; two-sided P1; pairing form | none | staked here |
Status: a synthesis, not a discovery. Component claims carry the prior evidence cited above; the composite frame, the taxonomy’s scoping of P1, and P2’s kernel remain unverified; the protocols that would test §9 remain unwritten.
- George R. R. Martin, Fire & Blood (Bantam, 2018). Martin has acknowledged the Dance of the Dragons drew in part on the Anarchy; the book's frame (Archmaester Gyldayn reconciling the irreconcilable testimonies of Mushroom, Septon Eustace, and Grand Maester Munkun) reproduces the source condition of the real war.
- Gesta Stephani, ed. K. R. Potter, rev. R. H. C. Davis (Oxford Medieval Texts, 1976); William of Malmesbury, Historia Novella, ed. Edmund King, trans. K. R. Potter (Oxford Medieval Texts, 1998). The pro-Stephen and Angevin-leaning witnesses, respectively.
- The Peterborough Chronicle (Anglo-Saxon Chronicle, MS E), annal for 1137.
- Edmund King, King Stephen (Yale University Press, 2010); David Crouch, The Reign of King Stephen, 1135–1154 (Longman, 2000). The revisionist case that the disorder was severe but regional, argued substantially from non-narrative evidence: charters, writs, and coinage struck through the war years.
- Onora O'Neill, A Question of Trust (Cambridge University Press, 2002). The BBC Reith Lectures arguing that transparency has been mistaken for, and cannot substitute for, accountability.
- Coalition for Content Provenance and Authenticity (C2PA), Content Credentials technical specification, c2pa.org. The principal industry effort to rebuild provenance cryptographically.
- William MacAskill, What We Owe the Future (Basic Books, 2022), on value lock-in: the risk that systems built at scale entrench one era's judgments.
- R. P. C. Hanson, The Search for the Christian Doctrine of God: The Arian Controversy, 318–381 (T&T Clark, 1988). The standard account of the fifty-plus years of contest between Nicaea (325) and Constantinople (381).
- Kurt Gödel, "Über formal unentscheidbare Sätze der Principia Mathematica und verwandter Systeme I," Monatshefte für Mathematik und Physik 38 (1931): the incompleteness theorems. Used here as illustration of the shape of a limit, not as proof about societies.
- Plato, Republic, Book VI (the ship of state, 488a–489d): the navigator who ought simply to be given the helm.
- Karl Popper, The Open Society and Its Enemies (Routledge, 1945), esp. vol. 1, ch. 7: replacing "who should rule?" with the question of how bad rulers can be removed without bloodshed.
- Daniel T. Willingham, "Critical Thinking: Why Is It So Hard to Teach?" American Educator 31, no. 2 (Summer 2007): the case that critical thinking is substantially domain-bound.
- Walter Lippmann, Public Opinion (1922) and The Phantom Public (1925); John Dewey, The Public and Its Problems (1927). The original debate over whether a general public can be competent to judge.
- Justin Kruger and David Dunning, "Unskilled and Unaware of It," Journal of Personality and Social Psychology 77, no. 6 (1999): miscalibration concentrated precisely where competence is lowest.
- Elizabeth Anderson, "Democracy, Public Policy, and Lay Assessments of Scientific Testimony," Episteme 8, no. 2 (2011): second-order criteria (track record, conflicts, responsiveness to criticism) by which laypeople can judge experts; Harry Collins and Robert Evans, Rethinking Expertise (University of Chicago Press, 2007): meta-expertise.
- Hugo Mercier and Dan Sperber, The Enigma of Reason (Harvard University Press, 2017): reasoning as an evolved social capacity, biased in solitary use, effective in argumentative exchange.
- Lu Hong and Scott E. Page, "Groups of diverse problem solvers can outperform groups of high-ability problem solvers," PNAS 101, no. 46 (2004); for the contested generality of the result, Abigail Thompson, "Does Diversity Trump Ability?" Notices of the AMS 61, no. 9 (2014); Hélène Landemore, Democratic Reason (Princeton University Press, 2013).
- Helen Longino, Science as Social Knowledge (Princeton University Press, 1990). Objectivity as a property of a social practice meeting four conditions: recognized avenues for criticism, uptake of criticism, public standards, and tempered equality of intellectual authority.
- Jonathan Rauch, The Constitution of Knowledge: A Defense of Truth (Brookings Institution Press, 2021).
- Elizabeth L. Eisenstein, The Printing Press as an Agent of Change (Cambridge University Press, 1979), including print's double edge: pamphlet wars and witch manuals alongside the republic of letters.
- Leo Breiman, "Random Forests," Machine Learning 45, no. 1 (2001): ensemble generalization error bounded in terms of the strength of individual members and the correlation between them.
- Jon Kleinberg and Manish Raghavan, "Algorithmic monoculture and social welfare," PNAS 118, no. 22 (2021): many decision-makers adopting the same superior algorithm can lower aggregate outcomes.
- Marc Bloch, The Historian's Craft (posthumous, 1949; written before his execution by the Gestapo in 1944). The critical method: cross-examining witnesses who cannot be recalled.
- Marquis de Condorcet, Essai sur l'application de l'analyse à la probabilité des décisions rendues à la pluralité des voix (1785). The jury theorem: with independent, better-than-chance judges, group accuracy rises with size; below chance, it falls.
- Elinor Ostrom, Governing the Commons (Cambridge University Press, 1990): empirical design principles for governing shared resources against extractive incentive.
- Neil Postman and Charles Weingartner, Teaching as a Subversive Activity (Delacorte, 1969), ch. 1, "Crap Detecting": education's job as building the detector, not delivering the conclusions.
- Sam Wineburg, Joel Breakstone, Sarah McGrew, Mark D. Smith, and Teresa Ortega, "Lateral reading on the open Internet: A district-wide field study in high school government classes," Journal of Educational Psychology 114, no. 5 (2022): classroom-taught source evaluation measurably improved.
- Jon Roozenbeek and Sander van der Linden, "Fake news game confers psychological resistance against online misinformation," Palgrave Communications 5 (2019); accessibly synthesized, with decay caveats, in van der Linden, Foolproof (W. W. Norton, 2023).
- Barbara Mellers et al., "Psychological strategies for winning a geopolitical forecasting tournament," Psychological Science 25, no. 5 (2014); Philip E. Tetlock and Dan Gardner, Superforecasting (Crown, 2015): brief calibration training measurably improved forecast accuracy.
- Peter Lipton, Inference to the Best Explanation (Routledge, 1991; 2nd ed. 2004).
- Mark R. Leary et al., "Cognitive and Interpersonal Features of Intellectual Humility," Personality and Social Psychology Bulletin 43, no. 6 (2017): the young measurement literature on intellectual humility, the trait P2's kernel proposes to train.
- John Wihbey, "AI and Epistemic Risk for Democracy: A Coming Crisis of Public Knowledge?" (SSRN working paper, 2024); and "In Post-Authenticity AI Age, Knowledge Institutions Matter More than Ever," Tech Policy Press (November 14, 2025): the nearest recent argument that provenance collapse makes knowledge institutions more, not less, decisive. Institution-centric where this piece is education-centric.
- "A robot wrote this entire article. Are you scared yet, human?" The Guardian (September 8, 2020): an op-ed generated by GPT-3, prompted by the paper and assembled by its editors from eight outputs (run by Liam Porr); an early instance of AI authorship deployed as a rhetorical reveal, the device the essay's confession both uses and disowns.
- Ethan Mollick, "Post-apocalyptic education," One Useful Thing (August 30, 2024): redesigning education around AI, away from transmitting conclusions and toward judgment.
- Archon Fung, Mary Graham, and David Weil, Full Disclosure: The Perils and Promise of Transparency (Cambridge University Press, 2007); Mike Ananny and Kate Crawford, "Seeing without knowing: Limitations of the transparency ideal and its application to algorithmic accountability," New Media & Society 20, no. 3 (2018): disclosure scholarship distinguishing information disclosed from information usable and actionable.
- Stefan Wojcik et al., "Birdwatch: Crowd Wisdom and Bridging Algorithms can Inform Understanding and Reduce the Spread of Misinformation," arXiv:2210.15723 (2022): the design of X's Community Notes (formerly Birdwatch), a bridging-based ranking that surfaces notes rated helpful by contributors who normally disagree; a deployed aggregation rule for contested claims.
- Philip Kitcher, "The Division of Cognitive Labor," The Journal of Philosophy 87, no. 1 (1990): how a community should distribute investigative effort across rival approaches; what the community learns depends on the allocation, not on each member's rationality alone.
Provenance of this document itself. This specification is the falsifiable half of a pair; the essay half is The Chronicler’s Problem, and the ledger of what the pair inherits and stakes lives there. The spec was drafted from the essay's argument, then revised under external review by differently instructed AIs. Every editorial call was human; added citations were verified before inclusion. If that chain changes your assessment, the pair is about why.