The Collective Evaluation Hypothesis.

A falsifiable companion to The Chronicler’s Problem.

Two complementary pieces · the spec

§1 Claim

When AI systems can produce plausible claims and indistinguishable forgeries faster than any individual can verify them, the decisive skill becomes evaluation, not access to information or its production. Evaluation cannot be centralized without recreating the problem it solves: no evaluator can fully verify its own output. It must therefore be distributed across many independent evaluators. The primary function of education should shift from teaching vetted conclusions to training people to evaluate.

Scope The claim is conditioned on claim-type (§6). Distribution is predicted to outperform lone expertise where independent errors can cancel (forecast-like empirical claims, contested judgments, historical reconstruction) and to underperform it on formal proofs and rare-skill determinate tasks; one competent radiologist beats a crowd reading an MRI.24

Sub-claims:

  1. Evaluation is several distinct methods, each with its own standard of validity and blind spot (§6); selecting the right one is a prerequisite skill.
  2. Distribution beats a single evaluator only when the evaluators are genuinely independent and individually better than chance; correlated evaluators produce agreement that proves nothing.21,24
  3. AI functions as an input to the process (retrieval, drafting, first-pass filtering), never the final authority; an evaluator that is not itself evaluated moves the problem up one level.
  4. Distributed evaluation reduces error relative to a single evaluator; it does not eliminate it, and it cannot verify its own foundational assumptions.

§2 Definitions

Access · Legibility · Power§2 · a three-way split
Accessthe information is available
Legibilityit can be understood
Powerit can be acted on, or reversed
A thousand-page bill is pure access: legible to almost no one, actionable by fewer still. O’Neill named the transparency–accountability gap;5 disclosure scholarship had already split usable from actionable.35 The vocabulary is adopted, not discovered.

§3 Problem

Four conditions motivate the claim:

  1. Provenance collapse. Forgery at near-zero cost breaks source-based verification; producing false claims scales, verifying them does not.
  2. Illegibility. Systems too complex for non-specialists to evaluate; accountability diffuses until no single actor is responsible.5
  3. Consent deficit. The systems that mediate public information were adopted without informed consent (unread terms, no alternative).
  4. Scaled canon formation. Model training fixes defaults for what is treated as true, set by no accountable or contestable process, a diffuse form of value lock-in.7

§4 Why single-authority solutions fail

Why a single authority fails§4 · the regress
A central arbiter decides what is true
↓  but who evaluates the arbiter?
A second arbiter checks the first
↓  and who evaluates that one?
… no stopping point
The escape: stop looking for a final evaluator. Distribute the checking across many independent nodes (§5). The network reduces dependence on any one authority; it does not end the regress. Whoever sets the weighting and the aggregation rule still evaluates (F8).

§5 Framework

Six operating principles, stated in the figure below. They substantially restate, for a general population, Longino’s conditions for objectivity in science18 (see §10). Competence boundaries carry their own failure mode (F7); continuous operation is what keeps the apparatus from hardening into another uncontestable authority.

The framework§5 · six operating principles
1 · Distributionno single evaluator holds final authority
2 · Method pluralismmatch each claim to the right method (§6)
3 · Method selectionthe master skill; teach it first
4 · Competence boundaryknow your own edge; defer past it
5 · Broad ownershipno small group controls the standards
6 · Continuous operationa living practice, not a fixed rule
AI’s role: an evaluated input, never the final evaluator.

§6 Evaluation methods

Evaluation methodsCollective Evaluation Hypothesis · §6

How a claim is judged, split by the kind of correctness available: whether a determinate answer exists to check against, or not. Choosing the right branch is the prerequisite skill; the common error is applying the wrong one.

DETERMINATEa fact of the matter exists: check correspondence
Formal
STANDARDprovable from stated rules or axioms
BLIND SPOTsilent on whether the rules fit the world
Empirical
STANDARDit replicates or predicts
BLIND SPOTmisses the unmeasured; can optimize a proxy
Pragmatic
STANDARDit works under real conditions
BLIND SPOT“works now” can hide deferred failure
Historical
STANDARDconvergence of independent witnesses; best explanation of surviving evidence, checked against non-testimonial residue (charters, coins, strata)
BLIND SPOTunderdetermined by evidence; survivorship bias in what remains
CONTESTEDno single right answer: judge defensibility
Aesthetic
STANDARDdefensibility of the judgment; consensus does not settle it
BLIND SPOTno determinate answer
Ethical
STANDARDreasoned justification
BLIND SPOTno determinate answer; irreducibly contested
APPLIED ACROSSrun within, or on top of, any method; not branches
Adversarial PROCEDURE
 attack the claim within any method; keep what survives
Calibration META-CHECK
 on the evaluator, not the claim: does stated confidence match observed accuracy
Execution axes. Any method can be run by a human or a machine, quantitatively or qualitatively, individually or collectively. These are ways to run a method, not methods.
Scope. A general taxonomy of how a claim is judged, not a map of the current AI-evaluation field, which addresses the narrower problem of testing machine outputs and is fragmenting rather than converging.
Boundaries. The methods blend at the edges: historical explanation runs on empirical inference, ethical claims carry empirical premises, and judging pragmatic success turns on a value-laden definition of “works.” The taxonomy is a selection heuristic (§5.3), not a partition.
Determinate: verify against a standard Contested: assess a judgment Cross-cutting: applies across both

The historical method was missing from the first version of this spec. Claims about the Anarchy, the war the companion essay descends from, are determinate (the war happened some way) but cannot be replicated, derived, or load-tested; the check historians developed is inference to the best explanation30 from convergent independent witnesses, disciplined by source criticism.23 It is the framework’s oldest running instance.

§7 Implications for education

The trained outcome is a person who can:

This does not replace conclusions with content-free “skills.” Evaluation is substantially domain-bound12; vetted conclusions change role rather than disappear, taught through evaluating claims within each domain. The untested kernel, after prior evidence is subtracted (§10), is whether competence-boundary recognition is trainable at population scale.14,31 Training sets the boundary; institutions make it usable by routing what lies past one person’s edge to evaluators inside theirs (F2).

§8 Failure conditions

Correlated vs. independent evaluators§8 · F1
Correlated · a mirror
→ → → → → →
Same source, same bias: one error, copied. Agreement feels like verification, but it is one witness wearing a million masks.
Independent · a check
↗ ↓ ← ↑ ↘ →
Errors point every way and cancel; the signal survives. The whole is more reliable than any node: the jury theorem, the random forest.
The fragile requirement: node independence is the property a few dominant machines erode by default. Rising error-correlation is apparatus failure even when accuracy still looks stable.

§9 Falsification

The hypothesis makes the following predictions; each is checkable, and it fails if one proves false.

The two-sided prediction§9 · P1
Distribution wins where independent errors can cancel:
  • forecast-like empirical claims
  • contested judgments
  • historical reconstruction
Lone expertise wins on rare-skill determinate tasks:
  • formal proofs
  • one radiologist reading an MRI
Either half failing falsifies. A frame that only ever wins is not a prediction; it is a story. This one names, in advance, the cases it should lose.

Status These predictions are checkable in principle, not yet in procedure. A real test requires named datasets, sampling procedures, comparison conditions, aggregation rules, numerical independence thresholds, minimum effect sizes, and time horizons; none are fixed here, and P3 says “pre-registered” where the numbers should stand. The forecast leg has a ready instrument in resolved forecasting-tournament question sets;29 the other legs do not. This section is falsifiability-conscious, not a pre-registration. Stating the debt is cheaper than pretending it is paid.

§10 Prior art & prior evidence

A specification claiming testability must say which of its predictions others already tested. What remains novel is the composite, the taxonomy’s scoping, and P2’s kernel.

Component Nearest prior work Status
Distributed error-correction over authority Popper (1945); Longino (1990); Rauch (2021)11,18,19 inherited frame
Independence condition & when crowds beat experts Condorcet (1785); Breiman (2001); Kleinberg & Raghavan (2021)24,21,22 formalized elsewhere; adopted
Benefit of cognitive diversity Hong & Page (2004; generality contested, Thompson 2014); Landemore (2013); Mercier & Sperber (2017)17,16 suggestive, contested
Teachability of evaluation components lateral reading (Wineburg et al. 2022); inoculation (Roozenbeek & van der Linden 2019); calibration training (Mellers et al. 2014)27,28,29 partially confirmed pre-spec
Domain-boundedness constraint Willingham (2007)12 constraint adopted
Deference criteria (F7 patches) Anderson (2011); Collins & Evans (2007)15 partial patches only
Building commons against capture (F3) Ostrom (1990)25 nearest playbook; unproven at this scale
Historical method as running instance source criticism (Bloch 1949); IBE (Lipton 1991)23,30 oldest confirmed instance
Method-selection as master skill; boundary-recognition kernel; two-sided P1; pairing form none staked here

Status: a synthesis, not a discovery. Component claims carry the prior evidence cited above; the composite frame, the taxonomy’s scoping of P1, and P2’s kernel remain unverified; the protocols that would test §9 remain unwritten.

The Companion The Chronicler’s Problem the argument behind this spec. Read the essay →
Sources References numbering is shared with the companion essay, so a citation means the same thing in both documents.
  1. George R. R. Martin, Fire & Blood (Bantam, 2018). Martin has acknowledged the Dance of the Dragons drew in part on the Anarchy; the book's frame (Archmaester Gyldayn reconciling the irreconcilable testimonies of Mushroom, Septon Eustace, and Grand Maester Munkun) reproduces the source condition of the real war.
  2. Gesta Stephani, ed. K. R. Potter, rev. R. H. C. Davis (Oxford Medieval Texts, 1976); William of Malmesbury, Historia Novella, ed. Edmund King, trans. K. R. Potter (Oxford Medieval Texts, 1998). The pro-Stephen and Angevin-leaning witnesses, respectively.
  3. The Peterborough Chronicle (Anglo-Saxon Chronicle, MS E), annal for 1137.
  4. Edmund King, King Stephen (Yale University Press, 2010); David Crouch, The Reign of King Stephen, 1135–1154 (Longman, 2000). The revisionist case that the disorder was severe but regional, argued substantially from non-narrative evidence: charters, writs, and coinage struck through the war years.
  5. Onora O'Neill, A Question of Trust (Cambridge University Press, 2002). The BBC Reith Lectures arguing that transparency has been mistaken for, and cannot substitute for, accountability.
  6. Coalition for Content Provenance and Authenticity (C2PA), Content Credentials technical specification, c2pa.org. The principal industry effort to rebuild provenance cryptographically.
  7. William MacAskill, What We Owe the Future (Basic Books, 2022), on value lock-in: the risk that systems built at scale entrench one era's judgments.
  8. R. P. C. Hanson, The Search for the Christian Doctrine of God: The Arian Controversy, 318–381 (T&T Clark, 1988). The standard account of the fifty-plus years of contest between Nicaea (325) and Constantinople (381).
  9. Kurt Gödel, "Über formal unentscheidbare Sätze der Principia Mathematica und verwandter Systeme I," Monatshefte für Mathematik und Physik 38 (1931): the incompleteness theorems. Used here as illustration of the shape of a limit, not as proof about societies.
  10. Plato, Republic, Book VI (the ship of state, 488a–489d): the navigator who ought simply to be given the helm.
  11. Karl Popper, The Open Society and Its Enemies (Routledge, 1945), esp. vol. 1, ch. 7: replacing "who should rule?" with the question of how bad rulers can be removed without bloodshed.
  12. Daniel T. Willingham, "Critical Thinking: Why Is It So Hard to Teach?" American Educator 31, no. 2 (Summer 2007): the case that critical thinking is substantially domain-bound.
  13. Walter Lippmann, Public Opinion (1922) and The Phantom Public (1925); John Dewey, The Public and Its Problems (1927). The original debate over whether a general public can be competent to judge.
  14. Justin Kruger and David Dunning, "Unskilled and Unaware of It," Journal of Personality and Social Psychology 77, no. 6 (1999): miscalibration concentrated precisely where competence is lowest.
  15. Elizabeth Anderson, "Democracy, Public Policy, and Lay Assessments of Scientific Testimony," Episteme 8, no. 2 (2011): second-order criteria (track record, conflicts, responsiveness to criticism) by which laypeople can judge experts; Harry Collins and Robert Evans, Rethinking Expertise (University of Chicago Press, 2007): meta-expertise.
  16. Hugo Mercier and Dan Sperber, The Enigma of Reason (Harvard University Press, 2017): reasoning as an evolved social capacity, biased in solitary use, effective in argumentative exchange.
  17. Lu Hong and Scott E. Page, "Groups of diverse problem solvers can outperform groups of high-ability problem solvers," PNAS 101, no. 46 (2004); for the contested generality of the result, Abigail Thompson, "Does Diversity Trump Ability?" Notices of the AMS 61, no. 9 (2014); Hélène Landemore, Democratic Reason (Princeton University Press, 2013).
  18. Helen Longino, Science as Social Knowledge (Princeton University Press, 1990). Objectivity as a property of a social practice meeting four conditions: recognized avenues for criticism, uptake of criticism, public standards, and tempered equality of intellectual authority.
  19. Jonathan Rauch, The Constitution of Knowledge: A Defense of Truth (Brookings Institution Press, 2021).
  20. Elizabeth L. Eisenstein, The Printing Press as an Agent of Change (Cambridge University Press, 1979), including print's double edge: pamphlet wars and witch manuals alongside the republic of letters.
  21. Leo Breiman, "Random Forests," Machine Learning 45, no. 1 (2001): ensemble generalization error bounded in terms of the strength of individual members and the correlation between them.
  22. Jon Kleinberg and Manish Raghavan, "Algorithmic monoculture and social welfare," PNAS 118, no. 22 (2021): many decision-makers adopting the same superior algorithm can lower aggregate outcomes.
  23. Marc Bloch, The Historian's Craft (posthumous, 1949; written before his execution by the Gestapo in 1944). The critical method: cross-examining witnesses who cannot be recalled.
  24. Marquis de Condorcet, Essai sur l'application de l'analyse à la probabilité des décisions rendues à la pluralité des voix (1785). The jury theorem: with independent, better-than-chance judges, group accuracy rises with size; below chance, it falls.
  25. Elinor Ostrom, Governing the Commons (Cambridge University Press, 1990): empirical design principles for governing shared resources against extractive incentive.
  26. Neil Postman and Charles Weingartner, Teaching as a Subversive Activity (Delacorte, 1969), ch. 1, "Crap Detecting": education's job as building the detector, not delivering the conclusions.
  27. Sam Wineburg, Joel Breakstone, Sarah McGrew, Mark D. Smith, and Teresa Ortega, "Lateral reading on the open Internet: A district-wide field study in high school government classes," Journal of Educational Psychology 114, no. 5 (2022): classroom-taught source evaluation measurably improved.
  28. Jon Roozenbeek and Sander van der Linden, "Fake news game confers psychological resistance against online misinformation," Palgrave Communications 5 (2019); accessibly synthesized, with decay caveats, in van der Linden, Foolproof (W. W. Norton, 2023).
  29. Barbara Mellers et al., "Psychological strategies for winning a geopolitical forecasting tournament," Psychological Science 25, no. 5 (2014); Philip E. Tetlock and Dan Gardner, Superforecasting (Crown, 2015): brief calibration training measurably improved forecast accuracy.
  30. Peter Lipton, Inference to the Best Explanation (Routledge, 1991; 2nd ed. 2004).
  31. Mark R. Leary et al., "Cognitive and Interpersonal Features of Intellectual Humility," Personality and Social Psychology Bulletin 43, no. 6 (2017): the young measurement literature on intellectual humility, the trait P2's kernel proposes to train.
  32. John Wihbey, "AI and Epistemic Risk for Democracy: A Coming Crisis of Public Knowledge?" (SSRN working paper, 2024); and "In Post-Authenticity AI Age, Knowledge Institutions Matter More than Ever," Tech Policy Press (November 14, 2025): the nearest recent argument that provenance collapse makes knowledge institutions more, not less, decisive. Institution-centric where this piece is education-centric.
  33. "A robot wrote this entire article. Are you scared yet, human?" The Guardian (September 8, 2020): an op-ed generated by GPT-3, prompted by the paper and assembled by its editors from eight outputs (run by Liam Porr); an early instance of AI authorship deployed as a rhetorical reveal, the device the essay's confession both uses and disowns.
  34. Ethan Mollick, "Post-apocalyptic education," One Useful Thing (August 30, 2024): redesigning education around AI, away from transmitting conclusions and toward judgment.
  35. Archon Fung, Mary Graham, and David Weil, Full Disclosure: The Perils and Promise of Transparency (Cambridge University Press, 2007); Mike Ananny and Kate Crawford, "Seeing without knowing: Limitations of the transparency ideal and its application to algorithmic accountability," New Media & Society 20, no. 3 (2018): disclosure scholarship distinguishing information disclosed from information usable and actionable.
  36. Stefan Wojcik et al., "Birdwatch: Crowd Wisdom and Bridging Algorithms can Inform Understanding and Reduce the Spread of Misinformation," arXiv:2210.15723 (2022): the design of X's Community Notes (formerly Birdwatch), a bridging-based ranking that surfaces notes rated helpful by contributors who normally disagree; a deployed aggregation rule for contested claims.
  37. Philip Kitcher, "The Division of Cognitive Labor," The Journal of Philosophy 87, no. 1 (1990): how a community should distribute investigative effort across rival approaches; what the community learns depends on the allocation, not on each member's rationality alone.

Provenance of this document itself. This specification is the falsifiable half of a pair; the essay half is The Chronicler’s Problem, and the ledger of what the pair inherits and stakes lives there. The spec was drafted from the essay's argument, then revised under external review by differently instructed AIs. Every editorial call was human; added citations were verified before inclusion. If that chain changes your assessment, the pair is about why.

← Back to Just in Time © 2026 Justin Gregoire · jtgregoire.com