0 cumulative citations
View corpus contextHigh‑stakes benchmarks can manufacture their own apparent correctness: when leaderboards and dispositive evaluations concentrate recognition, they channel R&D and human capital into what they reward, creating path dependence and blind spots; ensuring plural, contestable evaluation routes is essential to avoid systemic fragility.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Around the end of the first millennium China was, on many measures, the most technically accomplished society in the world; sustained, mechanised growth began some six centuries later in Britain. This article does not try to explain that reversal — the Needham Question — but uses it to isolate a mechanism the economic-history literature sets aside. Evaluative institutions do not merely recognise a distribution of capability that exists independently of them; where recognition is strongly coupled across otherwise distinct reward domains, they help to produce it, by drawing developmental effort in advance toward whatever a criterion can already see and reward highly. Capability that departs from the dominant criterion is then under-supplied; and, as the criterion reshapes the population against which its own validity is later judged, it comes to manufacture the evidence of its own correctness. This endogenous self-sealing — not the gaming that Goodhart’s law describes, but a criterion’s growing control over the evidence that might defeat it — is the article’s central proposition. The risk arises as cross-domain coupling and dispositive authority redirect developmental investment; where independent routes of correction are weak, the resulting capability distribution increasingly validates the criterion that produced it, eroding its answerability to contrary evidence. China’s imperial examination is the historical case and Europe an existence proof that misrecognition can be made non-final; contemporary artificial-intelligence benchmarks and safety evaluations are examined for evidence of the same structure, the danger lying less in the exclusion of rival evaluation than in the concentration of consequential recognition weight in a few gates. Innovation, on this account, requires not merely good standards but institutions under which standards can lose.
Summary
Main Finding
Evaluative institutions (tests, benchmarks, examinations) do not merely reveal an independent distribution of capabilities — when recognition and rewards across different domains are tightly coupled and concentrated in a few decisive gates, these institutions actively produce that distribution. By channeling attention and investment toward what they already reward, evaluative criteria can under-supply alternative capabilities and progressively "self‑seal": they reshape the population of efforts and evidence so that the criterion appears increasingly validated. This is distinct from simple gaming (Goodhart’s law); it is an endogenous construction of the evidence base that makes a criterion hard to refute. Historical (China’s imperial examination versus Europe) and contemporary (AI benchmarks and safety evaluations) examples illustrate the mechanism. The policy implication is that innovation requires not only good standards but institutions in which those standards can plausibly lose.
Key Points
- Evaluative institutions are formative, not just diagnostic: reward structures steer where developmental effort goes.
- Cross-domain coupling (when a single evaluation influences rewards in many domains) and dispositive authority (when that evaluation gate determines access, prestige, funding, or deployment) produce strong feedback that concentrates capability along the criterion.
- Endogenous self-sealing differs from gaming: it is the criterion’s increasing control over the available evidence and population that shields it from disconfirming evidence.
- Historical case: the imperial examination in China is used to show how a centralized, high‑stakes evaluative institution can redirect learning and careers and thereby alter the society’s capability distribution; Europe provides a counterexample where misrecognition was contested and overturned, showing the non-finality of such processes.
- Contemporary concern: AI benchmarks, leaderboards, and safety evaluation gates can play a similar formative role if consequential recognition weight is concentrated, potentially producing path dependence, blind spots, and fragility in the innovation ecosystem.
- Remedy is institutional: make standards contestable, maintain independent routes of correction, and avoid concentrating dispositive recognition in a few gates.
Data & Methods
- Methodological approach: conceptual/theoretical argument that isolates a mechanism (endogenous self-sealing) using comparative historical analysis and institutional interpretation rather than an econometric test.
- Historical case studies: analysis of late‑first‑millennium China and later British mechanised growth; in particular, the role and structure of the imperial examination system as an evaluative institution that shaped human capital allocation.
- Comparative evidence: Europe treated as an existence proof where different institutional trajectories allowed correction of misrecognition.
- Contemporary institutional analysis: examination of AI evaluation practices (benchmarks, leaderboards, safety assessments) for structural similarities — i.e., concentrated recognition weight, cross-domain coupling, and weak independent correction — rather than a statistical evaluation of outcomes.
- The article isolates mechanisms and plausibility pathways rather than claiming exhaustive empirical causation of macroeconomic divergence.
Implications for AI Economics
- Incentives and allocation: Concentrated, high‑stakes benchmarks/leaderboards will channel R&D investment toward optimizing for those metrics, potentially under-supplying capabilities not captured by the metric (safety, robustness, interpretability, domain‑specific competence).
- Path dependence and market structure: When recognition is dispositive, incumbents or early winners can capture rents and standardize practices that entrench particular capability profiles and business models.
- Fragility and systemic risk: Self‑sealed evaluative regimes reduce independent correction and increase the risk that blind spots (e.g., adversarial failure modes, deployment harms) go unaddressed until they become systemic.
- Policy and institutional design recommendations:
- Decentralize recognition weight: avoid a few dispositive gates; support multiple, domain‑specific, and complementary evaluations.
- Maintain independent routes of correction: fund and give weight to replication, field trials, external audits, red‑teaming, and adversarial evaluation outside the dominant benchmark regime.
- Encourage plural metrics and plural institutions: use ensembles of benchmarks (performance, safety, fairness, robustness), alternative evaluation providers, and rotating criteria to prevent lock‑in.
- Ensure contestability: create processes where standards can be challenged, refined, or replaced (open challenges, prize mechanisms, review bodies with veto power).
- Tie funding and deployment decisions to a broader set of evidentiary channels, not single tests or leaderboards.
- For economists: model the feedback from evaluative institutions into R&D choice and human-capital allocation; quantify how concentration of recognition alters rates of technical progress, specialization, and welfare; and assess optimal institutional mixes that balance coordination benefits of common standards with the need for contestability.
- Practical takeaway: Good benchmarks matter, but their social value depends on being embedded in institutional ecosystems that allow them to lose and be superseded — otherwise benchmarks can manufacture their own apparent correctness and misdirect the course of AI development.
Assessment
Claims (9)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Evaluative institutions are formative rather than merely diagnostic: their reward structures steer where developmental effort and investment are directed. Task Allocation | mixed | Allocation of developmental effort and investment across capabilities |
Reading fidelity
high
Study strength
low
|
not reported
|
| When recognition is coupled across domains and concentrated in dispositive evaluation gates, feedback effects concentrate capability development around the criteria used by those gates. Task Allocation | negative | Concentration and specialization of capabilities around dominant evaluation criteria |
Reading fidelity
high
Study strength
low
|
not reported
|
| Endogenous self-sealing differs from ordinary gaming because the evaluation criterion changes the population of efforts and evidence available to assess it, making the criterion increasingly difficult to refute. Governance And Regulation | negative | Independence and contestability of the evidence base used to validate an evaluation criterion |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| The imperial examination system in China redirected learning and careers and thereby altered the distribution of capabilities in society. Skill Acquisition | negative | Allocation of learning, careers, and human capital across capabilities |
Reading fidelity
high
Study strength
low
|
not reported
|
| Europe provides a contrasting institutional trajectory in which misrecognition could be contested and overturned, demonstrating that evaluative judgments need not be final. Governance And Regulation | positive | Institutional capacity to correct or overturn mistaken evaluations |
Reading fidelity
high
Study strength
low
|
not reported
|
| High-stakes AI benchmarks and leaderboards can channel research and development toward optimizing measured metrics while under-supplying capabilities that the metrics do not capture. Innovation Output | negative | Distribution of R&D effort across benchmarked and non-benchmarked capabilities |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| When recognition is dispositive, incumbents or early winners can capture rents and standardize practices that entrench particular capability profiles and business models. Market Structure | negative | Market entrenchment, rent capture, and persistence of particular business models |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| Self-sealed evaluation regimes reduce independent correction and increase the risk that blind spots and deployment harms remain unaddressed until they become systemic. Ai Safety And Ethics | negative | Independent error correction and risk of unresolved AI failure modes or deployment harms |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| Innovation is more robust when evaluation standards are contestable and embedded in institutions that allow them to be challenged, corrected, or superseded. Governance And Regulation | positive | Institutional contestability and capacity for corrective innovation |
Reading fidelity
high
Study strength
speculative
|
not reported
|