The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

AI benchmarks are not neutral yardsticks but power tools: weak validation, biased datasets and opportunistic gaming concentrate prestige and resources in wealthy labs, entrenching inequities and steering research away from socially beneficial directions.

The Benchmark Trap: Structures of Power and Injustice in AI Evaluations
Jason Branford, Angelie Kraft · August 15, 2026
arxiv theoretical n/a evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Jason Branford unresolved corpus identity
  2. Angelie Kraft unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Jason Branford provider ID
  2. Angelie Kraft provider ID
The paper argues that AI benchmarks are socio-technical artifacts that concentrate prestige and resources in well-funded labs, reproduce biases and mismeasurements, and thereby produce structural injustices that skew research priorities and harm marginalized stakeholders.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Artificial intelligence (AI) benchmarks are not neutral tools of evaluation but socio-technical artefacts that shape competition, power, and research priorities within AI. Benchmarks standardise the assessment of systems and facilitate the creation of leaderboards that reward state-of-the-art performance with prestige, citations, trust, and institutional influence. As the costs of developing competitive AI systems rise, these rewards increasingly concentrate among powerful, industry-funded labs. This paper situates these concerns within Iris Marion Young's theories of oppression and structural injustice. It argues that current benchmarking practices may perpetuate systematic harms affecting various actors in AI research, aligning with four of Young's "faces of oppression". Benchmarking culture is further framed as a source of structural injustice, as these harms emerge from normalised, individually defensible practices and network effects, even without explicit wrongdoing. By reinforcing existing power structures and narrowing possible research trajectories, benchmarking may in fact prevent the field from advancing in epistemically robust and socially beneficial ways.

Summary

Main Finding

AI benchmarks are socio-technical artefacts that do much more than measure model performance: they shape research agendas, concentrate prestige and resources, and can reproduce structural injustices. Because benchmarks become de facto standards (leaderboards, regulatory signals, media narratives), their conceptual and methodological flaws — lack of construct validity, dataset bias, naturalisation of "ground truths", and opportunities for gaming — translate into political‑economic harms (concentrated rents, exclusion of smaller labs and marginalized communities, labor exploitation, and epistemic narrowing). Framing these dynamics through Iris Marion Young’s theory of oppression, the authors show benchmarking culture can instantiate multiple “faces” of oppression and amount to avoidable structural injustice that ultimately harms the epistemic and social potential of AI.

Key Points

  • Benchmarks are not neutral measurement tools but socio-technical artefacts whose scores travel beyond their original contexts and acquire material force (leaderboards, procurement, regulation, media).
  • Principal methodological problems:
    • Lack of construct validity: many benchmarks do not define the concept they claim to measure or provide validation (systematic review: 445 LLM benchmarks; >50% had contested/no definitions; ~half lacked construct validation).
    • Dataset bias & misrepresentation: benchmark datasets often reflect Western, elite priorities and overrepresent specific demographics/domains, rewarding models tuned to those distributions and masking weaknesses elsewhere.
    • Naturalisation of “ground truths”: once adopted, benchmark labels/metrics (e.g., toxicity scoring via Perspective API) become treated as objective standards despite conceptual/implementation limits.
    • Gaming & contamination: incentives to top leaderboards drive strategic behavior (e.g., selective training, secret partnerships, leaked agreements) and data contamination that invalidates evaluations.
  • Stakeholders affected: researchers (esp. small or marginalized labs), annotators/data workers (precarious labor, extraction), industry actors (who benefit), policymakers/journalists (who rely on scores), and end‑users (disproportionately marginalized groups).
  • Political‑economic consequences:
    • Concentration of rewards (citations, funding, hiring pipelines) in well-resourced, industry‑backed labs — raising entry costs and producing winner-take-most dynamics.
    • Narrowing of research trajectories toward tasks/metrics that maximize leaderboard performance rather than social value or epistemic robustness.
    • Market and regulatory capture risks (benchmarks shape what regulators and purchasers treat as “safe” or “state‑of‑the‑art”).
  • Normative framing: Using Young’s faces of oppression, the paper argues benchmarks instantiate exploitation (capture of labor/benefits), marginalisation (exclusion of communities and research agendas), powerlessness (participation without authority over standards), and cultural imperialism (Western-centric definitions normalized). These are presented as structural injustices — not merely isolated bad acts — and, following McKeown, largely avoidable.

Data & Methods

  • Methodological approach: conceptual and normative analysis grounded in political philosophy (Iris Marion Young) combined with a synthesis of empirical and documentary evidence from recent AI literature and public cases.
  • Evidence types used in the paper:
    • Systematic review results (cited: Bean et al. 2025) of 445 LLM benchmarks assessing construct validity reporting.
    • Empirical dataset audits/analyses (e.g., Kraft, Simon, and Schimmler 2025 — 20 QA benchmarks showing Western/male overrepresentation).
    • Case vignettes and investigative reports (examples: Meta/Llama 4 admission of “fudged” results; FrontierMath/OpenAI–EpochAI partnership and secrecy; RealToxicityPrompts reliance on Perspective API).
    • Cited studies on contamination, benchmark overfitting, and marker effects (Sainz et al. 2023; Schaeffer et al. 2025; Uzunoglu et al. 2025; others).
    • Regulatory context references (EU AI Act citing benchmarking requirements).
  • Analysis method: applies Young’s five‑face framework to benchmarking practices and argues the harms are systemic; supplements normative claims with cited empirical patterns and concrete industry examples to support plausibility.

Implications for AI Economics

  • Incentives and concentration:
    • Benchmarks act as focal points that direct private and public R&D investment toward leaderboard-driven objectives, reinforcing scale advantages for well‑funded labs and creating winner‑take‑most market structures.
    • Prestige, citations, and ease of regulatory/market acceptance (via benchmark scores) act as quasi-rents that accrue disproportionately to incumbents.
  • Labor and externalities:
    • Creation and maintenance of benchmarks rely on often-precarious data labor (annotation, prompt/test case authorship); costs/social harms are externalised while benefits concentrate.
    • Intellectual property and scraping issues around benchmark datasets create legal and distributional externalities that affect smaller actors disproportionately.
  • Market design and regulation:
    • Overreliance on benchmark scores by purchasers and regulators risks locking in narrow standards that favor particular architectures/data regimes — raising the economic cost of compliance for diverse entrants and inhibiting innovation breadth.
    • Benchmark manipulation and secrecy undermine market transparency, complicating competition policy and procurement decisions.
  • Epistemic and allocative efficiency:
    • Steering research to beat benchmarks can produce inefficiencies: overinvestment in marginal leaderboard gains, underinvestment in societally valuable tasks that are hard to formalize as benchmarks (safety, fairness in diverse contexts, real-world robustness).
    • This misallocation can slow socially beneficial innovation and may increase systemic risk (if models are optimized to scored metrics rather than real-world safety).
  • Policy and governance measures (implied by the paper and feasible from an economics perspective):
    • Decouple regulatory and procurement decisions from single benchmark scores; require multi-dimensional, validated evidence of performance.
    • Fund public, independent benchmark creation and validation (to lower entry costs and reduce capture), and support replication/holdout test sets hosted by neutral parties.
    • Mandate transparency on benchmark provenance, funding, partnerships, and contamination risk; require disclosure of whether benchmarks were used in model training.
    • Require construct validation and documented purpose for any benchmark used in regulatory or procurement contexts.
    • Support open, pluralistic evaluation ecosystems (diverse benchmarks, domain‑specific tests, adversarial and deployment‑focused evaluations) to reduce monoculture risk.
    • Consider competition policy scrutiny when companies use proprietary benchmarking practices to erect entry barriers (e.g., exclusive datasets, secret contract evaluations).
  • Takeaway for economists and policymakers: benchmarks are not just technical instruments but economic institutions shaping incentives, rents, labor markets, and innovation pathways. Addressing benchmark‑driven injustices requires interventions in market design, public goods provision (open benchmark infrastructure), transparency rules, and regulatory practice to align incentives with socially beneficial, epistemically robust AI development.

Assessment

Paper Typetheoretical Evidence Strengthn/a — The paper synthesizes prior empirical critiques, documented industry incidents, and theoretical resources to make a normative argument rather than testing hypotheses; it therefore provides reasoned plausibility rather than empirical causal proof. Methods Rigormedium — Rigor stems from clear conceptual framing, engagement with prior literature (systematic review citations and documented case reports), and a structured application of Young's theory; it lacks pre-registered analysis, systematic case selection, or original empirical measurement that would raise rigor to 'high'. SampleA conceptual literature-based analysis using prior studies, systematic reviews, and documented industry cases (e.g., critiques of LLM benchmarks, Bean et al. 2025 systematic review of LLM benchmarks, journalistic reports of Llama 4 and FrontierMath, analyses of benchmark datasets and annotation labor). No original dataset or statistical sample is used. Themesgovernance innovation inequality org_design GeneralizabilityArguments are normative/conceptual and not empirically validated, limiting empirical generalizability., Illustrative examples are skewed toward high-profile LLMs and Western, well-resourced labs, so applicability to other AI subfields or geographic contexts may be limited., Temporal sensitivity: examples and dynamics reflect contemporary (mid-2020s) benchmarking culture and may evolve with new evaluation practices or regulation., Does not quantify magnitude of harms or economic effects, limiting translation into policy impact estimates.

Claims (13)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Current AI benchmarking culture may reinforce existing power structures by concentrating prestige, citations, trust, and institutional influence among powerful, industry-funded AI labs. Inequality negative Distribution of prestige, influence, and competitive advantage in AI research
Reading fidelity high
Study strength speculative
not reported
0.02
A systematic review found that more than half of 445 large language model benchmarks were designed using contested concept definitions or no definition at all. Decision Quality negative Presence and quality of construct definitions in LLM benchmarks
Reading fidelity high
Study strength high
n=445
more than half
0.2
In the same review of 445 LLM benchmarks, a bit less than half were published without reported construct validation through a justification rationale or empirical evidence. Decision Quality negative Reporting of construct-validation evidence for LLM benchmarks
Reading fidelity high
Study strength high
n=445
a bit less than half
0.2
Several question-answering benchmark datasets overrepresent questions about male individuals and Western locations. Inequality negative Demographic and geographic representation in benchmark questions
Reading fidelity high
Study strength medium
n=20
0.12
Popular AI benchmarks were predominantly created by Western, elite institutions. Inequality negative Institutional and geographic concentration of benchmark creation
Reading fidelity high
Study strength medium
predominantly
0.12
Benchmark datasets that are biased toward particular groups or regions can reward biased model outputs and contribute to algorithmic bias in AI systems. Ai Safety And Ethics negative Bias in model outputs and downstream AI systems
Reading fidelity high
Study strength medium
not reported
0.12
Benchmark incentives structure the AI community's innovation efforts by motivating researchers to design systems that compete well on established benchmarks. Innovation Output mixed Direction and diversity of AI research and innovation activity
Reading fidelity high
Study strength speculative
not reported
0.02
The paper argues that many highly visible benchmarks mislead the public into believing that AI systems are becoming increasingly human-like. Consumer Welfare negative Public understanding of AI capabilities
Reading fidelity high
Study strength speculative
not reported
0.02
Yann LeCun stated that the results for Meta's Llama 4 evaluations were 'fudged a little bit' and that different models were used for different benchmarks to obtain better results. Ai Safety And Ethics negative Integrity and validity of reported benchmark results
Reading fidelity high
Study strength medium
not reported
0.12
OpenAI funded the creation of the FrontierMath benchmark, and OpenAI and EpochAI agreed to keep their partnership secret until the release of OpenAI's o3 model. Governance And Regulation negative Transparency and independence of AI benchmark evaluation
Reading fidelity high
Study strength medium
not reported
0.12
Tailoring development approaches and training examples to the SWE-Bench benchmark can produce effects similar to memorisation and undermine the validity of the evaluation. Decision Quality negative Validity and generalisability of benchmark-based performance scores
Reading fidelity high
Study strength medium
not reported
0.12
The paper argues that benchmarking practices can produce structural injustices even when individual researchers and developers act without explicit wrongdoing, because harms arise from normalised practices and network effects. Inequality negative Distribution of participation, authority, and benefits across AI research stakeholders
Reading fidelity high
Study strength speculative
not reported
0.02
The creation of benchmark datasets often relies on human labour performed under precarious conditions, including low and unstable pay, gig work, isolation, and physical or psychological distress. Worker Satisfaction negative Working conditions and welfare of data authors and annotators
Reading fidelity high
Study strength medium
not reported
0.12

Notes