0 cumulative citations
View corpus contextAI benchmarks are not neutral yardsticks but power tools: weak validation, biased datasets and opportunistic gaming concentrate prestige and resources in wealthy labs, entrenching inequities and steering research away from socially beneficial directions.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Artificial intelligence (AI) benchmarks are not neutral tools of evaluation but socio-technical artefacts that shape competition, power, and research priorities within AI. Benchmarks standardise the assessment of systems and facilitate the creation of leaderboards that reward state-of-the-art performance with prestige, citations, trust, and institutional influence. As the costs of developing competitive AI systems rise, these rewards increasingly concentrate among powerful, industry-funded labs. This paper situates these concerns within Iris Marion Young's theories of oppression and structural injustice. It argues that current benchmarking practices may perpetuate systematic harms affecting various actors in AI research, aligning with four of Young's "faces of oppression". Benchmarking culture is further framed as a source of structural injustice, as these harms emerge from normalised, individually defensible practices and network effects, even without explicit wrongdoing. By reinforcing existing power structures and narrowing possible research trajectories, benchmarking may in fact prevent the field from advancing in epistemically robust and socially beneficial ways.
Summary
Main Finding
AI benchmarks are socio-technical artefacts that do much more than measure model performance: they shape research agendas, concentrate prestige and resources, and can reproduce structural injustices. Because benchmarks become de facto standards (leaderboards, regulatory signals, media narratives), their conceptual and methodological flaws — lack of construct validity, dataset bias, naturalisation of "ground truths", and opportunities for gaming — translate into political‑economic harms (concentrated rents, exclusion of smaller labs and marginalized communities, labor exploitation, and epistemic narrowing). Framing these dynamics through Iris Marion Young’s theory of oppression, the authors show benchmarking culture can instantiate multiple “faces” of oppression and amount to avoidable structural injustice that ultimately harms the epistemic and social potential of AI.
Key Points
- Benchmarks are not neutral measurement tools but socio-technical artefacts whose scores travel beyond their original contexts and acquire material force (leaderboards, procurement, regulation, media).
- Principal methodological problems:
- Lack of construct validity: many benchmarks do not define the concept they claim to measure or provide validation (systematic review: 445 LLM benchmarks; >50% had contested/no definitions; ~half lacked construct validation).
- Dataset bias & misrepresentation: benchmark datasets often reflect Western, elite priorities and overrepresent specific demographics/domains, rewarding models tuned to those distributions and masking weaknesses elsewhere.
- Naturalisation of “ground truths”: once adopted, benchmark labels/metrics (e.g., toxicity scoring via Perspective API) become treated as objective standards despite conceptual/implementation limits.
- Gaming & contamination: incentives to top leaderboards drive strategic behavior (e.g., selective training, secret partnerships, leaked agreements) and data contamination that invalidates evaluations.
- Stakeholders affected: researchers (esp. small or marginalized labs), annotators/data workers (precarious labor, extraction), industry actors (who benefit), policymakers/journalists (who rely on scores), and end‑users (disproportionately marginalized groups).
- Political‑economic consequences:
- Concentration of rewards (citations, funding, hiring pipelines) in well-resourced, industry‑backed labs — raising entry costs and producing winner-take-most dynamics.
- Narrowing of research trajectories toward tasks/metrics that maximize leaderboard performance rather than social value or epistemic robustness.
- Market and regulatory capture risks (benchmarks shape what regulators and purchasers treat as “safe” or “state‑of‑the‑art”).
- Normative framing: Using Young’s faces of oppression, the paper argues benchmarks instantiate exploitation (capture of labor/benefits), marginalisation (exclusion of communities and research agendas), powerlessness (participation without authority over standards), and cultural imperialism (Western-centric definitions normalized). These are presented as structural injustices — not merely isolated bad acts — and, following McKeown, largely avoidable.
Data & Methods
- Methodological approach: conceptual and normative analysis grounded in political philosophy (Iris Marion Young) combined with a synthesis of empirical and documentary evidence from recent AI literature and public cases.
- Evidence types used in the paper:
- Systematic review results (cited: Bean et al. 2025) of 445 LLM benchmarks assessing construct validity reporting.
- Empirical dataset audits/analyses (e.g., Kraft, Simon, and Schimmler 2025 — 20 QA benchmarks showing Western/male overrepresentation).
- Case vignettes and investigative reports (examples: Meta/Llama 4 admission of “fudged” results; FrontierMath/OpenAI–EpochAI partnership and secrecy; RealToxicityPrompts reliance on Perspective API).
- Cited studies on contamination, benchmark overfitting, and marker effects (Sainz et al. 2023; Schaeffer et al. 2025; Uzunoglu et al. 2025; others).
- Regulatory context references (EU AI Act citing benchmarking requirements).
- Analysis method: applies Young’s five‑face framework to benchmarking practices and argues the harms are systemic; supplements normative claims with cited empirical patterns and concrete industry examples to support plausibility.
Implications for AI Economics
- Incentives and concentration:
- Benchmarks act as focal points that direct private and public R&D investment toward leaderboard-driven objectives, reinforcing scale advantages for well‑funded labs and creating winner‑take‑most market structures.
- Prestige, citations, and ease of regulatory/market acceptance (via benchmark scores) act as quasi-rents that accrue disproportionately to incumbents.
- Labor and externalities:
- Creation and maintenance of benchmarks rely on often-precarious data labor (annotation, prompt/test case authorship); costs/social harms are externalised while benefits concentrate.
- Intellectual property and scraping issues around benchmark datasets create legal and distributional externalities that affect smaller actors disproportionately.
- Market design and regulation:
- Overreliance on benchmark scores by purchasers and regulators risks locking in narrow standards that favor particular architectures/data regimes — raising the economic cost of compliance for diverse entrants and inhibiting innovation breadth.
- Benchmark manipulation and secrecy undermine market transparency, complicating competition policy and procurement decisions.
- Epistemic and allocative efficiency:
- Steering research to beat benchmarks can produce inefficiencies: overinvestment in marginal leaderboard gains, underinvestment in societally valuable tasks that are hard to formalize as benchmarks (safety, fairness in diverse contexts, real-world robustness).
- This misallocation can slow socially beneficial innovation and may increase systemic risk (if models are optimized to scored metrics rather than real-world safety).
- Policy and governance measures (implied by the paper and feasible from an economics perspective):
- Decouple regulatory and procurement decisions from single benchmark scores; require multi-dimensional, validated evidence of performance.
- Fund public, independent benchmark creation and validation (to lower entry costs and reduce capture), and support replication/holdout test sets hosted by neutral parties.
- Mandate transparency on benchmark provenance, funding, partnerships, and contamination risk; require disclosure of whether benchmarks were used in model training.
- Require construct validation and documented purpose for any benchmark used in regulatory or procurement contexts.
- Support open, pluralistic evaluation ecosystems (diverse benchmarks, domain‑specific tests, adversarial and deployment‑focused evaluations) to reduce monoculture risk.
- Consider competition policy scrutiny when companies use proprietary benchmarking practices to erect entry barriers (e.g., exclusive datasets, secret contract evaluations).
- Takeaway for economists and policymakers: benchmarks are not just technical instruments but economic institutions shaping incentives, rents, labor markets, and innovation pathways. Addressing benchmark‑driven injustices requires interventions in market design, public goods provision (open benchmark infrastructure), transparency rules, and regulatory practice to align incentives with socially beneficial, epistemically robust AI development.
Assessment
Claims (13)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Current AI benchmarking culture may reinforce existing power structures by concentrating prestige, citations, trust, and institutional influence among powerful, industry-funded AI labs. Inequality | negative | Distribution of prestige, influence, and competitive advantage in AI research |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| A systematic review found that more than half of 445 large language model benchmarks were designed using contested concept definitions or no definition at all. Decision Quality | negative | Presence and quality of construct definitions in LLM benchmarks |
Reading fidelity
high
Study strength
high
|
n=445
more than half
|
| In the same review of 445 LLM benchmarks, a bit less than half were published without reported construct validation through a justification rationale or empirical evidence. Decision Quality | negative | Reporting of construct-validation evidence for LLM benchmarks |
Reading fidelity
high
Study strength
high
|
n=445
a bit less than half
|
| Several question-answering benchmark datasets overrepresent questions about male individuals and Western locations. Inequality | negative | Demographic and geographic representation in benchmark questions |
Reading fidelity
high
Study strength
medium
|
n=20
|
| Popular AI benchmarks were predominantly created by Western, elite institutions. Inequality | negative | Institutional and geographic concentration of benchmark creation |
Reading fidelity
high
Study strength
medium
|
predominantly
|
| Benchmark datasets that are biased toward particular groups or regions can reward biased model outputs and contribute to algorithmic bias in AI systems. Ai Safety And Ethics | negative | Bias in model outputs and downstream AI systems |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Benchmark incentives structure the AI community's innovation efforts by motivating researchers to design systems that compete well on established benchmarks. Innovation Output | mixed | Direction and diversity of AI research and innovation activity |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| The paper argues that many highly visible benchmarks mislead the public into believing that AI systems are becoming increasingly human-like. Consumer Welfare | negative | Public understanding of AI capabilities |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| Yann LeCun stated that the results for Meta's Llama 4 evaluations were 'fudged a little bit' and that different models were used for different benchmarks to obtain better results. Ai Safety And Ethics | negative | Integrity and validity of reported benchmark results |
Reading fidelity
high
Study strength
medium
|
not reported
|
| OpenAI funded the creation of the FrontierMath benchmark, and OpenAI and EpochAI agreed to keep their partnership secret until the release of OpenAI's o3 model. Governance And Regulation | negative | Transparency and independence of AI benchmark evaluation |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Tailoring development approaches and training examples to the SWE-Bench benchmark can produce effects similar to memorisation and undermine the validity of the evaluation. Decision Quality | negative | Validity and generalisability of benchmark-based performance scores |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The paper argues that benchmarking practices can produce structural injustices even when individual researchers and developers act without explicit wrongdoing, because harms arise from normalised practices and network effects. Inequality | negative | Distribution of participation, authority, and benefits across AI research stakeholders |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| The creation of benchmark datasets often relies on human labour performed under precarious conditions, including low and unstable pay, gig work, isolation, and physical or psychological distress. Worker Satisfaction | negative | Working conditions and welfare of data authors and annotators |
Reading fidelity
high
Study strength
medium
|
not reported
|