The EU AI Act's reliance on an undefined 'accuracy' pits legal fitness‑for‑purpose demands against ML's benchmarking culture, creating compliance uncertainty; the authors call for standardized reporting, interventional validation, and new tools to extend accuracy measurements to real-world deployments.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
No provider observation is available for this paper.
Missing data, not a zero citation count.
The machine learning community progresses (in part) by improving the "accuracy" of its systems. The EU AI Act explicitly refers to "accuracy" as part of its compliance measures for high-risk AI systems. Are we talking about the same thing? This work presents "accuracy" as a case-study for differing requirements of social worlds, the technological machine learning community and the legal community. While competition on accuracy contributes to technological development, machine learning scholars simultaneously recognize accuracy's shortcomings regarding the usefulness and effectiveness of machine learning systems. The legal counterpart embraces the vagueness of "accuracy," leaving interpretative flexibility for technological and societal changes. At the same time, accuracy is a core element of compliance within the EU AI Act. We elaborate on five main tensions, (a) nature of accuracy, (b) notion of performance, (c) scope of validity, (d) ends, and (e) statisticalness, to show that the two communities project disparate, and sometimes contradictory, expectations on accuracy. Both legal and technical communities lack precise understanding of "accuracy" beyond the contextual boundaries of their community. The resulting frictions, \eg, based on the empirical or normative understanding of accuracy, are symptoms of an unresolved (and unresolvable) debate on what accuracy is. We constructively use the frictions to recommend baselines and interventional studies in standardization, and demand for tools to extend the validity of accuracy measurements.
Summary
Main Finding
Accuracy is a boundary object whose technical (E-accuracy) and legal (L-accuracy) meanings diverge in ways that matter for competition, compliance costs, and welfare. The paper argues that neither machine‑learning communities nor legal doctrine can uniquely define “accuracy” across contexts; instead, frictions between E- and L-accuracy reveal both limits and actionable fixes. The authors distill five core tensions and give concrete recommendations (R1–R10) to reduce misalignment—calling for improved reporting, new technical tools to extend validity (especially toward individuals and deployment scenarios), and interventional studies to align measurement with intended social ends.
Key Points
- Two parallel concepts:
- E-accuracy: statistical/formal empirical measure used for benchmarking, algorithmic competition, and internal optimization (population‑level, comparative).
- L-accuracy: a teleological, normative compliance concept tied to fitness‑for‑purpose in the EU AI Act (prevention of harms, lifecycle consistency, transferability to deployment).
- Five tensions between E- and L-accuracy:
- Nature of accuracy: formal/empirical vs. normative/teleological.
- Notion of performance: accuracy as a component of performance vs. accuracy as a certification of performance.
- Scope of validity: validity needed for high-stakes deployment (transferability) vs. validity sufficient for benchmark comparisons.
- Ends: internal optimization (developer incentives) vs. external certification (regulatory aims).
- Statisticalness: population-aggregate focus vs. legal concern for individual-level harms and rights.
- Consequences of the mismatch:
- Measurement ambiguity fosters regulatory uncertainty and compliance risk.
- Benchmark-driven competition can incentivize models that optimize narrow metrics rather than social outcomes (misaligned incentives, potential "accuracy washing").
- Standardization delays (AI Act delegated details to standards bodies) leave providers to self-interpret compliance — creating heterogeneity and potential first-mover advantages/disadvantages.
- Recommendations (high level — R1–R10 summarized):
- Require richer accuracy reporting: justify metric choice, describe data, relate target constructs to intended purpose.
- Mandate baseline comparisons and context‑dependent interventional studies to show real-world performance.
- Develop technical methods to assess and report validity across deployment scenarios and toward individuals (e.g., local/ subgroup accuracy, uncertainty quantification).
- Accept and design for ongoing iterative interpretation (no final single notion of accuracy).
- Invest in standards and tooling that make accuracy measurements more transferable and legally actionable.
Data & Methods
- Conceptual, interdisciplinary analysis rather than empirical experimentation.
- Methods:
- Boundary object framing (Star & Griesemer) to explain cross-domain interpretative flexibility.
- Parallel dual-thread exposition: a legal thread (L-accuracy) using doctrinal analysis of the EU AI Act and methods of legal interpretation; a technical thread (E-accuracy) using ML epistemic practices and literature on benchmarking, evaluation, and socio-technical critiques.
- Literature review across ML, AI governance, and legal scholarship; appendices provide detailed doctrinal (Appendix L) and technical meta-reflections (Appendix E).
- Constructive policy and technical recommendations derived from the synthesized frictions.
- No original empirical datasets or econometric analysis; the paper’s claims are normative/conceptual and prescriptive.
Implications for AI Economics
- Market competition and benchmarking
- Benchmarks (E-accuracy) will continue to structure competition; but if legal compliance requires broader validity (L-accuracy), firms optimizing only benchmark metrics can face unexpected compliance costs or liability.
- Diverging accuracy definitions create rent opportunities: firms that invest earlier in L-aligned evaluation and documentation may gain market/contracting advantage.
- Investment and innovation incentives
- Uncertain regulatory requirements (pending standards) raise regulatory risk, affecting investment timing and the selection of projects (favoring products with clearer fitness-for-purpose).
- Developers may internalize higher evaluation costs (interventional studies, subgroup analyses), shifting R&D allocation from raw model improvement to evaluation, explainability, and monitoring tools.
- Compliance costs and market structure
- Small entrants face higher relative burden to produce legally‑robust accuracy evidence — potential consolidation toward larger firms that can absorb reporting and interventional study costs.
- Providers may face trade-offs: optimizing for legal-compliant accuracy vs. optimizing for marketplace performance/price — with implications for product variety and consumer surplus.
- Welfare, distributional effects, and externalities
- Aggregate accuracy metrics can mask subgroup harms; legal emphasis on fitness-for-purpose raises the need to value individual-level performance in welfare calculations.
- Misaligned incentives can generate negative externalities (e.g., deployment of systems that pass aggregate tests but harm vulnerable groups), suggesting a role for regulation to correct market failures.
- Policy and research agenda for economists
- Quantify compliance costs of richer accuracy reporting and interventional studies; model their impact on entry, competition, and innovation dynamics.
- Empirically measure how benchmark‑driven competitions affect downstream welfare when regulatory definitions of accuracy differ.
- Evaluate standard-setting timing and design as strategic complements/substitutes to firm investment in evaluation capability.
- Design instruments to estimate individual-level or subgroup-level accuracy externalities and compute socially optimal evaluation standards.
- Study contracting and certification markets where firms signal L-accuracy by third‑party audits, and assess potential for certification market failures (e.g., capture, low-quality audits).
- Practical recommendations for economic actors
- Firms: invest in evidence infrastructure (data provenance, subgroup metrics, lifecycle monitoring) early to reduce future compliance risk and gain competitive signaling.
- Regulators: prioritize clear guidance on what L-accuracy entails in deployment contexts; support standards and capacity-building to lower compliance costs for smaller firms.
- Researchers: run interventional field studies and cost-benefit analyses comparing narrow benchmark improvement vs. investments in deployment‑oriented validation.
Overall, the paper reframes “accuracy” as an economic coordination problem between markets (benchmarks and competition) and regulation (safety, rights, and validity). Economists can help quantify trade-offs, design incentives, and evaluate policies to align technical evaluation practices with social goals.
Assessment
Claims (10)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| The EU AI Act requires high-risk AI systems to achieve an appropriate level of accuracy throughout their lifecycle. Regulatory Compliance | positive | Compliance with the EU AI Act's accuracy requirement |
Reading fidelity
high
Study strength
high
|
not reported
|
| The EU AI Act does not explicitly define the term 'accuracy,' despite making it a central compliance requirement for high-risk AI systems. Governance And Regulation | negative | Precision and clarity of the legal accuracy requirement |
Reading fidelity
high
Study strength
high
|
not reported
|
| In machine learning, E-accuracy is presented as a formal and empirical measure of the relative performance of a machine-learning method compared with other methods. Output Quality | positive | Relative predictive performance of machine-learning methods |
Reading fidelity
high
Study strength
medium
|
not reported
|
| In the legal context, L-accuracy is a performance-compliance measure intended to certify that an AI system is fit for its intended purpose. Regulatory Compliance | positive | Fitness of an AI system for its intended purpose |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The legal meaning of accuracy is inherently normative and teleological because it is evaluated in relation to an AI system's intended purpose and the prevention or mitigation of harms. Ai Safety And Ethics | positive | Risk mitigation and protection against harms associated with AI-system use |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Accuracy has different and sometimes contradictory meanings and expectations in the machine-learning and legal communities. Governance And Regulation | mixed | Alignment between technical and legal interpretations of accuracy |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Neither the technical machine-learning community nor the legal community can authorize a single universal notion of accuracy. Governance And Regulation | negative | Existence of a universally accepted definition of accuracy |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The paper argues that machine-learning accuracy metrics have shortcomings regarding the usefulness and effectiveness of deployed machine-learning systems. Output Quality | negative | Usefulness and effectiveness of machine-learning systems beyond benchmark accuracy |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Accuracy measurements may have limited validity when transferred across individuals, deployment scenarios, or contexts different from those used for measurement. Output Quality | negative | Generalizability and validity of accuracy measurements |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The paper recommends that reports of AI accuracy justify the selected metric and data, explain the relationship between the AI system's target and the intended-purpose construct, include baseline comparisons, and use context-dependent interventional studies. Regulatory Compliance | positive | Transparency and contextual validity of accuracy reporting |
Reading fidelity
high
Study strength
low
|
not reported
|