1 cumulative citations
View corpus contextA single accuracy number hides an important risk: different but equally accurate AI models can disagree on individual cases, creating arbitrary outcomes that clash with the EU AI Act; providers should quantify and report per-person disagreement (individual conflict ratios and δ-ambiguity) and give deployers access to this information so they can judge whether outputs are reliable for high-risk decisions.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
When building AI systems for decision support, one often encounters the phenomenon of predictive multiplicity: a single best model does not exist; instead, one can construct many models with similar overall accuracy that differ in their predictions for individual cases. Especially when decisions have a direct impact on humans, this can be highly unsatisfactory. For a person subject to high disagreement between models, one could as well have chosen a different model of similar overall accuracy that would have decided the person's case differently. We argue that this arbitrariness conflicts with the EU AI Act, which requires providers of high-risk AI systems to report performance not only at the dataset level but also for specific persons. The goal of this paper is to put predictive multiplicity in context with the EU AI Act's provisions on accuracy and to subsequently derive concrete suggestions on how to evaluate and report predictive multiplicity in practice. Specifically: (1) We introduce the AI Act's accuracy provisions and argue that incorporating information about predictive multiplicity could serve compliance with specific provisions for providers. (2) Based on this legally rigorous analysis, we suggest individual conflict ratios and $δ$-ambiguity as tools to quantify the disagreement between models on individual cases and to help detect individuals subject to conflicting predictions. (3) Based on computational insights, we derive easy-to-implement rules on how model providers could evaluate predictive multiplicity in practice. (4) Ultimately, we suggest that information about predictive multiplicity should be made available to deployers under the AI Act, enabling them to judge whether system outputs for specific individuals are reliable enough for their use case.
Summary
Main Finding
Predictive multiplicity — the existence of many models with similar overall accuracy but differing individual predictions — creates an arbitrariness that conflicts with the EU AI Act’s accuracy and transparency obligations for high‑risk AI systems. The authors show that reporting dataset-level accuracy is insufficient: providers should measure and disclose disagreement between comparably good models at the individual level. They propose pragmatic, model‑agnostic metrics (individual conflict ratios and δ‑ambiguity) and an ad‑hoc computational workflow providers can use to detect persons whose cases are unreliable because models disagree, and recommend making that information available to deployers to trigger human oversight where needed.
Key Points
-
Legal framing
- The EU AI Act (notably Art. 15(1) on accuracy and Art. 11(1), Art. 13(3)(b)(v), Annex IV(3) on required reporting) requires providers of high‑risk AI systems to assess and report performance not only at dataset level but also for “specific persons” or groups the system is intended to serve.
- Predictive multiplicity means that even high overall accuracy can mask substantial individual‑level unreliability; the authors argue that this is material to compliance with the AI Act’s accuracy/transparency duties.
- Provider vs deployer roles: providers should assess and disclose individual‑level conflict information in instructions for use; deployers should use that information to ensure intended, representative use and to apply human oversight when appropriate.
-
Technical contribution
- Two pragmatic metrics to quantify individual‑level disagreement:
- Individual conflict ratio: for a given data point, the fraction (or proportion) of comparably accurate models that predict a different label/outcome than a reference model or than the majority; higher values flag an individual as “conflicting.”
- δ‑ambiguity: a thresholded ambiguity measure that defines the set of “comparably good” models by allowing up to δ additional loss relative to a reference best model; δ‑ambiguity quantifies label disagreement among models inside that performance window.
- Practical evaluation: exact Rashomon set enumeration is infeasible for complex models; the authors recommend an ad‑hoc, model‑agnostic approach (retrain many models by varying initialization, hyperparameters, training data splits, preprocessing choices) to approximate the set of comparably good models and estimate conflict metrics.
- Computational rules of thumb: generate diverse candidate models until disagreement statistics stabilize; set thresholds to flag individuals (e.g., conflict ratio exceeding a chosen level); surface flagged cases to deployers for human review or additional safeguards.
- Comparison to uncertainty quantification: predictive multiplicity is complementary and advantageous in that it is model‑agnostic and only requires retraining; it does not rely on probabilistic calibration assumptions and can be integrated into existing pipelines and audits.
- Two pragmatic metrics to quantify individual‑level disagreement:
-
Policy/recommendation
- Providers should intentionally search for comparably performing models (the Rashomon set approximation) to measure predictive multiplicity, report individual conflict information in conformity documents and instructions for use, and make conflict indicators available to deployers.
- Deployers should use conflict indicators to decide whether outputs are reliable for a given person and to trigger human oversight when conflict is high.
- Making predictive multiplicity information available promotes better compliance with the AI Act, improves transparency, and supports trust in high‑risk AI deployments.
Data & Methods
- Legal method
- Doctrinal (black‑letter) analysis of the AI Act text, recitals, structure and teleology; focus on accuracy/transparency provisions applicable to high‑risk AI systems; interpretation of provider/deployer obligations and intended purpose constraints.
- Technical method
- Conceptual framing around the Rashomon set (the set of models with comparably good performance).
- Definition and formalization of two metrics:
- Individual conflict ratio (local disagreement measure).
- δ‑ambiguity (ambiguity measured within a δ performance tolerance).
- Computational approach:
- Ad‑hoc model ensemble generation: train many candidate models using variations in seeds, hyperparameters, training subsamples, preprocessing choices, and potentially objectives; treat models whose overall performance falls within a chosen tolerance (δ) as comparably good.
- Compute per‑instance disagreement statistics across this ensemble to estimate conflict ratios and δ‑ambiguity.
- Practical recommendations for audits: stop sampling candidate models when per‑instance disagreement metrics stabilize; report both dataset‑level and individual conflict measures; use thresholds to flag high‑conflict individuals for human oversight.
- Limitations acknowledged
- Exact enumeration of the Rashomon set is computationally infeasible for complex model classes; the ad‑hoc sampling is an approximation and may under/overestimate true multiplicity depending on coverage of model space.
- Choice of δ (performance tolerance) and thresholds for conflict are context‑dependent and should be informed by intended use, harm severity, and state of the art.
Implications for AI Economics
- Compliance costs and business practices
- Measuring predictive multiplicity introduces additional development and audit costs: providers must retrain/ensemble many models, compute conflict metrics, and document results for regulatory compliance. These costs may be significant for complex, high‑risk models and could alter marginal costs of offering such systems.
- Providers may internalize these costs into pricing of high‑risk AI products; smaller firms might face higher relative burdens, affecting market structure and competition.
- Incentives for model selection and product design
- Regulatory pressure to report individual‑level reliability incentivizes providers to search the Rashomon set for models that minimize harmful disagreements (e.g., selecting models that reduce conflict for vulnerable groups) or to choose models that allow clearer human oversight.
- Predictive multiplicity can become a monetizable feature: providers who can credibly demonstrate low individual conflict may gain market advantage (higher trust, easier procurement by regulated deployers).
- Deployers and procurement decisions
- Deployers will demand conflict indicators as part of procurement and compliance checks. This changes the information flow in the market: deployers gain a decision‑relevant signal (per‑person reliability) beyond aggregate accuracy.
- Procurement models may incorporate the expected cost of human oversight for flagged individuals; systems with high multiplicity raise expected operating costs for deployers (more manual review, slower throughput), affecting adoption decisions.
- Risk allocation, liability, and insurance
- Disclosure of predictive multiplicity can shift liability expectations: providers who disclose high conflict may be seen as compliant but also as warning deployers of inherent model arbitrariness; deployers who nevertheless use the system may assume more operational responsibility (human oversight), affecting legal and insurance arrangements.
- Insurers and regulators may price coverage or require mitigation (e.g., stricter oversight) for systems with high predicted individual conflict.
- Welfare and distributional effects
- If conflict disproportionately affects certain groups or individuals, unaddressed multiplicity can exacerbate inequality or produce uneven welfare outcomes; measuring multiplicity helps detect such distributional risks and enables corrective model selection or policy interventions.
- Conversely, awareness of multiplicity allows targeted remedies (e.g., human review, alternative decision paths), potentially reducing wrongful harms and improving social welfare relative to opaque deployments.
- Innovation and standards
- Demand for standardized measures of individual conflict (e.g., harmonized δ choices or reporting formats) could spur industry standards and new auditing services, creating markets for third‑party validators and standardization bodies.
- Over time, the need to control multiplicity could encourage research into methods that produce more stable individual predictions or into optimization procedures that trade off aggregate accuracy for lower individual conflict — influencing the direction of ML innovation.
Overall, incorporating predictive multiplicity into compliance and reporting shifts both technical practice and economic incentives: it raises short‑term compliance costs but creates information that can reduce deployment risks, align incentives for safer design choices, and reshape markets for high‑risk AI through procurement, liability, and standardization dynamics.
Assessment
Claims (7)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Predictive multiplicity occurs in decision-support AI: a single best model does not exist; instead, one can construct many models with similar overall accuracy that differ in their predictions for individual cases. Decision Quality | negative | disagreement between models on individual-case predictions |
Reading fidelity
high
Study strength
medium
|
not reported
|
| When decisions have a direct impact on humans, predictive multiplicity is highly unsatisfactory because individuals subject to high disagreement could have had their case decided differently by another model of similar overall accuracy. Decision Quality | negative | reliability/fairness of individual-level decisions |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| This arbitrariness (predictive multiplicity) conflicts with the EU AI Act, which requires providers of high-risk AI systems to report performance not only at the dataset level but also for specific persons. Governance And Regulation | negative | compliance with AI Act reporting requirements regarding individual-level performance |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Incorporating information about predictive multiplicity could serve compliance with specific provisions of the EU AI Act for providers of high-risk systems. Governance And Regulation | positive | ability to demonstrate compliance with AI Act accuracy/reporting provisions |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| The paper proposes individual conflict ratios and δ-ambiguity as tools to quantify disagreement between models on individual cases and to help detect individuals subject to conflicting predictions. Decision Quality | positive | degree of disagreement/predictive ambiguity on individuals |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| Based on computational insights, the authors derive easy-to-implement rules on how model providers could evaluate predictive multiplicity in practice. Governance And Regulation | positive | practical evaluation procedures for predictive multiplicity |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Information about predictive multiplicity should be made available to deployers under the AI Act so they can judge whether system outputs for specific individuals are reliable enough for their use case. Governance And Regulation | positive | availability of predictive-multiplicity information to deployers / deployer ability to assess reliability for individuals |
Reading fidelity
high
Study strength
speculative
|
not reported
|