0 cumulative citations
View corpus contextWhich fairness metric an auditor picks can reverse an AI system’s bias verdict: across four systems and 91,572 defensible computations the regulatory pass/fail flips in every case, and the reported figure moves far more than sampling error. Published disclosures almost never document the choices, but requiring a stated reference computation and consistency disclosure would eliminate most of the arbitrariness.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Summary
Main Finding
Metric choice and other defensible specification decisions (the "Alternative Ethics Measures" problem) move reported fairness outcomes far more than sampling error and can flip regulatory verdicts. Enumerating 91,572 defensible specifications across four decision systems, the author shows that for every system–protected-attribute pairing examined the regulatory fairness verdict (under a real threshold) reverses within the defensible specification space, the identity of the disadvantaged group can reverse, and the reported figure shifts 6–30× the size of sampling error. Mandatory disclosure that requires (a) a stated reference computation and (b) consistency reconciliation would recover most—but not all—of that discretion (≈78% and ≈82% closed, respectively); a nontrivial measured residual would remain.
Key Points
- New construct: Alternative Ethics Measures — quantitative ethical indicators (e.g., fairness metrics) that organizations select, define, compute and disclose where no canonical standard exists. Modeled on financial “alternative performance measures.”
- Three observable points of discretion that affect reported figures:
- Selection — which metric family (e.g., error-rate parity, predictive parity).
- Definition — thresholds, decision cutoff, reference population, group definition/comparison.
- Consistency — whether choices change across reporting periods and whether changes are disclosed.
- Study 1 (multiverse audit):
- Examined one deployed instrument (COMPAS) plus three canonical benchmarks across criminal justice, income classification, consumer credit and employment.
- Enumerated 91,572 defensible specifications (jointly varying metric family, thresholds, reference populations, groupings, etc.).
- In all 7 system × protected-attribute pairings: the regulatory verdict flips somewhere in the defensible space; the disadvantaged group can reverse; the metric range is 6–30× sampling error.
- Validation reproduced COMPAS error-rate disputes and traced differences to an unstated reference population in the original analysis.
- Study 2 (disclosure read):
- Reviewed 6,797 published disclosure documents from two disclosure regimes (one mandatory, one voluntary).
- Only 38 documents reported any quantified fairness measure.
- None of those 38 (a) justified metric selection, (b) named alternatives not chosen, or (c) reported prior-period (reconciled) figures.
- Study 3 (policy counterfactual / recovery estimation):
- Requiring reconciliation to a stated reference computation would close ~78% of the discretion window.
- Adding consistency disclosure across periods would increase closure to ~82%.
- A measurable residual remains that disclosure alone cannot eliminate.
- Analogy to financial reporting: regulators in accounting govern alternative earnings measures via disclosure rules (define, reconcile, ensure consistency) rather than by prescribing one “correct” measure; a similar disclosure regime could substantially reduce metric-discretion harms in AI ethics without settling normative debates over the “right” metric.
Data & Methods
- Multiverse / specification enumeration (Study 1):
- Treated an ethics audit as a multiverse analysis and exhaustively enumerated all defensible combinations of metric family, thresholds/cutoffs, reference populations and subgroup definitions for four decision systems.
- Computed fairness metrics (and regulatory verdicts such as the four-fifths/impact-ratio threshold) across the full specification space (91,572 specifications).
- Performed decomposition to quantify which specification choices move the reported figure most and compared those magnitudes to sampling error.
- Reproduced historical COMPAS analyses to validate sensitivity and identify the role of unstated population choices.
- Disclosure corpus reading (Study 2):
- Manual/electronic review of 6,797 disclosure documents from two disclosure regimes (one mandatory, one voluntary).
- Coded documents for presence of quantified fairness measures, documentation of metric selection, naming of unchosen alternatives, and prior-period reconciliation.
- Counterfactual disclosure recovery (Study 3):
- Conditioned the multiverse results on hypothetical disclosure requirements (stated reference computation and consistency reporting).
- Measured how much of the specification-induced range would be eliminated under each disclosure step and their combination.
- Key quantitative outcomes reported: 91,572 specifications enumerated; verdict reversals in all 7 pairings; reported figure movement 6–30× sampling error; 6,797 documents read with only 38 quantified fairness reports; 78% and 82% containment via reconciliation and consistency disclosure.
Implications for AI Economics
- Measurement uncertainty matters for economic analysis of AI:
- Cost-benefit calculations, compliance-cost estimates, market valuations and regulatory risk assessments that rely on single reported fairness numbers can be severely biased by metric discretion.
- Estimates of social welfare impacts, distributional effects and externalities will vary depending on metric choice and specification; analysts must account for specification uncertainty.
- Incentives and firm behavior:
- In the absence of disclosure rules, firms can (intentionally or not) present the fairness computation that is most favorable, creating selection incentives analogous to non-GAAP earnings management.
- Requiring disclosure of metric, reference population, thresholds, and reconciliation reduces opportunistic selection opportunities and improves comparability across firms and over time.
- Policy design and regulatory economics:
- Regulators can materially reduce metric-driven variability without resolving normative disputes about the “correct” fairness metric by mandating disclosures (definition + reconciliation + consistency).
- Mandated disclosure is a cost-effective intervention: the paper estimates ~78–82% of the discretion window can be closed with modest reporting rules; residual uncertainty will remain and may need supplementary controls (auditor involvement, independent verification).
- Standardization vs. disclosure tradeoff: full metric standardization is not necessary to obtain most gains; disclosure gives regulators and market actors information to compare systems while preserving flexibility for context-specific normative choices.
- Market-level effects:
- Improved transparency will facilitate more accurate third-party assessments (investors, procurers, civil-society auditors), potentially shifting competition toward more robust, better-documented systems.
- Residual, disclosure-resistant variance implies demand for independent assurance (auditors), creating new service markets and signaling mechanisms—similar to how assurance markets evolved around non-GAAP and ESG reporting.
- Research and evaluation implications:
- Empirical studies and meta-analyses in AI economics should report sensitivity across defensible metric specifications (a multiverse approach) rather than relying on single-point estimates.
- Cost-of-compliance models and empirical estimates of regulatory impact should incorporate specification uncertainty and the potential for metric-driven strategic reporting.
Caveats and limits - The analysis focuses on fairness metrics where protected attributes are recorded; many real-world settings lack such attributes or use inferred proxies, which introduces further complications not measured here. - Disclosure reduces but does not eliminate discretion; remaining residual variation may require auditor involvement or substantive standard-setting for additional mitigation. - The exact regimes used in the disclosure corpus are described in the paper as “one mandatory and one voluntary”; the summary preserves that characterization rather than attributing specific laws beyond those discussed in the paper.
Bottom line: Metric discretion is large and economically consequential. Mandating clear disclosure of the metric family, computation choices (reference population, thresholds), and prior-period reconciliation would recover the bulk of the comparability lost to defensible specification choices without forcing a single canonical fairness standard.
Assessment
Claims (10)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Across 91,572 defensible specifications of four AI decision systems, the regulatory four-fifths fairness verdict flipped in all seven system–protected-attribute pairings examined. Ai Safety And Ethics | mixed | Whether the system passed or failed the regulatory four-fifths fairness rule |
Reading fidelity
high
Study strength
medium
|
n=91572
7 of 7 pairings
|
| The identity of the group classified as disadvantaged reversed across defensible fairness specifications. Ai Safety And Ethics | mixed | Which protected group was identified as disadvantaged |
Reading fidelity
high
Study strength
medium
|
n=7
the disadvantaged group reverses
|
| The variation in reported fairness figures across defensible specifications was 6 to 30 times larger than the sampling error reported by the audits. Ai Safety And Ethics | negative | Variability in the reported fairness measure |
Reading fidelity
high
Study strength
medium
|
n=91572
6 to 30 times the size of sampling error
|
| The study's validation exercise reproduced the error rates reported in the COMPAS fairness dispute, with the results depending on a reference population that the original analysis did not state. Ai Safety And Ethics | mixed | Race-specific COMPAS error rates |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Among 6,797 published disclosure documents from two regulatory regimes, only 38 reported a quantified fairness measure. Governance And Regulation | negative | Whether a disclosure document reported a quantified fairness measure |
Reading fidelity
high
Study strength
medium
|
n=6797
38 reporting a quantified fairness measure
|
| None of the 38 disclosure documents reporting a quantified fairness measure justified the metric selection, identified an alternative metric that was not chosen, or reported a prior-period figure. Governance And Regulation | negative | Completeness of fairness-measure disclosure |
Reading fidelity
high
Study strength
medium
|
n=38
0 of 38 disclosures
|
| Reconciling a disclosed fairness result to a stated reference computation would eliminate 78% of the discretion window. Governance And Regulation | positive | Reduction in the range of possible fairness results attributable to audit discretion |
Reading fidelity
high
Study strength
medium
|
n=91572
78 per cent of the discretion window
|
| Adding consistency disclosure to reconciliation would reduce the discretion window by 82% in total. Governance And Regulation | positive | Residual audit metric discretion after disclosure requirements |
Reading fidelity
high
Study strength
medium
|
n=91572
82 per cent
|
| A residual amount of metric discretion remains even after reconciliation and consistency disclosure. Governance And Regulation | negative | Residual variation in fairness findings after disclosure |
Reading fidelity
high
Study strength
medium
|
n=91572
18 per cent of the discretion window remains
|
| Current AI audit and disclosure regimes require quantified findings but do not specify which fairness metric, decision threshold, reference population, or comparison groups auditors must use. Governance And Regulation | negative | Specificity and standardization of regulatory fairness-measurement requirements |
Reading fidelity
high
Study strength
medium
|
n=3
|