The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Which fairness metric an auditor picks can reverse an AI system’s bias verdict: across four systems and 91,572 defensible computations the regulatory pass/fail flips in every case, and the reported figure moves far more than sampling error. Published disclosures almost never document the choices, but requiring a stated reference computation and consistency disclosure would eliminate most of the arbitrariness.

Same System, Opposite Verdicts: Metric Discretion in AI Ethics Audits and the Limits of Disclosure
Shay Tsaban · August 12, 2026 · Research Square
openalex descriptive high evidence 7/10 relevance Full text usable extracted full text DOI Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

OpenAlex

Latest observation:

  1. Shay Tsaban provider ID
Defensible choices about which fairness metric, threshold, reference population and comparison groups to report produce large variation in audit findings—across 91,572 specifications the regulatory verdict flips in every examined system—while real-world disclosures almost never document those choices, though simple disclosure requirements could recover most of the variation.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Summary

Main Finding

Metric choice and other defensible specification decisions (the "Alternative Ethics Measures" problem) move reported fairness outcomes far more than sampling error and can flip regulatory verdicts. Enumerating 91,572 defensible specifications across four decision systems, the author shows that for every system–protected-attribute pairing examined the regulatory fairness verdict (under a real threshold) reverses within the defensible specification space, the identity of the disadvantaged group can reverse, and the reported figure shifts 6–30× the size of sampling error. Mandatory disclosure that requires (a) a stated reference computation and (b) consistency reconciliation would recover most—but not all—of that discretion (≈78% and ≈82% closed, respectively); a nontrivial measured residual would remain.

Key Points

  • New construct: Alternative Ethics Measures — quantitative ethical indicators (e.g., fairness metrics) that organizations select, define, compute and disclose where no canonical standard exists. Modeled on financial “alternative performance measures.”
  • Three observable points of discretion that affect reported figures:
  • Selection — which metric family (e.g., error-rate parity, predictive parity).
  • Definition — thresholds, decision cutoff, reference population, group definition/comparison.
  • Consistency — whether choices change across reporting periods and whether changes are disclosed.
  • Study 1 (multiverse audit):
    • Examined one deployed instrument (COMPAS) plus three canonical benchmarks across criminal justice, income classification, consumer credit and employment.
    • Enumerated 91,572 defensible specifications (jointly varying metric family, thresholds, reference populations, groupings, etc.).
    • In all 7 system × protected-attribute pairings: the regulatory verdict flips somewhere in the defensible space; the disadvantaged group can reverse; the metric range is 6–30× sampling error.
    • Validation reproduced COMPAS error-rate disputes and traced differences to an unstated reference population in the original analysis.
  • Study 2 (disclosure read):
    • Reviewed 6,797 published disclosure documents from two disclosure regimes (one mandatory, one voluntary).
    • Only 38 documents reported any quantified fairness measure.
    • None of those 38 (a) justified metric selection, (b) named alternatives not chosen, or (c) reported prior-period (reconciled) figures.
  • Study 3 (policy counterfactual / recovery estimation):
    • Requiring reconciliation to a stated reference computation would close ~78% of the discretion window.
    • Adding consistency disclosure across periods would increase closure to ~82%.
    • A measurable residual remains that disclosure alone cannot eliminate.
  • Analogy to financial reporting: regulators in accounting govern alternative earnings measures via disclosure rules (define, reconcile, ensure consistency) rather than by prescribing one “correct” measure; a similar disclosure regime could substantially reduce metric-discretion harms in AI ethics without settling normative debates over the “right” metric.

Data & Methods

  • Multiverse / specification enumeration (Study 1):
    • Treated an ethics audit as a multiverse analysis and exhaustively enumerated all defensible combinations of metric family, thresholds/cutoffs, reference populations and subgroup definitions for four decision systems.
    • Computed fairness metrics (and regulatory verdicts such as the four-fifths/impact-ratio threshold) across the full specification space (91,572 specifications).
    • Performed decomposition to quantify which specification choices move the reported figure most and compared those magnitudes to sampling error.
    • Reproduced historical COMPAS analyses to validate sensitivity and identify the role of unstated population choices.
  • Disclosure corpus reading (Study 2):
    • Manual/electronic review of 6,797 disclosure documents from two disclosure regimes (one mandatory, one voluntary).
    • Coded documents for presence of quantified fairness measures, documentation of metric selection, naming of unchosen alternatives, and prior-period reconciliation.
  • Counterfactual disclosure recovery (Study 3):
    • Conditioned the multiverse results on hypothetical disclosure requirements (stated reference computation and consistency reporting).
    • Measured how much of the specification-induced range would be eliminated under each disclosure step and their combination.
  • Key quantitative outcomes reported: 91,572 specifications enumerated; verdict reversals in all 7 pairings; reported figure movement 6–30× sampling error; 6,797 documents read with only 38 quantified fairness reports; 78% and 82% containment via reconciliation and consistency disclosure.

Implications for AI Economics

  • Measurement uncertainty matters for economic analysis of AI:
    • Cost-benefit calculations, compliance-cost estimates, market valuations and regulatory risk assessments that rely on single reported fairness numbers can be severely biased by metric discretion.
    • Estimates of social welfare impacts, distributional effects and externalities will vary depending on metric choice and specification; analysts must account for specification uncertainty.
  • Incentives and firm behavior:
    • In the absence of disclosure rules, firms can (intentionally or not) present the fairness computation that is most favorable, creating selection incentives analogous to non-GAAP earnings management.
    • Requiring disclosure of metric, reference population, thresholds, and reconciliation reduces opportunistic selection opportunities and improves comparability across firms and over time.
  • Policy design and regulatory economics:
    • Regulators can materially reduce metric-driven variability without resolving normative disputes about the “correct” fairness metric by mandating disclosures (definition + reconciliation + consistency).
    • Mandated disclosure is a cost-effective intervention: the paper estimates ~78–82% of the discretion window can be closed with modest reporting rules; residual uncertainty will remain and may need supplementary controls (auditor involvement, independent verification).
    • Standardization vs. disclosure tradeoff: full metric standardization is not necessary to obtain most gains; disclosure gives regulators and market actors information to compare systems while preserving flexibility for context-specific normative choices.
  • Market-level effects:
    • Improved transparency will facilitate more accurate third-party assessments (investors, procurers, civil-society auditors), potentially shifting competition toward more robust, better-documented systems.
    • Residual, disclosure-resistant variance implies demand for independent assurance (auditors), creating new service markets and signaling mechanisms—similar to how assurance markets evolved around non-GAAP and ESG reporting.
  • Research and evaluation implications:
    • Empirical studies and meta-analyses in AI economics should report sensitivity across defensible metric specifications (a multiverse approach) rather than relying on single-point estimates.
    • Cost-of-compliance models and empirical estimates of regulatory impact should incorporate specification uncertainty and the potential for metric-driven strategic reporting.

Caveats and limits - The analysis focuses on fairness metrics where protected attributes are recorded; many real-world settings lack such attributes or use inferred proxies, which introduces further complications not measured here. - Disclosure reduces but does not eliminate discretion; remaining residual variation may require auditor involvement or substantive standard-setting for additional mitigation. - The exact regimes used in the disclosure corpus are described in the paper as “one mandatory and one voluntary”; the summary preserves that characterization rather than attributing specific laws beyond those discussed in the paper.

Bottom line: Metric discretion is large and economically consequential. Mandating clear disclosure of the metric family, computation choices (reference population, thresholds), and prior-period reconciliation would recover the bulk of the comparability lost to defensible specification choices without forcing a single canonical fairness standard.

Assessment

Paper Typedescriptive Evidence Strengthhigh — The paper triangulates its claim with three empirical studies: a comprehensive multiverse enumeration (91,572 defensible specifications) across four decision systems including a deployed instrument (COMPAS) and three benchmarks; a systematic read of 6,797 disclosure documents under two regimes; and a quantitative exercise estimating how disclosure would narrow the range. The large enumeration, replication/validation of the COMPAS error-rate dispute, and cross-regime disclosure read provide convergent empirical support for the main claims, though the scope is limited to systems and settings where protected attributes are recorded. Methods Rigorhigh — The design transparently operationalises the audit as a multiverse, enumerates defensible choices (metric family, threshold, reference population, comparison groups), quantifies movement relative to sampling error, and triangulates with a large corpus of real-world disclosures; includes validation of a known dispute (COMPAS) and an explicit decomposition of contribution of choice points. Potential limitations include reliance on the author's specification set (subjectivity in what counts as 'defensible') and restriction to four systems and available disclosures. SampleStudy 1: Four decision systems (one deployed instrument—COMPAS in Broward County—and three canonical benchmark tasks) evaluated across 91,572 defensible fairness-specification combinations spanning criminal justice, income classification, consumer credit, and employment. Study 2: Manual/automated read of 6,797 published disclosure documents from two disclosure regimes (one mandatory, one voluntary) searching for quantified fairness measures. Study 3: Counterfactual conditioning of the Study 1 specification space on progressively fuller disclosure rules to estimate how much of the specification range disclosure would recover. Themesgovernance inequality adoption GeneralizabilityAnalysis covers four systems and may not generalize to all AI systems, sectors, or model architectures., Requires recorded protected attributes—findings do not apply where such attributes are absent or imputed., The definition of what constitutes a 'defensible' specification set is partly subjective and may omit other reasonable choices in different legal or cultural contexts., Disclosures and regulatory regimes sampled are time- and jurisdiction-specific; results may vary under other rules or future standards.

Claims (10)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Across 91,572 defensible specifications of four AI decision systems, the regulatory four-fifths fairness verdict flipped in all seven system–protected-attribute pairings examined. Ai Safety And Ethics mixed Whether the system passed or failed the regulatory four-fifths fairness rule
Reading fidelity high
Study strength medium
n=91572
7 of 7 pairings
0.18
The identity of the group classified as disadvantaged reversed across defensible fairness specifications. Ai Safety And Ethics mixed Which protected group was identified as disadvantaged
Reading fidelity high
Study strength medium
n=7
the disadvantaged group reverses
0.18
The variation in reported fairness figures across defensible specifications was 6 to 30 times larger than the sampling error reported by the audits. Ai Safety And Ethics negative Variability in the reported fairness measure
Reading fidelity high
Study strength medium
n=91572
6 to 30 times the size of sampling error
0.18
The study's validation exercise reproduced the error rates reported in the COMPAS fairness dispute, with the results depending on a reference population that the original analysis did not state. Ai Safety And Ethics mixed Race-specific COMPAS error rates
Reading fidelity high
Study strength medium
not reported
0.18
Among 6,797 published disclosure documents from two regulatory regimes, only 38 reported a quantified fairness measure. Governance And Regulation negative Whether a disclosure document reported a quantified fairness measure
Reading fidelity high
Study strength medium
n=6797
38 reporting a quantified fairness measure
0.18
None of the 38 disclosure documents reporting a quantified fairness measure justified the metric selection, identified an alternative metric that was not chosen, or reported a prior-period figure. Governance And Regulation negative Completeness of fairness-measure disclosure
Reading fidelity high
Study strength medium
n=38
0 of 38 disclosures
0.18
Reconciling a disclosed fairness result to a stated reference computation would eliminate 78% of the discretion window. Governance And Regulation positive Reduction in the range of possible fairness results attributable to audit discretion
Reading fidelity high
Study strength medium
n=91572
78 per cent of the discretion window
0.18
Adding consistency disclosure to reconciliation would reduce the discretion window by 82% in total. Governance And Regulation positive Residual audit metric discretion after disclosure requirements
Reading fidelity high
Study strength medium
n=91572
82 per cent
0.18
A residual amount of metric discretion remains even after reconciliation and consistency disclosure. Governance And Regulation negative Residual variation in fairness findings after disclosure
Reading fidelity high
Study strength medium
n=91572
18 per cent of the discretion window remains
0.18
Current AI audit and disclosure regimes require quantified findings but do not specify which fairness metric, decision threshold, reference population, or comparison groups auditors must use. Governance And Regulation negative Specificity and standardization of regulatory fairness-measurement requirements
Reading fidelity high
Study strength medium
n=3
0.18

Notes