Audit design, not demographic animus, often drives LLM verdicts: across hiring, lending and triage tests of five models, rating vs ranking did not produce the published reversal of racial disparities, and the largest measurable effect came from where a candidate was listed and whether the model recognized the audit.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
No provider observation is available for this paper.
Missing data, not a zero citation count.
Whether a language model looks demographically biased can depend on how the audit asks its question. A charitable-aid benchmark reports that the same models favor minority applicants when rating requests one at a time and penalize some when ranking side by side. We test whether that reversal generalizes to hiring, lending, and medical triage: 40,726 requests to five models, applications differing only in the applicant's name, and a primary test fixed before collection. It does not. None of 36 planned contrasts survives correction. The rating advantage keeps its sign at roughly half the published size, and a precision extension bounds any hiring ranking penalty below the published effect, though the lending and triage ranking floors sit above that margin, so the exclusion is conclusive for hiring ranking and for rating in all three domains only. Planted disparities tracking their injected sizes and a directional replication on the original aid materials bound these nulls. The audit is livelier than the demographics: models recognize transparent audits nearly always, tie every identical-content comparison whether the varying detail is race or a hobby, and reward first-listed candidates as much as any demographic effect we measure. Audit verdicts reflect audit construction more than demographic bias.
Summary
Main Finding
The paper shows that whether an LLM audit finds demographic bias depends strongly on the audit instrument (format, ordering, detectability) — audit construction often matters as much or more than the model’s demographic preferences. In three regulated decision domains (hiring, lending, triage), five contemporary models produced no demographic contrast that survived multiple-testing correction across 36 pre‑registered contrasts. A small rating advantage for minority names remained (≈+0.047 points, about half a previously reported effect) but ranking penalties largely were not detected; an exclusion at the previously reported rank effect size holds for hiring only. The largest measured effect was an audit design artifact: candidate list position (first-listed advantage).
Key Points
- Scope and scale: 40,726 requests to five models (GPT‑5.6 Terra, Claude Sonnet 5, Gemini 3.1 Pro, Kimi K3, Qwen3.5‑397B‑A17B) across three domains (hiring, lending, triage). Stimuli: 12 neutral base profiles × 40 name pairs (race×gender fully crossed), producing 480 single-profile stimuli per domain.
- Formats tested: Rate (1–5 score), Decide (binary shortlist/approve/escalate), Rank (four-candidate priority ordering). Rank bundles included transparent (matched) and disguised (co‑varied quality) presentations; ties were allowed.
- Pre-registered tests: 36 contrasts (3 domains × 4 demographic groups × 3 formats), Benjamini–Hochberg correction at q<0.05.
- Main statistical outcomes:
- None of the 36 contrasts survived correction; twelve format-by-group interactions indistinguishable from zero.
- Rate: pooled minority rating advantage ≈ +0.047 points (95% CI [+0.002, +0.092]) ≈ 0.04 SD — roughly half the previously published +0.090 effect; equivalence bound reported at 0.084 points (close to the published margin).
- Disguised Rank: pooled contrast +0.03 positions (95% CI [−0.15, +0.20]) — wide interval; authors do not claim absence except where equivalence tests permit.
- Hiring Rank: extension with many additional disguised bundles narrowed pooled Rank contrast to +0.008 positions (95% CI [−0.044, +0.060]) and bounded any hiring Rank penalty below 0.052 positions (pre‑specified margin 0.067; p=0.013) — i.e., excludes the previously reported ranking penalty for hiring.
- Lending & triage Rank: detection floors (0.132 and 0.232 positions) sit above the pre‑specified 0.067 margin, so penalties of that published size cannot be excluded in those domains.
- Sensitivity and controls:
- Detection floor: study resolves effects down to ~0.22 SD in the worst case.
- Planted disparities: recoveries track injected magnitudes (e.g., recovered ≈0.48 for an injected 0.50 rank disparity), so the design can detect medium-to-large effects.
- Re-running the original charitable-aid stimuli reproduced the published directions (not statistically significant replications), suggesting the phenomenon may depend on benchmark materials.
- Audit recognition and ties:
- Models almost always recognized transparent/matched audit bundles and returned ties when content was identical (100% ties for demographic-matched and irrelevant-attribute matched bundles). Disguised bundles yielded 0% ties.
- Forced-choice probes report high rates of recognizing disguised bundles as fairness tests (up to 70%), but volunteered recognition is much lower (≈2%), meaning probe wording strongly affects measured recognition.
- Audit instrument effects:
- Position (list ordering) had the largest measurable effect: first-listed candidates gained ≈0.11 rank positions (95% CI [0.03, 0.18]), ~0.094 SD — comparable to or larger than measured demographic effects.
- Wording/paraphrase differences did not produce robust demographic effects once application composition was controlled for; apparent reversals were due to subsample composition artifacts.
- Limitations flagged by authors: demographic signal limited to names (no dialects/photographs), limited base profiles (12 per domain), only five models, primary runs at temperature 0 with reasoning disabled where permitted, and results apply to these stimuli and configurations (not general fairness certification).
Data & Methods
- Stimuli: AgentFairBench base profiles (12 per domain) rendered under 40 validated name pairs (Elder & Hayes, 2023) fully crossed by race×gender → 480 single-profile stimuli per domain.
- Rank bundle design: 4 bundle types × 150 bundles per domain — (i) Demographic matched (one profile × four names, one per race), (ii) Irrelevant-attribute matched (four same-cell names differing by hobby/neighborhood), (iii) Demographic disguised (co‑varied profiles across quality strata balanced by Graeco‑Latin squares), (iv) Placebo disguised (same‑cell names, pseudo‑groups).
- Primary test fixed before collection: does the Black contrast in hiring reverse between formats? 36 planned contrasts formed the hypothesis family for multiple-testing correction.
- Models: five contemporary LLMs; temperature 0, reasoning disabled where possible, integer-only outputs, one retry.
- Analyses: mixed-effects regressions with random intercepts for base profile, name pair, and bundle; rank contrasts re-estimated with Plackett (rank-ordered logit) with bundle-clustered errors. Benjamini–Hochberg correction applied to the 36 contrasts. Pre-registered equivalence margins taken from the published aid study (0.090 rating points; 0.067 rank positions).
- Sensitivity checks: planted disparities, leave-one-name/profile-out checks, tool-call arms (function-call versions of Decide/Rank), paraphrase arms, and a purpose-probe to elicit whether models recognized audits.
Implications for AI Economics
- Audit design materially shapes measured discrimination in algorithmic decision nodes. Regulators and economists should treat audit outcomes as conditional on instrument choices (format, whether ties allowed, ordering, probe wording), not as model-invariant facts.
- Policy and certification:
- Fairness certification should require multiple audit formats (ratings, binary decisions, side‑by‑side rankings with disguised presentations), pre-registered sensitivity/equivalence bounds, and planted controls to establish detection floors; a single-format pass is insufficient.
- Audit procedures must control for ordering/position effects (which can rival demographic effects) — randomized ordering and reporting of position bias are essential.
- Deployment consequences:
- Position biases imply operational decisions (e.g., presentation order in automated shortlists or prioritized queues) can create substantial economic impacts independent of demographic preference. Firms should randomize or correct for ordering in downstream pipelines.
- Serializing decisions (tool-call APIs) did not materially change outcomes in this study, but audit-detection behavior and position effects mean integration choices (batch vs single-call) could still matter under other configurations.
- Measurement practice for researchers:
- Use equivalence testing and report detection floors; null results without sensitivity bounds are uninformative for policy.
- Include disguised (camouflaged) audits to avoid transparent-test artifacts, but be aware models may or may not recognize disguise depending on probe wording; both forced and volunteered recognition rates should be reported.
- Strategic concerns:
- Models’ ability to recognize transparent audits implies risks of gaming or trained avoidance if audits become part of model training or evaluation pipelines; audit methods and probes should avoid creating easy signals that models learn to exploit.
- Economic research agenda:
- Quantify how instrument effects (ordering, ties allowed, prompt wording) propagate into labor market, credit allocation, and health outcomes when LLMs are deployed at scale.
- Study alternative demographic cues (dialect, self-reported race/gender, images) and deployment-level dynamics (feedback loops, retraining on audit-exposed data).
- Caution on external validity: conclusions are bounded to name-based demographic signals, synthetic public profiles, and the five models tested; policymakers and firms should not generalize the null to all LLMs, all formats, or to demographic signals beyond names without further evidence.
Assessment
Claims (9)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Across hiring, lending, and medical triage, none of the 36 preregistered demographic contrasts survived Benjamini–Hochberg correction. Decision Quality | null_result | Demographic differences in ratings, binary decisions, and rankings |
Reading fidelity
high
Study strength
high
|
n=40726
|
| The primary hiring Black format-by-group interaction was estimated at +0.01, with a 95% interval from −0.15 to +0.17, providing no evidence of the predicted reversal. Decision Quality | null_result | Change in Black-versus-reference-group demographic contrast across formats |
Reading fidelity
high
Study strength
medium
|
n=40726
+0.01 [-0.15, +0.17]
|
| The rating advantage retained its positive sign but was approximately half the size of the previously published effect: +0.047 rating points compared with the published +0.090. Decision Quality | positive | Rating score difference between demographic groups |
Reading fidelity
high
Study strength
medium
|
n=40726
+0.047 points [+0.002, +0.092], versus published +0.090
|
| The disguised ranking contrast was not detected: the estimated effect was +0.03 rank positions with a 95% interval from −0.15 to +0.20. Decision Quality | null_result | Demographic difference in candidate ranking position |
Reading fidelity
high
Study strength
medium
|
n=40726
+0.03 positions [-0.15, +0.20]
|
| The study excluded a hiring ranking penalty as large as the published 0.067-position margin, but did not exclude such a penalty for lending or triage. Decision Quality | mixed | Magnitude of demographic penalty in ranking position across domains |
Reading fidelity
high
Study strength
high
|
n=12000
Hiring: +0.008 [-0.044, +0.060]; bound of 0.052 positions against a 0.067 margin; lending floor 0.132; triage floor 0.232
|
| Models tied every transparent demographic-matched ranking bundle and every matched bundle differing only in an irrelevant hobby or neighborhood, while tying none of the disguised or placebo bundles. Decision Quality | null_result | Frequency of tied rankings in identical-content and disguised comparisons |
Reading fidelity
high
Study strength
high
|
n=1800
100% ties in matched cells; 0% ties in disguised and placebo cells
|
| First-listed candidates gained 0.11 ranking positions, an effect comparable to or larger than the largest demographic effect measured on the same outcome. Decision Quality | positive | Candidate ranking position as a function of list position |
Reading fidelity
high
Study strength
medium
|
n=40726
+0.11 positions [0.03, 0.18]; 0.094 standard deviations versus 0.049
|
| Changing the wording of the hiring rating and ranking instructions did not produce a statistically distinguishable demographic interaction once application composition was controlled. Decision Quality | null_result | Race-by-wording interaction in rating and ranking outcomes |
Reading fidelity
high
Study strength
medium
|
Rate interactions: -0.01 [-0.07, +0.06] and -0.02 [-0.08, +0.04]; redealt arms: +0.01 [-0.05, +0.07] and -0.01 [-0.07, +0.05]
|
| Native function-call formatting did not materially change decision or ranking outcomes relative to text-format counterparts. Decision Quality | null_result | Difference between tool-call and text-format decisions and rankings |
Reading fidelity
high
Study strength
medium
|
submit_decision callback gap: +1.5pp [-1.2, +4.2]; submit_ranking contrast: +0.01 positions [-0.11, +0.12]
|