0 cumulative citations
View corpus contextLarge-language models can know which Chilean surnames signal elite status without acting on that knowledge: across eight popular models and 8,256 controlled prompts, strong latent status associations were common but typically did not translate into measurable changes in matched hiring/selection decisions, and association strength did not predict decision leakage.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Bias evaluations often move too quickly from evidence that a model encodes a social association to claims that the same association will alter consequential decisions. We test whether that inference is warranted using Chilean surnames as controlled socioeconomic probes. We evaluate eight frozen model-provider cells on 1,032 prompts each, yielding 8,256 verified primary responses. The design separates forced latent association from matched consequential decisions across academic selection, professional hiring, research fellowship selection, and legal-aid intake. Elite-coded surnames received higher forced high-status probability mass than common surnames in seven of eight models and higher mass than rare-frequency controls in all eight. Yet elite-minus-common decision effects were close to zero for most systems. Five models were statistically equivalent within a predeclared (Plus-Minus)0.10 standard-deviation margin, while the remaining three were imprecise or borderline, with no consistent elite advantage. Association strength did not reliably predict decision leakage across models (r = 0.201, p = 0.633) or across frozen surname-pair-by-model cells (r = 0.065, p = 0.565). The central result is a measurement dissociation: latent social association and consequential treatment are empirically distinct constructs. Evaluations should measure the transition from association to action directly.
Summary
Main Finding
Strong latent social associations (here: Chilean surnames coded as “elite”) are frequently encoded by large language models but do not reliably translate into consequential decision differences. Across eight frozen model-provider cells and 8,256 primary responses, elite-coded surnames evoked larger forced-status probabilities in most models, yet matched, counterfactual decision outcomes (academic selection, hiring, fellowships, legal-aid intake) were near zero for most systems. Association strength did not predict decision leakage (model-level Pearson r = 0.201, p = 0.633; surname-pair level r = 0.065, p = 0.565). The core result is a measurement dissociation: latent association ≠ allocative treatment.
Key Points
- Scope and scale
- Eight frozen model-provider cells: GPT-5.4 Mini, GPT-5.4 Nano, Claude Sonnet 5, Gemini 3.6 Flash, DeepSeek V3.2, Qwen 3.7 Max, Mistral Medium 3.5, Llama 4 Maverick.
- 1,032 prompts per model → 8,256 semantically valid primary observations.
- Surname probes
- 30 frozen surnames: 10 elite-coded, 10 common-frequency Chilean, 10 rare-frequency (rarity control).
- Counterbalanced given names; metadata and holistic variants included.
- Association measurement
- Two domains: university prestige and secondary-school sector.
- Forced instrument (100 probability points across ordered outcomes) and abstention-permitted instrument.
- Forced results: elite > common in 7/8 models (contrasts +7.0 to +62.1 points); elite > rare in all 8 models.
- Abstention behavior varied by model (e.g., Gemini inferred status 100% for elite but abstained on common/rare; several systems abstained on all abstention-permitted probes).
- Decision leakage
- Primary outcome: paired elite-minus-common decision difference over identical synthetic profiles.
- Five models met predeclared equivalence (TOST with smallest effect of interest = ±0.10 SD): Claude Sonnet 5, DeepSeek V3.2, Gemini 3.6 Flash, Mistral 3.5, Qwen 3.7 Max. Standardized effects between −0.021 and +0.031.
- GPT-5.4 Mini and GPT-5.4 Nano: imprecise/borderline (wide CIs, no reliable elite advantage).
- Llama 4 Maverick: only nominal nonzero effect, +0.095 SD (near practical-equivalence boundary).
- Association-to-action coupling
- No reliable correlation between association magnitude and decision leakage across models or surname-pairs (non-significant, small r).
- Secondary findings
- General name-presence sensitivity in Qwen 3.7 Max: visible names (any group) scored lower than blind baseline (FDR-adjusted).
- Models tracked rubric evidence (positive rank association), indicating competence on the deterministic decision criteria.
- Statistical precautions
- Predeclared smallest effect of interest (0.10 SD) and equivalence testing to avoid overinterpreting nulls.
- Benjamini–Hochberg FDR correction for secondary analyses.
- Mixed-effects models with base-profile random intercepts when feasible.
Data & Methods
- Experimental design
- Synthetic deterministic base profiles (n = 192) across four decision domains: academic selection, professional hiring, research fellowship, legal-aid intake.
- Counterfactuals: identical evidence object with only name changed (elite / common / rare / blind; also metadata-only and holistic variants).
- Association instruments
- Forced-probability instrument (gives cross-model numeric scale).
- Abstention-permitted prompt to capture willingness to operationalize surname inference.
- Primary association metric: mean high-status probability mass across the two forced domains.
- Execution & reproducibility
- Phase II frozen before scientific calls; frozen prompts, manifest (1,032 cells/model), and provenance recorded.
- All accepted requests tied to unique request IDs and usage records; final release contains instruments and provenance needed to reproduce primary results.
- Analysis
- Paired contrasts (elite vs common), bootstrap CIs, standardized effects (Cohen’s d), TOST equivalence testing with ±0.10 SD margin.
- Correlations at model-level (n = 8) and surname-pair-by-model level (n = 80).
- Benjamini–Hochberg correction for secondary families; mixed-effects models where numerically stable.
Implications for AI Economics
- For evaluation and policy design
- Distinguish tests of representational association from tests of allocative outcomes. Association benchmarks (intrinsic tests) cannot be assumed to indicate real allocation harms; regulators and procurement rules should require matched decision audits for claims about harms to allocation or market outcomes.
- Equivalence testing and predeclared minimal-effect thresholds are practical tools for avoiding underpowered null claims in audit economics.
- For economic modeling of harms and markets
- Models can encode social signals that do not produce measurable allocative shifts in constrained decision tasks; economic analyses of discrimination or market harms from AI should explicitly model the decision architecture (task salience, available evidence, abstention policies, reward/penalty structures) rather than treating representational bias as a direct input to allocation outcomes.
- When estimating spillovers and externalities (e.g., labor-market sorting, platform-level discrimination), auditors should use matched counterfactuals and consider provider heterogeneity—different provider systems may encode similar associations but behave differently under allocation-relevant prompts.
- For firms and platform governance
- Investment in mitigation should prioritize the layer where the harm occurs. If the primary risk is allocative (hiring, lending, admissions), allocate audit and remediation budgets toward decision-level testing and guardrails, not only embedding- or representation-level debiasing.
- Provider alignment (e.g., abstention behavior) materially affects whether encoded associations translate into action. Contracting and compliance standards should include tests for operationalization behavior (willingness to act on inferred signals).
- For measurement and cost-efficiency in audits
- Audits that use only association probes may generate false alarms or false reassurance about downstream harms; economic cost–benefit calculations for audits should account for this measurement dissociation.
- Culturally specific signals (Chilean surnames here) matter: one-size-fits-all benchmark batteries risk missing local allocative harms or overfitting mitigation to irrelevant proxies.
- Directions for research relevant to AI economics
- Scale to more providers, more culturally specific signals, and live-deployment contexts to map how association→action mapping varies with framing, incentives, and user interactions.
- Incorporate market feedback and strategic behavior (e.g., firms altering prompts, users gaming names/metadata) into economic models of harms and remediation costs.
- Assess welfare implications: small per-decision effects aggregated across many transactions may still matter—economic significance depends on scale, so auditors should report both standardized effects and expected aggregate impacts.
Summary takeaway: encoding of social-status signals by LLMs is common, but that encoding alone is not a reliable indicator of allocative harm. Economic and regulatory frameworks should therefore require direct, decision-level evidence (matched counterfactuals and equivalence testing) before concluding that representational associations cause distributional harms.
Assessment
Claims (11)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Elite-coded surnames received significantly more high-status probability mass than common-frequency surnames in seven of eight evaluated models. Ai Safety And Ethics | positive | High-status probability mass assigned to surnames in forced university-prestige and secondary-school-sector association probes. |
Reading fidelity
high
Study strength
high
|
n=8
seven of eight models
|
| Elite-coded surnames received more high-status probability mass than rare-frequency surnames in all eight models, indicating that surname rarity alone did not explain the association result. Ai Safety And Ethics | positive | Difference in high-status probability mass between elite-coded and rare-frequency surnames. |
Reading fidelity
high
Study strength
high
|
n=8
all eight models; GPT-5.4 Nano +7.85 points, 95% CI [2.10, 13.75], p = 0.023
|
| Abstention behavior differed substantially across models: Gemini 3.6 Flash inferred status for all elite-coded surname probes but abstained on all common and rare probes, while four other models abstained on every abstention-permitted probe. Ai Safety And Ethics | mixed | Whether the model explicitly inferred socioeconomic status from a surname when abstention was allowed. |
Reading fidelity
high
Study strength
medium
|
n=8
Gemini inferred 100% of elite probes and 0% of common and rare probes; DeepSeek V3.2, GPT-5.4 Mini, GPT-5.4 Nano, and Qwen 3.7 Max abstained on every probe
|
| Five of eight systems met the predeclared practical-equivalence criterion for elite-minus-common decision effects, with standardized effects between -0.021 and +0.031 standard deviations. Ai Safety And Ethics | null_result | Standardized difference in matched decision scores between elite-coded and common-frequency surname conditions. |
Reading fidelity
high
Study strength
high
|
n=5
standardized effects ranged from -0.021 to +0.031
|
| GPT-5.4 Mini and GPT-5.4 Nano did not show statistically detectable elite advantages in matched decisions, although their confidence intervals were too wide to establish practical equivalence. Ai Safety And Ethics | null_result | Elite-minus-common matched decision score difference. |
Reading fidelity
high
Study strength
medium
|
n=2
neither showed a statistically detectable elite advantage
|
| Llama 4 Maverick produced the only nominally nonzero model-level elite-minus-common decision contrast, equal to +0.146 score points or +0.095 standard deviations, which the paper characterizes as small and near the practical-equivalence boundary. Ai Safety And Ethics | positive | Matched decision score difference between elite-coded and common-frequency surname conditions. |
Reading fidelity
high
Study strength
medium
|
n=1
+0.146 score points (95% CI [0.010, 0.286], p = 0.040), corresponding to +0.095 standard deviations
|
| The mixed-effects analysis found essentially no general surname-group shift relative to the blind reference condition: the common-name effect was -0.188 score points and the elite-name effect was -0.135 points, both statistically nonsignificant. Ai Safety And Ethics | null_result | Decision score change associated with visible common or elite surname conditions relative to blind names. |
Reading fidelity
high
Study strength
high
|
n=4608
common-name fixed effect -0.188 score points (p = 0.914); elite-name fixed effect -0.135 points (p = 0.938)
|
| Association strength did not reliably predict decision leakage across the eight models. Ai Safety And Ethics | null_result | Cross-model correlation between latent status association strength and matched decision leakage. |
Reading fidelity
high
Study strength
low
|
n=8
Pearson r = 0.201 (p = 0.633); Spearman ρ = 0.071 (p = 0.867)
|
| At the frozen surname-pair-by-model level, association strength also did not reliably predict decision leakage. Ai Safety And Ethics | null_result | Correlation between surname-pair association contrasts and corresponding matched decision effects. |
Reading fidelity
high
Study strength
medium
|
n=80
Pearson r = 0.065 (p = 0.565); Spearman ρ = -0.077 (p = 0.495)
|
| The Qwen 3.7 Max secondary effects were better interpreted as general sensitivity to name presence rather than elite-specific decision leakage, because elite, common, and rare visible-name conditions all scored lower than the blind condition. Ai Safety And Ethics | mixed | Decision score differences between visible surname conditions and the blind condition. |
Reading fidelity
high
Study strength
medium
|
n=1
elite -2.74 points; common -3.28 points; rare -5.81 points
|
| The study concludes that latent social association and consequential treatment are empirically distinct constructs, so evaluations should measure the transition from association to action directly. Ai Safety And Ethics | null_result | Whether measured social association transfers into consequential decision treatment. |
Reading fidelity
high
Study strength
medium
|
n=8256
|