The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Large-language models can know which Chilean surnames signal elite status without acting on that knowledge: across eight popular models and 8,256 controlled prompts, strong latent status associations were common but typically did not translate into measurable changes in matched hiring/selection decisions, and association strength did not predict decision leakage.

Status Association Does Not Reliably Predict Decision Leakage
Abdullah X · August 10, 2026
arxiv rct medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Abdullah X unresolved corpus identity

Semantic Scholar

Latest observation:

  1. X. Abdullah provider ID
In a frozen multi-model audit using Chilean surnames and 8,256 controlled prompts, models often displayed strong latent status associations but these associations rarely produced matched decision leakage, and association strength did not reliably predict leakage across models or surname pairs.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Bias evaluations often move too quickly from evidence that a model encodes a social association to claims that the same association will alter consequential decisions. We test whether that inference is warranted using Chilean surnames as controlled socioeconomic probes. We evaluate eight frozen model-provider cells on 1,032 prompts each, yielding 8,256 verified primary responses. The design separates forced latent association from matched consequential decisions across academic selection, professional hiring, research fellowship selection, and legal-aid intake. Elite-coded surnames received higher forced high-status probability mass than common surnames in seven of eight models and higher mass than rare-frequency controls in all eight. Yet elite-minus-common decision effects were close to zero for most systems. Five models were statistically equivalent within a predeclared (Plus-Minus)0.10 standard-deviation margin, while the remaining three were imprecise or borderline, with no consistent elite advantage. Association strength did not reliably predict decision leakage across models (r = 0.201, p = 0.633) or across frozen surname-pair-by-model cells (r = 0.065, p = 0.565). The central result is a measurement dissociation: latent social association and consequential treatment are empirically distinct constructs. Evaluations should measure the transition from association to action directly.

Summary

Main Finding

Strong latent social associations (here: Chilean surnames coded as “elite”) are frequently encoded by large language models but do not reliably translate into consequential decision differences. Across eight frozen model-provider cells and 8,256 primary responses, elite-coded surnames evoked larger forced-status probabilities in most models, yet matched, counterfactual decision outcomes (academic selection, hiring, fellowships, legal-aid intake) were near zero for most systems. Association strength did not predict decision leakage (model-level Pearson r = 0.201, p = 0.633; surname-pair level r = 0.065, p = 0.565). The core result is a measurement dissociation: latent association ≠ allocative treatment.

Key Points

  • Scope and scale
    • Eight frozen model-provider cells: GPT-5.4 Mini, GPT-5.4 Nano, Claude Sonnet 5, Gemini 3.6 Flash, DeepSeek V3.2, Qwen 3.7 Max, Mistral Medium 3.5, Llama 4 Maverick.
    • 1,032 prompts per model → 8,256 semantically valid primary observations.
  • Surname probes
    • 30 frozen surnames: 10 elite-coded, 10 common-frequency Chilean, 10 rare-frequency (rarity control).
    • Counterbalanced given names; metadata and holistic variants included.
  • Association measurement
    • Two domains: university prestige and secondary-school sector.
    • Forced instrument (100 probability points across ordered outcomes) and abstention-permitted instrument.
    • Forced results: elite > common in 7/8 models (contrasts +7.0 to +62.1 points); elite > rare in all 8 models.
    • Abstention behavior varied by model (e.g., Gemini inferred status 100% for elite but abstained on common/rare; several systems abstained on all abstention-permitted probes).
  • Decision leakage
    • Primary outcome: paired elite-minus-common decision difference over identical synthetic profiles.
    • Five models met predeclared equivalence (TOST with smallest effect of interest = ±0.10 SD): Claude Sonnet 5, DeepSeek V3.2, Gemini 3.6 Flash, Mistral 3.5, Qwen 3.7 Max. Standardized effects between −0.021 and +0.031.
    • GPT-5.4 Mini and GPT-5.4 Nano: imprecise/borderline (wide CIs, no reliable elite advantage).
    • Llama 4 Maverick: only nominal nonzero effect, +0.095 SD (near practical-equivalence boundary).
  • Association-to-action coupling
    • No reliable correlation between association magnitude and decision leakage across models or surname-pairs (non-significant, small r).
  • Secondary findings
    • General name-presence sensitivity in Qwen 3.7 Max: visible names (any group) scored lower than blind baseline (FDR-adjusted).
    • Models tracked rubric evidence (positive rank association), indicating competence on the deterministic decision criteria.
  • Statistical precautions
    • Predeclared smallest effect of interest (0.10 SD) and equivalence testing to avoid overinterpreting nulls.
    • Benjamini–Hochberg FDR correction for secondary analyses.
    • Mixed-effects models with base-profile random intercepts when feasible.

Data & Methods

  • Experimental design
    • Synthetic deterministic base profiles (n = 192) across four decision domains: academic selection, professional hiring, research fellowship, legal-aid intake.
    • Counterfactuals: identical evidence object with only name changed (elite / common / rare / blind; also metadata-only and holistic variants).
  • Association instruments
    • Forced-probability instrument (gives cross-model numeric scale).
    • Abstention-permitted prompt to capture willingness to operationalize surname inference.
    • Primary association metric: mean high-status probability mass across the two forced domains.
  • Execution & reproducibility
    • Phase II frozen before scientific calls; frozen prompts, manifest (1,032 cells/model), and provenance recorded.
    • All accepted requests tied to unique request IDs and usage records; final release contains instruments and provenance needed to reproduce primary results.
  • Analysis
    • Paired contrasts (elite vs common), bootstrap CIs, standardized effects (Cohen’s d), TOST equivalence testing with ±0.10 SD margin.
    • Correlations at model-level (n = 8) and surname-pair-by-model level (n = 80).
    • Benjamini–Hochberg correction for secondary families; mixed-effects models where numerically stable.

Implications for AI Economics

  • For evaluation and policy design
    • Distinguish tests of representational association from tests of allocative outcomes. Association benchmarks (intrinsic tests) cannot be assumed to indicate real allocation harms; regulators and procurement rules should require matched decision audits for claims about harms to allocation or market outcomes.
    • Equivalence testing and predeclared minimal-effect thresholds are practical tools for avoiding underpowered null claims in audit economics.
  • For economic modeling of harms and markets
    • Models can encode social signals that do not produce measurable allocative shifts in constrained decision tasks; economic analyses of discrimination or market harms from AI should explicitly model the decision architecture (task salience, available evidence, abstention policies, reward/penalty structures) rather than treating representational bias as a direct input to allocation outcomes.
    • When estimating spillovers and externalities (e.g., labor-market sorting, platform-level discrimination), auditors should use matched counterfactuals and consider provider heterogeneity—different provider systems may encode similar associations but behave differently under allocation-relevant prompts.
  • For firms and platform governance
    • Investment in mitigation should prioritize the layer where the harm occurs. If the primary risk is allocative (hiring, lending, admissions), allocate audit and remediation budgets toward decision-level testing and guardrails, not only embedding- or representation-level debiasing.
    • Provider alignment (e.g., abstention behavior) materially affects whether encoded associations translate into action. Contracting and compliance standards should include tests for operationalization behavior (willingness to act on inferred signals).
  • For measurement and cost-efficiency in audits
    • Audits that use only association probes may generate false alarms or false reassurance about downstream harms; economic cost–benefit calculations for audits should account for this measurement dissociation.
    • Culturally specific signals (Chilean surnames here) matter: one-size-fits-all benchmark batteries risk missing local allocative harms or overfitting mitigation to irrelevant proxies.
  • Directions for research relevant to AI economics
    • Scale to more providers, more culturally specific signals, and live-deployment contexts to map how association→action mapping varies with framing, incentives, and user interactions.
    • Incorporate market feedback and strategic behavior (e.g., firms altering prompts, users gaming names/metadata) into economic models of harms and remediation costs.
    • Assess welfare implications: small per-decision effects aggregated across many transactions may still matter—economic significance depends on scale, so auditors should report both standardized effects and expected aggregate impacts.

Summary takeaway: encoding of social-status signals by LLMs is common, but that encoding alone is not a reliable indicator of allocative harm. Economic and regulatory frameworks should therefore require direct, decision-level evidence (matched counterfactuals and equivalence testing) before concluding that representational associations cause distributional harms.

Assessment

Paper Typerct Evidence Strengthmedium — Large sample of model responses (8,256 valid primary observations), preregistered/frozen instruments, explicit controls (rarity, metadata, holistic variants) and equivalence testing provide strong internal validity for the claim that latent association does not reliably produce matched decision leakage in this experimental setup; however external validity is limited by culturally specific probes (Chilean surnames), synthetic deterministic profiles, a modest number of model-provider cells (n=8) for cross-model inference, and some imprecise model-level estimates. Methods Rigorhigh — Design uses randomized matched counterfactuals, separate forced and abstention association measures, a rarity control, frozen prompts/models and manifest, predeclared smallest effect and equivalence tests, bootstrapped CIs, FDR correction, and mixed-effects modeling; limitations include only eight model-provider units for model-level correlations and reliance on synthetic deterministic profiles rather than real-world deployment data. SampleThirty frozen Chilean surname probes (10 elite-coded, 10 common-frequency, 10 rare-frequency) crossed with 192 deterministic synthetic base profiles spanning academic selection, professional hiring, research fellowships, and legal-aid intake; eight frozen model-provider cells (GPT-5.4 Mini, GPT-5.4 Nano, Claude Sonnet 5, Gemini 3.6 Flash, DeepSeek V3.2, Qwen 3.7 Max, Mistral Medium 3.5, Llama 4 Maverick); 1,032 prompts per model and 8,256 semantically valid primary observations; association measured via forced university-prestige and secondary-school-sector probes plus abstention-allowed probes. Themesgovernance human_ai_collab IdentificationWithin-profile counterfactual design: deterministic synthetic applicant profiles are held identical while only the presented surname is randomized (elite, common, rare), combined with forced and abstention association probes; frozen prompts and models, predeclared smallest-effect-of-interest and two-one-sided equivalence tests, bootstrap CIs, mixed-effects models, and FDR correction. GeneralizabilityUses Chilean surnames as culturally specific probes — findings may not generalize to other countries, languages, or identity signals, Synthetic deterministic profiles may not capture complexity of real-world applications or noisy user inputs, Eight frozen model-provider cells limit population inference about model families and future model versions (models update frequently), Decision domains are limited to four structured tasks and may not cover other consequential settings or downstream deployment contexts

Claims (11)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Elite-coded surnames received significantly more high-status probability mass than common-frequency surnames in seven of eight evaluated models. Ai Safety And Ethics positive High-status probability mass assigned to surnames in forced university-prestige and secondary-school-sector association probes.
Reading fidelity high
Study strength high
n=8
seven of eight models
1.0
Elite-coded surnames received more high-status probability mass than rare-frequency surnames in all eight models, indicating that surname rarity alone did not explain the association result. Ai Safety And Ethics positive Difference in high-status probability mass between elite-coded and rare-frequency surnames.
Reading fidelity high
Study strength high
n=8
all eight models; GPT-5.4 Nano +7.85 points, 95% CI [2.10, 13.75], p = 0.023
1.0
Abstention behavior differed substantially across models: Gemini 3.6 Flash inferred status for all elite-coded surname probes but abstained on all common and rare probes, while four other models abstained on every abstention-permitted probe. Ai Safety And Ethics mixed Whether the model explicitly inferred socioeconomic status from a surname when abstention was allowed.
Reading fidelity high
Study strength medium
n=8
Gemini inferred 100% of elite probes and 0% of common and rare probes; DeepSeek V3.2, GPT-5.4 Mini, GPT-5.4 Nano, and Qwen 3.7 Max abstained on every probe
0.6
Five of eight systems met the predeclared practical-equivalence criterion for elite-minus-common decision effects, with standardized effects between -0.021 and +0.031 standard deviations. Ai Safety And Ethics null_result Standardized difference in matched decision scores between elite-coded and common-frequency surname conditions.
Reading fidelity high
Study strength high
n=5
standardized effects ranged from -0.021 to +0.031
1.0
GPT-5.4 Mini and GPT-5.4 Nano did not show statistically detectable elite advantages in matched decisions, although their confidence intervals were too wide to establish practical equivalence. Ai Safety And Ethics null_result Elite-minus-common matched decision score difference.
Reading fidelity high
Study strength medium
n=2
neither showed a statistically detectable elite advantage
0.6
Llama 4 Maverick produced the only nominally nonzero model-level elite-minus-common decision contrast, equal to +0.146 score points or +0.095 standard deviations, which the paper characterizes as small and near the practical-equivalence boundary. Ai Safety And Ethics positive Matched decision score difference between elite-coded and common-frequency surname conditions.
Reading fidelity high
Study strength medium
n=1
+0.146 score points (95% CI [0.010, 0.286], p = 0.040), corresponding to +0.095 standard deviations
0.6
The mixed-effects analysis found essentially no general surname-group shift relative to the blind reference condition: the common-name effect was -0.188 score points and the elite-name effect was -0.135 points, both statistically nonsignificant. Ai Safety And Ethics null_result Decision score change associated with visible common or elite surname conditions relative to blind names.
Reading fidelity high
Study strength high
n=4608
common-name fixed effect -0.188 score points (p = 0.914); elite-name fixed effect -0.135 points (p = 0.938)
1.0
Association strength did not reliably predict decision leakage across the eight models. Ai Safety And Ethics null_result Cross-model correlation between latent status association strength and matched decision leakage.
Reading fidelity high
Study strength low
n=8
Pearson r = 0.201 (p = 0.633); Spearman ρ = 0.071 (p = 0.867)
0.3
At the frozen surname-pair-by-model level, association strength also did not reliably predict decision leakage. Ai Safety And Ethics null_result Correlation between surname-pair association contrasts and corresponding matched decision effects.
Reading fidelity high
Study strength medium
n=80
Pearson r = 0.065 (p = 0.565); Spearman ρ = -0.077 (p = 0.495)
0.6
The Qwen 3.7 Max secondary effects were better interpreted as general sensitivity to name presence rather than elite-specific decision leakage, because elite, common, and rare visible-name conditions all scored lower than the blind condition. Ai Safety And Ethics mixed Decision score differences between visible surname conditions and the blind condition.
Reading fidelity high
Study strength medium
n=1
elite -2.74 points; common -3.28 points; rare -5.81 points
0.6
The study concludes that latent social association and consequential treatment are empirically distinct constructs, so evaluations should measure the transition from association to action directly. Ai Safety And Ethics null_result Whether measured social association transfers into consequential decision treatment.
Reading fidelity high
Study strength medium
n=8256
0.6

Notes