The Commonplace
Home Three-study pilot Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Open-weight LLMs used for recruitment respond to job-ad language in ways that disadvantage protected groups: agentic wording substantially lowers recommendations for female personas (rrb ≈ 0.31), while coded-exclusion language markedly reduces both recruiter scores and expressed interest for non-White personas (rrb ≈ 0.65–0.76). Label-ablation and embedding tests implicate explicit demographic labels and encoded representations, suggesting practical pre-deployment posting-language audits can flag adverse impact under regulatory thresholds.

Linguistic Triggers of Gender and Racial Bias in Open-Weight LLMs Applied to Recruitment
Kitahara, Kosuke, Yamaguchi, Nobuhiro · September 16, 2026 · arXiv (Cornell University)
openalex quasi_experimental medium evidence 8/10 relevance Full text usable extracted full text DOI Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

OpenAlex

Latest observation:

  1. Kitahara, Kosuke provider ID
  2. Yamaguchi, Nobuhiro provider ID

Semantic Scholar

Latest observation:

  1. Kosuke Kitahara provider ID
  2. Nobuhiro Yamaguchi provider ID
A multi-model audit finds that agentic (male-coded) posting language reduces recruiter recommendation scores for female personas and that coded-exclusion posting language sharply suppresses non-White recruiter scores and non-White job-seeker interest, with label-ablation and WEAT indicating explicit persona labels and representational embeddings drive much of the effect.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Open-weight large language models are rapidly entering hiring pipelines, yet their discriminatory failure modes -- and the regulatory exposure these create under the EU AI Act high-risk classification (Annex III) and U.S. EEOC adverse-impact analysis -- remain poorly understood. We present the first systematic, multi-model audit of open-weight LLMs that treats job-posting language as the primary experimental variable, evaluating six models (Llama 3.2, Mistral, Gemma 3, Qwen 3, Phi 3, DeepSeek-R1) across four controlled experiments that jointly probe recruiter-simulation and job-seeker-simulation tasks. We find that (1) agentic posting language depresses recruiter recommendation scores for female candidates (r_rb = 0.309, p_Bonf = 7x10^-5; model-fixed-effects r_rb = 0.448), while communal language partially reverses the penalty; and (2) coded-exclusion language suppresses non-White recruiter scores at large effect sizes (r_rb = 0.646-0.758) and, on the job-seeker side, selectively deters non-White personas from expressing interest -- operationalizing a chilling-effect mechanism at scale. A label-ablation experiment isolates the explicit demographic persona label as the primary causal driver, and Word Embedding Association Tests corroborate these findings at the representational level (d = 1.01-1.45 under Caliskan et al.'s multi-word gender attribute lists). We translate these results into a concrete pre-deployment audit protocol -- posting-vocabulary scoring, persona-conditioned LLM probing, and adverse-impact flagging against the four-fifths threshold -- that operationalizes the documentation and risk-management obligations Annex III imposes on high-risk AI in recruitment.

Summary

Main Finding

Open-weight LLMs used in hiring pipelines exhibit systematic, large biases triggered by job-posting language. Agentic (masculine/achievement) vocabulary reduces recruiter recommendation scores for female candidates, while communal language mitigates that penalty. Coded-exclusion (cultural-match) language in postings sharply suppresses recruiter scores for non-White candidates and also reduces non-White personas’ self-reported interest—producing both evaluator-side and applicant-side (chilling) effects. These behavioral patterns are causally linked to explicit demographic labels and are reflected in model embeddings. The authors propose a pre-deployment audit protocol (posting-vocabulary scoring + persona-conditioned probing + four-fifths adverse-impact flagging) to operationalize EU AI Act Annex III and EEOC requirements.

Key Points

  • Two primary linguistic triggers:
    • Agentic vs. communal vocabulary (gender): Agentic postings depress female candidate recommendations (rrb = 0.309, pBonf = 7×10−5; model-fixed-effects rrb = 0.448). Communal language partially reverses the penalty.
    • Inclusive vs. coded-exclusion vocabulary (race): Coded-exclusion language strongly reduces recruiter scores for non-White candidates (rrb = 0.646–0.758) and lowers non-White job-seeker interest—evidence of a chilling effect.
  • Causal and representational evidence:
    • Label-ablation shows explicit demographic persona labels are the primary causal driver of differential outputs.
    • WEAT on model representations corroborates gendered associations (d = 1.01–1.45 using Caliskan et al. multi-word lists).
  • Dual-perspective design: experiments simulate both recruiter (outbound job-matching/recommender) and job-seeker perspectives, revealing asymmetric effects on evaluator outputs and applicant interest.
  • Practical outcome: a concrete, deployment-oriented audit protocol mapping directly onto EU AI Act Annex III documentation obligations and the EEOC four-fifths adverse-impact rule.

Data & Methods

  • Models audited: six open-weight LLMs run deterministically (temperature = 0 where possible) via Ollama:
    • Llama 3.2 (3B), Mistral 7B v0.3, Gemma 3 4B, Qwen 3 8B, Phi 3 3.8B Mini, DeepSeek-R1 7B distill.
  • Stimuli:
    • Gender experiments: 40 synthetic job postings (agentic vs. communal), with job titles co-varied to reflect ecological co-occurrence of titles and wording.
    • Race experiments: 20 matched job-posting pairs with identical titles but bodies varying between inclusive vs. coded-exclusion (cultural-matching) registers.
  • Roles and tasks:
    • Recruitment-agent role: model acts as recruiter and assigns recommendation strength (1–10) for a candidate profile; probes outbound matching/exportable recommender behavior.
    • Job-seeker role: model acts as job-seeker with specified demographic label and rates its own interest (1–10).
  • Candidate attributes:
    • Gender: Male vs Female (other qualifications held constant).
    • Race/ethnicity: White, Black/African American, Asian American, Hispanic American.
  • Output controls:
    • Structured JSON schema enforced to standardize outputs.
  • Statistical analysis:
    • Mann–Whitney U tests (two-sided) with Bonferroni correction within experiment families.
    • Effect sizes reported as rank-biserial correlations (rrb); group comparisons also report mean score differences.
    • Model-fixed-effects analyses: mean-centering within model to account for score calibration differences; pooling across models yields larger n (e.g., n=120 per group for gender pooled).
  • Embedding validation:
    • Word Embedding Association Test (WEAT) using Caliskan et al.’s multi-word male/female attribute lists and the agentic/communal target sets.
  • Key robustness/ablation checks:
    • Description-body-only vs. co-varied title+body ablations to isolate components.
    • Label-ablation to isolate the role of explicit demographic labels.

Implications for AI Economics

  • Labor-market efficiency and distributional effects:
    • Linguistic-triggered evaluator bias plus chilling effects on applicant interest can reduce the quantity and alter the composition of applicant pools for given jobs. This produces allocative inefficiencies (missed matches), reinforces occupational segregation, and may depress wages or career mobility for affected groups.
    • Small changes in posting vocabulary—cheap for employers—can generate systematic frictions that aggregate into sizable labor-market distortions when LLMs are widely deployed.
  • Regulatory and compliance economics:
    • The paper’s audit protocol operationalizes regulatory obligations under the EU AI Act (Annex III) and EEOC adverse-impact analysis. Compliance will impose measurable costs on HR tech vendors and employers (pre-deployment audits, logging, remediation) but also creates a demand for audit tooling and certification services.
    • The four-fifths rule as an operational threshold gives a binary, monetizable compliance metric (flag/no-flag) attractive to regulators, insurers, and compliance vendors.
  • Market and product incentives:
    • HR platforms and LLM providers face incentives to reduce posting-triggered biases (product differentiation). Firms that certify low adverse-impact scores can gain market share; conversely, firms failing audits risk legal exposure, reputational costs, and possible loss of talent.
    • New markets are likely to emerge: posting-vocabulary scorers, persona-conditioned audit-as-a-service, and liability insurance priced on audit outcomes.
  • Externalities and social costs:
    • Chilling effects (reduced application likelihood from disadvantaged groups) create negative externalities not internalized by individual employers, justifying regulation or industry-standard audits.
    • Persistent representational biases in widespread models can amplify historical inequality; economic models of technology adoption should account for distributional impacts, not just aggregate productivity gains.
  • Measurement and policy implications:
    • The study suggests concrete, operational metrics (rrb effect sizes; four-fifths selection-rate comparisons; WEAT effect sizes) that economists and policymakers can incorporate into empirical analyses of AI-driven hiring.
    • Cost–benefit assessments of AI in recruitment must include remediation and monitoring costs plus potential litigation and labor-market welfare effects.
  • Suggestions for economic modeling and further work:
    • Quantify macro impacts: estimate how given effect sizes (e.g., rrb ranges observed) would change application rates, hiring probabilities, and downstream earnings distribution across occupations at scale.
    • Calibration of regulatory thresholds: analyze trade-offs (type I vs type II errors) when using four-fifths or other statistical thresholds to trigger remediation.
    • Market structure analysis: assess competition among HR vendors when audit compliance becomes a differentiator; study whether incumbent platforms internalize biases less or more readily than newer entrants.

Limitations to consider when applying these results: stimuli were synthetic and designed to isolate linguistic registers; audited models were small-to-medium open-weight models (not all-production-grade proprietary models); demographics were provided as explicit labels (real-world implicit cues may interact differently). Nonetheless, the experiments produce clear, deployment-relevant signals that posting vocabulary materially affects LLM-mediated hiring outcomes and that pre-deployment linguistic audits are both feasible and economically consequential.

Assessment

Paper Typequasi_experimental Evidence Strengthmedium — The paper uses well-controlled, multi-model experiments, ablation, and embedding-level validation that produce consistent, sizable effects; however, external validity is limited by synthetic stimuli, a modest number of postings (40/20 pairs), explicit persona labels (which may overstate effects relative to implicit cues), and a set of relatively small open-weight models that may not generalize to larger proprietary systems or deployed pipelines. Methods Rigormedium — Strengths: multi-model design, deterministic structured outputs, pre-registered-like controlled stimuli, ablation and model-fixed-effects, nonparametric tests, and WEAT triangulation; Weaknesses: stimuli partly LLM-generated and author-edited (possible construction bias), modest sample of postings, limited model family (3–8B open models), no human-ground-truth validation of scores, and single API/platform (Ollama). SampleSix open-weight LLMs run via Ollama (Llama 3.2 3B, Mistral 7B v0.3, Gemma 3 4B, Qwen 3 8B, Phi 3 3.8B Mini, DeepSeek-R1 7B distill). Gender experiments: 40 synthetic job postings categorized agentic vs. communal; candidate gender Male/Female; roles: recruiter (recommendation strength 1–10) and job-seeker (interest 1–10); n=20 job-posting observations per cell, pooled across six models yielding n=120 per group when pooled. Race experiments: 20 job-posting pairs (inclusive vs. coded-exclusion) with four race/ethnicity personas (White, Black/African American, Asian American, Hispanic American); same two roles and sampling. Analyses include label-ablation, model-fixed-effects (mean-centering), Mann–Whitney U tests with Bonferroni correction, effect sizes reported as rank-biserial correlations; WEAT applied to embedding representations using multi-word attribute lists. Themeslabor_markets governance IdentificationControlled manipulation of job-posting language (agentic vs. communal; inclusive vs. coded-exclusion) crossed with explicit synthetic candidate persona labels (gender or race/ethnicity); role-based probes (recruiter recommendation and job-seeker interest) collected deterministically (temperature=0) across six open-weight LLMs; label-ablation isolates the effect of explicit demographic labels; model-fixed-effects (mean-centering by model) control for cross-model score calibration; representational validation via Word Embedding Association Tests (WEAT); group differences tested with Mann–Whitney U and Bonferroni correction. GeneralizabilityFindings drawn from 3–8B open-weight models may not generalize to larger proprietary LLMs (e.g., GPT-4o, Gemini Ultra) or vendor fine-tuned systems., Synthetic job postings (LLM-drafted then author-edited) may not capture full complexity or distribution of real-world job ads., Explicit demographic labels in prompts may exaggerate discrimination relative to realistic, more implicit candidate signals., Deterministic, temperature=0 querying and constrained JSON outputs differ from many production settings that use stochastic sampling or human-in-the-loop calibration., Cultural and jurisdictional differences — stimuli and race categories are U.S.-centric; results may not transfer to non-U.S. labor markets or non-English job postings., Evaluation uses model-generated 1–10 scores rather than human-validated outcomes (hire/interview), limiting behavioral external validity.

Claims (7)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Agentic job-posting language depresses recruiter recommendation scores for female candidates. Hiring negative Recruiter recommendation strength on a 1–10 scale
Reading fidelity high
Study strength medium
n=240
rrb = 0.309, pBonf = 7×10−5; model-fixed-effects rrb = 0.448
0.48
Communal job-posting language partially reverses the recruiter recommendation penalty observed for female candidates. Hiring positive Recruiter recommendation strength on a 1–10 scale
Reading fidelity high
Study strength medium
n=240
0.48
Coded-exclusion language suppresses recruiter recommendation scores for non-White candidates. Hiring negative Recruiter recommendation strength on a 1–10 scale
Reading fidelity high
Study strength medium
n=480
rrb = 0.646–0.758
0.48
Coded-exclusion language selectively deters non-White job-seeker personas from expressing interest in job postings. Hiring negative Job-seeker interest score on a 1–10 scale
Reading fidelity high
Study strength medium
n=480
0.48
The explicit demographic persona label is the primary causal driver of the observed bias in the label-ablation experiment. Ai Safety And Ethics mixed Differences in recruiter or job-seeker model scores under demographic-label ablation
Reading fidelity high
Study strength low
not reported
0.24
Word Embedding Association Tests corroborate the gender-bias findings at the representational level. Ai Safety And Ethics positive Association between gendered attribute terms and agentic versus communal job-posting representations
Reading fidelity high
Study strength medium
n=40
d = 1.01–1.45
0.48
The study proposes a pre-deployment audit protocol combining posting-vocabulary scoring, persona-conditioned LLM probing, and adverse-impact flagging against the four-fifths threshold for recruitment systems. Governance And Regulation positive Pre-deployment identification and documentation of potential demographic disparities in AI-assisted recruitment
Reading fidelity high
Study strength speculative
not reported
0.08

Notes