0 cumulative citations
View corpus contextOpen-weight LLMs used for recruitment respond to job-ad language in ways that disadvantage protected groups: agentic wording substantially lowers recommendations for female personas (rrb ≈ 0.31), while coded-exclusion language markedly reduces both recruiter scores and expressed interest for non-White personas (rrb ≈ 0.65–0.76). Label-ablation and embedding tests implicate explicit demographic labels and encoded representations, suggesting practical pre-deployment posting-language audits can flag adverse impact under regulatory thresholds.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
0 cumulative citations
View corpus contextOpen-weight large language models are rapidly entering hiring pipelines, yet their discriminatory failure modes -- and the regulatory exposure these create under the EU AI Act high-risk classification (Annex III) and U.S. EEOC adverse-impact analysis -- remain poorly understood. We present the first systematic, multi-model audit of open-weight LLMs that treats job-posting language as the primary experimental variable, evaluating six models (Llama 3.2, Mistral, Gemma 3, Qwen 3, Phi 3, DeepSeek-R1) across four controlled experiments that jointly probe recruiter-simulation and job-seeker-simulation tasks. We find that (1) agentic posting language depresses recruiter recommendation scores for female candidates (r_rb = 0.309, p_Bonf = 7x10^-5; model-fixed-effects r_rb = 0.448), while communal language partially reverses the penalty; and (2) coded-exclusion language suppresses non-White recruiter scores at large effect sizes (r_rb = 0.646-0.758) and, on the job-seeker side, selectively deters non-White personas from expressing interest -- operationalizing a chilling-effect mechanism at scale. A label-ablation experiment isolates the explicit demographic persona label as the primary causal driver, and Word Embedding Association Tests corroborate these findings at the representational level (d = 1.01-1.45 under Caliskan et al.'s multi-word gender attribute lists). We translate these results into a concrete pre-deployment audit protocol -- posting-vocabulary scoring, persona-conditioned LLM probing, and adverse-impact flagging against the four-fifths threshold -- that operationalizes the documentation and risk-management obligations Annex III imposes on high-risk AI in recruitment.
Summary
Main Finding
Open-weight LLMs used in hiring pipelines exhibit systematic, large biases triggered by job-posting language. Agentic (masculine/achievement) vocabulary reduces recruiter recommendation scores for female candidates, while communal language mitigates that penalty. Coded-exclusion (cultural-match) language in postings sharply suppresses recruiter scores for non-White candidates and also reduces non-White personas’ self-reported interest—producing both evaluator-side and applicant-side (chilling) effects. These behavioral patterns are causally linked to explicit demographic labels and are reflected in model embeddings. The authors propose a pre-deployment audit protocol (posting-vocabulary scoring + persona-conditioned probing + four-fifths adverse-impact flagging) to operationalize EU AI Act Annex III and EEOC requirements.
Key Points
- Two primary linguistic triggers:
- Agentic vs. communal vocabulary (gender): Agentic postings depress female candidate recommendations (rrb = 0.309, pBonf = 7×10−5; model-fixed-effects rrb = 0.448). Communal language partially reverses the penalty.
- Inclusive vs. coded-exclusion vocabulary (race): Coded-exclusion language strongly reduces recruiter scores for non-White candidates (rrb = 0.646–0.758) and lowers non-White job-seeker interest—evidence of a chilling effect.
- Causal and representational evidence:
- Label-ablation shows explicit demographic persona labels are the primary causal driver of differential outputs.
- WEAT on model representations corroborates gendered associations (d = 1.01–1.45 using Caliskan et al. multi-word lists).
- Dual-perspective design: experiments simulate both recruiter (outbound job-matching/recommender) and job-seeker perspectives, revealing asymmetric effects on evaluator outputs and applicant interest.
- Practical outcome: a concrete, deployment-oriented audit protocol mapping directly onto EU AI Act Annex III documentation obligations and the EEOC four-fifths adverse-impact rule.
Data & Methods
- Models audited: six open-weight LLMs run deterministically (temperature = 0 where possible) via Ollama:
- Llama 3.2 (3B), Mistral 7B v0.3, Gemma 3 4B, Qwen 3 8B, Phi 3 3.8B Mini, DeepSeek-R1 7B distill.
- Stimuli:
- Gender experiments: 40 synthetic job postings (agentic vs. communal), with job titles co-varied to reflect ecological co-occurrence of titles and wording.
- Race experiments: 20 matched job-posting pairs with identical titles but bodies varying between inclusive vs. coded-exclusion (cultural-matching) registers.
- Roles and tasks:
- Recruitment-agent role: model acts as recruiter and assigns recommendation strength (1–10) for a candidate profile; probes outbound matching/exportable recommender behavior.
- Job-seeker role: model acts as job-seeker with specified demographic label and rates its own interest (1–10).
- Candidate attributes:
- Gender: Male vs Female (other qualifications held constant).
- Race/ethnicity: White, Black/African American, Asian American, Hispanic American.
- Output controls:
- Structured JSON schema enforced to standardize outputs.
- Statistical analysis:
- Mann–Whitney U tests (two-sided) with Bonferroni correction within experiment families.
- Effect sizes reported as rank-biserial correlations (rrb); group comparisons also report mean score differences.
- Model-fixed-effects analyses: mean-centering within model to account for score calibration differences; pooling across models yields larger n (e.g., n=120 per group for gender pooled).
- Embedding validation:
- Word Embedding Association Test (WEAT) using Caliskan et al.’s multi-word male/female attribute lists and the agentic/communal target sets.
- Key robustness/ablation checks:
- Description-body-only vs. co-varied title+body ablations to isolate components.
- Label-ablation to isolate the role of explicit demographic labels.
Implications for AI Economics
- Labor-market efficiency and distributional effects:
- Linguistic-triggered evaluator bias plus chilling effects on applicant interest can reduce the quantity and alter the composition of applicant pools for given jobs. This produces allocative inefficiencies (missed matches), reinforces occupational segregation, and may depress wages or career mobility for affected groups.
- Small changes in posting vocabulary—cheap for employers—can generate systematic frictions that aggregate into sizable labor-market distortions when LLMs are widely deployed.
- Regulatory and compliance economics:
- The paper’s audit protocol operationalizes regulatory obligations under the EU AI Act (Annex III) and EEOC adverse-impact analysis. Compliance will impose measurable costs on HR tech vendors and employers (pre-deployment audits, logging, remediation) but also creates a demand for audit tooling and certification services.
- The four-fifths rule as an operational threshold gives a binary, monetizable compliance metric (flag/no-flag) attractive to regulators, insurers, and compliance vendors.
- Market and product incentives:
- HR platforms and LLM providers face incentives to reduce posting-triggered biases (product differentiation). Firms that certify low adverse-impact scores can gain market share; conversely, firms failing audits risk legal exposure, reputational costs, and possible loss of talent.
- New markets are likely to emerge: posting-vocabulary scorers, persona-conditioned audit-as-a-service, and liability insurance priced on audit outcomes.
- Externalities and social costs:
- Chilling effects (reduced application likelihood from disadvantaged groups) create negative externalities not internalized by individual employers, justifying regulation or industry-standard audits.
- Persistent representational biases in widespread models can amplify historical inequality; economic models of technology adoption should account for distributional impacts, not just aggregate productivity gains.
- Measurement and policy implications:
- The study suggests concrete, operational metrics (rrb effect sizes; four-fifths selection-rate comparisons; WEAT effect sizes) that economists and policymakers can incorporate into empirical analyses of AI-driven hiring.
- Cost–benefit assessments of AI in recruitment must include remediation and monitoring costs plus potential litigation and labor-market welfare effects.
- Suggestions for economic modeling and further work:
- Quantify macro impacts: estimate how given effect sizes (e.g., rrb ranges observed) would change application rates, hiring probabilities, and downstream earnings distribution across occupations at scale.
- Calibration of regulatory thresholds: analyze trade-offs (type I vs type II errors) when using four-fifths or other statistical thresholds to trigger remediation.
- Market structure analysis: assess competition among HR vendors when audit compliance becomes a differentiator; study whether incumbent platforms internalize biases less or more readily than newer entrants.
Limitations to consider when applying these results: stimuli were synthetic and designed to isolate linguistic registers; audited models were small-to-medium open-weight models (not all-production-grade proprietary models); demographics were provided as explicit labels (real-world implicit cues may interact differently). Nonetheless, the experiments produce clear, deployment-relevant signals that posting vocabulary materially affects LLM-mediated hiring outcomes and that pre-deployment linguistic audits are both feasible and economically consequential.
Assessment
Claims (7)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Agentic job-posting language depresses recruiter recommendation scores for female candidates. Hiring | negative | Recruiter recommendation strength on a 1–10 scale |
Reading fidelity
high
Study strength
medium
|
n=240
rrb = 0.309, pBonf = 7×10−5; model-fixed-effects rrb = 0.448
|
| Communal job-posting language partially reverses the recruiter recommendation penalty observed for female candidates. Hiring | positive | Recruiter recommendation strength on a 1–10 scale |
Reading fidelity
high
Study strength
medium
|
n=240
|
| Coded-exclusion language suppresses recruiter recommendation scores for non-White candidates. Hiring | negative | Recruiter recommendation strength on a 1–10 scale |
Reading fidelity
high
Study strength
medium
|
n=480
rrb = 0.646–0.758
|
| Coded-exclusion language selectively deters non-White job-seeker personas from expressing interest in job postings. Hiring | negative | Job-seeker interest score on a 1–10 scale |
Reading fidelity
high
Study strength
medium
|
n=480
|
| The explicit demographic persona label is the primary causal driver of the observed bias in the label-ablation experiment. Ai Safety And Ethics | mixed | Differences in recruiter or job-seeker model scores under demographic-label ablation |
Reading fidelity
high
Study strength
low
|
not reported
|
| Word Embedding Association Tests corroborate the gender-bias findings at the representational level. Ai Safety And Ethics | positive | Association between gendered attribute terms and agentic versus communal job-posting representations |
Reading fidelity
high
Study strength
medium
|
n=40
d = 1.01–1.45
|
| The study proposes a pre-deployment audit protocol combining posting-vocabulary scoring, persona-conditioned LLM probing, and adverse-impact flagging against the four-fifths threshold for recruitment systems. Governance And Regulation | positive | Pre-deployment identification and documentation of potential demographic disparities in AI-assisted recruitment |
Reading fidelity
high
Study strength
speculative
|
not reported
|