0 cumulative citations
View corpus contextAn audit of Barcelona’s public employment platform shows apparent gender parity masks deep exclusions: women face lower shortlisting rates in mid-salary roles and non-binary applicants are shortlisted at under one-third the rate of men, while candidates over 55 are effectively absent from the pipeline.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Algorithmic fairness evaluation commonly assesses AI systems as bounded technical components, abstracting away the organizational context in which they operate. We present, to our knowledge, the first independent end-to-end fairness audit of a semi-automated hiring system operated by Barcelona Activa, a public employment agency using the third-party TalentClue platform for candidate search and shortlisting. We analyze approximately 497,000 candidate-vacancy pipeline entries from September 2017 to September 2022, covering seven pipeline stages that span automated processing, human discretion, candidate data, and employer decisions. Aggregate outcomes across binary genders are statistically indistinguishable, yet this parity masks substantial disparities by salary level, age, and gender identity. Women face adverse impact in mid-salary shortlisting (DIR = 0.786, p < 0.001), alongside salary disparities in 15 of 20 sectors and a compounded disadvantage for women aged 46-55 (DIR = 0.77). Non-binary candidates are shortlisted at less than one third the rate of men (DIR = 0.295), although this estimate rests on a small sample (N = 285). Candidates aged 55 and over are entirely absent from the pipeline despite comprising 15.6% of Barcelona's labor force. The gender gap in shortlisting narrows over time, from 6.5 percentage points in 2017 to 1.3 in 2022. The audit further reveals a vendor-deployer information asymmetry: Barcelona Activa lacks access to key information about TalentClue's matching logic and evaluation. Fairness outcomes can thus arise from interactions among automated processing, human discretion, data quality, vendor opacity, and pipeline structure. We build on prior calls for sociotechnical, end-to-end fairness evaluation, showing empirically why model-level assessment alone can be insufficient for understanding fairness in deployed systems.
Summary
Main Finding
An independent end-to-end audit of Barcelona Activa’s semi-automated hiring pipeline (using the TalentClue platform) over ~497,000 candidate–vacancy entries (Sep 2017–Sep 2022) shows that apparent aggregate gender parity conceals meaningful, persistent disparities emerging from interactions across multiple pipeline stages (automated matching, analyst filtering, and employer decisions). Disparities are concentrated by salary band, age, gender identity, and sector, and they change over time. Vendor opacity and missing operational data constrain deployer governance and auditing.
Key Points
-
Scope and dataset
- 497,000 candidate–vacancy pipeline entries (Sep 2017–Sep 2022).
- Seven-stage pipeline audited: vacancy receipt → analyst assignment → keyword/filters → platform search → filtered results → analyst pre-selection → employer handoff.
- Critical missing fields: 24.0% missing gender; 14.5% missing country of origin; no candidate qualification/experience fields; no vendor matching logic; no systematic post-shortlist outcome tracking.
-
Aggregate vs disaggregated outcomes
- Aggregate shortlisting outcomes across binary genders were statistically indistinguishable, but disaggregation reveals substantive inequalities.
- Women: adverse impact in mid-salary shortlisting (DIR = 0.786, p < 0.001); compounded disadvantage for women aged 46–55 (DIR = 0.77).
- Non-binary/other: substantially lower shortlisting (DIR = 0.295 vs men), but estimate limited by small sample (N = 285).
- Age exclusion: candidates 55+ were effectively absent from the pipeline despite representing 15.6% of Barcelona’s labor force.
- Sectoral/salary disparities: persistent salary gaps across 15 of 20 sectors.
-
Temporal dynamics
- Gender shortlisting gap narrowed over time: 6.5 percentage points in 2017 → 1.3 percentage points in 2022, demonstrating changing fairness properties and the need for ongoing monitoring.
-
Governance and transparency failures
- Vendor–deployer information asymmetry: Barcelona Activa lacked access to TalentClue’s internal matching/ranking logic and any vendor evaluation artifacts.
- Operational process gaps: no audit trail for keyword/filter selection by analysts; no structured review criteria in manual steps; no downstream feedback loop from employers.
Data & Methods
-
Audit framework
- Applied Eticas’ post-deployment algorithmic impact assessment and AI Risk Taxonomy; scope limited to Bias & Fairness (other risks noted but not quantitatively evaluated).
- End-to-end, sociotechnical framing: analyzed interactions across automated components, human discretion, and organizational procedures.
-
Representativeness benchmarking
- Compared candidate pool composition to Barcelona active labor force using Spanish Labor Force Survey (EPA) data (2019–2021), disaggregated by gender, age, and origin.
-
Quantitative measures
- Primary metric: Disparate Impact Ratio (DIR = selection rate_protected / selection rate_reference); 0.80 threshold used as a practitioner benchmark (EEOC guideline).
- Statistical testing: chi-square or Fisher’s exact tests at α = 0.05.
- Stratified analyses: by sector, salary band, contract type, education level, and age group to detect conditional disparities.
- Intersectional analysis: sex × age, sex × education, sex × origin where subgroup sizes permitted reliable inference.
- Temporal analysis: annual shortlisting rates (2017–2022) to assess stability/trends.
-
Key limitations
- Missing and coarse data (no skills/experience), substantial missingness in key demographics, small-N subgroups (non-binary), and inability to inspect vendor internals or the excluded candidate pool returned by the platform.
- No employer-level post-shortlist outcomes to trace final hiring disparities.
Implications for AI Economics
-
Measurement and misallocation risks
- Pipelines with opaque automated components and undocumented human steps can produce hidden selection distortions that misallocate labor—both excluding qualified workers (notably older workers and non-binary candidates) and biasing the effective supply presented to employers.
- Aggregate parity measures can mask redistributive effects across wages, sectors, and age cohorts; economic evaluations of AI in labor markets must use disaggregated, intersectional metrics.
-
Market structure and externalities
- Vendor opacity creates informational externalities: deployers (public agencies) cannot validate match quality or fairness, impairing competition on fairness and increasing regulatory and monitoring costs.
- Small public purchasers (local agencies) are particularly exposed because they lack leverage to demand transparency; this can lock in unequal access patterns across local labor markets.
-
Regulatory and compliance cost trade-offs
- EU AI Act (high-risk classification for employment systems) and analogous rules (e.g., NYC Local Law 144) will increase compliance burdens (data access, monitoring, documentation). The audit shows those obligations must extend beyond component-level model checks to pipeline- and organization-level controls to be effective.
- Continuous monitoring (rather than point-in-time audits) is economically justified: fairness outcomes here evolved materially over five years, implying that one-off audits understate ongoing risks.
-
Policy and procurement recommendations (economic levers)
- Contract clauses requiring vendor transparency on matching logic, evaluation metrics, and logs of excluded results; enforceable data-access provisions for deployers and independent auditors.
- Procure for auditable pipelines: require instrumentation of analyst decisions (keywords, rationale, review protocols) and post-shortlist outcome reporting from employers to enable closed-loop accountability.
- Subsidize or mandate continuous, disaggregated monitoring (by salary band, age, gender identity, sector) for public-sector deployments to internalize governance externalities.
- Adjust economic models that evaluate AI adoption in HR to include expected costs of governance, data collection, and remediation, and to value avoided labor-market exclusion externalities (e.g., lost income, reduced labor force participation for affected groups).
-
Research and modeling implications
- Economists modeling AI impacts on employment should represent hiring as a multi-stage pipeline with compounded selection probabilities, rather than a single-stage probabilistic filter.
- Evaluate distributional impacts (wage and sectoral composition) and dynamic effects (how narrowing or widening gaps alter labor supply and incentives over time).
- Treat small-subgroup estimates (e.g., non-binary) carefully; there is both a statistical-power issue and an equity imperative to collect better demographic data.
Overall, the paper empirically demonstrates that fairness in hiring is a sociotechnical, dynamic problem with measurable economic consequences. Effective economic policy and procurement must account for pipeline structure, vendor transparency, continuous monitoring costs, and intersectional distributional impacts when evaluating or deploying AI-assisted hiring systems.
Assessment
Claims (10)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Aggregate outcomes across binary genders in the Barcelona Activa hiring pipeline are statistically indistinguishable. Hiring | null_result | Aggregate gender differences in pipeline outcomes |
Reading fidelity
high
Study strength
medium
|
n=497000
|
| Women experience adverse impact in shortlisting for mid-salary vacancies, with a disparate impact ratio of 0.786. Hiring | negative | Shortlisting rate for women in mid-salary vacancies |
Reading fidelity
high
Study strength
high
|
n=497000
DIR = 0.786, p < 0.001
|
| Gender-related salary disparities persist across 15 of the 20 sectors examined. Hiring | mixed | Gender disparities in hiring-pipeline outcomes across sectors and salary levels |
Reading fidelity
high
Study strength
medium
|
n=497000
15 of 20 sectors
|
| Women aged 46–55 face compounded disadvantage in the hiring pipeline, with a disparate impact ratio of 0.77. Hiring | negative | Hiring-pipeline selection or shortlisting rate for women aged 46–55 |
Reading fidelity
high
Study strength
medium
|
n=497000
DIR = 0.77
|
| Non-binary candidates are shortlisted at less than one-third the rate of men. Hiring | negative | Shortlisting rate by gender identity |
Reading fidelity
high
Study strength
medium
|
n=285
DIR = 0.295
|
| Candidates aged 55 and over are entirely absent from the hiring pipeline, despite representing 15.6% of Barcelona's labor force. Hiring | negative | Representation and participation of older candidates in the hiring pipeline |
Reading fidelity
high
Study strength
medium
|
n=497000
15.6% of Barcelona's labor force
|
| The gender gap in shortlisting narrowed over time, from 6.5 percentage points in 2017 to 1.3 percentage points in 2022. Hiring | positive | Gender gap in annual shortlisting rates |
Reading fidelity
high
Study strength
medium
|
n=497000
from 6.5 percentage points in 2017 to 1.3 percentage points in 2022
|
| Barcelona Activa lacks access to key information about TalentClue's matching logic and evaluation, creating a vendor–deployer information asymmetry. Governance And Regulation | negative | Deployer ability to independently assess and govern the hiring system |
Reading fidelity
high
Study strength
high
|
not reported
|
| Model-level assessment alone can be insufficient for understanding fairness in deployed hiring systems because disparities can arise through interactions among automated processing, human discretion, data quality, vendor opacity, and pipeline structure. Ai Safety And Ethics | negative | Ability of component-level audits to identify fairness disparities in deployed hiring |
Reading fidelity
high
Study strength
medium
|
n=497000
|
| Point-in-time fairness assessments may miss changing fairness outcomes, supporting continuous monitoring of deployed hiring systems. Governance And Regulation | positive | Temporal stability of gender fairness outcomes |
Reading fidelity
high
Study strength
medium
|
n=497000
gender gap narrowed from 6.5 to 1.3 percentage points
|