1 cumulative citations
View corpus contextA large language model ranks freelancer–project fit largely by skills and experience but also reads extra signals into profiles; while average bias against minority groups is small, intersectional differences mean productivity cues are weighted differently across demographic subgroups.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
General-purpose Large Language Models (LLMs) show significant potential in recruitment applications, where decisions require reasoning over unstructured text, balancing multiple criteria, and inferring fit and competence from indirect productivity signals. Yet, it is still uncertain how LLMs assign importance to each attribute and whether such assignments are in line with economic principles, recruiter preferences or broader societal norms. We propose a framework to evaluate an LLM's decision logic in recruitment, by drawing on established economic methodologies for analyzing human hiring behavior. We build synthetic datasets from real freelancer profiles and project descriptions from a major European online freelance marketplace and apply a full factorial design to estimate how a LLM weighs different match-relevant criteria when evaluating freelancer-project fit. We identify which attributes the LLM prioritizes and analyze how these weights vary across project contexts and demographic subgroups. Finally, we explain how a comparable experimental setup could be implemented with human recruiters to assess alignment between model and human decisions. Our findings reveal that the LLM weighs core productivity signals, such as skills and experience, but interprets certain features beyond their explicit matching value. While showing minimal average discrimination against minority groups, intersectional effects reveal that productivity signals carry different weights between demographic groups.
Summary
Main Finding
A large general-purpose LLM (Gemini 2.0 Flash), when placed in the role of a recruiter and evaluated on a fully factorial set of synthetic freelancer profiles and project briefs from a European freelance platform, largely behaves like an economic hiring evaluator: it places strongest weight on productivity signals (experience, reputation, skill match) and much weaker average weight on socio‑demographic attributes. However, intersectional heterogeneity shows the model uses demographic cues to modulate how it interprets productivity signals, creating subtle channels for unequal treatment even when average discrimination is small.
Key Points
- Research questions addressed:
- RQ1: Which profile/job attributes does the LLM emphasize or ignore?
- RQ2: Does the LLM’s decision logic vary across socio‑demographic groups?
- RQ3: How to compare LLM and human recruiter decision logic experimentally?
- Experimental scale and outputs:
- 10,800 synthetic freelancer profiles × 16 synthetic briefs → 172,800 profile–brief pairs.
- Gemini 2.0 Flash provided hiring probabilities on a 10‑point scale; each pair scored 3 times and averaged.
- Score distribution: mean ≈ 6.5, SD ≈ 1.3, observed range ~3–9.8; only 2.48% of pairs varied across reruns (max ±1 point).
- Attribute importance (OLS summary of implicit weights; coefficients interpreted as causal marginal effects due to fully randomized factorial design):
- Largest penalties/rewards: Experience (largest effect magnitude, e.g., underqualified penalized ≈ −2.25 pts), Reputation (≈ −1.19 pts for low reputation), Skill match (≈ −0.69 pts for distant mismatch).
- Other notable effects: Daily rate (≈ −0.52), Remote preference (≈ −0.50). Part‑time, firm size, industry, education had small effects.
- Socio‑demographic attributes (perceived gender/ethnicity, education) had minimal average direct effects (ranked low), but heterogeneity analyses show group‑dependent weighting of productivity signals (intersectional effects).
- Fairness and alignment:
- On average, minimal direct discrimination by gender/ethnicity in the model’s scores.
- Intersectional differences: the same productivity signal (e.g., experience, reputation) can affect scores differently depending on perceived demographic attributes — a subtler, context‑dependent source of unequal outcomes.
- Methodological contribution:
- A fully factorial synthetic design adapted from correspondence studies enables causal attribution of LLM decision weights and can be reused to compare LLMs with human recruiters using identical stimuli.
Data & Methods
- Context and scope:
- Data and attribute distributions are drawn from a large European freelance marketplace.
- Primary occupational focus: full‑stack developers in France (high male name prevalence); robustness checks on SEO content writers reported in appendix.
- Synthetic data construction:
- Full factorial design for profiles and briefs to break correlations and enable causal identification.
- Profiles (10,800): varied on skills (exact / close substitute / distant), experience (1 / 5 / 9 years), work arrangement preferences (remote/on‑site, full/part‑time), reputation (projects 0/1/5; rating 0/5; badge yes/no), daily rate (300–500€), prior employer (SME vs large; e‑commerce vs banking), and socio‑demographics signaled by first names (European male, European female, Arabic male); education (BSc vs MSc).
- Briefs (16): varied recruiter gender (name), firm size (SME vs large), remote vs onsite, full vs part time; some dimensions held constant (e.g., required duration, tech stack, required experience).
- LLM elicitation:
- Model: Gemini 2.0 Flash, prompts in French.
- Prompt asked for a hiring probability (10‑point scale) and a brief explanation of reasoning; no fairness correction instruction.
- Each profile–brief pair scored three times; mean used.
- Statistical analysis:
- OLS regressions of average LLM score on binary indicators for attribute levels; standard errors clustered at brief level.
- Given orthogonalized design, OLS coefficients interpreted as marginal causal effects.
- Reported model summary: R^2 ≈ 0.90 — linear additive approximation explains most score variation in this experimental context.
- Limitations noted by authors:
- Full factorial design trades external representativeness for causal identification.
- Prompt framing (hiring probability) and profile formatting may influence model reasoning; results may depend on prompt, language, or model family.
- Single LLM and occupational focus limit generalizability; real‑world downstream outcomes not observed.
Implications for AI Economics
- Theoretical alignment: LLMs can internalize economic signaling logic. In this experiment the LLM rewarded close matches and trust signals (experience, reputation) in ways consistent with labor‑economics models of signaling and productivity inference.
- Measurement & auditing tools: Fully factorial synthetic designs adapted from correspondence studies provide a rigorous, causal way to estimate implicit attribute weights in LLM decision‑making and to surface interaction (intersectional) effects that simple attribution methods might miss.
- Policy and platform implications:
- Even when average demographic effects are small, intersectional heterogeneity can produce unequal outcomes; audits should measure both average and subgroup‑conditional effects.
- Platforms using LLMs for candidate ranking or screening should monitor how the model reweights productivity signals across groups and consider calibration, counterfactual testing, or constraint mechanisms to prevent unintended disparate impacts.
- Regulatory auditing frameworks could adopt similar experimental stimuli to evaluate vendor claims about fairness and alignment.
- Research directions:
- Direct human–AI comparison: deploy the identical synthetic stimuli to human recruiters to quantify alignment/misalignment and the welfare implications of replacing or assisting human screening with LLMs.
- External validity tests: evaluate other LLMs, prompt framings, languages, and occupational contexts; link model scores to downstream hiring outcomes to measure market‑level effects.
- Dynamic and systemic effects: study how LLM recommendations change demand‑side behavior (which candidates are contacted) and feedback loops that could amplify inequalities over time.
- Practical takeaway for economists and practitioners: LLMs can be powerful, interpretable evaluators of unstructured candidate information, but rigorous causal audits (including intersectional analyses) are necessary before large‑scale deployment in hiring to detect subtle disparate‑impact mechanisms that simple averages hide.
Assessment
Claims (9)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| General-purpose Large Language Models (LLMs) show significant potential in recruitment applications, where decisions require reasoning over unstructured text, balancing multiple criteria, and inferring fit and competence from indirect productivity signals. Adoption Rate | positive | potential suitability of LLMs for recruitment tasks (conceptual capability) |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| We propose a framework to evaluate an LLM's decision logic in recruitment by drawing on established economic methodologies for analyzing human hiring behavior. Research Productivity | positive | existence of an evaluation framework for LLM decision logic in recruitment |
Reading fidelity
high
Study strength
medium
|
not reported
|
| We build synthetic datasets from real freelancer profiles and project descriptions from a major European online freelance marketplace. Research Productivity | positive | availability of synthetic dataset derived from marketplace data |
Reading fidelity
high
Study strength
medium
|
not reported
|
| We apply a full factorial design to estimate how a LLM weighs different match-relevant criteria when evaluating freelancer-project fit. Research Productivity | positive | LLM-assigned weights to match-relevant criteria (estimated via factorial design) |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The LLM prioritizes core productivity signals, such as skills and experience, when evaluating freelancer-project fit. Task Allocation | positive | relative importance (weights) assigned by the LLM to 'skills' and 'experience' attributes |
Reading fidelity
medium
Study strength
medium
|
not reported
|
| The LLM interprets certain features beyond their explicit matching value. Task Allocation | mixed | interpretation/use of profile/project features by the LLM beyond explicit matching signals |
Reading fidelity
medium
Study strength
medium
|
not reported
|
| The LLM shows minimal average discrimination against minority groups. Hiring | null_result | average difference in LLM evaluations across demographic (minority) groups |
Reading fidelity
medium
Study strength
medium
|
not reported
|
| Intersectional effects reveal that productivity signals carry different weights between demographic groups. Hiring | mixed | variation in attribute weights (productivity signals) across intersectional demographic subgroups |
Reading fidelity
medium
Study strength
medium
|
not reported
|
| We explain how a comparable experimental setup could be implemented with human recruiters to assess alignment between model and human decisions. Research Productivity | positive | feasibility and design of comparable human-recruiter experiments |
Reading fidelity
high
Study strength
medium
|
not reported
|