The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

A large language model ranks freelancer–project fit largely by skills and experience but also reads extra signals into profiles; while average bias against minority groups is small, intersectional differences mean productivity cues are weighted differently across demographic subgroups.

Evaluating LLM Behavior in Hiring: Implicit Weights, Fairness Across Groups, and Alignment with Human Preferences
Morgane Hoffmann, Emma Jouffroy, Warren Jouanneau, Marc Palyart, Charles Pebereau · January 16, 2026
arxiv quasi_experimental medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Morgane Hoffmann unresolved corpus identity
  2. Emma Jouffroy unresolved corpus identity
  3. Warren Jouanneau unresolved corpus identity
  4. Marc Palyart unresolved corpus identity
  5. Charles Pebereau unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Morgane Hoffmann provider ID
  2. Emma Jouffroy provider ID
  3. Warren Jouanneau provider ID
  4. Marc Palyart provider ID
  5. Charles Pébereau provider ID
Using a full-factorial vignette experiment built from real marketplace profiles, the LLM places greatest weight on core productivity signals like skills and experience but also over-interprets peripheral features, producing minimal average demographic bias yet revealing intersectional differences in attribute weighting.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

General-purpose Large Language Models (LLMs) show significant potential in recruitment applications, where decisions require reasoning over unstructured text, balancing multiple criteria, and inferring fit and competence from indirect productivity signals. Yet, it is still uncertain how LLMs assign importance to each attribute and whether such assignments are in line with economic principles, recruiter preferences or broader societal norms. We propose a framework to evaluate an LLM's decision logic in recruitment, by drawing on established economic methodologies for analyzing human hiring behavior. We build synthetic datasets from real freelancer profiles and project descriptions from a major European online freelance marketplace and apply a full factorial design to estimate how a LLM weighs different match-relevant criteria when evaluating freelancer-project fit. We identify which attributes the LLM prioritizes and analyze how these weights vary across project contexts and demographic subgroups. Finally, we explain how a comparable experimental setup could be implemented with human recruiters to assess alignment between model and human decisions. Our findings reveal that the LLM weighs core productivity signals, such as skills and experience, but interprets certain features beyond their explicit matching value. While showing minimal average discrimination against minority groups, intersectional effects reveal that productivity signals carry different weights between demographic groups.

Summary

Main Finding

A large general-purpose LLM (Gemini 2.0 Flash), when placed in the role of a recruiter and evaluated on a fully factorial set of synthetic freelancer profiles and project briefs from a European freelance platform, largely behaves like an economic hiring evaluator: it places strongest weight on productivity signals (experience, reputation, skill match) and much weaker average weight on socio‑demographic attributes. However, intersectional heterogeneity shows the model uses demographic cues to modulate how it interprets productivity signals, creating subtle channels for unequal treatment even when average discrimination is small.

Key Points

  • Research questions addressed:
    • RQ1: Which profile/job attributes does the LLM emphasize or ignore?
    • RQ2: Does the LLM’s decision logic vary across socio‑demographic groups?
    • RQ3: How to compare LLM and human recruiter decision logic experimentally?
  • Experimental scale and outputs:
    • 10,800 synthetic freelancer profiles × 16 synthetic briefs → 172,800 profile–brief pairs.
    • Gemini 2.0 Flash provided hiring probabilities on a 10‑point scale; each pair scored 3 times and averaged.
    • Score distribution: mean ≈ 6.5, SD ≈ 1.3, observed range ~3–9.8; only 2.48% of pairs varied across reruns (max ±1 point).
  • Attribute importance (OLS summary of implicit weights; coefficients interpreted as causal marginal effects due to fully randomized factorial design):
    • Largest penalties/rewards: Experience (largest effect magnitude, e.g., underqualified penalized ≈ −2.25 pts), Reputation (≈ −1.19 pts for low reputation), Skill match (≈ −0.69 pts for distant mismatch).
    • Other notable effects: Daily rate (≈ −0.52), Remote preference (≈ −0.50). Part‑time, firm size, industry, education had small effects.
    • Socio‑demographic attributes (perceived gender/ethnicity, education) had minimal average direct effects (ranked low), but heterogeneity analyses show group‑dependent weighting of productivity signals (intersectional effects).
  • Fairness and alignment:
    • On average, minimal direct discrimination by gender/ethnicity in the model’s scores.
    • Intersectional differences: the same productivity signal (e.g., experience, reputation) can affect scores differently depending on perceived demographic attributes — a subtler, context‑dependent source of unequal outcomes.
  • Methodological contribution:
    • A fully factorial synthetic design adapted from correspondence studies enables causal attribution of LLM decision weights and can be reused to compare LLMs with human recruiters using identical stimuli.

Data & Methods

  • Context and scope:
    • Data and attribute distributions are drawn from a large European freelance marketplace.
    • Primary occupational focus: full‑stack developers in France (high male name prevalence); robustness checks on SEO content writers reported in appendix.
  • Synthetic data construction:
    • Full factorial design for profiles and briefs to break correlations and enable causal identification.
    • Profiles (10,800): varied on skills (exact / close substitute / distant), experience (1 / 5 / 9 years), work arrangement preferences (remote/on‑site, full/part‑time), reputation (projects 0/1/5; rating 0/5; badge yes/no), daily rate (300–500€), prior employer (SME vs large; e‑commerce vs banking), and socio‑demographics signaled by first names (European male, European female, Arabic male); education (BSc vs MSc).
    • Briefs (16): varied recruiter gender (name), firm size (SME vs large), remote vs onsite, full vs part time; some dimensions held constant (e.g., required duration, tech stack, required experience).
  • LLM elicitation:
    • Model: Gemini 2.0 Flash, prompts in French.
    • Prompt asked for a hiring probability (10‑point scale) and a brief explanation of reasoning; no fairness correction instruction.
    • Each profile–brief pair scored three times; mean used.
  • Statistical analysis:
    • OLS regressions of average LLM score on binary indicators for attribute levels; standard errors clustered at brief level.
    • Given orthogonalized design, OLS coefficients interpreted as marginal causal effects.
    • Reported model summary: R^2 ≈ 0.90 — linear additive approximation explains most score variation in this experimental context.
  • Limitations noted by authors:
    • Full factorial design trades external representativeness for causal identification.
    • Prompt framing (hiring probability) and profile formatting may influence model reasoning; results may depend on prompt, language, or model family.
    • Single LLM and occupational focus limit generalizability; real‑world downstream outcomes not observed.

Implications for AI Economics

  • Theoretical alignment: LLMs can internalize economic signaling logic. In this experiment the LLM rewarded close matches and trust signals (experience, reputation) in ways consistent with labor‑economics models of signaling and productivity inference.
  • Measurement & auditing tools: Fully factorial synthetic designs adapted from correspondence studies provide a rigorous, causal way to estimate implicit attribute weights in LLM decision‑making and to surface interaction (intersectional) effects that simple attribution methods might miss.
  • Policy and platform implications:
    • Even when average demographic effects are small, intersectional heterogeneity can produce unequal outcomes; audits should measure both average and subgroup‑conditional effects.
    • Platforms using LLMs for candidate ranking or screening should monitor how the model reweights productivity signals across groups and consider calibration, counterfactual testing, or constraint mechanisms to prevent unintended disparate impacts.
    • Regulatory auditing frameworks could adopt similar experimental stimuli to evaluate vendor claims about fairness and alignment.
  • Research directions:
    • Direct human–AI comparison: deploy the identical synthetic stimuli to human recruiters to quantify alignment/misalignment and the welfare implications of replacing or assisting human screening with LLMs.
    • External validity tests: evaluate other LLMs, prompt framings, languages, and occupational contexts; link model scores to downstream hiring outcomes to measure market‑level effects.
    • Dynamic and systemic effects: study how LLM recommendations change demand‑side behavior (which candidates are contacted) and feedback loops that could amplify inequalities over time.
  • Practical takeaway for economists and practitioners: LLMs can be powerful, interpretable evaluators of unstructured candidate information, but rigorous causal audits (including intersectional analyses) are necessary before large‑scale deployment in hiring to detect subtle disparate‑impact mechanisms that simple averages hide.

Assessment

Paper Typequasi_experimental Evidence Strengthmedium — High internal validity for identifying how this specific LLM maps controlled attribute variation to output because of the full factorial experimental design, but limited external validity: results depend on the particular model version, prompt formulation, synthetic vignette construction, and marketplace data source, and do not directly observe real hiring outcomes. Methods Rigormedium — The use of a full factorial design and regression/interaction analyses is methodologically sound and appropriate for isolating attribute effects; however, rigor depends on implementation details not reported here (e.g., number of vignettes, randomization checks, multiple prompts/model seeds, robustness to prompt phrasing, pre-registration, and calibration of synthetic profiles), and there are potential measurement issues in inferred demographics and vignette realism. SampleSynthetic dataset of freelancer profiles and project descriptions constructed from real profiles and project ads drawn from a large European online freelance marketplace; attributes manipulated include skills, years of experience, hourly rate, portfolio strength, client ratings/reviews, and inferred demographic indicators; full factorial combinations of these attributes are used to generate vignettes evaluated by the LLM under a fixed prompt. Themeslabor_markets human_ai_collab IdentificationConstructs synthetic freelancer–project vignettes by fully factorially varying observable attributes (skills, experience, rates, portfolio signals, inferred demographics, etc.), feeds each vignette to the LLM under a fixed prompt, and estimates attribute weights by regressing the model's fit/suitability scores on the manipulated attributes (including interaction terms for project context and demographic subgroups); uses factorial contrasts and ANOVA-style decomposition to recover average marginal effects and heterogeneous effects. GeneralizabilityFindings apply to the specific LLM version, prompt phrasing, and inference settings used; other models or prompts may behave differently., Data come from a single European freelance marketplace — results may not generalize to other platforms, regions, sectors, or labor market institutions., Synthetic vignettes may omit contextual cues and complex realism present in true hiring interactions, limiting external validity to real recruiter decisions., Demographic indicators are inferred (or simulated) and may not reflect real-world identity presentation or intersectional dynamics accurately., LLM behavior can change with model updates, temperature/decoder settings, or few-shot examples; results capture a snapshot rather than durable model properties.

Claims (9)

ClaimDirectionOutcomeConfidence & EvidenceDetails
General-purpose Large Language Models (LLMs) show significant potential in recruitment applications, where decisions require reasoning over unstructured text, balancing multiple criteria, and inferring fit and competence from indirect productivity signals. Adoption Rate positive potential suitability of LLMs for recruitment tasks (conceptual capability)
Reading fidelity high
Study strength speculative
not reported
0.08
We propose a framework to evaluate an LLM's decision logic in recruitment by drawing on established economic methodologies for analyzing human hiring behavior. Research Productivity positive existence of an evaluation framework for LLM decision logic in recruitment
Reading fidelity high
Study strength medium
not reported
0.48
We build synthetic datasets from real freelancer profiles and project descriptions from a major European online freelance marketplace. Research Productivity positive availability of synthetic dataset derived from marketplace data
Reading fidelity high
Study strength medium
not reported
0.48
We apply a full factorial design to estimate how a LLM weighs different match-relevant criteria when evaluating freelancer-project fit. Research Productivity positive LLM-assigned weights to match-relevant criteria (estimated via factorial design)
Reading fidelity high
Study strength medium
not reported
0.48
The LLM prioritizes core productivity signals, such as skills and experience, when evaluating freelancer-project fit. Task Allocation positive relative importance (weights) assigned by the LLM to 'skills' and 'experience' attributes
Reading fidelity medium
Study strength medium
not reported
0.29
The LLM interprets certain features beyond their explicit matching value. Task Allocation mixed interpretation/use of profile/project features by the LLM beyond explicit matching signals
Reading fidelity medium
Study strength medium
not reported
0.29
The LLM shows minimal average discrimination against minority groups. Hiring null_result average difference in LLM evaluations across demographic (minority) groups
Reading fidelity medium
Study strength medium
not reported
0.29
Intersectional effects reveal that productivity signals carry different weights between demographic groups. Hiring mixed variation in attribute weights (productivity signals) across intersectional demographic subgroups
Reading fidelity medium
Study strength medium
not reported
0.29
We explain how a comparable experimental setup could be implemented with human recruiters to assess alignment between model and human decisions. Research Productivity positive feasibility and design of comparable human-recruiter experiments
Reading fidelity high
Study strength medium
not reported
0.48

Notes