The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests About 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

At a large recruitment firm, AI voice interviewers raised offer rates by 12% and increased hires’ starts and retention by roughly 17–18% while leaving measured productivity unchanged, with gains driven by more standardized and comparable interviews.

Voice AI in Firms: A Natural Field Experiment on Automated Job Interviews
Brian Jabarian, Luca Henkel · July 30, 2026
arxiv rct high evidence 9/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Brian Jabarian unresolved corpus identity
  2. Luca Henkel unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Brian Jabarian provider ID
  2. Luca Henkel provider ID
In a pre-registered natural field experiment with ~70,000 applicants, AI-led voice interviews raised offer rates by 12% and increased job starts and retention by ~17–18% without detectable declines in hired-worker productivity, driven by more structured, consistent information collection.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

This paper studies whether AI automation can improve organizational outcomes by reducing variance when collecting information. We conducted a large-scale natural field experiment in which 70,000 job applicants were randomly assigned to be interviewed by human recruiters or AI voice agents. In both conditions, human recruiters evaluate the interviews and make hiring decisions. Applicants interviewed by AI agents are 12% more likely to receive job offers, and these gains translate into higher job starts and worker retention, with no decline in the productivity of hired workers. Analyzing interview transcripts reveals that AI voice agents achieve controlled variance: their interviews are more structured and consistent while remaining responsive to individual applicants, which is associated with more hiring-relevant information collected. These results demonstrate that automating information collection with AI can enhance decision quality through standardization.

Summary

Main Finding

Randomized assignment of 70,884 applicants to AI voice–led versus human-led interviews increased hiring and retention without reducing worker productivity. Applicants interviewed by an AI voice agent were 12% more likely to receive job offers (8.70% → 9.73%), and those randomized to AI interviews were ~18% more likely to start the job and to remain employed for at least one month (effects persist through 2–4 months). Productivity measures for hires do not decline. The mechanism is “controlled variance”: AI-led interviews are more structured and consistent while remaining responsive, yielding richer and more comparable hiring-relevant signals.

Key Points

  • Experiment design

    • Natural field experiment in partnership with PSG Global Solutions.
    • Sample: 70,884 applications for entry-level customer service jobs.
    • Random assignment after pre-screen to one of: Human interviewer, AI voice interviewer, or Choice condition.
    • All hiring decisions are made by human recruiters (AI automates information collection only).
    • Study pre-registered and IRB-approved.
  • Main quantitative outcomes (intent-to-treat)

    • Offer rate: Human 8.70% → AI 9.73% (12% relative increase).
    • Job start probability: AI condition +18% (p < 0.001).
    • ≥1 month retention: AI condition +18% (p < 0.001); similar relative increases at 2–4 months (16–17%).
    • Conditional on accepted offers: AI → +7% starts (p = 0.003) and +6% ≥1 month retention (p = 0.025).
  • Productivity and separations

    • No statistically or economically meaningful differences in productivity metrics observed for hires (time handling customers, customer satisfaction scores, quality assurance scores).
    • No difference in distribution of voluntary vs involuntary separations between AI and human hires.
  • Mechanism: controlled variance / standardization

    • AI interviews more consistently follow topic order, cover a more consistent set of topics, and use more standardized wording and greater lexical richness.
    • Transcript/NLP evidence: AI interviews elicit linguistic features associated with higher offers (e.g., sustained conversational exchange) and fewer features associated with lower offers (e.g., backchannel signals, applicant-posed questions).
    • Recruiter scoring shifts: recruiters give higher interview scores to AI-interviewed applicants (shift from low → medium), and qualitative comments are more positive.
    • Recruiters change weighting: when evaluating AI-interviewed applicants, they put relatively less weight on interview scores and more on standardized language test scores.
  • Applicant responses and preferences

    • No evidence of backlash: offer acceptance, NPS (Net Promoter Score), and survey measures of stress/comfort are similar across treatments.
    • AI interviews perceived as less natural; reported gender-based discrimination nearly halved under AI (3.30% vs 5.98%, p = 0.02).
    • Choice condition: 78% of applicants choose the AI interviewer. However, applicants who choose AI have lower language and analytical test scores (negative sorting).
  • Implementation limits and frictions

    • 5% of applicants ended the interview because they refused to speak with an AI.
    • 7% of AI interviews experienced technical difficulties.
    • Recruitment sample: 131 recruiters evaluated interviews, with a core group of 43 handling most cases.
  • Transparency and potential conflicts

    • Study funded and supported by multiple institutions; the partner firm supplied data but was not involved in analyses.
    • Pre-registered; IRB approvals cited.
    • After data collection, one author later accepted an unpaid Chief Economist role at the partner firm; this is disclosed.

Data & Methods

  • Experimental setup

    • Large-scale randomized controlled trial embedded in a real hiring funnel at a recruitment process outsourcing firm.
    • Randomization occurs after pre-screening; treatments: Human interviewer, AI interviewer, Choice.
    • Recruiters follow the same set of interview guidelines in both human and AI treatments; the AI agent implements the same guidelines programmatically.
  • Outcomes measured

    • Hiring outcomes: offer rates, offer acceptance, job starts, retention at 1–4 months.
    • Productivity outcomes: time spent handling customers, customer satisfaction scores, quality assurance scores.
    • Behavioral measures: applicant surveys (candidate experience, NPS), choice behavior in the Choice arm.
    • Recruiter behavior: numerical interview scores and qualitative comments; recruiter survey on signal importance.
    • Operational frictions: technical failures, applicant refusals.
  • Analytical tools

    • Intent-to-treat comparisons across randomized arms (statistical significance reported).
    • Transcript analysis using natural language processing to (i) validate that interview signals predict offers, (ii) quantify structure and lexical patterns, and (iii) link linguistic features to outcome differences.
    • Sentiment analysis of recruiter comments.
    • Heterogeneity and conditional analyses (e.g., conditional on accepted offer; choice sorting).
  • AI system described

    • AI voice agent architecture: ASR (speech-to-text), LLM-based text generation, text-to-speech synthesis; guardrails to reduce hallucinations and stay on-topic.
    • Implementation challenges noted: accent/noise/ASR errors, hallucinations, applicant gaming, and real-world robustness.

Implications for AI Economics

  • Mechanism insight: “controlled variance” as a value proposition

    • The primary economic gain arises from reducing interviewer-driven noise in information collection while preserving responsiveness to the applicant—improving the signal quality recruiters use to make decisions.
    • This suggests AI can enhance firm outcomes not only by automating tasks or cutting costs but by improving the reliability and comparability of inputs to human decision-makers.
  • Human–AI division of labor

    • Evidence supports a model where AI automates structured information collection and humans retain evaluative judgment. This can shift human labor toward higher-order evaluation and decision tasks, consistent with theories of labor reallocation in the presence of AI.
  • Effects on bias and fairness

    • Reported gender-discrimination complaints fell under AI-led interviews in this setting; AI standardization may reduce some forms of interviewer-driven bias. However, negative sorting into AI (lower test scores among AI choosers) and potential algorithmic biases warrant careful monitoring.
  • Adoption considerations

    • High applicant willingness to use AI (78% when given choice) and no observed negative effects on acceptance or satisfaction indicate favorable uptake in similar populations.
    • Nontrivial frictions exist (5% refusal; 7% technical issues), so deployment requires operational robustness, fallback protocols, and user-choice options where appropriate.
  • Productivity vs. quality tradeoffs

    • Hiring gains did not come at observable productivity cost in this entry-level, high-turnover customer-service context. This suggests automation of information collection can raise effective match quality without reducing worker output in comparable tasks.
  • External validity and limits

    • Results are from entry-level customer service hires at a recruitment outsourcing firm; effects may differ in higher-skilled or less standardized hiring contexts where interpersonal nuance is more central.
    • The AI system, interview guidelines, recruiter incentives, and labor market context (easy replacement; low switching costs) shape outcomes and generalizability.
  • Research and policy directions

    • Further work should test controlled variance across other domains (e.g., managerial promotion panels, professional hiring) and examine long-run effects on workforce composition and wage outcomes.
    • Regulators and firms should consider transparency, auditability, and fairness evaluations when deploying AI for information collection in personnel decisions.

If you want, I can produce a one-page infographic-style brief with the key numbers and mechanism flow (AI → controlled variance → richer signals → higher offers & retention) or extract additional robustness checks and heterogeneity results from the appendix.

Assessment

Paper Typerct Evidence Strengthhigh — Large-scale randomized design (≈70,884 applicants), pre-registration and IRB approval, direct measurement of economically meaningful outcomes (offers, starts, retention, productivity), and complementary mechanism evidence from transcripts, surveys, and recruiter behavior strongly support causal claims; external-validity caveats remain (single firm, single job type, specific AI implementation). Methods Rigorhigh — Well-powered randomized assignment with pre-registration, multiple objective outcomes (offers, starts, retention), supplementary analyses linking interview transcripts to decision outcomes, recruiter surveys, and operational metrics; potential concerns include single-partner setting, some technical failures and applicant refusals, and limited detail here on balance/attrition diagnostics and blinding of recruiters to treatment status. Sample≈70,884 applications for entry-level customer service roles at PSG Global Solutions (a Teleperformance subsidiary); applicants who passed an initial pre-screen were randomized to AI-led interview, human-led interview, or choice; 131 recruiters evaluated interviews (core group of 43 handled majority); outcomes include offer rates, job starts, retention up to four months, productivity measures for a subset (handle time, customer satisfaction, QA scores), interview transcripts for NLP analysis, and candidate/recruiter surveys. Themeslabor_markets human_ai_collab org_design productivity IdentificationRandomized assignment (natural field experiment) of applicants who passed initial pre-screening to be interviewed by an AI voice agent, a human recruiter, or given a choice; all subsequent hiring decisions were made by human recruiters—intention-to-treat estimates identify the causal effect of automating the information-collection stage on offers, starts, retention, and productivity; pre-registered on AEA RCT Registry. GeneralizabilitySingle-firm sample (PSG Global Solutions/Teleperformance) may not represent other firms or sectors, Focus on entry-level customer-service jobs in a high-turnover market — results may not generalize to skilled, managerial, or different occupational settings, Applicants were pre-screened and thus not representative of all job seekers, Specific AI voice-agent architecture, prompts, and deployment choices may drive results and may not generalize to other AI systems, Recruiters retained final hiring authority; outcomes may differ where AI has evaluation or decision power, Technical failure (≈7%) and refusal-to-participate (≈5%) rates may vary across contexts and affect external validity, Results are time- and model-version dependent given rapid evolution of LLM/voice systems

Claims (12)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Applicants interviewed by an AI voice agent were 12% more likely to receive a job offer than applicants interviewed by human recruiters. Hiring positive Likelihood of receiving a job offer
Reading fidelity high
Study strength high
n=70884
12% higher likelihood of receiving a job offer
1.0
Applicants interviewed by AI were 18% more likely to start their jobs than applicants interviewed by human recruiters. Employment positive Likelihood of starting the job
Reading fidelity high
Study strength high
n=70884
18% higher likelihood of starting their job (p < 0.001)
1.0
Applicants interviewed by AI were 18% more likely to have an employment spell lasting at least one month than applicants interviewed by human recruiters. Turnover positive Likelihood of remaining employed for at least one month
Reading fidelity high
Study strength high
n=70884
18% higher likelihood of having an employment spell lasting at least one month (p < 0.001)
1.0
The positive employment-retention effect of AI interviews persisted through four months after hiring. Turnover positive Likelihood of still being employed after two, three, or four months
Reading fidelity high
Study strength high
n=70884
17% higher likelihood after two months; 16% after three months; 17% after four months
1.0
AI interviews did not reduce the productivity of hired workers. Firm Productivity null_result Customer-handling time, customer satisfaction, and employer quality-assurance scores
Reading fidelity high
Study strength medium
Neither statistically significant nor economically meaningful differences
0.6
AI-led interviews were more structured and consistent than human-led interviews. Organizational Efficiency positive Interview structure and consistency
Reading fidelity high
Study strength medium
not reported
0.6
AI-led interviews were associated with linguistic features that predict higher offer rates and fewer features associated with lower offer rates in human-led interviews. Decision Quality positive Hiring-relevant information and offer-predictive interview features
Reading fidelity high
Study strength medium
not reported
0.6
Five percent of applicants ended their interview because they were unwilling to speak to an AI, and the AI voice agent experienced technical difficulties in 7% of cases. Error Rate negative Interview termination due to unwillingness to use AI and technical difficulties
Reading fidelity high
Study strength medium
5% ended their interview; 7% involved technical difficulties
0.6
AI-led interviews nearly halved reported gender-based discrimination relative to human-led interviews. Ai Safety And Ethics negative Reported gender-based discrimination
Reading fidelity high
Study strength medium
3.30% vs 5.98% (p = 0.02)
0.6
When applicants were given a choice, 78% chose the AI voice agent rather than a human recruiter. Worker Satisfaction positive Applicant preference for AI versus human interviewer
Reading fidelity high
Study strength medium
78% chose the AI voice agent
0.6
Applicants who chose the AI interviewer had significantly lower language and analytical scores than applicants who chose a human recruiter. Skill Acquisition negative Language and analytical test scores among applicants selecting each interviewer type
Reading fidelity high
Study strength medium
not reported
0.6
Human recruiters assigned significantly higher interview scores to applicants interviewed by AI than to applicants they interviewed themselves. Decision Quality positive Recruiter-assigned interview scores
Reading fidelity high
Study strength medium
n=131
0.6

Notes