The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests About 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Generative-AI assistants speed up after-sales chat and boost customer ratings overall, but benefits concentrate among lower-performing agents while top agents become more distracted and sometimes worsen outcomes.

Generative AI in Action: Field Experimental Evidence from Alibaba's Customer Service Operations
Xiao Ni, Yiwei Wang, Tianjun Feng, Lauren Xiaoyan Lu, Yitong Wang, Congyi Zhou · February 08, 2026
arxiv rct high evidence 10/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Xiao Ni unresolved corpus identity
  2. Yiwei Wang unresolved corpus identity
  3. Tianjun Feng unresolved corpus identity
  4. Lauren Xiaoyan Lu unresolved corpus identity
  5. Yitong Wang unresolved corpus identity
  6. Congyi Zhou unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Xiaojie Ni provider ID
  2. Yiwei Wang provider ID
  3. Tianjun Feng provider ID
  4. L. Lu provider ID
  5. Yitong Wang provider ID
  6. Congyi Zhou provider ID
A randomized field experiment at Alibaba finds that giving chat agents access to a generative-AI assistant speeds up issue identification and chat handling and raises customer ratings overall, with the largest gains for low-performing agents while top performers slow down and see declines in quality due to increased multitasking.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

In collaboration with Alibaba, this study leverages a large-scale field experiment to assess the impact of a generative AI assistant on worker performance in e-commerce after-sales service. Human agents providing digital chat support were randomly assigned with access to a gen AI assistant that offered two core functions: diagnosis of customer issues and solution proposals, presented as text messages. Agents retained discretion to adopt, modify, or disregard AI-generated messages. To evaluate gen AI's impact, we estimate both the intention-to-treat (ITT) effect of gen AI access and the local average treatment effect (LATE) of gen AI usage. Results show that gen AI significantly improved service speed, measured by issue identification time and chat duration. Gen AI also improved subjective service quality reflected in customer ratings and dissatisfaction rates, but it had no significant effect on objective service quality indicated by customer retrial rates. The performance improvements stemmed not only from automation but also from changes in the dynamics of agent-customer interactions: agent communication became more informative and efficient, while customers experienced reduced communication burdens. Low performers achieved the greatest improvements in both service speed and quality, narrowing the performance gap. In contrast, top-performing agents showed little improvement in service speed but experienced declines in both subjective and objective service quality. Evidence suggests that this decline results from increased multitasking tendency, proxied by longer shift-away times across concurrent chats, which slowed customer responses and raised abandonment and retrial rates. These findings suggest that gen AI reshapes work, demanding tailored deployment strategies.

Summary

Main Finding

Access to a generative-AI assistant in Alibaba’s after-sales chat support causally increased agent speed and raised customers’ subjective satisfaction, but did not improve objective resolution rates on average. Effects were heterogeneous: low-performing agents gained the most on both speed and perceived quality (narrowing the performance gap), while top-performing agents experienced declines in both subjective and objective quality driven by increased multitasking.

Key Points

  • Intervention: Agents were randomized to access a gen-AI assistant that produced two text suggestions per chat (issue diagnosis and solution); agents could copy, edit, or ignore suggestions.
  • Aggregate effects (ITT / LATE):
    • Faster service: shorter issue-identification time and reduced chat duration.
    • Higher subjective quality: increased customer ratings and lower dissatisfaction rates.
    • No significant improvement in objective quality overall: customer retrial rates (a proxy for unresolved issues) were unchanged on average.
  • Mechanisms:
    • Agent behavior: responses became faster, more numerous, and linguistically more informative (higher lexical density, clarity, diversity).
    • Customer behavior: reduced communication burden (fewer supporting pictures uploaded, simpler language), implying improved perceived clarity/effort.
    • Both automation (time savings) and altered interaction dynamics (more informative agent messages, less customer effort) drove improvements in perceived service.
  • Heterogeneity by pre-treatment performance:
    • Low performers (bottom quintile) showed the largest improvements in speed and subjective quality.
    • Top performers (top quintile) showed little speed gain and a deterioration in both subjective (ratings, dissatisfaction) and objective (retrial) outcomes.
    • For top performers, process data indicate increased shift-away times and multitasking across concurrent chats, slowing customer responses, increasing abandonment, and triggering immediate retrials.
  • Adoption and usage:
    • Average adherence to AI suggestions was low and varied by skill level; the authors estimate causal LATE of actual AI usage in addition to ITT of access.
    • Top agents were more skeptical and used AI suggestions less frequently, often verifying outputs more.

Data & Methods

  • Setting: Large-scale randomized field experiment in Alibaba/Taobao after-sales chat support (order-related returns/exchanges/refunds/repairs).
  • Randomization: Agents randomly assigned to treatment (AI access) or control (no AI access); assignment remained as intended.
  • Intervention: A gen-AI assistant (Alibaba’s Qwen LLMs fine-tuned on service chats and order data) provided two text suggestions per chat: (1) diagnosis, (2) proposed solution. Suggestions shown in chat sidebar; agent discretion preserved.
  • Outcomes measured from fine-grained process logs and post-chat surveys:
    • Service speed: issue identification time, total chat duration, response latency.
    • Subjective quality: customer post-chat rating (1–5), dissatisfaction indicator.
    • Objective quality: customer retrial rates (repeat contacts), abandonment rates.
    • Interaction measures: number of messages, lexical density/clarity/diversity of agent text, customer uploads (images), shift-away time (time agents spend away from a chat while handling others).
  • Identification strategy:
    • Intention-to-treat (ITT) estimated the effect of access to AI.
    • Local average treatment effect (LATE) estimated causal effect of actual AI usage (instrumenting usage with assignment).
    • Heterogeneity analyzed by pre-treatment performance quintiles.
  • Supplementary evidence: textual analysis of messages, process-level timing metrics, and internal agent survey data used to probe mechanisms (use patterns, attitudes).

Implications for AI Economics

  • Human-AI complementarity is nuanced: gen-AI can speed up tasks and improve perceived service without necessarily improving objective task success. Economists should distinguish productivity gains in speed/effort from gains in underlying task effectiveness.
  • Distributional effects:
    • Gen-AI can reduce within-firm performance dispersion by disproportionately helping low performers — a potential equalizing force that may affect returns to skill and wage dispersion in service sectors.
    • Conversely, it can harm top performers’ realized performance through workflow changes (e.g., induced multitasking), suggesting non-monotonic effects across the skill distribution.
  • Behavioral channels matter: productivity effects arise through changes in interaction dynamics (message informativeness, customer effort), not only automation. Modeling frameworks should incorporate behavioral responses (multitasking, verification, trust/aversion) to AI tools.
  • Measurement caution: reliance on subjective satisfaction metrics alone can overstate welfare gains if objective resolution remains unchanged. Policy and management evaluations should include objective outcomes (retrial/resolution rates) and process-level indicators.
  • Managerial and policy recommendations:
    • Tailor deployment: prioritize AI support for lower-skilled workers to maximize gains and reduce variance.
    • Monitor unintended workflow effects: design UI and incentives to limit harmful multitasking (e.g., discourage excessive shift-away behavior, prioritize single-chat focus when needed).
    • Training and role design: provide training to integrate AI suggestions productively and preserve tasks that keep top performers engaged, avoiding demotivation or deskilling.
    • Performance metrics: combine speed, satisfaction, and objective resolution in KPIs to avoid perverse incentives (e.g., rewarding fast but ineffective handling).
  • Research directions: study longer-run effects on skill development and labor outcomes (wages, retention), general equilibrium impacts across firms/sectors, and design interventions (interface, incentives, verification protocols) that mitigate adverse effects for high performers while amplifying benefits for low performers.

Assessment

Paper Typerct Evidence Strengthhigh — Causal identification rests on a large-scale randomized field experiment with ITT estimates and LATE via instrumenting usage with assignment; outcomes include both objective (timestamps, retrial rates) and subjective (customer ratings) measures and documented heterogeneous and mechanism analyses, yielding strong internal validity. Methods Rigorhigh — Rigorous experimental design (randomization), complementary ITT/LATE estimation to handle noncompliance, multiple objective and subjective outcome measures, and plausible mechanism tests (communication content, shift-away times, heterogeneity by baseline skill). Potential limitations (spillovers, long-term effects, exact compliance rates) are not specified in the prompt but do not materially weaken the core randomized identification. SampleHuman customer-service agents on Alibaba's e-commerce after-sales digital chat platform participating in a large-scale field experiment; randomization at the agent level with observed chat-level outcomes drawn from platform logs (chat timestamps, agent activity traces), customer ratings and dissatisfaction flags, and customer retrial/abandonment indicators; exact sample sizes and time window not reported in the summary. Themesproductivity human_ai_collab IdentificationRandomized assignment of human after-sales chat agents to access a generative-AI assistant (agent-level randomization). Estimation of intention-to-treat (ITT) compares outcomes by assignment; local average treatment effects (LATE) for AI usage are recovered by instrumenting observed usage with random assignment (to address partial compliance). Mechanisms explored via agent activity logs (e.g., shift-away time) and heterogeneity analysis by pre-treatment performance. GeneralizabilitySingle firm/platform (Alibaba) — effects may differ in other firms or platform structures, Single task/context: after-sales customer chat support — not necessarily generalizable to other occupations or offline service tasks, Cultural/language context (likely Chinese-language e-commerce) may affect interaction patterns and customer responses, Results depend on the specific AI model, its UI, and how suggestions were presented — different models/UIs could change effects, Short- to medium-term experiment — long-run effects (learning, task reallocation, staffing changes) are not observed, Agents retained discretion to use suggestions — effects may differ under more prescriptive or fully automated deployments

Claims (8)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Access to the generative AI assistant significantly improved service speed, measured by faster issue identification time and shorter chat duration. Task Completion Time positive service speed (issue identification time and chat duration)
Reading fidelity high
Study strength high
not reported
1.0
Generative AI access improved subjective service quality, as reflected in higher customer ratings and lower customer dissatisfaction rates. Output Quality positive customer ratings and dissatisfaction rates
Reading fidelity high
Study strength high
not reported
1.0
Generative AI access had no significant effect on objective service quality as measured by customer retrial rates. Output Quality null_result customer retrial rates
Reading fidelity high
Study strength high
not reported
1.0
Performance improvements from generative AI stemmed not only from automation but also from changes in agent-customer interaction dynamics: agents' messages became more informative and efficient, and customers experienced reduced communication burdens. Organizational Efficiency mixed informativeness/efficiency of agent communication; customer communication burden
Reading fidelity medium
Study strength medium
not reported
0.36
Lower-performing agents achieved the greatest improvements in both service speed and service quality after receiving access to the generative AI assistant, narrowing the performance gap with higher-performing agents. Task Completion Time positive service speed and service quality by baseline agent performance subgroup
Reading fidelity high
Study strength medium
not reported
0.6
Top-performing agents showed little improvement in service speed but experienced declines in both subjective and objective service quality after receiving generative AI access. Output Quality negative service speed, customer ratings, and objective quality (retrial/abandonment) for top performers
Reading fidelity high
Study strength medium
not reported
0.6
Evidence suggests the decline in top performers' quality is driven by increased multitasking: longer 'shift-away' times across concurrent chats slowed customer responses and raised abandonment and retrial rates. Task Allocation negative shift-away time as proxy for multitasking; customer response time, abandonment and retrial rates
Reading fidelity medium
Study strength medium
not reported
0.36
These findings imply that generative AI reshapes work and that deployment should be tailored to worker heterogeneity and task dynamics. Governance And Regulation mixed implication for deployment strategy (qualitative recommendation)
Reading fidelity high
Study strength speculative
not reported
0.1

Notes