0 cumulative citations
View corpus contextGenerative-AI assistants speed up after-sales chat and boost customer ratings overall, but benefits concentrate among lower-performing agents while top agents become more distracted and sometimes worsen outcomes.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
In collaboration with Alibaba, this study leverages a large-scale field experiment to assess the impact of a generative AI assistant on worker performance in e-commerce after-sales service. Human agents providing digital chat support were randomly assigned with access to a gen AI assistant that offered two core functions: diagnosis of customer issues and solution proposals, presented as text messages. Agents retained discretion to adopt, modify, or disregard AI-generated messages. To evaluate gen AI's impact, we estimate both the intention-to-treat (ITT) effect of gen AI access and the local average treatment effect (LATE) of gen AI usage. Results show that gen AI significantly improved service speed, measured by issue identification time and chat duration. Gen AI also improved subjective service quality reflected in customer ratings and dissatisfaction rates, but it had no significant effect on objective service quality indicated by customer retrial rates. The performance improvements stemmed not only from automation but also from changes in the dynamics of agent-customer interactions: agent communication became more informative and efficient, while customers experienced reduced communication burdens. Low performers achieved the greatest improvements in both service speed and quality, narrowing the performance gap. In contrast, top-performing agents showed little improvement in service speed but experienced declines in both subjective and objective service quality. Evidence suggests that this decline results from increased multitasking tendency, proxied by longer shift-away times across concurrent chats, which slowed customer responses and raised abandonment and retrial rates. These findings suggest that gen AI reshapes work, demanding tailored deployment strategies.
Summary
Main Finding
Access to a generative-AI assistant in Alibaba’s after-sales chat support causally increased agent speed and raised customers’ subjective satisfaction, but did not improve objective resolution rates on average. Effects were heterogeneous: low-performing agents gained the most on both speed and perceived quality (narrowing the performance gap), while top-performing agents experienced declines in both subjective and objective quality driven by increased multitasking.
Key Points
- Intervention: Agents were randomized to access a gen-AI assistant that produced two text suggestions per chat (issue diagnosis and solution); agents could copy, edit, or ignore suggestions.
- Aggregate effects (ITT / LATE):
- Faster service: shorter issue-identification time and reduced chat duration.
- Higher subjective quality: increased customer ratings and lower dissatisfaction rates.
- No significant improvement in objective quality overall: customer retrial rates (a proxy for unresolved issues) were unchanged on average.
- Mechanisms:
- Agent behavior: responses became faster, more numerous, and linguistically more informative (higher lexical density, clarity, diversity).
- Customer behavior: reduced communication burden (fewer supporting pictures uploaded, simpler language), implying improved perceived clarity/effort.
- Both automation (time savings) and altered interaction dynamics (more informative agent messages, less customer effort) drove improvements in perceived service.
- Heterogeneity by pre-treatment performance:
- Low performers (bottom quintile) showed the largest improvements in speed and subjective quality.
- Top performers (top quintile) showed little speed gain and a deterioration in both subjective (ratings, dissatisfaction) and objective (retrial) outcomes.
- For top performers, process data indicate increased shift-away times and multitasking across concurrent chats, slowing customer responses, increasing abandonment, and triggering immediate retrials.
- Adoption and usage:
- Average adherence to AI suggestions was low and varied by skill level; the authors estimate causal LATE of actual AI usage in addition to ITT of access.
- Top agents were more skeptical and used AI suggestions less frequently, often verifying outputs more.
Data & Methods
- Setting: Large-scale randomized field experiment in Alibaba/Taobao after-sales chat support (order-related returns/exchanges/refunds/repairs).
- Randomization: Agents randomly assigned to treatment (AI access) or control (no AI access); assignment remained as intended.
- Intervention: A gen-AI assistant (Alibaba’s Qwen LLMs fine-tuned on service chats and order data) provided two text suggestions per chat: (1) diagnosis, (2) proposed solution. Suggestions shown in chat sidebar; agent discretion preserved.
- Outcomes measured from fine-grained process logs and post-chat surveys:
- Service speed: issue identification time, total chat duration, response latency.
- Subjective quality: customer post-chat rating (1–5), dissatisfaction indicator.
- Objective quality: customer retrial rates (repeat contacts), abandonment rates.
- Interaction measures: number of messages, lexical density/clarity/diversity of agent text, customer uploads (images), shift-away time (time agents spend away from a chat while handling others).
- Identification strategy:
- Intention-to-treat (ITT) estimated the effect of access to AI.
- Local average treatment effect (LATE) estimated causal effect of actual AI usage (instrumenting usage with assignment).
- Heterogeneity analyzed by pre-treatment performance quintiles.
- Supplementary evidence: textual analysis of messages, process-level timing metrics, and internal agent survey data used to probe mechanisms (use patterns, attitudes).
Implications for AI Economics
- Human-AI complementarity is nuanced: gen-AI can speed up tasks and improve perceived service without necessarily improving objective task success. Economists should distinguish productivity gains in speed/effort from gains in underlying task effectiveness.
- Distributional effects:
- Gen-AI can reduce within-firm performance dispersion by disproportionately helping low performers — a potential equalizing force that may affect returns to skill and wage dispersion in service sectors.
- Conversely, it can harm top performers’ realized performance through workflow changes (e.g., induced multitasking), suggesting non-monotonic effects across the skill distribution.
- Behavioral channels matter: productivity effects arise through changes in interaction dynamics (message informativeness, customer effort), not only automation. Modeling frameworks should incorporate behavioral responses (multitasking, verification, trust/aversion) to AI tools.
- Measurement caution: reliance on subjective satisfaction metrics alone can overstate welfare gains if objective resolution remains unchanged. Policy and management evaluations should include objective outcomes (retrial/resolution rates) and process-level indicators.
- Managerial and policy recommendations:
- Tailor deployment: prioritize AI support for lower-skilled workers to maximize gains and reduce variance.
- Monitor unintended workflow effects: design UI and incentives to limit harmful multitasking (e.g., discourage excessive shift-away behavior, prioritize single-chat focus when needed).
- Training and role design: provide training to integrate AI suggestions productively and preserve tasks that keep top performers engaged, avoiding demotivation or deskilling.
- Performance metrics: combine speed, satisfaction, and objective resolution in KPIs to avoid perverse incentives (e.g., rewarding fast but ineffective handling).
- Research directions: study longer-run effects on skill development and labor outcomes (wages, retention), general equilibrium impacts across firms/sectors, and design interventions (interface, incentives, verification protocols) that mitigate adverse effects for high performers while amplifying benefits for low performers.
Assessment
Claims (8)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Access to the generative AI assistant significantly improved service speed, measured by faster issue identification time and shorter chat duration. Task Completion Time | positive | service speed (issue identification time and chat duration) |
Reading fidelity
high
Study strength
high
|
not reported
|
| Generative AI access improved subjective service quality, as reflected in higher customer ratings and lower customer dissatisfaction rates. Output Quality | positive | customer ratings and dissatisfaction rates |
Reading fidelity
high
Study strength
high
|
not reported
|
| Generative AI access had no significant effect on objective service quality as measured by customer retrial rates. Output Quality | null_result | customer retrial rates |
Reading fidelity
high
Study strength
high
|
not reported
|
| Performance improvements from generative AI stemmed not only from automation but also from changes in agent-customer interaction dynamics: agents' messages became more informative and efficient, and customers experienced reduced communication burdens. Organizational Efficiency | mixed | informativeness/efficiency of agent communication; customer communication burden |
Reading fidelity
medium
Study strength
medium
|
not reported
|
| Lower-performing agents achieved the greatest improvements in both service speed and service quality after receiving access to the generative AI assistant, narrowing the performance gap with higher-performing agents. Task Completion Time | positive | service speed and service quality by baseline agent performance subgroup |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Top-performing agents showed little improvement in service speed but experienced declines in both subjective and objective service quality after receiving generative AI access. Output Quality | negative | service speed, customer ratings, and objective quality (retrial/abandonment) for top performers |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Evidence suggests the decline in top performers' quality is driven by increased multitasking: longer 'shift-away' times across concurrent chats slowed customer responses and raised abandonment and retrial rates. Task Allocation | negative | shift-away time as proxy for multitasking; customer response time, abandonment and retrial rates |
Reading fidelity
medium
Study strength
medium
|
not reported
|
| These findings imply that generative AI reshapes work and that deployment should be tailored to worker heterogeneity and task dynamics. Governance And Regulation | mixed | implication for deployment strategy (qualitative recommendation) |
Reading fidelity
high
Study strength
speculative
|
not reported
|