4 cumulative citations
View corpus contextHuman-AI teams perform better when humans can learn from outcomes: experiments that provide outcome feedback—especially paired with AI explanations—show positive synergy, whereas explanations alone often harm joint performance.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
The collaboration between humans and artificial intelligence (AI) holds the promise of achieving superior outcomes compared to either acting alone-a phenomenon called human-AI synergy. Nevertheless, our understanding of the conditions that facilitate such human-AI synergy when humans are advised by AI remains limited. A recent meta-analysis showed that, on average, human-AI combinations do not outperform the better individual agent. We argue that this pessimistic conclusion arises from insufficient attention to human learning in the experimental designs. To substantiate this claim, we re-analyzed all 74 studies included in the original meta-analysis, yielding two new findings. First, most previous research overlooked design features that foster human learning, such as providing outcome feedback to participants. Second, our re-analysis demonstrated that studies providing outcome feedback show tentatively higher synergy than those without outcome feedback. Crucially, feedback paired with AI explanations tends to yield positive synergy, while explanations without feedback were linked to negative synergy-indicating that explanations increase synergy only when humans can learn to verify the AI's reliability through feedback. We conclude that the current literature underestimates the potential of human-AI collaboration because it predominantly relies on paradigms that do not facilitate human learning, thus hindering humans from effectively adapting their collaboration strategies. We therefore advocate for a paradigm shift in human-AI interaction research that explicitly addresses human learning and thus enhances our understanding of and support for successful human-AI collaboration.
Summary
Main Finding
Re-analysis of 74 human–AI interaction studies (370 experimental conditions) shows that many prior experiments systematically undercut human learning opportunities (e.g., no trial-by-trial outcome feedback). Studies that provided outcome feedback—especially when paired with AI explanations—tended to report higher human–AI synergy. Explanations without feedback were associated with negative synergy. The authors conclude that neglecting human learning leads to an underestimate of the potential for productive human–AI collaboration.
Key Points
- Dataset: 74 studies synthesized by Vaccaro et al. (370 conditions). The authors recoded design features relevant to learning.
- Prevalence of feedback: Only 10 of 74 studies (14%) included experimental conditions with outcome feedback; these provided 98 of 370 effect sizes (26%).
- Synergy metric: Human–AI synergy measured as Hedges’ g comparing the human–AI combination to the better individual agent.
- Main quantitative results:
- Without feedback: Hedges’ g = −0.17 (95% CI −0.37 to 0.00), BFinclusion = 1.4 (tendency toward negative synergy).
- With feedback: Hedges’ g = 0.12 (95% CI −0.28 to 0.34), BFinclusion = 0.92.
- Contrast (with vs without feedback): g = 0.34 (95% CI −0.01 to 0.69), posterior probability PD = 84% (tentative evidence of higher synergy with feedback).
- Explanations × feedback:
- Explanations + feedback: g = 0.30 (95% CI −0.18 to 0.48), BFinclusion = 4.02 (positive tendency).
- Explanations without feedback: g = −0.31 (95% CI −0.48 to −0.11), BFinclusion = 17.14 (clear negative effect).
- Contrast when explanations present (feedback vs no feedback): g = 0.60 (95% CI 0 to 0.94), PD = 97% (strong evidence feedback improves synergy when explanations are provided).
- Interpretation: AI explanations appear beneficial only when participants can validate AI reliability via outcome feedback; explanations without feedback can worsen human–AI performance.
- Robustness checks: Adding moderators used by Vaccaro et al. did not change the directional pattern.
- Caveats:
- Small fraction of studies with feedback and limited trial-level reporting.
- Most studies report average performance across trials, obscuring within-subject learning dynamics.
- Feedback usefulness depends on task structure (timeliness, noise, feature visibility).
Data & Methods
- Source data: Vaccaro et al.’s systematic review and meta-analysis (74 studies). Authors merged their new codes with Vaccaro et al.’s open data.
- Coding: Six authors coded learning-relevant design features (e.g., outcome feedback, practice trials); each condition double-coded, disagreements adjudicated.
- Meta-analytic method: Robust Bayesian Model Averaging (RoBMA) with a three-level hierarchical structure (effect sizes nested in experiments nested in studies). Default RoBMA priors used.
- Estimation details: Spike-and-slab algorithm; at least 45k posterior samples after burn-in; convergence monitored by Gelman–Rubin ˆR < 1.05. Contrasts derived from conditional marginal posterior distributions.
- Data & code availability: Openly available (OSF link provided in paper).
- Supplemental analyses: Limited trial-level analyses in three studies with available time-series data showed participants could learn to align with the better-performing AI when informative signals were present.
Implications for AI Economics
- Reassess empirical estimates of complementarities: Economists modeling human–AI complementarities (e.g., productivity gains, task allocation) should account for learning dynamics. Cross-sectional estimates from studies lacking feedback likely understate potential gains from human–AI teams.
- Importance of feedback infrastructure: In organizational settings, investing in outcome feedback mechanisms (rapid, accurate performance feedback) can unlock complementarities between workers and AI systems. Cost–benefit analyses should include the value of increased synergy due to learning.
- Design of field experiments and lab-in-the-field studies: When evaluating AI interventions in markets or firms, randomize feedback and explanation treatments and collect trial-level/longitudinal data to capture dynamic adaptation and learning effects.
- Incentive and training policies: Policy and management interventions (training, performance transparency, explicit explanations of model behavior) should be paired with feedback loops to ensure workers can calibrate reliance on AI and improve joint performance.
- Measurement and evaluation: Standardize performance metrics and report participant-level and time-resolved outcomes. Economists estimating returns to AI adoption should use measures that reflect performance after learning (not only initial trials).
- Regulation and consumer protection: Regulators evaluating AI-assisted decision systems (e.g., credit scoring, health diagnostics) should consider how feedback availability and explainability affect downstream human behavior and aggregate outcomes; lack of feedback may produce systematically suboptimal human reliance.
- Research agenda for applied economists: Incorporate models of endogenous learning and adaptive reliance into structural models of human–AI interaction; estimate how learning frictions and feedback technologies change equilibrium adoption, task assignment, wages, and welfare.
Actionable recommendations for researchers and practitioners: - In experiments and deployments, provide timely, accurate outcome feedback and pair it with interpretable explanations. - Collect and report trial-level and participant-level data to observe learning trajectories. - When assessing human–AI complementarities, evaluate performance after sufficient exposure (not only initial interactions). - Include feedback and explanation design as key policy levers in cost–benefit and regulatory analyses.
Assessment
Claims (8)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| A recent meta-analysis showed that, on average, human-AI combinations do not outperform the better individual agent. Decision Quality | null_result | performance of human-AI combinations relative to the better individual agent (human or AI) |
Reading fidelity
high
Study strength
high
|
n=74
|
| The pessimistic conclusion from the prior meta-analysis arises from insufficient attention to human learning in experimental designs. Decision Quality | negative | estimated human-AI synergy as biased by experimental design (i.e., underestimation due to lack of human-learning-facilitating features) |
Reading fidelity
high
Study strength
speculative
|
n=74
|
| Most previous research overlooked design features that foster human learning, such as providing outcome feedback to participants. Skill Acquisition | negative | presence of outcome feedback (design feature) in human-AI experimental studies |
Reading fidelity
high
Study strength
high
|
n=74
|
| Studies that provided outcome feedback show tentatively higher human-AI synergy than those without outcome feedback. Decision Quality | positive | human-AI synergy (performance advantage of human-AI pairings) |
Reading fidelity
high
Study strength
medium
|
n=74
|
| Feedback paired with AI explanations tends to yield positive synergy, while explanations without feedback were linked to negative synergy. Decision Quality | mixed | human-AI synergy as a function of (a) outcome feedback and (b) presence of AI explanations |
Reading fidelity
high
Study strength
medium
|
n=74
|
| The current literature underestimates the potential of human-AI collaboration because it predominantly relies on paradigms that do not facilitate human learning, which hinders humans from effectively adapting their collaboration strategies. Decision Quality | negative | estimated potential (performance) of human-AI collaboration as reported in the literature |
Reading fidelity
high
Study strength
medium
|
n=74
|
| Human-AI explanations increase synergy only when humans can learn to verify the AI's reliability through feedback; explanations without feedback do not increase (and may decrease) synergy. Decision Quality | mixed | effect of AI explanations on human-AI synergy conditional on presence of outcome feedback |
Reading fidelity
high
Study strength
medium
|
n=74
|
| Research in human-AI interaction should shift paradigms to explicitly address human learning in order to better understand and support successful human-AI collaboration. Research Productivity | positive | research practices and resulting quality of evidence about human-AI collaboration |
Reading fidelity
high
Study strength
speculative
|
n=74
|