The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Human-AI teams perform better when humans can learn from outcomes: experiments that provide outcome feedback—especially paired with AI explanations—show positive synergy, whereas explanations alone often harm joint performance.

Fostering human learning is crucial for boosting human-AI synergy
Julian Berger, Jason W. Burton, Ralph Hertwig, Thomas Kosch, Ralf H. J. M. Kurvers, Benito Kurzenberger, Christopher Lazik, Linda Onnasch, Tobias Rieger, Anna I. Thoma, Dirk U. Wulff, Stefan M. Herzog · December 15, 2025
arxiv review_meta medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Julian Berger unresolved corpus identity
  2. Jason W. Burton unresolved corpus identity
  3. Ralph Hertwig unresolved corpus identity
  4. Thomas Kosch unresolved corpus identity
  5. Ralf H. J. M. Kurvers unresolved corpus identity
  6. Benito Kurzenberger unresolved corpus identity
  7. Christopher Lazik unresolved corpus identity
  8. Linda Onnasch unresolved corpus identity
  9. Tobias Rieger unresolved corpus identity
  10. Anna I. Thoma unresolved corpus identity
  11. Dirk U. Wulff unresolved corpus identity
  12. Stefan M. Herzog unresolved corpus identity

Semantic Scholar

Latest observation:

  1. J. Berger provider ID
  2. Jason W. Burton provider ID
  3. Ralph Hertwig provider ID
  4. Thomas Kosch provider ID
  5. R. Kurvers provider ID
  6. Benito Kurzenberger provider ID
  7. C. Lazik provider ID
  8. L. Onnasch provider ID
  9. Tobias Rieger provider ID
  10. A. I. Thoma provider ID
  11. Dirk U. Wulff provider ID
  12. Stefan M. Herzog provider ID
A re-analysis of 74 experiments finds that design features that enable human learning—particularly outcome feedback combined with AI explanations—are associated with positive human–AI synergy, while explanations without feedback tend to reduce synergy.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

The collaboration between humans and artificial intelligence (AI) holds the promise of achieving superior outcomes compared to either acting alone-a phenomenon called human-AI synergy. Nevertheless, our understanding of the conditions that facilitate such human-AI synergy when humans are advised by AI remains limited. A recent meta-analysis showed that, on average, human-AI combinations do not outperform the better individual agent. We argue that this pessimistic conclusion arises from insufficient attention to human learning in the experimental designs. To substantiate this claim, we re-analyzed all 74 studies included in the original meta-analysis, yielding two new findings. First, most previous research overlooked design features that foster human learning, such as providing outcome feedback to participants. Second, our re-analysis demonstrated that studies providing outcome feedback show tentatively higher synergy than those without outcome feedback. Crucially, feedback paired with AI explanations tends to yield positive synergy, while explanations without feedback were linked to negative synergy-indicating that explanations increase synergy only when humans can learn to verify the AI's reliability through feedback. We conclude that the current literature underestimates the potential of human-AI collaboration because it predominantly relies on paradigms that do not facilitate human learning, thus hindering humans from effectively adapting their collaboration strategies. We therefore advocate for a paradigm shift in human-AI interaction research that explicitly addresses human learning and thus enhances our understanding of and support for successful human-AI collaboration.

Summary

Main Finding

Re-analysis of 74 human–AI interaction studies (370 experimental conditions) shows that many prior experiments systematically undercut human learning opportunities (e.g., no trial-by-trial outcome feedback). Studies that provided outcome feedback—especially when paired with AI explanations—tended to report higher human–AI synergy. Explanations without feedback were associated with negative synergy. The authors conclude that neglecting human learning leads to an underestimate of the potential for productive human–AI collaboration.

Key Points

  • Dataset: 74 studies synthesized by Vaccaro et al. (370 conditions). The authors recoded design features relevant to learning.
  • Prevalence of feedback: Only 10 of 74 studies (14%) included experimental conditions with outcome feedback; these provided 98 of 370 effect sizes (26%).
  • Synergy metric: Human–AI synergy measured as Hedges’ g comparing the human–AI combination to the better individual agent.
  • Main quantitative results:
    • Without feedback: Hedges’ g = −0.17 (95% CI −0.37 to 0.00), BFinclusion = 1.4 (tendency toward negative synergy).
    • With feedback: Hedges’ g = 0.12 (95% CI −0.28 to 0.34), BFinclusion = 0.92.
    • Contrast (with vs without feedback): g = 0.34 (95% CI −0.01 to 0.69), posterior probability PD = 84% (tentative evidence of higher synergy with feedback).
    • Explanations × feedback:
      • Explanations + feedback: g = 0.30 (95% CI −0.18 to 0.48), BFinclusion = 4.02 (positive tendency).
      • Explanations without feedback: g = −0.31 (95% CI −0.48 to −0.11), BFinclusion = 17.14 (clear negative effect).
      • Contrast when explanations present (feedback vs no feedback): g = 0.60 (95% CI 0 to 0.94), PD = 97% (strong evidence feedback improves synergy when explanations are provided).
  • Interpretation: AI explanations appear beneficial only when participants can validate AI reliability via outcome feedback; explanations without feedback can worsen human–AI performance.
  • Robustness checks: Adding moderators used by Vaccaro et al. did not change the directional pattern.
  • Caveats:
    • Small fraction of studies with feedback and limited trial-level reporting.
    • Most studies report average performance across trials, obscuring within-subject learning dynamics.
    • Feedback usefulness depends on task structure (timeliness, noise, feature visibility).

Data & Methods

  • Source data: Vaccaro et al.’s systematic review and meta-analysis (74 studies). Authors merged their new codes with Vaccaro et al.’s open data.
  • Coding: Six authors coded learning-relevant design features (e.g., outcome feedback, practice trials); each condition double-coded, disagreements adjudicated.
  • Meta-analytic method: Robust Bayesian Model Averaging (RoBMA) with a three-level hierarchical structure (effect sizes nested in experiments nested in studies). Default RoBMA priors used.
  • Estimation details: Spike-and-slab algorithm; at least 45k posterior samples after burn-in; convergence monitored by Gelman–Rubin ˆR < 1.05. Contrasts derived from conditional marginal posterior distributions.
  • Data & code availability: Openly available (OSF link provided in paper).
  • Supplemental analyses: Limited trial-level analyses in three studies with available time-series data showed participants could learn to align with the better-performing AI when informative signals were present.

Implications for AI Economics

  • Reassess empirical estimates of complementarities: Economists modeling human–AI complementarities (e.g., productivity gains, task allocation) should account for learning dynamics. Cross-sectional estimates from studies lacking feedback likely understate potential gains from human–AI teams.
  • Importance of feedback infrastructure: In organizational settings, investing in outcome feedback mechanisms (rapid, accurate performance feedback) can unlock complementarities between workers and AI systems. Cost–benefit analyses should include the value of increased synergy due to learning.
  • Design of field experiments and lab-in-the-field studies: When evaluating AI interventions in markets or firms, randomize feedback and explanation treatments and collect trial-level/longitudinal data to capture dynamic adaptation and learning effects.
  • Incentive and training policies: Policy and management interventions (training, performance transparency, explicit explanations of model behavior) should be paired with feedback loops to ensure workers can calibrate reliance on AI and improve joint performance.
  • Measurement and evaluation: Standardize performance metrics and report participant-level and time-resolved outcomes. Economists estimating returns to AI adoption should use measures that reflect performance after learning (not only initial trials).
  • Regulation and consumer protection: Regulators evaluating AI-assisted decision systems (e.g., credit scoring, health diagnostics) should consider how feedback availability and explainability affect downstream human behavior and aggregate outcomes; lack of feedback may produce systematically suboptimal human reliance.
  • Research agenda for applied economists: Incorporate models of endogenous learning and adaptive reliance into structural models of human–AI interaction; estimate how learning frictions and feedback technologies change equilibrium adoption, task assignment, wages, and welfare.

Actionable recommendations for researchers and practitioners: - In experiments and deployments, provide timely, accurate outcome feedback and pair it with interpretable explanations. - Collect and report trial-level and participant-level data to observe learning trajectories. - When assessing human–AI complementarities, evaluate performance after sufficient exposure (not only initial interactions). - Include feedback and explanation design as key policy levers in cost–benefit and regulatory analyses.

Assessment

Paper Typereview_meta Evidence Strengthmedium — Uses a large set of experimental studies and systematic re-coding to detect consistent patterns, which gives suggestive evidence that feedback matters for synergy; however, comparisons across heterogeneous studies are observational (no random assignment of feedback/explanation at the meta-study level), so confounding by task type, AI quality, sample composition, or other design choices could drive results. Methods Rigormedium — Rigorous re-analysis and subgroup/meta-regression methods applied to an existing meta-analytic sample; strengths include systematic coding of design features and use of quantitative synthesis, but rigor is limited by reliance on the original studies' reporting, heterogeneity in outcomes/measures, potential coding subjectivity, and lack of pre-registered experimental variation on the key moderators. SampleThe dataset comprises the 74 human–AI interaction experiments included in the prior meta-analysis—laboratory and online behavioral studies across multiple tasks and domains where humans received AI advice; studies varied in participant pools (students, crowdworkers, sometimes professionals), task types, AI accuracy and interface features, and whether they provided outcome feedback or explanations. Themeshuman_ai_collab skills_training IdentificationMeta-analysis / re-analysis of 74 experimental studies: the authors coded study-level design features (e.g., presence of outcome feedback, presence of AI explanations) and used subgroup comparisons and meta-regressions to associate those features with measured human–AI synergy; identification is associational across studies rather than from randomized variation in these features. GeneralizabilityBased on lab/online experiments rather than field or firm-level productivity data, Heterogeneous tasks, AI systems, and outcome measures reduce direct extrapolation to specific real-world settings, Comparisons of studies with vs. without feedback are observational and may be confounded by other study characteristics, Many original studies likely use convenience samples (students, MTurk) limiting population external validity, Possible publication/reporting bias and small-sample studies among the 74 could bias aggregate estimates

Claims (8)

ClaimDirectionOutcomeConfidence & EvidenceDetails
A recent meta-analysis showed that, on average, human-AI combinations do not outperform the better individual agent. Decision Quality null_result performance of human-AI combinations relative to the better individual agent (human or AI)
Reading fidelity high
Study strength high
n=74
0.4
The pessimistic conclusion from the prior meta-analysis arises from insufficient attention to human learning in experimental designs. Decision Quality negative estimated human-AI synergy as biased by experimental design (i.e., underestimation due to lack of human-learning-facilitating features)
Reading fidelity high
Study strength speculative
n=74
0.04
Most previous research overlooked design features that foster human learning, such as providing outcome feedback to participants. Skill Acquisition negative presence of outcome feedback (design feature) in human-AI experimental studies
Reading fidelity high
Study strength high
n=74
0.4
Studies that provided outcome feedback show tentatively higher human-AI synergy than those without outcome feedback. Decision Quality positive human-AI synergy (performance advantage of human-AI pairings)
Reading fidelity high
Study strength medium
n=74
0.24
Feedback paired with AI explanations tends to yield positive synergy, while explanations without feedback were linked to negative synergy. Decision Quality mixed human-AI synergy as a function of (a) outcome feedback and (b) presence of AI explanations
Reading fidelity high
Study strength medium
n=74
0.24
The current literature underestimates the potential of human-AI collaboration because it predominantly relies on paradigms that do not facilitate human learning, which hinders humans from effectively adapting their collaboration strategies. Decision Quality negative estimated potential (performance) of human-AI collaboration as reported in the literature
Reading fidelity high
Study strength medium
n=74
0.24
Human-AI explanations increase synergy only when humans can learn to verify the AI's reliability through feedback; explanations without feedback do not increase (and may decrease) synergy. Decision Quality mixed effect of AI explanations on human-AI synergy conditional on presence of outcome feedback
Reading fidelity high
Study strength medium
n=74
0.24
Research in human-AI interaction should shift paradigms to explicitly address human learning in order to better understand and support successful human-AI collaboration. Research Productivity positive research practices and resulting quality of evidence about human-AI collaboration
Reading fidelity high
Study strength speculative
n=74
0.04

Notes