Evidence (320 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
20058 claims
Filter claims →
Productivity
17184 claims
Filter claims →
Governance
16099 claims
Filter claims →
Human-AI Collaboration
16034 claims
Filter claims →
Innovation
10501 claims
Filter claims →
Org Design
10496 claims
Filter claims →
Labor Markets
6444 claims
Filter claims →
Skills & Training
5385 claims
Filter claims →
Inequality
4148 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 1820 | 479 | 278 | 1820 | 4588 |
| Organizational Efficiency | 2711 | 616 | 401 | 173 | 3922 |
| Governance & Regulation | 2075 | 886 | 459 | 246 | 3714 |
| Technology Adoption Rate | 1467 | 530 | 258 | 206 | 2488 |
| Decision Quality | 1281 | 496 | 289 | 152 | 2228 |
| Output Quality | 1227 | 447 | 207 | 138 | 2025 |
| AI Safety & Ethics | 634 | 754 | 207 | 83 | 1688 |
| Research Productivity | 826 | 241 | 114 | 422 | 1624 |
| Firm Productivity | 1052 | 154 | 163 | 66 | 1441 |
| Task Allocation | 685 | 211 | 331 | 99 | 1335 |
| Market Structure | 433 | 423 | 242 | 46 | 1150 |
| Innovation Output | 639 | 91 | 105 | 34 | 871 |
| Task Completion Time | 476 | 113 | 43 | 36 | 672 |
| Firm Revenue | 445 | 126 | 58 | 25 | 656 |
| Skill Acquisition | 364 | 119 | 109 | 34 | 626 |
| Consumer Welfare | 288 | 167 | 104 | 31 | 592 |
| Employment Level | 214 | 140 | 174 | 50 | 582 |
| Error Rate | 230 | 251 | 35 | 16 | 535 |
| Fiscal & Macroeconomic | 268 | 136 | 71 | 50 | 532 |
| Inequality Measures | 100 | 307 | 96 | 12 | 515 |
| Worker Satisfaction | 221 | 173 | 60 | 30 | 484 |
| Automation Exposure | 155 | 138 | 65 | 36 | 398 |
| Regulatory Compliance | 171 | 120 | 30 | 13 | 335 |
| Developer Productivity | 222 | 58 | 27 | 13 | 321 |
| Team Performance | 188 | 56 | 50 | 24 | 320 |
| Wages & Compensation | 146 | 104 | 46 | 16 | 312 |
| Training Effectiveness | 207 | 41 | 21 | 26 | 298 |
| Job Displacement | 23 | 153 | 52 | 4 | 232 |
| Hiring & Recruitment | 102 | 57 | 30 | 11 | 202 |
| Skill Obsolescence | 16 | 102 | 24 | 6 | 148 |
| Creative Output | 71 | 42 | 23 | 6 | 143 |
| Social Protection | 57 | 30 | 11 | 3 | 101 |
| Labor Share of Income | 29 | 42 | 24 | 2 | 97 |
| Worker Turnover | 43 | 29 | 6 | 4 | 82 |
| Industry | — | — | — | 1 | 1 |
Extending personalization from individuals to teams creates asymmetric desiderata: feasibility is based on the union of members' capabilities, novelty on the union of prior work, and alignment closer to the intersection of members' communities.
Conceptual team formulation replacing individual researcher u with team T and defining team-level context and utility functions.
The performance of a human–AI combination cannot be inferred mechanically from the independent performance of the human and the AI.
The paper summarizes meta-analytic evidence from Vaccaro et al. (2024); the supplied text does not report the meta-analysis sample size or effect estimate.
Managerial enforcement increased Grok's cooperation from 16% at baseline to 100%, while Qwen's cooperation increased only to 56–76% and its deception remained present.
Manager-Type experiment comparing no manager with fixed, elected, and rotating managers in homogeneous groups.
Without a manager, one defector could be contained by cooperative peers, but groups containing both Qwen and Grok experienced late-game collapse toward zero contribution.
Mixed-NoManager experiment rerunning heterogeneous compositions without enforcement; comparisons included four Claude agents with one Qwen and groups containing both defectors.
At baseline, Claude, GPT-4o, and Gemini were highly cooperative, DeepSeek was conditionally cooperative, and Grok and Qwen were defectors.
Baseline public-goods-game condition with six LLM families in homogeneous groups; cooperation was measured across 20 rounds with five agents per group.
A fully codifiable research team reaches its peak size when its effective automated task share reaches the threshold s* = (β − r)/(β(1 − r)), where r is the AI-to-member task-cost ratio.
Analytical derivation in Corollary 1 and Appendix A.3.
In the model, research-team size is quasi-concave in AI capability: it can rise at most once and fall at most once, with a single turning point from increasing to decreasing.
Theoretical result (Theorem 1), derived from the model of leader attention, execution costs, and member employment.
Only the web_shop and travel comparisons between specialized agents and the best general baseline were statistically significant at p < 0.05.
Table IX reports significance results for five specialized-agent versus best-baseline comparisons; retail and code_gen confidence intervals included zero, while web_shop and travel were marked significant.
For identical tasks and models, changing the organizational topology shifts performance scores by more than 30 points and can double wall-clock time.
The abstract reports results across 100 runs while holding tasks and models fixed and varying collaboration topology.
The results also reveal clear divergences in human-AI interaction.
Empirical claim from the confirmatory user study reporting systematic differences between human-AI and human-human collaboration patterns; no quantified divergences or sample size given in the excerpt.
Pull request presentation and reviewer interaction dynamics play a role in human-AI collaborative software development.
Synthesis/interpretation of empirical findings from analysis of PR descriptions, reviewer engagement, response timing, and merge outcomes for five AI coding agents in the AIDev dataset.
Pull request description styles are associated with differences in reviewer engagement.
Correlation/association analysis between PR description characteristics and reviewer interaction metrics in the AIDev dataset of PRs from five AI coding agents.
Within hybrid groups, both human and AI agents systematically adjust their strategies relative to single-agent (human-only or AI-only) conditions, indicating higher-order interaction effects where agents adapt to each other's presence.
Behavioral comparison within the controlled experiment between agents' actions in hybrid groups versus their behavior in single-agent conditions (methods indicate observation of strategy adjustments; no sample size or statistical details provided in abstract).
In human-AI edit chains, AI introduces high-throughput changes while humans act as security gatekeepers.
Edit-chain analysis comparing volume/rate of AI-originated edits to human edits and observation of human reviews acting as gatekeeping steps.
AI teammates with distinct communicative personas (supportive or contrarian) exerted robust social effects on collaboration.
Experimental manipulation of AI persona (supportive vs contrarian) and reported effects on collaboration outcomes across tasks.
Two stable human-AI collaboration patterns emerged from production deployment: one-click rollout for high-confidence changes (60% of cases) and commandeer-revise for complex decisions (40%).
Paper reports emergent collaboration patterns and their relative frequencies (60% vs 40%) observed during production deployment.
Communication mediates the relationship between team organization and performance outcomes (productivity and quality) — this was the study hypothesis.
Stated hypothesis in the paper (abstract).
In a proportional large-market limit the five-protocol ordering survives exactly: consensus discovery vanishes while blind, market, private, and portfolio search converge to 0.500, 0.547, 0.847, and 0.874.
Asymptotic analysis (proportional large-market limit) of the model yielding closed-form limiting discovery probabilities for each protocol and the statement that consensus goes to zero.
The centralized planner gain rises strictly with copying, and in the canonical environment the symmetric market overtakes decentralized report-following at copying probability c = 0.788462.
Analytic comparative statics in the latent common-cue (copying) model applied to the canonical instance (16 boxes, 1 target, 8 searchers), yielding threshold c = 0.788462 where symmetric market > decentralized report-following.
In the canonical instance the anonymous symmetric equilibrium achieves discovery probability 0.5991: strictly above consensus, but below both private search and the planner.
Computed equilibrium performance for the canonical model (16 boxes, 1 target, 8 searchers) giving discovery probability 0.5991 and comparison to consensus (0.3835), private search (0.8322), and planner (0.8594).
The cooperative effects of the prosocial AI interventions were short-lived, fading after the first few rounds.
Temporal analysis of contributions over rounds in the iterated game showing decay of the prosocial AI effect after the initial rounds (reported in the experiment with N = 1,283).
Personality effects depend critically on task structure.
Authors compared the impact of personality manipulation across three distinct task domains (structured coding, open-ended research collaboration, competitive bargaining) and report differing outcomes by domain. Abstract does not provide numeric sample sizes or statistical details.
Du et al. (2026) find that information-based team faultlines can enhance proactive behavior via deep information processing, while AI adoption moderates and mitigates the negative effects of social-based faultlines on team cooperation.
Information-processing theoretical framing and empirical analysis reported in the paper (study type and sample size not specified in the excerpt).
Human–AI complementarity in finance is conditional rather than automatic, depending on task structure, private information, feedback quality, incentives, explanation design, and governance.
Synthesis of literature from finance, management, HCI, and AI showing moderating factors for complementarity (conceptual integration; no unified empirical sample size reported).
There is a suggestive non-linear relationship between embodiment and team performance.
Analysis reported in the paper indicating a non-linear (not strictly monotonic) association between degree of agent embodiment (Box, Avatar, humanoid) and measured team performance; described as 'suggestive' in the abstract, without quantified functional form or statistics included there.
Artificial agents have an uneven impact on team outcomes, with some mixed human–AI teams performing exceptionally well and others markedly worse.
Observed performance outcomes across mixed human–AI teams in the escape room experiment, showing high between-team variability; exact sample size and statistical details not provided in the abstract.
cBCI synergy is heavily contingent on the temporal dynamics of trust, providing a critical framework for designing dynamically gated Human-AI systems.
Interpretive/concluding claim based on experimental results (timing-dependent failure modes, Oracle gating, Hybrid Fusion effects) reported in the study.
AI timing dictates the mechanism of team failure: high-speed AI interventions risk inducing reflexive blind compliance while delayed interventions can induce ambiguous cognitive conflict.
Synthesis claim derived from experimental contrasts between Fast/Less-Accurate and Slow/Accurate AI conditions and observed human/team behaviors (blind compliance vs. delayed conflict).
GenAI enables small teams to expand capacity while creating new dependencies and coordination logics.
Empirical finding from 17 interviews indicating both expanded capacity and emergent dependencies/coordination needs.
Across studies, causal modeling reveals that cognitive alignment systematically drives attentional coordination in successful collaboration, while mismatches between effort and attention characterize unproductive regulation.
Synthesis of causal inference results from the three studies using time-series measures (JME, JVA) and episode-based analyses across the pooled dataset (182 dyads total).
Augmentation is bounded rather than linear (i.e., human-AI augmentation shows diminishing or negative returns past a balanced zone).
Synthesis of interview themes across 34 cases producing the bounded-augmentation / curvilinear conceptualization.
Mediators such as trust, cohesion and accountability are reshaped when AI-generated contributions enter collaboration.
Thematic evidence from interviews indicating changes in trust, cohesion and accountability dynamics associated with the introduction of AI outputs into team collaboration.
Social (leadership engagement, trust, ownership, mediation and alignment) and technical (automation, creation, reliability, distraction and integration) subsystems combine to enable or erode team effectiveness, summarized in an e-leadership–AI orientation matrix.
Analytic synthesis from thematic coding (Gioia-informed) of interview data producing a conceptual matrix mapping social and technical factors to outcomes.
Analysis identifies a curvilinear pattern of bounded augmentation, where effectiveness peaks in a zone of balanced use but declines under under-use and over-reliance.
Thematic (Gioia-informed) analysis of 34 semi-structured interviews with project managers across five UK industries; pattern emerges from cross-case coding and synthesis.
We identify significant differences between human and AI negotiation behaviors, finding that humans favor lower-complexity deals and are significantly less reliable partners compared to LM-based agents.
Results from the user study comparing human vs LM-based agent negotiation behavior (statements in the results section).
The authors identify ten evaluation practices that teams use, ranging from lightweight interpretive checks to formal organizational processes (examples: qualitative user reviews, red-team testing, A/B experiments, telemetry/log analysis, structured annotation, governance/meta-evaluation).
Thematic coding of 19 interview transcripts produced a taxonomy enumerating ten practices (paper reports the taxonomy as an outcome).
Teamwork partner type moderates the effect of service empathy on collaboration proficiency (i.e., the impact of service empathy on proficiency differs by human vs AI partner).
Reported interaction/moderated-mediation analyses from the online experiment (n = 861) indicating a significant partner-type × service-empathy interaction predicting collaboration proficiency.
Employees' emotional state significantly moderates the relationship between partner type (human vs AI) and collaboration proficiency.
Moderation analyses reported from the same online experimental dataset (n = 861), testing interaction terms between partner type and measured employee emotion on collaboration proficiency; authors report a significant moderating effect.
AI excels at hypothesis generation but cannot replace scientific reasoning and experimental validation; human expertise remains essential.
Argument and case examples in the paper showing AI-generated hypotheses requiring human-led experimental design, interpretation, and validation.
Misrouted failure attribution can cause unnecessary components to be reopened and can regress working code.
An incident in which a two-component fault woke five components and three bystanders regressed working code; another incident routed a service-owned stale-state failure to two components that only declared the scenario.
The false progress-breaker trip drove one run from all six components being complete or progressing to only three components remaining.
Scheduler summary lines and measured repair-round evidence; the leading component was killed while failing tests decreased from five to two and three tests newly passed.
Among users who used Copilot at least 100 times, adoption was associated with decreases in small-group emails sent, unique recipients, and rounds of email conversation.
Analysis of a separate dataset tracking emails sent to fewer than 10 recipients; the paper reports statistically significant decreases for all three measures at p<0.05.
Role ambiguity negatively predicts employee-AI collaboration.
Structural equation modeling of survey responses from 541 employees in Chinese high-technology firms.
Similarity-based cooperation becomes more difficult in games with more than two players, such as the Public Goods game.
The authors varied the number of players and compared cooperation across multiple cooperation problems, including the Public Goods game.
Private-only communication produced lower average cooperation than public or full communication, and reduced Gemini's cooperation from 55.5% without communication to 31.5%.
Communication-sweep comparison across four communication conditions: none, public-only, private-only, and full communication.
Illusion of alignment arises routinely in human collaboration: across 18 real meetings, participants confirmed an average of 2.89 hidden disagreements per meeting that they had not articulated during the original discussions.
Real-user study involving 18 meetings and 43 participants; participants reviewed detector-generated questions and confirmed previously unvoiced disagreements.
After adding an adoption-and-output lag of 0.5 to 2.0 years, the forecasted median peak date shifts to 2028.5, with a 90% interval from 2027.4 to 2029.6.
Monte Carlo forecast conditional on eventual full adoption, with an assumed diffusion lag uniformly spanning 0.5 to 2.0 years.
Under the paper's capability-to-task benchmark and priors, a fully codifiable research team is forecast to peak at a median date of 2027.2 on the capability clock, with a 90% interval of 2026.4 to 2028.1.
Monte Carlo simulation with 20,000 draws, using the METR software time-horizon trend, a task-duration prior, and priors for β and r.
Communication inflexibility in AI, limited shared mental models, and trust miscalibration are recurring barriers to reliable human-agent collaboration.
Narrative synthesis of 192 articles focused on communication, coordination, trust calibration, metacognition, and human-agent teaming outcomes.
Direct human-human interaction accounted for 32.4% of completed tasks under No-CA but only 11.6% under CA.
Classification of completed tasks into four interaction modes based on whether tasks crossed developer boundaries and whether a coding agent participated.