Evidence (143 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
20058 claims
Filter claims →
Productivity
17184 claims
Filter claims →
Governance
16099 claims
Filter claims →
Human-AI Collaboration
16034 claims
Filter claims →
Innovation
10501 claims
Filter claims →
Org Design
10496 claims
Filter claims →
Labor Markets
6444 claims
Filter claims →
Skills & Training
5385 claims
Filter claims →
Inequality
4148 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 1820 | 479 | 278 | 1820 | 4588 |
| Organizational Efficiency | 2711 | 616 | 401 | 173 | 3922 |
| Governance & Regulation | 2075 | 886 | 459 | 246 | 3714 |
| Technology Adoption Rate | 1467 | 530 | 258 | 206 | 2488 |
| Decision Quality | 1281 | 496 | 289 | 152 | 2228 |
| Output Quality | 1227 | 447 | 207 | 138 | 2025 |
| AI Safety & Ethics | 634 | 754 | 207 | 83 | 1688 |
| Research Productivity | 826 | 241 | 114 | 422 | 1624 |
| Firm Productivity | 1052 | 154 | 163 | 66 | 1441 |
| Task Allocation | 685 | 211 | 331 | 99 | 1335 |
| Market Structure | 433 | 423 | 242 | 46 | 1150 |
| Innovation Output | 639 | 91 | 105 | 34 | 871 |
| Task Completion Time | 476 | 113 | 43 | 36 | 672 |
| Firm Revenue | 445 | 126 | 58 | 25 | 656 |
| Skill Acquisition | 364 | 119 | 109 | 34 | 626 |
| Consumer Welfare | 288 | 167 | 104 | 31 | 592 |
| Employment Level | 214 | 140 | 174 | 50 | 582 |
| Error Rate | 230 | 251 | 35 | 16 | 535 |
| Fiscal & Macroeconomic | 268 | 136 | 71 | 50 | 532 |
| Inequality Measures | 100 | 307 | 96 | 12 | 515 |
| Worker Satisfaction | 221 | 173 | 60 | 30 | 484 |
| Automation Exposure | 155 | 138 | 65 | 36 | 398 |
| Regulatory Compliance | 171 | 120 | 30 | 13 | 335 |
| Developer Productivity | 222 | 58 | 27 | 13 | 321 |
| Team Performance | 188 | 56 | 50 | 24 | 320 |
| Wages & Compensation | 146 | 104 | 46 | 16 | 312 |
| Training Effectiveness | 207 | 41 | 21 | 26 | 298 |
| Job Displacement | 23 | 153 | 52 | 4 | 232 |
| Hiring & Recruitment | 102 | 57 | 30 | 11 | 202 |
| Skill Obsolescence | 16 | 102 | 24 | 6 | 148 |
| Creative Output | 71 | 42 | 23 | 6 | 143 |
| Social Protection | 57 | 30 | 11 | 3 | 101 |
| Labor Share of Income | 29 | 42 | 24 | 2 | 97 |
| Worker Turnover | 43 | 29 | 6 | 4 | 82 |
| Industry | — | — | — | 1 | 1 |
Increased individual performance from generative-AI assistance can coincide with reduced similarity-adjusted diversity of population-level outputs.
The paper cites an online short-story experiment by Doshi and Hauser (2024), reporting improved third-party evaluations of individual stories while AI-assisted stories became more similar to one another.
Mathematical creativity comprises at least four mechanistically distinct modes: reflexive mathematics, analogical mathematics, problem-driven mathematics, and bridging distant domains.
Conceptual taxonomy developed from historical case studies and an architecture-level discussion of transformer-based systems; the paper explicitly states that the list is not exhaustive.
The two AI-AI co-creation configurations showed minimal differences in creativity and novelty, except that complementary roles outperformed identical roles on creativity in task 3.
Pairwise comparisons between complementary-role and identical-role AI-AI conditions found no significant differences in creativity for tasks 1 and 2 or in novelty for any task; complementary roles were superior on creativity in task 3 (p < .01).
Suno and Lyria 3 do not converge on a common acoustic profile; each system is more acoustically distant from the other than two random human subsamples would typically be.
Cross-system comparison in the 72-feature MIR space, benchmarked against distances between random human subsamples.
Cognitive and logical interaction styles are associated with greater novelty, whereas logical and informative interaction styles are associated with greater usefulness.
Human–ChatGPT dialogues were coded along cognitive, logical, and informative dimensions, and the resulting posts were evaluated independently for novelty and usefulness.
The benefits of AI-assisted creativity vary according to individual skill level, with moderate-skill individuals deriving the greatest creative benefit from AI support.
The first empirical component examined creative outputs from 90 UK participants balanced across three education-defined skill groups, with and without optional ChatGPT assistance.
Informative human–AI interactions are positively associated with the usefulness of creative outputs, but not with their novelty.
Recorded dialogues were coded for the volume and relevance of information provided to ChatGPT, and posts were independently rated on novelty and usefulness.
Cognitive human–AI interactions are positively associated with the novelty of creative outputs, but not with their usefulness.
Recorded human–ChatGPT dialogues were coded for cognitive interaction, while the resulting social-media posts were independently evaluated on novelty and usefulness using post-level ratings.
The structured review found generally positive associations between AI capability and organizational creativity and firm performance, but measurement approaches vary substantially across studies.
The thematic synthesis reports prior empirical applications with positive associations while noting differences in the number and labeling of dimensions and in reflective versus formative measurement models.
Increasing sampling temperature increased model output diversity only up to a plateau, and the configuration that exceeded the plateau did so with substantial gibberish and poor format adherence.
Simulation sweep over sampling temperatures from 0.5 to 2.0. Diversity plateaued around 0.50; GPT-4o mini at temperature 2.0 generated 14.3% non-dictionary output words and had format adherence below 75%.
Prompting simulated L2 personas in their native languages increased simulated diversity but did not close the gap with human pools.
Simulation comparison across three models using English-only output versus native-language generation followed by translation. For GPT-4o mini, mean centroid similarity changed from 0.698 to 0.606, Δ = -0.092, p < 0.001, but the best native-language simulation remained less diverse than the human AI-ideation pool.
AI refinement preserved diversity in the metaphors themselves but made the explanations more similar to one another.
Separate embedding analysis of metaphors alone versus metaphors combined with explanations. When explanations were included, AI refinement shifted toward the centroid relative to human-only by b = 0.046, p < 0.001.
Looking for solutions together encourages members to examine more alternatives but narrows the range of these ideas, while working independently generates greater variety but fewer alternatives being examined.
Synthesis of the empirical patterns observed in the online competition data: joint search associated with higher count of attempts but lower exploration/variety, independent search associated with greater variety but fewer attempts (details/samples not provided in abstract).
Among top-selling books, those with substantial AI text draw on more distinctive language from existing books than do books with no AI text; for AI-text books, overlap rises with revenue, a gradient not detected for non-AI books.
Language-overlap analysis comparing top-selling books classified by detected AI content (>25%) to existing books; correlation between overlap measure and revenue for AI-text books reported, with no similar correlation for non-AI books.
Approach motivation (BAS Drive) moderates whether interactive partnership benefits originality.
Moderation analysis reported from the pilot (N = 62) showing interaction between BAS Drive (a measured personality/motivation scale) and the effect of interactive partnership on originality.
Reasoning models roam a wider hypothesis space, yet no model class spontaneously proposes null hypotheses — a move humans make more freely.
Model-output analysis comparing 'reasoning' vs 'non-reasoning' classes on hypothesis-space breadth and presence/absence of null hypotheses; human responses used as comparison.
AI advances science through structurally distinct creative pathways rather than a single mechanism; the creative pathway depends on how AI is incorporated into the research process.
Interpretation synthesized from observed heterogeneity in creativity outcomes across classified AI research modes (Tool-oriented vs Adaptation-oriented) in the >1M publication analysis.
Through a pre-registered randomized control trial, we show that incentives mediate AI's homogenizing force in a creative writing task where participants can use AI interactively.
Pre-registered randomized controlled trial (experimental design) conducted on a creative writing task with interactive AI use (details such as sample size not provided in excerpt).
The effect of increasing the share of AI-automated R&D tasks is non-monotonic: firms initially target more radical innovations, but beyond a threshold of human-AI complementarity, they shift the focus toward incremental innovations.
Analytical comparative-statics in the theoretical model: varying the fraction of R&D tasks performable by AI yields a non-monotonic relationship between AI task-share and optimal recombination distance, with a threshold determined by human-AI complementarity.
Higher AI productivity encourages more distant recombinations, if the direct facilitation effect is stronger than the indirect effect due to intensified competition from rivals.
Comparative-static result from the analytical model: the paper derives a condition comparing the direct facilitation effect of AI on accessing distant knowledge and the indirect effect from increased competition; when the former dominates, equilibrium recombination distance increases with AI productivity.
AI usage has dual effects on employees: it can both enhance innovative behavior and predict disengagement, as revealed by a dual-path (SOR-based) model.
Interpretation/synthesis from the four-stage longitudinal study of 285 finance professionals using a dual-path model based on SOR theory (combining the mediation and moderation results).
Current systems stall when they must select a new type of mathematical object or practice to formalize, even if they can search and score candidates once the candidate type is specified.
Comparison of systems such as AlphaEvolve and FunSearch, which search spaces whose object types are fixed by the problem statement, with historical examples of Turing, Boole, Gentzen, and Gödel selecting practices as objects of formalization.
Current transformer-based systems concentrate their mathematical competence in recombination and search over existing building blocks, rather than in inventing genuinely new conceptual primitives.
The paper infers this from the predominance of existential results in autonomous AI mathematical discovery and from a mechanistic study finding narrow heuristics in language-model arithmetic; it explicitly characterizes the architectural account as a conjecture.
The four proposed mechanisms of mathematical creativity are likely non-substitutable, so competence in existing modes such as search, deduction, and straightforward cross-domain transfer does not transfer to modes requiring genuinely new conceptual primitives.
Theoretical argument based on the distinction between recombination over a fixed candidate type and invention of a new candidate type; the paper acknowledges that the claim is not proved outright.
Stepwise Guidance reduced new idea generation relative to the other conditions.
Researchers counted unique idea units generated across two iterations and compared quantitative design behavior changes across conditions.
The observed homogenization patterns are more consistent with learned system priors than with prompt constraints.
The paper used genre-name-only null prompts to isolate system-level priors from prompt-engineering effects and reports that both systems exhibited homogenizing tendencies despite minimal prompting.
Suno collapses acoustic distinctions between genres without compressing within-genre acoustic spread.
Comparison of Suno-generated and human-produced tracks across four genres using 72 MIR features and diagnostics of within-genre dispersion and between-genre separability.
Lyria 3 reduces within-genre acoustic diversity relative to human-produced music.
Black-box audit comparing 100 Lyria 3 tracks per genre with matched human corpora across 72 MIR features and diagnostics of dispersion, redundancy, and separability, covering Afrobeats, K-pop, Dance Pop, and Heavy Metal.
Across three additional AI-creativity studies and multiple diversity metrics, AI-generated ideas were significantly less diverse than human-generated ideas.
Study 2b conducted a meta-analysis of data from Stevenson et al. (2022), Hubert et al. (2024), and Boussioux et al. (2024), applying multiple diversity measures.
AI-generated idea sets were substantially less diverse than human-generated idea sets, meaning that the AI ideas were more similar to one another.
Study 2a compared the diversity of human- and AI-generated ideas from Study 1, operationalizing diversity as semantic distance between ideas in a set.
AI-generated ideas had slightly lower novelty than human-generated ideas, but the decrease was small.
Study 1c used human judges to rate the novelty of AI-generated and human-generated product ideas.
The strongest reported simulation, Claude Sonnet 4.6, recovered only 75.3% of the diversity of the human-only pool.
Comparison of the simulated metaphor pools with human pools across the three model families.
Every LLM-simulated metaphor pool was less diverse than every human pool, including the most homogenized human AI-ideation pool.
Non-preregistered adversarial simulation using personas based on participants' backgrounds and three model families: GPT-4o mini, Claude Sonnet 4.6, and Qwen3.5-27b. All comparisons were reported as p < 0.001.
AI ideation compressed both L1 and L2 writers' metaphor pools and made the L2 collective-diversity advantage statistically undetectable.
Comparison of L1 and L2 pools under AI ideation. The group means were 0.457 for L1 and 0.451 for L2, with Δ = 0.006 and p = 0.481.
AI ideation produced the most homogeneous human metaphor pools, whereas AI refinement produced pools that were statistically indistinguishable from the human-only condition.
Multilevel regression of collective metaphor diversity measured by mean centroid similarity. AI ideation differed from human-only by b = 0.036, p < 0.001; AI refinement differed by b = 0.007, p = 0.516. Higher centroid similarity indicates lower diversity.
At the collective level, LLMs aggregate knowledge into a unified distribution rather than exhibiting the knowledge partitioning inherent to human populations (where individuals occupy distinct regions of knowledge space).
Theoretical account plus empirical analyses comparing distributions of ideas across multiple independent LLM samples versus multiple human participants, showing greater overlap / less partitioning in LLM outputs.
At the individual level, LLMs exhibit fixation: early outputs constrain subsequent ideation.
Theoretical argument plus empirical tests reported in the paper showing within-generator sequences where early outputs reduce later idea novelty/diversity for LLMs.
Ideas generated by independent samples of humans tend to be more diverse than ideas generated from independent LLM samples.
Empirical comparison reported across the paper's studies (four studies) that measure idea diversity between independently sampled human participants and independently sampled LLM outputs.
Reliance on generative AI can reduce cultural variance and diversity, especially in creative work.
The paper asserts this as a motivating premise and studies it using an agent-based model and evolutionary game theory (theoretical/modeling evidence).
Compared to their counterfactuals that search apart, groups searching together exhibit less exploration in their search outcomes.
Empirical analysis of naturally occurring data from a strongly incentivized online competition platform; groups searching together were compared to counterfactual groups that searched apart (method details and sample size not provided in abstract).
Traditional packaging design models based on experience, templates, functionality, or imitation produce homogenized designs, long development cycles, high costs, and limited ability to meet contemporary demands for personalization, sustainability, and e-commerce.
Paper's comparative background and motivation section describing limitations of traditional approaches; positioned against the study's empirical comparison.
Self-reported cognitive outsourcing predicts lower originality specifically in human-human dyads.
Correlation / regression result from the in-person pilot (N = 62) reporting that self-reported cognitive outsourcing is associated with lower originality in human-human dyads but not in other conditions.
More innovative creators are especially harmed under the strong-IP regime — a phenomenon the paper terms the "originality penalty."
Analytical result derived from the static game model in the paper highlighting differential effects by creator innovativeness; theoretical characterization labeled "originality penalty."
A regime of strong intellectual property rights, modeled as a static Stackelberg game, also fails to provide adequate creative incentives (it underpowers creative incentives).
Theoretical analysis using a static Stackelberg-game model developed in the paper; analytical results show reduced creator incentives under this regime.
Non-reasoning LLMs collapse into a narrow 'hivemind' of similar ideas.
Comparative analysis of idea outputs from different LLM classes showing reduced diversity/similarity concentration for non-reasoning models (as described in results).
GenAI usage significantly decreased creativity-relevant skills.
Experiment with 82 participants reported in the paper; authors report a statistically significant decrease in measures of creativity-relevant skills for participants using GenAI.
In deployed settings, the effects of AI systems on human agency, creativity, and institutional well-being emerge over time, shaped by repeated interaction, reuse, and integration into real-world workflows, and these dynamics are rarely visible through pre-deployment evaluation or isolated prompt–response analysis.
Argumentative observation based on conceptual reasoning; no empirical data or sample size reported.
Generated ideas often degrade after implementation.
Paper statement about the gap between idea generation and implemented results reported in the Creation-phase analysis; no quantified follow-up study reported in the excerpt.
Using LLMs led to fewer creative moments observed in participants (p=0.002).
Within-subject comparison between LLM-assisted and unassisted conditions with reported p-value p=0.002. Study sample N=20.
Patent text similarity analysis confirms a 'homogenization trap' (AI-associated increases in patent-text similarity).
Text-similarity analysis of patent documents reported in the paper showing increased patent similarity associated with AI use.