Evidence (7560 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
9875 claims
Filter claims →
Productivity
8807 claims
Filter claims →
Governance
7870 claims
Filter claims →
Human-AI Collaboration
7560 claims
Filtered →
Org Design
4892 claims
Filter claims →
Innovation
4781 claims
Filter claims →
Labor Markets
4004 claims
Filter claims →
Skills & Training
3308 claims
Filter claims →
Inequality
2332 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 870 | 233 | 116 | 1066 | 2363 |
| Governance & Regulation | 976 | 451 | 218 | 133 | 1809 |
| Organizational Efficiency | 949 | 224 | 144 | 88 | 1416 |
| Technology Adoption Rate | 764 | 287 | 141 | 122 | 1325 |
| Research Productivity | 501 | 152 | 74 | 362 | 1101 |
| Output Quality | 542 | 216 | 69 | 69 | 896 |
| Decision Quality | 387 | 198 | 94 | 54 | 740 |
| Firm Productivity | 513 | 67 | 101 | 27 | 714 |
| AI Safety & Ethics | 249 | 303 | 73 | 36 | 667 |
| Market Structure | 190 | 192 | 134 | 27 | 548 |
| Task Allocation | 243 | 77 | 91 | 36 | 452 |
| Innovation Output | 291 | 33 | 55 | 20 | 401 |
| Skill Acquisition | 206 | 72 | 65 | 21 | 364 |
| Employment Level | 133 | 63 | 115 | 22 | 335 |
| Fiscal & Macroeconomic | 153 | 79 | 52 | 32 | 323 |
| Task Completion Time | 206 | 37 | 12 | 15 | 272 |
| Firm Revenue | 179 | 52 | 29 | 5 | 266 |
| Consumer Welfare | 130 | 76 | 47 | 13 | 266 |
| Inequality Measures | 48 | 137 | 51 | 6 | 242 |
| Worker Satisfaction | 101 | 81 | 25 | 13 | 220 |
| Error Rate | 84 | 110 | 11 | 5 | 210 |
| Wages & Compensation | 98 | 47 | 30 | 10 | 185 |
| Regulatory Compliance | 88 | 73 | 17 | 7 | 185 |
| Automation Exposure | 66 | 64 | 33 | 16 | 182 |
| Team Performance | 105 | 29 | 30 | 11 | 176 |
| Training Effectiveness | 109 | 22 | 14 | 21 | 168 |
| Developer Productivity | 114 | 21 | 14 | 8 | 158 |
| Job Displacement | 12 | 90 | 24 | 1 | 127 |
| Hiring & Recruitment | 57 | 9 | 9 | 5 | 80 |
| Skill Obsolescence | 6 | 56 | 9 | 1 | 72 |
| Social Protection | 43 | 17 | 8 | 2 | 70 |
| Creative Output | 35 | 21 | 9 | 4 | 70 |
| Labor Share of Income | 18 | 21 | 17 | 1 | 57 |
| Worker Turnover | 15 | 16 | — | 4 | 35 |
| Industry | — | — | — | 1 | 1 |
Human Ai Collab
Remove filter
Current evidence does not support the simple claim that autonomous code generation automatically improves engineering outcomes.
Synthesis of mixed results from controlled studies, meta-analyses, and benchmarks reported in the paper (no single sample size given in abstract).
However, the exoplanet workflow is effectively tied with a strong combined-summary baseline, showing that decomposition does not always improve top-line performance.
Reported comparison between the coordinated workflow and a strong combined-summary baseline for exoplanet vetting indicating no meaningful improvement.
The paper draws on empirical studies from 2024–2026.
Methodological statement in the paper specifying the time window of empirical studies used in the analysis.
We compared the traits causing the incidents with the traits that 197 developers building AI systems for those tasks would have preferred.
Study design: comparison between trait set responsible for incidents (from incident reports) and stated developer preferences collected from a sample of 197 developers working on those tasks.
We compared the extracted traits with the traits that 202 workers highly familiar with those tasks would have preferred.
Study design: a comparison between LLM-extracted traits from incident reports and stated preferences from a sample of 202 workers familiar with the tasks.
We used an LLM-as-an-expert approach to extract the main traits of the AI systems involved in those incidents using an established framework of twelve traits.
Methods statement: applied a Large Language Model to code/extract AI system traits from the incident reports using an established 12-trait framework.
We analyzed 1,524 reports of incidents in which AI systems were used to perform 171 occupational tasks across 12 industry sectors.
Descriptive statement in paper: dataset comprised 1,524 incident reports, covering 171 occupational tasks and 12 industry sectors (dataset construction / corpus used for analysis).
The study uses PyQu to quantify changes across five quality attributes for Python code.
Methodological description: application of PyQu (an ML-based quality assessment tool for Python) to measure five quality attributes before and after refactoring edits.
From the observed diffs, we derive a taxonomy of 24 recurring change operations.
Manual/automated analysis of diffs from the studied agentic refactoring PRs to identify and categorize recurring change operations into a 24-item taxonomy.
Across 660 trials with Claude Code, code cleanliness does not change the agent's pass rate.
Empirical evaluation: 660 trials run using Claude Code on the minimal-pair repos with hidden tests; reported comparison of pass rates between clean and messy repo variants showing no change.
Each output is scored with a unified rubric covering task completion, correctness, compliance, and clarity.
Measurement approach stated in the abstract (unified rubric with listed dimensions).
The study uses three LLM systems: ChatGPT, Claude, and Grok.
Method description in the paper's abstract naming the three LLMs evaluated.
The evaluation covers four task types: summarization, planning, explanation, and coding.
Method description in the paper's abstract listing the four task types used for evaluation.
The study compares three prompt conditions: a raw prompt, a checklist-improved prompt, and a clarifying-question prompt.
Experimental design described in the paper (three prompt conditions stated in the abstract).
Large language models (LLMs) are widely used for open-ended tasks.
Stated as background/context in the paper's introduction; no quantitative data reported in the abstract.
This study used a controlled mixed-design experiment with 60 participants who completed analytical survival ranking tasks in multi-turn human–AI collaborations, with pre/post measurements and two types of prompting training (general or sycophancy-focused).
Methodological description in the paper's abstract/summary.
Reported empirical values are transformed through transparent indicators such as relative growth, CAGR, growth multipliers, stock-flow ratios, concentration ratios, and HHI.
Methodological description and application in the paper listing these specific indicators used to summarize public data on AI investment, adoption, robots, compute, and labour-market reallocation.
The study uses a conceptual-empirical quantitative diagnostic design rather than a causal econometric model.
Explicit methodological statement in the paper describing the design choice and rejecting causal econometric modeling in favor of diagnostics using public institutional data and transparent indicators.
The agentic economy is not yet a completed global order, but its transition pressure is measurable enough to require a distinct economic vocabulary, reproducible diagnostics, and future sector-level measurement.
Synthesis of diagnostic indicators (AI investment/adoption trends, robot stock, compute-energy coupling, labour reallocation measures) showing measurable transition pressures; conclusion drawn from the conceptual-empirical diagnostic.
Following PRISMA 2020 guidelines, searches across Google Scholar, Web of Science, Scopus, ScienceDirect, and CNKI yielded 1,562 initial records, of which 21 studies published between 2019 and 2026 met inclusion criteria.
Methodological description of the systematic literature review reported in the paper: initial records = 1,562; included studies = 21; publication years 2019–2026.
Small and medium-sized enterprises (SMEs) constitute over 98.5% of businesses in many economies including China.
Descriptive statistic reported in the paper's background/intro; source of the statistic not specified within the summary provided.
This study analyzes developments through April 2026.
Explicit timeframe statement in the paper's summary/introduction.
We conducted a randomized controlled experiment in which participants—analogs of early-career knowledge workers—were assigned to self-study a technical domain using either traditional resources or large-language-model (LLM) assistance.
Statement of experimental design in the paper (randomized controlled experiment assigning participants to either traditional resources or LLM assistance; participants described as analogs of early-career knowledge workers).
The formal semantics and proof-checked admission model are specified and under active development, with evaluation of the verified core reserved for future work.
Author statement in the paper about the current development status and that evaluation of the verified core is deferred to future work.
Skills can be mapped into three categories: those AI is absorbing, those needed to work alongside AI today, and those that make humans irreplaceable tomorrow.
Conceptual taxonomy offered in the chapter, based on labour market data and workplace evidence; presented as an analytical framework rather than a quantified finding.
Fear and hype about technological transitions are temporary.
One of five lessons drawn from historical analogy and labour market history as presented in the chapter.
Virtually every job is being touched by AI.
Stated in chapter summary; claimed on the basis of labour market data and emerging workplace evidence (no numeric sample given in excerpt).
Only 9% of jobs are fully automatable.
Reported directly in chapter; based on labour market data (specific data source and sample size not stated in the excerpt).
AI automates tasks, not jobs.
Conceptual argument in chapter drawing on labour market data and historical analogy; presented as a framing claim rather than a specific empirical estimate.
The study used a structured questionnaire (five-point Likert) administered to employees in AI-enabled organizations across various sectors and analyzed the data using SPSS (descriptive statistics, reliability analysis, correlation analysis, regression analysis).
Methods section summary provided in the paper (survey instrument description and analytical techniques).
We evaluate PRISM across 35 enterprise conversational agents over a three-week deployment period on the Yellow.ai V3 platform.
Statement in abstract: evaluation across 35 agents over a three-week deployment on Yellow.ai V3 platform (empirical deployment described).
AI deployment has limited effects on retrial rates.
Same randomized field experiment; retrial rates (repeat customer contacts) were measured and reported as showing limited/no substantive change under AI deployment.
Five structural characteristics define the Metis AI zone: consequential irreversibility, relational irreducibility, normative open texture, adversarial co-evolution, and accountability anchoring.
Theoretical specification and definition of five characteristics grounded in social science, philosophy, and humanitarian practice; no empirical prevalence or measurement reported.
The dominant discourse on AI limitations frames the boundary of AI capability as a divide between digital tasks (where AI excels) and physical tasks (where embodiment is required).
Statement in paper framing prevailing discourse; conceptual observation rather than empirical test (literature critique). No sample size reported.
Neither survey nor transcript-based measures of participation equity improved under LLM facilitation (an "illusion of inclusion").
Quantitative survey measures and transcript-based analyses of participation equity (e.g., measures of turn-taking, speaking/typing share) showed no improvement in equity metrics for facilitated conditions compared to controls across the experiments.
Across both studies, LLM facilitation did not significantly improve group consensus.
Experimental comparison across the two studies (total N=879) measuring agreement/consensus metrics for groups randomized to LLM facilitation versus other facilitators or no facilitation; reported null effect on consensus.
Study 2 (N=675) compares facilitator strategies against a no-facilitation baseline.
Study 2 comprised N=675 participants (groups of three) randomized to different LLM facilitation strategies and a no-facilitation control.
Study 1 (N=204) compares three frontier LLMs as facilitators.
Study 1 comprised N=204 participants (groups of three) randomized to facilitator conditions comparing three frontier language models.
We present two empirical studies (N=879) of real-time, text-based group deliberation in an incentive-compatible charity allocation task with real financial stakes ($7,200 USD).
Two online experiments involving real-time, text-based group deliberation. Total participants N=879 in groups of three; total monetary stakes for the charity allocation task equal $7,200 USD.
The study used a qualitative interpretivist research design drawing on semistructured interviews with 28 managers and professionals from 12 organizations across technology, finance and knowledge-intensive service sectors in Europe and Asia, using thematic and interpretive analysis supported by organizational document review.
Methodology statement from the paper (explicit description of sample, sectors, regions and analytic approach).
AI should be conceptualized as a co-evolving organizational capability rather than a deterministic technology.
Argument developed from interpretive analysis of interview data (n=28), literature engagement and organizational document review.
The study develops an emergent framework of AI–human co-adaptation comprising three interrelated dimensions: technological alignment, cognitive calibration and ethical anchoring.
Framework derived from thematic/interpretive analysis of interview data (n=28) and supporting organizational documents.
The paper introduces the concept of 'augmented work agency' as a multi-level, interpretive form of human agency in algorithmically mediated environments.
Conceptual development within the paper grounded in literature review and qualitative interview data (28 participants) and organizational document review.
This study used a three-wave lagged survey design with 381 valid matched employees from knowledge-intensive firms in China.
Methods statement in paper reporting study design and sample composition: three-wave lagged survey and 381 valid matched employee responses from knowledge-intensive Chinese firms.
The overall impact of prompt design on readability remains limited.
Reported results from prompt-dimension experiments indicating that while some prompt elements influence readability, the aggregate effect size of prompt engineering on overall readability was limited.
Current LLMs produce code with overall readability comparable to human-written code.
Comparison of readability scores (from the paper's readability model) between LLM-generated code and human-written code across 5,869 scenarios; reported summary conclusion that overall readability is comparable.
We re-recruited 530 participants from 52 countries two years after they gave their preferences in the PRISM dataset to evaluate personalised and non-personalised language models in blinded multi-turn conversations (large-scale within-subject experiment).
Study methodology reported in paper: within-subject experiment, re-recruitment of 530 participants from 52 countries, blinded multi-turn conversations comparing models.
Personalisation is a standard feature of conversational AI systems used by millions; yet, the efficacy of personalisation methods is often evaluated in academic research using simulated users rather than real people.
Authors' literature and field observation stated in introduction; contextual claim about common practice in academic evaluations (no numeric experiment reported for this claim).
The study includes Natural Language Processing (NLP) analysis of 5 million consumer contacts.
Methodological statement in the paper specifying the NLP data volume.
The study includes surveys of 800 marketers.
Methodological statement in the paper specifying the survey sample size.