Evidence (8807 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
9875 claims
Filter claims →
Productivity
8807 claims
Filtered →
Governance
7870 claims
Filter claims →
Human-AI Collaboration
7560 claims
Filter claims →
Org Design
4892 claims
Filter claims →
Innovation
4781 claims
Filter claims →
Labor Markets
4004 claims
Filter claims →
Skills & Training
3308 claims
Filter claims →
Inequality
2332 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 870 | 233 | 116 | 1066 | 2363 |
| Governance & Regulation | 976 | 451 | 218 | 133 | 1809 |
| Organizational Efficiency | 949 | 224 | 144 | 88 | 1416 |
| Technology Adoption Rate | 764 | 287 | 141 | 122 | 1325 |
| Research Productivity | 501 | 152 | 74 | 362 | 1101 |
| Output Quality | 542 | 216 | 69 | 69 | 896 |
| Decision Quality | 387 | 198 | 94 | 54 | 740 |
| Firm Productivity | 513 | 67 | 101 | 27 | 714 |
| AI Safety & Ethics | 249 | 303 | 73 | 36 | 667 |
| Market Structure | 190 | 192 | 134 | 27 | 548 |
| Task Allocation | 243 | 77 | 91 | 36 | 452 |
| Innovation Output | 291 | 33 | 55 | 20 | 401 |
| Skill Acquisition | 206 | 72 | 65 | 21 | 364 |
| Employment Level | 133 | 63 | 115 | 22 | 335 |
| Fiscal & Macroeconomic | 153 | 79 | 52 | 32 | 323 |
| Task Completion Time | 206 | 37 | 12 | 15 | 272 |
| Firm Revenue | 179 | 52 | 29 | 5 | 266 |
| Consumer Welfare | 130 | 76 | 47 | 13 | 266 |
| Inequality Measures | 48 | 137 | 51 | 6 | 242 |
| Worker Satisfaction | 101 | 81 | 25 | 13 | 220 |
| Error Rate | 84 | 110 | 11 | 5 | 210 |
| Wages & Compensation | 98 | 47 | 30 | 10 | 185 |
| Regulatory Compliance | 88 | 73 | 17 | 7 | 185 |
| Automation Exposure | 66 | 64 | 33 | 16 | 182 |
| Team Performance | 105 | 29 | 30 | 11 | 176 |
| Training Effectiveness | 109 | 22 | 14 | 21 | 168 |
| Developer Productivity | 114 | 21 | 14 | 8 | 158 |
| Job Displacement | 12 | 90 | 24 | 1 | 127 |
| Hiring & Recruitment | 57 | 9 | 9 | 5 | 80 |
| Skill Obsolescence | 6 | 56 | 9 | 1 | 72 |
| Social Protection | 43 | 17 | 8 | 2 | 70 |
| Creative Output | 35 | 21 | 9 | 4 | 70 |
| Labor Share of Income | 18 | 21 | 17 | 1 | 57 |
| Worker Turnover | 15 | 16 | — | 4 | 35 |
| Industry | — | — | — | 1 | 1 |
Productivity
Remove filter
We performed an extensive evaluation of 37 state-of-the-art Vision-Language Models on MultihopSpatial.
Empirical evaluation described in the paper listing the number of models evaluated (37).
The paper treats data as a new type of production factor and endogenizes it within the production function.
Theoretical/methodological: the paper constructs a macro-level theoretical model that explicitly includes data as an endogenous input in the production function (no empirical/sample data).
Economic evaluations of GLAI should account for end-to-end risk externalities (error propagation, institutional trust, rights impacts), not only short-term productivity gains.
Methodological recommendation grounded in conceptual synthesis of technical, behavioral, and legal risks; normative argument rather than empirical result.
Generative Legal AI (GLAI) systems are built on token-prediction (LLM) architectures rather than formal legal-reasoning architectures.
Conceptual and technical analysis in the paper distinguishing GLAI from other legal-tech; literature synthesis on common LLM architectures. No original empirical dataset or sample size—qualitative/technical review.
The paper's formalism shows that prompt/system messages shape distributions over possible execution paths (indirect control) but do not evaluate actual partial paths at runtime.
Formal mapping in the paper that treats prompts as shaping prior over paths; conceptual argument and illustrative examples.
Through a thematic review of existing research, the authors identified recurring themes about incentive schemes: their components, how researchers manipulate them, and their impact on research outcomes.
Authors' stated method and findings: thematic review (the scope/number of reviewed papers not specified in excerpt).
A critical aspect of conducting human–AI decision-making studies is the role of participants, often recruited through crowdsourcing platforms.
Claim based on the authors' thematic literature review noting participant sourcing practices (specific studies and counts not given in excerpt).
Researchers conduct empirical studies investigating how humans use AI assistance for decision-making and how this collaboration impacts results.
Statement summarizing the research landscape; supported implicitly by the authors' thematic review of existing empirical studies (number of studies not specified in excerpt).
The study provides empirical evidence specific to a small open EU economy (Slovakia) on the relationship between AI adoption and labour productivity.
Use of harmonised Eurostat enterprise and productivity data for Slovakia and EU27 over 2021–2024, analysed with descriptive statistics, gap analysis, dynamics of change, correlation, and an illustrative regression model.
Returns to AI are heterogeneous across firms; estimating treatment effects requires attention to selection, complementarities, and dynamic adoption pipelines.
Methodological argument referencing treatment-effect literature and observed firm heterogeneity; supported by conceptual examples rather than a single empirical treatment-effect estimate.
Productivity effects at the aggregate (economy-wide) level are delayed relative to firm-level gains.
Cross-study synthesis noting temporal lags between observed firm-level productivity improvements and measurable aggregate effects in the literature included in the SLR.
The review followed the PRISMA protocol and synthesized 78 peer-reviewed studies and institutional reports published between 2015 and 2025.
Systematic Literature Review using PRISMA protocol; sample of 78 peer-reviewed studies and institutional reports (2015–2025) as described in the paper.
Human-only and AI-assisted teams performed similarly on most outcomes.
Comparison across outcome measures from the randomized experiment; summary statement indicates parity on most measured tasks except for detection of major coding errors.
We randomly assigned 288 researchers to 103 teams working under three conditions (human-only, AI-assisted, AI-led).
Experimental design reported in paper: randomized assignment of 288 researchers into 103 teams across three experimental conditions.
Green computing capacity is measured by a composite index covering computing infrastructure, green energy support, low-carbon operating efficiency, and computing–network coordination.
Method description in abstract listing the components of the composite green computing capacity index.
AI technological development is measured by city-level AI patent grants.
Method description in abstract stating AI is proxied by city-level AI patent grants.
Urban green productivity is measured by an undesirable output Super-SBM model.
Method description in abstract specifying the Super-SBM undesirable-output Data Envelopment Analysis model used to compute green productivity.
The study uses panel data for 287 Chinese prefecture-level and above cities from 2005 to 2023.
Statement in abstract describing data scope and timeframe; sample count explicitly given as 287 cities and years 2005–2023.
The study evaluates green productivity (GP) across three dimensions: labourers, means of labour, and objects of labour.
Paper states it measures GP along three dimensions (labourers, means of labour, objects of labour); methodological description in paper; measurement/construct definition rather than empirical test. No sample size reported in the summary.
An undirected configuration change in a prior iteration produced zero impact, illustrating the cost of iterating without diagnosis.
Reported observation from an earlier iteration in the case study where a non-diagnostic change had no measurable effect.
The paper combines findings from information systems research, organizational behavior studies, and artificial intelligence literature through its analysis of recent empirical and theoretical studies conducted between 2021 and 2026.
Methodological description provided in the paper (literature synthesis covering 2021–2026).
Using aggregate data, the study provides no evidence that AI benefits any particular group of workers — neither highly educated nor less-educated ones.
Authors' interaction analysis between AI adoption and human capital using aggregate panel data; reported null finding for differential benefits across education/skill groups (1995–2017, 35 OECD countries).
Results are robust to state-by-year and industry-by-year fixed effects.
Robustness checks reported in paper that include state-by-year and industry-by-year fixed effects with results stated to hold.
Where AI can perform tasks independently, we find no significant employment effect.
Heterogeneous DiD estimates showing null (statistically non-significant) employment coefficients for occupations/industries where AI can perform tasks independently.
We examine aggregate effects using administrative data covering essentially all U.S. employers in a difference-in-differences design exploiting occupational AI exposure across industries and states.
Statement in paper describing data and empirical strategy: administrative data covering essentially all U.S. employers; difference-in-differences design exploiting occupational AI exposure variation across industries and states.
We observe no differences in productivity across adoption levels.
Authors' empirical comparison reporting null differences in measured firm productivity across adoption categories (based on their matched data).
We observe no differences in capex across adoption levels.
Authors' empirical comparison reporting null differences in capital expenditures across adoption categories (based on firm financial data matched to adoption measure).
Under matched time limits, originality with a GPT-4 partner is statistically equivalent to that with a human partner.
Result from the in-person pilot (N = 62) comparing originality scores between participants partnered with GPT-4 versus human partners under matched time limits; reported as statistical equivalence in the paper.
We test this question across four lifecycle stages: data preparation, data extraction, statistical analysis, and reporting, using one generated skill per stage.
Experimental setup description specifying four task-lifecycle stages and use of one generated skill for each stage.
A supplemental token-matched control adds 1,512 runs and finds that Full skills perform similarly to task-irrelevant skill-formatted content.
Additional control experiment (token-matched content) reported in supplement, with run count and comparison results described.
The total spread across variants is only 1.2 percentage points.
Reported range/variation in performance metrics across all skill variants in the ablation experiment.
Compared with prompting using the task alone, neither the full generated skill nor any ablated skill variant significantly improves performance; all p-values are at least 0.396.
Statistical hypothesis tests comparing Full and ablated-skill variants to task-only prompting; reported minimum p-value threshold.
The main ablation covers 56 tasks, nine model configurations, and three providers, yielding 7,560 runs.
Description of experimental design and aggregate counts reported in the paper (tasks × model configurations × providers → total run count).
We find no reliable improvement from full generated skills over No-Skill prompting.
Empirical comparison between Full generated-skill prompting and No-Skill (task-only) prompting across the study's evaluation tasks and model configurations; statistical testing reported.
Average inference power is a model-intrinsic constant, invariant to input resolution, image complexity, and prompt type, with less than 5% variation across all conditions.
Systematic on-device energy profiling across five models spanning three architecture families, four input resolutions, and two hardware platforms (NVIDIA RTX 3070 and Jetson Orin NX); direct measurement of power draw across different inputs and prompts.
Organized search in SSRN, Google Scholar, Web of Science, Scopus, and government repositories resulted in 37 sources that satisfy pre-determined inclusion criteria and are rated in three levels of evidence.
Description of the review's search and screening process reported in the paper (databases searched listed, inclusion criteria applied, resulting count = 37, evidence rated into three levels).
The study used a sequential mixed-methods design (Qual → Quan) consisting of eight expert interviews analyzed with grounded-theory coding and a follow-up survey of 499 AI-aware consumers.
Methods reported in the paper: eight expert interviews (qualitative) and a survey with 499 respondents (quantitative).
From that coded sample the authors built a causal model of 26 constructs and 67 relationships (64 directed, 3 contested).
Reported model construction from the coded sample as stated in the abstract.
The authors filtered that corpus and coded a stratified random sample of 3,100 documents with an LLM-assisted pipeline.
Reported sampling and coding procedure stated in the abstract.
We collected 38,709 grey-literature documents (engineering blogs and Reddit threads) and filtered to those substantively about code review.
Reported data-collection procedure and corpus size stated in the abstract.
We isolate the orchestration layer with a controlled swap: 22 locked evaluation tasks, six foundation models, changing only the orchestration layer (a frozen conventional production loop versus the Writer Agent Harness).
Methodological statement describing experimental design reported in the paper: controlled swap with 22 tasks and six models.
Task-completion quality is at parity between harness and baseline (0.78->0.81), directional at this sample size.
Reported average/directional quality scores from the controlled swap across 22 tasks and six models, showing scores 0.78 (baseline) and 0.81 (harness).
The observed patterns in BEA–BLS data for 63 U.S. industries over 1997–2023 do not reflect cyclical variation but register a structural change in the system of factors of production.
Trend/structural analysis of BEA–BLS data for 63 U.S. industries (1997–2023) reported in the paper.
The digital sector comprises three times fewer industries than the physical sector.
Empirical statement based on BEA–BLS data covering 63 U.S. industries (1997–2023) as reported in the paper.
The periodization of US macroeconomic productivity cycles was refined by identifying the new stages 'pandemic and adaptation phase' and 'artificial intelligence phase'.
Calculation of AAPC indices for 1947–2025 and retrospective comparative analysis leading to refinement of periodization and naming of new stages.
Eight distinct macroeconomic cycles of productivity change in the United States from 1947 to 2025 are identified.
Secondary data analysis of aggregated US Bureau of Labor Statistics series for 1947–2025; long-term average annual rates of productivity change (AAPC) computed using the index method and geometric mean growth rate; comparative analysis to identify cycle breaks.
Survey data were collected from firms located in major Chinese cities (Beijing, Shenzhen, Xi’an, and Zhengzhou), resulting in 750 valid responses for analysis.
Reported survey sampling and data collection in the paper; explicit statement of cities sampled and number of valid responses (750).
Workers were assigned to no overrides, free overrides, or a two-per-machine limit on downward overrides.
Experimental design statement in paper: randomized assignment into three arms (no overrides, free overrides, constrained two-per-machine downward override limit).
We tested [the policy] through a randomized field experiment with 553 workers at a major Chinese smart vending machine retailer that manages more than 59,000 machines and 4,000 SKUs.
Randomized field experiment described in paper; sample stated as 553 workers and operational context (retailer with >59,000 machines and >4,000 SKUs).
The runs spanned several model generations, two agent harnesses, two reasoning effort levels, a testing tool, and two design oriented prompts.
Description of experimental conditions reported in the study (factors varied across the 90 runs).