Evidence (8974 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
10085 claims
Filter claims →
Productivity
8974 claims
Filtered →
Governance
8062 claims
Filter claims →
Human-AI Collaboration
7749 claims
Filter claims →
Org Design
5057 claims
Filter claims →
Innovation
4896 claims
Filter claims →
Labor Markets
4088 claims
Filter claims →
Skills & Training
3372 claims
Filter claims →
Inequality
2377 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 882 | 244 | 117 | 1097 | 2424 |
| Governance & Regulation | 1010 | 469 | 229 | 135 | 1875 |
| Organizational Efficiency | 977 | 235 | 149 | 90 | 1462 |
| Technology Adoption Rate | 781 | 299 | 143 | 128 | 1362 |
| Research Productivity | 506 | 155 | 74 | 363 | 1110 |
| Output Quality | 555 | 219 | 71 | 70 | 915 |
| Decision Quality | 395 | 200 | 95 | 54 | 751 |
| Firm Productivity | 523 | 67 | 101 | 27 | 724 |
| AI Safety & Ethics | 262 | 309 | 75 | 36 | 688 |
| Market Structure | 195 | 201 | 135 | 30 | 566 |
| Task Allocation | 248 | 77 | 96 | 38 | 464 |
| Innovation Output | 300 | 34 | 55 | 20 | 411 |
| Skill Acquisition | 207 | 75 | 65 | 21 | 368 |
| Employment Level | 138 | 67 | 119 | 24 | 350 |
| Fiscal & Macroeconomic | 156 | 80 | 53 | 33 | 329 |
| Task Completion Time | 211 | 38 | 13 | 16 | 280 |
| Firm Revenue | 183 | 52 | 29 | 5 | 270 |
| Consumer Welfare | 131 | 77 | 48 | 13 | 269 |
| Inequality Measures | 50 | 141 | 54 | 9 | 254 |
| Worker Satisfaction | 104 | 85 | 25 | 13 | 227 |
| Error Rate | 87 | 112 | 11 | 5 | 215 |
| Automation Exposure | 69 | 69 | 37 | 20 | 198 |
| Wages & Compensation | 102 | 49 | 31 | 11 | 193 |
| Team Performance | 115 | 30 | 30 | 11 | 187 |
| Regulatory Compliance | 88 | 74 | 17 | 7 | 186 |
| Training Effectiveness | 109 | 22 | 14 | 21 | 168 |
| Developer Productivity | 116 | 21 | 15 | 8 | 161 |
| Job Displacement | 12 | 92 | 26 | 1 | 131 |
| Hiring & Recruitment | 57 | 12 | 9 | 5 | 83 |
| Skill Obsolescence | 6 | 59 | 10 | 2 | 77 |
| Social Protection | 43 | 17 | 8 | 2 | 70 |
| Creative Output | 35 | 21 | 9 | 4 | 70 |
| Labor Share of Income | 18 | 23 | 17 | 1 | 59 |
| Worker Turnover | 15 | 16 | — | 4 | 35 |
| Industry | — | — | — | 1 | 1 |
Productivity
Remove filter
The authors build a dynamic model of public good provision in which agents contribute by solving problems posted on a public platform and accumulated solutions form a depreciating public archive.
Methodological claim in the paper — statement that a dynamic theoretical model is constructed; this is a description of the paper's method.
Injecting generic green language into prompts has no reliable effect.
Controlled prompting experiments reported in the benchmark comparing prompts with 'generic green language' to other prompt types; claim of no reliable effect on measured footprint (no numerical statistics given in abstract).
This paper has been accepted at PEARC 2026.
Statement in the paper indicating conference acceptance.
The University's GIS Center Ecological Archive (849 curated datasets) serves as a single-agent baseline deployment of EnviSmart.
Reported deployment dataset count provided in the paper: 849 curated datasets used as a single-agent baseline.
The experiment used stratified randomization across 32 strata with 255 treatment firms and 260 control firms; baseline characteristics are well balanced across groups.
Experimental design description: stratification by geography, traction score, and baseline AI use; reporting of allocation counts and balance tests in Table 2.
Attrition from the accelerator was low (1.6%, eight ventures) and balanced across treatment and control.
Program enrollment and retention records for the 515 firms in the randomized accelerator; 8 firms attrited.
The gains from treatment are broad-based: there are no significant differential effects by baseline firm performance or founder technical background.
Heterogeneity/subgroup analyses in the randomized sample (515 firms) comparing treatment effects across strata defined by baseline traction and founder technical background.
Treated firms' demand for labor remains unchanged.
RCT with 515 firms; firms reported labor demand/changes, comparison between treatment and control groups showed no significant change.
The authors measure the set of tasks that are automated in a given year via queries to ChatGPT's Deep Research (their 'heroic measurement').
Methodological statement in the introduction describing the measurement approach for identifying automated tasks.
When the automation process is continuous, firms switch from labor to capital at exactly the point where costs are equal; the switching process itself generates no productivity growth.
Theoretical result (Proposition in the model) derived in the task-based framework under competitive equilibrium and the no-de-automation assumption.
Despite substantial expected AI progress, most respondents do not forecast major departures from recent macroeconomic baselines, citing factors like historical base rates, adoption lags, demographic headwinds, policy responses, and infrastructure bottlenecks.
Qualitative summary of respondents' reasoning accompanying their unconditional forecasts (Key Findings and 1.2 description of survey elicitation).
These chats were committed to public repositories as part of routine development, capturing in-the-wild behavior.
Data collection method: analysis of chat transcripts that were committed to public repositories (authors state collected from repos and describe them as routine commits).
We analyze 74,998 developer messages from 11,579 chat sessions across 1,300 repositories and 899 developers using Cursor and GitHub Copilot.
Reported dataset counts in the paper (message, session, repository, developer counts) drawn from public commit histories of chats.
We find little evidence of crashing waves (in contrast to recent work by METR).
Analysis of the >3,000 tasks and >17,000 evaluations which reportedly do not show abrupt, concentrated surges in AI capability on small sets of tasks.
The evaluation is based on more than 17,000 evaluations by workers from these jobs.
Reported sample of >17,000 human evaluations of model outputs.
We test for these effects in preliminary evidence from an ongoing evaluation of AI capabilities across over 3,000 broad-based tasks derived from the U.S. Department of Labor O*NET categorization that are text-based and thus LLM-addressable.
Empirical study design reporting an ongoing evaluation covering >3,000 text-based tasks mapped from O*NET.
This paper employs a staggered difference-in-differences (DID) model using data from Chinese A-share listed manufacturing companies from 2012 to 2023 and uses the National Artificial Intelligence Innovative Application Pioneer Zone (AIIAPZ) policy as a quasi-natural experiment.
Staggered DID empirical design; sample described as Chinese A-share listed manufacturing firms, 2012–2023; AIIAPZ policy used as treatment assignment (quasi-natural experiment).
The study uses panel data from listed manufacturing firms in China and employs a quasi-natural experiment approach.
Statement in the abstract describing data source (panel of listed manufacturing firms in China) and empirical strategy (quasi-natural experiment).
Code authoring and review are only a small part of the larger software engineering process; the resulting code must also be maintained and updated over time.
Conceptual/argumentative claim presented in the paper to motivate longitudinal analysis (not presented as an empirical estimate from the dataset).
We offer several longitudinal estimates of survival and churn rates for agent-generated versus human-authored code.
Longitudinal analysis reported in the paper comparing survival and churn for agent-generated and human-authored code over time using the dataset (paper states these estimates were produced).
We compare five popular coding agents, including OpenAI Codex, Claude Code, GitHub Copilot, Google Jules, and Devin, examining how their usage differs in various development aspects such as merge frequency, edited file types, and developer interaction signals, including comments and reviews.
Comparative analysis across agents using the constructed dataset of ~110,000 PRs (paper states these five agents were compared on metrics like merge frequency, edited file types, and interaction signals).
We construct a novel dataset of approximately 110,000 open-source pull requests, including associated commits, comments, reviews, issues, and file changes, collectively representing millions of lines of source code.
Descriptive dataset construction reported in the paper (stated sample size ~110,000 PRs including commits, comments, reviews, issues, file changes; representing millions of lines of code).
The paper extends classical (Solow) and endogenous (Romer) growth models to incorporate TAI, producing a dynamic framework for analyzing AI-driven structural change.
Methodological claim: the authors explicitly state they build on Solow (1956) and Romer (1990) to develop an integrated dynamic model that incorporates TAI; evidence is described model extension and formalization within the paper.
The study is based on a qualitative analysis of recent academic literature, comparative analysis of sector-specific applications of Big Data technologies, and synthesis of empirical findings from international studies using a systemic and structural analysis approach.
Methodological statement within the paper describing data sources and analytic approach; not an empirical claim about outcomes.
The research documents a transition in the literature (2013–2025) from early 'risk-of-automation' evaluations toward task-based and firm-level econometric models.
Literature review/synthesis across the 2013–2025 body of research as described in the paper.
Society 5.0 and Industry 5.0 call for human-centric technology integration, but the concept lacks an operational definition that can be measured, optimized, or evaluated at the firm level.
Motivating claim grounded in literature gap analysis presented in the paper (argument that normative frameworks lack formal, operational metrics at firm level).
We propose the Workplace Augmentation Design Index (WADI), a 36-item theory-grounded instrument for diagnosing human-centricity at the firm level.
Instrument design/proposal presented in the paper (36 items mapped to the five workplace-design dimensions); no validation sample reported in the abstract.
We conducted a PRISMA-guided systematic review of 120 papers (screened from 6,096 records) to map the evidence base for each workplace-design dimension.
Systematic literature review using PRISMA protocol; final sample = 120 papers; initial records screened = 6,096.
Existing models of human-AI complementarity treat the augmentation function phi(D) as exogenous and thus ignore that two firms with identical technology investments can achieve radically different augmentation outcomes depending on workplace organization.
Argument based on literature review of prior models (the paper contrasts its approach with existing complementarity models). No new empirical sample reported for this specific claim.
The widening effect of AI adoption on the electricity output growth gap diminishes over time and becomes statistically insignificant after approximately three years.
Temporal (dynamic) empirical analysis / event-study-style estimation tracing the AI adoption effect over multiple years post-adoption; statistical significance reported to fade by year ~3. Sample size / exact time windows not provided in the summary.
The review employed a systematic analysis of multidisciplinary studies (qualitative, quantitative, and bibliometric) focused on agentic AI technologies in financial domains, covering literature published up to mid-2024.
Stated methodology of the paper (systematic review description).
The study adopted a positivist philosophy and a descriptive-correlational design.
Methods section statement in the paper describing the research philosophy and study design.
Data were collected from innovation-focused executives across 39 licensed Kenyan commercial banks.
Paper statement specifying sample source: 'Using data from innovation-focused executives across 39 licensed banks.'
Technological innovation was assessed via adoption of new systems, integration of digital channels, and use of Artificial Intelligence and data analytics.
Measurement description provided in the paper listing the components used to operationalize technological innovation.
Competitiveness in the study was measured through market share, return on equity and customer satisfaction.
Measurement description provided in the paper describing dependent variable operationalization (explicit list of three indicators).
The user study had N=50 participants.
Reported user study sample size (N=50) used to evaluate AI-assisted intent expansion in ecologically valid settings.
Under the current evaluation resolution, 5W3H, CO-STAR, and RISEN achieve similarly high goal-alignment scores, suggesting that dimensional decomposition itself is an important active ingredient.
Controlled comparison between three structured frameworks (5W3H, CO-STAR, RISEN) across the evaluated outputs, with no meaningful differences reported between them.
The study evaluated 3,240 model outputs (3 languages x 6 conditions x 3 models x 3 domains x 20 tasks) using an independent judge (DeepSeek-V3).
Reported experimental design and evaluation: 3 languages, 6 conditions, 3 models, 3 domains, 20 tasks; judged by DeepSeek-V3.
We implement a rigorously controlled execution-based testbed featuring Git worktree isolation and explicit global memory to evaluate agent coordination frameworks.
Methodological description in the paper indicating the testbed design choices (Git worktree isolation, explicit global memory) used to ensure controlled, reproducible execution of agent-generated code.
We benchmark a single-agent baseline against two multi-agent paradigms: a subagent architecture (parallel exploration with post-hoc consolidation) and an agent team architecture (experts with pre-execution handoffs) using a rigorously controlled, execution-based testbed.
Description of experimental setup in the paper: an execution-based testbed with Git worktree isolation and explicit global memory; experiments explicitly compare single-agent, subagent, and agent-team architectures under fixed computational time budgets.
Limitations: the Comscore data observe household internet activity on home (non-mobile) devices and do not capture offline or mobile device activities, so extrapolation to total at-home activities should be done with caution.
Authors' explicit limitation discussion in paper stating data do not include mobile devices or offline activities.
ChatGPT adoption leaves the total time spent on productive online activities (including any time spent using ChatGPT) unchanged.
Same IV long-difference estimates as above; authors state 'leaving time spent on productive digital tasks unchanged' and that total productive activity time does not decline significantly.
The analysis uses detailed Internet browsing microdata from over 200,000 U.S. households' home devices from 2021 to 2024.
Comscore web browsing panel described in paper; authors state dataset covers 'over 200,000 U.S. households' across 2021-2024; data provides timestamps, visit durations, URLs, demographic bins, etc.
We release the anonymized dataset and analysis with a new query intent taxonomy to inform future designs of real-world AI research assistants and to support realistic evaluation.
Paper states that the anonymized Asta Interaction Dataset, accompanying analysis, and a new query intent taxonomy are being released publicly.
The Asta Interaction Dataset comprises over 200,000 user queries and interaction logs from two deployed tools (a literature discovery interface and a scientific question-answering interface) within an LLM-powered retrieval-augmented generation platform.
Statement in paper describing dataset composition: >200,000 user queries and interaction logs collected from two deployed tools (literature discovery and scientific Q&A) within an RAG platform. Dataset release described in methods/dataset section.
Methods combine targeted literature synthesis, comparative conceptual analysis, and framework building (with recent scholarly and institutional sources reviewed).
Explicit methodological statement in the paper describing the review and analytic approach; no primary-data methods used.
AI coding assistants are a high-visibility class of corporate AI and are given special attention as an illustrative case in the paper.
Paper specifically calls out AI coding assistants as a focal example in the conceptual analysis and discussion; based on literature review rather than original measurement.
The Article translates these insights into risk-sensitive guideposts for modernizing governance of AI-enabled tools and emerging modalities, from agentic systems to blockchain-deployed smart contracts.
Prescriptive/conceptual policy guidance presented in the Article (normative recommendations; governance framework).
The Innovation Frontier traces LegalTech’s evolution from 2000s-vintage e-discovery to generative AI.
Historical/chronological analysis in the Article (literature review/history of LegalTech provided by authors).
The Legal Services Value Chain disaggregates the lifecycle of a legal matter into five distinct nodes of activity.
Model description in the Article (conceptual architecture; decomposition of legal work).