Evidence (7560 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
9875 claims
Filter claims →
Productivity
8807 claims
Filter claims →
Governance
7870 claims
Filter claims →
Human-AI Collaboration
7560 claims
Filtered →
Org Design
4892 claims
Filter claims →
Innovation
4781 claims
Filter claims →
Labor Markets
4004 claims
Filter claims →
Skills & Training
3308 claims
Filter claims →
Inequality
2332 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 870 | 233 | 116 | 1066 | 2363 |
| Governance & Regulation | 976 | 451 | 218 | 133 | 1809 |
| Organizational Efficiency | 949 | 224 | 144 | 88 | 1416 |
| Technology Adoption Rate | 764 | 287 | 141 | 122 | 1325 |
| Research Productivity | 501 | 152 | 74 | 362 | 1101 |
| Output Quality | 542 | 216 | 69 | 69 | 896 |
| Decision Quality | 387 | 198 | 94 | 54 | 740 |
| Firm Productivity | 513 | 67 | 101 | 27 | 714 |
| AI Safety & Ethics | 249 | 303 | 73 | 36 | 667 |
| Market Structure | 190 | 192 | 134 | 27 | 548 |
| Task Allocation | 243 | 77 | 91 | 36 | 452 |
| Innovation Output | 291 | 33 | 55 | 20 | 401 |
| Skill Acquisition | 206 | 72 | 65 | 21 | 364 |
| Employment Level | 133 | 63 | 115 | 22 | 335 |
| Fiscal & Macroeconomic | 153 | 79 | 52 | 32 | 323 |
| Task Completion Time | 206 | 37 | 12 | 15 | 272 |
| Firm Revenue | 179 | 52 | 29 | 5 | 266 |
| Consumer Welfare | 130 | 76 | 47 | 13 | 266 |
| Inequality Measures | 48 | 137 | 51 | 6 | 242 |
| Worker Satisfaction | 101 | 81 | 25 | 13 | 220 |
| Error Rate | 84 | 110 | 11 | 5 | 210 |
| Wages & Compensation | 98 | 47 | 30 | 10 | 185 |
| Regulatory Compliance | 88 | 73 | 17 | 7 | 185 |
| Automation Exposure | 66 | 64 | 33 | 16 | 182 |
| Team Performance | 105 | 29 | 30 | 11 | 176 |
| Training Effectiveness | 109 | 22 | 14 | 21 | 168 |
| Developer Productivity | 114 | 21 | 14 | 8 | 158 |
| Job Displacement | 12 | 90 | 24 | 1 | 127 |
| Hiring & Recruitment | 57 | 9 | 9 | 5 | 80 |
| Skill Obsolescence | 6 | 56 | 9 | 1 | 72 |
| Social Protection | 43 | 17 | 8 | 2 | 70 |
| Creative Output | 35 | 21 | 9 | 4 | 70 |
| Labor Share of Income | 18 | 21 | 17 | 1 | 57 |
| Worker Turnover | 15 | 16 | — | 4 | 35 |
| Industry | — | — | — | 1 | 1 |
Human Ai Collab
Remove filter
The divergence in collective outputs is not driven by participants abandoning AI, but by how participants use it.
Behavioral/usage data from the RCT indicating continued AI use across incentive conditions and differing usage patterns (no sample size or quantitative metrics provided in excerpt).
We evaluated EcoAssist through benchmarks of 500 websites and a controlled study with 20 developers.
Explicit methodological statement in paper: benchmark sample size = 500 websites; user study sample size = 20 developers.
Functional correctness (test-based correctness) exhibits negligible statistical association with design satisfaction.
Statistical analysis reported in experiments comparing test pass (functional correctness) and design-satisfaction labels produced by verifier; paper states negligible association.
The authors build a dynamic model of public good provision in which agents contribute by solving problems posted on a public platform and accumulated solutions form a depreciating public archive.
Methodological claim in the paper — statement that a dynamic theoretical model is constructed; this is a description of the paper's method.
We conducted a systematic review and bibliometric analysis of 627 articles.
Statement in abstract reporting a systematic review and bibliometric analysis; sample size explicitly given as 627 articles.
Injecting generic green language into prompts has no reliable effect.
Controlled prompting experiments reported in the benchmark comparing prompts with 'generic green language' to other prompt types; claim of no reliable effect on measured footprint (no numerical statistics given in abstract).
This paper has been accepted at PEARC 2026.
Statement in the paper indicating conference acceptance.
The University's GIS Center Ecological Archive (849 curated datasets) serves as a single-agent baseline deployment of EnviSmart.
Reported deployment dataset count provided in the paper: 849 curated datasets used as a single-agent baseline.
The study employed a mixed-methods approach: a quantitative survey of 150 leading Nigerian firms across finance, tech, and manufacturing, complemented by qualitative analysis of government policy and workforce interviews.
Methodological statement in the paper explicitly describing sample and methods (quantitative survey n=150; qualitative policy and interviews).
The governance calibration problem — balancing control with the autonomy that gives agentic AI its value — emerges as the STS joint optimization challenge: governance must simultaneously enable and constrain autonomous operation.
Authors' synthesis and theoretical claim based on STS analysis and identified tensions between autonomy benefits and control needs in the literature.
Agentic AI transformation barriers constitute an interdependent sociotechnical system rather than isolated obstacles.
Interpretive conclusion drawn from STS mapping and cross-barrier interaction analysis across the reviewed literature.
Governance serves as the social subsystem's primary mechanism for managing the technical subsystem.
Interpretation from STS analysis in the review: authors identify governance as the key social mechanism constraining/enabling technical subsystem behavior.
STS mapping based on root-cause analysis revealed that 12 barriers originate in the technical subsystem and 17 in the social subsystem.
Authors' STS mapping of the 29 barriers to subsystem origins (technical vs. social) as derived from their root-cause analysis of the coded literature.
Twenty-nine barriers were identified and classified into five dimensions: technological (7), organizational (7), human (6), governance and regulatory (4), and economic (5).
Results of inductive coding of the 30-source literature corpus yielding 29 distinct barriers and reported counts per dimension.
Sociotechnical Systems (STS) theory was applied as an interpretive lens to map dimensions onto social and technical subsystems and analyze cross-subsystem interactions.
Self-reported analytic approach: application of STS theory to the coded barriers to map origins and interactions across subsystems.
Barriers were identified inductively through open and axial coding.
Self-reported qualitative method: inductive thematic analysis using open and axial coding on the literature corpus.
A critical narrative literature review of 30 sources (2019–2026) was conducted.
Self-reported study method: critical narrative literature review; sample_size = 30 sources published between 2019 and 2026.
AIGC and HGC exhibit distinct creation behaviors and consumption behaviors.
Descriptive comparisons in the longitudinal dataset showing differences in production rates, content volumes, and consumption patterns between AIGC and HGC.
The paper uses a comprehensive longitudinal dataset comprising tens of millions of users from a leading Chinese video-sharing platform.
Statement in paper summarizing data source: a longitudinal dataset covering 'tens of millions of users' from a major Chinese video-sharing platform; used for descriptive and comparative analyses of creation and consumption behavior.
Increasing reasoning effort (low, medium, high) provides no consistent benefit to estimation performance.
Controlled variation of each model's reasoning effort (low/medium/high) while asking them to produce 95% credible intervals for population statistics.
These chats were committed to public repositories as part of routine development, capturing in-the-wild behavior.
Data collection method: analysis of chat transcripts that were committed to public repositories (authors state collected from repos and describe them as routine commits).
We analyze 74,998 developer messages from 11,579 chat sessions across 1,300 repositories and 899 developers using Cursor and GitHub Copilot.
Reported dataset counts in the paper (message, session, repository, developer counts) drawn from public commit histories of chats.
As advanced artificial systems become more autonomous participants in these processes, the resulting interaction space begins to resemble a new kind of ecosystem in which diverse agents exchange information, cooperate, compete, and jointly explore complex adaptive landscapes.
Conceptual argument presented in the paper drawing on theories of adaptive systems and collective intelligence; no empirical test or dataset reported.
Human and artificial agents are increasingly interacting within a shared informational environment that shapes economic activity, scientific discovery, governance, and collective decision making.
Statement in paper's introduction; based on observational/phenomenological claim and citation-less framing (conceptual assertion rather than empirical analysis). No sample or empirical method reported.
Conventional microeconomic models often treat interactions between algorithmic platforms and workers as static principal-agent problems.
Literature statement in paper (conceptual framing / literature review); no empirical sample reported.
We introduce the Agentic Task Exposure (ATE) score, a composite measure computed algorithmically from O*NET task data using calibrated adoption parameters (not a regression estimate), incorporating AI capability scores, workflow coverage factors, and logistic adoption velocity.
Methodological description in the paper; algorithmic construction from O*NET task data with specified calibrated adoption parameters and components (AI capability scores, workflow coverage, logistic adoption).
Code authoring and review are only a small part of the larger software engineering process; the resulting code must also be maintained and updated over time.
Conceptual/argumentative claim presented in the paper to motivate longitudinal analysis (not presented as an empirical estimate from the dataset).
We offer several longitudinal estimates of survival and churn rates for agent-generated versus human-authored code.
Longitudinal analysis reported in the paper comparing survival and churn for agent-generated and human-authored code over time using the dataset (paper states these estimates were produced).
We compare five popular coding agents, including OpenAI Codex, Claude Code, GitHub Copilot, Google Jules, and Devin, examining how their usage differs in various development aspects such as merge frequency, edited file types, and developer interaction signals, including comments and reviews.
Comparative analysis across agents using the constructed dataset of ~110,000 PRs (paper states these five agents were compared on metrics like merge frequency, edited file types, and interaction signals).
We construct a novel dataset of approximately 110,000 open-source pull requests, including associated commits, comments, reviews, issues, and file changes, collectively representing millions of lines of source code.
Descriptive dataset construction reported in the paper (stated sample size ~110,000 PRs including commits, comments, reviews, issues, file changes; representing millions of lines of code).
This study uses semi-structured interviews with 10 practitioners to examine perceptions of collaborating with human versus AI teammates.
Methods statement in the paper: semi-structured interviews; sample size explicitly reported as 10 practitioners.
Society 5.0 and Industry 5.0 call for human-centric technology integration, but the concept lacks an operational definition that can be measured, optimized, or evaluated at the firm level.
Motivating claim grounded in literature gap analysis presented in the paper (argument that normative frameworks lack formal, operational metrics at firm level).
We propose the Workplace Augmentation Design Index (WADI), a 36-item theory-grounded instrument for diagnosing human-centricity at the firm level.
Instrument design/proposal presented in the paper (36 items mapped to the five workplace-design dimensions); no validation sample reported in the abstract.
We conducted a PRISMA-guided systematic review of 120 papers (screened from 6,096 records) to map the evidence base for each workplace-design dimension.
Systematic literature review using PRISMA protocol; final sample = 120 papers; initial records screened = 6,096.
Existing models of human-AI complementarity treat the augmentation function phi(D) as exogenous and thus ignore that two firms with identical technology investments can achieve radically different augmentation outcomes depending on workplace organization.
Argument based on literature review of prior models (the paper contrasts its approach with existing complementarity models). No new empirical sample reported for this specific claim.
The review employed a systematic analysis of multidisciplinary studies (qualitative, quantitative, and bibliometric) focused on agentic AI technologies in financial domains, covering literature published up to mid-2024.
Stated methodology of the paper (systematic review description).
A subset of four datasets included settings in which the AI provided explanations of its decision.
Paper states that four of the datasets involved AI explanations (explicitly stated in abstract).
The study compared HCT to the AI-as-advisor approach using 10 datasets from various domains, including medical diagnostics and misinformation discernment.
Paper reports an empirical comparison across 10 datasets spanning multiple domains (explicitly stated in abstract).
The hybrid confirmation tree (HCT) elicits a human judgment and an AI judgment independently; if they agree that decision is accepted, and if they disagree a second human breaks the tie.
Description of the HCT method in the paper (procedural/design specification).
The user study had N=50 participants.
Reported user study sample size (N=50) used to evaluate AI-assisted intent expansion in ecologically valid settings.
Under the current evaluation resolution, 5W3H, CO-STAR, and RISEN achieve similarly high goal-alignment scores, suggesting that dimensional decomposition itself is an important active ingredient.
Controlled comparison between three structured frameworks (5W3H, CO-STAR, RISEN) across the evaluated outputs, with no meaningful differences reported between them.
The study evaluated 3,240 model outputs (3 languages x 6 conditions x 3 models x 3 domains x 20 tasks) using an independent judge (DeepSeek-V3).
Reported experimental design and evaluation: 3 languages, 6 conditions, 3 models, 3 domains, 20 tasks; judged by DeepSeek-V3.
The paper frames the LLM-politician relationship through principal-agent theory and bounded rationality, conceptualizing the legislator as a principal delegating advisory tasks to a boundedly rational agent under structural information asymmetry.
Explicit theoretical framing described in the introduction or theory section of the paper.
Model outputs were evaluated using a dual framework combining LLM-as-Judge semantic scoring and programmatic text similarity metrics.
Paper describes the evaluation methodology: semantic scoring via LLM-as-Judge plus programmatic text similarity measures applied to model-generated rationales vs official memoranda.
Six LLMs were evaluated: GPT-5-mini, GPT-5-chat (OpenAI), Claude Haiku 4.5 (Anthropic), and Llama 4 Maverick, Llama 3.3 70B, Llama 3.1 8B (Meta).
Paper explicitly lists the six evaluated models spanning three provider families and multiple capability tiers.
The study uses a dataset of 15 Romanian Senate law proposals paired with their official explanatory memoranda (expuneri de motive).
Explicit statement in the paper describing the dataset composition: 15 Romanian Senate law proposals each paired with its official explanatory memorandum.
Limitations: the Comscore data observe household internet activity on home (non-mobile) devices and do not capture offline or mobile device activities, so extrapolation to total at-home activities should be done with caution.
Authors' explicit limitation discussion in paper stating data do not include mobile devices or offline activities.
ChatGPT adoption leaves the total time spent on productive online activities (including any time spent using ChatGPT) unchanged.
Same IV long-difference estimates as above; authors state 'leaving time spent on productive digital tasks unchanged' and that total productive activity time does not decline significantly.
The analysis uses detailed Internet browsing microdata from over 200,000 U.S. households' home devices from 2021 to 2024.
Comscore web browsing panel described in paper; authors state dataset covers 'over 200,000 U.S. households' across 2021-2024; data provides timestamps, visit durations, URLs, demographic bins, etc.
We release the anonymized dataset and analysis with a new query intent taxonomy to inform future designs of real-world AI research assistants and to support realistic evaluation.
Paper states that the anonymized Asta Interaction Dataset, accompanying analysis, and a new query intent taxonomy are being released publicly.