Evidence (8974 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
10085 claims
Filter claims →
Productivity
8974 claims
Filtered →
Governance
8062 claims
Filter claims →
Human-AI Collaboration
7749 claims
Filter claims →
Org Design
5057 claims
Filter claims →
Innovation
4896 claims
Filter claims →
Labor Markets
4088 claims
Filter claims →
Skills & Training
3372 claims
Filter claims →
Inequality
2377 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 882 | 244 | 117 | 1097 | 2424 |
| Governance & Regulation | 1010 | 469 | 229 | 135 | 1875 |
| Organizational Efficiency | 977 | 235 | 149 | 90 | 1462 |
| Technology Adoption Rate | 781 | 299 | 143 | 128 | 1362 |
| Research Productivity | 506 | 155 | 74 | 363 | 1110 |
| Output Quality | 555 | 219 | 71 | 70 | 915 |
| Decision Quality | 395 | 200 | 95 | 54 | 751 |
| Firm Productivity | 523 | 67 | 101 | 27 | 724 |
| AI Safety & Ethics | 262 | 309 | 75 | 36 | 688 |
| Market Structure | 195 | 201 | 135 | 30 | 566 |
| Task Allocation | 248 | 77 | 96 | 38 | 464 |
| Innovation Output | 300 | 34 | 55 | 20 | 411 |
| Skill Acquisition | 207 | 75 | 65 | 21 | 368 |
| Employment Level | 138 | 67 | 119 | 24 | 350 |
| Fiscal & Macroeconomic | 156 | 80 | 53 | 33 | 329 |
| Task Completion Time | 211 | 38 | 13 | 16 | 280 |
| Firm Revenue | 183 | 52 | 29 | 5 | 270 |
| Consumer Welfare | 131 | 77 | 48 | 13 | 269 |
| Inequality Measures | 50 | 141 | 54 | 9 | 254 |
| Worker Satisfaction | 104 | 85 | 25 | 13 | 227 |
| Error Rate | 87 | 112 | 11 | 5 | 215 |
| Automation Exposure | 69 | 69 | 37 | 20 | 198 |
| Wages & Compensation | 102 | 49 | 31 | 11 | 193 |
| Team Performance | 115 | 30 | 30 | 11 | 187 |
| Regulatory Compliance | 88 | 74 | 17 | 7 | 186 |
| Training Effectiveness | 109 | 22 | 14 | 21 | 168 |
| Developer Productivity | 116 | 21 | 15 | 8 | 161 |
| Job Displacement | 12 | 92 | 26 | 1 | 131 |
| Hiring & Recruitment | 57 | 12 | 9 | 5 | 83 |
| Skill Obsolescence | 6 | 59 | 10 | 2 | 77 |
| Social Protection | 43 | 17 | 8 | 2 | 70 |
| Creative Output | 35 | 21 | 9 | 4 | 70 |
| Labor Share of Income | 18 | 23 | 17 | 1 | 59 |
| Worker Turnover | 15 | 16 | — | 4 | 35 |
| Industry | — | — | — | 1 | 1 |
Productivity
Remove filter
The tool's productivity effect decomposes into two channels: one independent of worker expertise and one that scales with worker expertise.
Analytical decomposition within the model (theoretical derivation described in the paper).
The authors develop a dynamic model in which a decision-maker chooses AI usage intensity for a worker over time, trading immediate productivity against the erosion of worker skill.
Analytical contribution: dynamic theoretical model described in the paper (model structure described; no empirical sample).
In the same graphify RAG pipeline, CAPC achieves 2.4x improvement vs cache-all on httpx.
Validation on the same RAG pipeline; reported 2.4x improvement relative to cache-all on the httpx codebase.
In a graphify knowledge-graph RAG pipeline across two codebases, CAPC achieves 9.3x improvement vs cache-all on FastAPI.
Validation on a production RAG pipeline across two codebases; reported 9.3x improvement relative to cache-all on the FastAPI codebase.
On an enterprise tool-using assistant with a 94k-token schema prefix, CAPC yields a 51.7% cost reduction at r=3.
Validation experiment on a production enterprise assistant workload with a 94k-token schema prefix; reported cost reduction at compression ratio r=3.
On LongBench-v2, CAPC achieves mean savings of 49% over cache-only, 64% over query-aware compression, and 90% over vanilla.
Aggregate experimental results reported on LongBench-v2 across the 16 configurations.
CAPC is the cheapest strategy in 16/16 configurations on LongBench-v2.
Empirical evaluation on the LongBench-v2 benchmark across 16 configurations reported in the paper.
Cache-Aware Prompt Compression (CAPC) pairs query-agnostic compression with explicit cache_control plus a tier-preserving ratio bound that prevents over-compression from pushing the cached prefix into the hot tier.
Method proposal in the paper and empirical evaluation demonstrating the mechanism; design + experimental validation claimed.
Under realistic rho, query-aware compression beats naive caching at high compression ratios (r >= 6).
Cost model predictions plus confirmatory experiments (details in paper); comparison across compression ratios and measured rho values.
The study recommends targeted reskilling and labor-augmenting innovation policies to manage heterogeneous firm-level outcomes in Italy’s digital transition.
Policy recommendation drawn from empirical findings (pooled and firm-specific fixed-effects regression results on 2005–2024 panel showing heterogeneous impacts of AI patent stock).
For UniCredit, increases in AI patent stock are associated with higher productivity growth.
Firm-specific fixed-effects regressions for UniCredit using the 2005–2024 panel, reporting a rise in productivity growth linked to AI patent stock.
In pooled models across firms, AI-related innovation (proxied by AI patent stock) has a positive and statistically significant effect on labor-market outcomes.
Fixed-effects regressions on a firm-level panel dataset covering 2005–2024; AI exposure measured by AI patent stock; pooled (multi‑firm) model estimates reported as statistically significant.
Studying AI teammate design rigorously demands infrastructure no existing tool provides (i.e., reproducible configuration of an AI teammate embedded in instrumented, real-time collaboration sustained over time).
Argument in the paper motivating TRAIL; asserts gap in existing tools rather than presenting comparative empirical evidence.
An AI teammate's design properties (personality, communication style, when it speaks) can shape a team's trust, coordination, and decisions.
Stated motivation and summary of experimental findings (single-blind persona manipulation in classroom deployment showing effects on contribution ratings, linguistic alignment, team climate, and over-reliance).
The socially-supportive agent produced lower over-reliance on the AI.
Single-blind persona manipulation in the six-session classroom deployment (~51 students); over-reliance metric (behavioral or self-report) was lower when the agent was socially supportive.
The socially-supportive agent produced a warmer team climate.
Single-blind persona manipulation in the classroom deployment (~51 students); team climate measures (survey or rating) showed warmer climate for the socially-supportive persona.
The cognitive-scaffolding agent produced closer linguistic alignment between humans and the AI.
Same single-blind persona manipulation in the six-session classroom deployment (~51 students); linguistic alignment was measured via export-driven text-similarity analysis and found to be closer for the cognitive-scaffolding persona.
A single blind persona change produced a design-consistent double dissociation: a cognitive-scaffolding agent drew stronger contribution ratings.
Single-blind manipulation of agent persona during the classroom deployment (~51 students); measured contribution ratings (presumably from participants or raters) showed higher ratings for the cognitive-scaffolding persona.
TRAIL enabled export-driven AI–human text-similarity analysis.
Platform capability described and used in the deployment to perform AI–human linguistic similarity analyses (export-ready analytics demonstrated in the study).
In that deployment, TRAIL held the AI to a stable minority of the conversation.
Empirical observation from the six-session classroom deployment (~51 students) reporting AI contribution share remained a stable minority of messages in the recorded conversations.
In a real six-session classroom deployment (about 51 students), TRAIL sustained longitudinal chaining.
Empirical deployment described in paper: six-session classroom deployment with approximately 51 students; outcome observed was the ability to chain experiments longitudinally across sessions.
The Team Research and AI Integration Lab (TRAIL) is a web platform that makes the AI teammate a configurable, reproducible design object by pairing a Big Five persona with a selective-participation message pipeline, dual memory, chained longitudinal experiments, and export-ready analytics.
System description and implementation presented in the paper (platform architecture and feature list). No numerical sample size; features illustrated by the implemented platform.
Concurrently, humans increasingly use these systems as cognitive extensions.
Statement in the paper referencing adoption/use patterns (offloading/cognitive extension) observed contemporaneously with capability increases; no usage rates, sample sizes, or specific field studies provided in the excerpt.
The length of tasks such systems can complete at 50% reliability doubled roughly every seven months.
Temporal trend reported in the paper indicating a doubling-in-length rate for tasks completed at the 50% reliability threshold between 2023 and 2026; no underlying dataset, time-series, or sample size provided in the excerpt.
Between 2023 and 2026, frontier AI systems crossed documented human expert baselines on a growing set of bounded, well-specified, evaluable cognitive tasks, including graduate-level science questions, competition mathematics, software-engineering benchmarks, and structured diagnostic reasoning.
Reported benchmark results and evaluations across multiple domains (graduate-level science questions, competition mathematics, software-engineering benchmarks, structured diagnostic reasoning) cited in the paper; no specific sample sizes or individual benchmark names/scores given in the excerpt.
Policy implication: Strengthening AI investment, promoting green innovation, maintaining resilient international trade systems, and adopting balanced trade policies will support sustainable long-run economic development under increasing geopolitical uncertainty.
Synthesis and policy recommendations derived from the ARDL empirical findings (positive AII and GI effects on SEG; trade openness benefits in China; negative tariff effects) over the 2015Q1–2025Q4 sample.
Green innovation (GI) promotes sustainable economic growth (SEG) in both China and the United States, with a relatively larger contribution in China.
Country-specific ARDL estimates from quarterly data (2015Q1–2025Q4, 44 quarters) showing positive and significant GI coefficients for both countries and a larger coefficient (or contribution) estimated for China.
AI investment (AII) has a positive and statistically significant effect on sustainable economic growth (SEG), with a stronger long-run impact in the United States than in China.
Country-specific ARDL estimation using quarterly data (2015Q1–2025Q4, 44 observations per country); reported coefficient signs and statistical significance indicating positive long-run AII→SEG effects and larger estimated long-run effect for the U.S. than China.
Larger organizations may rely on AI autonomy even under moderate uncertainty (organizational scale raises propensity to grant autonomy).
Model variants that include organizational scale, showing how firm size affects delegation incentives (theoretical analysis; no empirical sample).
Partial delegation (human-in-the-loop governance) can strictly dominate both full AI autonomy and full human control, providing a theoretical foundation for hybrid governance structures.
Analytical model comparing profits/performance across governance regimes (theoretical proof and comparative statics; no empirical sample).
There exists a demand-variance threshold above which delegating pricing and inventory decisions to agentic AI becomes optimal, even when the AI is imperfect.
Analytical model and comparative-static analysis presented in the paper (theoretical derivation; no empirical sample).
Successful integration of agent-generated contributions depends not only on advances in agent capabilities but also on the human and organizational processes that govern their use.
Interpretation and discussion based on empirical observations from the dataset (distributional patterns, project-level heterogeneity, and observed collaboration models); not a causal test.
Human-agent collaboration is dominated by a single-human oversight model, in which one developer reviews and/or modifies the agent's contributions, while multi-human collaboration patterns remain uncommon.
Analysis of collaboration patterns in 25,264 agentic PRs (e.g., counts of human reviewers or modifiers per agentic PR) showing prevalence of single-human oversight vs. multi-human involvement.
Small projects (1-5 contributors) exhibit higher participation ratios and average levels of agentic PR activity than medium-sized and large projects.
Subgroup comparison by project size (contributor count) across the dataset of 2,361 repositories, analyzing participation ratios and average agentic PR activity by size category.
We analyze 25,264 agentic PRs from 2,361 popular GitHub repositories.
Empirical dataset constructed and analyzed by the authors: 25,264 agentic pull requests collected from 2,361 popular GitHub repositories (three-month observation period).
Threshold sensitivity analysis confirms the gate decisions are robust to upward perturbations of at least 14% in three of four representative cases.
Sensitivity analysis performed on the protocol applied to representative role cases; abstract reports robustness quantified as >= 14% perturbation tolerated in 3 of 4 cases.
PHP-AIO (Protocol for Human Preservation in AI-Optimized Organizations) is a five-gate sequential decision protocol with a final composite check that quantifies these unpriced systemic risks at the role level and produces auditable automation decisions.
Methodological contribution described in the paper (protocol specification / model design); presented as a proposal rather than empirically validated in the abstract.
GNATprove proved the absence of run-time errors for the rest [of the code].
Authors' report that GNATprove was used to prove absence of run-time errors for non-selected components (tool verification output).
GNATprove established functional correctness for selected primitives.
Authors report GNATprove was used to discharge proofs that established functional correctness for selected cryptographic/primitive components (formal verification results).
GNATprove discharged 49,280 proof obligations.
Tool output reported in the paper: count of proof obligations discharged by GNATprove.
Under a verifier-driven loop, AI agents wrote and verified bare-metal security software in Ada/SPARK spanning classical and post-quantum cryptography, TLS 1.3, IKEv2, X.509, and a Matrix client.
Experimental report in paper describing a verifier-driven workflow where agents produced and verified code across the listed components (method: verifier-driven loop; qualitative description of scope).
AI coding agents produce code faster than humans can review it.
Author statement in paper (qualitative claim; no numerical sample or detailed empirical protocol provided in the excerpt).
Under complementarity, cheaper AI increases investment and capital accumulation.
Calibrated DSGE experiments indicate higher investment and capital accumulation in the complementarity case following AI price declines.
Under complementarity, cheaper AI raises wages.
Model results from the calibrated DSGE showing wage increases when AI and formal labor are complements and AI prices fall.
Under complementarity, cheaper AI amplifies aggregate output.
Calibrated DSGE model experiments showing higher output (aggregate production) when AI is complementary to formal labor and AI prices decline.
When AI complements formal workers, cheaper AI expands formal employment.
Model comparative statics and calibration: results show formal employment rises in the complementarity parameterization following an AI price decline.
Under substitution, cheaper AI increases the role of the informal sector as an employment buffer.
Calibrated DSGE model experiments indicating a reallocation of employment toward informal sector when AI substitutes formal labor and AI prices fall.
Human-centric approaches to AI diffusion within enterprise processes enable sustainable process evolution through human–AI collaboration.
Argument and illustrative case studies presented in the paper that link human-centric approaches and collaborative human-AI workflows to sustained process evolution; the summary indicates qualitative evidence rather than quantified longitudinal measures.
Organizational, technological and human factors determine successful AI adoption in process management.
Synthesis from the paper's systematic literature review and case study analysis identifying key factors across organizational, technological, and human domains; no quantified causal inference presented in the summary.
Process mining tools such as Signavio and Celonis are enhancing their capabilities with AI-driven process discovery, transformation, and optimization.
Descriptive evidence from the paper's systematic analysis and case studies referring to specific tools (Signavio, Celonis) and their AI-enabled features. No numerical adoption or performance statistics provided in the summary.