Evidence (2424 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
10085 claims
Filter claims →
Productivity
8974 claims
Filter claims →
Governance
8062 claims
Filter claims →
Human-AI Collaboration
7749 claims
Filter claims →
Org Design
5057 claims
Filter claims →
Innovation
4896 claims
Filter claims →
Labor Markets
4088 claims
Filter claims →
Skills & Training
3372 claims
Filter claims →
Inequality
2377 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 882 | 244 | 117 | 1097 | 2424 |
| Governance & Regulation | 1010 | 469 | 229 | 135 | 1875 |
| Organizational Efficiency | 977 | 235 | 149 | 90 | 1462 |
| Technology Adoption Rate | 781 | 299 | 143 | 128 | 1362 |
| Research Productivity | 506 | 155 | 74 | 363 | 1110 |
| Output Quality | 555 | 219 | 71 | 70 | 915 |
| Decision Quality | 395 | 200 | 95 | 54 | 751 |
| Firm Productivity | 523 | 67 | 101 | 27 | 724 |
| AI Safety & Ethics | 262 | 309 | 75 | 36 | 688 |
| Market Structure | 195 | 201 | 135 | 30 | 566 |
| Task Allocation | 248 | 77 | 96 | 38 | 464 |
| Innovation Output | 300 | 34 | 55 | 20 | 411 |
| Skill Acquisition | 207 | 75 | 65 | 21 | 368 |
| Employment Level | 138 | 67 | 119 | 24 | 350 |
| Fiscal & Macroeconomic | 156 | 80 | 53 | 33 | 329 |
| Task Completion Time | 211 | 38 | 13 | 16 | 280 |
| Firm Revenue | 183 | 52 | 29 | 5 | 270 |
| Consumer Welfare | 131 | 77 | 48 | 13 | 269 |
| Inequality Measures | 50 | 141 | 54 | 9 | 254 |
| Worker Satisfaction | 104 | 85 | 25 | 13 | 227 |
| Error Rate | 87 | 112 | 11 | 5 | 215 |
| Automation Exposure | 69 | 69 | 37 | 20 | 198 |
| Wages & Compensation | 102 | 49 | 31 | 11 | 193 |
| Team Performance | 115 | 30 | 30 | 11 | 187 |
| Regulatory Compliance | 88 | 74 | 17 | 7 | 186 |
| Training Effectiveness | 109 | 22 | 14 | 21 | 168 |
| Developer Productivity | 116 | 21 | 15 | 8 | 161 |
| Job Displacement | 12 | 92 | 26 | 1 | 131 |
| Hiring & Recruitment | 57 | 12 | 9 | 5 | 83 |
| Skill Obsolescence | 6 | 59 | 10 | 2 | 77 |
| Social Protection | 43 | 17 | 8 | 2 | 70 |
| Creative Output | 35 | 21 | 9 | 4 | 70 |
| Labor Share of Income | 18 | 23 | 17 | 1 | 59 |
| Worker Turnover | 15 | 16 | — | 4 | 35 |
| Industry | — | — | — | 1 | 1 |
The labor-market outcome of a decline in AI prices depends on the elasticity of substitution between imported AI capital and formal labor.
Analytical result from the DSGE model and its comparative-static analysis; central mechanism in model design.
AI's effects are often uneven and highly context-dependent.
Summary statement in the abstract based on the systematic review of 194 articles noting heterogeneity in AI impacts across contexts and dimensions.
Prior work shows that agents prompted with low agreeableness produce adversarial language, while those prompted with high agreeableness become cooperative.
Citation to prior literature (not specified in the abstract) reporting correlations/causal effects of agreeableness prompts on generated language (adversarial vs cooperative). No sample size or study details provided in the abstract.
Overall conclusion: AI plays a dual role — fostering productivity and inclusion (through employment and some gender balance gains) while posing risks of increased within-firm inequality.
Synthesis of empirical findings from fixed-effects regressions, mediation and moderation analyses on the firm panel showing employment and wage gains alongside increased pay dispersion.
AI’s impact on university-educated labour cannot be understood through technological capability alone; it requires analysing the rentier dynamics of contemporary capitalism.
Theoretical argument and conceptual framework drawing on political economy and sociology (no empirical sample reported).
The U-shaped relationship between AIIA and APCRS remains significantly U-shaped across grain strategic zones.
Subsample/region-specific tests reported in the paper showing the U-shaped relationship persists in grain strategic zones using the provincial panel.
The effect of AIIA on APCRS is more pronounced in regions with higher levels of marketization and industrialization.
Regional heterogeneity analysis in the paper comparing subsamples or interacting AIIA with measures of marketization and industrialization across the 30 provinces (2016–2024).
Agricultural labor productivity strengthens the curvature of the estimated nonlinear (U-shaped) relationship between AIIA and APCRS.
Heterogeneity/moderation tests reported in the paper indicating that higher agricultural labor productivity makes the U-shaped pattern more pronounced, based on the 30-province panel.
Artificial intelligence industry agglomeration (AIIA) has a U-shaped relationship with agricultural pollution–carbon reduction synergy (APCRS) in the full sample.
Full-sample empirical analysis using panel regressions on data for 30 provinces (2016–2024) showing a nonlinear (U-shaped) estimated relationship between AIIA and APCRS.
This article adopts a contextual approach to technology, considering it in conjunction with the social context in which it is situated.
Methodological statement made by the author about the approach taken in the paper (contextual rather than purely technical); not an empirical claim.
Longevity produces a short-run welfare loss that recedes as capital deepening raises wages, since households initially compress consumption and fertility to finance a longer retirement.
Model-derived welfare time path following a longevity shock showing initial welfare decline and subsequent recovery as aggregate capital deepens and wages rise; mechanism traced to household saving and fertility responses in simulations.
The two shocks move fertility in opposite directions: the AI shock raises fertility modestly through an income effect, while the longevity shock lowers fertility by strengthening life-cycle saving motives and increasing the cost of childrearing.
Endogenous-fertility overlapping-generations model with counterfactual simulations for AI and longevity shocks; comparative statics and simulation results regarding fertility responses and their mechanisms.
The empirical tests reported in the study use a sample of agricultural enterprises.
Paper text explicitly frames findings and implications for agricultural enterprises and states empirical tests were conducted on agri-business firms.
Defining query difficulty is one of the hardest problems in deployment engineering.
Statement/assertion in the paper (introductory claim); no specific empirical measurement in the abstract.
Current models achieve penetration success rates ranging from 10.7% to 69.3%.
Empirical results reported from evaluation of the 19 LLMs across the designed target servers (success-rate measurements).
However, evidence is uneven: many studies are simulation-based.
Review observation from the synthesis of the 35 included studies noting study designs (simulation prevalence noted but not numerically specified).
Coding agents already know how to navigate files, edit code, run commands, and repair outputs, but lack the simulator's executable contract (vocabulary, structural constraints, validation rules, termination conditions).
Framing/assumption presented in the paper motivating the approach (not an empirical claim).
Data contamination (training-data overlap) complicates interpretation of the models' performance.
Author notes the possibility that models' training data may have contained the target papers or related material, making results ambiguous.
The paper provides a consolidated, theory‑driven synthesis of the mechanisms through which AI‑mediated platforms simultaneously create opportunities and reproduce disadvantage for women.
Originality/value statement in the paper describing its contribution as a consolidated, theory‑driven synthesis and actionable insights for researchers, policymakers, and platform designers.
The paper analyzes the direct impact of artificial intelligence on employment structure, occupational tasks, and skill demand, as well as its indirect effects on job mobility, cross-border and industry differences, and policy interventions.
Descriptive claim of scope drawn from the systematic literature review conducted by the authors; no single empirical sample reported.
Most published twins are either coarse persona bots conditioned on a few demographic questions or detailed individual-level twins built on purpose-collected surveys and interview transcripts.
Author's literature summary / positioning statement in paper (qualitative assessment of existing published twins).
The paper's contribution includes an estimand distinction, an inspectable ABM/RL mechanism, and a reproducible artifact demonstrating that transparent behavioral assumptions are sufficient to generate gaming-like boundary dynamics without implying that computable regulation is inherently undesirable.
Author-stated contributions in the abstract describing methodological and reproducibility outputs (estimand distinction, inspectable model, reproducible artifact).
The study uses annual time-series data from 2003–2024 and the Autoregressive Distributed Lag (ARDL) modelling approach to estimate short- and long-run coefficients.
Explicit statement in the paper: annual time-series data 2003–2024 and ARDL modelling to simultaneously estimate short- and long-run coefficients.
We illustrate this transition through examples in consumer markets, education, news, and coding.
Authors state they use sectoral examples to illustrate the framework; this is a claim about the paper's contents rather than an empirical finding.
The UPCT framework offers a unified explanation for varied phenomena: pandemic resilience patterns, divergent digital transformation outcomes, and emerging risks of AI-driven organizational rigidity.
Synthesis claim by the author asserting explanatory scope of the theoretical framework; no empirical cross-case synthesis or formal validation included.
The paper's Universal Phase Crystallization Theory (UPCT) reconceptualizes organizations as recursive generative cycles (Φ→R→S→Φ′) and asserts organizational existence is better described as E = ΦR rather than E = S.
Theoretical/model claim introduced and developed in the paper; purely conceptual without empirical testing.
Modern retrieval agents expose many configuration choices -- LLM, retriever, number of documents, number of hops, and synthesis strategy -- each shaping both answer quality and serving cost.
Paper's conceptual description of retrieval pipelines and configuration dimensions (LLM, retriever, number of documents, number of hops, synthesis strategy). No empirical sample size reported for this descriptive claim.
The paper contributes by providing a structured synthesis that bridges efficiency-driven and labor-oriented perspectives on AI-driven manufacturing.
Authors' stated contribution in the paper: a structured thematic synthesis integrating two perspectives from the reviewed literature.
This study analyzes three key dimensions: labor displacement as a structural risk, the limitations of job transformation, and the emergence of human-centered AI.
Explicit methodological statement in the paper: systematic literature review and thematic synthesis focusing on three named dimensions.
We propose the Shannon Scaling Law, a unified theoretical framework that models LLM training as information transmission over a noisy channel, grounded in the Shannon-Hartley theorem, mapping model parameters to channel bandwidth and training tokens to signal power.
Theoretical formulation presented in the paper, grounded on Shannon-Hartley theorem and a mapping between model/data quantities and communication-theoretic quantities (bandwidth, signal power).
The authors construct a dynamic, posting-level measure of generative AI exposure using a two-stage large language model pipeline that identifies tasks in each posting and classifies the extent to which generative AI can perform or assist them.
Paper methodology description: two-stage LLM pipeline to identify tasks and classify generative AI perform/assist capacity at the posting level.
The study uses a nationwide dataset of job postings in the United States covering all sectors of the economy.
Paper statement: 'Using a nationwide dataset of job postings in the United States, covering all sectors of the economy.' (dataset description)
Much of the earlier provider spread came from end-to-end system behavior rather than planning alone.
Inference from the contrast between the cross-provider championship (end-to-end) where provider differences were observed and the planner bakeoff (standardized execution) where planners were near-equal.
AI platform conversation-log exposure scores partly measure the platform user base rather than the underlying workforce.
Comparative empirical analysis using AI platform conversation logs to construct occupation exposure scores; authors compare exposure measures across platforms and show variation attributable to platform user composition rather than labor-force composition.
AI functions both as a general-purpose technology and as an innovation in the method of innovation.
Conceptual/theoretical framing presented in the paper (the authors characterize AI as both a GPT and an innovation in methods of innovation).
The same observable behavioral signal can carry opposite meaning for different agent configurations.
Synthesis of the cross-configuration empirical findings (directional disagreements such as the error-rate example and other features).
Swapping the framework while the LLM is held fixed produces large behavioral differences in every action feature.
Comparative analysis across configurations holding LLM fixed; reported observation across action features.
Current LLM agents are proficient at calling isolated APIs but struggle with the "last mile" of commercial software automation.
Authors' comparative characterization based on literature context and their benchmark motivation; stated in introduction rather than a quantified experiment in the excerpt.
In operational meteorology, adjoint-based methods derive value from the forecast model itself but require full data assimilation infrastructure.
Technical background in paper describing adjoint-based methods and their infrastructural requirements (methodological literature references; no new empirical data).
Semiconductors are a representative case study for analyzing weaponized interdependence in advanced technology sectors.
Methodological claim in the paper: selection and focus on the semiconductor sector as illustrative of broader advanced-technology sector dynamics under export restraints and chokepoint activation.
The study was a preregistered experiment across seven leading LLMs and twelve investment scenarios covering legitimate, high-risk, and objectively fraudulent opportunities.
Methodological description in the paper stating preregistration, 7 LLMs, 12 scenarios; combined dataset included 3,360 AI advisory conversations and a 1,201-participant human benchmark.
By analyzing agents' reasoning text through a twenty-mechanism scoring framework, targeted prompt interventions causally amplify or suppress specific behavioral mechanisms.
Qualitative and quantitative analysis of agents' chain-of-thought / reasoning text using a 20-mechanism scoring framework; experimental manipulations of prompts reported to change mechanism scores (interpreted causally as interventions on prompts).
Formal network verification has made substantial progress in proving correctness properties but is typically applied in offline, pre-deployment settings and faces challenges in accommodating continuous changes and validating live production behavior.
Authors' summary of the state of the art in network verification (assertion in paper; no empirical data in abstract).
They can produce fluent outputs that resemble reflection, but lack temporal continuity, causal feedback, and anchoring in real-world interaction.
Descriptive claim made in the text contrasting surface-level fluency with missing properties; no empirical data or experiments provided.
User interactions in online recommendation platforms create interdependencies among content creators: feedback on one creator's content influences the system's learning and, in turn, the exposure of other creators' contents.
Conceptual/empirical motivation stated in the paper; motivates the multi-agent bandit modeling of creator interactions in recommender systems.
We ran two large preregistered experiments (N=17,950 responses from 14,779 people) using conversational AI models to persuade participants on a range of attitudinal and behavioural outcomes, including signing real petitions and donating money to charity.
Statement in paper reporting two preregistered experiments, sample sizes (17,950 responses; 14,779 people), use of conversational AI models, and target outcomes including petition signing and charitable donations.
By conceptualizing the emergence of a posthuman economy, this study contributes to interdisciplinary debates on artificial intelligence, digital capitalism, and the transformation of economic organization.
Author-stated contribution of the paper based on conceptual/theoretical work; no empirical validation reported.
There is a robust inverted U-shaped relationship between robotics manufacturing development and urban carbon emissions.
Panel data analysis using 277 Chinese prefecture-level cities from 2008 to 2019; econometric analysis reported in the paper finds an inverted U-shaped association and robustness checks are claimed.
Both rapid model improvement and benchmark quality issues contributed to underestimating agent capabilities.
Synthesis of results: improved LLM performance plus audit findings showing benchmark errors together explain the prior underestimation; based on the re-evaluation and audit described in the paper.
AI automation is a continuum between (i) crashing waves where AI capabilities surge abruptly over small sets of tasks, and (ii) rising tides where the increase in AI capabilities is more continuous and broad-based.
Conceptual framing proposed by the authors (theoretical proposition).