Evidence (4721 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
21267 claims
Filter claims →
Productivity
17978 claims
Filter claims →
Governance
17038 claims
Filter claims →
Human-AI Collaboration
16914 claims
Filter claims →
Org Design
11104 claims
Filter claims →
Innovation
11087 claims
Filter claims →
Labor Markets
6711 claims
Filter claims →
Skills & Training
5616 claims
Filter claims →
Inequality
4343 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 1880 | 496 | 296 | 1854 | 4721 |
| Organizational Efficiency | 2906 | 665 | 438 | 180 | 4210 |
| Governance & Regulation | 2162 | 929 | 480 | 247 | 3866 |
| Technology Adoption Rate | 1533 | 545 | 278 | 210 | 2593 |
| Decision Quality | 1391 | 534 | 321 | 173 | 2429 |
| Output Quality | 1298 | 472 | 231 | 145 | 2153 |
| AI Safety & Ethics | 682 | 821 | 230 | 90 | 1837 |
| Research Productivity | 855 | 253 | 121 | 425 | 1675 |
| Firm Productivity | 1105 | 171 | 175 | 73 | 1531 |
| Task Allocation | 735 | 229 | 361 | 99 | 1433 |
| Market Structure | 457 | 461 | 251 | 47 | 1222 |
| Innovation Output | 673 | 94 | 108 | 36 | 913 |
| Task Completion Time | 499 | 118 | 43 | 38 | 702 |
| Firm Revenue | 458 | 130 | 61 | 26 | 677 |
| Skill Acquisition | 381 | 122 | 113 | 34 | 650 |
| Consumer Welfare | 316 | 176 | 115 | 39 | 648 |
| Employment Level | 223 | 143 | 177 | 53 | 600 |
| Error Rate | 246 | 282 | 44 | 19 | 594 |
| Fiscal & Macroeconomic | 283 | 142 | 78 | 52 | 562 |
| Inequality Measures | 103 | 329 | 106 | 13 | 552 |
| Worker Satisfaction | 225 | 185 | 63 | 30 | 503 |
| Automation Exposure | 158 | 155 | 72 | 37 | 426 |
| Regulatory Compliance | 186 | 126 | 35 | 14 | 362 |
| Team Performance | 193 | 56 | 51 | 24 | 326 |
| Developer Productivity | 224 | 58 | 27 | 13 | 323 |
| Wages & Compensation | 148 | 108 | 50 | 17 | 323 |
| Training Effectiveness | 218 | 44 | 21 | 27 | 313 |
| Job Displacement | 23 | 159 | 53 | 5 | 240 |
| Hiring & Recruitment | 109 | 61 | 32 | 11 | 215 |
| Skill Obsolescence | 16 | 107 | 26 | 6 | 155 |
| Creative Output | 71 | 44 | 28 | 6 | 150 |
| Social Protection | 58 | 31 | 12 | 3 | 104 |
| Labor Share of Income | 29 | 43 | 25 | 2 | 99 |
| Worker Turnover | 45 | 29 | 6 | 4 | 84 |
| Industry | — | — | — | 1 | 1 |
Even under the paper's most favorable treatment of classification disagreements, the research design could not detect a difference in volatility smaller than approximately one-sixth.
Power and measurement-error analysis applied to the null result and the estimated classification reliability.
Accounting for the observed classification error leaves room for a true difference in volatility of approximately one quarter of the comparison group's average volatility.
Sensitivity analysis that propagates the out-of-sample classification reliability estimate through the estimator.
A language-model-based measure of token-level perplexity diverges from the paper's human-oriented processing-cost measure across annual reports.
The paper compares the two measures using 587 filings and reports their cross-sectional correlation.
AI-tilted portfolios do not uniformly outperform minimum-variance portfolios.
The portfolio analysis compares AI-tilted characteristic portfolios with minimum-variance benchmarks using in-sample and rolling out-of-sample performance across multiple market regimes.
Increasing the beta-estimation window reduces coefficient dispersion but produces only modest gains in statistical power.
Sensitivity analyses vary the length of the time-series beta-estimation window and assess the resulting coefficient dispersion and power.
The study's evidence does not establish causal effects or broad generalizability because it uses a purposive sample and primarily cross-sectional or qualitative primary data.
Study limitations stated in the methods and limitations description; evidence consists mainly of descriptive survey/interview responses and practitioner-reported outcomes.
Traditional absorptive capacity theory implicitly treats humans as the sole agents of knowledge absorption, an assumption challenged by the emergence of AI.
Conceptual mapping and comparative synthesis of absorptive-capacity and AI/ML literature.
The evaluation suites were small by benchmark standards, ranging from 9 to 100 cases, and unchanged-suite reruns typically varied by no more than one case.
Reported suite sizes and repeatability observation in the study setup.
AICOME is most useful for recovering a limited number of theoretically important constructs from rich existing datasets, rather than replacing direct survey measurement or entire batteries of jointly missing items.
This conclusion follows from the validation and boundary-condition analyses summarized in the abstract and introduction.
The reviewed qualitative evidence provides deep contextual insight into mechanisms, norms and power dynamics but has limited external generalizability and places less emphasis on causal identification and economy-wide quantification.
The paper's stated assessment of its evidence base, which consists primarily of ethnographies, interviews, participant observation, archival studies and comparative case analysis.
The economic effects of AI on expert work are heterogeneous and are shaped by tacit knowledge, institutional power, trust, market design and professional identity rather than by technical capability alone.
Cross-case comparison of ethnographic and grounded studies across sociology, management and STS.
On OSWorld-Verified, UI-Venus-2-27B achieves a score of 80.5, exceeding the listed Qwen-UI-Agent-27B, Seed-2.1-Pro, and GPT-5.5 systems but trailing Claude-Opus-4.8 at 83.4.
Figure 1 reports comparative OSWorld-Verified benchmark scores.
Holding β fixed while changing batch size changes the effective statistical estimator represented by an exponential moving average.
The survey derives that an EMA with decay β has nominal memory of approximately (1−β)⁻¹ steps, or approximately B/(1−β) tokens when each step contains B tokens.
The retrospective clinical evaluation covered 1,340 heterogeneous patient-trial pairs from 288 de-identified clinical-note cases across four oncology-specific workflows and one broader NIH referral workflow.
Study-design description of the multicenter retrospective evaluation spanning referral, intake, consultation, and tumor-board workflows.
Under the paper's model, expected victims per scam channel are approximately proportional to the scam conversion rate, inversely proportional to the reporting rate, and proportional to the ban threshold, with these effects multiplying.
Analytical Theorem 1 derived from the model of lure conversion, reporting, and channel banning.
The pilot dataset consisted of 18 recorded sales-pitch videos covering three product types, five buyer-persona groups, five pitch stages, and English, Hindi, Telugu, and code-mixed language conditions.
Dataset description in the pilot evaluation.
The paper provides descriptive and statistical associations rather than evidence establishing causal effects.
Cross-sectional observational design based on secondary survey data, with a moderate sample of 128 firms.
Online interactions reconfigure peer relationships, social norms, and identity construction among children and adolescents.
The article draws on qualitative research and theoretical and conceptual work about development and technology, alongside observational evidence.
Correspondence-based fit research is most compatible with surveys, alignment metrics, dyadic or market-matching measures, and causal estimation, while constitutive-fit research is more compatible with qualitative methods, longitudinal process tracing, and analysis of interpretation and affect.
Methodological implications derived from the distinct epistemological assumptions of the two ontologies.
Person–organization fit research rests on two distinct ontologies: a correspondence ontology that treats fit as alignment between person and organization attributes, and a constitutive ontology that treats fit as an enacted and interpretive accomplishment.
Conceptual and theoretical synthesis distinguishing two underlying ontologies; no new empirical data are reported.
Supervisor support was directly associated with higher occupational resilience and lower affective numbing.
Longitudinal panel; supervisor support was measured at T1, while occupational resilience and affective numbing were measured at T2.
Perceived AI autonomy was directly associated with higher occupational resilience and lower affective numbing.
Longitudinal panel; AI autonomy was measured at T1, while occupational resilience and affective numbing were measured at T2, controlling for AI work pressure and supervisor support where specified.
The cluster analysis identifies four country types arranged along a gradient from high digital maturity and high national readiness to low readiness and a low AI publication footprint.
Cluster analysis of country profiles combining AI publication footprint and national AI readiness.
The reviewed research is methodologically concentrated in quantitative, model-centric studies, while qualitative and mixed-method research is limited.
Methodological coding in the systematic review, including analysis of research designs and evaluation approaches.
The reviewed evidence is heterogeneous in its measures, contexts, and outcomes, and many underlying studies provide limited causal identification.
Limitations reported by the narrative review concerning the heterogeneity and methodological quality of the secondary literature.
The dynamics of CEV can be organized along two dimensions: velocity, referring to how quickly vulnerability emerges or changes, and duration, referring to how long vulnerability persists.
Conceptual dimensional framework proposed by the paper; the summary reports no empirical operationalization or quantitative validation.
Research attention is highly uneven across simulator capabilities: controllability, interaction, and stability receive substantially more attention than asset construction, the physics engine, and state feedback.
Cumulative paper counts for each capability in the 200-paper corpus, categorized by principal contribution.
A positive Moran spatial shift indicates that geographic covariates are more spatially clustered in the target domain than in the source domain, whereas a negative shift indicates that they are more spatially dispersed.
Interpretation of the defined Moran spatial shift, calculated as the target-domain Moran's I minus the source-domain Moran's I.
The proposed geographic domain-shift framework captures two complementary forms of difference between source and target regions: feature-distribution differences through mutual-information shift and spatial-structure differences through Moran spatial shift.
Definitions and methodological description of the two proposed domain-shift metrics.
The increase in reported GPU capacity was driven mainly by adoption of newer hardware generations and expansion of medium-scale multi-GPU configurations, rather than by a field-wide shift to very large GPU clusters.
Year-over-year descriptive comparisons of reported GPU counts and hardware generations; configurations using nine or more GPUs accounted for only 11.7% of papers with reported GPU counts in 2025.
The fraction of the real-event excess reproduced by known-null timestamps is highest at detectably active moments and declines to approximately zero at quiet moments: 0.56 for landmark-active anchors, 0.43 for sub-threshold anchors, and −0.04 for quiet anchors.
Stratified known-null experiments using identical six-hour clock-only matching across three pre-event activity classes.
Stages 1, 2, and 5 of KDAF were not exercised in the FinanceBench evaluation because the public benchmark did not provide domain experts or an organizational context.
Evaluation-scope statement in the framework description.
The evaluated CARP system is a hybrid retriever: graph structure controls eligibility, reachability, and provenance, while lexical evidence carries the largest single weight in final evidence ordering.
Description of the implemented CARP algorithm, including the additive composite ranking in which normalized lexical score has weight 0.45.
Latent kernels are not identified when observationally equivalent candidate types have different latent pairs; however, if the map from types to reduced forms is injective, the stored pair (πθ, σθ) can be recovered by dictionary lookup.
Proposition 4.3 and the discussion of observational equivalence in Sections 1 and 4.2.
LLM-based systems have three distinguishable information-retention layers: parametric memory in model weights, contextual memory within the current context window, and external product-level memory such as logs, profiles, knowledge bases, and vector stores.
Conceptual taxonomy of memory mechanisms, supported by cited surveys of memory and retrieval-augmented generation.
As of Q2/Q3 2026, five of six monitored public proxy metrics show stress signals, while SOFR–OIS remains green and appears to be a lagging indicator.
Ranking and monitoring of six public proxy metrics for sensitivity to λ*, reported using Q2/Q3 2026 observations.
Using the empirically observed median junior-tranche thickness δ = 0.15 leaves λ* unchanged but narrows the warning window between transition onset and the critical cliff.
A rerun of the simulation using the reported median junior-tranche thickness from Osberghaus and Schepens (2026).
Platform workers observed in the CSS2023 are younger, more educated, and more likely to belong to higher-income dual-earner households than the precarious-gig-worker profile commonly portrayed in qualitative research.
Descriptive comparison reported by the authors based on the CSS2023 sample; the supplied text does not report numerical differences or statistical tests for these demographic characteristics.
Authenticity measures depended on situational and contextual factors rather than showing a consistent destructive-leadership pattern.
Cross-case comparison of LIWC authenticity measures in the political and business corpora.
Surface indicators of linguistic simplicity were not robust across both political and business domains.
Comparison of surface language measures, including function-word use and sentence or word-length proxies, across the two case studies.
Clout and analytical-thinking measures were inconsistent or context-dependent across the political and business comparisons.
Cross-case comparison of LIWC clout and analytical-thinking scores, with replication assessed across the two domains.
The proposed pilot could serve as a natural experiment for estimating how increased worker agency affects firm behavior, wage outcomes, turnover, and technology adoption.
Evaluation opportunity identified by the proposal; the pilot is not reported as implemented and no estimates are provided.
The negative association with carbon emissions is most robust for the common intensity of AI task exposure, while the substitution-versus-empowerment composition margin is less precisely identified.
Robustness checks using fixed 2016 occupational recruitment weights, province-by-year and industry-by-year fixed effects, and carbon outcomes reconstructed from continuously observed listed firms.
The evidence base is dominated by cross-sectional studies using SEM or PLS-SEM, with 19 studies using SEM or PLS-SEM and few longitudinal or experimental designs.
Methodological coding of the included primary studies.
Only 2 primary studies directly examined Bangladesh, limiting the strength of Bangladesh-specific causal and contextual conclusions.
Systematic review of 40 included primary empirical studies; geographic coverage was assessed during coding.
The urban pollution–carbon reduction synergy outcome is constructed using the interaction between carbon emissions and an environmental pollution index based on industrial wastewater, sulfur dioxide in industrial exhaust, and industrial solid waste.
Variable-construction description in the research-design section; the pollution index is a prefecture-level composite reverse indicator.
Positive social and economic outcomes from DSI are conditional on deliberate attention to inclusion, governance, privacy, and ethics.
Synthesis of the main finding and recurring success factors across qualitative cases; causal inference is limited by the qualitative design.
The launch benchmark contained approximately 900 human-labeled explanations, with the two classes approximately balanced and a slight majority of failed examples at about 54%.
Description of the benchmark dataset in Phase I.
Five of the six tasks have at least 70% of agent-designed methods within algorithmic distance d ≤ 1 of a human method, with Weather the main exception at 48.3%.
Task-specific distributions of minimum algorithmic distance shown in Figure 3a.
Cultural coverage is increasing, but most resources in the cultural block are single-turn: 4 of the 5 listed entries do not test sustained interaction.
Review of the five cultural-resource entries in Table 2, four of which are explicitly marked as single-turn.