Evidence (17184 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
20058 claims
Filter claims →
Productivity
17184 claims
Filtered →
Governance
16099 claims
Filter claims →
Human-AI Collaboration
16034 claims
Filter claims →
Innovation
10501 claims
Filter claims →
Org Design
10496 claims
Filter claims →
Labor Markets
6444 claims
Filter claims →
Skills & Training
5385 claims
Filter claims →
Inequality
4148 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 1820 | 479 | 278 | 1820 | 4588 |
| Organizational Efficiency | 2711 | 616 | 401 | 173 | 3922 |
| Governance & Regulation | 2075 | 886 | 459 | 246 | 3714 |
| Technology Adoption Rate | 1467 | 530 | 258 | 206 | 2488 |
| Decision Quality | 1281 | 496 | 289 | 152 | 2228 |
| Output Quality | 1227 | 447 | 207 | 138 | 2025 |
| AI Safety & Ethics | 634 | 754 | 207 | 83 | 1688 |
| Research Productivity | 826 | 241 | 114 | 422 | 1624 |
| Firm Productivity | 1052 | 154 | 163 | 66 | 1441 |
| Task Allocation | 685 | 211 | 331 | 99 | 1335 |
| Market Structure | 433 | 423 | 242 | 46 | 1150 |
| Innovation Output | 639 | 91 | 105 | 34 | 871 |
| Task Completion Time | 476 | 113 | 43 | 36 | 672 |
| Firm Revenue | 445 | 126 | 58 | 25 | 656 |
| Skill Acquisition | 364 | 119 | 109 | 34 | 626 |
| Consumer Welfare | 288 | 167 | 104 | 31 | 592 |
| Employment Level | 214 | 140 | 174 | 50 | 582 |
| Error Rate | 230 | 251 | 35 | 16 | 535 |
| Fiscal & Macroeconomic | 268 | 136 | 71 | 50 | 532 |
| Inequality Measures | 100 | 307 | 96 | 12 | 515 |
| Worker Satisfaction | 221 | 173 | 60 | 30 | 484 |
| Automation Exposure | 155 | 138 | 65 | 36 | 398 |
| Regulatory Compliance | 171 | 120 | 30 | 13 | 335 |
| Developer Productivity | 222 | 58 | 27 | 13 | 321 |
| Team Performance | 188 | 56 | 50 | 24 | 320 |
| Wages & Compensation | 146 | 104 | 46 | 16 | 312 |
| Training Effectiveness | 207 | 41 | 21 | 26 | 298 |
| Job Displacement | 23 | 153 | 52 | 4 | 232 |
| Hiring & Recruitment | 102 | 57 | 30 | 11 | 202 |
| Skill Obsolescence | 16 | 102 | 24 | 6 | 148 |
| Creative Output | 71 | 42 | 23 | 6 | 143 |
| Social Protection | 57 | 30 | 11 | 3 | 101 |
| Labor Share of Income | 29 | 42 | 24 | 2 | 97 |
| Worker Turnover | 43 | 29 | 6 | 4 | 82 |
| Industry | — | — | — | 1 | 1 |
Productivity
Remove filter
The paper distinguishes substantive peer competition from symbolic peer competition by measuring the former with peers’ real digital-transformation investments and the latter with peers’ disclosures or announcements.
Operationalization of the two peer-competition measures in the empirical analysis.
The effects of peer-driven digital transformation vary across industries and market structures.
Reported heterogeneity and supplementary analyses using industry and market-structure splits.
Employee resistance to AI adoption is driven by fears of job displacement and skill obsolescence, and the paper argues that training programmes and transparent change management can help overcome this resistance.
Thematic synthesis and practical recommendations in the systematic literature review; the supplied text reports no employee survey sample or estimated intervention effect.
Visibility-enabled metrics such as typing activity, response times, and screen time may improve measurement precision while also encouraging gaming and performative labor that inflate measured productivity without corresponding improvements in output quality.
Conceptual implication for AI economics and productivity measurement; the paper does not report a causal test or quantify the gap between measured productivity and output quality.
Digital visibility has dual and divergent effects: it can increase felt recognition and relational connection while also creating performative pressures that weaken authentic belonging.
Conceptual theorization grounded in relational cohesion theory and sociomaterial perspectives; supported by literature synthesis rather than primary empirical data.
Twitter-derived signals and Google Trends signals provide complementary information for apparel demand forecasting; neither signal source uniformly dominates the other when used alone.
Pairwise head-to-head comparisons of signal sets, in which Twitter alone and Google Trends alone each won approximately 51%–57% of matchups.
Short-run disruption includes job churn and wage compression for affected groups, while long-run outcomes depend on reskilling, capital re-allocation, and institutions.
Asserted in the supplied example contribution; no longitudinal employment, wage, or reskilling evidence is provided.
Regions with higher human capital and adoption capacity capture more productivity gains, while disadvantaged regions face stagnation.
Presented as a regional heterogeneity claim; no regional panel, productivity measure, or comparative estimate is supplied.
AI substitutes for routine cognitive and manual tasks, shifting worker duties toward nonroutinized, interpersonal, and creative tasks.
Presented as a task-based displacement claim; no task-level dataset or estimates are supplied.
Net employment effects are modest short-run losses, with potential long-run gains if complementary skill investment and policy support occur.
Asserted in the supplied example contribution; the text provides no employment panel, identification strategy results, or quantified estimates.
High-skill cognitive tasks and complementary occupations gain earnings, while routine tasks and low-skill occupations face displacement and wage pressure.
Asserted in the supplied example contribution; no occupational employment or wage data are presented.
AI-driven automation accelerates occupational task reallocation, raising productivity but producing uneven wage effects.
Asserted in the supplied example contribution; no underlying paper, dataset, sample, or statistical analysis is provided.
Ichnology is a useful testbed for formalizing inference because trace evidence is highly ambiguous: one organism can produce diverse traces, while similar traces can result from different organisms or abiotic processes.
Conceptual analysis of ambiguity in trace-based interpretation and comparison of biotic and abiotic explanations.
Realizing a required distinction may require revealing additional information, creating a tradeoff between privacy and the ability to make the required judgment.
Conceptual analysis of realization and informational repair; illustrated across institutional and automated decision contexts.
For middle managers, AI has both positive and negative effects: it supports data analysis and managerial decision-making while creating concerns about automation of some managerial responsibilities.
Cross-study synthesis of findings differentiated by organizational level.
Leadership effects in studies of AI adoption and productivity are endogenous to firm selection and internal processes, so causal studies should use instruments or quasi-experimental designs.
The paper's methodological implication concerning endogeneity and causal inference in AI adoption and productivity research.
Digital-era leadership research increasingly frames leadership as distributed and technologically mediated, with AI systems serving as decision aids, communication intermediaries, or partial substitutes for leader tasks.
Synthesis of emerging e-leadership and digital leadership literatures.
The effects of transformational leadership operate through cognitive and affective mediators and vary according to contextual moderators.
Review synthesis of meta-analytic findings concerning mediators and boundary conditions.
Sales experts showed only partial agreement with the features identified by the model as important: some features with strong predictive power were regarded by experts as unintuitive or unimportant.
Five sales experts reviewed SHAP explanations for four correctly classified regional sales cases through a structured questionnaire, including ratings of feature importance and agreement with explanations.
Incorporating constitutive dynamics into matching theory implies that worker preferences and suitability may be endogenous and shaped by organizational experience rather than fully stable and observable in advance.
Theoretical implication for labor-market matching and market-design models; no equilibrium model or empirical estimate is presented.
Correspondence-based fit research is most compatible with surveys, alignment metrics, dyadic or market-matching measures, and causal estimation, while constitutive-fit research is more compatible with qualitative methods, longitudinal process tracing, and analysis of interpretation and affect.
Methodological implications derived from the distinct epistemological assumptions of the two ontologies.
The two ontologies imply different process explanations for organizational outcomes: correspondence models emphasize trait alignment leading to outcomes such as satisfaction and retention, whereas constitutive models emphasize interpretive work leading to meaningfulness and adjustment.
The article develops separate processual explanations for each ontology; the claim is conceptual rather than an estimate from observed data.
Person–organization fit research rests on two distinct ontologies: a correspondence ontology that treats fit as alignment between person and organization attributes, and a constitutive ontology that treats fit as an enacted and interpretive accomplishment.
Conceptual and theoretical synthesis distinguishing two underlying ontologies; no new empirical data are reported.
In the authors' illustrative macroeconomic exercise, the currently automatable share implies negligible aggregate effects of around 0.1 percentage points per year, Wave 1 implies about 0.9 percentage points, Wave 2 around 4 percentage points, and Wave 3 more than 20 percentage points in annual productivity and price effects.
Order-of-magnitude calculation assuming displaced employment is fully automated over roughly ten years at an even rate, with displaced labor redeployed; the authors explicitly state that these are not forecasts.
Occupations made automatable by Wave 1 show a slight employment decline of about 1%, while occupations first made automatable at Wave 2 or later show flat or rising employment.
Employment-weighted US OEWS changes for 2023–24 and 2024–25, grouped by the wave in which occupations first meet all nine AI capability requirements.
Supervisor support was directly associated with higher occupational resilience and lower affective numbing.
Longitudinal panel; supervisor support was measured at T1, while occupational resilience and affective numbing were measured at T2.
Perceived AI autonomy was directly associated with higher occupational resilience and lower affective numbing.
Longitudinal panel; AI autonomy was measured at T1, while occupational resilience and affective numbing were measured at T2, controlling for AI work pressure and supervisor support where specified.
The relationships between AI adoption orientation, entrepreneurial intention, entrepreneurial behaviour, and business performance vary by venture stage, supporting the characterization of AI adoption as a stage-contingent capability.
The study used measurement-invariance testing and PLS-SEM multi-group analysis to compare 108 new and 86 established women entrepreneurs.
AI perspective is primarily associated with entrepreneurial outcomes among established women entrepreneurs rather than equally across both venture stages.
Multi-group PLS-SEM analysis of new and established women entrepreneurs; the abstract reports that AI perspective becomes significant primarily in established ventures.
Organizational AI acculturation can generate positive private returns while also producing negative social externalities through propagation of biased or stale practices.
Conceptual analysis of the divergence between firm-level incentives and aggregate welfare; the paper proposes studying incentive and regulatory designs rather than reporting an empirical estimate.
Coding was the only reported domain showing mild forgetting for Thomson-1.0-Large relative to its base model, while general mathematical and abstract reasoning remained within the Qwen performance level.
Authors' interpretation of cross-domain results, including coding scores of 39.9% for Thomson-1.0-Large versus 40.9% for Qwen3.5-397B and reasoning scores of 68.4% versus 66.8%.
On 20 draw.io tasks, ASIL matches draw.io’s MCP content contract for GPT-5.4 but performs worse than it for sonnet4.6.
The paper reports matched native-interface comparisons on 20 draw.io tasks for two models.
In the 150-request Claude-MCP command benchmark, 73.3% of requests were executed directly, 22.7% required clarification, and 4.0% requested unavailable operations.
Command-profile analysis of the deployed interactive Claude-MCP agent.
GPT-5.6-Luna with a manager matched GPT-5.6-Terra's single-call accuracy while using 44% of the cost: 77.8% versus 77.0% at $1.50 versus $3.41 per pass.
Five-pass accuracy and cost comparison on 100 LiveCodeBench problems; the accuracy difference was not statistically significant (two-sided p = 0.76), while the cost difference was significant.
Qwen3.8-27B with a manager achieved 86.4% pass@1 versus 87.4% for single-call Claude Fable 5 while costing $9.36 less per 100-problem pass.
Five-pass managed Qwen results and five single-call Fable results on the LiveCodeBench benchmark; the accuracy difference was not statistically resolved (p = 0.73), while the cost saving was reported as p = 0.005.
AI enforcement may generate heterogeneous distributional effects across firms, individuals, large businesses, small businesses, and the informal sector.
Policy implication concerning differential exposure to AI-enabled tax enforcement; no distributional estimates are reported.
Automation of audit, risk-scoring, and tax-processing tasks is expected to reconfigure public-sector labor demand toward data-science and governance roles.
Economic interpretation of the likely labor-market effects of automating tax-administration tasks; no employment dataset or causal estimate is reported.
The paper concludes that access to AI technologies alone may be insufficient to generate sustained organizational value; firms also need organizational resources and capabilities to integrate AI into business processes.
Conceptual interpretation combining the RBV and four-layer AI framework with secondary evidence on Vietnamese enterprise readiness; the framework is explicitly not empirically validated.
In Vietnam, AI adoption is uneven and is concentrated primarily among large enterprises, financial institutions, and technology firms with relatively advanced digital infrastructure and financial resources.
Qualitative contextual analysis based on secondary sources and comparison of global best practices with Vietnam's adoption conditions; no representative enterprise survey is reported.
In the observed Verified comparisons, the treatment-control performance gap was largest at the 20,480-token window and closest to zero at the 262,144-token window.
Three successive Verified comparisons using the same 169 task IDs and fixed 480-second budget; the paper explicitly cautions that the cross-window ordering does not isolate window size from run-era change.
Successful AI implementation in S&OP depends on contextual conditions and mechanisms, and implementation involves both opportunities and barriers.
Paper IV’s analysis of existing AI implementation cases and Paper V’s CIMO-based conceptualization of AI implementation in supply-chain planning.
The first study finds that integration requirements differ across S&OP subprocesses and planning situations, indicating that a one-size-fits-all approach to S&OP is insufficient.
Paper I, described as an empirical multiple-case study examining integration requirements within S&OP subprocesses.
Deployment context can materially change the cost economics of reasoning workloads; the paper uses the Deployment Cost Multiplier (DCM) to compare cloud and on-premises inference costs.
DCM is computed as the ratio of total cloud inference cost to total on-premises inference cost; on-premises costs are estimated from self-run evaluations on an 8x NVIDIA B300 system using amortized capital and operating costs.
Adding GDB and Git MCP servers left ECC largely unchanged for the three tested models but increased extraneous valid calls.
The Easy-Noise and Hard-Noise suites interleaved external-server tasks with design tasks and were compared with clean matched suites for three representative models.
On longer Hard-Noise sessions, Plan-and-Act improved Gemma 4 26B coverage but increased extraneous valid calls.
A 60-task Hard-Noise session was evaluated under a fixed run-scope configuration, comparing ReAct with Plan-and-Act.
Performance gaps between models widen as tasks require more state recovery and dependency management.
ECC was compared across Easy, Medium, Hard, History, and Cross task suites, which progressively include longer dependencies and cross-task state.
Across HumanEval Plus, MBPP, MATH-500, and ASQA, the experiments report that PROGROUTER reduces operating cost relative to key baselines while maintaining strong task-solving performance.
The abstract states the cross-benchmark experimental finding; the reported benchmarks cover code generation, mathematical reasoning, and retrieval-augmented long-form question answering.
The World Economic Forum projects a net global gain of 78 million jobs by 2030, alongside 92 million job losses in more automatable categories, with software- and AI-related roles among the fastest-growing.
World Economic Forum Future of Jobs projection cited by the report; this is a forecast rather than an observed causal estimate.
Employment among developers aged 22–25 fell nearly 20% from its late-2022 peak between 2021 and mid-2025, while employment among more experienced developers grew by approximately 6–12%.
Payroll-based labor-market research covering 2021–2025; the paper explicitly characterizes these labor-market findings as correlational rather than causal.
Google's DORA survey found that individual developer effectiveness increased by 17%, while delivery stability declined by nearly 10%.
Industry-wide developer survey conducted by Google's DORA team.