Evidence (3922 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
20058 claims
Filter claims →
Productivity
17184 claims
Filter claims →
Governance
16099 claims
Filter claims →
Human-AI Collaboration
16034 claims
Filter claims →
Innovation
10501 claims
Filter claims →
Org Design
10496 claims
Filter claims →
Labor Markets
6444 claims
Filter claims →
Skills & Training
5385 claims
Filter claims →
Inequality
4148 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 1820 | 479 | 278 | 1820 | 4588 |
| Organizational Efficiency | 2711 | 616 | 401 | 173 | 3922 |
| Governance & Regulation | 2075 | 886 | 459 | 246 | 3714 |
| Technology Adoption Rate | 1467 | 530 | 258 | 206 | 2488 |
| Decision Quality | 1281 | 496 | 289 | 152 | 2228 |
| Output Quality | 1227 | 447 | 207 | 138 | 2025 |
| AI Safety & Ethics | 634 | 754 | 207 | 83 | 1688 |
| Research Productivity | 826 | 241 | 114 | 422 | 1624 |
| Firm Productivity | 1052 | 154 | 163 | 66 | 1441 |
| Task Allocation | 685 | 211 | 331 | 99 | 1335 |
| Market Structure | 433 | 423 | 242 | 46 | 1150 |
| Innovation Output | 639 | 91 | 105 | 34 | 871 |
| Task Completion Time | 476 | 113 | 43 | 36 | 672 |
| Firm Revenue | 445 | 126 | 58 | 25 | 656 |
| Skill Acquisition | 364 | 119 | 109 | 34 | 626 |
| Consumer Welfare | 288 | 167 | 104 | 31 | 592 |
| Employment Level | 214 | 140 | 174 | 50 | 582 |
| Error Rate | 230 | 251 | 35 | 16 | 535 |
| Fiscal & Macroeconomic | 268 | 136 | 71 | 50 | 532 |
| Inequality Measures | 100 | 307 | 96 | 12 | 515 |
| Worker Satisfaction | 221 | 173 | 60 | 30 | 484 |
| Automation Exposure | 155 | 138 | 65 | 36 | 398 |
| Regulatory Compliance | 171 | 120 | 30 | 13 | 335 |
| Developer Productivity | 222 | 58 | 27 | 13 | 321 |
| Team Performance | 188 | 56 | 50 | 24 | 320 |
| Wages & Compensation | 146 | 104 | 46 | 16 | 312 |
| Training Effectiveness | 207 | 41 | 21 | 26 | 298 |
| Job Displacement | 23 | 153 | 52 | 4 | 232 |
| Hiring & Recruitment | 102 | 57 | 30 | 11 | 202 |
| Skill Obsolescence | 16 | 102 | 24 | 6 | 148 |
| Creative Output | 71 | 42 | 23 | 6 | 143 |
| Social Protection | 57 | 30 | 11 | 3 | 101 |
| Labor Share of Income | 29 | 42 | 24 | 2 | 97 |
| Worker Turnover | 43 | 29 | 6 | 4 | 82 |
| Industry | — | — | — | 1 | 1 |
Organisational changes associated with AI adoption differ across phases of digital HR maturity.
The study interprets interview themes using Dave Ulrich's digital HR progression framework, comprising efficiency, innovation, information, and connection phases.
The effects of transformational leadership operate through cognitive and affective mediators and vary according to contextual moderators.
Review synthesis of meta-analytic findings concerning mediators and boundary conditions.
The relationships between AI adoption orientation, entrepreneurial intention, entrepreneurial behaviour, and business performance vary by venture stage, supporting the characterization of AI adoption as a stage-contingent capability.
The study used measurement-invariance testing and PLS-SEM multi-group analysis to compare 108 new and 86 established women entrepreneurs.
On 20 draw.io tasks, ASIL matches draw.io’s MCP content contract for GPT-5.4 but performs worse than it for sonnet4.6.
The paper reports matched native-interface comparisons on 20 draw.io tasks for two models.
In the 150-request Claude-MCP command benchmark, 73.3% of requests were executed directly, 22.7% required clarification, and 4.0% requested unavailable operations.
Command-profile analysis of the deployed interactive Claude-MCP agent.
AI-related workplace change simultaneously produces efficiency gains or improved work organisation and displacement-related effects for low-skilled workers.
Thematic coding of frontline-worker responses identified workflow optimisation, reduced repetitive tasks and reduced overtime alongside partial job substitution and income decline.
GPT-5.6-Luna with a manager matched GPT-5.6-Terra's single-call accuracy while using 44% of the cost: 77.8% versus 77.0% at $1.50 versus $3.41 per pass.
Five-pass accuracy and cost comparison on 100 LiveCodeBench problems; the accuracy difference was not statistically significant (two-sided p = 0.76), while the cost difference was significant.
Qwen3.8-27B with a manager achieved 86.4% pass@1 versus 87.4% for single-call Claude Fable 5 while costing $9.36 less per 100-problem pass.
Five-pass managed Qwen results and five single-call Fable results on the LiveCodeBench benchmark; the accuracy difference was not statistically resolved (p = 0.73), while the cost saving was reported as p = 0.005.
The organizational context affecting GM decisions should be modeled as co-constituted by GM actions and governance structures rather than treated solely as an exogenous moderator.
Cross-domain theoretical synthesis emphasizing governance, institutional constraints, stakeholder interactions, and operational feedback loops.
In hospitality, general managers often combine strategic and operational authority, so the relationship between leader characteristics and organizational outcomes is conditional on the GM's blended role.
Cross-domain synthesis of hospitality GM literature addressing the locational assumption and the convergence of strategic and operational authority.
Upper Echelons Theory does not function as a universal set of assumptions for hospitality general managers; its locational, dispositional, situational, and temporal assumptions operate as boundary conditions that require qualification in high-contact service settings.
Integrative review and cross-domain theoretical synthesis of 92 empirical and conceptual studies on hospitality general managers, covering GM characteristics, leadership styles, and succession.
The paper concludes that access to AI technologies alone may be insufficient to generate sustained organizational value; firms also need organizational resources and capabilities to integrate AI into business processes.
Conceptual interpretation combining the RBV and four-layer AI framework with secondary evidence on Vietnamese enterprise readiness; the framework is explicitly not empirically validated.
Successful AI implementation in S&OP depends on contextual conditions and mechanisms, and implementation involves both opportunities and barriers.
Paper IV’s analysis of existing AI implementation cases and Paper V’s CIMO-based conceptualization of AI implementation in supply-chain planning.
The first study finds that integration requirements differ across S&OP subprocesses and planning situations, indicating that a one-size-fits-all approach to S&OP is insufficient.
Paper I, described as an empirical multiple-case study examining integration requirements within S&OP subprocesses.
Deployment context can materially change the cost economics of reasoning workloads; the paper uses the Deployment Cost Multiplier (DCM) to compare cloud and on-premises inference costs.
DCM is computed as the ratio of total cloud inference cost to total on-premises inference cost; on-premises costs are estimated from self-run evaluations on an 8x NVIDIA B300 system using amortized capital and operating costs.
Across HumanEval Plus, MBPP, MATH-500, and ASQA, the experiments report that PROGROUTER reduces operating cost relative to key baselines while maintaining strong task-solving performance.
The abstract states the cross-benchmark experimental finding; the reported benchmarks cover code generation, mathematical reasoning, and retrieval-augmented long-form question answering.
The spatial spillover of AI changes over time from an initial negative inter-city factor-siphoning effect to a later positive green-technology radiation effect.
A time-varying spatial Durbin model is applied to the 282-city panel to estimate dynamic cross-city spillovers.
The paper links organizational exposure to cybersecurity incidents, business disruption, economic consequences, and resilience-based mitigation and recovery in an integrated causal framework.
Framework development mapping causal links among exposure, incidents, disruption, economic impacts, and resilience strategies.
The paper presents the framework as a scoping aid rather than a substitute for real-world piloting or probationary evaluation.
Explicit methodological limitation stated in the task-requirements and suitability-mapping discussion.
Digital transformation is negatively correlated with cost stickiness and positively correlated with total asset turnover.
Pearson pairwise correlations among the key variables in the 280-observation panel.
The multi-step-with-examples recipe was more accurate but more expensive per essay than the single-step recipes.
Reported recipe-level comparison of MAE and estimated API cost per grading attempt.
The value difference between productive evaluation of an active policy and restart evaluation of a frontier policy can be exactly decomposed into a restart-policy-quality term plus a productive-path-reuse term.
Equation 4.6 adds and subtracts restart evaluation of the active policy to produce an exact algebraic decomposition.
Organizational selection can shift toward platform-mediated commercialization before that route generates greater total surplus than incumbent-mediated collaboration.
Model comparison between evolutionary viability, determined by organizational-selection payoffs, and productive efficiency, determined by total surplus; incumbent outside-option requirements create a wedge between the two thresholds.
The economic benefit of healthcare AI adoption depends not only on algorithmic performance but also on patient volume, verification time, infrastructure choices, and workflow design.
The case study examined the economic consequences of integrating AI into a clinical workflow and included sensitivity analysis of annual patient volume; the framework also models activity durations, resource use, infrastructure, and lifecycle costs.
There is a nonempty trap region in which a weaker LLM tier uses less energy per satisfied answer but consumes more server time per satisfied answer than the strongest tier.
Proposition 1 derives the condition 1 < S_tilde_j/S_tilde_0 < Delta w_0/Delta w_j under the assumption that the weaker tier has lower per-slot power draw.
Supply chain digitalization functions as a contingent dynamic capability rather than a universally accessible technological fix.
This conclusion is based on the reported asymmetric subgroup effects and the paper's interpretation that resource constraints limit firms' ability to benefit from digitalization.
The average cost per BixBench3 task varied by a factor of 367 across models, from $0.35 to $129.14.
Observed model-level costs averaged across the 20 benchmark tasks; the paper separately discusses cache-adjusted costs for GLM 5.2.
The proposed architecture is intended to balance blockchain trust and verifiability with the flexibility, scalability, and expressiveness of centralized metadata services.
Architectural rationale for separating immutable blockchain records from dynamic centralized services; no measured latency, cost, scalability, or security results are reported.
CEOs most often used an acknowledge–reframe response pattern when they responded to the disruption.
Directed content analysis of coded CEO communicative moves across the sampled communications.
Among more than 5,000 customer-support personnel, AI use increased productivity by 34% for low-skilled employees, by an average of 14% for medium-skilled employees, and produced almost no increase for high-skilled employees.
Reported study of more than 5,000 customer-support personnel; the paper provides no further details about study design or estimation.
GameXpert-Bench evaluates coding agents across three stages of the user-facing game-development lifecycle: initial game generation, bug diagnosis and repair, and multi-turn optimization.
The authors qualitatively analyzed complete human–agent game-development trajectories and identified three recurring categories of artifact-changing interactions; the benchmark then operationalizes each category as a separate track.
The learning process is conditional and can stall at the individual or local level without producing organization-wide gains.
Theoretical argument that the staged transition from individual learning to organizational capability is not automatic; no empirical frequency or effect size is reported.
Low-code conversational platforms such as Copilot Studio provide governance and ease of deployment, but are weak orchestrators for reasoning across more than a handful of connected actions and expose an internal decision path that users cannot inspect or change.
Qualitative architectural assessment based on the authors' enterprise experience; no comparative benchmark or usability study is reported.
The cost of generating custom code with frontier models has fallen substantially, but the costs of reviewing, understanding, and maintaining that code have not fallen comparably.
Conceptual argument based on the authors' description of enterprise development practices; no quantitative cost study is reported.
The paper's five infrastructure cases—OCI, Kubernetes, OpenTelemetry, S3, and PostgreSQL—populate five distinct institutional outcome cells rather than producing a uniform portability result.
The abstract and introduction compare five infrastructure cases and assign each a different degree or layer of portability and substitution value.
PostgreSQL compatibility accelerates database evaluation and developer onboarding but does not, by itself, establish production substitutability or lift-and-shift portability.
The paper cites CockroachDB's documentation of supported and unsupported PostgreSQL features and YugabyteDB's explicit distinction between API compatibility and lift-and-shift migration, along with migration requirements involving schema, data, and application optimization.
S3 compatibility can reduce rewrite risk for workloads within a tested object-operation profile, but it does not establish an OCI- or Kubernetes-style neutral conformance regime in the cited public record.
The paper compares Amazon's S3 API reference with Ceph and MinIO documentation, including documented feature differences and the absence of a cited neutral conformance program.
OpenTelemetry can reduce lock-in at the instrumentation and telemetry-export layers, but it does not by itself make observability backend workflows portable.
The paper describes OpenTelemetry's APIs, SDKs, OTLP, semantic conventions, and Collector, while identifying non-portable backend elements such as dashboards, query languages, alert logic, retention policies, incident workflows, and historical-data migration.
Kubernetes conformance supports portability at the required API core but does not, by itself, establish portability of a production environment across managed Kubernetes services.
The paper cites CNCF's conformance program and distinguishes required API support from dependencies involving identity, storage, networking, ingress, add-ons, observability, upgrades, and operating models.
The paper identifies three corresponding organizational configurations: Strategic Transformation, Scalable Solutions, and Expertise Platform.
Theoretical typology developed through comparative case analysis of eight units in five countries.
Governments create dedicated in-house consulting units for three distinct reasons: retaining control over consequential work, capturing returns from problems that recur across government, and mobilizing expertise dispersed across public and adjacent institutions.
Comparative qualitative analysis and theoretical typology derived from documentary analysis of eight in-house consulting units across five countries.
Cloud computing expands provisioning and service options through elasticity and managed services, but the resulting value depends on workload fit, controls, organizational readiness, security, compatibility, provider dependence, cost, and skills.
Integrative synthesis of cloud architecture and adoption literature, including Omurgonulsen et al. (2021), Jayeola et al. (2022), and Liu et al. (2021).
Artistic research data (ARD) resist one-size-fits-all definitions because they are heterogeneous, context-dependent, and vary across artistic practices.
The publication synthesizes interdisciplinary dialogue, practice-based inquiry, workshops, group deliberation, and contributed texts from an Advanced Study Group involving artist-researchers and research support staff.
The effectiveness of human capital analytics varies by firm size, sector, data maturity, and regulatory environment.
The literature review identifies context sensitivity across organizational and institutional settings; it does not provide quantified subgroup estimates.
Artistic research data (ARD) resist one-size-fits-all definitions because they are heterogeneous, context-dependent, and vary across artistic practices.
The publication synthesizes interdisciplinary dialogue, practice-based inquiry, workshops, group deliberation, and contributed texts from an Advanced Study Group involving artist-researchers and research support staff.
Artistic research data (ARD) resist one-size-fits-all definitions because they are heterogeneous, context-dependent, and vary across artistic practices.
The publication synthesizes interdisciplinary dialogue, practice-based inquiry, workshops, group deliberation, and contributed texts from an Advanced Study Group involving artist-researchers and research support staff.
Artistic research data (ARD) resist one-size-fits-all definitions because they are heterogeneous, context-dependent, and vary across artistic practices.
The publication synthesizes interdisciplinary dialogue, practice-based inquiry, workshops, group deliberation, and contributed texts from an Advanced Study Group involving artist-researchers and research support staff.
Artistic research data (ARD) resist one-size-fits-all definitions because they are heterogeneous, context-dependent, and vary across artistic practices.
The publication synthesizes interdisciplinary dialogue, practice-based inquiry, workshops, group deliberation, and contributed texts from an Advanced Study Group involving artist-researchers and research support staff.
Automating transaction verification may lower marginal verification costs and some audit fees, but new assurance activities such as code and protocol review create fixed costs and expertise premiums, leaving the net effect on audit and compliance costs empirically unresolved.
Conceptual cost decomposition distinguishing automated transaction verification from newly required system-assurance activities; the article explicitly identifies the net effect as an empirical question.
Artistic research data (ARD) resist one-size-fits-all definitions because they are heterogeneous, context-dependent, and vary across artistic practices.
The publication synthesizes interdisciplinary dialogue, practice-based inquiry, workshops, group deliberation, and contributed texts from an Advanced Study Group involving artist-researchers and research support staff.