Evidence (8974 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
10085 claims
Filter claims →
Productivity
8974 claims
Filtered →
Governance
8062 claims
Filter claims →
Human-AI Collaboration
7749 claims
Filter claims →
Org Design
5057 claims
Filter claims →
Innovation
4896 claims
Filter claims →
Labor Markets
4088 claims
Filter claims →
Skills & Training
3372 claims
Filter claims →
Inequality
2377 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 882 | 244 | 117 | 1097 | 2424 |
| Governance & Regulation | 1010 | 469 | 229 | 135 | 1875 |
| Organizational Efficiency | 977 | 235 | 149 | 90 | 1462 |
| Technology Adoption Rate | 781 | 299 | 143 | 128 | 1362 |
| Research Productivity | 506 | 155 | 74 | 363 | 1110 |
| Output Quality | 555 | 219 | 71 | 70 | 915 |
| Decision Quality | 395 | 200 | 95 | 54 | 751 |
| Firm Productivity | 523 | 67 | 101 | 27 | 724 |
| AI Safety & Ethics | 262 | 309 | 75 | 36 | 688 |
| Market Structure | 195 | 201 | 135 | 30 | 566 |
| Task Allocation | 248 | 77 | 96 | 38 | 464 |
| Innovation Output | 300 | 34 | 55 | 20 | 411 |
| Skill Acquisition | 207 | 75 | 65 | 21 | 368 |
| Employment Level | 138 | 67 | 119 | 24 | 350 |
| Fiscal & Macroeconomic | 156 | 80 | 53 | 33 | 329 |
| Task Completion Time | 211 | 38 | 13 | 16 | 280 |
| Firm Revenue | 183 | 52 | 29 | 5 | 270 |
| Consumer Welfare | 131 | 77 | 48 | 13 | 269 |
| Inequality Measures | 50 | 141 | 54 | 9 | 254 |
| Worker Satisfaction | 104 | 85 | 25 | 13 | 227 |
| Error Rate | 87 | 112 | 11 | 5 | 215 |
| Automation Exposure | 69 | 69 | 37 | 20 | 198 |
| Wages & Compensation | 102 | 49 | 31 | 11 | 193 |
| Team Performance | 115 | 30 | 30 | 11 | 187 |
| Regulatory Compliance | 88 | 74 | 17 | 7 | 186 |
| Training Effectiveness | 109 | 22 | 14 | 21 | 168 |
| Developer Productivity | 116 | 21 | 15 | 8 | 161 |
| Job Displacement | 12 | 92 | 26 | 1 | 131 |
| Hiring & Recruitment | 57 | 12 | 9 | 5 | 83 |
| Skill Obsolescence | 6 | 59 | 10 | 2 | 77 |
| Social Protection | 43 | 17 | 8 | 2 | 70 |
| Creative Output | 35 | 21 | 9 | 4 | 70 |
| Labor Share of Income | 18 | 23 | 17 | 1 | 59 |
| Worker Turnover | 15 | 16 | — | 4 | 35 |
| Industry | — | — | — | 1 | 1 |
Productivity
Remove filter
We present the Governed AI-Assisted Engineering (GAIE) framework, a three-tier graduated human oversight model for agentic code generation in regulated domains.
Proposed framework described in the paper (framework design/contribution).
Digitalization enables service-sector expansion through fintech and e-commerce.
Empirical sectoral data and comparative case studies highlighting fintech and e-commerce impacts in services; policy analysis situates enabling conditions. No numeric sample size or quantified effect in summary.
Digitalization enhances competitiveness in manufacturing.
Empirical sectoral data and comparative case studies focused on manufacturing; China emphasized as a central case. No explicit sample size or quantified effect reported in the summary.
AI, the Internet of Things (IoT), and platform economies contribute to productivity gains across manufacturing, services, and (to a lesser extent) agriculture in emerging markets, with China as a central case.
Mixed-methods approach combining empirical sectoral data, policy analysis, and comparative case studies; China used as a central case. Sample size/quantitative scope not specified in summary.
Welfare analysis finds the AI shock welfare-improving under complementarity between labor and AI capital.
Model welfare calculations (household utility/welfare measures) under parameterizations that assume complementarity between labor and AI capital; numerical comparisons of welfare before and after the AI shock.
A longevity shock acts as a saving-supply disturbance: it deepens the aggregate capital stock.
Model simulation of an exogenous longevity shock (longer lifespans) in the overlapping-generations GE model, producing higher aggregate capital accumulation.
The AI shock produces a front-loaded output expansion that decays monotonically.
Model-implied output dynamics following an AI technology shock shown in numerical simulations.
An AI technology shock acts as a capital-demand disturbance: it raises all rates of return, most sharply the return to AI capital.
Theoretical dynamic overlapping-generations general equilibrium model with endogenous fertility; numerical simulation of an exogenous AI technology shock that increases returns to capital, with model-implied trajectories of rates of return reported.
This work presents a first-of-its-kind integration of OCR-driven document digitalization and LLM-based generative design specifically tailored for large-scale petroleum engineering, providing a robust, scalable solution for thousands of wells and establishing a new industry benchmark.
Authors' claim of novelty and comparative statement versus previous tools (which they state were limited to single-well or text-only data); presented as an asserted contribution rather than empirically benchmarked against other systems.
The study concludes that this scalable AI assistant is essential for large-scale operators to maintain design consistency, institutionalize knowledge from vast historical datasets, and achieve a significant reduction in both labor costs and operational risk.
Authors' conclusion based on deployment experience and reported productivity/quality improvements; no formal causal identification study reported.
A multi-layer auditing mechanism—covering rule-based, logical, and consistency checks—validates the drafts and the LLM refines the output, incorporating domain-specific optimization suggestions based on historical performance trends and regional geological constraints.
System design description detailing the auditing layers and LLM-based refinement incorporating domain-specific optimization using historical performance trends and geological constraints.
By combining user-defined prompts with regional historical databases and offset well statistics from thousands of operations, an AI agent generates comprehensive design drafts.
System architecture and data sources described in the paper; mentions use of regional historical databases and offset well statistics from 'thousands of operations' to generate design drafts.
Integration of OCR enabled the system to process a wide range of design types, including historical hard-copy records, thereby enriching the knowledge base with 'dark data.'
System pipeline description: OCR applied to legacy scanned logs and blueprints converted into structured datasets; claim about expanding the knowledge base with previously unstructured/hard-copy records.
The automated audit identified critical design conflicts in complex multi-well pads that typically elude human oversight.
Deployment report indicating detection of design conflicts in multi-well pad designs by the system's multi-layer auditing mechanism; no numeric count reported.
The automated audit process eliminated over 95% of clerical errors.
Reported measurement from the automated audit process during deployment; claim given as 'eliminated over 95% of clerical errors'.
The time required for generating standard design documents was reduced by approximately 75%, allowing engineering teams to focus on high-level strategy rather than clerical documentation.
Quantitative productivity impact reported in deployment across the stated well portfolio; reduction reported as 'approximately 75%'.
The system was deployed across a portfolio of over 1,000 wells.
Deployment statement in the paper: explicit claim that deployment covered over 1,000 wells.
The paper introduces a scalable, AI-driven assistant system designed to automate multi-disciplinary designs, including drilling, completion, and surface network designs.
System development and architecture described in the paper (NLP + computer vision + LLM pipeline); implementation details provided but no randomized evaluation reported.
By measuring and designing agent power distributions and response functions, it may be possible to better understand, predict, and optimize collective behavior and identify the conditions under which collective intelligence and optimal order emerge.
Proposal/suggestion based on the theoretical framework and analytical insights presented in the paper; no empirical tests or implementation reported.
A system-level utility function parameterized by a risk-appetite coefficient can be used to derive an optimal degree of order that balances productivity, stability, and adaptability.
Theoretical introduction of a system-level utility function and analytical derivation of an optimal order in terms of model parameters (including a risk-appetite coefficient); purely theoretical results, no empirical validation.
Macroscopic properties — including total power, useful power, entropy, order, fragility, and mobility — emerge from these two variables of heterogeneous agents.
Analytical derivations in the paper that connect agent-level variables (power and response functions) to system-level/macroscopic quantities; theoretical mathematical exposition, no empirical sample.
The framework is built on two fundamental agent-level variables: power, which measures agent influence on collective outcomes, and response functions, which determine how agents react to observations.
Theoretical model development and definition of core variables within the paper (analytical/mathematical framework); no empirical sample reported.
A 'favourable transmission path' exists in which AI-induced productivity strengthens purchasing power and effective demand.
Conceptual framework presented in the review (the paper characterises possible transmission paths).
How human-AI teams coordinate and integrate expertise matters as much as the capability available to them.
Overall conclusion drawn from experimental comparisons across different team structures, collaborator counts, and scaffolded vs. unscaffolded conditions.
The scaffolding's performance improvement is most clear in three-person teams.
Reported subgroup analysis by team size showing the largest gains for three-person teams.
A scaffolding that combines shared group memory with simulated human-in-the-loop (HITL) gates yields higher mean performance.
Experimental evaluation comparing teams with and without the proposed scaffolding in the Collaborative Gym environment; reported improvements in mean performance.
Simplicity can be evolutionarily favourable because the specialized decision-maker can capture private returns (rents, status, control, or superior information) that remove the volunteer's dilemma.
Theoretical reasoning within the model; assertion that private payoffs can be larger than payoffs under universal complexity, enabling stable specialization (no empirical sample reported).
The framework links bounded rationality, rational inattention, hierarchy, markets, and cultural evolution, suggesting that simplicity is not a failure of adaptation but a precondition for scalable social organization.
Conceptual claim based on synthesis of theoretical framework; paper presents a unified theoretical perspective rather than empirical validation.
Societies with simpler agents and a specialized decision-making centre can dominate when the costs saved by distributed simplicity exceed the utility lost through reduced individual autonomy and imperfect delegation.
Analytical condition derived from the paper's formal model comparing competing social organizations; model-based (no empirical sample).
The specialized decision-maker need not face a volunteer's dilemma, because its private payoff can exceed that available under universal complexity through rents, status, control or superior information.
Theoretical argument within the formal framework; no empirical tests or sample described in the provided text.
A cognitive division of labour reduces decision costs while preserving much of the value created by knowledge.
Result of the paper's formal comparison between societies with uniform complexity and societies with distributed simplicity plus a specialized decision centre (theoretical/model-based evidence; no empirical sample).
In complex environments, selection can favour heterogeneous populations: most individuals use low-cost heuristics and simplified choice architectures, whereas a minority of agents or institutions specialize in information processing.
Formal theoretical model developed in the paper comparing societies of uniformly complex agents with societies containing simpler agents plus a specialized decision-making centre; no empirical sample reported.
Information improves decisions, pushing the population forward as more information becomes available.
Stated as background motivation; general empirical literature is invoked but no specific study, sample, or data reported in the paper.
The proposed structured inference architecture and task decomposition strategy contribute to the design of LLM-based financial decision systems.
Conclusion/claim in the paper based on the preceding experimental results and analyses; this is a high-level contribution claim rather than a quantified empirical effect.
Standard portfolio optimization exploiting low correlation with the stock index and the variance of each system’s output achieves superior risk-adjusted performance.
Portfolio optimization experiments described in the paper that use system output variance and correlations with the index; evaluated with the same Japanese stock data backtesting framework (no numeric sample size or effect magnitudes given in the excerpt).
Semantic alignment between analytical outputs and downstream decision layers is a critical driver of system performance, consistent with a structured regularization interpretation.
Analysis of intermediate agent outputs reported in the paper (details of metrics, sample size, or statistical strength not provided in the excerpt).
A structured decision architecture that explicitly decomposes investment analysis into fine-grained, domain-informed tasks assigned to specialized inference modules (rather than abstract role-level instructions) improves inference quality and transparency.
Proposed framework described in the paper and evaluated empirically via the reported backtests; claim is primarily methodological and supported by the comparative experiments mentioned in the text.
Fine-grained task decomposition significantly improves risk-adjusted returns compared to conventional coarse-grained designs.
Backtesting experiments on Japanese stock data (prices, financial statements, news, macro information) under a leakage-controlled backtesting setting; results validated via bootstrap confidence intervals, multiple-testing corrections, and subperiod stability analysis (details such as exact sample size or time period not stated in provided text).
Organizations that deliberately architect human-AI relationships are 2.5 times more likely to report superior financial performance.
Reported association from Deloitte's 2026 Global Human Capital Trends survey analysis (paper states '2.5 times more likely').
Organizations that deliberately architect human-AI relationships are twice as likely to exceed AI investment returns.
Association reported in the paper based on analysis of Deloitte's 2026 Global Human Capital Trends survey (over 3,000 business leaders); specific comparative statistic 'twice as likely' reported.
Nearly 60% of workers intentionally use AI at work.
Reported descriptive statistic drawn from Deloitte's 2026 Global Human Capital Trends survey of over 3,000 business leaders across 15 countries (paper cites this survey as data source).
A robot can use prior team experience (externalized CPs) to become a better teammate in future interactions.
Empirical results from the experiment (20 participants, 160 round-level observations) showing improved rescue success and reduced task time when the robot is initialized with a selected prior CP.
People in the MATRX Urban Search and Rescue (USAR) environment can externalize collaboration patterns they discover during teamwork through a chat and reflection interface.
Study setup and observational data from 20 participants in the MATRX USAR environment where participants used a chat and reflection interface to externalize collaboration patterns (as described by authors).
The strongest gains from initializing the robot with prior CPs appear at the beginning of the interaction, suggesting reusable episodic memory helps robots enter collaboration with more effective task knowledge and support smoother early teamwork.
Analysis of performance over interaction progression reported in the study using the same experimental dataset (20 participants, 160 round-level observations); authors describe larger effects early in episodes.
Initializing the robot with a single automatically selected prior CP reduces average task time by 283 seconds.
Same experimental dataset (20 participants, 160 round-level observations); reported comparison of average task completion times between conditions with and without initialization.
Initializing the robot with a single automatically selected prior collaboration pattern (CP) increases rescue success from 25.7% to 41.3%.
Controlled experiment reported in the paper across 20 participants and 160 round-level observations; comparison of rescue success rates between runs with and without initialization using an automatically selected prior CP.
There is a marked geographical concentration of AI-related innovation within a limited number of technologically advanced economies.
Descriptive cross-country and spatial analysis using OECD Patents and Functional Urban Areas (FUAs) databases showing concentration of AI-related patents/innovation in a few advanced economies.
Intangible capital accumulation remains strongly linked to localized sectoral productivity gains.
Cross-country sectoral/localized analysis combining INTAN-Invest (intangible capital) with sectoral productivity measures from OECD STAN/FUAs, using descriptive analysis and panel/robust regressions reported in the paper.
We provide a Deep Learning model trained on this dataset, validated by field experts, and deployed in an industrial setting, serving as an initial benchmark for this public dataset.
Paper reports training of a DL model on the released dataset, expert validation of model outputs, and industrial deployment described by the authors as an initial benchmark.
We create and share the largest public steel microstructure segmentation dataset to date, available under an MIT License with a permanent DOI, contributing a fully annotated, high-resolution dataset to the field.
Dataset release described in the paper: authors state dataset is fully annotated, high-resolution, released under MIT license with DOI and claim it is the largest public dataset for this task.