Evidence (8974 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
10085 claims
Filter claims →
Productivity
8974 claims
Filtered →
Governance
8062 claims
Filter claims →
Human-AI Collaboration
7749 claims
Filter claims →
Org Design
5057 claims
Filter claims →
Innovation
4896 claims
Filter claims →
Labor Markets
4088 claims
Filter claims →
Skills & Training
3372 claims
Filter claims →
Inequality
2377 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 882 | 244 | 117 | 1097 | 2424 |
| Governance & Regulation | 1010 | 469 | 229 | 135 | 1875 |
| Organizational Efficiency | 977 | 235 | 149 | 90 | 1462 |
| Technology Adoption Rate | 781 | 299 | 143 | 128 | 1362 |
| Research Productivity | 506 | 155 | 74 | 363 | 1110 |
| Output Quality | 555 | 219 | 71 | 70 | 915 |
| Decision Quality | 395 | 200 | 95 | 54 | 751 |
| Firm Productivity | 523 | 67 | 101 | 27 | 724 |
| AI Safety & Ethics | 262 | 309 | 75 | 36 | 688 |
| Market Structure | 195 | 201 | 135 | 30 | 566 |
| Task Allocation | 248 | 77 | 96 | 38 | 464 |
| Innovation Output | 300 | 34 | 55 | 20 | 411 |
| Skill Acquisition | 207 | 75 | 65 | 21 | 368 |
| Employment Level | 138 | 67 | 119 | 24 | 350 |
| Fiscal & Macroeconomic | 156 | 80 | 53 | 33 | 329 |
| Task Completion Time | 211 | 38 | 13 | 16 | 280 |
| Firm Revenue | 183 | 52 | 29 | 5 | 270 |
| Consumer Welfare | 131 | 77 | 48 | 13 | 269 |
| Inequality Measures | 50 | 141 | 54 | 9 | 254 |
| Worker Satisfaction | 104 | 85 | 25 | 13 | 227 |
| Error Rate | 87 | 112 | 11 | 5 | 215 |
| Automation Exposure | 69 | 69 | 37 | 20 | 198 |
| Wages & Compensation | 102 | 49 | 31 | 11 | 193 |
| Team Performance | 115 | 30 | 30 | 11 | 187 |
| Regulatory Compliance | 88 | 74 | 17 | 7 | 186 |
| Training Effectiveness | 109 | 22 | 14 | 21 | 168 |
| Developer Productivity | 116 | 21 | 15 | 8 | 161 |
| Job Displacement | 12 | 92 | 26 | 1 | 131 |
| Hiring & Recruitment | 57 | 12 | 9 | 5 | 83 |
| Skill Obsolescence | 6 | 59 | 10 | 2 | 77 |
| Social Protection | 43 | 17 | 8 | 2 | 70 |
| Creative Output | 35 | 21 | 9 | 4 | 70 |
| Labor Share of Income | 18 | 23 | 17 | 1 | 59 |
| Worker Turnover | 15 | 16 | — | 4 | 35 |
| Industry | — | — | — | 1 | 1 |
Productivity
Remove filter
We propose a preliminary reliance-control framework where the level of control can be used to identify AI overreliance and underreliance.
Authors present a conceptual/framework contribution derived from analysis of the twenty-two interviews; this is a proposed (theoretical) framework rather than an experimentally validated one.
Strategic adoption of AI can significantly improve project outcomes and operational performance in the construction industry.
Synthesis of case study findings indicating improved scheduling, risk management, resource allocation, reduced delays and costs, and improved productivity; support is based on the analysed cases rather than a large-scale representative sample.
Artificial Neural Networks (ANN) and predictive modelling support data-driven decision-making in construction.
Paper highlights the use of ANN and predictive modelling in case studies and their role in supporting data-driven decision-making; the summary does not provide quantitative performance metrics for these models.
Quantitative results demonstrate notable improvements in productivity and time efficiency across the analysed cases.
Summary reports quantitative analyses across the case studies showing improvements in productivity and time efficiency; no explicit sample size, statistical significance values, or effect magnitudes provided in the summary.
These enhancements lead to measurable reductions in project delays, operational costs, and safety risks.
Authors state quantitative measurements from analysed cases indicate reductions in delays, costs, and safety risks attributable to AI-driven tools. The summary does not provide numeric magnitudes or sample counts.
AI-driven tools enhance project scheduling, risk management, and resource allocation.
Reported findings across multiple case studies (qualitative and quantitative analyses) where AI applications were applied to scheduling, risk management, and resource allocation tasks. Specific number of cases or statistical tests not provided in the summary.
Successful implementation requires tailored strategies that address contextual, technical, and human factors.
Authors' synthesis and recommendations based on patterns and barriers identified in the included studies.
Client technological readiness plays a positive role in remote auditing.
Reported moderating/mediating findings across included studies summarized in the review indicating client readiness supports remote audit processes.
Technologies such as big data analytics, artificial intelligence, and federated learning have a transformative impact on audit quality and efficiency.
Synthesis of findings from the 10 included empirical studies reporting effects of these technologies on auditing outcomes.
Policy implications: strengthening digital infrastructure, human capital, and innovation capacity is important to ensure inclusive productivity gains from the AI revolution in BRICS economies.
Normative recommendation derived from empirical findings that digital infrastructure complements AI-driven TC and EC and that differential AI effects are linked to country-level capacities; recommendation follows from observed divergence across economies.
The study contributes methodologically by providing a comparative, frontier‑based assessment of AI-driven productivity in emerging economies and by distinguishing innovation (frontier-shifting) and diffusion (efficiency) effects of AI.
Two-stage empirical approach combining Malmquist TFP decomposition (frontier analysis) with panel regressions linking TFP components to multiple AI penetration indicators (patents, investment, robot density, digital infrastructure) across BRICS, 2005–2023.
Digital infrastructure is a critical complementary factor influencing both efficiency improvements and frontier‑shifting technological change.
Regression analysis includes digital infrastructure indicators and reports that better digital infrastructure is associated with positive effects on both EC and TC (either directly or via interaction terms with AI indicators). Panel data over BRICS, 2005–2023.
Adoption-oriented AI indicators, including robot density, contribute to efficiency improvements (EC).
Panel regressions linking Efficiency Change (EC) to adoption-oriented indicators (robot density and similar diffusion measures) show positive associations, interpreted as diffusion improving efficiency rather than shifting the frontier.
Innovation-oriented AI activities (AI patents and research investment) are strongly associated with frontier‑shifting technological change (TC).
Second-stage panel regression analysis relating TC to AI penetration indicators (AI patents, AI research investment), using BRICS panel data (2005–2023). Reported statistically significant positive associations between patent/research investment indicators and TC.
China and India exhibit sustained productivity growth over 2005–2023 driven primarily by technological progress.
Malmquist Total Factor Productivity (TFP) index computed for BRICS and decomposed into Efficiency Change (EC) and Technological Change (TC); time series patterns show sustained TFP growth for China and India with TC as the dominant component. Panel covers BRICS economies (Brazil, Russia, India, China, South Africa) for 2005–2023.
Current LLMs are imperfect spatial reasoners, a problem that AADvark addresses by incorporating external constraint solver tools with a specialized visual feedback mechanism.
Diagnosis followed by methodological response: authors argue LLM spatial reasoning is imperfect and describe AADvark's use of external constraint solvers and visual feedback to mitigate this; empirical evidence not provided in this excerpt.
Unlike previous state-of-the-art systems, AADvark captures the dynamic part interactions with one or more degrees-of-freedom.
Design claim about the system's modeling of dynamic part interactions (method/architecture difference); supported by the authors' system design and comparison to prior state-of-the-art as asserted in the paper excerpt.
In this paper we present a prototype of AADvark, an agentic system designed for this task.
Statement of contribution: presentation of a prototype system (methodological contribution described in the paper); evidence would be the prototype and its implementation details (not provided here).
In order for Agent-Aided Design to make a real impact in industrial manufacturing, we need a system that is capable of generating such 3D assemblies.
Normative/argumentative claim by the authors that industrial impact requires capability to generate 3D assemblies with moving parts; no empirical test provided.
In the past year, researchers have started to create agentic systems that can design real-world CAD-style objects in a training-free setting, a new variety of system that we call Agent-Aided Design.
Literature/field observation asserted by the paper (statement of recent research trend); no sample size or empirical count provided in the excerpt.
Our results suggest that grounding reward design in empirical analysis of information impact and user answerability improves clarification efficiency.
Conclusion drawn from the paper's empirical work: identification of task relevance and user answerability properties, operationalization via RL rewards, and the CLARITI evaluation showing fewer questions for matched resolution rate; abstract does not report experimental details or metrics beyond the 41% reduction.
CLARITI is an 8B-parameter clarification module.
Model specification reported in the abstract; factual description of the trained model's scale (no further empirical detail provided in the abstract).
We operationalize these properties as multi-stage reinforcement learning rewards to train CLARITI, an 8B-parameter clarification module.
Methodological claim: the paper reports implementation of multi-stage RL rewards and training of a clarification model named CLARITI with 8 billion parameters (claim reported in abstract; no training dataset size reported).
Using Shapley attribution and distributional comparisons, we identify two key properties of effective clarification: task relevance (which information predicts success) and user answerability (what users can realistically provide).
Analytical methods reported in the paper: Shapley attribution and distributional comparisons applied to datasets of software engineering tasks and simulated user responses (abstract mentions these methods but gives no numeric sample size).
Humans often specify tasks incompletely, so assistants must know when and how to ask clarifying questions.
Background claim stated in the paper's introduction/abstract; likely supported by literature on underspecified task specifications and/or the authors' motivating examples (no specific sample size or experiment reported in the abstract).
The approach provides a practical path toward more transparent, controllable, and accountable AI use without requiring new model architectures.
Authors' asserted benefit of the proposed interaction-layer framework; no empirical demonstration that transparency, control, or accountability are achieved or that no architectural changes are required in practice.
The framework enables auditable reasoning traces and supports alignment with emerging governance standards, including the EU AI Act and ISO/IEC 42001.
Stated compliance/alignment claim linking the proposed interaction-layer approach to existing regulatory standards; no compliance testing or audit examples reported.
This reframes the question from whether the model can think to whether the human-AI system can reason.
Conceptual reframing stated in the paper; no empirical evidence required as it is a change of perspective.
We introduce 'The Architect's Pen' as a practical method where the human uses the model as an external medium for structured reflection by embedding phases of articulation, critique, and revision into human-AI interaction.
Method description / practical proposal included in the paper; no experimental evaluation, user study, or quantitative validation reported.
This perspective emphasizes collaborative intelligence, combining human judgment and contextual understanding with machine speed, memory, and associative capacity.
Theoretical claim about complementary strengths of humans and models within the proposed framework; presented without empirical tests.
Building on recent work on 'System-2' learning, reflective reasoning can be relocated to the interaction layer and framed as a cognitive protocol that can be structured, measured, and governed using existing systems.
Conceptual extension of prior literature ('System-2' learning) into an interaction-layer protocol; no empirical protocol testing or measurement evidence provided.
Reasoning should be treated as a relational process distributed between human and model rather than an internal capability of either.
Methodological proposal / theoretical framing presented by the authors; no empirical validation reported.
Large language models have advanced rapidly, from pattern recognition to emerging forms of reasoning.
Stated as an observational claim in the paper's introduction; no empirical evaluation or dataset provided.
This approach aligns with emerging compliance expectations, including the EU AI Act and ISO/IEC 42001, by making reasoning processes traceable under real conditions of use.
Claim of regulatory alignment made by the authors; presented as interpretive/legal/standards-relevant argument rather than supported by empirical analysis or legal review data in this excerpt.
Stabilising interaction makes uncertainty and drift visible before enforcement is applied, enabling more precise capability governance.
Normative/operational claim in the paper about the anticipated effect of the proposed interventions; no empirical test or measurement reported in this excerpt.
Together, these layers form a missing operational substrate for governance by increasing signal-to-noise at the point of use.
Argumentative claim from the paper proposing that the combined interventions improve the information available at the decision point; no empirical validation or sample size provided here.
This paper is the first in a five-paper research series on stabilising human-AI reasoning that proposes a two-layer approach: Parts II–IV introduce human-side mechanisms (uncertainty cues, conflict surfacing, auditable reasoning traces) and Part V develops a model-side Epistemic Control Loop (ECL) that detects instability and modulates generation.
Descriptive claim about the structure and scope of the paper series as stated by the authors; internal to the publication (no external dataset).
Large language models are increasingly integrated into decision-making in areas such as healthcare, law, finance, engineering, and government.
Statement in paper describing observed/adoptive trend; no empirical dataset, sample size, or quantitative analysis reported in the text.
The paper proposes a conceptual framework of the underlying mechanisms of the LLM fallacy and a typology of its manifestations across computational, linguistic, analytical, and creative domains.
Author(s) contribution described in the paper (framework and typology); no empirical testing reported in the abstract.
The rapid integration of large language models (LLMs) into everyday workflows has transformed how individuals perform cognitive tasks such as writing, programming, analysis, and multilingual communication.
Author(s) assertion based on literature review and conceptual overview; no empirical sample or experiment reported in the abstract.
This work contributes to the growing body of research on digital sovereignty and the political economy of AI in frontier markets.
Author's concluding claim about the study's contribution to literature.
Many advanced nations are already integrating AI into their core systems.
General descriptive statement in the paper's background/comparative context; no quantitative enumeration or country-sample provided in the excerpt.
To fund this transition, the paper introduces a blended finance structure designed to attract multilateral banks and private venture capital.
Policy/finance architecture proposed in the paper (design description); no funding rounds, commitments, or empirical investor responses reported in the excerpt.
With coordinated reform, AI could boost Cameroon’s long-term productivity by 1.5% to 2.8% annually.
Result reported from the paper's digital infrastructure modeling; no empirical field trial or sampled population reported in the excerpt.
This model draws on international standards from the OECD, UNESCO, and the African Union, alongside the NIST Risk Management Framework.
Paper text states the model's normative/standards sources; descriptive claim about frameworks referenced.
The study proposes a three-layer framework tailored to Cameroon’s specific political economy using comparative policy analysis and digital infrastructure modeling.
Methodological claim in the paper (description of what the study proposes); based on the authors' analytical work rather than reported empirical validation.
Cameroon should not view AI simply as modernization; it must be treated as a sovereign strategy built on institutional economics, deliberate governance, and a solid blended finance architecture.
Normative policy recommendation derived from the paper's comparative analysis and modeling; no empirical trial or longitudinal data reported in the excerpt.
Artificial Intelligence is ... a structural force that determines national competitiveness and economic resilience.
Author assertion supported by literature review and high-level argumentation (comparative policy analysis); no empirical sample or dataset reported in the excerpt.
A hybrid AI-human sprint planning framework should assign algorithmic tools to estimation and backlog formatting while mandating human deliberation for risk assessment and ambiguity resolution.
Theoretical framework proposed by the authors, motivated by the experimental findings (trade-offs observed between efficiency and risk capture/rework) and qualitative analysis.
Human-only planning excels at adaptability.
Controlled experiment comparing human-only, AI-only, and hybrid models with qualitative indicators of planning robustness and adaptability showing superior adaptability for human-only planning.