Evidence (8974 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
10085 claims
Filter claims →
Productivity
8974 claims
Filtered →
Governance
8062 claims
Filter claims →
Human-AI Collaboration
7749 claims
Filter claims →
Org Design
5057 claims
Filter claims →
Innovation
4896 claims
Filter claims →
Labor Markets
4088 claims
Filter claims →
Skills & Training
3372 claims
Filter claims →
Inequality
2377 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 882 | 244 | 117 | 1097 | 2424 |
| Governance & Regulation | 1010 | 469 | 229 | 135 | 1875 |
| Organizational Efficiency | 977 | 235 | 149 | 90 | 1462 |
| Technology Adoption Rate | 781 | 299 | 143 | 128 | 1362 |
| Research Productivity | 506 | 155 | 74 | 363 | 1110 |
| Output Quality | 555 | 219 | 71 | 70 | 915 |
| Decision Quality | 395 | 200 | 95 | 54 | 751 |
| Firm Productivity | 523 | 67 | 101 | 27 | 724 |
| AI Safety & Ethics | 262 | 309 | 75 | 36 | 688 |
| Market Structure | 195 | 201 | 135 | 30 | 566 |
| Task Allocation | 248 | 77 | 96 | 38 | 464 |
| Innovation Output | 300 | 34 | 55 | 20 | 411 |
| Skill Acquisition | 207 | 75 | 65 | 21 | 368 |
| Employment Level | 138 | 67 | 119 | 24 | 350 |
| Fiscal & Macroeconomic | 156 | 80 | 53 | 33 | 329 |
| Task Completion Time | 211 | 38 | 13 | 16 | 280 |
| Firm Revenue | 183 | 52 | 29 | 5 | 270 |
| Consumer Welfare | 131 | 77 | 48 | 13 | 269 |
| Inequality Measures | 50 | 141 | 54 | 9 | 254 |
| Worker Satisfaction | 104 | 85 | 25 | 13 | 227 |
| Error Rate | 87 | 112 | 11 | 5 | 215 |
| Automation Exposure | 69 | 69 | 37 | 20 | 198 |
| Wages & Compensation | 102 | 49 | 31 | 11 | 193 |
| Team Performance | 115 | 30 | 30 | 11 | 187 |
| Regulatory Compliance | 88 | 74 | 17 | 7 | 186 |
| Training Effectiveness | 109 | 22 | 14 | 21 | 168 |
| Developer Productivity | 116 | 21 | 15 | 8 | 161 |
| Job Displacement | 12 | 92 | 26 | 1 | 131 |
| Hiring & Recruitment | 57 | 12 | 9 | 5 | 83 |
| Skill Obsolescence | 6 | 59 | 10 | 2 | 77 |
| Social Protection | 43 | 17 | 8 | 2 | 70 |
| Creative Output | 35 | 21 | 9 | 4 | 70 |
| Labor Share of Income | 18 | 23 | 17 | 1 | 59 |
| Worker Turnover | 15 | 16 | — | 4 | 35 |
| Industry | — | — | — | 1 | 1 |
Productivity
Remove filter
Prior research has emphasized GenAI’s ability to enhance productivity and creative outcomes.
Literature review / background statements in the paper referencing prior studies (no sample size specified in the paper's statement).
ChatGPT Pro performed best among the tested models, occasionally constructing counterexamples and corrected proofs.
Qualitative comparison across models (Gemini, Refine, Claude, ChatGPT Pro) on the 4 papers showing ChatGPT Pro sometimes produced counterexamples and corrected proofs.
The deployed Archi instance offers retrieval and analysis capabilities by combining documentation, historical data, and live monitoring systems.
Paper's description of the deployed system's data sources and capabilities (documentation, historical data, live monitoring).
An instance of Archi has been deployed for the Computing Operations team of the CMS experiment at CERN's LHC since February 2026 as a support agent for technical operators.
Reported production deployment in the paper (deployment date, target team and site).
Archi is an open-source, end-to-end framework for scientific collaborations that combines the systematic ingestion and organization of heterogeneous data sources with the deployment of configurable, private, and extensible agents that retrieve and reason over them.
Paper's system description / implementation claims (architecture and feature list). No numeric evaluation provided for this descriptive claim.
The paper concludes with policy recommendations to foster a conducive environment for AI integration, positioning Algeria to leverage technological advances for sustainable economic growth.
Concluding statement in the paper summarizing recommended policy actions; framed as guidance rather than empirically tested interventions.
Targeted investments and policy reforms could accelerate AI adoption and productivity gains in Algeria.
Policy recommendation inferred from the study's comparative findings and supported by citations to Brynjolfsson, Rock, and Syverson (2017) and McKinsey & Company (2023); presented as a prospective/conditional claim rather than an empirically estimated causal effect within the paper.
Artificial intelligence (AI) is rapidly transforming global economies by enhancing productivity, enabling innovation, and reshaping labor markets.
Framing claim supported by citations to Agrawal, Gans, & Goldfarb (2019) and Acemoglu & Restrepo (2020) as described in the paper's introduction; no primary empirical estimate reported in this paper.
Through a case study on house price prediction, we find that AACT outperforms traditional AI-based decision-support in reducing over-reliance on AI.
Empirical comparison reported in a case study (house price prediction) between AACT and traditional AI decision-support; includes measured over-reliance and statistical comparison (sample size not reported in abstract).
We introduce the AI-Assisted Critical Thinking (AACT) framework, which leverages a domain-specific AI model’s counterfactual analysis of human decision to help decision-makers identify potential flaws in their decision argument and support the correction of them.
Paper presents a new framework (AACT) and describes its design; demonstrated via a case study (house price prediction).
A four-stage roadmap toward self-evolving agent ecosystems and concrete recommendations for practitioners can guide navigation of the transition to agentic systems.
Prescriptive contribution of the paper: a proposed four-stage roadmap and practitioner recommendations derived from the preceding analysis (theoretical/prescriptive; no empirical validation or sample size reported).
Agentic Engineering is an emergent discipline that is distinct from software engineering in its core object of study, control model, and human role.
Conceptual proposal and definitional work in the paper outlining new discipline characteristics (theoretical, no empirical testing or sample size reported).
The historical arc from licensed software to SaaS to what we term Agent-as-a-Service (AaaS) shows that each shift transferred additional complexity away from end-users.
Historical/architectural trend analysis presented in the paper (qualitative; references to industry evolution, no quantitative sample size reported).
The emergence of AI agents—systems where large language models serve as the primary reasoning engine, dynamically generating and discarding code as an instrumental resource—constitutes a fundamental restructuring of the software paradigm rather than an incremental improvement.
Argument based on first-principles analysis of complexity scaling and conceptual comparison between traditional software and agentic systems (theoretical analysis presented in the paper).
ALE is intended not merely as another leaderboard, but as an instrument for closing the gap between benchmark success and GDP-relevant impact.
Author-stated intent and high-level goal of the benchmark.
ALE is designed as a living benchmark: its task pool grows continuously as new workflows and industries are onboarded.
Design and maintenance policy described by the authors.
ALE was developed in collaboration with 250+ industry experts.
Author statement specifying collaborator count.
This paper introduces Agents' Last Exam (ALE), a benchmark designed to evaluate AI agents on long-horizon, economically valuable, real-world tasks with verifiable outcomes.
Description of benchmark introduced by the authors (design claim).
Recent AI systems have achieved strong results on a wide range of benchmarks.
Statement in paper (background/context); refers to existing benchmark results in the literature (no specific benchmarks or datasets named in this excerpt).
The open-source implementation includes audit trails and confidence scoring, providing a replicable foundation for LLM-based actuarial variable extraction in property-casualty insurance.
Authors state the released implementation is open-source and includes audit trail and confidence scoring features; presented as part of the contribution.
Integration with chain ladder reserving demonstrates practical actuarial value: severity-segmented analysis reduced reserve estimation error from 6.5% to 4.0%.
Applied the extracted severity segmentation to chain ladder reserving in an integration experiment; reported reserve estimation error decreased from 6.5% to 4.0%. Sample size/portfolio details not stated in the claim.
We validate 14 core variables using two independent clinical expert reviewers scoring 20 synthetic claims on a five-point Likert rubric, achieving mean scores above 4.0 and a weighted kappa of 0.53.
Validation experiment: two independent clinical expert reviewers scored 20 synthetic claims on a 5-point Likert scale for 14 core variables; reported metrics are mean Likert scores (>4.0) and weighted kappa = 0.53.
A modular four-script Python pipeline processes synthetic FHIR-based claims data and real claims documents, extracting 36 actuarial variables across reserving, ratemaking, and claims management categories.
Authors report implementation of a four-script Python pipeline applied to synthetic FHIR-based claims and real documents, with 36 target variables defined.
We present a proof-of-concept framework using large language models (LLMs) to extract structured actuarial variables from unstructured claims data.
Authors implemented a prototype framework described in the paper (implementation details and pipeline described).
The method is accessible to public entities under budget constraints because it used free AI models.
Author reports that the deployments used free AI models rather than paid services and were implemented within the budgets of the two public units.
The method operates within protocols designed to comply with international and national data-protection law and with the principles of public administration.
Author statement that the method used protocols designed for legal compliance; paper reports no detected incidents and claims protocol adherence.
The analysis is consistent with the hypothesis that the method is portable across agencies with distinct mandates.
Observed positive outcomes in two distinct public-sector units (SES/CONT and UCI/SEDET) after applying the same methodology; author frames this as consistency with portability hypothesis.
UCI/SEDET analyzed cases totaling USD 104.3 million in financial volume during the period examined.
Aggregate monetary total of cases analyzed reported from SEI-GDF official indicators in the paper.
UCI/SEDET issued 288 formal recommendations to public managers during the examined period.
Count of formal recommendations reported from SEI-GDF official indicators as presented in the paper.
UCI/SEDET recorded a 92% increase in technical-report production during the period examined.
Quantitative production figures from SEI-GDF official indicators reported in the paper for UCI/SEDET.
Official indicators from SEI-GDF recorded an average processing time fall of 50% at UCI/SEDET during the period examined.
Quantitative before–after measurement from SEI-GDF official indicators for UCI/SEDET as reported in the paper.
Official indicators from the Electronic Information System of the Federal District Government (SEI-GDF) recorded an average processing time fall of 18.2% at SES/CONT during the period examined.
Quantitative before–after measurement from SEI-GDF official indicators for SES/CONT as reported in the paper.
The method was applied in two distinct units: the Sectoral Internal Control Office of the Federal District Department of Health (SES/CONT) throughout 2024, and the Internal Control Unit of the Federal District Department of Economic Development, Labor and Income (UCI/SEDET) throughout 2025.
Paper reports implementation timelines and unit names; described as auditable cases.
The author developed a four-layer structured pedagogical methodology for teaching generative-AI use in the public sector.
Author description of the methodology in the paper; applied in two case units.
Work Flexibility is the strongest predictor of Employee Productivity (β = 0.562, p < 0.001), indicating flexible working conditions play an important role in improving employee performance and work efficiency.
Reported quantitative result from the study using PLS-SEM; β and p-value provided in the paper indicating the largest standardized effect among predictors. Sample size not reported in the excerpt.
Human-Centric AI Adoption has a positive and statistically significant effect on Employee Productivity (β = 0.263, p = 0.028).
Reported quantitative result from the study using Partial Least Squares Structural Equation Modeling (PLS-SEM); β and p-value provided in the paper. Sample size not reported in the excerpt.
The productivity-enhancing effect of fintech is stronger in regions with higher levels of economic development.
Heterogeneity/subsample analysis reported for regional economic development levels using the sample of Chinese A-share listed manufacturing firms (2015–2023); paper states fintech's effect on TFP is more pronounced in more economically developed regions (no subgroup sample sizes or quantitative estimates provided in the excerpt).
The productivity-enhancing effect of fintech is more pronounced in high-tech industries.
Heterogeneity/subsample analysis in the paper using the sample of Chinese A-share listed manufacturing firms (2015–2023); paper reports stronger fintech–TFP effects in high-tech industry subsample (no subgroup sample sizes or coefficients provided in the excerpt).
The positive effect of fintech on corporate total factor productivity operates primarily through the channels of supply chain finance and innovation effects.
Mediation/ mechanism analysis reported in the study using the same sample of Chinese A-share listed manufacturing firms (2015–2023); paper states supply chain finance and innovation as the primary channels (specific mediation estimates not provided in the excerpt).
Fintech development can significantly enhance corporate total factor productivity for Chinese A-share listed manufacturing firms.
Empirical analysis on a sample of Chinese A-share listed manufacturing enterprises covering 2015–2023; result described as statistically significant in the paper (specific estimation methods and sample size not provided in the excerpt).
The findings suggest that twin-based market research is no longer gated by data design, but by item volume, model selection, and a small set of construction-level decisions.
Interpretive conclusion based on empirical results across the construction-method grid and performance patterns (discussion/implication in paper).
Best-cell Fisher-z rank-order correlation reaches r = 0.590 on the SOEP held-out evaluation set.
Reported best-performing cell Fisher-z (or Fisher-transformed correlation) from held-out evaluation on SOEP.
Best-cell accuracy reaches 78.8% on the SOEP held-out evaluation set.
Reported best-performing cell accuracy from held-out evaluation on SOEP.
Switching the embedding from a narrative persona summary to a raw dialog history of past responses raises hold-out accuracy in every model-by-reasoning cell at the 100 percent depth.
Empirical comparison between two embedding methods at 100% information depth across all model-by-reasoning cells (reported in results).
Twin quality rises with information depth but with diminishing returns past the 75 percent entropy quartile, which acts as a cost-efficient Pareto point relative to the best-performing 100 percent cells.
Empirical evaluation across information-depth conditions, comparing hold-out performance by normalized Shannon entropy quartiles (reported in results).
Under linear local composition, every protocol tree defines a barycentric coordinate chart on the simplex of leaf weights; Tamari-cover reparameterizations of protocol trees preserve complementarity, and for N = 4 these reparameterizations satisfy the pentagon identity.
Mathematical construction and proofs in the paper linking protocol trees, barycentric coordinates, Tamari lattice reparameterizations, and the pentagon identity (theoretical work; no empirical sample).
For N = 2 in regression under squared loss, the optimal linear-pooling weight has a closed form and admits a residual-correction interpretation.
Closed-form derivation and interpretation provided in the paper (mathematical derivation; no empirical sample).
Across our large-scale empirical analysis, Parthenon substantially improves the performance of state-of-the-art models and harnesses on legal-matter tasks.
Reported evaluations in the paper comparing baseline state-of-the-art models/harnesses to the Parthenon framework across their empirical dataset (Harvey LAB), claiming substantial performance gains.
An anti-leakage learning loop converts scored failures into task-agnostic edits to skills, tools, and knowledge, letting the system improve with experience without touching model weights.
Paper describes a proposed/implemented learning loop (anti-leakage) that translates scored agent failures into edits to non-weight system components (skills, tools, knowledge) and claims this enables improvement without model weight updates.
We introduce Parthenon, a self-evolving legal-agent framework that factors Model, Harness, Agent roles, legal Knowledge, deterministic Tools, and procedural Skills into auditable surfaces for source traceability, date and number grounding, deliverable compliance, and issue closure.
Paper describes the design and implementation of the Parthenon framework and its modular decomposition into Model, Harness, Agent roles, Knowledge, Tools, and Skills, claiming these enable auditable traces and grounding.