Evidence (8974 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
10085 claims
Filter claims →
Productivity
8974 claims
Filtered →
Governance
8062 claims
Filter claims →
Human-AI Collaboration
7749 claims
Filter claims →
Org Design
5057 claims
Filter claims →
Innovation
4896 claims
Filter claims →
Labor Markets
4088 claims
Filter claims →
Skills & Training
3372 claims
Filter claims →
Inequality
2377 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 882 | 244 | 117 | 1097 | 2424 |
| Governance & Regulation | 1010 | 469 | 229 | 135 | 1875 |
| Organizational Efficiency | 977 | 235 | 149 | 90 | 1462 |
| Technology Adoption Rate | 781 | 299 | 143 | 128 | 1362 |
| Research Productivity | 506 | 155 | 74 | 363 | 1110 |
| Output Quality | 555 | 219 | 71 | 70 | 915 |
| Decision Quality | 395 | 200 | 95 | 54 | 751 |
| Firm Productivity | 523 | 67 | 101 | 27 | 724 |
| AI Safety & Ethics | 262 | 309 | 75 | 36 | 688 |
| Market Structure | 195 | 201 | 135 | 30 | 566 |
| Task Allocation | 248 | 77 | 96 | 38 | 464 |
| Innovation Output | 300 | 34 | 55 | 20 | 411 |
| Skill Acquisition | 207 | 75 | 65 | 21 | 368 |
| Employment Level | 138 | 67 | 119 | 24 | 350 |
| Fiscal & Macroeconomic | 156 | 80 | 53 | 33 | 329 |
| Task Completion Time | 211 | 38 | 13 | 16 | 280 |
| Firm Revenue | 183 | 52 | 29 | 5 | 270 |
| Consumer Welfare | 131 | 77 | 48 | 13 | 269 |
| Inequality Measures | 50 | 141 | 54 | 9 | 254 |
| Worker Satisfaction | 104 | 85 | 25 | 13 | 227 |
| Error Rate | 87 | 112 | 11 | 5 | 215 |
| Automation Exposure | 69 | 69 | 37 | 20 | 198 |
| Wages & Compensation | 102 | 49 | 31 | 11 | 193 |
| Team Performance | 115 | 30 | 30 | 11 | 187 |
| Regulatory Compliance | 88 | 74 | 17 | 7 | 186 |
| Training Effectiveness | 109 | 22 | 14 | 21 | 168 |
| Developer Productivity | 116 | 21 | 15 | 8 | 161 |
| Job Displacement | 12 | 92 | 26 | 1 | 131 |
| Hiring & Recruitment | 57 | 12 | 9 | 5 | 83 |
| Skill Obsolescence | 6 | 59 | 10 | 2 | 77 |
| Social Protection | 43 | 17 | 8 | 2 | 70 |
| Creative Output | 35 | 21 | 9 | 4 | 70 |
| Labor Share of Income | 18 | 23 | 17 | 1 | 59 |
| Worker Turnover | 15 | 16 | — | 4 | 35 |
| Industry | — | — | — | 1 | 1 |
Productivity
Remove filter
Authors who received LLM-generated feedback had a significantly higher likelihood of revising their manuscripts, corresponding to a 12.55% relative increase over the baseline revision rate.
Randomized field experiment comparing treatment (LLM feedback) vs control; sample described as >31,000 arXiv preprints and >45,000 researchers; reported comparative revision rate and statistical significance.
A human-centred approach underpinned by ongoing reskilling and ethical governance is vital for sustainable workforce evolution in the Indian IT sector.
Authors' policy/recommendation derived from their literature synthesis and thematic analysis (qualitative conclusion).
The paper introduces a conceptual framework for hybrid intelligence within the Indian IT sector.
Authors present a new conceptual framework as part of this qualitative research article (conceptual contribution).
Collaboration between humans and AI enhances decision-making, efficiency, and innovation.
Reported result from thematic evaluation of literature and secondary data (qualitative synthesis). No sample size or quantified effect provided.
AI improves overall organisational productivity.
Authors' synthesis of peer-reviewed studies and secondary data indicating productivity impacts (qualitative literature review). No quantitative sample size reported.
AI increases human capacities.
Conclusion from comprehensive analysis of peer-reviewed literature and thematic evaluation of secondary data (literature review). No primary sample size reported.
Time and effort dissociate: participants reported lower subjective effort with AI despite equivalent completion times.
Empirical result reported in the abstract: subjective effort ratings were lower for AI-assisted conditions even though measured completion times were equivalent (preregistered study, N = 1237).
Participants predicted AI to be significantly faster.
Empirical result reported in the abstract: participants' predicted completion times indicated AI-assisted completion would be faster than independent completion (statistical significance claimed). Sample from preregistered study (N = 1237).
Large language models (LLMs) have the potential to boost human productivity by speeding up task completion -- provided users know when to offload cognitive work to them.
Framing/introductory claim in the paper (theoretical/argumentative), no direct empirical evidence reported in the abstract.
The aim is to keep autonomous agency composable while keeping accountability non-negotiable, so that coordination itself can become shared infrastructure for a human-AI society that is open, pluralistic, and governable.
Stated design/ethical objective in the paper; normative claim about intended social and governance outcomes rather than an empirically validated result.
FP is designed to wrap and bridge existing protocols rather than replace them, enabling incremental adoption while reducing integration and governance overhead.
Design rationale/claim in the paper about interoperability and incremental adoption strategy; no empirical deployment, integration case studies, or measured overhead reductions presented.
FP treats policy, provenance, and audit as first-class concerns.
Design/architectural claim in the paper stating that policy, provenance, and audit are prioritized within FP; no empirical compliance or audit trials presented.
FP provides economic primitives for metering, receipts, and settlement.
Design claim in the paper listing economic primitives as part of FP; no deployment or economic experiments reported.
FP supports native multi-party organization and event-based collaboration.
Feature/architecture claim in the paper describing native support for multi-party organization and event-driven collaboration; no empirical evaluation or user studies provided.
FP unifies heterogeneous entities, including agents, tools, resources, humans, institutions, and organizations.
Design specification/feature claim in the paper describing FP's data and entity model; no empirical interoperability study reported.
This paper introduces the Foundation Protocol (FP), a graph-first coordination layer for an emerging human-AI society.
Claim of authorship/introduction in the paper; architectural/design proposal rather than an evaluated system.
Agents need to form reliable relationships, organize multi-agent work, exchange value, support an AI economy, and stay safe and accountable under real-world oversight.
Normative/requirements statement in the paper describing necessary capabilities for scaled multi-agent systems; no empirical validation or experimental data provided.
Autonomous agents are moving from tools into a layer of social infrastructure: they browse, purchase, deploy software, manage systems, and increasingly interact with one another.
Statement in the paper's introductory/abstract text presenting an observed trend; conceptual/qualitative claim without empirical data or measured sample.
Prior work has demonstrated that people generally find AI narrative explanations to be understandable, trustworthy, and convincing for changing beliefs and opinions.
Citation to prior literature reported in the paper (background literature review claiming general findings about perceptions of AI narrative explanations).
Narrative explanations increased reliance on the AI, both when the AI prediction was correct and when it was incorrect.
Findings from the paper's human behavioral experiment reporting increased reliance on AI with accompanying narratives under both correct and incorrect AI prediction conditions.
The development of LLM agents has led to a growing body of work on knowledge-work AI, including coding, research, and healthcare.
Statement grounded in observation of recent literature trends and the cited body of work on LLM agents applied to coding, research, and healthcare domains.
These cases show how benchmark design choices shape the strongest work claim a score can support, and where gaps arise between the benchmarked task, tested setting, scored product, and broader work claim.
Qualitative findings from the three case analyses demonstrating how different design choices limit or enable particular work claims and exposing gaps between task, setting, and scored product.
APEX-SWE [is] a software-engineering benchmark with executable scored products.
Description of the APEX-SWE benchmark in the paper's case analysis.
OfficeQA Pro [is] a grounded document-analysis benchmark scored by final answers.
Description of the OfficeQA Pro benchmark in the paper's case analysis.
GDPval [is] a non-code occupational deliverable benchmark.
Description of the GDPval benchmark in the paper's case analysis.
We demonstrate the approach through three benchmark case analyses: GDPval, OfficeQA Pro, and APEX-SWE.
Empirical/methodological demonstration reported in paper via three case analyses of existing benchmarks; the paper applies its three-step approach to each case.
To name the work activity being evaluated and distinguish it from common benchmark tasks, we derive an inventory of 18 work activities from the O*NET occupational task database.
Method described in paper: mapping/derivation from the O*NET occupational task database to produce an inventory of 18 work activities.
We translate these concerns into benchmark design and reporting guidance, covering how tasks should be mapped to work activities, how tested settings should specify materials, tools, roles, and constraints, and how scoring should focus on the work product left by the system.
Paper provides prescriptive guidance derived from conceptual analysis and the reviewed literature; guidance illustrated via application to case benchmarks.
We review work studies showing that knowledge work is organized through roles and responsibilities, local materials and tools, and artifacts that must remain usable in downstream workflows.
Literature review of work studies cited in the paper; synthesis of organizational features of knowledge work.
This paper contributes a three-step approach for making explicit how benchmarked tasks represent the work claims attached to their scores: defining the work activity under evaluation, specifying the tested setting, and scoring the appropriate work product.
Methodological contribution described in paper; approach presented and motivated, and later applied in case analyses (three benchmark case studies).
The paper proposes five evaluation dimensions for AutoResearch systems: novelty, validity, impact, reliability, and provenance.
Paper explicitly proposes these five dimensions as an evaluation rubric; conceptual proposal.
The field can be organized around five workflow conditions: literature and research grounding; hypothesis formation and planning; experimentation and tool use; feedback, validation, and review; and reporting and knowledge communication.
Authors propose this five-condition organizational framework as part of their survey and synthesis; conceptual contribution.
Vibe Research denotes the human-steered region of prompt-based assistance and human-verified execution within AutoResearch.
Paper-introduced terminology and conceptual delineation of a sub-region of the AutoResearch spectrum; definitional statement.
AutoResearch is defined as the developmental spectrum of AI-powered scientific workflow automation.
Paper provides an explicit definitional framing (terminology introduced by authors); conceptual contribution rather than empirical finding.
This shift marks a transition from task-level AI for science to workflow-level research automation.
Conceptual argument backed by literature survey and examples of systems that coordinate multiple research tasks; no single quantitative study reported.
Scientific research is being reshaped by AI systems that move beyond isolated assistance toward longer-horizon workflows spanning literature grounding, hypothesis generation, experimentation, validation, reporting, and revision.
Survey / conceptual synthesis of recent AI research systems and literature; paper presents this as an observed trend rather than reporting original empirical measurements.
The AI-driven econometric approach outperforms traditional approaches by delivering more accurate forecasting and more timely policy recommendations.
Explicit claim in the paper that the approach outperforms traditional methods by producing more accurate forecasts and timelier recommendations; the excerpt contains no quantitative comparison, performance metrics, statistical tests, or sample sizes.
The framework relies on distributed data processing and MLOps pipelines to enable system scalability and continuous model improvement.
System architecture description in the paper stating use of distributed processing and MLOps pipelines; no performance benchmarks, scalability tests, or deployment metrics are reported in the excerpt.
The proposed approach uses ensemble models and deep learning combined with econometric methods to ensure both model interpretability and robust findings.
Methodological claim in the paper describing use of ensemble and deep learning models integrated with econometric techniques; no reported evaluation metrics, interpretability measures, or robustness tests in the provided text.
Combining structured economic indicators with unstructured data from job postings and skill descriptions provides a real-time picture of employment patterns, wage changes, and skill requirements.
Paper describes integrating structured and unstructured data sources (economic indicators, job postings, skill descriptions) to produce a real-time view; no empirical metrics, evaluation sample, or quantitative validation given in the excerpt.
An AI-based econometric system that incorporates machine learning algorithms and extensible data processing can enhance labor market predictions and research compared with traditional econometric models.
Methodological description in the paper stating development of an AI-based econometric system that incorporates ML and extensible data processing; no sample size or empirical evaluation statistics provided in the text excerpt.
These findings challenge narratives that automation and digitalization induce net job loss in manufacturing.
Interpretation based on the paper's empirical results showing positive effects of digital transformation on labor demand and demand for skilled workers (Chinese A-share manufacturing firms, 2011–2024). (Sample size not stated in provided text.)
Digital transformation enhances employees' digital literacy.
Mechanism analysis reported in the paper using firm-level measures of employee digital skills/digital literacy as an intermediate outcome (Chinese A-share manufacturing firms, 2011–2024). (Sample size not stated in provided text.)
Increased total factor productivity (driven by digital transformation) promotes both the amount of labor demanded and the intensity of factor input.
Mechanism/mediation analysis linking digital transformation → TFP → labor demand and factor-input intensity in the firm-level regressions (Chinese A-share manufacturing firms, 2011–2024). (Sample size not stated in provided text.)
Digital transformation enhances firms' total factor productivity (TFP).
Mechanism analysis / mediation analysis reported in the paper using firm-level data (Chinese A-share manufacturing firms, 2011–2024). (Sample size not stated in provided text.)
Digital transformation increases firms' need (demand) for highly educated, high-skilled workers.
Regression analysis on Chinese A-share listed manufacturing firms (2011–2024); analysis of worker composition/skill-demand reported by the authors. (Sample size not stated in provided text.)
Digital transformation significantly increases the quantity of firm labor demand.
Regression analysis using data from Chinese A-share listed manufacturing firms between 2011 and 2024; mechanism and heterogeneity analyses reported in the paper. (Sample size not stated in provided text.)
IDS jointly and incrementally synthesizes implementation and proof, and learns from failed attempts to systematically try promising strategies.
Description of the IDS method and architecture presented in the paper (system design and algorithmic loop).
This paper presents the first effective approach to addressing the gap between LLM coding agents and mechanized formal verification for distributed systems (Inductive Deductive Synthesis, IDS).
Statement of novelty supported by the empirical claim that IDS succeeds on all 7 benchmark specs while prior SOTA agents did not; methodological description of IDS as a joint, incremental synthesis and learning system.
IDS further incorporates performance feedback into the same loop, yielding implementations up to 3x faster than published verified systems.
Empirical benchmarking of IDS-produced implementations against published verified systems, with performance (runtime) comparisons reporting up to a 3x speedup.