Evidence (3308 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
9875 claims
Filter claims →
Productivity
8807 claims
Filter claims →
Governance
7870 claims
Filter claims →
Human-AI Collaboration
7560 claims
Filter claims →
Org Design
4892 claims
Filter claims →
Innovation
4781 claims
Filter claims →
Labor Markets
4004 claims
Filter claims →
Skills & Training
3308 claims
Filtered →
Inequality
2332 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 870 | 233 | 116 | 1066 | 2363 |
| Governance & Regulation | 976 | 451 | 218 | 133 | 1809 |
| Organizational Efficiency | 949 | 224 | 144 | 88 | 1416 |
| Technology Adoption Rate | 764 | 287 | 141 | 122 | 1325 |
| Research Productivity | 501 | 152 | 74 | 362 | 1101 |
| Output Quality | 542 | 216 | 69 | 69 | 896 |
| Decision Quality | 387 | 198 | 94 | 54 | 740 |
| Firm Productivity | 513 | 67 | 101 | 27 | 714 |
| AI Safety & Ethics | 249 | 303 | 73 | 36 | 667 |
| Market Structure | 190 | 192 | 134 | 27 | 548 |
| Task Allocation | 243 | 77 | 91 | 36 | 452 |
| Innovation Output | 291 | 33 | 55 | 20 | 401 |
| Skill Acquisition | 206 | 72 | 65 | 21 | 364 |
| Employment Level | 133 | 63 | 115 | 22 | 335 |
| Fiscal & Macroeconomic | 153 | 79 | 52 | 32 | 323 |
| Task Completion Time | 206 | 37 | 12 | 15 | 272 |
| Firm Revenue | 179 | 52 | 29 | 5 | 266 |
| Consumer Welfare | 130 | 76 | 47 | 13 | 266 |
| Inequality Measures | 48 | 137 | 51 | 6 | 242 |
| Worker Satisfaction | 101 | 81 | 25 | 13 | 220 |
| Error Rate | 84 | 110 | 11 | 5 | 210 |
| Wages & Compensation | 98 | 47 | 30 | 10 | 185 |
| Regulatory Compliance | 88 | 73 | 17 | 7 | 185 |
| Automation Exposure | 66 | 64 | 33 | 16 | 182 |
| Team Performance | 105 | 29 | 30 | 11 | 176 |
| Training Effectiveness | 109 | 22 | 14 | 21 | 168 |
| Developer Productivity | 114 | 21 | 14 | 8 | 158 |
| Job Displacement | 12 | 90 | 24 | 1 | 127 |
| Hiring & Recruitment | 57 | 9 | 9 | 5 | 80 |
| Skill Obsolescence | 6 | 56 | 9 | 1 | 72 |
| Social Protection | 43 | 17 | 8 | 2 | 70 |
| Creative Output | 35 | 21 | 9 | 4 | 70 |
| Labor Share of Income | 18 | 21 | 17 | 1 | 57 |
| Worker Turnover | 15 | 16 | — | 4 | 35 |
| Industry | — | — | — | 1 | 1 |
Skills Training
Remove filter
Participants performed significantly better with GitHub Copilot than with their human teammate.
Experimental comparison of task performance between Copilot-assisted individual condition and human pair condition; statistical significance reported in results (sample size n=22).
Policymakers can reinforce these conditions by shifting from technology-neutral principles to auditable process standards that couple AI investment with reskilling and data-quality obligations.
Policy recommendation based on the study's findings and synthesis; presented as a normative implication rather than empirically tested within the study. (Sample size not reported.)
Leaders should fund training coverage and design (not just headline hours), equip non-specialists to interpret model outputs, pair performance artefacts with participatory routines, and treat explainability as a usability requirement to achieve durable, auditable value in safety-critical energy contexts.
Prescriptive recommendation based on a 'field-tested playbook' synthesised from the multi-case qualitative study (interviews, surveys, documents). The claim is drawn from authors' interpretation of cross-case patterns rather than causal inference. (Sample size not reported.)
Structured upskilling and precise recourse mechanisms are associated with higher confidence, productivity, and clearer sustainability pathways.
Observed association in multi-case qualitative data: interviews, staff/manager surveys, and policy documents; triangulated through thematic coding and cross-case synthesis. (Sample size not reported.)
A tight workflow fit that minimises cognitive overhead at the decision point accelerates legitimate use and strengthens links to emissions monitoring and predictive-maintenance outcomes.
Synthesised from interviews, Likert-scale surveys of technical staff and managers, and internal workflow/policy documents across multiple cases in the energy sector. (Sample size not reported.)
Communicative governance — e.g. model cards, bias tests, validation reports, and explicit appeal rights — earns trust, curbs shadow workarounds, and improves safety culture.
Reported from thematic coding of interviews, surveys of staff and managers, and documentary evidence across multiple cases; triangulation claimed. (Sample size not reported.)
Broad-based capability building beyond specialist teams prevents benefits from concentrating in expert enclaves and reduces brittle scale.
Derived from cross-case thematic synthesis of interviews, Likert surveys of mid-level managers and technical staff, and internal policy/strategy document analysis (multi-case qualitative evidence). (Sample size not reported.)
Three reinforcing levers shape adoption outcomes: (1) broad-based capability building beyond specialist teams, (2) communicative governance that couples transparency with contestability, and (3) a tight workflow fit that minimises cognitive overhead at the decision point.
Qualitative, multi-case design triangulating a semi-structured interview with a senior manager, Likert-scale surveys of mid-level managers and technical staff, and analysis of internal policies and strategy documents; thematic coding with intercoder reliability and cross-case synthesis. (Sample size not reported.)
From synthesis of results, we suggest three practices that focus on preserving agency in software engineering for coding, learning, and mentorship, especially as AI grows increasingly autonomous.
Authors' prescriptive recommendations derived from the paper's qualitative synthesis; presented as proposed practices rather than empirically tested interventions.
Seniors leverage pre-AI foundational instincts to steer modern tools and possess valuable perspectives for mentoring juniors in their early AI-encouraged career development.
Qualitative accounts from senior participants in the Delphi/ACTA process and blind reviews showing seniors reference pre-AI practices and see mentoring value.
Juniors enter as AI‑natives, seniors adapted mid‑career.
Authors' synthesis from a three-phase mixed-methods study: ACTA combined with a Delphi process (5 seniors), an AI-assisted debugging task (10 juniors), and blind reviews of junior prompt histories by 5 additional seniors.
Policy proposals including universal basic income, portable benefits, retraining programs, and AI taxation are viable mechanisms to manage the socio-economic transition associated with AI, and the paper assesses these proposals.
Paper states it evaluates these policy proposals drawing on empirical studies, reports, and historical analysis; the abstract does not report empirical tests or effectiveness estimates for these policies.
The distributional consequences of AI adoption will be shaped primarily by institutional factors—including labor market regulation, education policy, and corporate governance structures—rather than by the technology itself.
Argument based on a literature review drawing on recent empirical studies, industry reports, and historical analyses of past technological transitions; no new empirical estimate or sample size provided in the abstract.
AI differs from previous automation technologies in its capacity to perform cognitive and creative tasks.
Paper's conceptual claim supported by references to recent empirical studies and industry reports on generative AI and large language models; no specific sample size or quantified effect reported in the abstract.
Our paper contributes to the emerging discourse on AI overreliance and provides an understanding of the appropriate degree of reliance as essential to developers making the most of these powerful technologies.
Authors' claimed contribution based on synthesis of themes from twenty-two interviews and presentation of the reliance-control framework.
The reliance-control framework can be used to recommend future research to explore different control levels supported by current and emergent LLM-driven tools.
Paper explicitly uses the framework to motivate and recommend directions for future research; based on qualitative interview findings (n=22) and authors' synthesis.
We propose a preliminary reliance-control framework where the level of control can be used to identify AI overreliance and underreliance.
Authors present a conceptual/framework contribution derived from analysis of the twenty-two interviews; this is a proposed (theoretical) framework rather than an experimentally validated one.
The model's contribution lies in integrating four interdependent governance layers—technical, organizational, workforce, and regulatory—within a single labor-market framework.
Paper's stated conceptual contribution describing the four-layer governance model derived from the evidence map and synthesis.
Based on an evidence map of the included studies, we propose a hybrid governance model combining technical and organizational audits, inclusive upskilling/reskilling, participatory regulation, and responsible HR policies to align AI innovation with decent and inclusive work.
Conceptual proposal grounded in the paper's evidence map and qualitative synthesis of the 19 studies; model components explicitly listed in the text.
The evidence indicates that AI can support inclusion through assistive technologies and improved matching in labor-market settings.
Synthesis claim based on thematic analysis of the 19 included peer-reviewed studies (qualitative evidence across the corpus pointing to assistive technologies and improved matching as inclusion-supporting mechanisms).
The positive effect of supply chain digitalization on human capital structure is stronger for enterprises located in the eastern region of China.
Heterogeneity analysis in the paper using the DID framework on A-share listed companies (2013–2022); regional subsample analysis shows a larger effect in eastern China.
The positive effect of supply chain digitalization on human capital structure is stronger for enterprises operating in more competitive industries.
Heterogeneity analysis reported in the paper using DID on A-share listed firms (2013–2022); industry competition intensity is used to split sample and examine differential effects.
The positive effect of supply chain digitalization on optimizing human capital structure is stronger for enterprises facing higher external environmental uncertainty.
Heterogeneity analysis in the paper using the DID sample of A-share listed firms (2013–2022); authors report the effect is more pronounced under higher environmental uncertainty.
Supply chain digitalization enhances enterprises' capacity to absorb high-skilled labor by promoting the accumulation of digital intangible assets.
Mechanism analysis in the paper using DID on A-share listed companies (2013–2022); accumulation of digital intangible assets is cited as a channel increasing firms' demand/ability to hire high-skilled workers.
Supply chain digitalization enhances enterprises' capacity to absorb high-skilled labor by alleviating financing constraints.
Mechanism analysis reported in the paper using the quasi-natural experiment and DID approach on A-share listed firms; easing financing constraints is presented as one channel.
Supply chain digitalization enhances enterprises' capacity to absorb high-skilled labor by boosting public trust in brands.
Mechanism analysis in the paper using the DID design on A-share listed firms (2013–2022); brand/public trust is reported as a mediating channel.
Supply chain digitalization enhances enterprises' capacity to absorb high-skilled labor by increasing firms' market attention.
Mechanism analysis reported in the paper using the same DID framework and sample (A-share listed firms 2013–2022); market attention is listed as an identified channel through which digitalization affects human capital.
Supply chain digitalization drives the optimization of the human capital structure of enterprises.
Empirical analysis on A-share listed companies on the Shanghai and Shenzhen Stock Exchanges from 2013 to 2022; authors treat pilots of supply chain innovation and application as a quasi-natural experiment and employ a difference-in-differences (DID) approach to identify the effect.
Successful implementation requires tailored strategies that address contextual, technical, and human factors.
Authors' synthesis and recommendations based on patterns and barriers identified in the included studies.
Client technological readiness plays a positive role in remote auditing.
Reported moderating/mediating findings across included studies summarized in the review indicating client readiness supports remote audit processes.
Technologies such as big data analytics, artificial intelligence, and federated learning have a transformative impact on audit quality and efficiency.
Synthesis of findings from the 10 included empirical studies reporting effects of these technologies on auditing outcomes.
The approach provides a practical path toward more transparent, controllable, and accountable AI use without requiring new model architectures.
Authors' asserted benefit of the proposed interaction-layer framework; no empirical demonstration that transparency, control, or accountability are achieved or that no architectural changes are required in practice.
The framework enables auditable reasoning traces and supports alignment with emerging governance standards, including the EU AI Act and ISO/IEC 42001.
Stated compliance/alignment claim linking the proposed interaction-layer approach to existing regulatory standards; no compliance testing or audit examples reported.
This reframes the question from whether the model can think to whether the human-AI system can reason.
Conceptual reframing stated in the paper; no empirical evidence required as it is a change of perspective.
We introduce 'The Architect's Pen' as a practical method where the human uses the model as an external medium for structured reflection by embedding phases of articulation, critique, and revision into human-AI interaction.
Method description / practical proposal included in the paper; no experimental evaluation, user study, or quantitative validation reported.
This perspective emphasizes collaborative intelligence, combining human judgment and contextual understanding with machine speed, memory, and associative capacity.
Theoretical claim about complementary strengths of humans and models within the proposed framework; presented without empirical tests.
Building on recent work on 'System-2' learning, reflective reasoning can be relocated to the interaction layer and framed as a cognitive protocol that can be structured, measured, and governed using existing systems.
Conceptual extension of prior literature ('System-2' learning) into an interaction-layer protocol; no empirical protocol testing or measurement evidence provided.
Reasoning should be treated as a relational process distributed between human and model rather than an internal capability of either.
Methodological proposal / theoretical framing presented by the authors; no empirical validation reported.
Large language models have advanced rapidly, from pattern recognition to emerging forms of reasoning.
Stated as an observational claim in the paper's introduction; no empirical evaluation or dataset provided.
For settings with multiple interventions, a tractable approximation that prioritizes interventions based on the magnitude of the policy-value discrepancy is effective.
Proposed algorithm/approximation in the paper (methodological contribution); evaluated empirically in simulations and experiments described in the paper.
In the single-intervention regime, the optimal strategy is to recommend the action that maximizes the human value function.
Theoretical result derived in the paper within a Markov decision process model for single-intervention settings.
Policy-value inconsistencies naturally identify opportunities for intervention.
Analytical/formal argument within a Markov decision process framework showing that when human policy-value consistency fails, discrepancies indicate intervention opportunities.
The paper proposes a conceptual framework of the underlying mechanisms of the LLM fallacy and a typology of its manifestations across computational, linguistic, analytical, and creative domains.
Author(s) contribution described in the paper (framework and typology); no empirical testing reported in the abstract.
The rapid integration of large language models (LLMs) into everyday workflows has transformed how individuals perform cognitive tasks such as writing, programming, analysis, and multilingual communication.
Author(s) assertion based on literature review and conceptual overview; no empirical sample or experiment reported in the abstract.
Linking these measures to administrative data from 2012 to 2023 shows a broad shift from manual and digital toward frontier skills across occupations.
Longitudinal analysis linking OTSS to administrative labor market data covering 2012–2023, showing temporal changes in skill composition toward frontier skills.
We compute OTSS for all occupations in the German labour market.
Paper reports application of the OTSS metric across the set of occupations covering the German labour market.
Using natural language processing, generative AI and supervised machine learning, we develop an AI‐powered skill classification that enriches occupation‐linked skill labels with standardised GenAI‐generated descriptions and structured indicators of technological content, enabling transparent classification by technology intensity.
Paper describes methodological approach combining NLP, generative AI and supervised ML to create the skill classification and enriched labels.
This paper introduces a novel skill‐based measure of occupational technology intensity – the occupational technology skill share (OTSS) – that distinguishes between manual, digital and frontier technologies, including artificial intelligence (AI).
Paper statement of contribution / methodological development (description of new measure OTSS).
Successful AI implementation in auditing requires an integrated framework that aligns technological readiness, auditor acceptance, and innovation diffusion to sustainably improve audit quality in Indonesia.
Authors' conclusion and recommendation derived from thematic synthesis of reviewed literature and comparative findings.
Comparative analysis indicates Indonesia remains at the early majority stage of AI adoption in auditing.
Authors' comparative synthesis of the reviewed literature and country-specific discussion classifying Indonesia's adoption stage as early majority.