Evidence (313 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
21267 claims
Filter claims →
Productivity
17978 claims
Filter claims →
Governance
17038 claims
Filter claims →
Human-AI Collaboration
16914 claims
Filter claims →
Org Design
11104 claims
Filter claims →
Innovation
11087 claims
Filter claims →
Labor Markets
6711 claims
Filter claims →
Skills & Training
5616 claims
Filter claims →
Inequality
4343 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 1880 | 496 | 296 | 1854 | 4721 |
| Organizational Efficiency | 2906 | 665 | 438 | 180 | 4210 |
| Governance & Regulation | 2162 | 929 | 480 | 247 | 3866 |
| Technology Adoption Rate | 1533 | 545 | 278 | 210 | 2593 |
| Decision Quality | 1391 | 534 | 321 | 173 | 2429 |
| Output Quality | 1298 | 472 | 231 | 145 | 2153 |
| AI Safety & Ethics | 682 | 821 | 230 | 90 | 1837 |
| Research Productivity | 855 | 253 | 121 | 425 | 1675 |
| Firm Productivity | 1105 | 171 | 175 | 73 | 1531 |
| Task Allocation | 735 | 229 | 361 | 99 | 1433 |
| Market Structure | 457 | 461 | 251 | 47 | 1222 |
| Innovation Output | 673 | 94 | 108 | 36 | 913 |
| Task Completion Time | 499 | 118 | 43 | 38 | 702 |
| Firm Revenue | 458 | 130 | 61 | 26 | 677 |
| Skill Acquisition | 381 | 122 | 113 | 34 | 650 |
| Consumer Welfare | 316 | 176 | 115 | 39 | 648 |
| Employment Level | 223 | 143 | 177 | 53 | 600 |
| Error Rate | 246 | 282 | 44 | 19 | 594 |
| Fiscal & Macroeconomic | 283 | 142 | 78 | 52 | 562 |
| Inequality Measures | 103 | 329 | 106 | 13 | 552 |
| Worker Satisfaction | 225 | 185 | 63 | 30 | 503 |
| Automation Exposure | 158 | 155 | 72 | 37 | 426 |
| Regulatory Compliance | 186 | 126 | 35 | 14 | 362 |
| Team Performance | 193 | 56 | 51 | 24 | 326 |
| Developer Productivity | 224 | 58 | 27 | 13 | 323 |
| Wages & Compensation | 148 | 108 | 50 | 17 | 323 |
| Training Effectiveness | 218 | 44 | 21 | 27 | 313 |
| Job Displacement | 23 | 159 | 53 | 5 | 240 |
| Hiring & Recruitment | 109 | 61 | 32 | 11 | 215 |
| Skill Obsolescence | 16 | 107 | 26 | 6 | 155 |
| Creative Output | 71 | 44 | 28 | 6 | 150 |
| Social Protection | 58 | 31 | 12 | 3 | 104 |
| Labor Share of Income | 29 | 43 | 25 | 2 | 99 |
| Worker Turnover | 45 | 29 | 6 | 4 | 84 |
| Industry | — | — | — | 1 | 1 |
The most frequently reported frontline-worker concerns about AI were the need for skills training, the need for policy support and labour-market reshuffling, each reported by 22 of 150 respondents (14.70%).
Frequency distribution of eight recurring categories in 150 usable responses to Q19.
When retained inputs are available, teacher relabeling can recover target-base specialists without new gold annotation, although it is not always less computationally expensive.
Refresh-distillation experiments using retained inputs and labels generated by the old specialist, compared with full retraining.
China’s vocational AI curricula only partially match employer demand.
Comparison of 498,331 Chinese job advertisements from 2020–2024 with 46 institutional vocational AI training plans using a demand-weighted coverage index.
The tutor role is the only role currently implemented; the avatar lecturer is in design, the design consultant is in design/pilot, and the meta-agent is a research frontier.
Table 1 explicitly reports the implementation status of all four cumulative roles.
Together, these results point to forces pre-dating generative AI (i.e., labor-market shifts began before ChatGPT) rather than being caused solely by ChatGPT's release.
Synthesis of findings from UI records, LinkedIn profiles, and university syllabi showing changes and cohort gaps beginning in early 2022 and prior to ChatGPT's launch.
Data-centric methods refine training data without altering models, offering cost-effectiveness but struggling with data sparsity and noise.
Conceptual claim in paper motivating the data-centric approach; summary of limitations (data sparsity, noise).
The article develops a conceptual framework linking GenAI use in higher education to knowledge transformation, critical thinking, ethical judgment, digital capability, managerial decision-making, business ethics, workforce readiness, and organizational readiness.
Presentation of a conceptual framework by the authors as part of the review (theoretical/conceptual work; no empirical validation reported).
Given the results, educators should revisit pair programming as an educational tool in addition to embracing modern AI.
Authors' recommendation in the paper's conclusion based on experimental findings (performance, workload, emotion, retention outcomes).
The practical burden of scaling depends on how efficiently real resources are converted into that (logical) compute.
Argument in the paper linking conceptual 'logical compute' to real-world conversion efficiency (qualitative claim; no empirical sample in excerpt).
Participant targeting: 44% of programs targeted doctors and 44% targeted medical students (with possible overlap), and 56% targeted entry‑to‑practice career stages.
Participant audience and career-stage data extracted from the 27 included programs; proportions reported in the review.
Most programs were delivered in academic settings: 56% of evaluated programs reported an academic setting.
Setting information extracted from the 27 included programs, with 56% reported as delivered in academic settings.
A plurality of programs were short in duration: 44% of programs were categorized as short courses.
Extraction of program length from the 27 included studies; 44% were classified as short courses per the review's categorization.
Most programs were introductory in content: 67% of included programs taught introductory AI concepts rather than advanced/technical AI skills.
Program content extraction across the 27 included studies yielded that 67% were classified as teaching introductory AI.
RAD requires estimating cost distributions and choosing a reference policy and quantile-weighting function; these choices determine the method's conservatism and sample efficiency.
Methodological and practical considerations discussed in the paper; noted dependency on estimation and design choices (no quantitative sample-efficiency results provided in the summary).
Evaluation of the equivalency system should use metrics such as concordance between claimed competencies and verified inputs, predictive validity versus labor-market integration outcomes, and false positive/negative rates in automated decisions.
Methodological recommendation in the paper outlining specific evaluation metrics; this is a prescriptive claim (no empirical implementation reported).
In verl's AgentLoop at 7B, 45 of 115 generations contained a complete tool call, but none were accepted, executed, or followed by an observation; thus no tool-mediated trajectory was present in the sampled training experience.
Training-loop instrumentation of 115 generations at 7B; the paper separately reports that the same zero outcome at 1.5B is over-determined.
The cancer baseline could not be substantially rescued by the low-compute adaptation procedure, achieving 0.19 balanced accuracy with the ten-example-per-class probe versus 0.21 with the full probe.
SCIN few-shot adaptation experiment using frozen representations and linear probes.
The paper identifies a sequence-starvation effect in hybrid recommendation rankers, in which strong non-sequential features can shortcut the main discriminative objective and leave the sequence module weakly supervised.
The paper's conceptual diagnosis of hybrid ranker training, supported by its description of sparse ranking labels and the auxiliary-supervision ablation.
Regular organizational AI training and transparent algorithm-governance policies weaken the relationship between AI adoption and employees' job insecurity, thereby reducing adverse psychological effects.
A total of 56 reviewed studies introduced organizational support as a moderating variable, focusing primarily on AI training and transparent algorithm governance.
The authors conclude that retraining alone is unlikely to be sufficient if the displacement implied by Waves 2 and 3 occurs at the projected scale and speed.
Policy interpretation of the projected automation shares, especially the possibility that many workers could be displaced too rapidly to move into new roles while AI capabilities continue to advance.
High turnover leads hotel operators to underinvest in the analytical capabilities needed to make machine-learning scheduling systems effective, even when the systems are deployed.
Theoretical argument applying Becker's Human Capital Theory to hotel scheduling; the claim is not tested with original data.
A majority of surveyed Indian law students, 71.1%, reported receiving no formal training on the ethical use of AI.
Descriptive results from a structured survey of 380 undergraduate law students recruited using purposive convenience sampling.
Soft skills such as teamwork and communication are often taught implicitly rather than embedded as explicit, assessed curriculum components.
Content analysis of 46 institutional vocational AI training plans using a 198-keyword bilingual taxonomy covering hard and soft skills.
The small expert panel may constrain the generalizability of the framework, which reflects expert judgment and weighting rather than experimentally measured program impacts.
The study used a panel of 10 experts and did not experimentally evaluate outcomes such as skills, employability, earnings, or innovation.
Only 5 of 15 participants reported that their organizations provided adequate support during restructuring transitions; the other 10 described organizational support as reactive rather than proactive.
Participant reports from semi-structured interviews, summarized in the thematic analysis and Table 2.
Many micro-credentials lack transparent standards for assessment validity and comparability, creating concerns about assessment rigour, proctoring, and fraud.
Review synthesis of studies addressing quality assurance and assessment practices; no pooled fraud or validity estimate is reported.
A pervasive shortage of labeled fraud instances constrains supervised-learning development, model evaluation, and comparability across studies.
Cross-study synthesis identifying limited labeled datasets as a recurring methodological limitation, alongside heterogeneous fraud definitions and evaluation protocols.
Training needs and organizational resistance are substantial barriers to AI adoption for sustainable decision-making.
Thirty-four of the 119 reviewed articles discussed training and organizational resistance as adoption barriers.
The FinanceGym question pool was reduced from 29,669 unconstrained generations to 2,078 quality-filtered questions before expert review.
A multi-stage LLM quality-filtering pipeline assessing feasibility, institutional relevance, coherence, groundedness, and balance.
The agents did not effectively respond to negative research feedback: despite dozens of AI-review rounds identifying problems later raised by human reviewers, neither agent received an acceptance and the agents mainly added caveats while continuing unpromising research directions.
The agents submitted their papers to subagents and external AI review tools across dozens of revision rounds; the researchers compared the resulting revisions with later human expert reviews.
AI agents precipitate a pedagogical crisis for training researchers (pedagogical crisis).
Conceptual argument in the paper that current teaching and apprenticeship models may be disrupted by agents; no empirical study of educational outcomes presented.
Parallel rollouts (increasing width) provide only logarithmic relief: the effective number of independent samples is capped by correlation, so width gives only a logarithmic improvement in sample complexity.
Mathematical analysis/proof in the paper (Width Limits result) showing how correlation among parallel rollouts limits the benefit from increasing width; theoretical derivation rather than empirical measurement.
The signal connecting early steps to final outcomes decays exponentially with depth, creating a critical horizon beyond which reliable learning from endpoint data alone requires exponentially many samples.
Analytical proof within an information-theoretic model of sequential processes; the paper states a formal bound (Signal Decay Bound) derived from assumptions about signal attenuation across intervening steps. No empirical sample size—this is a theoretical result based on the model's assumptions.
Training large language models (LLMs) is computationally expensive, partly because the loss exhibits slow power-law convergence whose origin remains debatable.
Empirical observation cited by the paper and general background motivation; the paper states this as an observed phenomenon motivating the study (no sample size or specific experiments reported in the abstract).
Traditional approaches to digital or AI literacy are insufficient for the challenges posed by generative AI.
Argumentative claim in the paper asserting insufficiency of existing literacy approaches; appears to be based on conceptual analysis rather than an empirical evaluation (no sample size provided).
Sensemaking (the ability to co-construct causal explanations, surface uncertainties, and adapt goals) is the key capability that current training pipelines do not explicitly develop or evaluate.
Authoritative claim defining 'sensemaking' as essential and asserting a gap in current training/evaluation practices; no empirical evaluation or dataset cited in the excerpt.
Many formal NAACLS-accredited cytogenetics programs have collapsed, contributing to reduced pipeline capacity for cytogenetics specialists.
Cited counts/observations of NAACLS program closures and professional white papers discussed in the review.
Diagnostic heuristic: if letting AI in makes the task feel effortless, it is in the wrong place.
Authors' heuristic for educators (conceptual guidance; no empirical test reported in the excerpt).
The architecture of the undergraduate degree is structurally incapable of replacing the informal post-degree apprenticeship system through curricular revision alone.
Argument presented in the paper, supported by the systematic review of eighteen peer-reviewed studies and labor-market analyses cited in the abstract.
Higher education has misdiagnosed the resulting challenge as curriculum misalignment—a content problem assumed to be solvable through revised syllabi, AI electives, and marginal expansions of experiential learning.
Argument presented in the paper, supported by the paper's systematic review of eighteen peer-reviewed studies and labor-market analyses (as described in the abstract).
Existing LLM4Rec paradigms are bottlenecked by the difficulty of measuring and improving chain-of-thought (CoT) quality in open-domain recommendation during supervised fine-tuning (SFT).
Author assertion about limitations of prior LLM4Rec paradigms (literature/diagnosis in the paper).
Existing AI education, AI literacy, and human-AI collaboration frameworks remain centred on prompting, task execution, and productivity support and are poorly equipped to address this tacit layer of expert cognition.
Argumentative critique in the paper drawing on conceptual analysis and review of prevailing frameworks; no empirical evaluation or sample reported.
AI adoption presents workforce adaptation challenges.
Reported in the study's literature synthesis and thematic analysis of secondary sources (qualitative review). No sample size reported.
Process-based supervision introduces challenges regarding the sustainability of human-in-the-loop feedback loops.
Socio-technical argumentation in the paper—concern raised about ongoing human verification burden; no longitudinal or empirical data on human labor sustainability provided.
This directional skew is not eliminated by one-shot in-context prompting.
Intervention of one-shot in-context prompting applied to models; evaluation shows the intervention-oriented error skew persists despite one-shot prompting.
The policy and research challenge posed by platform-mediated automation is not merely job quantity (technological unemployment) but institutional continuity — how societies reproduce practical competence when platforms optimize for efficiency rather than formation.
Normative and conceptual claim developed through literature synthesis (institutional economics, platform governance, workforce development); presented as an analytical reframing rather than an empirically tested hypothesis.
Limited reskilling coverage constrains workers' ability to adapt to AI-driven changes.
Paper reviews official reports and secondary data (2020–2024) indicating low coverage/uptake of reskilling programs in India and links this to limited adaptation capacity.
Across heterogeneous learners, a common broadcast curriculum can be slower than personalized instruction by a factor linear in the number of learner types.
Theoretical comparative result in the model (analysis of broadcast vs personalized curricula across heterogeneous learner types; abstract states factor linear in number of types).
No evaluated program reported Kirkpatrick‑Barr level‑4 outcomes (organizational change, patient outcomes, or sustained metacognitive mastery).
Reviewers mapped reported outcomes from all 27 included programs and found none that demonstrated organizational-level impacts or patient‑level outcomes (level 4).
Implementing this framework requires significant resources and continuous updating.
Stated explicitly under Main Finding and Disadvantages/Risks; paper lists cost/time metrics to track (cost-per-curriculum, time-to-update) and highlights resource intensity. Support is descriptive/analytic rather than empirical.