Evidence (155 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
21267 claims
Filter claims →
Productivity
17978 claims
Filter claims →
Governance
17038 claims
Filter claims →
Human-AI Collaboration
16914 claims
Filter claims →
Org Design
11104 claims
Filter claims →
Innovation
11087 claims
Filter claims →
Labor Markets
6711 claims
Filter claims →
Skills & Training
5616 claims
Filter claims →
Inequality
4343 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 1880 | 496 | 296 | 1854 | 4721 |
| Organizational Efficiency | 2906 | 665 | 438 | 180 | 4210 |
| Governance & Regulation | 2162 | 929 | 480 | 247 | 3866 |
| Technology Adoption Rate | 1533 | 545 | 278 | 210 | 2593 |
| Decision Quality | 1391 | 534 | 321 | 173 | 2429 |
| Output Quality | 1298 | 472 | 231 | 145 | 2153 |
| AI Safety & Ethics | 682 | 821 | 230 | 90 | 1837 |
| Research Productivity | 855 | 253 | 121 | 425 | 1675 |
| Firm Productivity | 1105 | 171 | 175 | 73 | 1531 |
| Task Allocation | 735 | 229 | 361 | 99 | 1433 |
| Market Structure | 457 | 461 | 251 | 47 | 1222 |
| Innovation Output | 673 | 94 | 108 | 36 | 913 |
| Task Completion Time | 499 | 118 | 43 | 38 | 702 |
| Firm Revenue | 458 | 130 | 61 | 26 | 677 |
| Skill Acquisition | 381 | 122 | 113 | 34 | 650 |
| Consumer Welfare | 316 | 176 | 115 | 39 | 648 |
| Employment Level | 223 | 143 | 177 | 53 | 600 |
| Error Rate | 246 | 282 | 44 | 19 | 594 |
| Fiscal & Macroeconomic | 283 | 142 | 78 | 52 | 562 |
| Inequality Measures | 103 | 329 | 106 | 13 | 552 |
| Worker Satisfaction | 225 | 185 | 63 | 30 | 503 |
| Automation Exposure | 158 | 155 | 72 | 37 | 426 |
| Regulatory Compliance | 186 | 126 | 35 | 14 | 362 |
| Team Performance | 193 | 56 | 51 | 24 | 326 |
| Developer Productivity | 224 | 58 | 27 | 13 | 323 |
| Wages & Compensation | 148 | 108 | 50 | 17 | 323 |
| Training Effectiveness | 218 | 44 | 21 | 27 | 313 |
| Job Displacement | 23 | 159 | 53 | 5 | 240 |
| Hiring & Recruitment | 109 | 61 | 32 | 11 | 215 |
| Skill Obsolescence | 16 | 107 | 26 | 6 | 155 |
| Creative Output | 71 | 44 | 28 | 6 | 150 |
| Social Protection | 58 | 31 | 12 | 3 | 104 |
| Labor Share of Income | 29 | 43 | 25 | 2 | 99 |
| Worker Turnover | 45 | 29 | 6 | 4 | 84 |
| Industry | — | — | — | 1 | 1 |
The predicted direction and magnitude of cognitive-competence change are not determined by the bifurcation itself; the paper's decline result depends on an illustrative assumption that substitutive use lowers competence.
The authors explicitly state that Γu = 1, Γc = 0.5, and Γd = 0.1 are illustrative modeling assumptions and that scaffolded or augmentative use could leave competence unchanged or increase it.
Human-AI interaction creates feedback loops in which human inputs shape models, models alter the value of human skills, and organizational routines evolve.
Thematic analysis of qualitative interviews combined with conceptual theory construction.
Coding was the only reported domain showing mild forgetting for Thomson-1.0-Large relative to its base model, while general mathematical and abstract reasoning remained within the Qwen performance level.
Authors' interpretation of cross-domain results, including coding scores of 39.9% for Thomson-1.0-Large versus 40.9% for Qwen3.5-397B and reasoning scores of 68.4% versus 66.8%.
Over four generations, the retrained intent-classification reference score changes by only 0.2 points while the zero-shot score rises by 21 points.
Evaluation of retrained references and zero-shot adoption scores across four Qwen generations.
The value of labor is expected to shift from time-efficiency skills toward attention-management and coordination skills, with possible polarization if attention-intensive roles become concentrated in a small number of firms.
Conceptual labor-market implication derived from the paper’s attention, asynchronous-work, and AI-complementarity frameworks.
Employers and platforms should weigh short-term productivity gains against possible erosion of worker skills and long-run employability.
Normative labor and human-capital implication derived from the proposed capability-dependency framework.
Employees with low critical-evaluation ability who frequently use AI may experience higher short-term productivity but are likely to become overly reliant on AI output and fall behind without intervention.
Conceptual employee typology proposed by the authors; no longitudinal evidence or productivity measurement is provided.
GenAI may increase the value of expertise in critical evaluation, domain knowledge, oversight, validation, and refinement, while routine or initial theorizing tasks may be delegated and reduce demand for some junior interpretive labor.
Conceptual labor-market implication concerning complementarity between AI tools and supervisory expertise, alongside a stated deskilling and substitution risk; no wage or hiring data are reported.
It also drives a shift from skill-biased to measurability-biased technical change.
Theoretical claim in the paper arguing that incentives and returns will favor tasks that are easy to measure/verify; presented as a consequence of the model rather than demonstrated with empirical data.
Accounting must adapt to rapidly changing technological environments.
Assertion in the review linking historical challenges to the contemporary need for technological adaptation (narrative literature review, no empirical sample reported).
Artificial intelligence differs from human intelligence on dimensions including learning, consciousness, embodiment, and ethical reasoning.
Conceptual comparison and argumentation in the chapter (philosophical and cognitive-science framing; references to expert views), not an empirically quantified comparison.
Overqualification can serve as an early warning indicator of unequal adaptation, discrimination risks, and insufficient recognition of human competencies in artificial intelligence-intensive labor markets.
Interpretation based on the study's observed positive associations between AI vibrancy (and some IHCAI components) and national overqualification rates in the panel analysis for 18 European countries (2017–2024).
Workers' beliefs about the future labor market impact of emerging digital technologies (AI/LLMs) vary substantially within and across professions (law and management consulting).
Survey (and associated randomized experiment) administered to workers in two professional fields; findings reported as part of pilot/preliminary data (sample size not provided in chapter summary).
AI integration produces both job obsolescence risks and opportunities for upskilling and task augmentation in sectors like manufacturing, healthcare, and logistics.
Drawn from the paper's sectoral case studies and empirical labor-market evidence (specific sample sizes and effect estimates not provided in the abstract).
The vulnerability of highly skilled workers to AI substitution depends on a skill threshold that is shaped by organizational depth, baseline costs, and risk differentials.
Theoretical derivation in the HAT model establishing a skill threshold concept and showing how model parameters (organizational depth, costs, risk differentials) shift that threshold; no empirical evidence provided.
The collapse threshold depends jointly on the competence a user brings to a task and on the tool's transparency, the fraction of its working a user can reconstruct.
Model parameter analysis in paper identifying dependence of bifurcation/threshold on initial competence and a parameter representing tool transparency.
Early field evidence is consistent with such costs, though the largest meta-analytic evidence on prior technologies points the other way, and the question of whether generative AI differs is open.
Mixed evidence: (a) unspecified early field studies that the paper says are consistent with offloading costs, (b) reference to a 'largest meta-analytic evidence' on prior technologies showing the opposite effect; no sample sizes, study names, or effect estimates provided in the excerpt.
The effect of AI development on firms' labor educational structure is substantially larger in high-technology industries: the effect in high-technology industries is approximately 2.5 times as large as that in non-high-technology industries.
Industry heterogeneity analysis reported in the paper comparing coefficients for high-technology vs. non-high-technology industry subsamples using firm-level data (Chinese A-share firms, 2014–2024); reported ratio ≈ 2.5.
The substitution (for low-educated labor) and complementarity (with high-educated labor) effects of AI on firms' labor educational structure exhibit significant regional heterogeneity: the substitution effect is stronger in developed regions, while the complementarity effect is more pronounced in less developed regions.
Subgroup/heterogeneity analysis across regions using the firm-level panel (Chinese A-share firms, 2014–2024); reported differences in coefficients by regional development level.
Firms' technological innovation capability significantly mediates the effect of AI development on labor educational structure: by enhancing technological innovation capability, AI reduces demand for low-educated labor and increases demand for high-educated labor.
Mediation/causal pathway analysis reported in the study using firm-level data and mediation regressions on Chinese A-share listed firms (2014–2024); the paper reports that technological innovation capability is a significant mediating variable linking AI development to changes in labor education composition.
AI is changing skill requirements—some skills become obsolete and new skills are required.
Paper identifies changing skill requirements as a key area of examination (abstract). This is stated as an asserted trend based on the paper's review rather than a quantified empirical finding in the provided text.
The CAW result generalizes through CES aggregation and, when tasks are separated into substitutable versus complementary, yields a directional inversion of skill-biased technical change.
Theoretical extension of the core model using CES (constant elasticity of substitution) aggregation and task decomposition in the paper; the claim arises from model generalization and comparative-static reasoning. No empirical validation provided in the excerpt.
Routine automation primarily dismantles specialised physical skills, enhancing mobility only within homogeneous manual clusters.
Simulation results distinguishing effects of the routine-task automation exposure measure vs. AI exposure; analysis of which skill types are eroded and resulting changes in mobility within occupational clusters.
AI plays a dual role as enhancer and eroder, simultaneously strengthening performance while eroding underlying expertise (the 'AI-as-Amplifier Paradox').
Framing claim presented in the paper's conceptual argument and grounded by the paper's stated year-long empirical study among cancer specialists (no numerical sample size reported in abstract).
These productivity gains are most pronounced for lower-skilled workers, producing a pattern the authors call “skill compression.”
Cross-study pattern reported in the literature review: comparative evidence across worker-skill strata in multiple empirical papers showing larger relative gains for lower-skilled/junior workers; specific underlying studies and sample sizes are not enumerated in the brief.
Senior HR professionals reported a gap between the competencies required for human–AI collaboration and existing HR skills and roles.
Semi-structured interviews with senior HR professionals and thematic analysis in the qualitative phase.
Under the paper's illustrative substitutive-use assumptions, increasing LLM transmission pressure causes an abrupt loss of average human cognitive competence when the autonomous attractor loses stability.
Model assigns Γu = 1, Γc = 0.5, and Γd = 0.1 to uncoupled, autonomous-coupled, and dependent users, respectively, then derives the equilibrium average competence and its bifurcation.
Beyond the tipping point, the model predicts a rapid increase in both regular LLM use and persistent cognitive dependency, accompanied by a sharp decrease in the uncoupled population fraction.
Equilibrium bifurcation curves for U*, C*, and D* derived from the population model.
An observational study cited by the paper found that experienced endoscopists who regularly used AI-assisted polyp detection had their unassisted adenoma detection rates fall from 28% to 22% over three months.
Secondary report of an observational study attributed to Budzyń et al. (2025); the paper does not state the number of endoscopists.
The paper characterizes AI dependency as involving invisible damage to cognitive autonomy, irreversible lock-in, and no discrete moment of crisis.
Conceptual and structural comparison between the proposed game and the classical Tragedy of the Commons; this is not supported by an empirical sample.
AI-assisted development can create skill-formation risks: using AI beyond one's current understanding may crowd out learning, while relying heavily on AI for already-mastered tasks may erode hands-on skills.
Research on skill formation and reliance on AI-assisted programming, synthesized in the report.
Workers who spend years performing low-skill repair and validation may lose existing craft skills and fail to build new expertise, leaving them less professionally marketable when end-to-end AI agents become widespread.
This is a proposed downstream mechanism in the position paper, based on the author's synthesis of the producing-to-validating shift rather than a reported longitudinal test.
The World Economic Forum projects that 39% of core workforce skills will change by 2030.
The paper cites the World Economic Forum's Future of Jobs Report 2025.
On documented OLMo continuation edges, adapter retention falls from 0.88–0.99 after a 46B-token continuation to zero after a 2.9T-token continuation.
Prospectively specified OLMo distance ablation using five documented checkpoint edges, including short and long pure-continuation targets.
For text-to-SQL, the frozen specialist can forfeit up to 59% of the attainable gain on a single upgrade hop.
Comparison of frozen specialists with retrained references on the text-to-SQL task across upgrade hops.
Automation can erode tacit knowledge by reducing skills and judgment when tasks are codified.
The paper identifies erosion of tacit knowledge under automation as a recurrent sociotechnical dynamic and failure mode.
People using LLM tools show weaker engagement in occipito-parietal and prefrontal brain regions than people using search-and-exploration methods.
The paper reports neuroscience experiments comparing LLM-tool users with users of search-and-exploration methods.
Daily LLM usage is associated with a reduction in independent thinking.
The paper cites a study of 670 participants examining daily LLM use and independent thinking.
When artists are excluded from AI deployment, previously skilled tasks are more likely to be automated or substituted, while responsibilities become fragmented and work pace and pressure increase.
The paper's contingent model and discussion of exclusionary deployments identify skill substitution, degraded skills, fragmented responsibilities, and intensified work as consequences of worker exclusion.
By 2030, approximately 39% of core worker skills may transform or become obsolete as a result of technological integration and demographic shifts.
Projection attributed to the World Economic Forum's Future of Jobs Report (2025); this paper does not describe the forecasting model or provide a sample size.
An estimated 56% of graduates rate themselves as deficient in job-specific skills.
Self-assessment statistic attributed to the Cengage Group Report (2025); the paper does not report the report's sample size or methodology.
A learning trap can emerge when easy-task learning is relatively efficient or when the distribution of user types is highly right-skewed.
Theoretical monopoly analysis and parametric examples. Under these conditions, improvements in easy-task capability can increase AI's comparative advantage in easy tasks and lead the provider to restrict access to users with few complex tasks.
Insufficient AI literacy is associated with over-reliance on generative AI, which can undermine critical thinking and long-term cognitive development.
Narrative synthesis of educational literature discussing over-reliance on generative algorithms and so-called e-cheating; no pooled effect estimate is reported.
Augmentation is not inherently safe because overreliance on AI agents can gradually erode workers’ skills and oversight capabilities.
The claim is based on the framework’s prospective risk scenarios and their taxonomy-level analysis; the excerpt does not report a causal experiment or longitudinal worker study.
Erroneous Agent Actions and Human Capability Erosion together represented 51.9% of all risk scenarios and occurred mostly under augmentation.
Combined descriptive analysis of the taxonomy categories and deployment modes across the anticipated risk-scenario corpus.
Human Capability Erosion accounted for 21.3% of all risk scenarios, making it the second-largest taxonomy category.
Descriptive analysis of the 8,356 anticipated risk scenarios, including risks involving loss of workers’ skills and oversight capabilities.
The quality-of-life benefits of AI may weaken or reverse when cognitive offloading, opacity, and repeated substitution produce dependency or skill erosion.
Conceptual Proposition 1, supported by the paper's synthesis of literature on automation, deskilling, and capacity-hostile environments; no new causal estimate is reported.
Twelve of 15 participants expressed significant concern that AI and automation could undermine their professional relevance, with concerns focused on skill obsolescence, rapid digital upskilling requirements, and uncertainty about which roles would remain.
Qualitative interviews and reflexive thematic analysis; 12 participants reported significant AI-related concern.
In the cited multicentre observational study, adenoma detection in non-AI colonoscopy fell from 28.4% to 22.4% after endoscopists were exposed to AI assistance.
Multicentre observational study cited as calibrating evidence; the paper notes that the causal interpretation is contested and that the study included physicians with more than 2,000 procedures each.
Within the tested policy class of the agent-based market simulation, every occupation–jurisdiction cell with an empty signaling window converges to zero engagement and complete fallback-skill collapse.
Agent-based simulation with mobile workers, a Builder and Free-Rider firm structure, and clients forming beliefs from a censored public record.