Evidence (83 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
10085 claims
Filter claims →
Productivity
8974 claims
Filter claims →
Governance
8062 claims
Filter claims →
Human-AI Collaboration
7749 claims
Filter claims →
Org Design
5057 claims
Filter claims →
Innovation
4896 claims
Filter claims →
Labor Markets
4088 claims
Filter claims →
Skills & Training
3372 claims
Filter claims →
Inequality
2377 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 882 | 244 | 117 | 1097 | 2424 |
| Governance & Regulation | 1010 | 469 | 229 | 135 | 1875 |
| Organizational Efficiency | 977 | 235 | 149 | 90 | 1462 |
| Technology Adoption Rate | 781 | 299 | 143 | 128 | 1362 |
| Research Productivity | 506 | 155 | 74 | 363 | 1110 |
| Output Quality | 555 | 219 | 71 | 70 | 915 |
| Decision Quality | 395 | 200 | 95 | 54 | 751 |
| Firm Productivity | 523 | 67 | 101 | 27 | 724 |
| AI Safety & Ethics | 262 | 309 | 75 | 36 | 688 |
| Market Structure | 195 | 201 | 135 | 30 | 566 |
| Task Allocation | 248 | 77 | 96 | 38 | 464 |
| Innovation Output | 300 | 34 | 55 | 20 | 411 |
| Skill Acquisition | 207 | 75 | 65 | 21 | 368 |
| Employment Level | 138 | 67 | 119 | 24 | 350 |
| Fiscal & Macroeconomic | 156 | 80 | 53 | 33 | 329 |
| Task Completion Time | 211 | 38 | 13 | 16 | 280 |
| Firm Revenue | 183 | 52 | 29 | 5 | 270 |
| Consumer Welfare | 131 | 77 | 48 | 13 | 269 |
| Inequality Measures | 50 | 141 | 54 | 9 | 254 |
| Worker Satisfaction | 104 | 85 | 25 | 13 | 227 |
| Error Rate | 87 | 112 | 11 | 5 | 215 |
| Automation Exposure | 69 | 69 | 37 | 20 | 198 |
| Wages & Compensation | 102 | 49 | 31 | 11 | 193 |
| Team Performance | 115 | 30 | 30 | 11 | 187 |
| Regulatory Compliance | 88 | 74 | 17 | 7 | 186 |
| Training Effectiveness | 109 | 22 | 14 | 21 | 168 |
| Developer Productivity | 116 | 21 | 15 | 8 | 161 |
| Job Displacement | 12 | 92 | 26 | 1 | 131 |
| Hiring & Recruitment | 57 | 12 | 9 | 5 | 83 |
| Skill Obsolescence | 6 | 59 | 10 | 2 | 77 |
| Social Protection | 43 | 17 | 8 | 2 | 70 |
| Creative Output | 35 | 21 | 9 | 4 | 70 |
| Labor Share of Income | 18 | 23 | 17 | 1 | 59 |
| Worker Turnover | 15 | 16 | — | 4 | 35 |
| Industry | — | — | — | 1 | 1 |
When candidate quality is heterogeneous, prompt injection is less effective on average, but can occasionally allow lower-quality candidates to outrank higher-quality ones, raising fairness concerns.
Controlled experiments comparing homogeneous vs. heterogeneous candidate quality conditions and tracking ranking outcomes; specific experimental counts not included in the abstract.
The LLM fallacy has implications for education, hiring, and AI literacy.
Implications and argumentation presented in the paper; these are prospective and conceptual rather than supported by empirical data in the abstract.
Small differences in managerial incentives can determine which skill path a worker takes (whether they realize full potential or deskill).
Comparative statics / theoretical sensitivity analysis in the dynamic model indicating tipping behavior based on managerial incentives.
Demand for labor will shift toward data scientists, ML engineers, and interdisciplinary scientists, while wet-lab expertise and translational teams remain crucial.
Workforce trend analysis and employer hiring patterns summarized in the paper; interviews/case studies indicating changes in team composition.
The hiring effect strengthens and turns significant on a higher-quality subsample (β2 = −0.039, p < 0.05).
Estimated coefficient reported for a higher-quality subsample in the paper (β2 = −0.039 with p < 0.05).
In the raw data, high-exposure firms cut annual net hiring from 8.3% to 3.1%, while low-exposure engineering–R&D firms held near 7.7%.
Descriptive pre/post comparison of raw firm-level hiring rates reported by the author (percentages before vs after 2022 shock for high- and low-exposure groups).
After the 2022 LLM shock, more-exposed firms slowed net hiring (β2 = −0.004 log points per standard deviation of exposure).
Estimated coefficient from the continuous-treatment DiD model on the full panel (author reports β2 = −0.004 log points per SD of exposure).
The effectiveness of prompt injection rapidly diminishes as more candidates inject, collapsing when manipulation becomes widespread.
Controlled experiments that vary the share of candidates performing prompt injection and observe changes in manipulation effectiveness; exact sample size not provided in the abstract.
Xie et al. (2026) show experimentally that job candidates are less satisfied with firms using AI evaluators than with human experts due to perceived loss of control; the negative effect is stronger for individuals with an internal locus of control.
Experimental study on recruitment using control theory as described (sample size not provided).
Evidence from online labor markets shows a 2%–21% reduction in posting volumes for automatable creative tasks following ChatGPT's release.
Empirical analyses of online labor market posting volumes reported in multiple studies included in the review; range reported across studies.
Across synthesized studies, there was a 14–41% reduction in postings for entry- and mid-level software development and content-creation roles in high-income economies between 2022 and 2024 (range across individual studies: −14% to −41%; median: −23%).
Synthesis of empirical studies retained in the systematic review (numerical range and median reported across non-overlapping study designs and geographies); no pooled meta-analytic estimate provided.
Workers acquire skills through generative AI tools but lack credible ways to signal or validate these skills in competitive freelance markets (a structural challenge the paper terms 'invisible competencies').
Reported finding and conceptual contribution based on the paper's mixed-methods study (survey + semi-structured interviews).
In fixed-unit subsets where complexity rose (Python on the cognitive metric, and all languages on the cyclomatic metric), newcomer participation does not decline.
Subgroup (fixed-unit) analyses that split units by whether complexity rose; DiD estimates within subsets show no decline in newcomer participation despite increases in complexity.
We find no evidence of crowding-out: across estimators newcomer inflow shows no significant decline after adoption (point estimates run from a small increase to, under the most conservative trend specification, a slight and insignificant dip).
Difference-in-differences analysis against matched non-adopting controls, applied to 603 adopters with pre-adoption periods; multiple estimators and trend specifications reported.
The study maps employment channels for AI-competent graduates and documents the most frequent job titles/roles and associated wage levels.
Descriptive analysis of employer channels, occupational role frequencies, and wage data compiled in the monitoring dataset covering graduates and alternative-route entrants.
LLM-based screening is most vulnerable when manipulation is rare and candidate quality differences are small.
Synthesis of experimental results across conditions varying prevalence of manipulation and magnitude of candidate quality differences; sample size not specified in the abstract.
Prompt injection reliably improves applicant rankings when résumé quality is homogeneous and few candidates inject.
Controlled experiments reported in the paper that vary résumé quality homogeneity and fraction of candidates using prompt injection; exact sample size not stated in the abstract.
Spending more time viewing resumes corresponds to candidates' selection chance increasing by 3-4% if they are not recommended.
Experimental analysis of participants' resume-viewing time and selection decisions in a biased AI resume-screening study; comparison conditional on whether AI recommendation was present (text reports a 3–4% increase for non-recommended candidates). Sample size not stated in the provided excerpt.
The most valuable asset a university can offer students in a post-AI economy is credible endorsement—the capacity of a trusted faculty member, advisor, or other mentor to vouch with specificity for a student's character, competence, and potential.
Normative/analytical claim in the essay based on social capital and mentoring research; presented as the author's recommended institutional response rather than empirically validated evidence.
Aggregate employment gains from robot exposure accrue through firm expansion and new worker entry, rather than through intensive-margin expansion of incumbent workers.
Combination of district-level employment growth results and worker-level cohort evidence showing reductions in incumbent worker intensive margins, implying expansion occurs via firm growth and new hires (administrative employer-employee data and industry robot stocks, 2014-2021).
Job-posting analysis shows that approximately 44% of engineering-related positions in the wind sector require advanced digital skills.
Quantified result reported from the paper's job-posting analysis; the summary gives the percentage but does not report the number of job postings analysed.
The Cognitive Operations Manager is proposed as a prototype AI-native professional role for coordinating tacit signal modelling, semantic modelling, AI system calibration, expert validation, and ethical governance.
Proposal of a new professional role in the paper (conceptual/visionary; no pilot study, job analysis, or workforce data reported).
AI-powered EPM helps identify potential leaders.
Summarized outcome across empirical studies in the scoping review (n=29).
The increase in hiring probability is driven by entry-level hires.
Subgroup/heterogeneity analysis within the LinkedIn/GitHub observational data showing the hiring increase concentrated among entry-level SWE hires.
GHC adoption is associated with around a 3%–5% higher monthly probability of hiring SWEs.
Observational analysis using LinkedIn and GitHub data comparing firms that adopted GitHub Copilot (GHC) to firms that did not; association measured as change in firms' monthly probability of hiring software engineers.
There exist successful initiatives, organizational strategies, and policy interventions that have enhanced women’s inclusion, career progression, and representation in emerging tech roles.
Paper reports examples from the reviewed literature and policy analyses that are characterized as 'successful initiatives'; the abstract does not list specific programs, evaluation designs, or sample sizes.
Appointment-level recommendations placed both bots at or above Senior Lecturer level in the Australian university system.
Authors state that appointment-level syntheses from assessors recommended both scholar-bots at or above the Senior Lecturer rank (Australian system); based on the experts' syntheses.
The evidence indicates that AI can support inclusion through assistive technologies and improved matching in labor-market settings.
Synthesis claim based on thematic analysis of the 19 included peer-reviewed studies (qualitative evidence across the corpus pointing to assistive technologies and improved matching as inclusion-supporting mechanisms).
An empirical study revealed that active and targeted individual adaptation can effectively avoid the negative impact of algorithmic bias and significantly improve the overall job search success rates of different groups.
Statement in abstract reporting results of an empirical study conducted by the authors; however, the abstract does not report sample size, experimental design, statistical significance levels, or effect sizes.
Human resources applications of AI focus on recruitment and workforce planning.
Specific thematic finding reported in the abstract from the literature synthesis of included studies.
For high-performing BDA adopters, employee growth is even more pronounced.
Heterogeneity analysis in the paper indicating stronger employee growth among high-performing BDA adopters in the German start-up sample.
Conditional on survival, BDA adopters show stronger employee growth.
Paper reports greater employee growth for surviving BDA adopters compared with non-adopters based on empirical data from German start-ups.
Poaching employees is an inherent aspect of competition for highly qualified talent and is particularly pronounced among tech giants.
Statement in abstract; general observation supported by literature/case-law references implied in paper (no specific empirical sample or quantitative method reported in abstract).
Organizations can design more effective recruitment strategies by signaling AI adoption to increase attractiveness to prospective applicants.
Practical implication drawn from the combined experimental findings (Study 1 N = 145; Study 2 N = 240; total N = 385) showing AI-adoption signals increase organizational attractiveness via perceived innovation ability, particularly for applicants with high AI self-efficacy.
The positive indirect effect of AI-adoption signals on organizational attractiveness via perceived innovation ability is stronger for job seekers with high AI self-efficacy (Study 2 moderated mediation).
Study 2: moderated mediation model showing AI self-efficacy moderates the mediated relationship; sample size N = 240; participants were active job seekers.
Perceived innovation ability mediates the positive association between AI-adoption signals and organizational attractiveness (Study 2).
Study 2: moderated mediation analysis in an experiment recruiting active job seekers; sample size N = 240; mediation of AI-signal -> perceived innovation ability -> organizational attractiveness was validated.
AI-adoption signals are significantly positively associated with organizational attractiveness (Study 1).
Study 1: scenario-based experiment comparing AI-adoption signal vs no-signal conditions; sample size N = 145.
The evaluation compared models on multiple metrics (accuracy, precision, recall, F1, AUC) across repeated trials and cross-company tests, and reported gains for AI methods across these metrics.
Evaluation protocol described: repeated trials, cross-validation, holdout sets, cross-company tests; reported performance improvements for AI models on the listed metrics.
Ensemble methods and deep learning models show the largest and most consistent improvements in predictive performance relative to classic statistical models.
Aggregate results across repeated trials and evaluation metrics indicate Random Forests and Gradient Boosting (ensembles) and deep neural networks outperform linear/logistic regression and other baselines on the publicly available datasets used.
Modern AI-driven prediction methods (especially ensemble models and deep neural networks) systematically outperform traditional statistical approaches at predicting job performance in publicly available workforce datasets.
Direct model comparison reported in the paper: baseline statistical models (linear/logistic regression) versus machine learning models (Random Forest, Gradient Boosting, SVM, deep neural networks) evaluated on multiple publicly available workforce datasets using cross-validation and holdout sets; performance reported on accuracy, precision, recall, F1, and AUC across repeated trials.
The model was prompted to suggest jobs to 24 simulated candidate profiles balanced in terms of gender, age, experience and professional field.
Methods reported in the paper: experimental prompting of GPT-5 with N=24 simulated profiles, balanced across specified attributes.
This study evaluates how a state-of-the-art generative model (GPT-5) suggests occupations based on gender and work experience background for under-35-year-old Italian graduates.
Study design described in the paper: targeted population (under-35 Italian graduates), model used (GPT-5) and evaluation focus (occupation suggestions).
A subset of universities performs markedly better on employment effectiveness, graduate wages, and placement into popular AI roles (i.e., identifiable high-performing institutions).
Comparative analysis across the 191 universities, including employment rates, observed wage outcomes, and placement distributions; identification and reporting of key/high-performing institutions and their metrics.
Adoption of AI in pharma will increase demand for computational biologists, ML engineers, and data scientists and may displace or redefine some traditional bench roles.
Labor-market trend reports and organizational case studies included in the review noting hiring patterns and role changes; qualitative synthesis rather than comprehensive labor-market study.
DAR implies changes to labor and contracting: reversible AI leadership reshapes task boundaries, demand for oversight skills, and should be reflected in contracts and procurement with explicit authority-reversal rules and audit obligations.
Theoretical/ normative argument in implications section; no empirical labor or contract data included.
The paper presents hypothesis tests assessing whether university status (and Alliance ranking) and the presence of specialized AI programs affect graduate employment effectiveness, and reports identification of key/high-performing universities.
Statement of empirical approach: hypothesis testing on effects of university status/Alliance ranking and specialized programs using the monitoring dataset; results and significance levels are reported in the full article.
Analyses of online job postings indicate significant declines in demand for highly automatable and entry-level roles.
Empirical studies using online job-posting data described in the paper (methods: job-posting frequency/trend analysis; sample size/timeframe not specified in the excerpt).
Traditional IT service hiring will be displaced by expansion of product-focused roles and Global Capability Centres (GCCs).
Synthesis of industry reports and workforce data indicating shifts in hiring patterns; the abstract does not report sample sizes or exact metrics.
Despite high overall employment (80% for ages 25–54), nurseries reported they were prevented from hiring new workers due to high wages and unqualified workers.
Reported responses from nurseries (survey/industry responses) referenced in the paper; sample size and survey details not provided in the excerpt.
Higher non-wage costs and higher formalization costs create barriers to creating formal salaried employment and alter firms’ hiring and investment decisions.
Theoretical and policy interpretation based on measured NWC and CFIL levels in the 19-country sample and economic reasoning about how employer cost structure affects hiring and investment incentives; no firm-level causal estimation reported.