The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests About 🎲 Workforce Futures
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Evidence (168 claims)

Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.

The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).

Browse by theme

Nine broad, paper-level topics. Click one to filter the claims below.

Adoption
10085 claims
Filter claims →
Productivity
8974 claims
Filter claims →
Governance
8062 claims
Filter claims →
Human-AI Collaboration
7749 claims
Filter claims →
Org Design
5057 claims
Filter claims →
Innovation
4896 claims
Filter claims →
Labor Markets
4088 claims
Filter claims →
Skills & Training
3372 claims
Filter claims →
Inequality
2377 claims
Filter claims →

Claims by outcome category

Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.

Outcome Positive Negative Mixed Null Total
Other 882 244 117 1097 2424
Governance & Regulation 1010 469 229 135 1875
Organizational Efficiency 977 235 149 90 1462
Technology Adoption Rate 781 299 143 128 1362
Research Productivity 506 155 74 363 1110
Output Quality 555 219 71 70 915
Decision Quality 395 200 95 54 751
Firm Productivity 523 67 101 27 724
AI Safety & Ethics 262 309 75 36 688
Market Structure 195 201 135 30 566
Task Allocation 248 77 96 38 464
Innovation Output 300 34 55 20 411
Skill Acquisition 207 75 65 21 368
Employment Level 138 67 119 24 350
Fiscal & Macroeconomic 156 80 53 33 329
Task Completion Time 211 38 13 16 280
Firm Revenue 183 52 29 5 270
Consumer Welfare 131 77 48 13 269
Inequality Measures 50 141 54 9 254
Worker Satisfaction 104 85 25 13 227
Error Rate 87 112 11 5 215
Automation Exposure 69 69 37 20 198
Wages & Compensation 102 49 31 11 193
Team Performance 115 30 30 11 187
Regulatory Compliance 88 74 17 7 186
Training Effectiveness 109 22 14 21 168
Developer Productivity 116 21 15 8 161
Job Displacement 12 92 26 1 131
Hiring & Recruitment 57 12 9 5 83
Skill Obsolescence 6 59 10 2 77
Social Protection 43 17 8 2 70
Creative Output 35 21 9 4 70
Labor Share of Income 18 23 17 1 59
Worker Turnover 15 16 4 35
Industry 1 1
The article develops a conceptual framework linking GenAI use in higher education to knowledge transformation, critical thinking, ethical judgment, digital capability, managerial decision-making, business ethics, workforce readiness, and organizational readiness.
Presentation of a conceptual framework by the authors as part of the review (theoretical/conceptual work; no empirical validation reported).
high mixed Instructing Higher Education in the Era of Generative AI: Im... conceptual linkage among educational inputs and downstream capabilities (knowled...
Given the results, educators should revisit pair programming as an educational tool in addition to embracing modern AI.
Authors' recommendation in the paper's conclusion based on experimental findings (performance, workload, emotion, retention outcomes).
high mixed Fast and Forgettable: A Controlled Study of Novices' Perform... educational practice recommendation (pair programming vs AI-assisted instruction...
The practical burden of scaling depends on how efficiently real resources are converted into that (logical) compute.
Argument in the paper linking conceptual 'logical compute' to real-world conversion efficiency (qualitative claim; no empirical sample in excerpt).
high mixed The Unreasonable Effectiveness of Scaling Laws in AI efficiency of converting real resources into logical compute
Participant targeting: 44% of programs targeted doctors and 44% targeted medical students (with possible overlap), and 56% targeted entry‑to‑practice career stages.
Participant audience and career-stage data extracted from the 27 included programs; proportions reported in the review.
high mixed Assessing the effectiveness of artificial intelligence educa... target audience (doctors, medical students) and career stage distribution (entry...
Most programs were delivered in academic settings: 56% of evaluated programs reported an academic setting.
Setting information extracted from the 27 included programs, with 56% reported as delivered in academic settings.
high mixed Assessing the effectiveness of artificial intelligence educa... program delivery setting (academic vs non-academic)
A plurality of programs were short in duration: 44% of programs were categorized as short courses.
Extraction of program length from the 27 included studies; 44% were classified as short courses per the review's categorization.
high mixed Assessing the effectiveness of artificial intelligence educa... program duration (short vs longer formats)
Most programs were introductory in content: 67% of included programs taught introductory AI concepts rather than advanced/technical AI skills.
Program content extraction across the 27 included studies yielded that 67% were classified as teaching introductory AI.
high mixed Assessing the effectiveness of artificial intelligence educa... program content focus (introductory vs advanced/technical AI skills)
RAD requires estimating cost distributions and choosing a reference policy and quantile-weighting function; these choices determine the method's conservatism and sample efficiency.
Methodological and practical considerations discussed in the paper; noted dependency on estimation and design choices (no quantitative sample-efficiency results provided in the summary).
high mixed Safe RLHF Beyond Expectation: Stochastic Dominance for Unive... method conservatism (relative safety level) and sample efficiency (amount of dat...
Evaluation of the equivalency system should use metrics such as concordance between claimed competencies and verified inputs, predictive validity versus labor-market integration outcomes, and false positive/negative rates in automated decisions.
Methodological recommendation in the paper outlining specific evaluation metrics; this is a prescriptive claim (no empirical implementation reported).
high mixed Establishes a technical and academic bridge between the educ... concordance rate, predictive validity (e.g., accuracy, AUC), false positive/nega...
Diagnostic heuristic: if letting AI in makes the task feel effortless, it is in the wrong place.
Authors' heuristic for educators (conceptual guidance; no empirical test reported in the excerpt).
high negative The Effortless Trap: Productive Struggle, AI, and the Illusi... perceived effort during task when AI is allowed
The architecture of the undergraduate degree is structurally incapable of replacing the informal post-degree apprenticeship system through curricular revision alone.
Argument presented in the paper, supported by the systematic review of eighteen peer-reviewed studies and labor-market analyses cited in the abstract.
high negative Apprenticeship after AI: Bridging Gaps in Early-Career Knowl... capacity of undergraduate curricular revisions to substitute for post-degree app...
Higher education has misdiagnosed the resulting challenge as curriculum misalignment—a content problem assumed to be solvable through revised syllabi, AI electives, and marginal expansions of experiential learning.
Argument presented in the paper, supported by the paper's systematic review of eighteen peer-reviewed studies and labor-market analyses (as described in the abstract).
high negative Apprenticeship after AI: Bridging Gaps in Early-Career Knowl... adequacy of curricular fixes (revised syllabi, AI electives, marginal experienti...
Existing LLM4Rec paradigms are bottlenecked by the difficulty of measuring and improving chain-of-thought (CoT) quality in open-domain recommendation during supervised fine-tuning (SFT).
Author assertion about limitations of prior LLM4Rec paradigms (literature/diagnosis in the paper).
high negative Taiji: Pareto Optimal Policy Optimization with Semantics-IDs... ability to measure and improve CoT quality during SFT
Existing AI education, AI literacy, and human-AI collaboration frameworks remain centred on prompting, task execution, and productivity support and are poorly equipped to address this tacit layer of expert cognition.
Argumentative critique in the paper drawing on conceptual analysis and review of prevailing frameworks; no empirical evaluation or sample reported.
high negative Tacit Signal Infrastructure: Towards AI Systems that Model E... effectiveness-of-current-training-and-collaboration-frameworks-for-tacit-cogniti...
AI adoption presents workforce adaptation challenges.
Reported in the study's literature synthesis and thematic analysis of secondary sources (qualitative review). No sample size reported.
high negative Human–AI Collaboration in the Indian IT Industry: A Qualitat... workforce adaptation / need for retraining
Process-based supervision introduces challenges regarding the sustainability of human-in-the-loop feedback loops.
Socio-technical argumentation in the paper—concern raised about ongoing human verification burden; no longitudinal or empirical data on human labor sustainability provided.
high negative Optimizing Process Based Reward Models through Reinforcement... sustainability of human-in-the-loop feedback (human labor burden / scalability o...
This directional skew is not eliminated by one-shot in-context prompting.
Intervention of one-shot in-context prompting applied to models; evaluation shows the intervention-oriented error skew persists despite one-shot prompting.
high negative Ideological Bias in LLMs' Economic Causal Reasoning effectiveness of one-shot in-context prompting at reducing ideological direction...
The policy and research challenge posed by platform-mediated automation is not merely job quantity (technological unemployment) but institutional continuity — how societies reproduce practical competence when platforms optimize for efficiency rather than formation.
Normative and conceptual claim developed through literature synthesis (institutional economics, platform governance, workforce development); presented as an analytical reframing rather than an empirically tested hypothesis.
high negative When Platforms Replace the Pipeline: AI, Labor Erosion, and ... institutional continuity and human capital reproduction (quality of workforce fo...
Limited reskilling coverage constrains workers' ability to adapt to AI-driven changes.
Paper reviews official reports and secondary data (2020–2024) indicating low coverage/uptake of reskilling programs in India and links this to limited adaptation capacity.
high negative Artificial Intelligence and labour market polarisation in In... coverage/effectiveness of reskilling and workers' adaptive capacity
Across heterogeneous learners, a common broadcast curriculum can be slower than personalized instruction by a factor linear in the number of learner types.
Theoretical comparative result in the model (analysis of broadcast vs personalized curricula across heterogeneous learner types; abstract states factor linear in number of types).
high negative A Mathematical Theory of Understanding speed of instruction / time to learn under broadcast curriculum vs personalized ...
No evaluated program reported Kirkpatrick‑Barr level‑4 outcomes (organizational change, patient outcomes, or sustained metacognitive mastery).
Reviewers mapped reported outcomes from all 27 included programs and found none that demonstrated organizational-level impacts or patient‑level outcomes (level 4).
high negative Assessing the effectiveness of artificial intelligence educa... Kirkpatrick‑Barr level‑4 outcomes (organizational impact, patient outcomes, meta...
Implementing this framework requires significant resources and continuous updating.
Stated explicitly under Main Finding and Disadvantages/Risks; paper lists cost/time metrics to track (cost-per-curriculum, time-to-update) and highlights resource intensity. Support is descriptive/analytic rather than empirical.
high negative Curriculum engineering: organisation, orientation, and manag... resource intensity (cost-per-curriculum), time-to-update, maintenance burden
The paper provides lessons for scaling regression automation and enabling effective human-AI teaming in Agile settings.
Stated contribution of the paper (synthesis of lessons from the industrial case study).
high neutral Human-AI Collaboration for Scaling Agile Regression Testing:... availability of lessons and guidance
The authors filtered that corpus and coded a stratified random sample of 3,100 documents with an LLM-assisted pipeline.
Reported sampling and coding procedure stated in the abstract.
high null result 3100 Opinions on Code Review in an AI World: Building Causal... coded sample size using LLM-assisted pipeline
ATHENA is not presented as a validated measurement instrument; rather, it is a conceptual and methodological scaffold for empirical validation and responsible organizational experimentation.
Explicit qualification in the paper that ATHENA is a conceptual scaffold and has not been validated as a measurement instrument (stated limitation).
high null result Reconceptualizing Competence through Facets: ATHENA as a Str... validation status of ATHENA as a measurement instrument
The study contributes a taxonomy of AI workforce impact, a Workforce Resilience Readiness Score (WRRS), an AI Workforce Trust Index (AWTI), an Ethical Automation Boundary concept, and a pilot empirical validation design.
Declared methodological and conceptual contributions in the paper (these are presented as deliverables of the study; no validated results reported in the excerpt).
high null result From Automation Panic to Workforce Resilience: A Governance ... new measurement/conceptual tools (taxonomy, WRRS, AWTI, Ethical Automation Bound...
The paper introduces the 'Retrainability Index' to measure program outcomes using post-intervention wage recovery and shifts in Routine Task Intensity (RTI).
Methodological contribution described in the paper: formulation of a composite index (Retrainability Index) combining wage recovery and occupation RTI change to evaluate WIOA outcomes.
high null result Did US Worker Retraining Reduce Participant Automation Expos... Retrainability Index (composite of wage recovery and RTI shifts)
The paper evaluates 'Spec Kit' and 'TDAD' as instantiations of the SGM via a four-month pilot study.
Empirical pilot evaluation reported in the paper; duration specified as four months. Sample size or number of teams/participants in pilot not specified in the summary.
high null result The Productivity-Reliability Paradox: Specification-Driven G... evaluation of SGM instantiations (Spec Kit, TDAD) over four months
Self-concordance did not mediate the AI-over-questionnaire effect on goal progress.
Preplanned mediation model reported in the paper found no evidence that self-concordance mediated the AI vs questionnaire effect on goal progress; reported as non-significant in the preregistered analysis.
high null result AI-Assisted Goal Setting Improves Goal Progress Through Soci... goal progress (mediator tested: self-concordance, self-report)
Compared with the matched written-reflection questionnaire, the AI did not significantly improve overall goal progress.
Preplanned comparison within the preregistered RCT; reported non-significant difference between AI and written-reflection condition on overall goal progress at two-week follow-up (no significant p-value reported in the summary).
high null result AI-Assisted Goal Setting Improves Goal Progress Through Soci... goal progress (self-reported goal progress at two-week follow-up)
We conducted a preregistered three-arm randomized controlled trial (RCT) comparing an AI career coach ('Leon,' powered by Claude Sonnet), a matched structured written questionnaire, and a no-support control.
Preregistered RCT reported in the paper; three arms as described; total sample size N = 517; participants randomized to AI coach, written-reflection questionnaire, or no-support control; outcomes assessed at two-week follow-up.
high null result AI-Assisted Goal Setting Improves Goal Progress Through Soci... trial design / allocation and follow-up measurement of goal-related outcomes at ...
Operationalizing DSS requires building domain ontologies/knowledge graphs, designing synthetic curricula, training compact domain models, benchmarking against monolithic LLMs, and measuring total cost-of-ownership (energy, latency, bandwidth, infrastructure).
Paper's recommended experimental and measurement agenda (procedural/methodological prescriptions); this is a proposed research plan rather than an empirical result.
high null result An Alternative Trajectory for Generative AI validation metrics proposed by the paper (benchmark performance, energy/inferenc...
Fine-tuning was done parameter-efficiently: only 0.5% of the Qwen2.5-Coder-7B parameters were trained using GRPO.
Methods section: GRPO-based reinforcement learning fine-tuning, with parameter-efficient update covering 0.5% of model parameters.
high null result Learning to Present: Inverse Specification Rewards for Agent... Proportion of model parameters updated during training (0.5%)
Evaluations reporting outcomes predominantly relied on learner surveys, knowledge/skill tests, or self‑reported behavior change measures.
Methods of evaluation extracted from the included studies: most used surveys, tests, or self-report measures to assess Kirkpatrick‑Barr levels 1–3.
high null result Assessing the effectiveness of artificial intelligence educa... evaluation methods (surveys, tests, self-report behavior change)
Suggested evaluation metrics include placement rates, wage premiums, competency attainment, compliance scores, cost per qualification, and update latency.
Paper's recommended evaluation metrics (prescriptive).
high null result Curriculum engineering: organisation, orientation, and manag... placement rates, wage premiums, competency attainment, compliance scores, cost p...
Recommended analysis methods are qualitative (semi-structured interviews, focus groups, document review) and quantitative (surveys, competency mapping, statistical analysis of outcomes), plus systematic audit methods including traceability checks.
Paper's methods section (methodological specification).
high null result Curriculum engineering: organisation, orientation, and manag... use of specified qualitative, quantitative, and audit methods
Data inputs for the framework should include competency taxonomies, labor-market signals, regulatory requirements, learner assessment results, and stakeholder interviews.
Paper's data-input specification (descriptive).
high null result Curriculum engineering: organisation, orientation, and manag... presence and use of specified data inputs
Research and audit should emphasise validity, reliability, and compliance using mixed methods (qualitative interviews/focus groups; quantitative surveys/statistics) and systematic curriculum audits.
Recommended research & audit approach in paper (methodological guidance).
high null result Curriculum engineering: organisation, orientation, and manag... application of mixed-methods and systematic audits to assess validity/reliabilit...
Tools recommended include logigrams (visual decision/compliance flows) and algorigram (algorithmic step-flows for planning, assessment, audit).
Tool definitions and recommendations in paper (descriptive).
high null result Curriculum engineering: organisation, orientation, and manag... adoption of logigrams and algorigrams in curricula tooling
Core components of the framework are inputs (learner needs, industry requirements, regulatory standards), processes (curriculum mapping, competency alignment, career assessment), and outputs (structured lesson plans, compliance-ready frameworks, career-path documentation).
Framework component list provided in paper (descriptive).
high null result Curriculum engineering: organisation, orientation, and manag... presence and completeness of inputs/processes/outputs in implementation
Scope of the program includes curriculum design, organisational management, career-alignment, and audit/compliance processes.
Explicit scope statement in paper (descriptive).
high null result Curriculum engineering: organisation, orientation, and manag... inclusion of specified scope elements in program design
The framework foregrounds logical modelling (logigrams, algorigrams) and mixed-methods data analysis to support design, auditability, and alignment with industry and regulatory standards.
Paper's methodological design and tool recommendations (conceptual). No empirical implementation data reported.
high null result Curriculum engineering: organisation, orientation, and manag... use of logical modelling tools and mixed-methods analysis in curriculum design
The program offers a comprehensive curriculum-engineering framework linking organizational orientation, management systems, lesson planning, and career assessment into traceable, compliance-ready curriculum products.
Paper's program description and framework specification (conceptual); no empirical evaluation or sample size reported.
high null result Curriculum engineering: organisation, orientation, and manag... availability of traceable, compliance-ready curriculum products (framework prese...
DPS uses the inferred per-prompt state distributions as a predictive prior to select prompts estimated to be most informative, avoiding exhaustive candidate rollouts for filtering.
Method and selection mechanism described: predictive prior ranking/filtering replaces rollout-heavy candidate evaluation. (Procedure described in paper; empirical comparisons reported.)
high null result Dynamics-Predictive Sampling for Active RL Finetuning of Lar... selection of prompts (number of candidate rollouts avoided)
The ARDL framework captures the gradual adjustment process and allows incorporation of human capital by interacting it with AI to assess whether AI benefits differ across skill levels.
Methodological claim in paper describing the advantages of the chosen ARDL specification and the use of interactions with human capital.
high positive Economic Growth, AI Adoption and Human Capital Across the OE... ability to model gradual adjustment and interaction effects between AI adoption ...
Structured reskilling programs, human-centric system design, deliberate role enrichment, and participatory governance are strategic recommendations to address workforce transformation in AI-driven logistics environments.
Conclusions and recommendations from the paper's secondary data review of peer-reviewed research and industry evidence (2022–2026). These are prescriptive recommendations rather than outcomes from a new empirical test; no sample size provided.
high positive Redefining warehouse workforce competencies and roles throug... effectiveness of recommended strategies for addressing workforce transformation ...
As a secondary contribution, the authors offer the underlying LLM-assisted, grey-literature theory-building method as a scalable template for software-engineering research, with a public implementation.
Paper reports method and claims a public implementation; stated in abstract.
high positive 3100 Opinions on Code Review in an AI World: Building Causal... method availability and scalability
Workforce development should be grounded in systems design principles, constraint reduction, and continuous evaluation (i.e., key design principles for workforce development are proposed grounded in systems design).
Prescriptive recommendation emerging from the paper's systems-oriented analysis and synthesis of adult learning theory and organizational design (no empirical evaluation reported).
high positive Optimizing Human Capital in AI-Enabled Architectures: A Syst... effectiveness of workforce development / training approaches for AI-enabled work...
These five dimensions are decomposed into nineteen sub-dimensions and sixty facets, each interpretable through four progressive mastery levels.
Framework taxonomy and granularity as specified in the paper (explicit counts given in the conceptual model; no validation sample).
high positive Reconceptualizing Competence through Facets: ATHENA as a Str... granularity of the facet taxonomy (number of sub-dimensions, facets, mastery lev...
The ATHENA framework is organized around five interdependent dimensions: cognition, conation, knowledge, emotion, and sensorimotor resources.
Descriptive specification of the framework's structure as presented in the paper (conceptual delineation; no empirical measurement reported).
high positive Reconceptualizing Competence through Facets: ATHENA as a Str... dimensions composing the conceptual model of competence