The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests About 🎲 Workforce Futures
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Evidence (8974 claims)

Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.

The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).

Browse by theme

Nine broad, paper-level topics. Click one to filter the claims below.

Adoption
10085 claims
Filter claims →
Productivity
8974 claims
Filtered →
Governance
8062 claims
Filter claims →
Human-AI Collaboration
7749 claims
Filter claims →
Org Design
5057 claims
Filter claims →
Innovation
4896 claims
Filter claims →
Labor Markets
4088 claims
Filter claims →
Skills & Training
3372 claims
Filter claims →
Inequality
2377 claims
Filter claims →

Claims by outcome category

Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.

Outcome Positive Negative Mixed Null Total
Other 882 244 117 1097 2424
Governance & Regulation 1010 469 229 135 1875
Organizational Efficiency 977 235 149 90 1462
Technology Adoption Rate 781 299 143 128 1362
Research Productivity 506 155 74 363 1110
Output Quality 555 219 71 70 915
Decision Quality 395 200 95 54 751
Firm Productivity 523 67 101 27 724
AI Safety & Ethics 262 309 75 36 688
Market Structure 195 201 135 30 566
Task Allocation 248 77 96 38 464
Innovation Output 300 34 55 20 411
Skill Acquisition 207 75 65 21 368
Employment Level 138 67 119 24 350
Fiscal & Macroeconomic 156 80 53 33 329
Task Completion Time 211 38 13 16 280
Firm Revenue 183 52 29 5 270
Consumer Welfare 131 77 48 13 269
Inequality Measures 50 141 54 9 254
Worker Satisfaction 104 85 25 13 227
Error Rate 87 112 11 5 215
Automation Exposure 69 69 37 20 198
Wages & Compensation 102 49 31 11 193
Team Performance 115 30 30 11 187
Regulatory Compliance 88 74 17 7 186
Training Effectiveness 109 22 14 21 168
Developer Productivity 116 21 15 8 161
Job Displacement 12 92 26 1 131
Hiring & Recruitment 57 12 9 5 83
Skill Obsolescence 6 59 10 2 77
Social Protection 43 17 8 2 70
Creative Output 35 21 9 4 70
Labor Share of Income 18 23 17 1 59
Worker Turnover 15 16 4 35
Industry 1 1
Clear
Productivity Remove filter
We disaggregate and then substantially reorganize the approximately 20K activities in the US Department of Labor's O*NET occupational database to produce a comprehensive ontology of work activities.
Methodological: authors report transforming the O*NET activity taxonomy (~20,000 activity-level records) by disaggregation and reorganization into a new ontology.
high positive Where can AI be used? Insights from a deep ontology of work ... creation of a comprehensive ontology of work activities
Models trained in EnterpriseLab remain robust across diverse enterprise benchmarks, including EnterpriseBench (+10%) and CRMArena (+10%).
Benchmark evaluations reported in the paper showing reported +10% improvements on EnterpriseBench and CRMArena relative to baseline; exact baselines, statistical tests, and sample sizes are not specified in the abstract.
high positive EnterpriseLab: A Full-Stack Platform for developing and depl... benchmark performance on EnterpriseBench and CRMArena
8B-parameter models trained in EnterpriseLab reduce inference costs by 8-10x compared to frontier models (implied GPT-4o).
Empirical cost comparison reported in the paper; the abstract states an 8-10x reduction in inference costs for the 8B models trained in EnterpriseLab versus the referenced frontier model(s). Detailed cost accounting and sample sizes not provided in the abstract.
8B-parameter models trained within EnterpriseLab match GPT-4o's performance on complex enterprise workflows.
Empirical evaluation reported in the paper comparing 8B-parameter models trained in EnterpriseLab to GPT-4o on complex enterprise workflows; specific benchmark tests and metrics are referenced but details (sample sizes, exact metrics) are not provided in the abstract.
high positive EnterpriseLab: A Full-Stack Platform for developing and depl... model performance on complex enterprise workflows (task success/quality)
We validate the platform through EnterpriseArena, an instantiation with 15 applications and 140+ tools across IT, HR, sales, and engineering domains.
Reported instantiation/experimental setup in the paper: EnterpriseArena contains 15 applications and 140+ tools spanning specified domains.
high positive EnterpriseLab: A Full-Stack Platform for developing and depl... scope/scale of experimental validation (number of applications and tools)
EnterpriseLab provides integrated training pipelines with continuous evaluation.
System/design claim in paper describing integrated training and evaluation tooling as part of the platform.
high positive EnterpriseLab: A Full-Stack Platform for developing and depl... availability of integrated training pipelines and continuous evaluation
EnterpriseLab includes automated trajectory synthesis that programmatically generates training data from environment schemas.
System/design claim described in paper; supported by the authors' description of an automated data-generation component.
high positive EnterpriseLab: A Full-Stack Platform for developing and depl... automated generation of training trajectories from environment schemas
EnterpriseLab provides a modular environment exposing enterprise applications via a Model Context Protocol, enabling seamless integration of proprietary and open-source tools.
Feature/design claim in paper; supported by implementation details of the 'Model Context Protocol' and reported integration capabilities in the platform description.
high positive EnterpriseLab: A Full-Stack Platform for developing and depl... tool/application integration capability
We introduce EnterpriseLab, a full-stack platform that unifies tool integration, data generation, and training into a closed-loop framework.
System/design claim describing the contribution of the paper (platform implementation and architecture); supported by the paper's implementation description rather than independent validation.
high positive EnterpriseLab: A Full-Stack Platform for developing and depl... existence and integration of a unified development pipeline (tool integration, d...
The paper reports details from a 100% deployment of DRL with policy regularizations on Alibaba's e-commerce platform, Tmall.
Direct statement in the abstract claiming full deployment across Tmall; implies a real-world, company-scale deployment but the abstract provides no operational metrics or counts.
high positive DeepStock: Reinforcement Learning with Policy Regularization... deployment/adoption of the DRL-with-regularization system
Imposing policy regularizations improves the final performance of several DRL methods for inventory management.
Empirical claim supported by the paper's synthetic experiments and reported production deployment on Alibaba/Tmall (as stated in the abstract); no quantitative effect sizes provided in the abstract.
high positive DeepStock: Reinforcement Learning with Policy Regularization... final performance (policy quality) of DRL inventory methods
Imposing policy regularizations, grounded in classical inventory concepts such as 'Base Stock', can significantly accelerate hyperparameter tuning for DRL methods.
Paper reports synthetic experiments and a production deployment (Alibaba/Tmall) where policy regularizations were applied; abstract claims acceleration in hyperparameter tuning but does not report numeric tuning-time metrics in the abstract.
high positive DeepStock: Reinforcement Learning with Policy Regularization... speed/efficiency of hyperparameter tuning
Deep Reinforcement Learning (DRL) provides a general-purpose methodology for training inventory policies that can leverage big data and compute.
Argument/assertion made in the paper's introduction/abstract (conceptual claim about DRL capabilities); no empirical sample or quantitative test reported in the abstract.
high positive DeepStock: Reinforcement Learning with Policy Regularization... ability to train inventory policies using large data and compute
Human-replacing technologies have a strategic role in enhancing industrial productivity and ensuring the long-term resilience of Ukraine’s mining and metallurgical sector amid workforce shortages and structural labour-market changes due to war and demographic decline.
Integrated sectoral assessment in the paper combining current context (workforce shortages, structural changes), literature on technology-driven productivity/resilience, and industry-specific considerations; presented as a high-level conclusion.
high positive Human-replacing technologies as a driver of labour productiv... industrial productivity and sectoral resilience
Integrating ergonomic assessments and human–systems–interaction approaches into automation projects is important to prevent cognitive overload, occupational stress and operational risks for control‑room operators.
Recommendation and emphasis in the paper, supported by references to ergonomics and human-factors literature; presented as a preventive/mitigative approach rather than a quantified empirical result for the sector.
high positive Human-replacing technologies as a driver of labour productiv... cognitive overload, occupational stress, operational risk (errors/incidents)
Successful technological modernization requires continuous investment in human capital, reskilling and the development of digital and engineering competencies.
Policy/recommendation based on the paper's synthesis of the sector analysis and literature on skill requirements and technology adoption; not presented as an original empirical estimate in the summary.
high positive Human-replacing technologies as a driver of labour productiv... effectiveness of modernization efforts via training/reskilling investments
Higher robot density is associated with productivity gains, particularly in low-robotized sectors such as Ukraine’s mining and metallurgical industry.
Empirical evidence cited from international and industry-specific studies reviewed in the paper (literature review/meta-analytic style evidence); no Ukraine-specific causal estimate with sample size reported in the summary.
high positive Human-replacing technologies as a driver of labour productiv... productivity (associated gains)
Human-replacing technologies also have an indirect impact on productivity by increasing total factor productivity (TFP).
Analytical argumentation in the paper supported by references to empirical studies showing TFP effects of automation/digitalization; literature synthesis rather than a new econometric estimate presented for Ukraine.
high positive Human-replacing technologies as a driver of labour productiv... total factor productivity
Human-replacing technologies (mechanization, automation, robotization, digitalization and AI-augmentation) make a direct contribution to labour productivity growth in Ukraine's mining and metallurgical sector.
Sectoral analysis and synthesis in the paper drawing on empirical international and industry-specific studies; literature review of productivity impacts of mechanization/automation/robotization/digitalization/AI in industrial contexts.
Industrial intelligence and the digital economy can be leveraged as a 'dual engine' to boost regional TFCP and advance high-quality green and low-carbon economic development, supporting differentiated regional coordination policies.
Synthesis/implication drawn from the paper's empirical findings (SDM results on 30 provinces, 2010–2023) showing positive total/spillover effects and regional heterogeneity.
high positive Study on the impact of industrial intelligence and the digit... total factor carbon productivity (TFCP)
Green finance has an insignificant positive effect on regional TFCP.
Coefficient on green finance control variable in the Spatial Durbin Model (30 provinces, 2010–2023) is positive but not statistically significant.
high positive Study on the impact of industrial intelligence and the digit... total factor carbon productivity (TFCP)
The digital economy presents different regional driving patterns: a 'local-spillover dual drive' in the east, a 'local-dominated drive' in the central region, and a 'spillover-dominated drive' in the west.
Regional/subsample Spatial Durbin Model estimates for digital economy variables across east, central, and west subsamples (30 provinces, 2010–2023) with reported direct and indirect effects.
high positive Study on the impact of industrial intelligence and the digit... total factor carbon productivity (TFCP)
The digital economy exerts a significantly positive direct effect on local TFCP and a strong positive spatial spillover effect, forming a 'local driving + spatial radiation' promotion pattern.
Spatial Durbin Model estimates on panel data (30 provinces, 2010–2023) showing statistically significant positive direct and indirect (spillover) coefficients for digital economy variables.
high positive Study on the impact of industrial intelligence and the digit... total factor carbon productivity (TFCP)
Regional TFCP shows significant positive spatial autocorrelation.
Spatial analysis (Spatial Durbin Model and spatial statistics) applied to panel of 30 provincial-level regions; reported significant spatial autocorrelation (e.g., positive Moran's I implied).
high positive Study on the impact of industrial intelligence and the digit... total factor carbon productivity (TFCP)
Across 378 hardware validated experiments, concise human-expert skills with structured expert knowledge enable near-perfect success rates across platforms.
Reported experimental results: 378 hardware-validated experiments across platforms comparing agent configurations; finding reported that human-expert skills produce near-perfect success rates (no numeric success rate provided in excerpt).
high positive Skilled AI Agents for Embedded and IoT Systems Development task success rate (hardware-validated)
Large language models (LLMs) and agentic systems have shown promise for automated software development.
Statement in paper referencing prior successes of LLMs and agentic systems for automated software development (no empirical data reported in this excerpt).
high positive Skilled AI Agents for Embedded and IoT Systems Development automation-assisted software development capability
The largest gains appear when AI is embedded in an orchestrated workflow rather than deployed as an isolated coding assistant.
Central thesis supported by comparisons across five delivery configurations (traditional baseline and V1–V4) in a retrospective longitudinal field study of the Chiron platform applied to three real software modernization programs; authors observe greater portfolio-level improvements when AI is integrated into coordinated workflows.
high positive Orchestrating Human-AI Software Delivery: A Retrospective Lo... aggregate team/organizational performance (speed, coverage, issue load) when AI ...
V3 and V4 add acceptance-criteria validation, repository-native review, and hybrid human-agent execution, simultaneously improving speed, coverage, and issue load.
Observed differences across the five delivery configurations (baseline, V1–V4) in the field study of three modernization programs; authors link feature additions in V3/V4 to measured improvements in stage durations, coverage, and validation issues.
high positive Orchestrating Human-AI Software Delivery: A Retrospective Lo... stage durations (speed), first-release coverage, validation-stage issue load
First-release coverage rises from 77.0% to 90.5% across the portfolio as platform versions progress.
Observed first-release coverage measured in the retrospective longitudinal field study of three real modernization programs, reported as percentages across delivery configurations.
high positive Orchestrating Human-AI Software Delivery: A Retrospective Lo... first-release coverage (percent of tasks covered on first release)
Validation-stage issue load falls from 8.03 to 2.09 issues per 100 tasks across the portfolio as platform versions progress.
Observed outcomes from the retrospective field study on three programs; validation-stage issues counted and normalized per 100 tasks across delivery configurations.
high positive Orchestrating Human-AI Software Delivery: A Retrospective Lo... validation-stage issues per 100 tasks
Modeled senior-equivalent effort falls from 1080.0 to 139.5 SEE-days under the platform configurations studied.
Modeled senior-equivalent effort computed from the study's staffing scenarios and observed outputs across the three real programs.
high positive Orchestrating Human-AI Software Delivery: A Retrospective Lo... senior-equivalent effort (SEE-days)
Modeled raw effort falls from 1080.0 to 232.5 person-days under the platform configurations studied (baseline -> V4 aggregate).
Modeled outcomes computed from observed task volumes and explicit staffing scenarios in the retrospective longitudinal field study covering three real programs.
Portfolio totals move from 36.0 to 9.3 summed project-weeks under baseline staffing assumptions (across the three studied programs and five delivery configurations).
Retrospective longitudinal field study of the Chiron platform applied to three real software modernization programs (COBOL banking migration ~30k LOC, accounting modernization ~400k LOC, .NET/Angular mortgage modernization ~30k LOC); observed and modeled outcomes were aggregated to produce portfolio totals under explicit staffing scenarios.
high positive Orchestrating Human-AI Software Delivery: A Retrospective Lo... summed project-weeks (portfolio time)
Reinforcement learning (post-training) on our corpus improves downstream embodied manipulation performance.
Downstream evaluation described in the paper showing improved performance on embodied manipulation tasks after RL post-training on MultihopSpatial-Train.
high positive MultihopSpatial: Multi-hop Compositional Spatial Reasoning B... embodied manipulation task performance
Reinforcement learning (post-training) on our MultihopSpatial-Train corpus enhances intrinsic VLM spatial reasoning.
Experimental intervention: RL-based post-training on the authors' training corpus followed by evaluation on intrinsic spatial reasoning benchmarks (described in the paper).
high positive MultihopSpatial: Multi-hop Compositional Spatial Reasoning B... intrinsic spatial reasoning performance of VLMs
We provide MultihopSpatial-Train, a dedicated large-scale training corpus intended to foster spatial intelligence in VLMs.
Dataset/resource contribution described in the paper (existence and intended use of MultihopSpatial-Train).
high positive MultihopSpatial: Multi-hop Compositional Spatial Reasoning B... training resource availability for spatial intelligence
We propose Acc@50IoU, a complementary metric that simultaneously evaluates reasoning and visual grounding by requiring both answer selection and precise bounding box prediction.
Methodological contribution in the paper defining the Acc@50IoU metric and its intended use to measure combined answer correctness and bounding-box IoU >= 0.5.
high positive MultihopSpatial: Multi-hop Compositional Spatial Reasoning B... combined answer accuracy and box localization (reasoning + visual grounding)
We introduce MultihopSpatial, a comprehensive benchmark designed for multi-hop and compositional spatial reasoning, featuring 1- to 3-hop complex queries across diverse spatial perspectives.
Dataset/benchmark construction described in the paper (design and scope of MultihopSpatial).
high positive MultihopSpatial: Multi-hop Compositional Spatial Reasoning B... ability to evaluate multi-hop and compositional spatial reasoning
Spatial reasoning is foundational for Vision-Language Models (VLMs), particularly when deployed as Vision-Language-Action (VLA) agents in physical environments.
Conceptual/introductory statement in the paper motivating the work (literature-based argument about VLMs and VLA agents).
high positive MultihopSpatial: Multi-hop Compositional Spatial Reasoning B... spatial reasoning capability as a foundational requirement
The findings position AI not merely as an operational tool but as a strategic orchestrator of regenerative production systems, offering a clear roadmap for accelerating circular transitions in line with the Sustainable Development Goals.
Conclusions drawn from the mixed-methods review (bibliometric analysis of 196 articles and systematic review of 104 studies) as reported in the abstract.
high positive Artificial intelligence as a catalyst for the circular econo... role of AI in enabling/regenerating production systems and accelerating circular...
Artificial intelligence is emerging as a powerful driver of the circular economy (CE), enabling production systems to become more resource-efficient, less waste-intensive and strategically aligned with sustainability goals.
Mixed-methods assessment combining bibliometric network analysis (196 peer-reviewed articles, 2023–2024) and a systematic review of 104 studies, as reported in the abstract.
high positive Artificial intelligence as a catalyst for the circular econo... resource efficiency and waste intensity of production systems
AI can reduce production scrap by as much as 30% in documented cases.
Systematic review of studies (paper reports a systematic review of 104 studies); the abstract cites documented cases showing up to 30% reduction in production scrap.
high positive Artificial intelligence as a catalyst for the circular econo... production scrap (waste generated during production)
AI can increase resource-efficiency metrics by up to 25% in documented cases.
Systematic review of studies (paper reports a systematic review of 104 studies); the abstract states documented cases showing up to 25% increases in resource-efficiency metrics.
high positive Artificial intelligence as a catalyst for the circular econo... resource-efficiency metrics
GenAI implementations that are strategically deployed in managed Azure cloud infrastructure provide a positive ROI over time when aligned with business processes, enterprise architecture, and performance metrics.
Conclusion drawn from the paper's mixed-method analysis (quantitative ROI modelling, cost–benefit analysis, and case study synthesis).
high positive Measuring Business ROI of Generative AI Adoption on Azure Cl... Return on Investment (ROI) over time
Close coupling among Azure OpenAI Service, Azure Machine Learning, and cost governance tooling (FinOps) significantly decreases overall cost of ownership and enhances scalability and compliance.
Architectural analysis of Azure-native GenAI services and cost/governance tooling reported in the paper.
high positive Measuring Business ROI of Generative AI Adoption on Azure Cl... overall cost of ownership, scalability, compliance
Measurable ROI from GenAI on Azure is mainly driven by improvements in productivity, optimization of operational costs, faster decision making, and increased speed of innovation across business functions.
Reported results from the paper's mixed-method study combining quantitative ROI modelling and cost–benefit analysis plus qualitative synthesis of secondary enterprise case studies.
high positive Measuring Business ROI of Generative AI Adoption on Azure Cl... business Return on Investment (ROI) driven by productivity, cost optimization, d...
Microsoft Azure has become one of the first enterprise-scale platforms facilitating GenAI-driven change.
Statement in the paper's abstract asserting Azure's market position as an early enterprise-scale platform for GenAI.
high positive Measuring Business ROI of Generative AI Adoption on Azure Cl... enterprise-scale platform adoption
Our empirics demonstrate that self-evolving AI offers a scalable and interpretable paradigm.
Empirical results on the U.S. equity market are cited as evidence; the paper claims scalability and interpretability based on those empirical demonstrations and the architecture of the system.
high positive Beyond Prompting: An Autonomous Framework for Systematic Fac... scalability and interpretability of the AI-driven investing approach
Applying this methodology to the U.S. equity market, long-short portfolios formed on the simple linear combination of signals deliver a return of 59.53% (annualized).
Empirical backtest/application to the U.S. equity market reported in the paper; specific annualized return percentage is provided. Sample period, universe, and number of observations not stated in the excerpt.
high positive Beyond Prompting: An Autonomous Framework for Systematic Fac... annualized portfolio return
Applying this methodology to the U.S. equity market, long-short portfolios formed on the simple linear combination of signals deliver an annualized Sharpe ratio of 3.11.
Empirical backtest/application to the U.S. equity market reported in the paper; specific performance metric (annualized Sharpe) is provided. Sample period, universe, and number of observations not stated in the excerpt.