The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests About 🎲 Workforce Futures
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Evidence (8807 claims)

Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.

The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).

Browse by theme

Nine broad, paper-level topics. Click one to filter the claims below.

Adoption
9875 claims
Filter claims →
Productivity
8807 claims
Filtered →
Governance
7870 claims
Filter claims →
Human-AI Collaboration
7560 claims
Filter claims →
Org Design
4892 claims
Filter claims →
Innovation
4781 claims
Filter claims →
Labor Markets
4004 claims
Filter claims →
Skills & Training
3308 claims
Filter claims →
Inequality
2332 claims
Filter claims →

Claims by outcome category

Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.

Outcome Positive Negative Mixed Null Total
Other 870 233 116 1066 2363
Governance & Regulation 976 451 218 133 1809
Organizational Efficiency 949 224 144 88 1416
Technology Adoption Rate 764 287 141 122 1325
Research Productivity 501 152 74 362 1101
Output Quality 542 216 69 69 896
Decision Quality 387 198 94 54 740
Firm Productivity 513 67 101 27 714
AI Safety & Ethics 249 303 73 36 667
Market Structure 190 192 134 27 548
Task Allocation 243 77 91 36 452
Innovation Output 291 33 55 20 401
Skill Acquisition 206 72 65 21 364
Employment Level 133 63 115 22 335
Fiscal & Macroeconomic 153 79 52 32 323
Task Completion Time 206 37 12 15 272
Firm Revenue 179 52 29 5 266
Consumer Welfare 130 76 47 13 266
Inequality Measures 48 137 51 6 242
Worker Satisfaction 101 81 25 13 220
Error Rate 84 110 11 5 210
Wages & Compensation 98 47 30 10 185
Regulatory Compliance 88 73 17 7 185
Automation Exposure 66 64 33 16 182
Team Performance 105 29 30 11 176
Training Effectiveness 109 22 14 21 168
Developer Productivity 114 21 14 8 158
Job Displacement 12 90 24 1 127
Hiring & Recruitment 57 9 9 5 80
Skill Obsolescence 6 56 9 1 72
Social Protection 43 17 8 2 70
Creative Output 35 21 9 4 70
Labor Share of Income 18 21 17 1 57
Worker Turnover 15 16 4 35
Industry 1 1
Clear
Productivity Remove filter
The causal effect of adoption on architectural smell density (ASD) was estimated using a staggered difference-in-differences design and the Borusyak imputation estimator.
Methodological claim describing the causal identification and estimation strategy applied to the 151-repository panel.
high null result Mining Architectural Quality Under Agentic AI Adoption: A Ca... architectural smell density (ASD)
We mined 151 open-source Java repositories, 74 with detectable agentic AI adoption (identified via configuration files and Co-Authored-By commit trailers) and 77 propensity-matched controls, across a 13-month per-repository window yielding 1,811 monthly Arcan snapshots.
Descriptive dataset and methods statement in paper: 151 repositories (74 treated, 77 matched controls), 13-month windows, producing 1,811 monthly Arcan snapshots.
high null result Mining Architectural Quality Under Agentic AI Adoption: A Ca... dataset composition (number of repositories, treated vs control, snapshots)
Causal evidence on the effect of AI coding tool adoption on software architecture is scarce; prior causal work has focused on code-level outcomes (complexity, static analysis warnings) and whether such degradation propagates to architecture-level outcomes remains unknown.
Literature/background statement in abstract asserting gaps in prior work; not an empirical result from this paper's dataset.
high null result Mining Architectural Quality Under Agentic AI Adoption: A Ca... state of the literature regarding causal evidence on architecture-level effects
The superior zero-shot forecasting accuracy of foundation models does not inherently translate into better decision utility for resource consolidation.
Empirical analysis in the paper mapping forecasting outputs to downstream consolidation decisions and utility metrics, showing lack of improvement in decision utility despite better forecast accuracy.
high null result CloudCons: A Comprehensive End-to-End Benchmark for Cloud Re... decision utility in resource consolidation (trade-off between resource efficienc...
We analyze 15,549 agentic PRs from 148 projects in the AIDev dataset.
Descriptive statement of dataset and sample used in the study (paper reports analysis of 15,549 agentic PRs from 148 projects).
high null result Toward Instructions-as-Code: Understanding the Impact of Ins... number of agentic pull requests and projects analyzed
From 2024 to 2026, more than 130 articles were submitted to this Special Issue (SI), and only 18 papers were accepted after rigorous peer review.
Editorial report in the paper describing CFP submissions and acceptance counts.
high null result Guest editorial: Digital age wisdom in Chinese management: a... number of submissions and acceptances for the SI
We conduct a qualitative study on a representative sample of 306 non-merged pull requests created or co-authored by the agents mentioned earlier, followed by a quantitative analysis of the reasons for rejection.
Authors' reported methods: qualitative study of a sample of 306 non-merged PRs and subsequent quantitative analysis.
high null result Understanding the Rejection of Fixes Generated by Agentic Pu... qualitative and quantitative characterization of non-merged PRs
Professional radiologists analyzed chest X-rays with access to state-of-the-art machine learning predictions in this replication setting.
Description of the experimental/contextual setting in the paper: professional radiologists using ML predictions on chest X-rays drawn from Collab-CXR.
high null result Revisiting the ABCs of Working with AI: A Replication with R... task context — radiology readings with ML assistance
The replication uses radiologist assessments from repeated-case designs, which include 68 radiologists and 11,420 paired radiologist–patient–pathology observations.
Direct reporting of study sample and design in the paper (repeated-case design; counts of radiologists and paired observations).
high null result Revisiting the ABCs of Working with AI: A Replication with R... sample composition (number of radiologists and paired observations)
This note leverages the public Collab-CXR data repository described by Moehring et al. (2025) and first analyzed for human-AI collaboration by Agarwal et al. (2023).
Explicit statement in the paper identifying the data source and prior analyses (references provided).
high null result Revisiting the ABCs of Working with AI: A Replication with R... dataset used for replication (Collab-CXR)
In a production switchback experiment, the offline-trained policy reduces courier-side time costs without degrading customer-facing delivery quality.
Empirical claim supported by production switchback experiment described in the paper; asserts no degradation in customer-facing delivery quality concurrent with courier-side time improvements (no numerical metrics or sample sizes provided in excerpt).
high null result Multi-Agent Reinforcement Learning from Delayed Marketplace ... customer-facing delivery quality
The study uses a qualitative, mixed-methods design combining a systematic literature review, secondary evidence from an industry MRO digital survey, five semi-structured expert interviews, and two technical case studies (neural networks for aircraft retirement and an AI-based digital twin for a Power Electronics Cooling System).
Methods description provided in the paper (explicit counts: 5 interviews, 2 case studies); method = author-reported study design.
high null result Aviation 4.0: the impacts of digital transformation on the a... study design and methods employed
After screening, 35 studies were included in the thematic synthesis and supplemented by official regulatory and industry documents.
Review screening result reported in the paper: number of included studies = 35; supplementation by regulatory and industry documents stated.
high null result Artificial Intelligence-Driven Optimization in Pharmacy Inve... number of included studies and supplementary documents
A structured search protocol was designed for Scopus, Web of Science, PubMed, IEEE Xplore, and Google Scholar covering January 2016 to May 2026, English-language records only.
Methods statement in the review describing the databases, date range, and language restriction used for the systematic search.
high null result Artificial Intelligence-Driven Optimization in Pharmacy Inve... search protocol (databases, date range, language)
The implementation literature on AI for pharmacy inventory and pharmaceutical supply chains remains dispersed across pharmacy operations, operations research, health informatics, and supply chain analytics.
The review's thematic synthesis of the searched literature (review methods described below) identified studies across these disciplinary areas.
high null result Artificial Intelligence-Driven Optimization in Pharmacy Inve... disciplinary distribution of implementation literature
Devil's Advocate (DA) is an AI assistant that critiques the human's initial ideas, whereas Dialectical Inquiry (DI) provides alternatives and synthesizes a resolution.
Conceptual/definitional claim in the paper describing the operationalization of DA and DI for the experiments.
high null result Shaping The Tool Or Shaping The Mind: An Investigation Of Du... operational definition of AI-supported conflict techniques
This research empirically compares DA and DI in AI contexts.
Paper reports experimental comparison between AI behaviors implementing Devil's Advocate (DA) and Dialectical Inquiry (DI) across the studies.
high null result Shaping The Tool Or Shaping The Mind: An Investigation Of Du... comparative effects of DA vs DI on SDM outcomes
Both studies examine benefit (information elaboration) and cost (cognitive load) pathways when AI supports SDM.
Paper explicitly frames both studies to measure information elaboration as a benefit pathway and cognitive load as a cost pathway; stated measurement plan in methods.
high null result Shaping The Tool Or Shaping The Mind: An Investigation Of Du... information elaboration and cognitive load
Study 2 tests mind-shaping interventions through user strategy training.
Study design described in the paper: a second experiment (Study 2) manipulating user strategy training (mind-shaping) to evaluate effects on SDM processes and outcomes.
high null result Shaping The Tool Or Shaping The Mind: An Investigation Of Du... effects of user strategy training on information elaboration and cognitive load
Study 1 tests tool-shaping interventions by comparing three AI bot prototype conditions (Information-only, DA, DI) against a control treatment.
Study design described in the paper: randomized/controlled experiment (Study 1) with four conditions (three AI prototype conditions plus control).
high null result Shaping The Tool Or Shaping The Mind: An Investigation Of Du... effects of AI prototype conditions on information elaboration and cognitive load
The 'do no harm' property is confirmed empirically.
Abstract states empirical confirmation in simulations and applications; specifics (e.g., datasets, sample sizes) not included in abstract.
high null result AI-Assisted Variance Reduction in Randomized Experiments empirical verification that adjusted estimator does not worsen performance when ...
Including AI predictions as covariates has a 'do no harm' property: the adjusted estimator reverts to the unadjusted difference in means when predictions are uninformative.
Stated theoretical property in the paper and described as empirically confirmed in simulations and applications (per abstract).
high null result AI-Assisted Variance Reduction in Randomized Experiments bias/consistency and non-worsening of estimator when predictions uninformative
Raw blind-panel decision quality is similar for A and B (7.01 vs. 6.96).
Blind-panel scoring of generated reports from agents A and B; panel size and panel methodology not specified in abstract.
high null result AI Scientists Are Only as Good as Their Evidence: A Stratifi... raw blind-panel decision-quality score
Self-evaluated creative performance remained unchanged when using GenAI.
Same experiment with 82 participants; authors report no significant difference in self-evaluated creative performance between GenAI users and controls.
high null result When Ai Sparks Less: Generative Ai And The Decline Of Self-P... self-evaluated creative performance
Each of the four published papers used in the experiments contained an error that I helped identify or correct.
Author statement that the 4 papers each contained an error; author involvement in identification/correction is asserted.
high null result Can AI Refute Economic Theory? Evidence from Beyond the Know... presence of errors in the 4 target papers
I conducted experiments in which I asked several AI models (Gemini, Refine, Claude, and ChatGPT) to check the correctness of four published papers in economic theory.
Author reports running direct experiments: prompted listed models to check 4 published economic-theory papers.
high null result Can AI Refute Economic Theory? Evidence from Beyond the Know... existence of experiments using specified models on 4 papers
We evaluate the system on operator feedback and a question set collected from production usage, graded by human and automated panels.
Paper's stated evaluation methodology: operator feedback + production question set, graded by humans and automated panels.
high null result Archi: Agentic Operations at the CMS Experiment evaluation methodology (feedback and graded question set)
This study benchmarks Algeria’s readiness to adopt AI against Morocco, Egypt, and Turkey using data from the World Bank (2022), the Oxford Insights Government AI Readiness Index, and sector-specific studies.
Methodological statement in the paper specifying data sources used for the comparative assessment (World Bank 2022, Oxford Insights index, sector studies).
high null result Artificial Intelligence and Economic Productivity: A Compara... AI readiness / readiness indicators
The article aims to provide systematic literature support for subsequent research and adaptive policy formulation.
Statement of the paper's stated objective; methodological and policy-intent claim from the authors.
high null result Influence of Artificial Intelligence in the Labor Market policy formulation support
This article is based on a systematic literature review and summarizes the four core theoretical mechanisms of substitution, complementarity, new task creation, and skill mismatch.
Methodological claim from the paper: the authors conducted a systematic literature review and identified these four theoretical mechanisms.
high null result Influence of Artificial Intelligence in the Labor Market theoretical mechanisms
Traditional software and agentic systems are distinct: in traditional software code is the carrier of decision logic, whereas in agentic systems code is ephemeral tooling used by an LLM-driven reasoning loop.
Formalization and conceptual definitions developed in the paper (first-principles formal distinction; no empirical sample size reported).
high null result The End of Software Engineering: How AI Agents Are Fundament... architectural role of code (carrier of logic vs ephemeral tool)
For over half a century, software engineering has operated on a foundational premise: human engineers decompose problems, encode decision logic into static code, and manually adapt that code as requirements evolve.
Historical/descriptive claim presented in the paper's framing and literature review; citation of longstanding software engineering practices (qualitative, no empirical sample size reported).
high null result The End of Software Engineering: How AI Agents Are Fundament... software development practice (human-driven decomposition and static code mainte...
We implement a two-stage processing architecture separating document-level extraction (Stage 1) from claim-level synthesis (Stage 2).
Implementation description in paper: architecture design and pipeline stages described by the authors.
high null result Leveraging LLMs for Unstructured Claims Data Analysis system architecture (document-level vs claim-level processing)
In neither unit did internal control mechanisms identify any information-security incident, sensitive-data leakage, or formal compliance challenge from external oversight bodies during the period examined.
Author reports absence of recorded incidents in internal control mechanisms and no external oversight challenges for both units over the study period; based on internal records and SEI-GDF auditable indicators.
high null result The Main Barrier to AI Adoption in the Public Sector is Lack... information-security incidents / sensitive-data leakage / formal compliance chal...
The research is grounded in the Resource-Based View (RBV) and Dynamic Capabilities Theory (DCT) to explain how technological and managerial resources contribute to organizational performance.
Author statement in the paper describing the theoretical framework (RBV and DCT) used to frame the study.
The study adopts a quantitative research design and analyzes collected data using Partial Least Squares Structural Equation Modeling (PLS-SEM).
Author statement in the paper describing research design and analytical method.
Digital Leadership did not demonstrate a statistically significant direct effect on Employee Productivity (β = -0.094, p = 0.275).
Reported quantitative result from the study using PLS-SEM; β and p-value provided in the paper showing a non-significant direct effect. Sample size not reported in the excerpt.
We scored over 2.1 million twin responses on 500 participants and 183 held-out questions.
Reported evaluation counts in the paper: 2.1M responses, 500 participants, 183 held-out questions.
high null result Synthetic Personalities: How Well Can LLMs Mimic Individual ... number of evaluated twin responses / evaluation scale
The construction-method grid covers three open-weight LLMs, five cumulative information depths ranked by normalized Shannon entropy, two embedding methods, and two reasoning modes.
Paper's experimental design specification (methods section).
high null result Synthetic Personalities: How Well Can LLMs Mimic Individual ... experimental factorization of model types, information depths, embedding methods...
We construct detailed individual-level twins from the German Socio-Economic Panel (SOEP) and evaluate them across a 3 × 5 × 2 × 2 construction-method grid.
Methodological description of the study: experimental construction and evaluation on SOEP data.
high null result Synthetic Personalities: How Well Can LLMs Mimic Individual ... feasibility of constructing and evaluating detailed individual-level twins from ...
A large-scale empirical study on Harvey LAB used 12,510 agent trajectories.
Paper states an empirical study run on Harvey LAB with a sample described as 12,510 agent trajectories.
high null result Parthenon Law: A Self-Evolving Legal-Agent Framework agent trajectories (dataset size)
The paper analyzes multiple dimensions of scientific creativity and impact, specifically recombinant novelty, object novelty, 3-year short-run citation impact, and 10-year long-run citation impact.
Methodological description in paper listing the specific dependent variables and time horizons used to measure novelty and impact.
high null result Does Artificial Intelligence Advance Science? measures used (recombinant novelty, object novelty, 3-year citations, 10-year ci...
The analysis draws on over one million publications from OpenAlex.
Descriptive statement in paper specifying dataset source (OpenAlex) and sample size of publications used for analysis.
high null result Does Artificial Intelligence Advance Science? sample of publications (dataset size)
Experts rated 24 AI risks on harm probability and severity, sector and actor vulnerability, actor responsibility, and overall concern.
Study design described in paper: set of 24 defined AI risks rated across several dimensions by Delphi panel participants (n=272).
high null result Prioritization of Risks from Artificial Intelligence: A Delp... risk ratings across multiple dimensions (probability, severity, vulnerability, r...
We conducted a three-round Delphi study conducted late 2025 with 272 international AI experts.
Methodological description in the paper: three-round Delphi study, timing reported as late 2025, sample size reported as 272 international AI experts.
high null result Prioritization of Risks from Artificial Intelligence: A Delp... study_participation / sample characterization
Total (aggregate) unemployment is statistically insignificant in explaining sustainable development, indicating aggregate measures mask critical distributional differences across skill groups.
ARDL estimation results reported in the paper showing an insignificant coefficient for total unemployment; discussion emphasizing distributional masking.
high null result Artificial Intelligence, Disaggregated Unemployment, And Sus... sustainable development (effect of total unemployment)
The empirical analysis is based on panel data of new energy vehicle firms in the Yangtze River Delta from 2001 to 2023.
Dataset description provided in the paper's abstract/introduction indicating the time span and regional coverage.
R&D expenditure does not constitute a significant mediating channel between artificial intelligence and firms' new quality productive forces.
Mediation analysis using the panel data and constructed indicators; reported nonsignificant mediation effect of R&D expenditure (no sample size or statistics reported in excerpt).
high null result Mechanisms and Effects of Artificial Intelligence on New Qua... new quality productive forces (mediating role of R&D expenditure)
The system was evaluated on OMH-Polyglot, a multilingual coding benchmark spanning Turkish, Arabic, Chinese, and code-switched specifications.
Experimental evaluation reported in the paper using the OMH-Polyglot benchmark.
high null result Cross-Lingual Token Arbitrage: Optimizing Code Agent Context... benchmark evaluation on OMH-Polyglot (coverage of languages and code-switched sp...
The study developed a manufacturing value chain resilience (MVCR) index system based on three dimensions: Readiness, Response, and Recovery, using the CSMAR database.
Methodological description: construction of MVCR index using CSMAR microdata and a three-dimension framework (Readiness, Response, Recovery).
high null result Industrial Robot Application and the Manufacturing Value Cha... manufacturing value chain resilience (MVCR) index