The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests About 🎲 Workforce Futures
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Evidence (8974 claims)

Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.

The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).

Browse by theme

Nine broad, paper-level topics. Click one to filter the claims below.

Adoption
10085 claims
Filter claims →
Productivity
8974 claims
Filtered →
Governance
8062 claims
Filter claims →
Human-AI Collaboration
7749 claims
Filter claims →
Org Design
5057 claims
Filter claims →
Innovation
4896 claims
Filter claims →
Labor Markets
4088 claims
Filter claims →
Skills & Training
3372 claims
Filter claims →
Inequality
2377 claims
Filter claims →

Claims by outcome category

Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.

Outcome Positive Negative Mixed Null Total
Other 882 244 117 1097 2424
Governance & Regulation 1010 469 229 135 1875
Organizational Efficiency 977 235 149 90 1462
Technology Adoption Rate 781 299 143 128 1362
Research Productivity 506 155 74 363 1110
Output Quality 555 219 71 70 915
Decision Quality 395 200 95 54 751
Firm Productivity 523 67 101 27 724
AI Safety & Ethics 262 309 75 36 688
Market Structure 195 201 135 30 566
Task Allocation 248 77 96 38 464
Innovation Output 300 34 55 20 411
Skill Acquisition 207 75 65 21 368
Employment Level 138 67 119 24 350
Fiscal & Macroeconomic 156 80 53 33 329
Task Completion Time 211 38 13 16 280
Firm Revenue 183 52 29 5 270
Consumer Welfare 131 77 48 13 269
Inequality Measures 50 141 54 9 254
Worker Satisfaction 104 85 25 13 227
Error Rate 87 112 11 5 215
Automation Exposure 69 69 37 20 198
Wages & Compensation 102 49 31 11 193
Team Performance 115 30 30 11 187
Regulatory Compliance 88 74 17 7 186
Training Effectiveness 109 22 14 21 168
Developer Productivity 116 21 15 8 161
Job Displacement 12 92 26 1 131
Hiring & Recruitment 57 12 9 5 83
Skill Obsolescence 6 59 10 2 77
Social Protection 43 17 8 2 70
Creative Output 35 21 9 4 70
Labor Share of Income 18 23 17 1 59
Worker Turnover 15 16 4 35
Industry 1 1
Clear
Productivity Remove filter
Raw blind-panel decision quality is similar for A and B (7.01 vs. 6.96).
Blind-panel scoring of generated reports from agents A and B; panel size and panel methodology not specified in abstract.
high null result AI Scientists Are Only as Good as Their Evidence: A Stratifi... raw blind-panel decision-quality score
Self-evaluated creative performance remained unchanged when using GenAI.
Same experiment with 82 participants; authors report no significant difference in self-evaluated creative performance between GenAI users and controls.
high null result When Ai Sparks Less: Generative Ai And The Decline Of Self-P... self-evaluated creative performance
Each of the four published papers used in the experiments contained an error that I helped identify or correct.
Author statement that the 4 papers each contained an error; author involvement in identification/correction is asserted.
high null result Can AI Refute Economic Theory? Evidence from Beyond the Know... presence of errors in the 4 target papers
I conducted experiments in which I asked several AI models (Gemini, Refine, Claude, and ChatGPT) to check the correctness of four published papers in economic theory.
Author reports running direct experiments: prompted listed models to check 4 published economic-theory papers.
high null result Can AI Refute Economic Theory? Evidence from Beyond the Know... existence of experiments using specified models on 4 papers
We evaluate the system on operator feedback and a question set collected from production usage, graded by human and automated panels.
Paper's stated evaluation methodology: operator feedback + production question set, graded by humans and automated panels.
high null result Archi: Agentic Operations at the CMS Experiment evaluation methodology (feedback and graded question set)
This study benchmarks Algeria’s readiness to adopt AI against Morocco, Egypt, and Turkey using data from the World Bank (2022), the Oxford Insights Government AI Readiness Index, and sector-specific studies.
Methodological statement in the paper specifying data sources used for the comparative assessment (World Bank 2022, Oxford Insights index, sector studies).
high null result Artificial Intelligence and Economic Productivity: A Compara... AI readiness / readiness indicators
The article aims to provide systematic literature support for subsequent research and adaptive policy formulation.
Statement of the paper's stated objective; methodological and policy-intent claim from the authors.
high null result Influence of Artificial Intelligence in the Labor Market policy formulation support
This article is based on a systematic literature review and summarizes the four core theoretical mechanisms of substitution, complementarity, new task creation, and skill mismatch.
Methodological claim from the paper: the authors conducted a systematic literature review and identified these four theoretical mechanisms.
high null result Influence of Artificial Intelligence in the Labor Market theoretical mechanisms
Traditional software and agentic systems are distinct: in traditional software code is the carrier of decision logic, whereas in agentic systems code is ephemeral tooling used by an LLM-driven reasoning loop.
Formalization and conceptual definitions developed in the paper (first-principles formal distinction; no empirical sample size reported).
high null result The End of Software Engineering: How AI Agents Are Fundament... architectural role of code (carrier of logic vs ephemeral tool)
For over half a century, software engineering has operated on a foundational premise: human engineers decompose problems, encode decision logic into static code, and manually adapt that code as requirements evolve.
Historical/descriptive claim presented in the paper's framing and literature review; citation of longstanding software engineering practices (qualitative, no empirical sample size reported).
high null result The End of Software Engineering: How AI Agents Are Fundament... software development practice (human-driven decomposition and static code mainte...
We implement a two-stage processing architecture separating document-level extraction (Stage 1) from claim-level synthesis (Stage 2).
Implementation description in paper: architecture design and pipeline stages described by the authors.
high null result Leveraging LLMs for Unstructured Claims Data Analysis system architecture (document-level vs claim-level processing)
In neither unit did internal control mechanisms identify any information-security incident, sensitive-data leakage, or formal compliance challenge from external oversight bodies during the period examined.
Author reports absence of recorded incidents in internal control mechanisms and no external oversight challenges for both units over the study period; based on internal records and SEI-GDF auditable indicators.
high null result The Main Barrier to AI Adoption in the Public Sector is Lack... information-security incidents / sensitive-data leakage / formal compliance chal...
The research is grounded in the Resource-Based View (RBV) and Dynamic Capabilities Theory (DCT) to explain how technological and managerial resources contribute to organizational performance.
Author statement in the paper describing the theoretical framework (RBV and DCT) used to frame the study.
The study adopts a quantitative research design and analyzes collected data using Partial Least Squares Structural Equation Modeling (PLS-SEM).
Author statement in the paper describing research design and analytical method.
Digital Leadership did not demonstrate a statistically significant direct effect on Employee Productivity (β = -0.094, p = 0.275).
Reported quantitative result from the study using PLS-SEM; β and p-value provided in the paper showing a non-significant direct effect. Sample size not reported in the excerpt.
We scored over 2.1 million twin responses on 500 participants and 183 held-out questions.
Reported evaluation counts in the paper: 2.1M responses, 500 participants, 183 held-out questions.
high null result Synthetic Personalities: How Well Can LLMs Mimic Individual ... number of evaluated twin responses / evaluation scale
The construction-method grid covers three open-weight LLMs, five cumulative information depths ranked by normalized Shannon entropy, two embedding methods, and two reasoning modes.
Paper's experimental design specification (methods section).
high null result Synthetic Personalities: How Well Can LLMs Mimic Individual ... experimental factorization of model types, information depths, embedding methods...
We construct detailed individual-level twins from the German Socio-Economic Panel (SOEP) and evaluate them across a 3 × 5 × 2 × 2 construction-method grid.
Methodological description of the study: experimental construction and evaluation on SOEP data.
high null result Synthetic Personalities: How Well Can LLMs Mimic Individual ... feasibility of constructing and evaluating detailed individual-level twins from ...
A large-scale empirical study on Harvey LAB used 12,510 agent trajectories.
Paper states an empirical study run on Harvey LAB with a sample described as 12,510 agent trajectories.
high null result Parthenon Law: A Self-Evolving Legal-Agent Framework agent trajectories (dataset size)
The paper analyzes multiple dimensions of scientific creativity and impact, specifically recombinant novelty, object novelty, 3-year short-run citation impact, and 10-year long-run citation impact.
Methodological description in paper listing the specific dependent variables and time horizons used to measure novelty and impact.
high null result Does Artificial Intelligence Advance Science? measures used (recombinant novelty, object novelty, 3-year citations, 10-year ci...
The analysis draws on over one million publications from OpenAlex.
Descriptive statement in paper specifying dataset source (OpenAlex) and sample size of publications used for analysis.
high null result Does Artificial Intelligence Advance Science? sample of publications (dataset size)
Experts rated 24 AI risks on harm probability and severity, sector and actor vulnerability, actor responsibility, and overall concern.
Study design described in paper: set of 24 defined AI risks rated across several dimensions by Delphi panel participants (n=272).
high null result Prioritization of Risks from Artificial Intelligence: A Delp... risk ratings across multiple dimensions (probability, severity, vulnerability, r...
We conducted a three-round Delphi study conducted late 2025 with 272 international AI experts.
Methodological description in the paper: three-round Delphi study, timing reported as late 2025, sample size reported as 272 international AI experts.
high null result Prioritization of Risks from Artificial Intelligence: A Delp... study_participation / sample characterization
Total (aggregate) unemployment is statistically insignificant in explaining sustainable development, indicating aggregate measures mask critical distributional differences across skill groups.
ARDL estimation results reported in the paper showing an insignificant coefficient for total unemployment; discussion emphasizing distributional masking.
high null result Artificial Intelligence, Disaggregated Unemployment, And Sus... sustainable development (effect of total unemployment)
The empirical analysis is based on panel data of new energy vehicle firms in the Yangtze River Delta from 2001 to 2023.
Dataset description provided in the paper's abstract/introduction indicating the time span and regional coverage.
R&D expenditure does not constitute a significant mediating channel between artificial intelligence and firms' new quality productive forces.
Mediation analysis using the panel data and constructed indicators; reported nonsignificant mediation effect of R&D expenditure (no sample size or statistics reported in excerpt).
high null result Mechanisms and Effects of Artificial Intelligence on New Qua... new quality productive forces (mediating role of R&D expenditure)
The system was evaluated on OMH-Polyglot, a multilingual coding benchmark spanning Turkish, Arabic, Chinese, and code-switched specifications.
Experimental evaluation reported in the paper using the OMH-Polyglot benchmark.
high null result Cross-Lingual Token Arbitrage: Optimizing Code Agent Context... benchmark evaluation on OMH-Polyglot (coverage of languages and code-switched sp...
The study developed a manufacturing value chain resilience (MVCR) index system based on three dimensions: Readiness, Response, and Recovery, using the CSMAR database.
Methodological description: construction of MVCR index using CSMAR microdata and a three-dimension framework (Readiness, Response, Recovery).
high null result Industrial Robot Application and the Manufacturing Value Cha... manufacturing value chain resilience (MVCR) index
The study constructed indices of industrial robot application at the enterprise-industry-year level by matching industry-level industrial robot data published by the IFR with microdata from Chinese A-share listed companies.
Methodological description in the paper: matching IFR industry-level industrial robot data to microdata from Chinese A-share listed firms to build enterprise-industry-year robot-application indices.
high null result Industrial Robot Application and the Manufacturing Value Cha... index of industrial robot application (enterprise-industry-year)
The experiment was run twice: a first run with unrealistically loud injections, and a second run with signals rescaled to a physically motivated SNR range.
Protocol described in paper explicitly states two runs with different injection SNR scalings (one 'unrealistically loud', one physically motivated).
Both agents received identical written specifications and identical compute resources.
Methodological statement in paper specifying that both agents were given the same written spec and the same shared computing infrastructure.
The pipeline comprised power spectral density estimation from raw Einstein Telescope simulated noise, geometric template bank generation, matched filter recovery of 100 binary black hole signal injections, automated results generation, and large language model-assisted production of a manuscript formatted in the style of Physical Review D.
Protocol description in paper; matched filter recovery included 100 injected signals (explicitly stated).
We compared two state-of-the-art agentic AI systems, Claude Code (Anthropic) and Codex (OpenAI), tasked with autonomously executing a simple end-to-end gravitational wave data analysis pipeline on a shared computing infrastructure without human intervention.
Experimental design described in paper: two named agents were given identical written specifications and identical compute resources and executed the full pipeline autonomously.
Greater frontier-level compute does not consistently translate to better performance.
Empirical observation in the paper's findings: increasing compute capacity at the Pareto frontier did not uniformly improve task performance across evaluated tasks.
high null result When Cloud Agents Meet Device Agents: Lessons from Hybrid Mu... task performance as a function of available compute at the frontier
The distinction matters: debt is a stock of design and governance liability, while the tax is a flow of operating cost that arises because stochastic agents act through tools and workflows.
Conceptual argument in the paper articulating difference between two defined concepts (Agentic Technical Debt vs Stochastic Tax); no empirical demonstration.
high null result Governing Technical Debt in Agentic AI Systems conceptual distinction between liability (stock) and operating cost (flow)
Stochastic Tax is the recurring operating burden of keeping probabilistic agent behavior within acceptable bounds.
Paper provides a formal definition / conceptual framing of 'Stochastic Tax'; stated as an operational concept (no empirical quantification provided).
high null result Governing Technical Debt in Agentic AI Systems operating burden from probabilistic agent behavior
Agentic Technical Debt is the accumulated liability created when prompts, memory, tool schemas, orchestration graphs, control policies, and observability routines are patched together faster than they can be validated, standardized, and governed.
Paper provides a formal definition / conceptual framing of 'Agentic Technical Debt'; presented as a definitional contribution rather than an empirically measured quantity.
high null result Governing Technical Debt in Agentic AI Systems conceptual definition of a technical/governance liability
Agentic AI systems reason over multiple steps, call tools, act through workflows, and adapt through memory and feedback.
Descriptive/definitional statement in the paper; presented as characteristics of agentic systems rather than supported by empirical measurement.
high null result Governing Technical Debt in Agentic AI Systems architectural/behavioral characteristics of agentic AI systems
Agentic AI systems are increasingly being explored as production infrastructure.
Stated as an observation in the paper's introduction/abstract; no empirical data, sample, or formal measurement provided (conceptual/observational claim).
high null result Governing Technical Debt in Agentic AI Systems exploration/adoption of agentic AI as production infrastructure
The paper evaluates the proposed architecture using the outcome metric 'time-to-insight'.
Methodological statement in the paper listing evaluation metrics.
high null result Beyond the Data Mesh Illusion: Designing Modern AI-augmented... time-to-insight (time required to generate actionable insight from data)
The paper evaluates the proposed architecture using the outcome metric 'time-to-find'.
Methodological statement in the paper listing evaluation metrics.
high null result Beyond the Data Mesh Illusion: Designing Modern AI-augmented... time-to-find (time required to locate relevant data/products)
The paper evaluates the proposed architecture using the outcome metric 'data product adoption'.
Methodological statement in the paper listing evaluation metrics.
We distill our findings into a meta-design and four design principles (DPs), grounded in kernel theories, for systems where human contextual intelligence and algorithmic recognition must coexist.
Design contribution presented in the paper (meta-design artifact and four DPs derived from the study).
high null result Schnitzel-Prediction: Designing Human-Ai Collaboration For C... design principles and meta-design artifact
We developed a collaborative forecasting system that leverages semantic processing using large language models (LLMs) to solve the 'cold-start' problem for novel menu items while preserving human agency via override mechanisms.
Description of system design and implementation produced during the ADR project (practice-driven abductive approach).
high null result Schnitzel-Prediction: Designing Human-Ai Collaboration For C... resolution of cold-start forecasting for novel menu items; preservation of human...
This paper reports on a 9-month action design research (ADR) project at a German financial services firm.
Explicit methodological description in the paper (study duration and organizational context).
high null result Schnitzel-Prediction: Designing Human-Ai Collaboration For C... study duration and setting
We examined how different degrees of embodiment affect team performance and conversational dynamics in a real-life escape room; teams were composed of either three humans or two humans and an artificial agent (a Box, an Avatar, or a hyper-realistic humanoid).
Experimental field study reported in the paper: a real-life escape room experiment comparing team compositions (3 humans vs. 2 humans + agent of three embodiment types). Sample size not reported in the provided text.
high null result Teaming Up with Artificial Agents in Non-routine Analytical ... team composition / experimental manipulation (embodiment)
The empirical strategy uses panel local projections to estimate the dynamic effects of AI adoption.
Methodological statement in the paper: application of panel local projections to panel data of industries/establishments over 2017-2025.
high null result AI Adoption and Labor Market Responses: Evidence from Job Po... estimation method / dynamic impulse responses
AI adoption is measured using the share of establishment-level job postings that explicitly require AI-related skills across 13 industries over 2017-2025.
Study design / data description: share of establishment-level job postings requiring AI skills; coverage across 13 industries for years 2017-2025.
high null result AI Adoption and Labor Market Responses: Evidence from Job Po... AI adoption (share of job postings requiring AI skills)
Estimation accuracy depended only weakly on message volume, indicating that more text alone does not guarantee better inference.
Analysis reported in the paper examining the relationship between message volume and estimation accuracy; described as a weak dependency.
high null result Can AI Guess What You Know? Performance Comparison of Large ... relationship between message volume (amount of text) and model estimation accura...
The model introduces the 'Sciencepreneur' as the central human archetype in agentic R&D.
Conceptual/design claim within the HARMONY artifact presented in the paper.
high null result From Replacement to Orchestration: A Socio-Technical Archite... role definition and skill profile for human operators in agentic R&D