The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests About 🎲 Workforce Futures
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Evidence (8807 claims)

Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.

The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).

Browse by theme

Nine broad, paper-level topics. Click one to filter the claims below.

Adoption
9875 claims
Filter claims →
Productivity
8807 claims
Filtered →
Governance
7870 claims
Filter claims →
Human-AI Collaboration
7560 claims
Filter claims →
Org Design
4892 claims
Filter claims →
Innovation
4781 claims
Filter claims →
Labor Markets
4004 claims
Filter claims →
Skills & Training
3308 claims
Filter claims →
Inequality
2332 claims
Filter claims →

Claims by outcome category

Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.

Outcome Positive Negative Mixed Null Total
Other 870 233 116 1066 2363
Governance & Regulation 976 451 218 133 1809
Organizational Efficiency 949 224 144 88 1416
Technology Adoption Rate 764 287 141 122 1325
Research Productivity 501 152 74 362 1101
Output Quality 542 216 69 69 896
Decision Quality 387 198 94 54 740
Firm Productivity 513 67 101 27 714
AI Safety & Ethics 249 303 73 36 667
Market Structure 190 192 134 27 548
Task Allocation 243 77 91 36 452
Innovation Output 291 33 55 20 401
Skill Acquisition 206 72 65 21 364
Employment Level 133 63 115 22 335
Fiscal & Macroeconomic 153 79 52 32 323
Task Completion Time 206 37 12 15 272
Firm Revenue 179 52 29 5 266
Consumer Welfare 130 76 47 13 266
Inequality Measures 48 137 51 6 242
Worker Satisfaction 101 81 25 13 220
Error Rate 84 110 11 5 210
Wages & Compensation 98 47 30 10 185
Regulatory Compliance 88 73 17 7 185
Automation Exposure 66 64 33 16 182
Team Performance 105 29 30 11 176
Training Effectiveness 109 22 14 21 168
Developer Productivity 114 21 14 8 158
Job Displacement 12 90 24 1 127
Hiring & Recruitment 57 9 9 5 80
Skill Obsolescence 6 56 9 1 72
Social Protection 43 17 8 2 70
Creative Output 35 21 9 4 70
Labor Share of Income 18 21 17 1 57
Worker Turnover 15 16 4 35
Industry 1 1
Clear
Productivity Remove filter
Ninety independent agent runs built the same application from one detailed specification, each scored on a fixed 14-criterion functional rubric (42 point maximum) and a visual quality review.
Experimental study described in the paper: 90 independent agent runs, single specification, evaluated on a 14-criterion rubric (42-point max) plus visual quality review.
high null result Reasoning effort, not tool access, buys first-try reliabilit... functional score (14-criterion rubric) and visual quality review
Merge and revert rates held steady.
Analysis of merge rates and revert rates over the study period showing no meaningful change after adoption/mandate.
high null result AI Writes Faster Than Humans Can Review: A Longitudinal Stud... merge rate and revert rate on pull requests
The analysis panel comprises 802 developers and 196,212 pull requests spanning January 2024–April 2026.
Data description in paper: longitudinal panel of developers and PRs with specified date range and counts.
high null result AI Writes Faster Than Humans Can Review: A Longitudinal Stud... sample size and time coverage
The study relies on secondary evidence from the U.S. Census Bureau, U.S. Bureau of Labor Statistics, OECD, IMF, Stanford AI Index, McKinsey Global Institute, NBER, and recent experimental research published from 2020 onward.
Explicit methodological statement in the paper.
high null result Effect of Artificial Intelligence Adoption on Labour Product... data sources and evidence base
Wages of labor that is used only in final goods production and is not displaced by AI increase in line with overall GDP.
Analytical economic model result indicating proportional wage growth for final-goods-only labor relative to GDP. No empirical sample reported.
high null result The Economic Benefits and Costs of AI and Policies to Mitiga... wages of final-goods-only labor
Traditional approaches had 75% accuracy in risk prediction (as reported in the paper).
Observational comparison reported in the study between AI modelling and traditional approaches; reported accuracy for traditional approaches provided but no sample size or methodological details in the supplied text.
high null result Leveraging AI-Driven Decision-Making to Enhance Corporate Go... risk prediction accuracy (traditional methods)
The AI premium is not present for loadings on casual use or open-weight (open-source) model use.
Decomposition analysis showing null or weaker relation between AI premium and loadings on casual usage metrics and open-weight/open-source model consumption.
high null result AI Premium AI premium related to casual or open-weight model use
We construct a high-frequency AI Factor from growth in tokens, dollars, and users, estimate firm-level AI Betas from stock return comovement, and characterize the AI Premium.
Methodological claim based on constructing a factor (AI Factor) using metrics of tokens, dollars, and users; estimating firm-level betas via stock return comovement.
high null result AI Premium AI Factor, firm-level AI Betas, characterization of AI Premium
The analysis uses 380 trillion tokens of realized AI consumption across more than four hundred large language models from the licensed proprietary OpenRouter dataset covering approximately 2 percent of current global monthly AI token consumption.
Descriptive statement about the dataset used: OpenRouter licensed proprietary dataset; 380 trillion tokens; >400 LLMs; coverage ≈2% of global monthly AI token consumption.
high null result AI Premium scale and coverage of AI consumption data (tokens, model count, % global consump...
New generation panel data methods were applied, taking into account cross-sectional dependence and heterogeneity across countries.
Methodological description in the paper indicating use of advanced panel techniques that account for cross-sectional dependence (common shocks/spillovers) and heterogeneity (institutional/structural differences).
high null result AI Readiness, Renewable Energy, and Industrial Development: ... methodological approach (panel estimation accounting for cross-sectional depende...
This study examines the determinants of economic growth in the 27 countries with the highest GDP for the period 2008–2020.
Study sample and period explicitly stated in the paper: the 27 highest-GDP countries, years 2008–2020.
high null result AI Readiness, Renewable Energy, and Industrial Development: ... economic growth (country-level)
This study systematically reviewed 194 peer-reviewed articles published between 2011 and 2025.
Statement in the paper's abstract describing a systematic review of 194 peer-reviewed articles (2011–2025).
high null result Artificial Intelligence and Economic Development: A Systemat... number_of_studies_reviewed
Comments detected as likely to be generated by LLMs remain relatively stable over time.
Detector-based proxy analysis on repository comments across 2021–2025.
high null result An Exploratory Study on LLM-Generated Code and Comments in C... proportion of comments detected as likely LLM-generated over time
We analyze more than 930,000 agent-authored pull requests.
Descriptive statement about the dataset used for the study: an analysis of >930,000 pull requests authored by autonomous coding agents.
high null result Govern the Repository, Not the Agent: Measuring Ecosystem-Le... number of agent-authored pull requests analyzed
Inequality-adjusted income measures are themselves not new.
Literature/contextual claim made by the authors (statement that prior measures exist; no new empirical evidence required).
high null result GAGI: A Gini-Adjusted GDP-per-Capita Index for Distribution-... existence of prior inequality-adjusted income measures
The empirical analysis uses data on industrial robots and Chinese listed companies covering the years 2006–2019.
Data description reported in the paper: matched dataset of industrial robot measures and observations for Chinese listed firms spanning 2006 to 2019.
high null result The application of industrial robots, capital distortion, an... data/sample coverage (industrial robots and Chinese listed companies, 2006–2019)
We hired 49 programmers to interact with GitHub Copilot to assess 148 HIPAA-derived NFRs against the iTrust codebase across three dimensions: requirement satisfaction level, reasoning, and code localization.
Study design reported in the paper: recruitment of 49 programmers, 148 NFR assessments, use of GitHub Copilot and iTrust codebase, and three specified assessment dimensions.
high null result Accuracy and Satisfaction in Multi-Turn LLM Dialogues for NF... study sample and experimental setup
Evaluating how well LLM-based dialogue systems support collaborative reasoning about NFRs requires methods that go beyond single-turn accuracy to capture both the correctness of system outputs and the quality of multi-turn interaction.
Methodological argument presented by the authors to motivate multi-turn study design (conceptual / methodological claim).
high null result Accuracy and Satisfaction in Multi-Turn LLM Dialogues for NF... evaluation methodology adequacy for NFR assessment
Non-Functional Requirements (NFRs) are inherently vague, context-dependent, and involve many parts of a program, making them difficult to assess with single-turn correctness benchmarks.
Conceptual claim motivating the study, based on properties of NFRs discussed in the paper (no empirical measurement reported for this claim).
high null result Accuracy and Satisfaction in Multi-Turn LLM Dialogues for NF... characteristics of NFRs (vagueness, context-dependence, broad code impact)
LLM-based dialogue assistants have become mainstream tools for software developers, yet current evaluation benchmarks focus exclusively on functional correctness.
Positioning statement in the paper's introduction / literature overview (no new empirical data reported for this claim).
high null result Accuracy and Satisfaction in Multi-Turn LLM Dialogues for NF... scope of evaluation benchmarks (functional correctness focus)
We study 21 LLMs across 16 widely used benchmarks spanning coding, reasoning, medicine, factuality, instruction following, and agentic tasks.
Description of empirical study scope in the paper (listed number of models and benchmarks and task domains).
high null result The Capability Frontier: Benchmarks Miss 82% of Model Perfor... study/sample scope (number of models and benchmarks)
Using panel data for 30 Chinese provinces from 2016 to 2024, this study constructs a marginal cost-based indicator of agricultural pollution–carbon reduction synergy (APCRS).
Methodological description in the paper: panel data covering 30 Chinese provinces for 2016–2024 and construction of a marginal cost-based APCRS indicator.
high null result How Does Artificial Intelligence Industry Agglomeration Affe... construction of APCRS (agricultural pollution–carbon reduction synergy) indicato...
A Clopper–Pearson bound on beta gives a finite-sample certificate on the largest gain any router, vote, or cascade could deliver before training a router.
Application of the Clopper–Pearson (binomial) confidence interval to estimate a finite-sample bound on beta; methodological proposal demonstrated in the paper.
high null result When Does Combining Language Models Help? A Co-Failure Ceili... bound on maximum possible ensemble gain (via beta)
Average pairwise error correlation rho cannot identify beta: error laws with identical marginals and pairwise correlations can have different all-wrong rates.
Theoretical/statistical construction and argument in the paper showing different joint error distributions can have same marginals and pairwise correlations but different all-wrong (beta) probabilities.
The study uses panel data covering 30 Chinese provinces from 2010 to 2022 to analyze AI's impact on value chain upgrading.
Stated dataset description in the paper (30 provinces, 2010–2022); used for the econometric analyses reported.
high null result The impact of artificial intelligence on value chain upgradi... dataset/sample description
The study contrasts usage across three populations: external personal-account users, external organizational-account users, and workers within OpenAI using an automated, privacy-protecting pipeline.
Methodological description in the paper stating the populations compared and the privacy-preserving data pipeline used.
high null result The Shift to Agentic AI: Evidence from Codex methodological scope (populations compared)
The position taken in the paper is platform-agnostic.
Author statement in the paper explicitly describing the position as platform-agnostic.
high null result Grounded Scaling: Why Agentic AI Needs Deterministic Environ... applicability across AI platforms
This study used survey data from 426 AI-adopting Chinese manufacturing firms and analyzed hypothesized relationships using hierarchical regression to isolate moderating effects of supply chain integration.
Methodological statement reported in the paper.
high null result Aligning customer and supplier integration for AI-enabled su... study methodology (survey + hierarchical regression)
The benchmark includes controlled evaluation settings for local improvement, cross-task transfer, cross-role transfer, and cross-model generalization.
Paper description of benchmark design and experimental protocol specifying controlled evaluation settings for local improvement, cross-task transfer, cross-role transfer, and cross-model generalization.
high null result Managing Procedural Memory in LLM Agents: Control, Adaptatio... availability of controlled evaluation settings
We introduce AFTER, a benchmark of 382 realistic enterprise tasks spanning six professional roles and 22 procedural skills.
Description of benchmark dataset introduced in the paper; the paper reports the benchmark contains 382 tasks, covers six professional roles and 22 procedural skills (dataset construction / annotation process described in methods).
high null result Managing Procedural Memory in LLM Agents: Control, Adaptatio... benchmark size and coverage (number of tasks, roles, skills)
When restricted to multi-commit PRs, the Copilot within-repo effect dissolves to +4.8 percentage points (p = 0.59).
Subset analysis limited to multi-commit PRs for Copilot in the AIDev dataset; reported point estimate and p-value.
high null result Beyond Simpson's Paradox: A Cascade of Confounders in AI Age... PR merge rate (Copilot co-authorship effect in multi-commit PR subset)
Within-repository controls eliminate Devin's co-authorship gap, reducing it from +33.5 percentage points to +1.6 pp (p = 0.73).
Within-repo controlled regression/analysis on Devin PRs in the AIDev dataset showing adjusted effect and p-value.
high null result Beyond Simpson's Paradox: A Cascade of Confounders in AI Age... PR merge rate (Devin co-authorship effect before and after within-repo control)
The argument that automation is leading to a general decline in employment opportunities is not supported by actual facts and trends; rather, it is a product of the pervasive influence of technological fetishism.
Author's evaluative conclusion based on the reviewed theoretical and empirical studies; the excerpt provides no specific datasets or statistical tests supporting this rebuttal.
high null result New Technologies and Increase in Employment general trend in employment opportunities in relation to automation
A conceptual framework is developed showing how digital infrastructure and institutional support mediate sectoral transformation.
Paper presents a conceptual framework (theoretical/modeling component) derived from empirical findings and policy analysis; this is descriptive rather than a quantified empirical result.
high null result How to Utilize New Technologies to Improve Productivity role of digital infrastructure and institutions in mediating transformation
We conclude by outlining implications for designing and evaluating human-AI teams as socio-technical systems and for prioritizing longitudinal and in-context studies that capture how teaming evolves over time.
Authors' conclusions and recommendations based on the systematic review and observed gaps in the literature (noted need for longitudinal, in-context studies).
high null result From testbeds to high-stakes work: a review of Human-AI team... research_design_priorities (longitudinal and in-context evaluation)
Bibliometric patterns suggest a shift since 2020 from foundational demonstrations in controlled settings toward applied, higher-stakes contexts where trust dynamics, communication, and ethical accountability more directly shape adoption and sustained performance.
Bibliometric analysis of the 104 studies showing temporal trends (pre- vs post-2020) in research contexts and topics.
high null result From testbeds to high-stakes work: a review of Human-AI team... temporal_shift_in_research_contexts (prevalence of applied/higher-stakes context...
Across studies, performance was the most frequently examined aspect, followed by trust, explainability and transparency, decision-making, and team processes.
Synthesis and frequency coding of outcomes/measured constructs across the 104 included empirical studies.
high null result From testbeds to high-stakes work: a review of Human-AI team... performance (and ranked prevalence of constructs like trust, explainability, dec...
Gaming and entertainment, aviation, military and defense operations, emergency response and public safety, and healthcare also represented substantial portions of the literature.
Domain breakdown from the systematic review of 104 empirical studies (frequency counts by domain reported in Results).
high null result From testbeds to high-stakes work: a review of Human-AI team... study_domain_representation_by_industry
Cross-domain and interdisciplinary studies were the largest category, representing broad workplace or team-based investigations not tied to a single industry and instead focused on general collaboration issues such as communication, teamwork, coordination, and coworker interaction.
Categorization / coding of the 104 included empirical studies; frequency counts by study domain reported in review.
high null result From testbeds to high-stakes work: a review of Human-AI team... study_domain_prevalence
We conducted a PRISMA-guided systematic review with bibliometric analysis of 104 peer-reviewed empirical studies published between 2015 and 2025 and identified through Engineering Village, IEEE Xplore, PubMed, ScienceDirect, and Web of Science.
Methods reported in paper: PRISMA-guided systematic review and bibliometric analysis; explicit statement of 104 peer-reviewed empirical studies and databases searched (Engineering Village, IEEE Xplore, PubMed, ScienceDirect, Web of Science).
high null result From testbeds to high-stakes work: a review of Human-AI team... number_of_studies_reviewed
Order, entropy, information, and useful energy are task-dependent and system-relative concepts whose meanings depend on the objectives of the system.
Conceptual argument and discussion in the paper about the context-dependence of informational and energetic notions within the proposed framework; no empirical evidence provided.
The paper's contribution is theoretical: it reframes the AI productivity debate beyond automation anxiety by linking technological change, income distribution and effective demand in a single analytical framework.
Author-stated contribution and framing in the conceptual review (description of scope and aim).
high null result Artificial Intelligence, Labour Income and Effective Demand:... conceptual framing of the AI productivity debate (qualitative contribution)
The review proposes possible indicators for future empirical research, including the productivity–real labour income gap and an absorption tension indicator.
Paper's methodological/propositional content (explicitly proposes indicators for empirical work).
high null result Artificial Intelligence, Labour Income and Effective Demand:... productivity–real labour income gap (indicator of distributive transmission) and...
The paper defines the 'Distributional Absorption Threshold of AI-Induced Productivity' as the point beyond which productivity gains associated with AI are no longer accompanied by proportionate increases in broadly distributed real purchasing power and household consumption.
Textual/definitional content of the conceptual review (the paper introduces and defines this concept).
high null result Artificial Intelligence, Labour Income and Effective Demand:... threshold at which productivity gains cease to be matched by distributed purchas...
We study shared-workspace human-AI teams using the Collaborative Gym environment with DiscoveryBench tasks.
Methodological description in the paper stating the experimental environment and task suite used.
high null result Searching for Synergy in Shared Workspace Human-AI Collabora... experimental platform and task set used
We ran 1,482 sessions in our experiments.
Statement in the paper reporting total experimental sessions (Collaborative Gym + DiscoveryBench).
high null result Searching for Synergy in Shared Workspace Human-AI Collabora... number of experimental sessions
Experimental results were validated through bootstrap confidence intervals, multiple-testing corrections, and subperiod stability analysis.
Methods statement in the paper indicating use of bootstrap confidence intervals, multiple-testing corrections, and subperiod stability analysis to validate backtest results.
high null result Toward Expert Investment Teams: A Multi-Agent LLM System wit... statistical robustness of experimental results
The article presents a comprehensive framework spanning macro-level governance principles and micro-level interaction typologies, illustrated through case examples from telecommunications, retail, insurance, and consumer products sectors.
Descriptive/methodological claim in the paper indicating the presence of a framework and sectoral case examples (no quantitative sample sizes reported for cases).
high null result Designing Human-Machine Collaboration: Strategic Imperatives... presence of a multi-level framework and illustrative case examples across specif...
The article draws on Deloitte's 2026 Global Human Capital Trends survey of over 3,000 business leaders across 15 countries.
Methodological/data source statement explicitly provided in the paper.
high null result Designing Human-Machine Collaboration: Strategic Imperatives... survey scope (sample size and country coverage)
The robot is initialized with the selected episodic memory before a new collaboration episode begins (i.e., memory reuse occurs at episode start).
Design and procedure reported in paper: authors state that the robot is initialized with the selected memory prior to new episodes as part of the intervention tested in the experiment.
high null result Improving Human-Robot Teamwork in Urban Search and Rescue Th... initialization procedure (methodological step)