The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests About 🎲 Workforce Futures
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Evidence (8807 claims)

Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.

The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).

Browse by theme

Nine broad, paper-level topics. Click one to filter the claims below.

Adoption
9875 claims
Filter claims →
Productivity
8807 claims
Filtered →
Governance
7870 claims
Filter claims →
Human-AI Collaboration
7560 claims
Filter claims →
Org Design
4892 claims
Filter claims →
Innovation
4781 claims
Filter claims →
Labor Markets
4004 claims
Filter claims →
Skills & Training
3308 claims
Filter claims →
Inequality
2332 claims
Filter claims →

Claims by outcome category

Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.

Outcome Positive Negative Mixed Null Total
Other 870 233 116 1066 2363
Governance & Regulation 976 451 218 133 1809
Organizational Efficiency 949 224 144 88 1416
Technology Adoption Rate 764 287 141 122 1325
Research Productivity 501 152 74 362 1101
Output Quality 542 216 69 69 896
Decision Quality 387 198 94 54 740
Firm Productivity 513 67 101 27 714
AI Safety & Ethics 249 303 73 36 667
Market Structure 190 192 134 27 548
Task Allocation 243 77 91 36 452
Innovation Output 291 33 55 20 401
Skill Acquisition 206 72 65 21 364
Employment Level 133 63 115 22 335
Fiscal & Macroeconomic 153 79 52 32 323
Task Completion Time 206 37 12 15 272
Firm Revenue 179 52 29 5 266
Consumer Welfare 130 76 47 13 266
Inequality Measures 48 137 51 6 242
Worker Satisfaction 101 81 25 13 220
Error Rate 84 110 11 5 210
Wages & Compensation 98 47 30 10 185
Regulatory Compliance 88 73 17 7 185
Automation Exposure 66 64 33 16 182
Team Performance 105 29 30 11 176
Training Effectiveness 109 22 14 21 168
Developer Productivity 114 21 14 8 158
Job Displacement 12 90 24 1 127
Hiring & Recruitment 57 9 9 5 80
Skill Obsolescence 6 56 9 1 72
Social Protection 43 17 8 2 70
Creative Output 35 21 9 4 70
Labor Share of Income 18 21 17 1 57
Worker Turnover 15 16 4 35
Industry 1 1
Clear
Productivity Remove filter
We performed an extensive evaluation of 37 state-of-the-art Vision-Language Models on MultihopSpatial.
Empirical evaluation described in the paper listing the number of models evaluated (37).
high neutral MultihopSpatial: Multi-hop Compositional Spatial Reasoning B... benchmark coverage across models evaluated
The paper treats data as a new type of production factor and endogenizes it within the production function.
Theoretical/methodological: the paper constructs a macro-level theoretical model that explicitly includes data as an endogenous input in the production function (no empirical/sample data).
high neutral Study on the impact of big data sharing on individuals’ welf... inclusion of data as a production factor (model specification)
Economic evaluations of GLAI should account for end-to-end risk externalities (error propagation, institutional trust, rights impacts), not only short-term productivity gains.
Methodological recommendation grounded in conceptual synthesis of technical, behavioral, and legal risks; normative argument rather than empirical result.
high neutral Why Avoid Generative Legal AI Systems? Hallucination, Overre... comprehensiveness of economic evaluations (inclusion of externalities vs. narrow...
Generative Legal AI (GLAI) systems are built on token-prediction (LLM) architectures rather than formal legal-reasoning architectures.
Conceptual and technical analysis in the paper distinguishing GLAI from other legal-tech; literature synthesis on common LLM architectures. No original empirical dataset or sample size—qualitative/technical review.
high neutral Why Avoid Generative Legal AI Systems? Hallucination, Overre... underlying model architecture type (token-prediction vs. formal-reasoning)
The paper's formalism shows that prompt/system messages shape distributions over possible execution paths (indirect control) but do not evaluate actual partial paths at runtime.
Formal mapping in the paper that treats prompts as shaping prior over paths; conceptual argument and illustrative examples.
high neutral Runtime Governance for AI Agents: Policies on Paths degree of control over execution path (distributional shaping vs. path-specific ...
Through a thematic review of existing research, the authors identified recurring themes about incentive schemes: their components, how researchers manipulate them, and their impact on research outcomes.
Authors' stated method and findings: thematic review (the scope/number of reviewed papers not specified in excerpt).
high neutral Incentive-Tuning: Understanding and Designing Incentives for... themes in incentive design practices and reported impacts on empirical study out...
A critical aspect of conducting human–AI decision-making studies is the role of participants, often recruited through crowdsourcing platforms.
Claim based on the authors' thematic literature review noting participant sourcing practices (specific studies and counts not given in excerpt).
high neutral Incentive-Tuning: Understanding and Designing Incentives for... participant recruitment source (e.g., crowdsourcing) and its influence on study ...
Researchers conduct empirical studies investigating how humans use AI assistance for decision-making and how this collaboration impacts results.
Statement summarizing the research landscape; supported implicitly by the authors' thematic review of existing empirical studies (number of studies not specified in excerpt).
high neutral Incentive-Tuning: Understanding and Designing Incentives for... human behaviour and decision outcomes when assisted by AI (empirical study outco...
The study provides empirical evidence specific to a small open EU economy (Slovakia) on the relationship between AI adoption and labour productivity.
Use of harmonised Eurostat enterprise and productivity data for Slovakia and EU27 over 2021–2024, analysed with descriptive statistics, gap analysis, dynamics of change, correlation, and an illustrative regression model.
high neutral Artificial Intelligence Adoption and Labour Productivity in ... Empirical characterization of AI adoption and labour productivity relationship f...
Returns to AI are heterogeneous across firms; estimating treatment effects requires attention to selection, complementarities, and dynamic adoption pipelines.
Methodological argument referencing treatment-effect literature and observed firm heterogeneity; supported by conceptual examples rather than a single empirical treatment-effect estimate.
high neutral Modern Management in the Age of Artificial Intelligence: Str... heterogeneity in returns to AI adoption (firm-level productivity or performance ...
Productivity effects at the aggregate (economy-wide) level are delayed relative to firm-level gains.
Cross-study synthesis noting temporal lags between observed firm-level productivity improvements and measurable aggregate effects in the literature included in the SLR.
high null result Artificial Intelligence and the Digital Economy: Impact on E... timing of aggregate productivity effects
The review followed the PRISMA protocol and synthesized 78 peer-reviewed studies and institutional reports published between 2015 and 2025.
Systematic Literature Review using PRISMA protocol; sample of 78 peer-reviewed studies and institutional reports (2015–2025) as described in the paper.
Human-only and AI-assisted teams performed similarly on most outcomes.
Comparison across outcome measures from the randomized experiment; summary statement indicates parity on most measured tasks except for detection of major coding errors.
high null result AI-assisted teams outperform AI-led teams but not human-only... multiple reproduction-related outcomes (overall performance parity)
We randomly assigned 288 researchers to 103 teams working under three conditions (human-only, AI-assisted, AI-led).
Experimental design reported in paper: randomized assignment of 288 researchers into 103 teams across three experimental conditions.
high null result AI-assisted teams outperform AI-led teams but not human-only... experimental_assignment
Green computing capacity is measured by a composite index covering computing infrastructure, green energy support, low-carbon operating efficiency, and computing–network coordination.
Method description in abstract listing the components of the composite green computing capacity index.
high null result Artificial Intelligence and Urban Green Productivity in Chin... green computing capacity (measurement method)
AI technological development is measured by city-level AI patent grants.
Method description in abstract stating AI is proxied by city-level AI patent grants.
high null result Artificial Intelligence and Urban Green Productivity in Chin... AI technological development (measurement method)
Urban green productivity is measured by an undesirable output Super-SBM model.
Method description in abstract specifying the Super-SBM undesirable-output Data Envelopment Analysis model used to compute green productivity.
high null result Artificial Intelligence and Urban Green Productivity in Chin... green productivity (measurement method)
The study uses panel data for 287 Chinese prefecture-level and above cities from 2005 to 2023.
Statement in abstract describing data scope and timeframe; sample count explicitly given as 287 cities and years 2005–2023.
high null result Artificial Intelligence and Urban Green Productivity in Chin... sample_frame / dataset
The study evaluates green productivity (GP) across three dimensions: labourers, means of labour, and objects of labour.
Paper states it measures GP along three dimensions (labourers, means of labour, objects of labour); methodological description in paper; measurement/construct definition rather than empirical test. No sample size reported in the summary.
high null result The Synergistic Effect of Digital Industry Agglomeration and... green productivity (GP) measured across labourers, means of labour, objects of l...
An undirected configuration change in a prior iteration produced zero impact, illustrating the cost of iterating without diagnosis.
Reported observation from an earlier iteration in the case study where a non-diagnostic change had no measurable effect.
high null result EvalLoop: A Methodology for Evaluation-Driven Iterative Impr... overall system performance change after undirected configuration change
The paper combines findings from information systems research, organizational behavior studies, and artificial intelligence literature through its analysis of recent empirical and theoretical studies conducted between 2021 and 2026.
Methodological description provided in the paper (literature synthesis covering 2021–2026).
high null result Human–AI Collaborative Systems for Workflow Optimization: A ... scope and sources of literature reviewed
Using aggregate data, the study provides no evidence that AI benefits any particular group of workers — neither highly educated nor less-educated ones.
Authors' interaction analysis between AI adoption and human capital using aggregate panel data; reported null finding for differential benefits across education/skill groups (1995–2017, 35 OECD countries).
high null result Economic Growth, AI Adoption and Human Capital Across the OE... differential benefits of AI by worker education/skill groups
Results are robust to state-by-year and industry-by-year fixed effects.
Robustness checks reported in paper that include state-by-year and industry-by-year fixed effects with results stated to hold.
high null result AI, Output, and Employment robustness of estimated effects to alternative fixed-effects specifications
Where AI can perform tasks independently, we find no significant employment effect.
Heterogeneous DiD estimates showing null (statistically non-significant) employment coefficients for occupations/industries where AI can perform tasks independently.
high null result AI, Output, and Employment employment (in occupations/industries with independent AI exposure)
We examine aggregate effects using administrative data covering essentially all U.S. employers in a difference-in-differences design exploiting occupational AI exposure across industries and states.
Statement in paper describing data and empirical strategy: administrative data covering essentially all U.S. employers; difference-in-differences design exploiting occupational AI exposure variation across industries and states.
high null result AI, Output, and Employment data_coverage_and_design (administrative data, DiD)
We observe no differences in productivity across adoption levels.
Authors' empirical comparison reporting null differences in measured firm productivity across adoption categories (based on their matched data).
high null result AI Adoption in S&P 500 Firms firm productivity
We observe no differences in capex across adoption levels.
Authors' empirical comparison reporting null differences in capital expenditures across adoption categories (based on firm financial data matched to adoption measure).
high null result AI Adoption in S&P 500 Firms capital expenditures (capex)
Under matched time limits, originality with a GPT-4 partner is statistically equivalent to that with a human partner.
Result from the in-person pilot (N = 62) comparing originality scores between participants partnered with GPT-4 versus human partners under matched time limits; reported as statistical equivalence in the paper.
We test this question across four lifecycle stages: data preparation, data extraction, statistical analysis, and reporting, using one generated skill per stage.
Experimental setup description specifying four task-lifecycle stages and use of one generated skill for each stage.
high null result Do LLM-Generated Skills Make Better AI Data Scientists? A Co... coverage of lifecycle stages tested
A supplemental token-matched control adds 1,512 runs and finds that Full skills perform similarly to task-irrelevant skill-formatted content.
Additional control experiment (token-matched content) reported in supplement, with run count and comparison results described.
high null result Do LLM-Generated Skills Make Better AI Data Scientists? A Co... task performance (Full skills vs token-matched irrelevant content)
The total spread across variants is only 1.2 percentage points.
Reported range/variation in performance metrics across all skill variants in the ablation experiment.
high null result Do LLM-Generated Skills Make Better AI Data Scientists? A Co... task performance (difference between best and worst variant)
Compared with prompting using the task alone, neither the full generated skill nor any ablated skill variant significantly improves performance; all p-values are at least 0.396.
Statistical hypothesis tests comparing Full and ablated-skill variants to task-only prompting; reported minimum p-value threshold.
high null result Do LLM-Generated Skills Make Better AI Data Scientists? A Co... task performance (accuracy / success rate)
The main ablation covers 56 tasks, nine model configurations, and three providers, yielding 7,560 runs.
Description of experimental design and aggregate counts reported in the paper (tasks × model configurations × providers → total run count).
high null result Do LLM-Generated Skills Make Better AI Data Scientists? A Co... number of experimental runs
We find no reliable improvement from full generated skills over No-Skill prompting.
Empirical comparison between Full generated-skill prompting and No-Skill (task-only) prompting across the study's evaluation tasks and model configurations; statistical testing reported.
high null result Do LLM-Generated Skills Make Better AI Data Scientists? A Co... task performance (accuracy / success rate)
Average inference power is a model-intrinsic constant, invariant to input resolution, image complexity, and prompt type, with less than 5% variation across all conditions.
Systematic on-device energy profiling across five models spanning three architecture families, four input resolutions, and two hardware platforms (NVIDIA RTX 3070 and Jetson Orin NX); direct measurement of power draw across different inputs and prompts.
high null result Seeing is Free, Speaking is Not: Uncovering the True Energy ... average inference power
Organized search in SSRN, Google Scholar, Web of Science, Scopus, and government repositories resulted in 37 sources that satisfy pre-determined inclusion criteria and are rated in three levels of evidence.
Description of the review's search and screening process reported in the paper (databases searched listed, inclusion criteria applied, resulting count = 37, evidence rated into three levels).
high null result From Compliance to Intelligence: Integrating AI and Predicti... number of included sources and evidence-level rating
The study used a sequential mixed-methods design (Qual → Quan) consisting of eight expert interviews analyzed with grounded-theory coding and a follow-up survey of 499 AI-aware consumers.
Methods reported in the paper: eight expert interviews (qualitative) and a survey with 499 respondents (quantitative).
high null result Conceptualization of causes and implications of AI adoption ... study design / methodological procedure
From that coded sample the authors built a causal model of 26 constructs and 67 relationships (64 directed, 3 contested).
Reported model construction from the coded sample as stated in the abstract.
high null result 3100 Opinions on Code Review in an AI World: Building Causal... causal model complexity (constructs and relationships)
The authors filtered that corpus and coded a stratified random sample of 3,100 documents with an LLM-assisted pipeline.
Reported sampling and coding procedure stated in the abstract.
high null result 3100 Opinions on Code Review in an AI World: Building Causal... coded sample size using LLM-assisted pipeline
We collected 38,709 grey-literature documents (engineering blogs and Reddit threads) and filtered to those substantively about code review.
Reported data-collection procedure and corpus size stated in the abstract.
high null result 3100 Opinions on Code Review in an AI World: Building Causal... grey-literature corpus size
We isolate the orchestration layer with a controlled swap: 22 locked evaluation tasks, six foundation models, changing only the orchestration layer (a frozen conventional production loop versus the Writer Agent Harness).
Methodological statement describing experimental design reported in the paper: controlled swap with 22 tasks and six models.
high null result The Harness Effect: How Orchestration Design Sets the Token ... experimental design / method
Task-completion quality is at parity between harness and baseline (0.78->0.81), directional at this sample size.
Reported average/directional quality scores from the controlled swap across 22 tasks and six models, showing scores 0.78 (baseline) and 0.81 (harness).
high null result The Harness Effect: How Orchestration Design Sets the Token ... task-completion quality (score)
The observed patterns in BEA–BLS data for 63 U.S. industries over 1997–2023 do not reflect cyclical variation but register a structural change in the system of factors of production.
Trend/structural analysis of BEA–BLS data for 63 U.S. industries (1997–2023) reported in the paper.
high null result 250 years of Smith’s work: How digital platforms bring us ba... structural change in factors of production
The digital sector comprises three times fewer industries than the physical sector.
Empirical statement based on BEA–BLS data covering 63 U.S. industries (1997–2023) as reported in the paper.
high null result 250 years of Smith’s work: How digital platforms bring us ba... count of industries in digital vs physical sector
The periodization of US macroeconomic productivity cycles was refined by identifying the new stages 'pandemic and adaptation phase' and 'artificial intelligence phase'.
Calculation of AAPC indices for 1947–2025 and retrospective comparative analysis leading to refinement of periodization and naming of new stages.
high null result Analysis of labor productivity in the context of technologic... identification of new macroeconomic stages in productivity cycles
Eight distinct macroeconomic cycles of productivity change in the United States from 1947 to 2025 are identified.
Secondary data analysis of aggregated US Bureau of Labor Statistics series for 1947–2025; long-term average annual rates of productivity change (AAPC) computed using the index method and geometric mean growth rate; comparative analysis to identify cycle breaks.
high null result Analysis of labor productivity in the context of technologic... macroeconomic cycles of aggregate labor productivity change
Survey data were collected from firms located in major Chinese cities (Beijing, Shenzhen, Xi’an, and Zhengzhou), resulting in 750 valid responses for analysis.
Reported survey sampling and data collection in the paper; explicit statement of cities sampled and number of valid responses (750).
Workers were assigned to no overrides, free overrides, or a two-per-machine limit on downward overrides.
Experimental design statement in paper: randomized assignment into three arms (no overrides, free overrides, constrained two-per-machine downward override limit).
high null result A Simple Solution to Improving Human Supervision of Algorith... Treatment assignment (experimental arms)
We tested [the policy] through a randomized field experiment with 553 workers at a major Chinese smart vending machine retailer that manages more than 59,000 machines and 4,000 SKUs.
Randomized field experiment described in paper; sample stated as 553 workers and operational context (retailer with >59,000 machines and >4,000 SKUs).
high null result A Simple Solution to Improving Human Supervision of Algorith... Experimental implementation / sample and setting description
The runs spanned several model generations, two agent harnesses, two reasoning effort levels, a testing tool, and two design oriented prompts.
Description of experimental conditions reported in the study (factors varied across the 90 runs).
high null result Reasoning effort, not tool access, buys first-try reliabilit... experimental condition coverage (model generation, harness, effort level, testin...