The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests About 🎲 Workforce Futures
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Evidence (7560 claims)

Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.

The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).

Browse by theme

Nine broad, paper-level topics. Click one to filter the claims below.

Adoption
9875 claims
Filter claims →
Productivity
8807 claims
Filter claims →
Governance
7870 claims
Filter claims →
Human-AI Collaboration
7560 claims
Filtered →
Org Design
4892 claims
Filter claims →
Innovation
4781 claims
Filter claims →
Labor Markets
4004 claims
Filter claims →
Skills & Training
3308 claims
Filter claims →
Inequality
2332 claims
Filter claims →

Claims by outcome category

Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.

Outcome Positive Negative Mixed Null Total
Other 870 233 116 1066 2363
Governance & Regulation 976 451 218 133 1809
Organizational Efficiency 949 224 144 88 1416
Technology Adoption Rate 764 287 141 122 1325
Research Productivity 501 152 74 362 1101
Output Quality 542 216 69 69 896
Decision Quality 387 198 94 54 740
Firm Productivity 513 67 101 27 714
AI Safety & Ethics 249 303 73 36 667
Market Structure 190 192 134 27 548
Task Allocation 243 77 91 36 452
Innovation Output 291 33 55 20 401
Skill Acquisition 206 72 65 21 364
Employment Level 133 63 115 22 335
Fiscal & Macroeconomic 153 79 52 32 323
Task Completion Time 206 37 12 15 272
Firm Revenue 179 52 29 5 266
Consumer Welfare 130 76 47 13 266
Inequality Measures 48 137 51 6 242
Worker Satisfaction 101 81 25 13 220
Error Rate 84 110 11 5 210
Wages & Compensation 98 47 30 10 185
Regulatory Compliance 88 73 17 7 185
Automation Exposure 66 64 33 16 182
Team Performance 105 29 30 11 176
Training Effectiveness 109 22 14 21 168
Developer Productivity 114 21 14 8 158
Job Displacement 12 90 24 1 127
Hiring & Recruitment 57 9 9 5 80
Skill Obsolescence 6 56 9 1 72
Social Protection 43 17 8 2 70
Creative Output 35 21 9 4 70
Labor Share of Income 18 21 17 1 57
Worker Turnover 15 16 4 35
Industry 1 1
Clear
Human Ai Collab Remove filter
We evaluate four mechanisms to enable cooperation: (1) repeating the game for many rounds, (2) reputation systems, (3) third-party mediators to delegate decision making to, and (4) contract agreements for outcome-conditional payments between players.
Description of experimental design / mechanisms evaluated in the study across four social dilemmas; details on implementation and sample sizes not provided in the excerpt.
high neutral CoopEval: Benchmarking Cooperation-Sustaining Mechanisms and... comparative effectiveness of four cooperation mechanisms
The paper provides lessons for scaling regression automation and enabling effective human-AI teaming in Agile settings.
Stated contribution of the paper (synthesis of lessons from the industrial case study).
high neutral Human-AI Collaboration for Scaling Agile Regression Testing:... availability of lessons and guidance
The Copilot was integrated with Hacon's CI pipelines and operates asynchronously as a 'silent AI teammate', producing candidate scripts for human review.
System integration and deployment description within the case study (implementation detail reported in the paper).
high neutral Human-AI Collaboration for Scaling Agile Regression Testing:... operational mode and integration with CI (asynchronous candidate generation for ...
We conducted an exploratory industrial case study of the Hacon Test Automation Copilot, an agentic AI system that generates system-level regression test scripts from validated specifications using retrieval-augmented generation and a multi-agent workflow.
Methodological claim: description of the study design and the system; the paper reports a single industrial case study at Hacon (a Siemens company).
high neutral Human-AI Collaboration for Scaling Agile Regression Testing:... capability to generate system-level regression test scripts
SAFI measures LLM performance on text-based representations of skills, not full occupational execution.
Methodological caveat stated by the authors clarifying the scope and limits of SAFI.
high neutral The AI Skills Shift: Mapping Skill Obsolescence, Emergence, ... scope of SAFI measure (text-based representations vs full job execution)
We propose an AI Impact Matrix that positions skills into four quadrants: High Displacement Risk, Upskilling Required, AI-Augmented, and Lower Displacement Risk.
Conceptual/interpretive framework introduced by the authors; described in text as proposed by the paper.
high neutral The AI Skills Shift: Mapping Skill Obsolescence, Emergence, ... interpretive classification of skills into four impact quadrants
Legitimate accountability is axiomatized through four minimal properties: Attributability (responsibility requires causal contribution), Foreseeability Bound (responsibility cannot exceed predictive capacity), Non-Vacuity (at least one agent bears non-trivial responsibility), and Completeness (all responsibility must be fully allocated).
Paper presents an explicit axiomatization listing these four properties as definitions/axioms forming the normative criteria for legitimate accountability.
high neutral The Accountability Horizon: An Impossibility Theorem for Gov... formal criteria for legitimate accountability
Collective behaviour is characterised through interaction graphs and joint action spaces.
Paper specifies interaction graphs and joint action spaces as part of the formal model (definitions and formal structure).
high neutral The Accountability Horizon: An Impossibility Theorem for Gov... formal representation of collective behaviour
Autonomy is characterised through a four-dimensional information-theoretic profile (epistemic, executive, evaluative, social).
Paper defines autonomy as a 4-dimensional information-theoretic profile (conceptual/mathematical definition within the formal model).
high neutral The Accountability Horizon: An Impossibility Theorem for Gov... measure/characterisation of agent autonomy
Using a strictly algorithmic baseline (mathematical bottleneck aggregation), we calculate Relative Occupational Automation Indices (OAI) for the U.S. labor market based on the DWA-level scores.
Method and calculation claim: algorithmic baseline aggregation applied across the 923 occupations / 2,087 DWAs to produce OAIs mapped to the U.S. labor market. Specific aggregation formula referenced but not numerically detailed in the excerpt.
high neutral Bounded by Risk, Not Capability: Quantifying AI Occupational... Relative Occupational Automation Index (OAI)
We deconstructed 923 occupations into 2,087 Detailed Work Activities (DWAs).
Explicit data processing claim in the paper: mapping of 923 occupations to 2,087 DWAs for analysis.
high neutral Bounded by Risk, Not Capability: Quantifying AI Occupational... coverage of occupations and DWAs used for analysis
A life insurance system integrated into an industry partner mobile app was tested in two experiments.
Paper reports two experiments running the ARQuest-enabled life insurance system inside a partner mobile app; experimental setup is stated though sample sizes are not provided in the excerpt.
high neutral AI in Insurance: Adaptive Questionnaires for Improved Risk P... experimental evaluation of system in partner app
We evaluate the architecture through a controlled experiment (600 runs across five industries: FinTech, Insurance, Healthcare, Vietnamese Banking, and Vietnamese Insurance).
Controlled experiment reported in the paper: 600 runs across five named industries (experimental setup reported in abstract).
high neutral Ontology-Constrained Neural Reasoning in Enterprise Agentic ... experimental performance of ontology-coupled vs ungrounded agents across industr...
The framework is calibrated with O*NET task data, a survey of 3,778 domain experts, and GPT-4o-derived task decompositions, and implemented in computer vision.
Calibration and empirical implementation using O*NET, a domain expert survey (n=3,778), and GPT-4o task decompositions; applied to computer vision tasks.
high neutral Economics of Human and AI Collaboration: When is Partial Aut... validity of calibration / empirical grounding of the framework
We introduce an entropy-based measure of task complexity that maps model accuracy into a labor substitution ratio, quantifying human labor displacement at each accuracy level.
New metric proposed in the paper (entropy-based task complexity) and mapping procedure from accuracy to substitution ratio; implemented in the framework.
high neutral Economics of Human and AI Collaboration: When is Partial Aut... labor substitution ratio (human labor displaced per unit accuracy)
Costinot and Werning (2023) develop a sufficient-statistic approach and find optimal technology taxes of 1–3.7% on robots.
Citation reported in the paper summarizing Costinot and Werning (2023)'s quantitative sufficient-statistic estimate.
high neutral NBER WORKING PAPER SERIES optimal robot tax rate
Guerreiro et al. (2022) characterize optimal Mirrleesian tax system with automation and find that robot taxes should be transitional—high when incumbent workers cannot retrain, converging to zero as new cohorts adjust skill investments.
Citation reported in the paper summarizing Guerreiro et al. (2022)'s theoretical result on transitional robot taxes.
high neutral NBER WORKING PAPER SERIES optimal robot tax path over time
If labor becomes economically redundant, the policy focus shifts from steering innovation to redesigning public finance and redistribution (e.g., new tax instruments, redistribution mechanisms).
Theoretical scenario analysis in the paper with references to related works (Korinek and Juelfs 2024; Korinek and Lockwood 2026).
high neutral NBER WORKING PAPER SERIES policy priority shift (steering -> public finance/redistribution)
Evaluation is carried out under three frozen context configurations (diff only: config_A; diff with file content: config_B; full context: config_C) enabling systematic ablation of context provision strategies.
Methodological description: three fixed context configurations defined and used for ablation experiments.
high neutral SWE-PRBench: Benchmarking AI Code Review Quality Against Pul... effect of context-provision design on model performance
We performed an extensive evaluation of 37 state-of-the-art Vision-Language Models on MultihopSpatial.
Empirical evaluation described in the paper listing the number of models evaluated (37).
high neutral MultihopSpatial: Multi-hop Compositional Spatial Reasoning B... benchmark coverage across models evaluated
Economic evaluations of GLAI should account for end-to-end risk externalities (error propagation, institutional trust, rights impacts), not only short-term productivity gains.
Methodological recommendation grounded in conceptual synthesis of technical, behavioral, and legal risks; normative argument rather than empirical result.
high neutral Why Avoid Generative Legal AI Systems? Hallucination, Overre... comprehensiveness of economic evaluations (inclusion of externalities vs. narrow...
Generative Legal AI (GLAI) systems are built on token-prediction (LLM) architectures rather than formal legal-reasoning architectures.
Conceptual and technical analysis in the paper distinguishing GLAI from other legal-tech; literature synthesis on common LLM architectures. No original empirical dataset or sample size—qualitative/technical review.
high neutral Why Avoid Generative Legal AI Systems? Hallucination, Overre... underlying model architecture type (token-prediction vs. formal-reasoning)
Through a thematic review of existing research, the authors identified recurring themes about incentive schemes: their components, how researchers manipulate them, and their impact on research outcomes.
Authors' stated method and findings: thematic review (the scope/number of reviewed papers not specified in excerpt).
high neutral Incentive-Tuning: Understanding and Designing Incentives for... themes in incentive design practices and reported impacts on empirical study out...
A critical aspect of conducting human–AI decision-making studies is the role of participants, often recruited through crowdsourcing platforms.
Claim based on the authors' thematic literature review noting participant sourcing practices (specific studies and counts not given in excerpt).
high neutral Incentive-Tuning: Understanding and Designing Incentives for... participant recruitment source (e.g., crowdsourcing) and its influence on study ...
Researchers conduct empirical studies investigating how humans use AI assistance for decision-making and how this collaboration impacts results.
Statement summarizing the research landscape; supported implicitly by the authors' thematic review of existing empirical studies (number of studies not specified in excerpt).
high neutral Incentive-Tuning: Understanding and Designing Incentives for... human behaviour and decision outcomes when assisted by AI (empirical study outco...
Returns to AI are heterogeneous across firms; estimating treatment effects requires attention to selection, complementarities, and dynamic adoption pipelines.
Methodological argument referencing treatment-effect literature and observed firm heterogeneity; supported by conceptual examples rather than a single empirical treatment-effect estimate.
high neutral Modern Management in the Age of Artificial Intelligence: Str... heterogeneity in returns to AI adoption (firm-level productivity or performance ...
Human-only and AI-assisted teams performed similarly on most outcomes.
Comparison across outcome measures from the randomized experiment; summary statement indicates parity on most measured tasks except for detection of major coding errors.
high null result AI-assisted teams outperform AI-led teams but not human-only... multiple reproduction-related outcomes (overall performance parity)
We randomly assigned 288 researchers to 103 teams working under three conditions (human-only, AI-assisted, AI-led).
Experimental design reported in paper: randomized assignment of 288 researchers into 103 teams across three experimental conditions.
high null result AI-assisted teams outperform AI-led teams but not human-only... experimental_assignment
An undirected configuration change in a prior iteration produced zero impact, illustrating the cost of iterating without diagnosis.
Reported observation from an earlier iteration in the case study where a non-diagnostic change had no measurable effect.
high null result EvalLoop: A Methodology for Evaluation-Driven Iterative Impr... overall system performance change after undirected configuration change
The paper combines findings from information systems research, organizational behavior studies, and artificial intelligence literature through its analysis of recent empirical and theoretical studies conducted between 2021 and 2026.
Methodological description provided in the paper (literature synthesis covering 2021–2026).
high null result Human–AI Collaborative Systems for Workflow Optimization: A ... scope and sources of literature reviewed
Results are robust to state-by-year and industry-by-year fixed effects.
Robustness checks reported in paper that include state-by-year and industry-by-year fixed effects with results stated to hold.
high null result AI, Output, and Employment robustness of estimated effects to alternative fixed-effects specifications
Where AI can perform tasks independently, we find no significant employment effect.
Heterogeneous DiD estimates showing null (statistically non-significant) employment coefficients for occupations/industries where AI can perform tasks independently.
high null result AI, Output, and Employment employment (in occupations/industries with independent AI exposure)
We examine aggregate effects using administrative data covering essentially all U.S. employers in a difference-in-differences design exploiting occupational AI exposure across industries and states.
Statement in paper describing data and empirical strategy: administrative data covering essentially all U.S. employers; difference-in-differences design exploiting occupational AI exposure variation across industries and states.
high null result AI, Output, and Employment data_coverage_and_design (administrative data, DiD)
Under matched time limits, originality with a GPT-4 partner is statistically equivalent to that with a human partner.
Result from the in-person pilot (N = 62) comparing originality scores between participants partnered with GPT-4 versus human partners under matched time limits; reported as statistical equivalence in the paper.
We test this question across four lifecycle stages: data preparation, data extraction, statistical analysis, and reporting, using one generated skill per stage.
Experimental setup description specifying four task-lifecycle stages and use of one generated skill for each stage.
high null result Do LLM-Generated Skills Make Better AI Data Scientists? A Co... coverage of lifecycle stages tested
A supplemental token-matched control adds 1,512 runs and finds that Full skills perform similarly to task-irrelevant skill-formatted content.
Additional control experiment (token-matched content) reported in supplement, with run count and comparison results described.
high null result Do LLM-Generated Skills Make Better AI Data Scientists? A Co... task performance (Full skills vs token-matched irrelevant content)
The total spread across variants is only 1.2 percentage points.
Reported range/variation in performance metrics across all skill variants in the ablation experiment.
high null result Do LLM-Generated Skills Make Better AI Data Scientists? A Co... task performance (difference between best and worst variant)
Compared with prompting using the task alone, neither the full generated skill nor any ablated skill variant significantly improves performance; all p-values are at least 0.396.
Statistical hypothesis tests comparing Full and ablated-skill variants to task-only prompting; reported minimum p-value threshold.
high null result Do LLM-Generated Skills Make Better AI Data Scientists? A Co... task performance (accuracy / success rate)
The main ablation covers 56 tasks, nine model configurations, and three providers, yielding 7,560 runs.
Description of experimental design and aggregate counts reported in the paper (tasks × model configurations × providers → total run count).
high null result Do LLM-Generated Skills Make Better AI Data Scientists? A Co... number of experimental runs
We find no reliable improvement from full generated skills over No-Skill prompting.
Empirical comparison between Full generated-skill prompting and No-Skill (task-only) prompting across the study's evaluation tasks and model configurations; statistical testing reported.
high null result Do LLM-Generated Skills Make Better AI Data Scientists? A Co... task performance (accuracy / success rate)
We survey 1,250 arXiv papers (2024-2026).
Systematic literature survey conducted by the paper; explicit statement of sample size and date range in the abstract.
high null result Recursive Self-Improvement in AI: From Bounded Self-Refineme... coverage of literature
From that coded sample the authors built a causal model of 26 constructs and 67 relationships (64 directed, 3 contested).
Reported model construction from the coded sample as stated in the abstract.
high null result 3100 Opinions on Code Review in an AI World: Building Causal... causal model complexity (constructs and relationships)
The authors filtered that corpus and coded a stratified random sample of 3,100 documents with an LLM-assisted pipeline.
Reported sampling and coding procedure stated in the abstract.
high null result 3100 Opinions on Code Review in an AI World: Building Causal... coded sample size using LLM-assisted pipeline
We collected 38,709 grey-literature documents (engineering blogs and Reddit threads) and filtered to those substantively about code review.
Reported data-collection procedure and corpus size stated in the abstract.
high null result 3100 Opinions on Code Review in an AI World: Building Causal... grey-literature corpus size
The study uses difference-in-differences linear regressions on 2023 and 2024 KBO season data to identify the causal impact of ABS adoption by player status.
Methods statement in paper: difference-in-differences linear regressions; sample sizes reported as n = 148 batters and n = 112 pitchers.
high null result Technology adoption and bias in officiating: automated Ball-... methodological approach (DiD estimation)
High-status pitchers' performance remains unaffected by ABS adoption.
Difference-in-differences linear regressions using KBO 2023 and 2024 season data for pitchers (n = 112); paper reports no detectable change for high-status pitchers.
high null result Technology adoption and bias in officiating: automated Ball-... Pitcher performance (aggregate; specific metrics not listed in summary)
The Korea Baseball Organization (KBO) officially implemented the Automated Ball-Strike System (ABS) in 2024.
Paper statement of policy change and use of 2023 and 2024 KBO season data; presented as factual background to the natural experiment.
high null result Technology adoption and bias in officiating: automated Ball-... ABS implementation / adoption
The periodization of US macroeconomic productivity cycles was refined by identifying the new stages 'pandemic and adaptation phase' and 'artificial intelligence phase'.
Calculation of AAPC indices for 1947–2025 and retrospective comparative analysis leading to refinement of periodization and naming of new stages.
high null result Analysis of labor productivity in the context of technologic... identification of new macroeconomic stages in productivity cycles
Eight distinct macroeconomic cycles of productivity change in the United States from 1947 to 2025 are identified.
Secondary data analysis of aggregated US Bureau of Labor Statistics series for 1947–2025; long-term average annual rates of productivity change (AAPC) computed using the index method and geometric mean growth rate; comparative analysis to identify cycle breaks.
high null result Analysis of labor productivity in the context of technologic... macroeconomic cycles of aggregate labor productivity change
Survey data were collected from firms located in major Chinese cities (Beijing, Shenzhen, Xi’an, and Zhengzhou), resulting in 750 valid responses for analysis.
Reported survey sampling and data collection in the paper; explicit statement of cities sampled and number of valid responses (750).