The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests About 🎲 Workforce Futures
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Evidence (8974 claims)

Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.

The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).

Browse by theme

Nine broad, paper-level topics. Click one to filter the claims below.

Adoption
10085 claims
Filter claims →
Productivity
8974 claims
Filtered →
Governance
8062 claims
Filter claims →
Human-AI Collaboration
7749 claims
Filter claims →
Org Design
5057 claims
Filter claims →
Innovation
4896 claims
Filter claims →
Labor Markets
4088 claims
Filter claims →
Skills & Training
3372 claims
Filter claims →
Inequality
2377 claims
Filter claims →

Claims by outcome category

Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.

Outcome Positive Negative Mixed Null Total
Other 882 244 117 1097 2424
Governance & Regulation 1010 469 229 135 1875
Organizational Efficiency 977 235 149 90 1462
Technology Adoption Rate 781 299 143 128 1362
Research Productivity 506 155 74 363 1110
Output Quality 555 219 71 70 915
Decision Quality 395 200 95 54 751
Firm Productivity 523 67 101 27 724
AI Safety & Ethics 262 309 75 36 688
Market Structure 195 201 135 30 566
Task Allocation 248 77 96 38 464
Innovation Output 300 34 55 20 411
Skill Acquisition 207 75 65 21 368
Employment Level 138 67 119 24 350
Fiscal & Macroeconomic 156 80 53 33 329
Task Completion Time 211 38 13 16 280
Firm Revenue 183 52 29 5 270
Consumer Welfare 131 77 48 13 269
Inequality Measures 50 141 54 9 254
Worker Satisfaction 104 85 25 13 227
Error Rate 87 112 11 5 215
Automation Exposure 69 69 37 20 198
Wages & Compensation 102 49 31 11 193
Team Performance 115 30 30 11 187
Regulatory Compliance 88 74 17 7 186
Training Effectiveness 109 22 14 21 168
Developer Productivity 116 21 15 8 161
Job Displacement 12 92 26 1 131
Hiring & Recruitment 57 12 9 5 83
Skill Obsolescence 6 59 10 2 77
Social Protection 43 17 8 2 70
Creative Output 35 21 9 4 70
Labor Share of Income 18 23 17 1 59
Worker Turnover 15 16 4 35
Industry 1 1
Clear
Productivity Remove filter
The positive effect of industrial robots on firm-level TFP is statistically significant among firms that face low government subsidies.
Heterogeneity/subsample analysis in the paper using firm-level data (2006–2019) that splits or interacts robot effects with measures of government subsidies.
high positive The application of industrial robots, capital distortion, an... total factor productivity (TFP) at the firm level (conditional on low government...
The application of industrial robots has a positive effect on firm-level total factor productivity (TFP).
Empirical analysis using matched data on industrial robots and Chinese listed companies for 2006–2019; the paper reports econometric tests estimating the relationship between robot application/adoption and firm-level TFP.
high positive The application of industrial robots, capital distortion, an... total factor productivity (TFP) at the firm level
Our findings provide insights for designing LLM-based dialogue systems that support NFR assessment.
Synthesis of empirical results (agreement, low accuracy vs. experts, and user-satisfaction models) leading to design recommendations discussed in the paper.
high positive Accuracy and Satisfaction in Multi-Turn LLM Dialogues for NF... design guidance for LLM dialogue systems
Proactive interactions (by the system) positively affect user satisfaction.
Same user-satisfaction model reported in the paper showing positive association between proactive system behavior and satisfaction across participant sessions.
Developers tend to agree with LLM assessments of NFRs.
Empirical analysis of the 49 programmers' interactions and judgments compared to the LLM's assessments across 148 HIPAA-derived NFRs (agreement rates computed in the paper).
high positive Accuracy and Satisfaction in Multi-Turn LLM Dialogues for NF... agreement between developers and LLM assessments
Higher query topic entropy produces a near-monotonic increase in the performance gap between oracle routing (Capability Frontier) and the best single model (shown via controlled probabilistic simulations).
Controlled probabilistic simulations reported in the paper; demonstration of relationship between query topic entropy and performance gap between oracle routing and best single model. (Simulation parameters not specified in the excerpt.)
high positive The Capability Frontier: Benchmarks Miss 82% of Model Perfor... performance gap between oracle routing and best single model (as a function of q...
Using the Capability Frontier, state-of-the-art (SOTA) accuracy is matched at an 85% cost reduction.
Empirical cost-performance comparison across the 21 LLMs and 16 benchmarks showing the Capability Frontier achieves SOTA accuracy at substantially lower cost.
high positive The Capability Frontier: Benchmarks Miss 82% of Model Perfor... accuracy at matched cost (cost reduction required to match SOTA accuracy)
Additionally correcting for single runs (i.e., sampling multiple generations and selective retention) yields an 82% improvement.
Empirical comparison using the Capability Frontier with both corrections (across 21 LLMs and 16 benchmarks) versus single-run, single-model benchmark evaluations.
high positive The Capability Frontier: Benchmarks Miss 82% of Model Perfor... performance (improvement over baseline single-run single-model evaluation)
Correcting for single-model evaluation yields a 54% error rate reduction.
Empirical evaluation comparing the Capability Frontier to each benchmark's top-performing model at matched cost across 21 LLMs and 16 benchmarks (coding, reasoning, medicine, factuality, instruction following, agentic).
AI-driven agglomeration can support green agricultural transformation.
Overall interpretation and policy implication derived from empirical findings (U-shaped relationship, mediation by technological progress, heterogeneity and threshold results) based on the provincial panel analysis.
high positive How Does Artificial Intelligence Industry Agglomeration Affe... green agricultural transformation (proxied by APCRS)
The dynamic (U-shaped) process is more likely to emerge when public innovation investment and rural household income exceed critical thresholds.
Threshold analysis reported in the paper indicating the nonlinear relationship becomes apparent when public innovation investment and rural household income pass estimated threshold values, based on the 30-province panel.
high positive How Does Artificial Intelligence Industry Agglomeration Affe... APCRS (agricultural pollution–carbon reduction synergy)
Agricultural socialized services and rural industrial integration buffer the initial negative association between AIIA and APCRS.
Moderation/interaction analysis in the paper showing those regional/institutional factors mitigate the early negative effect of AIIA on APCRS using the 30-province panel.
high positive How Does Artificial Intelligence Industry Agglomeration Affe... APCRS (agricultural pollution–carbon reduction synergy)
Technological progress partially mediates the relationship between AIIA and APCRS.
Mediation analysis reported in the paper using the same 30-province panel (2016–2024) indicating a partial mediation effect of technological progress.
high positive How Does Artificial Intelligence Industry Agglomeration Affe... APCRS (agricultural pollution–carbon reduction synergy)
Together, these contributions present a more rigorous alternative to the dominant accuracy-centric evaluation paradigm.
Synthesis of the paper's contributions (case study, new benchmarks, empirical analyses, randomized experiment).
high positive Life After Benchmark Saturation: A Case Study of CORE-Bench evaluation rigor (methodological robustness of multidimensional evaluation vs ac...
The observed ~2x speedup is likely underestimated because one-fifth of human-only reproductions reached the time limit before completing.
Observed censoring/timeouts in the randomized experiment (authors report ~20% of human-only runs hit the time limit).
high positive Life After Benchmark Saturation: A Case Study of CORE-Bench proportion of human-only reproductions that timed out (censoring) and its effect...
The randomized experiment found a statistically significant speedup by about a factor of two for human-agent collaboration compared to human-only on computational reproducibility tasks.
Small-scale randomized experiment reported in the paper (statistical test reported as significant; exact n not stated in the abstract).
high positive Life After Benchmark Saturation: A Case Study of CORE-Bench task completion time (speedup)
Despite accuracy saturation, CORE-Bench v1.1 remains useful for measuring reliability and model versus scaffold performance.
Empirical analyses on CORE-Bench v1.1 separating model contribution from scaffold/infrastructure contribution and reporting reliability-related metrics.
high positive Life After Benchmark Saturation: A Case Study of CORE-Bench reliability / relative contribution of model vs scaffold
Despite accuracy saturation, CORE-Bench v1.1 remains useful for measuring efficiency.
Empirical measurements on CORE-Bench v1.1 reported in the paper (benchmark timing/efficiency analyses).
high positive Life After Benchmark Saturation: A Case Study of CORE-Bench efficiency (time to complete reproducibility tasks)
We introduce an improved benchmark, CORE-Bench v1.1, and an out-of-distribution task suite, CORE-Bench OOD.
Development and release of two benchmark artifacts described in the paper (methodological contribution).
high positive Life After Benchmark Saturation: A Case Study of CORE-Bench existence and availability of CORE-Bench v1.1 and CORE-Bench OOD
Measuring agents along the six non-accuracy dimensions yields meaningful insights into agent performance even after accuracy saturates.
Empirical case study using CORE-Bench Hard and the authors' derived benchmarks and analyses.
high positive Life After Benchmark Saturation: A Case Study of CORE-Bench informative value of non-accuracy evaluation metrics
For Chinese firms, productivity gains from GenAI are most likely when adoption is supported by cloud infrastructure, data readiness, skilled labor, workflow redesign, and strong digital ecosystems.
Synthesis of China-focused digital transformation studies and literature incorporated into the review; no new China-specific empirical analysis in this paper.
high positive Generative AI, Digital Infrastructure, and Firm Productivity... likelihood and magnitude of productivity gains at firm-level in Chinese firms
Existing studies show that GenAI can improve software-development tasks.
Synthesis of empirical studies and task-level experiments (e.g., developer assistance tools) reviewed in the paper.
high positive Generative AI, Digital Infrastructure, and Firm Productivity... software development productivity (coding speed, bug rates, developer time saved...
Existing studies show that GenAI can improve consulting tasks.
Cited task-level studies and applied examples in advisory/consulting work synthesized in the review.
high positive Generative AI, Digital Infrastructure, and Firm Productivity... consulting task performance / decision quality
Existing studies show that GenAI can improve customer support tasks.
Review of task-level experiments and applied studies in customer support settings reported across the literature synthesized in the paper.
high positive Generative AI, Digital Infrastructure, and Firm Productivity... customer support task performance (response time, resolution quality, throughput...
Existing studies show that GenAI can improve writing tasks.
Synthesis of task-level productivity experiments and prior empirical studies on GenAI-assisted writing (literature reviewed in the paper).
high positive Generative AI, Digital Infrastructure, and Firm Productivity... writing task performance (speed, quality)
AIGP simultaneously provides interpretable and transparent pricing rationales.
Author claim in paper that the LLM-based approach produces interpretable pricing rationales; no details or quantified user/analyst evaluation provided in the excerpt.
high positive AIGP: An LLM-Based Framework for Long-Term Value Alignment i... interpretability / transparency of pricing rationales
In large-scale online A/B tests on Tao Factory, AIGP achieved +8.20% in milestone achievement rate over 14 days compared to the production baseline.
Reported result from 'large-scale online A/B tests on Tao Factory' comparing AIGP to production baseline over a 14-day period; exact experiment sample size not provided in the excerpt.
high positive AIGP: An LLM-Based Framework for Long-Term Value Alignment i... milestone achievement rate
In large-scale online A/B tests on Tao Factory, AIGP achieved +7.59% in Return on Investment (ROI) over 14 days compared to the production baseline.
Reported result from 'large-scale online A/B tests on Tao Factory' comparing AIGP to production baseline over a 14-day period; exact experiment sample size not provided in the excerpt.
high positive AIGP: An LLM-Based Framework for Long-Term Value Alignment i... Return on Investment (ROI)
In large-scale online A/B tests on Tao Factory, AIGP achieved +13.21% in Gross Merchandise Value (GMV) over 14 days compared to the production baseline.
Reported result from 'large-scale online A/B tests on Tao Factory' comparing AIGP to production baseline over a 14-day period; exact experiment sample size not provided in the excerpt.
high positive AIGP: An LLM-Based Framework for Long-Term Value Alignment i... Gross Merchandise Value (GMV)
The Long-Term Value Estimator (LTVE), trained via offline reinforcement learning on historical data, serves as a reward model to score candidate pricing actions and select preference pairs for Direct Preference Optimization (DPO), thereby aligning the pricing policy with long-term business objectives.
Methodological description in paper: LTVE trained by offline RL on historical data and used as reward model for scoring and DPO preference selection; described as central component of framework. No sample size or quantitative validation details in the excerpt beyond its described role.
high positive AIGP: An LLM-Based Framework for Long-Term Value Alignment i... alignment of pricing policy with long-term business objectives
For efficient deployment while maintaining high-quality outputs, AIGP employs supervised fine-tuning for knowledge distillation.
Methodological claim in paper describing use of supervised fine-tuning (SFT) / knowledge distillation as part of model training and deployment strategy; no quantitative performance numbers tied specifically to SFT in the provided excerpt.
high positive AIGP: An LLM-Based Framework for Long-Term Value Alignment i... output quality (model output quality) and deployment efficiency
AIGP is a novel framework that leverages a Large Language Model (LLM) prompted with domain knowledge, structured data and textual context to make interpretable, knowledge-aware pricing decisions.
Methodological description in the paper (proposed system architecture/design); no quantitative evaluation detail provided in the excerpt beyond assertions.
high positive AIGP: An LLM-Based Framework for Long-Term Value Alignment i... ability to make interpretable, knowledge-aware pricing decisions
AI-driven outcomes depend less on the technology itself and more on complementary conditions—human capital formation, digital and data infrastructure, institutional coordination, and governance capacity—that enable effective diffusion.
Thematic synthesis of reviewed literature (2015–2025) highlighting repeated findings that complementarities (human capital, infrastructure, institutions) mediate AI diffusion and impacts.
high positive The Impact of Artificial Intelligence as a General-Purpose T... AI-driven growth outcomes (magnitude/direction conditional on complementarities)
The research aims to inform the design of responsible sociotechnical systems that preserve meaningful human involvement while retaining the efficiency benefits of AI-enabled decision support.
Stated normative aim/purpose in the abstract (intended policy/design impact). No empirical evidence or evaluation reported in the provided text.
high positive Strategic Adoption of AI-Enabled Decision-Making Systems: De... preservation of meaningful human involvement; retention of efficiency benefits f...
Uncertainty-aware emulation transforms mechanistic crop simulation from a computational bottleneck into an on-demand discovery engine capable of interrogating the full genotype, environment and management space at a scale no process-based model can match.
Synthesis claim based on method development, large training dataset (2 million simulations), speed improvements, and demonstrated large-scale experiments (100k trait configurations).
high positive From Simulation to Discovery: AI Enabled Probabilistic Emula... ability to perform large-scale genotype × environment × management exploration (...
Radiation use efficiency and temperature-driven root dynamics are dominant drivers of yield resilience.
Analysis of trait importance or sensitivity within the emulator experiments indicating RUE and temperature-driven root dynamics as dominant factors influencing yield resilience.
high positive From Simulation to Discovery: AI Enabled Probabilistic Emula... relative importance of specific traits (radiation use efficiency and temperature...
We identify 181 maize trait combinations that consistently maintain high yield across all tested conditions.
Results of the large-scale emulator exploration across trait configurations, soils, and climate scenarios reporting identification of 181 trait combinations meeting the 'consistently high yield' criterion.
high positive From Simulation to Discovery: AI Enabled Probabilistic Emula... count of trait combinations maintaining high yield across tested conditions
Applying the framework across 100,000 trait configurations, six soil environments in Iowa and Illinois, and climate projections through the year 2100 under two emissions scenarios enables large-scale exploration.
Application/experiment described using 100,000 trait configurations, six soil environments (Iowa and Illinois), and climate projections to 2100 under two emissions scenarios.
high positive From Simulation to Discovery: AI Enabled Probabilistic Emula... scale of exploration (number of trait configurations, environments, and climate ...
The framework provides calibrated predictive uncertainty without costly Bayesian inference.
Modeling approach described as probabilistic, yielding calibrated predictive uncertainty; claim that this is achieved without expensive Bayesian methods.
high positive From Simulation to Discovery: AI Enabled Probabilistic Emula... calibration of predictive uncertainty
The framework is augmented with a convolutional synthetic weather generator that produces physically consistent climate sequences.
Methodological statement describing augmentation of the emulator with a convolutional synthetic weather generator intended to generate physically consistent climate sequences.
high positive From Simulation to Discovery: AI Enabled Probabilistic Emula... quality/physical consistency of synthetic weather sequences
The emulator was trained on two million simulations spanning diverse genetic, soil, and management conditions.
Method description stating training dataset size of two million APSIM simulations covering varied genetics, soils, and management.
high positive From Simulation to Discovery: AI Enabled Probabilistic Emula... training dataset size / coverage of genotype × environment × management space
The emulator reduces simulation time by several orders of magnitude compared to the mechanistic APSIM model.
Reported comparison of simulation runtimes between the probabilistic neural emulator and APSIM (statement of 'several orders of magnitude' speedup).
high positive From Simulation to Discovery: AI Enabled Probabilistic Emula... simulation runtime / task completion time
We develop a probabilistic neural emulator of APSIM that reproduces key maize growth processes across 13 outputs with high fidelity (with R^2 of 0.93).
Model evaluation comparing emulator outputs to APSIM across 13 output variables; reported R^2 = 0.93. (Training and test data drawn from simulations used to train the emulator.)
high positive From Simulation to Discovery: AI Enabled Probabilistic Emula... fidelity of emulator predictions to process-based model outputs (R^2 across 13 o...
Production scale positively moderates both the human capital mechanism and the product R&D mechanism through which AI promotes value chain upgrading.
Moderation analysis in the panel econometric framework using 30-province data (2010–2022) showing that larger production scale strengthens the mediating effects of human capital and product R&D.
high positive The impact of artificial intelligence on value chain upgradi... strength of mediated effect (via human capital and product R&D) on value chain u...
AI facilitates value chain upgrading in the equipment manufacturing industry through two channels: enhancing human capital levels and driving product R&D.
Mechanism tests (mediation/ channel analysis) conducted on the 30-province panel (2010–2022) showing empirical support for human capital and product R&D as mediators of AI's effect on upgrading.
high positive The impact of artificial intelligence on value chain upgradi... mediated effect on value chain upgrading via human capital and product R&D
AI significantly enhances value chain upgrading in capital-intensive and technology-intensive equipment manufacturing industries.
Industry-type heterogeneity analysis within the 30-province panel (2010–2022) comparing capital-intensive and technology-intensive subsectors; reported statistically significant positive coefficients for these subsectors.
high positive The impact of artificial intelligence on value chain upgradi... value chain upgrading in equipment manufacturing (by industry intensity type)
The positive effect of AI on value chain upgrading remains robust after a series of stability tests and when addressing endogeneity concerns.
Stability/robustness tests and endogeneity discussions reported in the paper applied to the same 30-province panel (2010–2022); unspecified robustness procedures and endogeneity treatments mentioned.
high positive The impact of artificial intelligence on value chain upgradi... value chain upgrading in the equipment manufacturing industry (robustness of est...
AI promotes value chain upgrading in the equipment manufacturing industry.
Panel econometric analysis using data from 30 Chinese provinces over 2010–2022; models report a statistically significant positive coefficient on AI measures; robustness checks reported.
high positive The impact of artificial intelligence on value chain upgradi... value chain upgrading in the equipment manufacturing industry
In June 2026, the median researcher generated more than 50 times as many monthly output tokens across Codex and ChatGPT as they did in November 2025.
Internal monthly output-token counts by role (researcher) from November 2025 and June 2026 in Codex and ChatGPT usage logs.
high positive The Shift to Agentic AI: Evidence from Codex monthly output tokens generated by median researcher
In June 2026, the median OpenAI employee in a legal role generated 13 times more monthly output tokens across Codex and ChatGPT than they did in November 2025.
Internal monthly output-token counts by role (legal) from November 2025 and June 2026 in Codex and ChatGPT usage logs.
high positive The Shift to Agentic AI: Evidence from Codex monthly output tokens generated by median legal-role employee