The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests About 🎲 Workforce Futures
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Evidence (8974 claims)

Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.

The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).

Browse by theme

Nine broad, paper-level topics. Click one to filter the claims below.

Adoption
10085 claims
Filter claims →
Productivity
8974 claims
Filtered →
Governance
8062 claims
Filter claims →
Human-AI Collaboration
7749 claims
Filter claims →
Org Design
5057 claims
Filter claims →
Innovation
4896 claims
Filter claims →
Labor Markets
4088 claims
Filter claims →
Skills & Training
3372 claims
Filter claims →
Inequality
2377 claims
Filter claims →

Claims by outcome category

Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.

Outcome Positive Negative Mixed Null Total
Other 882 244 117 1097 2424
Governance & Regulation 1010 469 229 135 1875
Organizational Efficiency 977 235 149 90 1462
Technology Adoption Rate 781 299 143 128 1362
Research Productivity 506 155 74 363 1110
Output Quality 555 219 71 70 915
Decision Quality 395 200 95 54 751
Firm Productivity 523 67 101 27 724
AI Safety & Ethics 262 309 75 36 688
Market Structure 195 201 135 30 566
Task Allocation 248 77 96 38 464
Innovation Output 300 34 55 20 411
Skill Acquisition 207 75 65 21 368
Employment Level 138 67 119 24 350
Fiscal & Macroeconomic 156 80 53 33 329
Task Completion Time 211 38 13 16 280
Firm Revenue 183 52 29 5 270
Consumer Welfare 131 77 48 13 269
Inequality Measures 50 141 54 9 254
Worker Satisfaction 104 85 25 13 227
Error Rate 87 112 11 5 215
Automation Exposure 69 69 37 20 198
Wages & Compensation 102 49 31 11 193
Team Performance 115 30 30 11 187
Regulatory Compliance 88 74 17 7 186
Training Effectiveness 109 22 14 21 168
Developer Productivity 116 21 15 8 161
Job Displacement 12 92 26 1 131
Hiring & Recruitment 57 12 9 5 83
Skill Obsolescence 6 59 10 2 77
Social Protection 43 17 8 2 70
Creative Output 35 21 9 4 70
Labor Share of Income 18 23 17 1 59
Worker Turnover 15 16 4 35
Industry 1 1
Clear
Productivity Remove filter
Under this distinct problem geometry (l1-stationarity, l_infty-smoothness, separable noise), we derive matched upper and lower bounds for SignSGD and explicitly characterize the problem class in which SignSGD provably dominates SGD.
Theoretical derivation of both upper bounds (for SignSGD) and matching lower bounds (for the problem class) presented in the paper; proofs establishing tightness.
high positive When and Why SignSGD Outperforms SGD: A Theoretical Study Ba... convergence bounds (upper and lower) for SignSGD under specified assumptions
By analyzing sign-based optimizers under l1-norm stationarity, l_infty-smoothness, and a separable noise model, we can better capture the coordinate-wise nature of signed updates and overcome the barrier that prevents sign-based methods from outperforming SGD in standard settings.
Theoretical analysis in the paper introducing these alternative geometric/assumption settings (l1-stationarity, l_infty-smoothness, separable noise) and deriving results under these assumptions.
high positive When and Why SignSGD Outperforms SGD: A Theoretical Study Ba... applicability of sign-based optimizer analysis and potential for improved conver...
Post-hoc SHAP attribution reveals that complaint recurrence and neighborhood-level statistics are stronger predictors of actionable violations than raw complaint volume.
Empirical claim based on post-hoc SHAP feature-attribution analysis applied to the paper's models; the excerpt reports a relative feature importance finding but provides no numeric effect sizes or sample counts.
high positive Scaling the Queue: Reinforcement Learning for Equitable Call... predictive importance for actionable violations (feature importance)
We formalize each domain as a Markov Decision Process (MDP) in which equitable classification coverage is a first-class reward objective.
Methodological specification in the paper asserting each operational domain was modeled as an MDP with equity-aware reward structure. No further empirical details in the excerpt.
high positive Scaling the Queue: Reinforcement Learning for Equitable Call... equitable classification coverage (as a modeled reward)
The proposed technique is designed to maximize throughput, minimize misclassification cost, and actively narrow historical equity gaps in service delivery.
Stated design objectives of the RL approach in the paper. No quantified outcomes or evaluation reported in the provided text.
high positive Scaling the Queue: Reinforcement Learning for Equitable Call... throughput; misclassification cost; historical equity gaps in service delivery
Rather than replacing human classifiers, our agents act as intelligent intake routers that learn to assign incoming complaints to action categories: escalate, batch, defer, inspect now.
Descriptive claim of agent behavior and intended design; asserts agents perform routing decisions into four action categories. No empirical performance numbers provided in the excerpt.
high positive Scaling the Queue: Reinforcement Learning for Equitable Call... complaint routing action assignment
We develop an equity-centered reinforcement learning (RL) framework that augments call classification capacity across six New York City Department of Buildings operational domains (boiler safety, crane and derrick oversight, heat and hot water, housing complaint triage, scaffold safety, and Natural Area District protection).
Methodological development described in the paper; claimed application domain spans six named DOB operational areas. No evaluation metrics or sample sizes provided in the excerpt.
high positive Scaling the Queue: Reinforcement Learning for Equitable Call... call classification capacity / intake routing capability
Design principle: effective AI assistance should clear a quality threshold suited to the target content, rather than simply be present.
Authors' proposed design principle based on empirical and qualitative results from their study.
high positive Making AI Drafts Count: A Quality Threshold in Audio Descrip... design guidance for AI assistance effectiveness
Qualitative findings suggest the required quality threshold for helpful AI drafts is content-dependent; as visual complexity increases, the quality needed from AI drafts increases.
Authors' qualitative analysis from the study (no numeric measures provided in the excerpt).
high positive Making AI Drafts Count: A Quality Threshold in Audio Descrip... relationship between visual complexity and required AI draft quality
There is a minimum quality threshold for AI drafts to be effective; simple presence of AI assistance is insufficient.
Synthesis of empirical results and comparisons between GenAD and baseline drafts reported by the authors (stated as an interpretation of the findings).
high positive Making AI Drafts Count: A Quality Threshold in Audio Descrip... effectiveness of AI assistance (dependent on draft quality)
Baseline drafts generated from simple, unguided prompts offered only modest benefits compared to authoring from scratch.
Empirical comparison reported in the within-subjects study contrasting GenAD drafts and baseline (unguided-prompt) drafts; no numeric effect sizes or sample sizes provided in the excerpt.
high positive Making AI Drafts Count: A Quality Threshold in Audio Descrip... benefit/effectiveness of baseline AI drafts (e.g., quality or efficiency gains)
GenAD drafts significantly reduced cognitive load.
Result reported from the within-subjects study (authors state a significant reduction in cognitive load when using GenAD drafts); specific measure, statistical values, and sample size not provided in the excerpt.
GenAD drafts cut completion time by more than half.
Result reported from the within-subjects study comparing completion time when using GenAD drafts versus authoring from scratch; exact sample size and numeric reduction not provided in the excerpt.
Recent work has shown that giving novice describers an AI-generated draft to start from helps produce higher-quality audio description (AD) and lowers the barrier to entry.
Statement refers to prior published work (no specific study, sample size, or citation provided in the excerpt).
high positive Making AI Drafts Count: A Quality Threshold in Audio Descrip... AD quality / barrier to entry for novice describers
AI supports data-driven decision-making processes in logistics operations.
Qualitative structured literature review of 31 scholarly sources synthesized by the study.
high positive Evaluating the Role of Artificial Intelligence in Optimizing... data-driven decision-making
AI optimizes transportation routes (route optimization), improving logistics performance.
Qualitative structured literature review of 31 scholarly sources synthesized by the study.
AI enhances logistics performance by improving forecasting accuracy.
Qualitative structured literature review of 31 scholarly sources (journal articles and related publications) synthesized by the study.
The driving effect of digital technology integration on low-carbon transformation is more prominent for firms located in central and western regions.
Regional heterogeneity analysis reported in the paper shows stronger effects for firms in central and western regions of China within the panel sample (2009–2021).
high positive The Impact of Digital Technology Integration on Low-Carbon T... low-carbon transformation progress (heterogeneous effect by region)
The driving effect of digital technology integration on low-carbon transformation is more prominent in non-state-owned enterprises.
Heterogeneity analysis in the paper finds larger estimated effects for non-state-owned firms compared with state-owned firms, using the same panel sample and methods.
high positive The Impact of Digital Technology Integration on Low-Carbon T... low-carbon transformation progress (heterogeneous effect by ownership type)
The driving effect of digital technology integration on low-carbon transformation is more prominent in large-scale firms.
Heterogeneity analysis reported in the paper shows stronger estimated effects for larger firms in the panel sample (energy-intensive A-share listed firms, 2009–2021).
high positive The Impact of Digital Technology Integration on Low-Carbon T... low-carbon transformation progress (heterogeneous effect by firm size)
Digital technology integration boosts low-carbon transformation mainly by reducing operating costs.
Mechanism analysis in the paper identifies reduced operating costs as a key channel through which digital technology integration promotes low-carbon transformation; based on panel regressions and mediation-style tests using the same sample.
high positive The Impact of Digital Technology Integration on Low-Carbon T... operating costs (as a mediating mechanism for low-carbon transformation)
Digital technology integration boosts low-carbon transformation mainly by enhancing corporate R&D innovation capacity.
Mechanism analysis in the paper finds that an increase in R&D innovation capacity is a primary transmission channel linking digital technology integration to low-carbon transformation; based on the same panel and empirical methods.
high positive The Impact of Digital Technology Integration on Low-Carbon T... R&D innovation capacity (as a mediating mechanism for low-carbon transformation)
The positive effect of digital technology integration on low-carbon transformation remains valid after a series of robustness tests.
Authors report that the main finding holds after conducting multiple robustness checks (details not provided in the summary); same panel sample and measurement approach as main analysis.
high positive The Impact of Digital Technology Integration on Low-Carbon T... low-carbon transformation progress (robustness of main effect)
Digital technology integration significantly promotes the low-carbon transformation of energy-intensive enterprises.
Panel data of A-share listed firms in energy-intensive industries (2009–2021); measure of corporate digital technology integration from text analysis (frequency of digital-technology-related words in annual reports); low-carbon transformation measured using the LTFP method; empirical regression tests reported in the paper.
high positive The Impact of Digital Technology Integration on Low-Carbon T... low-carbon transformation progress
The evaluation covers multiple collaborative tasks and a variety of base LLM models.
Paper states experiments were run across multiple collaborative tasks and a variety of base models (breadth of evaluation).
high positive Improving the Efficiency of Language Agent Teams with Adapti... evaluation breadth (number/types of tasks and models)
The LATTE protocol maintains consistency under partial observability and communication constraints while enabling dynamic allocation and adaptation.
Design claim supported by the protocol description and reported empirical results demonstrating consistent coordination under constrained conditions.
high positive Improving the Efficiency of Language Agent Teams with Adapti... consistency of coordination under partial observability/communication constraint...
LATTE empowers agents to dynamically allocate work, adapt coordination, and discover new tasks.
Claim supported by the framework design and demonstrations in the paper (agents use the coordination graph to reassign and discover tasks during execution).
high positive Improving the Efficiency of Language Agent Teams with Adapti... dynamic task allocation / discovery
LATTE matches or exceeds the accuracy of standard designs including MetaGPT, decentralized teams, top-down Leader-Worker hierarchies, and static decompositions.
Reported accuracy comparisons from empirical experiments across several collaborative tasks and base models.
high positive Improving the Efficiency of Language Agent Teams with Adapti... accuracy (output quality)
LATTE reduces coordination failures such as file conflicts and redundant outputs.
Empirical evaluation comparing incidence of coordination failures between LATTE and baseline team coordination approaches.
high positive Improving the Efficiency of Language Agent Teams with Adapti... coordination failures (file conflicts, redundant outputs)
LATTE reduces communication (and communication overhead) compared to standard designs.
Empirical comparisons reported across multiple collaborative tasks and base models, measuring communication and coordination metrics.
high positive Improving the Efficiency of Language Agent Teams with Adapti... communication / communication overhead
LATTE reduces wall-clock time compared to standard designs.
Empirical evaluation across multiple collaborative tasks and various base models with time measurements reported in comparisons to baselines.
high positive Improving the Efficiency of Language Agent Teams with Adapti... wall-clock time (task completion time)
LATTE reduces token usage compared to standard designs (including MetaGPT, decentralized teams, top-down Leader-Worker hierarchies, and static decompositions).
Empirical evaluation across multiple collaborative tasks and a variety of base models, comparing LATTE to listed baseline designs.
In LATTE, a team of agents collaboratively construct and maintain a shared, evolving coordination graph which encodes sub-task dependencies, individual agent assignment, and the current state of sub-task progress.
Paper describes the protocol and its components (design/specification); supported by implementation details in the paper.
high positive Improving the Efficiency of Language Agent Teams with Adapti... task_allocation and coordination state (coordination graph)
We introduce Language Agent Teams for Task Evolution (LATTE), a framework for coordinating LLM teams inspired by distributed systems.
Paper describes the LATTE framework as a proposed coordination protocol (design/conceptual contribution).
high positive Improving the Efficiency of Language Agent Teams with Adapti... framework introduction / coordination protocol
Olava Extract reduced inference cost by 78% to 97% compared with the frontier models tested.
Reported cost comparison (inference cost) versus the five frontier models evaluated in the study; percentage reductions presented in the paper.
Olava Extract achieved the strongest aggregate performance in the study, with a macro F1 of 0.812 and a micro F1 of 0.842.
Reported evaluation results comparing Olava Extract to five frontier models on structured contract extraction; explicit macro and micro F1 scores presented in the paper.
high positive A Few Good Clauses: Comparing LLMs vs Domain-Trained Small L... F1 score (macro and micro)
The main finding (that the reform increases grain yield) is robust to multiple checks, including parallel trend tests, placebo tests, propensity score matching DID (PSM-DID), and exclusion of special samples.
Battery of robustness tests reported in the paper: parallel trend tests, placebo tests, PSM-DID estimation, and analyses excluding special samples.
The grain-yield-enhancing effect is stronger in areas with stronger environmental regulation intensity.
Heterogeneity analysis in the paper comparing regions by environmental regulation intensity.
The grain-yield-enhancing effect is stronger in regions with higher levels of digital economy development.
Heterogeneity analysis dividing sample by regional digital economy development level.
The grain-yield-enhancing effect of the water resource tax reform is more pronounced in non-major grain-producing areas.
Heterogeneity analysis reported in the study comparing effects across major vs. non-major grain-producing regions.
The reform enhances regional green innovation, which contributes to higher grain yield by strengthening water-use efficiency and agricultural productivity.
Mechanism analysis presented in the study showing increases in measures of regional green innovation after the tax reform.
The water resource tax reform significantly increases grain yield.
Quasi-natural experiment using the pilot 'fee-to-tax' reform; panel dataset of Chinese prefecture-level cities, 2013–2019; multi-period difference-in-differences (DID) estimation supplemented by double machine learning and multiple robustness tests.
On missing value reconstruction, Schema-1 achieves lower reconstruction error than all classical statistical methods and frontier large language models on mean performance across conditions.
Empirical missing-value imputation experiments comparing Schema-1 to classical statistical methods and large language models, reporting lower mean reconstruction error across tested conditions (specific methods, conditions, and metrics not provided in the abstract).
high positive Data Language Models: A New Foundation Model Class for Tabul... missing value reconstruction error (imputation error)
Schema-1 outperforms gradient-boosted ensembles, AutoML stacks, and the tabular foundation models we evaluate on established row-level prediction benchmarks.
Empirical evaluation on established row-level prediction benchmarks comparing Schema-1 to gradient-boosted ensembles, AutoML stacks, and evaluated tabular foundation models (benchmarks and numerical results not detailed in the abstract).
high positive Data Language Models: A New Foundation Model Class for Tabul... row-level prediction performance (benchmark predictive performance)
Schema-1 is the first DLM: a 140M parameter model trained on more than 2.3M synthetic and real-world tabular datasets.
Model specification reported in the paper (explicit parameter count and training dataset count).
high positive Data Language Models: A New Foundation Model Class for Tabul... model size and training data scale (number of datasets)
A Data Language Model (DLM) understands tables the way a language model understands sentences: natively, without serialization or preprocessing, directly from raw cell values.
Model design and description presented in the paper (Schema-1 is given as an instance of a DLM); claimed capability based on architecture and training approach.
high positive Data Language Models: A New Foundation Model Class for Tabul... ability to consume tabular data directly from raw cell values without preprocess...
Under conditions of strong productivity growth, high-skill complementarity, low obsolescence, and broad ownership, automation raises output, capital, and consumption.
Comparative-static results from the heterogeneous-agent general-equilibrium model calibrated/analyzed under parameter configurations (strong productivity growth, high-skill complementarity, low obsolescence, broad ownership).
high positive The Demand Externality of Automation aggregate output, capital stock, aggregate consumption
Automation raises productivity.
Analytical results from a theoretical framework: a static benchmark and a stationary heterogeneous-agent general equilibrium model in which firms choose automation from a profit function and final-good production is Cobb–Douglas.
high positive The Demand Externality of Automation productivity (aggregate output per input)
Human performance on the benchmark is 80.7%.
Human baseline reported in the paper (same evaluation/rubrics as agents).
high positive Workspace-Bench 1.0: Benchmarking AI Agents on Workspace Tas... benchmark score (human performance)
We provide Workspace-Bench-Lite, a 100-task subset that preserves the benchmark distribution while reducing evaluation costs by about 70%.
Description of a reduced-size benchmark split (100 tasks) and reported cost reduction (~70%) in the paper.
high positive Workspace-Bench 1.0: Benchmarking AI Agents on Workspace Tas... evaluation cost (and distributional fidelity of the subset)