The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests About 🎲 Workforce Futures
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Evidence (8974 claims)

Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.

The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).

Browse by theme

Nine broad, paper-level topics. Click one to filter the claims below.

Adoption
10085 claims
Filter claims →
Productivity
8974 claims
Filtered →
Governance
8062 claims
Filter claims →
Human-AI Collaboration
7749 claims
Filter claims →
Org Design
5057 claims
Filter claims →
Innovation
4896 claims
Filter claims →
Labor Markets
4088 claims
Filter claims →
Skills & Training
3372 claims
Filter claims →
Inequality
2377 claims
Filter claims →

Claims by outcome category

Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.

Outcome Positive Negative Mixed Null Total
Other 882 244 117 1097 2424
Governance & Regulation 1010 469 229 135 1875
Organizational Efficiency 977 235 149 90 1462
Technology Adoption Rate 781 299 143 128 1362
Research Productivity 506 155 74 363 1110
Output Quality 555 219 71 70 915
Decision Quality 395 200 95 54 751
Firm Productivity 523 67 101 27 724
AI Safety & Ethics 262 309 75 36 688
Market Structure 195 201 135 30 566
Task Allocation 248 77 96 38 464
Innovation Output 300 34 55 20 411
Skill Acquisition 207 75 65 21 368
Employment Level 138 67 119 24 350
Fiscal & Macroeconomic 156 80 53 33 329
Task Completion Time 211 38 13 16 280
Firm Revenue 183 52 29 5 270
Consumer Welfare 131 77 48 13 269
Inequality Measures 50 141 54 9 254
Worker Satisfaction 104 85 25 13 227
Error Rate 87 112 11 5 215
Automation Exposure 69 69 37 20 198
Wages & Compensation 102 49 31 11 193
Team Performance 115 30 30 11 187
Regulatory Compliance 88 74 17 7 186
Training Effectiveness 109 22 14 21 168
Developer Productivity 116 21 15 8 161
Job Displacement 12 92 26 1 131
Hiring & Recruitment 57 12 9 5 83
Skill Obsolescence 6 59 10 2 77
Social Protection 43 17 8 2 70
Creative Output 35 21 9 4 70
Labor Share of Income 18 23 17 1 59
Worker Turnover 15 16 4 35
Industry 1 1
Clear
Productivity Remove filter
AI integration reduces compliance costs.
Empirical results reported in the paper indicate a reduction in compliance costs associated with AI integration, based on System GMM estimation and FE/RE robustness checks (no numeric cost savings provided in the supplied text).
AI integration significantly enhances customer satisfaction.
Paper reports statistically significant positive association between AI integration and customer satisfaction using System GMM and robustness checks (no details on customer satisfaction measurement or sample size in the supplied text).
AI integration significantly enhances risk-adjusted returns.
Reported empirical results using System GMM with FE and RE robustness checks; the paper states statistical significance but does not provide effect magnitudes in the supplied summary.
AI integration significantly enhances operational efficiency.
Same empirical analysis using System GMM, with FE and RE models for robustness (no sample size or numeric estimates provided in the supplied text).
AI integration significantly enhances return on assets (ROA).
Empirical analysis reported in the paper using System Generalized Method of Moments (System GMM) estimator, with Fixed Effects (FE) and Random Effects (RE) models used as robustness checks. (No sample size or test statistics provided in the text supplied.)
The index diverges sharply from existing AI exposure measures for specific occupation groups: power plant operators, railroad conductors, and aircraft cargo handling supervisors score high on RL feasibility but low on general AI exposure.
Empirical comparison between the RL Feasibility Index and existing AI-exposure measures, with named occupation groups showing opposite rankings.
high positive What Jobs Can AI Learn? Measuring Exposure by Reinforcement ... relative RL feasibility vs. general AI exposure for named occupations
Using LLM annotators guided by a rubric developed with RL experts and validated against confirmed deployment cases, we score all 17,951 O*NET tasks for training feasibility and aggregate to the occupation level, producing an RL Feasibility Index.
Empirical method described in paper: LLM-based annotation process guided by expert-developed rubric; validation against confirmed deployment cases; explicit enumeration of 17,951 O*NET tasks scored and aggregated into an index.
high positive What Jobs Can AI Learn? Measuring Exposure by Reinforcement ... training feasibility of O*NET tasks; RL Feasibility Index at task and occupation...
We examine this for every occupation in the US economy.
Statement of study scope in the paper (methodological claim about coverage).
high positive What Jobs Can AI Learn? Measuring Exposure by Reinforcement ... coverage of US occupations in the RL feasibility analysis
The no-talk baseline establishes that communication is necessary.
Experimental no-talk baseline showing worse coordination without communication between agents.
high positive Talk is Cheap, Communication is Hard: Dynamic Grounding Fail... coordination performance with vs without communication
These results highlight dynamic grounding as a critical and understudied axis of multi-agent coordination.
Synthesis/interpretation of the experimental findings reported in the paper.
high positive Talk is Cheap, Communication is Hard: Dynamic Grounding Fail... importance of dynamic grounding for multi-agent coordination
We introduce an iterated, multi-turn negotiation game in which two agents allocate shared resources toward private projects with verifiable jointly optimal outcomes.
Methodological contribution described in the paper (design of a new multi-turn negotiation game).
high positive Talk is Cheap, Communication is Hard: Dynamic Grounding Fail... existence of a multi-turn negotiation benchmark with verifiable optimal outcomes
Grounding is the collaborative process of establishing mutual belief sufficient for the current communicative purpose.
Conceptual/definitional statement presented by the authors (no empirical data reported).
Policy recommendations: invest in digital infrastructure, human capital development, and inclusive technology diffusion strategies to ensure more equitable distribution of AI-driven economic value.
Policy implications drawn from study findings (heterogeneous effects and mediation by structural conditions).
high positive The Economic Value of Agentic AI: A Comparative Analysis of ... equitable distribution of AI-driven economic value (policy interventions)
The magnitude of AI's growth effects varies across economic contexts: developed economies experience substantially stronger growth impacts (approximately 0.33) than emerging economies (approximately 0.15).
Heterogeneity analysis / subgroup comparisons (developed vs emerging economies) using the panel data regressions and/or quantile regressions on the 2015–2024 dataset; exact sample sizes per subgroup not reported.
high positive The Economic Value of Agentic AI: A Comparative Analysis of ... economic growth (heterogeneous treatment effects by country group)
AI adoption has a comparatively weaker direct effect on economic growth (direct effect β = 0.09).
Mediation/structural decomposition from the paper showing direct (non-mediated) coefficient from AI adoption to growth.
high positive The Economic Value of Agentic AI: A Comparative Analysis of ... economic growth (direct effect)
Agentic AI influences economic growth primarily through a productivity channel (mediated effect β = 0.35, p < 0.01).
Mediation analysis (panel data) estimating indirect effect of AI adoption on GDP growth via measured productivity channel; data sources: World Bank and OECD indicators, 2015–2024.
high positive The Economic Value of Agentic AI: A Comparative Analysis of ... economic growth (mediated via productivity)
AI adoption significantly improves firm-level productivity (β = 0.18, p < 0.01).
Fixed-effects panel regression using an AI Adoption Index as predictor on firm-level productivity; data drawn from World Bank (World Development Indicators and Enterprise Surveys) and OECD AI indicators for 2015–2024 (sample size not reported in text).
Agentic AI has strong potential to boost productivity and growth.
Statement in paper motivated by literature review and the study's empirical results linking AI adoption to productivity and growth.
high positive The Economic Value of Agentic AI: A Comparative Analysis of ... productivity and economic growth (general)
HAAS can serve as a pre-deployment workbench for comparing and inspecting human–AI allocation policies before organisational commitment.
Claim about intended use and demonstration of HAAS as an implemented tool; based on the framework implementation and benchmark experiments reported. No deployment-scale evaluation or sample sizes provided in the excerpt.
high positive HAAS: A Policy-Aware Framework for Adaptive Task Allocation ... ability to compare and inspect allocation policies prior to deployment
In manufacturing, stronger governance can improve operational performance and reduce fatigue simultaneously — a workload-buffering effect.
Domain-specific empirical result reported for the manufacturing benchmark in the paper, comparing operational performance and fatigue under different governance strengths. No numeric sample size or effect sizes provided in the excerpt.
high positive HAAS: A Policy-Aware Framework for Adaptive Task Allocation ... operational performance and worker fatigue
Task–agent fit is represented through five auditable cognitive dimensions and a five-mode autonomy spectrum (from human-only to fully autonomous) embedded in a reproducible benchmark spanning software engineering and manufacturing.
Design and benchmark description within the paper; specification of five cognitive dimensions and a five-mode autonomy spectrum and a reproducible benchmark across two domains. No numeric sample size provided.
high positive HAAS: A Policy-Aware Framework for Adaptive Task Allocation ... representation of task–agent fit and benchmarking across domains
HAAS combines a rule-based expert system that enforces governance constraints before any learning occurs, and a contextual-bandit learner that selects among feasible collaboration modes from outcome feedback.
Descriptive claim about the implemented HAAS framework as presented in the paper; method description of system architecture (rule-based expert system + contextual-bandit learner). No sample size reported.
high positive HAAS: A Policy-Aware Framework for Adaptive Task Allocation ... mechanism for adaptive task allocation (selected collaboration mode)
The field's near-term research agenda should explicitly include collecting and using triadic data.
Normative recommendation in the paper; presented as the authors' advised research priority rather than empirically justified within the excerpt.
high positive The Conversations Beneath the Code: Triadic Data for Long-Ho... inclusion of triadic data collection/use in near-term research agendas in the SW...
This data is the empirical key to four open questions in agent training.
Argumentative claim in the paper asserting centrality of triadic data to addressing unspecified four open research questions; no empirical demonstration included in the excerpt.
high positive The Conversations Beneath the Code: Triadic Data for Long-Ho... resolvability of four open questions in agent training using triadic data
This triadic data is capturable in 12-18 months with methods already mature in adjacent fields.
Claim in the paper based on authors' assessment of methodological maturity in adjacent fields; no empirical project timeline or pilot data is provided in the excerpt.
high positive The Conversations Beneath the Code: Triadic Data for Long-Ho... time required to collect a triadic dataset using existing methods
Any such corpus -- triadic or otherwise -- must justify its quality to a fine-tuning researcher through a four-tier evidence framework: mechanical verification, statistical corpus characterization, probe experiments, and pre-registered blind evaluation.
Methodological proposal in the paper outlining a four-tier evidence framework; presented as normative guidance rather than validated by application to a corpus in the excerpt.
high positive The Conversations Beneath the Code: Triadic Data for Long-Ho... quality and trustworthiness of fine-tuning corpora as judged by the four-tier fr...
The canonical instantiation of triadic data is two complementary products: long-horizon expert trajectories captured under stimulated-recall protocols, and simulated cross-functional companies -- instrumented teams of senior engineers, product managers, designers, and data scientists working through ambiguous deliverables on shared infrastructure.
Prescriptive specification in the paper proposing two concrete dataset types as canonical instantiations; presented as design/recommendation rather than empirically tested.
high positive The Conversations Beneath the Code: Triadic Data for Long-Ho... availability and suitability of dataset modalities (stimulated-recall expert tra...
The substrate for the next generation of software-engineering (SWE) agents is neither larger GitHub scrapes nor more solo-agent trajectories nor -- sufficient by itself -- open human-AI dialogue logs; it is triadic data: synchronized capture of the human-human conversations where engineering context is formed, the human-AI sessions where that context is partially consumed, and the multi-week cross-functional work that surrounds both.
Argument and conceptual proposal in the paper; no empirical validation or comparative experiments are provided in the excerpt.
high positive The Conversations Beneath the Code: Triadic Data for Long-Ho... effectiveness of training data substrates for improving agent performance on lon...
SCDPs are a useful framework for policy simulation for the digital economy, mechanism design for information systems, and digital twin modeling of cyberinfrastructure.
Paper posits these applications as prospective uses of the framework (argumentative/speculative; no empirical evaluation reported in abstract).
high positive The Design and Composition of Structural Causal Decision Pro... usefulness for policy simulation, mechanism design, and digital twin modeling
SCDPs are capable of modeling variable discounting, a tool used widely in social scientific modeling.
Paper states the capability as part of SCDP definition and examples (theoretical claim).
high positive The Design and Composition of Structural Causal Decision Pro... modeling of variable discounting
An SCDP can endogenously model the memory-formation process and is thus useful for modeling resource‑rational agents in dynamic settings.
Paper asserts SCDP can represent memory-formation endogenously and discusses application to resource-rational agents (theoretical modeling capability).
high positive The Design and Composition of Structural Causal Decision Pro... ability to model endogenous memory formation / resource-rational agents
SCDPs are strictly more expressive than POMDPs because they do not assume rational belief formation.
Comparative expressiveness claim stated in the paper; supported by theoretical argument or formal separation result (paper text states the claim explicitly).
high positive The Design and Composition of Structural Causal Decision Pro... expressiveness relative to POMDPs (ability to represent non-rational belief form...
SCDPs inherit the composition properties of SCDMs (i.e., SCDPs benefit from SCDM composability).
Logical consequence argued in the paper from SCDP being constructed from SCDMs; likely supported by formal argumentation in the text.
high positive The Design and Composition of Structural Causal Decision Pro... inheritance of composability by SCDPs
A Structural Causal Decision Process (SCDP) is defined as a recurring SCDM with a discount variable.
Formal definition introduced in the paper (theoretical definition).
high positive The Design and Composition of Structural Causal Decision Pro... definition of SCDP as recurring SCDM with discounting
SCDMs have a well-defined and computationally useful property of composability.
Paper states and demonstrates ("We show") composability property — presumably via formal proofs or constructive arguments in the text (theoretical proofs/exposition).
high positive The Design and Composition of Structural Causal Decision Pro... composability of causal decision models
SCDMs can have open root variables for which no probability distribution or structural equation is given.
Model definitions in the paper explicitly allow open root variables (theoretical description).
high positive The Design and Composition of Structural Causal Decision Pro... support for open root variables in model formalism
In SCDMs, agent decisions can be constrained by their causal antecedents (i.e., decisions can be constrained by their causal parents).
Model specification and definitions in the paper describing constraints on decisions as part of SCDM structure (theoretical construction).
high positive The Design and Composition of Structural Causal Decision Pro... decision constraints by causal antecedents
Structural Causal Decision Models (SCDMs) expand on Structural Causal Influence Models by explicitly representing the causal relationships between model variables and the payoffs of agent decisions.
Formal model development and comparison to existing SCIMs provided in the paper (theoretical definitions and arguments).
high positive The Design and Composition of Structural Causal Decision Pro... explicit representation of causal relationships between variables and payoffs
We present two new classes of causal models of decision-making agents: Structural Causal Decision Models (SCDMs) and Structural Causal Decision Processes (SCDPs).
Paper introduces formal definitions for two model classes and describes their properties in the text (theoretical exposition).
high positive The Design and Composition of Structural Causal Decision Pro... introduction of new model classes (SCDMs and SCDPs)
We propose PAEF (Production Agentic Evaluation Framework), a five-dimension evaluation framework with an open-source reference implementation, designed for continuous evaluation on production traffic rather than episodic benchmark runs.
Author contribution: design and open-source implementation of PAEF described in the paper.
high positive Evaluating Agentic AI in the Wild: Failure Modes, Drift Patt... provision of a continuous, production-focused evaluation framework (PAEF)
The taxonomy and its failure modes are grounded in observations from systems operating at billion-event scale.
Author statement that observations underlying the taxonomy come from systems operating at billion-event scale.
high positive Evaluating Agentic AI in the Wild: Failure Modes, Drift Patt... empirical grounding (scale) of observations used to derive the taxonomy
This paper presents a taxonomy of seven failure modes unique to production agentic systems.
Author contribution: taxonomy presented in the paper (count = seven failure modes).
high positive Evaluating Agentic AI in the Wild: Failure Modes, Drift Patt... cataloging of distinct failure modes in production agentic systems
These findings provide insights for designing flexible yet reliable constraint-based workflows.
Synthesis and discussion of study results and technical evaluation in paper's conclusion.
high positive U-Define: Designing User Workflows for Hard and Soft Constra... design guidance for constraint-based workflows
User-defined constraint types improve user satisfaction.
Reported user study measures showing higher satisfaction for participants using U-Define compared to baselines (no sample size or numeric effects provided).
high positive U-Define: Designing User Workflows for Hard and Soft Constra... user satisfaction (self-reported)
User-defined constraint types improve performance.
Reported results from user studies and/or technical evaluation indicating better task performance when users can set hard/soft constraint types (no numeric effect size or sample size in excerpt).
high positive U-Define: Designing User Workflows for Hard and Soft Constra... performance (task success / quality of generated plans)
User-defined constraint types improve perceived usefulness.
Results from the reported user studies comparing U-Define (user-defined constraint types) to baselines; based on participant responses and measures of perceived usefulness (sample sizes/details not provided in excerpt).
high positive U-Define: Designing User Workflows for Hard and Soft Constra... perceived usefulness (user-reported)
U-Define verifies hard constraints using formal model checking and verifies soft constraints using an LLM-as-judge evaluation.
Description of the complementary verification methods employed in the U-Define system (technical design/implementation).
high positive U-Define: Designing User Workflows for Hard and Soft Constra... verification of constraint types (hard via model checking, soft via LLM evaluati...
We present U-Define, a system that lets users define constraints in natural language and categorize them as either hard rules that must not be violated or soft preferences that allow flexibility.
System implementation and description in paper (design and implementation of U-Define).
high positive U-Define: Designing User Workflows for Hard and Soft Constra... ability to specify constraints (natural-language input and categorization into h...
Evaluating AI applications in actual multi-turn interactions with human users, looking at usability and satisfaction besides accuracy, provides added value compared to focusing on benchmark performance only.
Argument/interpretation in the paper based on the study's multi-turn human-in-the-loop evaluation showing differences between objective performance gains and participant perceptions.
high positive Seeking Information with RAG-Assistants: Does Model Size Mat... evaluation methodology value (usability, satisfaction, accuracy)
Hybrid systems (human + RAG assistant) are beneficial in information-seeking scenarios.
Conclusion drawn from the experiment showing human-AI collaboration outperforms model-only baselines across model sizes in a realistic multi-turn information-seeking task with N=112 participants.
high positive Seeking Information with RAG-Assistants: Does Model Size Mat... task performance in information-seeking