The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests About 🎲 Workforce Futures
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Evidence (8974 claims)

Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.

The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).

Browse by theme

Nine broad, paper-level topics. Click one to filter the claims below.

Adoption
10085 claims
Filter claims →
Productivity
8974 claims
Filtered →
Governance
8062 claims
Filter claims →
Human-AI Collaboration
7749 claims
Filter claims →
Org Design
5057 claims
Filter claims →
Innovation
4896 claims
Filter claims →
Labor Markets
4088 claims
Filter claims →
Skills & Training
3372 claims
Filter claims →
Inequality
2377 claims
Filter claims →

Claims by outcome category

Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.

Outcome Positive Negative Mixed Null Total
Other 882 244 117 1097 2424
Governance & Regulation 1010 469 229 135 1875
Organizational Efficiency 977 235 149 90 1462
Technology Adoption Rate 781 299 143 128 1362
Research Productivity 506 155 74 363 1110
Output Quality 555 219 71 70 915
Decision Quality 395 200 95 54 751
Firm Productivity 523 67 101 27 724
AI Safety & Ethics 262 309 75 36 688
Market Structure 195 201 135 30 566
Task Allocation 248 77 96 38 464
Innovation Output 300 34 55 20 411
Skill Acquisition 207 75 65 21 368
Employment Level 138 67 119 24 350
Fiscal & Macroeconomic 156 80 53 33 329
Task Completion Time 211 38 13 16 280
Firm Revenue 183 52 29 5 270
Consumer Welfare 131 77 48 13 269
Inequality Measures 50 141 54 9 254
Worker Satisfaction 104 85 25 13 227
Error Rate 87 112 11 5 215
Automation Exposure 69 69 37 20 198
Wages & Compensation 102 49 31 11 193
Team Performance 115 30 30 11 187
Regulatory Compliance 88 74 17 7 186
Training Effectiveness 109 22 14 21 168
Developer Productivity 116 21 15 8 161
Job Displacement 12 92 26 1 131
Hiring & Recruitment 57 12 9 5 83
Skill Obsolescence 6 59 10 2 77
Social Protection 43 17 8 2 70
Creative Output 35 21 9 4 70
Labor Share of Income 18 23 17 1 59
Worker Turnover 15 16 4 35
Industry 1 1
Clear
Productivity Remove filter
The paper evaluates 'Spec Kit' and 'TDAD' as instantiations of the SGM via a four-month pilot study.
Empirical pilot evaluation reported in the paper; duration specified as four months. Sample size or number of teams/participants in pilot not specified in the summary.
high null result The Productivity-Reliability Paradox: Specification-Driven G... evaluation of SGM instantiations (Spec Kit, TDAD) over four months
The paper identifies two amplifying mechanisms for PRP: the code review bottleneck and the context window constraint.
Theoretical argumentation in the paper naming two mechanisms that amplify the PRP phenomenon (qualitative explanation).
high null result The Productivity-Reliability Paradox: Specification-Driven G... mechanisms amplifying productivity-reliability trade-off
The paper formally defines PRP with three moderating variables: task abstraction, codebase maturity, and developer experience.
Theoretical/formal definition presented in the paper identifying three moderators; claim is descriptive of the paper's conceptual model.
high null result The Productivity-Reliability Paradox: Specification-Driven G... presence/definition of moderating variables for PRP
This paper conducted a multivocal literature review of 67 sources spanning 2022–2026.
Statement of method in the paper describing the literature review (count of sources = 67).
high null result The Productivity-Reliability Paradox: Specification-Driven G... study corpus size (number of sources reviewed)
Telemetry across 10,000+ developers shows flat delivery metrics (no improvement in delivery outcomes) despite changes in PR and review behavior.
Observational telemetry across >10,000 developers reported in the paper; described result is no meaningful change in delivery metrics (e.g., delivery throughput, lead time) despite increases in PRs and longer reviews.
high null result The Productivity-Reliability Paradox: Specification-Driven G... delivery metrics (throughput/lead time)
Public inference benchmarks compare AI systems at the model and provider level, but the unit at which deployment decisions are actually made is the endpoint: the (provider, model, stock-keeping-unit) tuple at which a specific quantization, decoding strategy, region, and serving stack is exposed.
Author assertion / methodological observation about how public benchmarks report results versus how deployments are decided; no empirical test reported in the excerpt.
high null result Token Arena: A Continuous Benchmark Unifying Energy and Cogn... granularity of benchmarking vs. deployment decision unit (endpoint = provider, m...
A symbolic lifting operator translates simulator trajectories into qualitative descriptors, motion labels, temporal predicates, and structural diagnostics that models interpret across iterative design cycles.
Architectural/methodological contribution described in the paper: a symbolic lifting operator that converts simulator trajectories into higher-level symbolic diagnostics used by LM agents during iterative refinement.
high null result Language Models Refine Mechanical Linkage Designs Through Sy... representation of simulator output (symbolic descriptors)
Language model agents explore discrete topologies while numerical optimisers fit continuous parameters.
Methodological description of the system architecture in the paper (division of labor between LM agents for discrete topology search and numerical optimisers for continuous parameter fitting).
high null result Language Models Refine Mechanical Linkage Designs Through Sy... task allocation between symbolic/discrete search and numerical optimisation
The review uses a collection of qualitative and quantitative approaches (i.e., it synthesizes both qualitative and quantitative studies).
Explicit methodological description in the abstract indicating mixed-methods literature synthesis.
high null result A Comprehensive Review of Technology Adoption and Its Impact... review methodology (use of qualitative and quantitative approaches)
A collection of qualitative and quantitative approaches reveals predictors of technological integration, including organisational preparedness, economic factors, policies, and human capital.
Statement about the review's synthesized findings from multiple qualitative and quantitative studies identifying these predictors; method = mixed-methods literature synthesis.
high null result A Comprehensive Review of Technology Adoption and Its Impact... predictors of technological integration
The primary technologies covered in this review are Electronic Health Records (EHR), telemedicine, artificial intelligence (AI), and the Internet of Things (IoT).
Explicit topical scope statement in the paper (description of review subjects); based on the paper's own selection of topics for review.
high null result A Comprehensive Review of Technology Adoption and Its Impact... topics covered (EHR, telemedicine, AI, IoT)
There is little empirical exploration of how professionals making high-stakes decisions perceive their agency and level of control when working with genAI systems.
Statement about a gap in the existing literature made by the authors (literature review / framing); no sample size (gap claim).
high null result Resume-ing Control: (Mis)Perceptions of Agency Around GenAI ... availability of empirical research on professionals' perceptions of agency/contr...
AI adoption has no detectable effects on overall employment.
Difference-in-differences estimates using administrative employment totals linked to survey-reported adoption show no statistically significant change in total employment.
As of 2024, AI adoption remains limited: about 10 per cent of firms report current use.
Newly collected firm-level survey data linked to administrative balance sheet and employer–employee records; prevalence reported in 2024 survey.
high null result The economic impact of artificial intelligence: evidence fro... current AI adoption rate
The empirical analysis uses panel data from 3,515 Chinese A-share listed firms, totaling 20,076 firm-year observations covering 2014–2022.
Statement of data and sample in the paper (sample frame and time period explicitly given).
high null result A Data-Driven Evaluation Framework for Quantifying the Impac... sample coverage / dataset size
The literature review employs the PRISMA model to screen, identify, and synthesize available literature on AI, Machine Learning and Deep Learning in promoting managerial productivity and task efficiency.
Methodological statement in the paper's abstract (explicitly states use of PRISMA for screening and synthesis).
high null result Artificial intelligence, machine learning, and deep learning... literature search and synthesis method (PRISMA use)
Methodological basis: the study used analysis of aggregated industry data and a scenario approach; information sources were Russian-language materials including the Ministry of Digital Development, HSE, the Autonomous Non-Profit Organization 'Digital Economy', and analytical reviews.
Explicit methodological and data-source statements in the paper.
high null result THE IMPACT OF AI ON POTENTIAL GDP AND LONG-TERM ECONOMIC GRO... methodological approach and data sources
All artifacts associated with this study are publicly available at https://zenodo.org/records/18489222.
Statement in the paper providing a Zenodo link to artifacts.
high null result The Impact of LLM-Assistants on Software Developer Productiv... availability of study artifacts
This review identifies key research gaps and provides recommendations for future research and practice.
Authors' discussion and conclusion sections synthesizing gaps and offering recommendations based on the mapping results.
high null result The Impact of LLM-Assistants on Software Developer Productiv... research gaps and recommendations (qualitative synthesis)
Satisfaction, Performance, and Efficiency are the most frequently investigated SPACE dimensions, whereas Communication and Activity remain underexplored.
Frequency counts and synthesis across the 39 included studies mapped to SPACE dimensions as reported by the authors.
high null result The Impact of LLM-Assistants on Software Developer Productiv... frequency of SPACE dimensions studied
Only 15% of the reviewed studies extend beyond three SPACE dimensions.
Authors' coding of included studies against the SPACE framework with reported proportion.
high null result The Impact of LLM-Assistants on Software Developer Productiv... proportion of studies examining >3 SPACE dimensions
90% of the reviewed studies adopt a multi-dimensional perspective by examining at least two SPACE dimensions.
Authors' coding of included studies against the SPACE framework, yielding the reported proportion.
high null result The Impact of LLM-Assistants on Software Developer Productiv... proportion of studies examining >=2 SPACE dimensions
This paper is a systematic review and mapping of 39 peer-reviewed studies published between January 2014 and December 2024 that examine the impact of LLM-assistants on software developer productivity.
Authors conducted a systematic review and mapping exercise covering peer-reviewed studies within the stated date range; the paper reports the count of included studies as 39.
high null result The Impact of LLM-Assistants on Software Developer Productiv... scope of literature reviewed (count of studies)
Cross-stage correlations are very weak: parsing->retrieval r = 0.14, parsing->generation r = 0.17, retrieval->generation r = 0.02.
Reported Pearson (or Spearman) correlation coefficients between stage-level metrics in the benchmark; exact correlation method not specified in excerpt.
high null result Benchmarking Complex Multimodal Document Processing Pipeline... correlation between stage-level quality metrics
We evaluate SecMate in a controlled study with 144 participants and 711 conversations.
Reported experimental study sample and conversation counts in the paper.
high null result SecMate: Multi-Agent Adaptive Cybersecurity Troubleshooting ... study sample size and conversation count
Our architecture combines a two-layer Graph Convolutional Network (GCN) encoder, twin critics, and a value network that drives the adversary.
Model architecture description in the paper specifying a 2-layer GCN encoder, twin critics, and a value network used for adversary control.
high null result Semi-Markov Reinforcement Learning for City-Scale EV Ride-Ha... model architecture components (2-layer GCN encoder, twin critics, adversary-driv...
The robust backup uses the Kantorovich--Rubinstein dual, a projected subgradient inner loop, and a primal--dual risk-budget update.
Algorithmic description in the paper detailing the robust backup solver components (Kantorovich--Rubinstein dual, projected subgradient, primal-dual update).
high null result Semi-Markov Reinforcement Learning for City-Scale EV Ride-Ha... robust backup algorithm design and optimization procedure
To mitigate distributional shifts, we optimize a Soft Actor--Critic (SAC) agent against a Wasserstein-1 ambiguity set with a graph-aligned Mahalanobis ground metric that captures spatial correlations.
Methodological description of a robust training objective: SAC optimized under a Wasserstein-1 ambiguity set using a graph-aligned Mahalanobis metric to encode spatial correlations.
high null result Semi-Markov Reinforcement Learning for City-Scale EV Ride-Ha... robustness to distributional shift via Wasserstein-1 ambiguity set with graph-al...
These intentions are projected at every decision step through a time-limited rolling mixed-integer linear program (MILP) that strictly enforces state-of-charge, port, and feeder constraints.
Method/algorithm description in the paper: a rolling MILP projection component implemented to enforce physical constraints (state-of-charge, charger port limits, feeder limits) at each decision step.
high null result Semi-Markov Reinforcement Learning for City-Scale EV Ride-Ha... constraint compliance via MILP projection (state-of-charge, port, feeder constra...
The policy learns over high-level intentions produced by a masked, temperature-annealed actor.
Method/algorithm description in the paper describing the actor design (masked, temperature-annealed) and the high-level intentions used for policy learning.
high null result Semi-Markov Reinforcement Learning for City-Scale EV Ride-Ha... policy representation (high-level intentions from masked, temperature-annealed a...
We formulate the problem as a hex-grid semi-Markov decision process (semi-MDP) with mixed actions -- discrete actions for serving, repositioning, and charging, together with continuous charging power -- and variable action durations.
Methodological description in the paper presenting the model formulation (hex-grid semi-MDP) and action space design; no external dataset required.
high null result Semi-Markov Reinforcement Learning for City-Scale EV Ride-Ha... problem formulation (hex-grid semi-MDP with mixed and continuous actions and var...
The analysis employs rigorous econometric methods including difference-in-differences estimation and propensity score matching to control for confounding variables across industry (NAICS 2-digit), firm size, geographic location, occupation-level characteristics, and macroeconomic conditions.
Methodological description in the paper specifying DiD and propensity score matching and listed covariates/controls.
high null result The Generative AI Revolution: Early Evidence of Structural T... methodological controls / identification strategy
The study uses U.S. Census Bureau Business Trends and Outlook Survey data tracking over 1.2 million businesses.
Paper statement that it incorporates the Census Bureau Business Trends and Outlook Survey covering >1,200,000 businesses.
high null result The Generative AI Revolution: Early Evidence of Structural T... business-level observations (adoption/behavior)
The analysis integrates the Anthropic Economic Index capturing approximately one million AI usage interactions.
Paper statement that the Anthropic Economic Index was used and captures ~1,000,000 AI usage interactions.
high null result The Generative AI Revolution: Early Evidence of Structural T... AI usage interactions (adoption/usage)
Semantic search maintained comparable inter-rater agreement while reducing chart abstraction time.
Clinical utility evaluation reports that inter-rater agreement was comparable between semantic-search-assisted abstraction and clinician-performed chart review.
The authors optimized embedding model and chunking strategy using a physician-authored benchmark dataset.
Methods: experiment described as optimization of embedding model and chunking using a physician-authored benchmark dataset.
high null result Health System Scale Semantic Search Across Unstructured Clin... model_and_chunking_configuration
The system uses instruction-tuned qwen3-embedding-0.6B embeddings, stores vectors in a managed database with storage-optimized indexing, maintains full-text metadata in a low-latency key-value store, and operates within a HIPAA-compliant governance framework.
Methods description of system architecture and governance provided in the paper.
high null result Health System Scale Semantic Search Across Unstructured Clin... system_architecture / governance_compliance
We deployed a semantic search system indexing 166 million clinical notes (484 million vectors) from 1.68 million patients.
Paper reports a production deployment at a large children's hospital and gives exact index counts: 166 million clinical notes, 484 million vectors, 1.68 million patients.
high null result Health System Scale Semantic Search Across Unstructured Clin... number_of_notes_indexed / index_size
Through a rigorous sensitivity analysis of resource scarcity and temporal dominance, we quantify the coordination gap.
Methodological description in the paper indicating the authors performed a systematic sensitivity analysis across environmental parameters (resource scarcity and temporal dominance) to measure performance differences between training modalities.
On the n=11 subset with published SWE-bench scores, composite and benchmark-only rankings are nearly uncorrelated (ρ_s=0.25).
Spearman rank correlation between composite rankings and benchmark-only rankings on an 11-agent subset that has published SWE-bench scores; reported correlation.
high null result AgentPulse: A Continuous Multi-Signal Framework for Evaluati... rank correlation between composite ranking and benchmark-only ranking
We instrument ITAS, a four-agent tutoring system built on Gemini 2.5 Flash and Google Vertex AI, across three throughput tiers (Standard PayGo, Priority PayGo, and Provisioned Throughput) and eleven concurrency levels up to 50 simultaneous users, producing over 3,000 requests drawn from a live graduate STEM deployment.
Methods statement in paper describing experimental setup: four-agent ITAS built on Gemini 2.5 Flash and Google Vertex AI; three throughput tiers; eleven concurrency levels up to 50; over 3,000 requests from a live graduate STEM deployment.
high null result Latency and Cost of Multi-Agent Intelligent Tutoring at Scal... instrumented request sample (number of requests and concurrency levels)
The welfare consequences of genAI can be organized by a two-dimensional taxonomy: the strength of the incentive to perform the task without AI, and the severity of model collapse.
Analytical organization derived from the theoretical model presented in the paper (conceptual taxonomy based on model parameters; no empirical sample reported in abstract).
high null result Generative artificial intelligence reduces social welfare th... social welfare outcomes as a function of incentive strength and model collapse s...
We develop a parsimonious model of behavior in collaborative interactions in which individuals can either exert human effort, rely on genAI, or refrain from work altogether.
Methodological claim: authors present a formal theoretical model with the specified choice set (model description in paper; no empirical sample reported in abstract).
high null result Generative artificial intelligence reduces social welfare th... choice among effort modalities (human effort, genAI reliance, abstention)
Predictive performance exhibits saturation beyond a certain context length.
Experiments varying the context (input) length in foundation models and observing changes in forecasting performance; reported saturation effect in analyses.
high null result FETS Benchmark: Foundation Models Outperform Dataset-specifi... change in forecast accuracy as context length increases
Task difficulty rated by human experts only weakly aligns with actual token costs, revealing a fundamental gap between human-perceived complexity and the computational effort agents actually expend.
Analysis comparing human expert difficulty ratings to measured token costs for tasks in SWE-bench Verified; weak alignment reported in the paper between ratings and token consumption.
high null result How Do AI Agents Spend Your Money? Analyzing and Predicting ... correspondence/alignment between human-rated task difficulty and measured token ...
Higher token usage does not translate into higher accuracy; accuracy often peaks at intermediate cost and saturates at higher costs.
Comparison of accuracy (task success) versus total token usage across runs/trajectories in the agentic coding experiments on SWE-bench Verified; reported observed relationship (peak at intermediate costs and saturation thereafter).
high null result How Do AI Agents Spend Your Money? Analyzing and Predicting ... task accuracy as a function of token usage
Learning-based control offers a more adaptive alternative, but it remains unclear whether such methods... can sustain hours of reliable operation, deliver consistent quality, and behave safely around people on a live production line.
Framing of a research gap in the paper's introduction; no primary experimental data presented here (statement of uncertainty motivating the study).
high null result Learning-augmented robotic automation for real-world manufac... operational reliability, product quality consistency, safety around people for l...
Die Studie basiert auf einer wiederholten Querschnittsbefragung lizenzierter Beschäftigter einer außeruniversitären Forschungseinrichtung.
Autorenangabe im Abstract: wiederholte Querschnittsbefragung (survey) unter lizenzieren Beschäftigten der untersuchten Forschungseinrichtung; methodische Beschreibung im Abstract.
high null result Generative KI in der Wissensarbeit: Wahrnehmung, Nutzen und ... Studiendesign / Datengrundlage (repeated cross-sectional survey)
An exploratory evaluation compared unstructured vibe coding, structured prompt engineering, and the Shift-Up approach in the development of a web application.
Paper reports an exploratory evaluation / comparative study described in the abstract; the task context is a web application development exercise comparing three approaches (no sample size reported in abstract).
high null result Shift-Up: A Framework for Software Engineering Guardrails in... comparative evaluation of development approaches
This review was conducted following the guidelines of the Preferred Reporting of Items in a Systematic Review and Meta-Analysis (PRISMA).
Methodological statement in the paper's abstract indicating PRISMA adherence; no further protocol details or study counts provided in the abstract.
high null result Artificial Intelligence, Public Policy and Governance - impl... methodological adherence to PRISMA reporting standards