The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests About 🎲 Workforce Futures
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Evidence (8974 claims)

Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.

The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).

Browse by theme

Nine broad, paper-level topics. Click one to filter the claims below.

Adoption
10085 claims
Filter claims →
Productivity
8974 claims
Filtered →
Governance
8062 claims
Filter claims →
Human-AI Collaboration
7749 claims
Filter claims →
Org Design
5057 claims
Filter claims →
Innovation
4896 claims
Filter claims →
Labor Markets
4088 claims
Filter claims →
Skills & Training
3372 claims
Filter claims →
Inequality
2377 claims
Filter claims →

Claims by outcome category

Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.

Outcome Positive Negative Mixed Null Total
Other 882 244 117 1097 2424
Governance & Regulation 1010 469 229 135 1875
Organizational Efficiency 977 235 149 90 1462
Technology Adoption Rate 781 299 143 128 1362
Research Productivity 506 155 74 363 1110
Output Quality 555 219 71 70 915
Decision Quality 395 200 95 54 751
Firm Productivity 523 67 101 27 724
AI Safety & Ethics 262 309 75 36 688
Market Structure 195 201 135 30 566
Task Allocation 248 77 96 38 464
Innovation Output 300 34 55 20 411
Skill Acquisition 207 75 65 21 368
Employment Level 138 67 119 24 350
Fiscal & Macroeconomic 156 80 53 33 329
Task Completion Time 211 38 13 16 280
Firm Revenue 183 52 29 5 270
Consumer Welfare 131 77 48 13 269
Inequality Measures 50 141 54 9 254
Worker Satisfaction 104 85 25 13 227
Error Rate 87 112 11 5 215
Automation Exposure 69 69 37 20 198
Wages & Compensation 102 49 31 11 193
Team Performance 115 30 30 11 187
Regulatory Compliance 88 74 17 7 186
Training Effectiveness 109 22 14 21 168
Developer Productivity 116 21 15 8 161
Job Displacement 12 92 26 1 131
Hiring & Recruitment 57 12 9 5 83
Skill Obsolescence 6 59 10 2 77
Social Protection 43 17 8 2 70
Creative Output 35 21 9 4 70
Labor Share of Income 18 23 17 1 59
Worker Turnover 15 16 4 35
Industry 1 1
Clear
Productivity Remove filter
Qwen3 embeddings with 300-token chunk size achieved 94.6% accuracy on a clinical question-answering benchmark.
Optimization experiment on a physician-authored clinical question-answering benchmark; best-performing configuration reported as qwen3 embeddings with 300-token chunks and 94.6% accuracy.
high positive Health System Scale Semantic Search Across Unstructured Clin... accuracy_on_clinical_question_answering_benchmark
The system delivers sub-second query latency: median 237 ms single-user, 451 ms at 20-user concurrency.
Full-scale performance characterization reported exact median latencies for single-user and 20-user concurrency.
Multi-agent reinforcement learning has emerged as a promising approach for the combined scheduling of production and transportation tasks in decentralized factories.
Literature-context claim in the paper's introduction summarizing prior research trends and motivations for applying multi-agent RL to integrated production-transport scheduling.
high positive An Analysis of the Coordination Gap between Joint and Modula... potential improvement in scheduling/operational efficiency
Modular training represents a viable alternative in environments where a single scheduling task dominates.
Empirical findings from the paper's experiments/sensitivity analysis indicating modularly trained agents perform comparably to joint training when one scheduling task (either production or transportation) is temporally dominant.
high positive An Analysis of the Coordination Gap between Joint and Modula... relative scheduling performance (modular vs joint training)
Joint training can produce superior performance compared to the best-performing combinations of dispatching rules and modular training.
Empirical evaluation reported in the study comparing joint-training multi-agent RL against modular training and dispatching-rule baselines across simulated job-shop scheduling environments with transportation resources (sensitivity analysis over resource scarcity and temporal dominance). Specific sample size / number of scenarios not stated in the abstract.
high positive An Analysis of the Coordination Gap between Joint and Modula... scheduling performance (e.g., makespan / throughput / overall schedule quality)
Taken together, these insights provide theoretical clarity and practical guidance for responsible GenAI integration into creative work.
Authors' stated contribution and practical recommendations derived from the conceptual framework; no empirical evaluation of guidance effectiveness provided.
high positive Beyond the Creativity Paradox: A Theory-informed Framework f... theoretical clarity and practical guidance for responsible GenAI integration
The study reinterprets process-oriented creativity theories through structural parallels with GenAI.
Conceptual reanalysis and theoretical reinterpretation based on literature synthesis (paper's theoretical contribution).
high positive Beyond the Creativity Paradox: A Theory-informed Framework f... process-oriented creativity theory reinterpretation
The authors propose a role-based integration model that aligns GenAI capabilities with key creative functions: idea generation, synthesis, strategic framing, and facilitation.
Presentation of a novel conceptual model / framework in the paper (theoretical design); no empirical validation or measured outcomes reported.
high positive Beyond the Creativity Paradox: A Theory-informed Framework f... alignment of GenAI capabilities with creative functions (idea generation, synthe...
The paper repositions GenAI as a cognitive collaborator rather than merely a productivity tool.
Argumentative / conceptual claim supported by the proposed theoretical reframing and role-based model in the paper; no empirical testing reported.
high positive Beyond the Creativity Paradox: A Theory-informed Framework f... role of GenAI in organizational workflows (cognitive collaborator vs productivit...
There are structural parallels between GenAI architectures and human cognition—such as heuristic search, divergent thinking, and iterative refinement.
Conceptual mapping and theoretical comparison between GenAI architecture characteristics and cognitive/creativity constructs presented in the paper (literature synthesis / theoretical argument).
high positive Beyond the Creativity Paradox: A Theory-informed Framework f... structural parallels between GenAI architectures and human cognition (heuristic ...
The study revisits foundational creativity theories to develop a framework for integrating GenAI into creative workflows.
Paper describes a conceptual review and theoretical synthesis of foundational creativity theories leading to a proposed integration framework; methodological (theoretical / conceptual) contribution rather than empirical validation.
high positive Beyond the Creativity Paradox: A Theory-informed Framework f... framework for integrating GenAI into creative workflows
Generative Artificial Intelligence (GenAI) is reshaping organisational creativity by emulating cognitive processes traditionally associated with human innovation.
Paper's theoretical argument and literature-grounded conceptual claims (conceptual analysis / literature review); no empirical sample or quantitative data reported.
high positive Beyond the Creativity Paradox: A Theory-informed Framework f... organisational creativity
There is an open opportunity to support collaborative construction where users and AI jointly develop an evolving knowledge representation.
Paper's stated research opportunity and motivation based on gaps identified in prior tools and systems (conceptual argument).
high positive MindTrellis: Co-Creating Knowledge Structures with AI throug... potential benefits of joint user-AI collaborative knowledge representation (prop...
In a user study where 12 participants created slide decks, MindTrellis outperformed retrieval-only baselines in knowledge organization and cognitive load, as measured by expert ratings of content coverage and structural quality.
Controlled user study reported in the paper: N = 12 participants performing slide-deck creation tasks; outcomes assessed via expert ratings of content coverage and structural quality (comparison to retrieval-only baseline).
high positive MindTrellis: Co-Creating Knowledge Structures with AI throug... knowledge organization and cognitive load (operationalized via expert ratings of...
MindTrellis is an interactive visual system where users and AI collaboratively build a dynamic knowledge graph; users can query the graph for document-grounded information and contribute by introducing new concepts, modifying relationships, and reorganizing the hierarchy.
System design and implementation described in the paper (feature description and demonstration).
high positive MindTrellis: Co-Creating Knowledge Structures with AI throug... system capability to support collaborative construction and manipulation of a dy...
The framework, all collected signals, scoring outputs, and evaluation harness are released under CC BY 4.0.
Statement of data and code release policy in the paper.
high positive AgentPulse: A Continuous Multi-Signal Framework for Evaluati... availability/license of framework and data
AgentPulse surfaces deployment signal absent from benchmarks; it is a methodology, not a ground-truth ranking.
Conceptual and empirical argument in the paper supported by the analyses described (correlations with external adoption proxies and divergence from benchmark-only rankings).
high positive AgentPulse: A Continuous Multi-Signal Framework for Evaluati... presence of deployment/adoption signals not captured by standard benchmarks
The Benchmark+Sentiment sub-composite correlates with VS Code installs (ρ_s=0.44, p<0.05), reported as illustrative given that only 11 of 35 agents have non-zero installs.
Spearman correlation between Benchmark+Sentiment sub-composite and VS Code installs on the 35-agent sample, with a caveat that installs are non-zero for only 11 agents; reported correlation and p-value.
high positive AgentPulse: A Continuous Multi-Signal Framework for Evaluati... VS Code installs (IDE install counts as adoption proxy)
The Benchmark+Sentiment sub-composite predicts Stack Overflow question volume (ρ_s=0.49, p<0.01) in the circularity-controlled test (n=35).
Circularity-controlled Spearman correlation between Benchmark+Sentiment sub-composite and Stack Overflow question volume on 35 agents; reported correlation and p-value.
high positive AgentPulse: A Continuous Multi-Signal Framework for Evaluati... Stack Overflow question volume (external adoption/engagement proxy)
A circularity-controlled test (n=35) shows the Benchmark+Sentiment sub-composite, which contains no GitHub-derived signals, predicts external adoption proxies it does not aggregate: GitHub stars (ρ_s=0.52, p<0.01).
Circularity-controlled correlation test (Spearman) between Benchmark+Sentiment sub-composite and GitHub stars on a 35-agent sample; reported Spearman correlation and p-value.
high positive AgentPulse: A Continuous Multi-Signal Framework for Evaluati... GitHub stars (external adoption proxy)
We introduce AgentPulse, a continuous evaluation framework scoring 50 agents across 10 workload categories along four factors (Benchmark Performance, Adoption Signals, Community Sentiment, and Ecosystem Health) aggregated from 18 real-time signals across GitHub, package registries, IDE marketplaces, social platforms, and benchmark leaderboards.
Methodological description in the paper; reported sample of 50 agents and use of 18 signals from enumerated sources.
high positive AgentPulse: A Continuous Multi-Signal Framework for Evaluati... AgentPulse composite and factor scores (Benchmark Performance, Adoption Signals,...
These results provide concrete tier-selection guidance across deployment scales from a single seminar to a university-wide rollout.
Concluding claim in paper based on empirical latency and cost comparisons across throughput tiers and concurrency up to 50 users from a live deployment.
high positive Latency and Cost of Multi-Agent Intelligent Tutoring at Scal... tier-selection guidance for deployment scale decision-making
Provisioned Throughput, expensive under continuous provisioning, becomes cost-competitive for institutions that can predict and concentrate their traffic toward high utilization.
Cost modeling and comparison across provisioning modes in the paper showing trade-offs conditional on utilization; based on the instrumented tiers and assumed usage patterns.
high positive Latency and Cost of Multi-Agent Intelligent Tutoring at Scal... cost competitiveness (cost per unit of usage vs utilization)
Cost analysis places both pay-per-token tiers well below the price of a STEM textbook per student per semester under a worst-case usage ceiling.
Cost analysis presented in the paper comparing per-student per-semester costs under a stated worst-case usage ceiling to the price of a STEM textbook; exact numeric assumptions not provided in the excerpt.
high positive Latency and Cost of Multi-Agent Intelligent Tutoring at Scal... cost per student per semester
Priority PayGo maintains flat sub-4-second response times across the full load range.
Empirical latency measurements from the ITAS instrumentation described in the paper (over 3,000 requests across concurrency levels up to 50 and three throughput tiers).
Multi-agent LLM tutoring systems improve response quality through agent specialization.
Statement in paper describing design rationale; no quantitative quality comparison or metrics provided in the excerpt.
Generative artificial intelligence (genAI) is rapidly reshaping how knowledge and culture are produced and consumed.
Author's descriptive statement based on observed changes in production/consumption patterns (no empirical sample reported in paper abstract).
high positive Generative artificial intelligence reduces social welfare th... production and consumption of knowledge and culture
The paper provides a structured overview of energy forecasting use cases along three main dimensions: stakeholders, attributes, and data categories.
Methodological contribution described in the paper (framework/overview component).
high positive FETS Benchmark: Foundation Models Outperform Dataset-specifi... existence of a structured overview framework
The FETS benchmark collects and analyzes 54 datasets across 9 data categories guided by typical stakeholder interests.
Descriptive claim about the benchmark dataset compilation reported in the paper (dataset count and category count).
high positive FETS Benchmark: Foundation Models Outperform Dataset-specifi... breadth of dataset coverage (count and categories)
Foundation models show improved performance at higher aggregation levels such as national load, district heating, and power grid data.
Subset analyses within the benchmark comparing performance across aggregation levels (e.g., national load, district heating, power grid) showing better results for aggregated data.
high positive FETS Benchmark: Foundation Models Outperform Dataset-specifi... forecast accuracy stratified by aggregation level
Foundation models outperform classical machine learning approaches despite the latter having seen the full historic target data during training.
Benchmark setup where classical ML baselines had access to full historic target data during training while foundation models were pretrained/generalized; empirical comparison across the dataset collection.
high positive FETS Benchmark: Foundation Models Outperform Dataset-specifi... forecasting accuracy when classical ML had access to full historic targets
Covariate-informed foundation models achieve the strongest performance.
Benchmark experiments that compare foundation model variants, including those that incorporate covariates, across the collected datasets; reported as a comparative finding in the benchmark results.
high positive FETS Benchmark: Foundation Models Outperform Dataset-specifi... predictive performance of covariate-informed vs non-covariate models
Foundation models consistently outperform dataset-specific optimized machine learning approaches across all settings and data categories.
Empirical benchmark comparing foundation models vs classical dataset-specific ML approaches across multiple forecasting settings and data categories; reported analysis uses the collection of datasets assembled for the FETS benchmark.
high positive FETS Benchmark: Foundation Models Outperform Dataset-specifi... predictive performance of time series forecasts (forecast accuracy/output qualit...
This paper presents the first systematic study of token consumption patterns in agentic coding tasks, analyzing trajectories from eight frontier LLMs on SWE-bench Verified and evaluating models' ability to predict their own token costs before task execution.
Stated scope and methodology in the paper: dataset is SWE-bench Verified, eight frontier LLMs were analyzed, and experiments included model self-prediction evaluation.
high positive How Do AI Agents Spend Your Money? Analyzing and Predicting ... scope of study (presence of systematic analysis and self-prediction evaluation)
Reducing variability in solder-joint quality and cycle time.
Abstract statement that variability in solder-joint quality and cycle time was reduced during the deployment (no quantitative variability metrics provided in the abstract).
high positive Learning-augmented robotic automation for real-world manufac... variability of solder-joint quality; variability of cycle time
It maintained near-human takt time.
Abstract claim comparing the system's cycle/takt time to human performance during the deployment (no numeric takt-time comparison provided in the abstract).
high positive Learning-augmented robotic automation for real-world manufac... takt time (cycle time) relative to human workers
Achieving a 99.4% pass rate on product-level quality-control tests.
Reported QC pass rate from the production run in the abstract (presumably based on the produced motors).
high positive Learning-augmented robotic automation for real-world manufac... product-level quality-control pass rate
Operating without physical fencing.
Abstract statement that the run occurred "without physical fencing" (implying operation around people without traditional fences).
high positive Learning-augmented robotic automation for real-world manufac... use of physical fences for safety (absent)
Produced 108 motors.
Count of products produced during the continuous run reported in the abstract.
high positive Learning-augmented robotic automation for real-world manufac... number of motors produced during the run
The system operated continuously for 5 h 10 min.
Reported continuous operation duration from the production run described in the abstract.
high positive Learning-augmented robotic automation for real-world manufac... continuous operational time without interruption
Less than 20 min of real-world data per task.
Reported training data requirement for the deployed tasks in the authors' field experiment (abstract statement).
high positive Learning-augmented robotic automation for real-world manufac... amount of real-world training data per task
With less than 20 min of real-world data per task, the system operated continuously for 5 h 10 min, producing 108 motors without physical fencing and achieving a 99.4% pass rate on product-level quality-control tests.
Single field deployment / production run reported in the paper; numbers reported in the abstract (training data time, continuous operation duration, number of motors produced, fencing status, QC pass rate).
high positive Learning-augmented robotic automation for real-world manufac... training data required; continuous operational duration; production quantity; pr...
We deployed the system on an electric-motor production line to automate deformable cable insertion and soldering under real manufacturing constraints, a step previously performed manually by human workers.
Field deployment on an actual electric-motor production line described by the authors (deployment + task specification).
high positive Learning-augmented robotic automation for real-world manufac... automation of previously manual deformable cable insertion and soldering tasks
We present Learning-Augmented Robotic Automation, a hybrid system that integrates learned task controllers and a neural 3D safety monitor into conventional industrial workflows.
Description of the system developed by the authors (system design/development reported in the paper).
high positive Learning-augmented robotic automation for real-world manufac... integration of learned controllers and 3D safety monitoring
Self-correction should be treated not as a default behavior, but as a control decision governed by measurable error dynamics.
Synthesis of theoretical framing (Markov model and diagnostic inequality) and empirical results across multiple models/datasets showing thresholds and promptability of EIR.
high positive When Does LLM Self-Correction Help? A Control-Theoretic Mark... policy/recommendation about when to enable iterative self-correction to improve ...
A 'verify-first' prompt ablation on GPT-4o-mini reduces EIR from 2% to 0% and turns -6.2 pp degradation into +0.2 pp (paired McNemar p < 10^-4).
A prompt-ablation experiment reported for GPT-4o-mini showing EIR dropping from 2% to 0% and the observed accuracy change flipping from -6.2 percentage points to +0.2 percentage points; statistical significance assessed with a paired McNemar test (p < 10^-4).
high positive When Does LLM Self-Correction Help? A Control-Theoretic Mark... EIR and accuracy change from self-correction after prompt modification
In this framework, EIR functions as a stability margin and prompting functions as lightweight controller design.
Conceptual framing in the paper (cybernetic feedback loop where the same language model is controller and plant), supported by associated experiments showing prompt changes affect EIR and outcomes.
high positive When Does LLM Self-Correction Help? A Control-Theoretic Mark... stability of iterative refinement (EIR) and resulting accuracy
Iterate only when ECR/EIR > Acc/(1 - Acc).
The paper frames self-correction as a two-state Markov model over {Correct, Incorrect} and derives this deployment diagnostic analytically from that model.
high positive When Does LLM Self-Correction Help? A Control-Theoretic Mark... whether iterative self-correction is expected to improve accuracy
Im Forschungskontext sind kontextbezogene Schulungs- und Begleitmaßnahmen entscheidend für den Erfolg der Copilot-Einführung.
Schlussfolgerung der Autoren aus den Befunden zur zeitlichen Entwicklung der Bewertungen wissenschaftlicher Mitarbeitender und zu unterschiedlichen Nutzenwahrnehmungen (im Abstract genannt).
high positive Generative KI in der Wissensarbeit: Wahrnehmung, Nutzen und ... Bedeutung von Schulungs- und Begleitmaßnahmen für Erfolg/Adoption
Die Untersuchung zeigt, dass Microsoft 365 Copilot insbesondere im administrativen Bereich Effizienzgewinne ermöglicht.
Selbstberichtete Einschätzungen der Beschäftigten (speziell Verwaltungsmitarbeitende) in der wiederholten Querschnittsbefragung; Autoren ziehen daraus praktische Relevanz im administrativen Bereich (Abstract).
high positive Generative KI in der Wissensarbeit: Wahrnehmung, Nutzen und ... Wahrgenommene Effizienzgewinne im administrativen Bereich