The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests About 🎲 Workforce Futures
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Evidence (8974 claims)

Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.

The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).

Browse by theme

Nine broad, paper-level topics. Click one to filter the claims below.

Adoption
10085 claims
Filter claims →
Productivity
8974 claims
Filtered →
Governance
8062 claims
Filter claims →
Human-AI Collaboration
7749 claims
Filter claims →
Org Design
5057 claims
Filter claims →
Innovation
4896 claims
Filter claims →
Labor Markets
4088 claims
Filter claims →
Skills & Training
3372 claims
Filter claims →
Inequality
2377 claims
Filter claims →

Claims by outcome category

Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.

Outcome Positive Negative Mixed Null Total
Other 882 244 117 1097 2424
Governance & Regulation 1010 469 229 135 1875
Organizational Efficiency 977 235 149 90 1462
Technology Adoption Rate 781 299 143 128 1362
Research Productivity 506 155 74 363 1110
Output Quality 555 219 71 70 915
Decision Quality 395 200 95 54 751
Firm Productivity 523 67 101 27 724
AI Safety & Ethics 262 309 75 36 688
Market Structure 195 201 135 30 566
Task Allocation 248 77 96 38 464
Innovation Output 300 34 55 20 411
Skill Acquisition 207 75 65 21 368
Employment Level 138 67 119 24 350
Fiscal & Macroeconomic 156 80 53 33 329
Task Completion Time 211 38 13 16 280
Firm Revenue 183 52 29 5 270
Consumer Welfare 131 77 48 13 269
Inequality Measures 50 141 54 9 254
Worker Satisfaction 104 85 25 13 227
Error Rate 87 112 11 5 215
Automation Exposure 69 69 37 20 198
Wages & Compensation 102 49 31 11 193
Team Performance 115 30 30 11 187
Regulatory Compliance 88 74 17 7 186
Training Effectiveness 109 22 14 21 168
Developer Productivity 116 21 15 8 161
Job Displacement 12 92 26 1 131
Hiring & Recruitment 57 12 9 5 83
Skill Obsolescence 6 59 10 2 77
Social Protection 43 17 8 2 70
Creative Output 35 21 9 4 70
Labor Share of Income 18 23 17 1 59
Worker Turnover 15 16 4 35
Industry 1 1
Clear
Productivity Remove filter
DOMUS is replicable digital public infrastructure: a modular, cloud-native Software-as-a-Service architecture that can be deployed across other UK boroughs and adapted to other public administration tasks characterised by scarcity, rule-bound eligibility, and high stakes.
Design and architectural claim in the paper about modular cloud-native SaaS architecture and intended replicability; this is an assertion about generalisability rather than evidence from multi-site deployments.
high positive Optimising Temporary Accommodation Placement Across London w... replicability and adaptability of system
The deployment maintained statutory compliance and role-based accountability.
Paper asserts that DOMUS operations preserved statutory compliance and role-based accountability during the pilot (claimed based on system design and pilot evaluation; no specific compliance audits or metrics provided in the provided text).
high positive Optimising Temporary Accommodation Placement Across London w... statutory compliance and role-based accountability
Results indicate high staff satisfaction.
User feedback / staff satisfaction reported from the pilot deployment (paper states high satisfaction but does not report sample size or survey statistics in the provided text).
high positive Optimising Temporary Accommodation Placement Across London w... staff satisfaction with system
Results indicate improved adherence to key placement constraints.
Pilot evaluation results reporting better adherence to placement constraints (e.g., bedroom need, affordability, accessibility) under DOMUS compared to manual workflows (no quantitative metrics provided in the provided text).
high positive Optimising Temporary Accommodation Placement Across London w... adherence to placement constraints
Results indicate substantial reductions in search time.
Findings from the pilot deployment comparing DOMUS-assisted search time to manual search workflows (paper reports reductions but does not provide numeric effect size in the provided text).
Household and property attributes are encoded into policy-consistent representations prior to AI-assisted ranking and explanation.
Technical design detail in the paper describing preprocessing/encoding pipeline used by DOMUS before AI ranking and explanation.
high positive Optimising Temporary Accommodation Placement Across London w... encoding of attributes into policy-consistent representations
The system combines transparent, rule-based filtering with large language model-assisted search to standardise the application of bedroom need, affordability thresholds, geographic preferences, and accessibility requirements, while preserving officer discretion and audibility.
Design and functionality description; statement that DOMUS uses rule-based filtering plus LLM-assisted search to standardise applications of policy rules and preserve discretion/audibility.
high positive Optimising Temporary Accommodation Placement Across London w... standardisation of policy application and preservation of discretion/audibility
DOMUS integrates household case records, policy-constrained affordability and suitability rules, and live private-rental listings within a single governance-aligned workflow.
System architecture and functionality described in the paper; design claim about integrated data sources and workflow alignment.
high positive Optimising Temporary Accommodation Placement Across London w... integration of data sources into workflow
The paper documents the creation and use of DOMUS, a cloud-based, AI-enabled decision-support system built from scratch at the University of East London and customised for the needs of London Borough of Newham to support statutory Temporary accommodation placement.
Implementation and deployment description in the paper (system development and customised pilot deployment for Newham).
high positive Optimising Temporary Accommodation Placement Across London w... existence and deployment of DOMUS
LabOSBench provides a reproducible, low-cost testbed for advancing computer-using agents toward scientific-instrument control.
Authors' claim about the benchmark's intended properties (reproducibility, low cost) based on its web-based simulated design and browser operation; no cost comparison or reproducibility study reported in the excerpt.
high positive LabOSBench: Benchmarking Computer Use Agents for Scientific ... testbed_reproducibility_and_cost
LabOSBench is a challenging benchmark for multimodal GUI agents built on a suite of web-based scientific-instrument simulators that operates directly via a browser and thus avoids resource-heavy OS virtualization while supporting flexible task configuration and execution-based evaluation.
System and benchmark description in paper; architectural/design claim (web-based simulators, browser operation, avoidance of OS virtualization, supports flexible configuration and execution-based evaluation).
The authors release their code and agent trajectories to support future research.
Statement in the paper indicating that code and recorded agent trajectories are publicly released.
high positive CoffeeBench: Benchmarking Long-Horizon LLM Agents in Heterog... availability of code and trajectories
Most evaluated models achieve positive net income in the CoffeeBench simulation.
Reported experimental results stating that the majority of evaluated LLMs produced positive cumulative net income.
Across several recent open-weight and proprietary LLMs, all models outperform a passive baseline that takes no actions.
Empirical evaluation comparing multiple recent open-weight and proprietary LLMs to a passive baseline (passive baseline defined as taking no actions); paper reports that every tested model outperformed that baseline.
high positive CoffeeBench: Benchmarking Long-Horizon LLM Agents in Heterog... performance relative to passive baseline (net income / cumulative reward)
We introduce CoffeeBench, a benchmark for evaluating LLM agents in a long-horizon multi-agent economy composed of heterogeneous firms.
Design and release described in the paper: authors state they developed CoffeeBench as a benchmark for long-horizon multi-agent economic evaluation.
high positive CoffeeBench: Benchmarking Long-Horizon LLM Agents in Heterog... benchmark availability / capability to evaluate agents
An agnostic model is formed (combining theoretical constructs, empirical evidence, and practical applications) enabling organizations to have the accountability, operational and human oversight needed to embrace responsible AI-enabled automation of enterprise systems and processes.
Paper proposes an agnostic model based on synthesis of theory, empirical evidence, and practice; the excerpt describes the model conceptually without presenting evaluation metrics or sample-based validation.
high positive Prompt-Driven Integration Workflow Generation: A Technical A... availability of accountability, operational oversight, and human oversight via p...
Artificial intelligence integration requires established governance frameworks, human-in-the-loop verification, and explainable artificial intelligence to ensure compliance with the organization's values and legislation.
Prescriptive recommendation in the paper combining theoretical constructs and practical application guidance; no empirical validation provided in the excerpt.
high positive Prompt-Driven Integration Workflow Generation: A Technical A... compliance with organizational values and legislation via governance mechanisms
Systematic AI integration can produce meaningful productivity gains across engineering design, content generation, multimedia creation, and scientific experimentation.
Paper combines theoretical, empirical, and practical examples to claim productivity gains; excerpt does not provide study methodology or sample sizes.
Systematic AI integration can produce meaningful improvements in configuration accuracy.
Asserted by paper based on examples from engineering, content, multimedia, and scientific workflows; excerpt contains no measurement details or sample sizes.
Modern engineering design, content generation, multimedia creation, and scientific experimentation conducted by organizations show that meaningful savings in development time can be realized by the systematic integration of artificial intelligence technologies into existing quality management systems.
Paper synthesizes theoretical constructs, empirical evidence, and practical applications to assert time-savings across multiple domains; no specific study design, sample size, or quantified effect provided in the excerpt.
Measurable system reliability improvements have been achieved using large language model capabilities with structured enterprise integration platforms.
Paper reports improvements in system reliability associated with LLM integration; excerpt lacks details on measurement approach, sample, or magnitude.
Measurable defect improvements (reduction in defects) have been achieved using large language model capabilities with structured enterprise integration platforms.
Paper claims empirical reductions in defects linked to LLM integration; no methodological details or sample size provided in the excerpt.
high positive Prompt-Driven Integration Workflow Generation: A Technical A... defect rates / defect reduction
Measurable productivity improvements have been achieved using large language model capabilities with structured enterprise integration platforms.
Paper asserts empirical/measurable productivity improvements attributable to LLMs integrated with enterprise platforms; the excerpt provides no details on study design, measurement method, or sample size.
Organizations should consider governance and quality management when introducing generative AI.
Normative recommendation by the paper, presented as best-practice guidance; based on theoretical constructs and practical applications described, not an empirical test in the excerpt.
high positive Prompt-Driven Integration Workflow Generation: A Technical A... adoption of governance and quality management practices
Automated orchestration, code generation, and generative creativity have been introduced as well.
Descriptive claim in paper indicating the introduction/adoption of specific generative-AI capabilities; no empirical details or sample size given in the excerpt.
high positive Prompt-Driven Integration Workflow Generation: A Technical A... introduction/adoption of specific AI capabilities (automated orchestration, code...
Generative AI systems have been incorporated into innovation and process optimization in organizations.
Stated in paper as an observed trend / descriptive claim; no specific study design, sample size, or empirical method reported in the excerpt.
high positive Prompt-Driven Integration Workflow Generation: A Technical A... incorporation of generative AI into organizational innovation and process optimi...
The authors release an updated version of the WorkBench benchmark with data and code quality improvements, new model scores, and analysis of agent progress on WorkBench since 2024.
Statement of release in the paper (announcement of updated benchmark contents); verifiable by checking the authors' repository or supplementary materials referenced in the paper.
high positive WorkBench Revisited: Workplace Agents Two Years On availability of updated benchmark (data, code, scores, analysis)
The rise of open-weight models has drastically lowered costs for a performance level that was previously only accessible to proprietary models.
Comparative cost analysis reported by authors linking open-weight model availability to lower costs for a given performance level; the excerpt does not include specific cost figures, sample sizes, or methodology.
high positive WorkBench Revisited: Workplace Agents Two Years On costs required to attain a given model performance level
Several classes of error have been totally eliminated from frontier agents on WorkBench.
Authors' qualitative analysis of error types observed in WorkBench evaluations between 2024 and 2026; the excerpt does not list which error classes, how elimination was measured, or sample sizes.
high positive WorkBench Revisited: Workplace Agents Two Years On elimination of specific error classes
On WorkBench, capability and safety go together rather than trade off: models that finish the most tasks also do the least unintended damage.
Comparative analysis of agent performance on the WorkBench benchmark showing association between higher task completion rates and lower rates of unintended harmful actions; specific statistical measures and sample size not provided in the excerpt.
high positive WorkBench Revisited: Workplace Agents Two Years On association between task completion rate and rate of unintended harmful actions
In June 2026 the best agent to date, Claude Opus 4.8, completes 89% of tasks on WorkBench.
Reported evaluation result on the WorkBench benchmark (June 2026) comparing agent task completion rates; exact sample size not stated in the excerpt.
high positive WorkBench Revisited: Workplace Agents Two Years On task completion rate (percentage of tasks completed)
In March 2024 the best agent on WorkBench, GPT-4, completed 43% of tasks.
Reported evaluation result on the WorkBench benchmark (March 2024) comparing agent task completion rates; exact sample size not stated in the excerpt.
high positive WorkBench Revisited: Workplace Agents Two Years On task completion rate (percentage of tasks completed)
Moving beyond experimental phases requires high-impact use cases and decentralized governance.
Paper emphasizes (argues) that scaling AI past experiments depends on choosing high-impact use cases and adopting decentralized governance; presented as recommendations rather than empirically validated findings in the summary.
high positive Zombie Ai Investments: From Technical Success To Business Fa... successful scaling / transition from experiments to production and business impa...
The research offers guidance for bridging the gap between technical success and business impact through operational mitigation strategies.
Paper provides proposed operational strategies and guidance (prescriptive content); no evidence of empirical testing given in the summary.
high positive Zombie Ai Investments: From Technical Success To Business Fa... organizational ability to realize business value from AI / operational effective...
The verifier does buy sound coverage: covering all unseen valuable statements while asserting only valid ones is possible with it, impossible without it; it relocates unavoidable errors from false to trivial.
Formal proofs within the paper's model showing existence (with verifier) and impossibility (without verifier) results; analysis of error types shifting from false statements to trivial (verifiable-but-not-valuable) statements.
high positive Flood and Harvest: The Provable Necessity of Trivia for Gene... coverage of unseen valuable statements and type/location of inevitable errors
A data-driven culture supports both AI deployment depth and breadth.
Empirical association estimated in staged OLS models using archival microdata from 770 large Spanish firms (data-driven culture measured and shown to be associated with both depth and breadth).
high positive Beyond AI Adoption: An Empirical Study on the Antecedents an... AI deployment depth and breadth
Digital infrastructure drives AI deployment breadth.
Empirical association estimated in staged OLS models using archival microdata from 770 large Spanish firms (digital infrastructure included as predictor of the breadth dimension).
AI-skilled human capital drives AI deployment depth.
Empirical association estimated in staged OLS models using archival microdata from 770 large Spanish firms (human capital measures included as predictors of the depth dimension).
Firm performance is positively associated with the interaction between AI deployment depth and AI deployment breadth (consistent with a complementarity logic).
Empirical analysis using archival microdata from 770 large Spanish firms; staged OLS regression models testing the interaction between measured AI deployment depth and breadth on firm performance.
The contribution is a falsifiable program centered on minimum functional description length and verified-change cost.
Authors' stated research contribution and framing in the paper.
high positive No Accidental Software Agent First Canonical Code for Human ... existence of a falsifiable research program around minimum functional descriptio...
Preliminary QLoRA experiments on Qwen2.5-Coder-14B show that 64,088 canonical trajectories are learnable and suppress tested forbidden-language markers.
Reported preliminary experiment using QLoRA on Qwen2.5-Coder-14B with 64,088 canonical trajectories; empirical claim within the paper.
high positive No Accidental Software Agent First Canonical Code for Human ... number of learnable canonical trajectories and suppression of tested forbidden-l...
For supported routine-product distributions, this approach gives a defensible planning target near 100-fold all-in cost reduction (amortized cost per verified correct change), though this is a hypothesis and not a guarantee for all software.
Stated projected limit/planning target in the paper; authors explicitly note these reduction bands are hypotheses and not measured frontier results.
high positive No Accidental Software Agent First Canonical Code for Human ... amortized all-in cost per verified correct change
Quotienting software by behavior equivalence under a declared oracle can collapse equivalent encodings into governed representatives with explicit evidence and proof obligations (core hypothesis).
Central theoretical hypothesis / conceptual claim in the paper (no empirical test reported).
high positive No Accidental Software Agent First Canonical Code for Human ... degree of encoding collapse / reduction of accidental representations
We propose an agent-first canonical code, a proof-carrying substrate that rewrites routine product software into canonical behavior profiles, typed change algebra, proof lanes, constrained edit grammars, semantic patch cells, runtime negative memory, and proof-carrying change objects.
Design/proposal presented by the authors (architectural proposal, not an evaluated system).
high positive No Accidental Software Agent First Canonical Code for Human ... feasibility of a proof-carrying canonical-code substrate
Human code repositories contain valuable signals such as tests, incidents, migrations, edge cases, product judgment, and operational history.
Asserted observation in the paper describing repository contents and their value; no quantified empirical study provided.
high positive No Accidental Software Agent First Canonical Code for Human ... value of repository signals for software understanding / maintenance
Emotional AI systems may improve organisational productivity when implemented within ethically grounded and transparent frameworks.
Authors' synthesis from the systematic review of the literature (claimed association based on reviewed studies); no quantitative pooled effect size reported in the abstract.
high positive Emotional AI in the Workplace: Systematic Review of Effects ... organisational productivity
Emotional AI systems may enhance employee engagement when implemented within ethically grounded and transparent frameworks.
Synthesis claim from the systematic review (authors conclude this based on aggregated findings across the surveyed literature); specific studies and counts not reported in the abstract.
high positive Emotional AI in the Workplace: Systematic Review of Effects ... employee engagement (attitudes such as job satisfaction, motivation, adaptabilit...
Data exhibits spillovers such that data generated by one task can augment the productivity of another task.
Model assumption and formalization of cross-task data spillovers in the analytical framework (theoretical derivation and model structure).
high positive Data-Driven Automation cross-task productivity augmentation via data spillovers
Brick outperforms all tested routers on the evaluated benchmark.
Empirical comparison on the 5,504-query benchmark where Brick's max-quality accuracy (76.98%) is reported as higher than other routers.
high positive Brick: Spatial Capability Routing for the Mixture-of-Models ... accuracy compared to alternative routing approaches
Median latency drops from 51.2s to 22.8s when using Brick.
Empirical measurement on the benchmark reporting median latency before and after applying Brick.