The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests About 🎲 Workforce Futures
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Evidence (8974 claims)

Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.

The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).

Browse by theme

Nine broad, paper-level topics. Click one to filter the claims below.

Adoption
10085 claims
Filter claims →
Productivity
8974 claims
Filtered →
Governance
8062 claims
Filter claims →
Human-AI Collaboration
7749 claims
Filter claims →
Org Design
5057 claims
Filter claims →
Innovation
4896 claims
Filter claims →
Labor Markets
4088 claims
Filter claims →
Skills & Training
3372 claims
Filter claims →
Inequality
2377 claims
Filter claims →

Claims by outcome category

Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.

Outcome Positive Negative Mixed Null Total
Other 882 244 117 1097 2424
Governance & Regulation 1010 469 229 135 1875
Organizational Efficiency 977 235 149 90 1462
Technology Adoption Rate 781 299 143 128 1362
Research Productivity 506 155 74 363 1110
Output Quality 555 219 71 70 915
Decision Quality 395 200 95 54 751
Firm Productivity 523 67 101 27 724
AI Safety & Ethics 262 309 75 36 688
Market Structure 195 201 135 30 566
Task Allocation 248 77 96 38 464
Innovation Output 300 34 55 20 411
Skill Acquisition 207 75 65 21 368
Employment Level 138 67 119 24 350
Fiscal & Macroeconomic 156 80 53 33 329
Task Completion Time 211 38 13 16 280
Firm Revenue 183 52 29 5 270
Consumer Welfare 131 77 48 13 269
Inequality Measures 50 141 54 9 254
Worker Satisfaction 104 85 25 13 227
Error Rate 87 112 11 5 215
Automation Exposure 69 69 37 20 198
Wages & Compensation 102 49 31 11 193
Team Performance 115 30 30 11 187
Regulatory Compliance 88 74 17 7 186
Training Effectiveness 109 22 14 21 168
Developer Productivity 116 21 15 8 161
Job Displacement 12 92 26 1 131
Hiring & Recruitment 57 12 9 5 83
Skill Obsolescence 6 59 10 2 77
Social Protection 43 17 8 2 70
Creative Output 35 21 9 4 70
Labor Share of Income 18 23 17 1 59
Worker Turnover 15 16 4 35
Industry 1 1
Clear
Productivity Remove filter
On BioML-Bench, spanning biomedical imaging, protein engineering, single-cell omics, and drug discovery, AutoScientists achieves a mean leaderboard percentile of 74.4% across 24 tasks, improving over the strongest AI agent by +8.33%.
Empirical evaluation on the BioML-Bench benchmark (24 tasks); reported mean leaderboard percentile and comparative improvement versus the strongest baseline agent.
high positive AutoScientists: Self-Organizing Agent Teams for Long-Running... leaderboard percentile across benchmark tasks
Under matched experimental budgets, AutoScientists improves over prior AI agents across biomedical machine learning, language-model training optimization, and protein fitness prediction.
Empirical comparisons reported in paper across multiple benchmark suites and tasks (BioML-Bench, GPT training optimization experiments, ProteinGym).
high positive AutoScientists: Self-Organizing Agent Teams for Long-Running... overall performance across multiple benchmarks
AutoScientists is a decentralized team of AI agents that interpret a shared experimental state, self-organize into teams around promising hypotheses, critique proposals before using experimental compute, and share successes and failures to reduce redundant exploration.
System design and implementation described in the paper (architecture and agent protocols); qualitative description of agent behaviors and coordination mechanisms; demonstrated in experiments.
high positive AutoScientists: Self-Organizing Agent Teams for Long-Running... agent coordination and information sharing (qualitative description)
We describe the benchmark design, evaluation protocol, and quality-control pipeline, and position OR-Space as a benchmark for studying the reliability, failure modes, and practical readiness of LLM agents in industrial OR workflows.
Statement of the paper's contributions and contents (methodological description of what the paper includes).
high positive OR-Space: A Full-Lifecycle Workspace Benchmark for Industria... capability to study reliability, failure modes, and readiness of LLM agents
By combining persistent workspaces with lifecycle-oriented tasks, OR-Space evaluates whether agents can perform reliable optimization work beyond end-to-end text generation.
Stated objective/claim in the paper about the benchmark's purpose and what it measures (conceptual/goal-oriented statement).
high positive OR-Space: A Full-Lifecycle Workspace Benchmark for Industria... reliability of LLM agents in performing optimization work (beyond text generatio...
OR-Space defines an Explain task mode, where agents answer grounded questions about solutions, constraints, and business implications using evidence spread across workspace artifacts.
Definition of the Explain task mode provided in the paper (design/specification).
high positive OR-Space: A Full-Lifecycle Workspace Benchmark for Industria... ability to generate grounded explanations using workspace evidence
OR-Space defines a Revise task mode, where agents modify existing models under changing requirements or solver feedback while preserving valid prior logic.
Definition of the Revise task mode in the benchmark design (descriptive claim in the paper).
high positive OR-Space: A Full-Lifecycle Workspace Benchmark for Industria... ability to revise models while preserving prior logic
OR-Space defines three task modes: Build, where agents construct solver-ready optimization models from heterogeneous artifacts.
Definition of one of the benchmark's task modes as described in the paper (method/design description).
high positive OR-Space: A Full-Lifecycle Workspace Benchmark for Industria... ability to construct solver-ready models
Each instance is an executable workspace containing business documents, structured data, optional code artifacts, solver outputs, and task-specific evaluators distributed across interdependent files.
Design specification of OR-Space provided in the paper (descriptive claim about benchmark instance structure).
high positive OR-Space: A Full-Lifecycle Workspace Benchmark for Industria... complexity and composition of benchmark instances
We introduce OR-Space, a full-lifecycle workspace benchmark for evaluating industrial optimization agents across model construction, model revision, and grounded explanation.
Paper presents and names a new benchmark (methodological contribution described directly in the text).
high positive OR-Space: A Full-Lifecycle Workspace Benchmark for Industria... capability of benchmarks to evaluate OR agents across lifecycle tasks
Large language model (LLM) agents are increasingly used to assist with operations research (OR) modeling.
Statement in the paper asserting an observed trend; likely based on literature/context motivating the work (no empirical sample or quantitative citation provided in the excerpt).
high positive OR-Space: A Full-Lifecycle Workspace Benchmark for Industria... LLM agent adoption in OR workflows
A recommended organizational design for the AI era is the 'resonance protocol enterprise' in which structures are temporary crystallizations, AI governance protects adaptive openness, and legitimacy derives from sustaining recursive renewal.
Normative/proposal in the paper outlining a new organizational design paradigm; presented as conceptual design without empirical pilot or evaluation.
high positive The Lantern in the Vault: AI, Crisis, and the Ontology of Or... organizational design aimed at sustaining adaptive renewal and legitimacy under ...
Digital transformation initially enhanced adaptability by fluidifying information flows and expanding relational connectivity, thereby improving some organizations' adaptability.
Theoretical claim supported by qualitative interpretation of digital transformation phenomena; no systematic measurement or reported sample.
high positive The Lantern in the Vault: AI, Crisis, and the Ontology of Or... organizational adaptability associated with digital transformation practices
Organizations capable of rapid relational reconfiguration, customer reconnection, and generative experimentation often proved more resilient during the pandemic.
Illustrative/theoretical interpretation of pandemic cases offered in the paper; no quantified sample or formal empirical evidence reported.
high positive The Lantern in the Vault: AI, Crisis, and the Ontology of Or... organizational resilience as a function of relational reconfiguration and experi...
Although AI creates obstacles, it also has the potential to be an important tool for creating innovative opportunities and continued growth if managed with sound practices.
Concluding statement in the paper's abstract presenting a normative/conditional conclusion based on the paper's evaluation and synthesis of evidence (no primary quantified results provided in the supplied text).
high positive Impact of Artificial Intelligence on Employment and Society innovation opportunities and continued economic/organizational growth under soun...
AI leads to the creation of new jobs.
The paper explicitly states it examines the creation of new jobs as a ramification of AI (abstract); claim presented qualitatively without reported sample sizes or quantified effect in the provided text.
high positive Impact of Artificial Intelligence on Employment and Society creation of new jobs / net employment effects
GENESIS is built on three composable primitives (agents, skills, hooks) and a knowledge layer (SYNAPSE) that doubles as the source of ground truth and the recipient of every artifact the framework produces, making capabilities compound across runs.
Architectural description in the paper; claim about knowledge base acting as ground truth and enabling capability compounding (design-level claim). No quantitative evaluation given in the abstract.
high positive GENESIS: Harnessing AI Agents for Autonomous 6G RAN Synthesi... accumulation/compounding of capabilities across runs (longitudinal improvement o...
GENESIS is an agentic AI framework that converts intents (e.g., a specification clause, a telemetry anomaly, or a research hypothesis) into solutions validated with over-the-air experiments, fed back into a persistent knowledge base.
System design / implementation claim presented in the paper (description of proposed framework). The abstract does not report empirical evaluation metrics or sample size.
high positive GENESIS: Harnessing AI Agents for Autonomous 6G RAN Synthesi... ability to produce solutions validated by over-the-air experiments (end-to-end R...
Large Language Models (LLMs) have compressed comparable R&D work in general software engineering from days to minutes.
Paper's stated comparison/claim (likely based on prior reports or authors' experience); no experimental details or sample size provided in the abstract.
high positive GENESIS: Harnessing AI Agents for Autonomous 6G RAN Synthesi... time to complete R&D/software engineering tasks
Agentic Technical Debt and Stochastic Tax are related but distinct: debt can amplify the tax.
Theoretical relationship asserted in the structural model; the note states debt can amplify the recurring Stochastic Tax and provides model expressions and discussion (and illustrative simulation) to substantiate the relationship.
high positive Modeling Agentic Technical Debt and Stochastic Tax: A Standa... impact of accumulated Agentic Technical Debt on the magnitude of Stochastic Tax ...
Combining both levers yields a 502% improvement on single-cell RNA denoising over the initial baseline.
Reported experimental result in the paper comparing SIA to the initial baseline on the single-cell RNA denoising task (denoising metric unspecified in abstract).
high positive SIA: Self Improving AI with Harness & Weight Updates denoising performance for single-cell RNA data
Combining both levers yields a 91.9% runtime reduction on GPU kernels over the initial baseline.
Reported experimental result in the paper comparing SIA to the initial baseline on the low-level GPU kernel optimisation task (runtime measured).
high positive SIA: Self Improving AI with Harness & Weight Updates runtime for GPU kernels
Combining both levers yields a 56.6% gain on LawBench (Chinese legal charge classification) over the initial baseline.
Reported experimental result in the paper comparing SIA to the initial baseline on the LawBench task.
high positive SIA: Self Improving AI with Harness & Weight Updates task performance on LawBench (unspecified metric in abstract)
Combining both levers (harness updates and weight updates) outperforms scaffold iteration alone on all three benchmarks.
Empirical comparison reported in the paper: experiments across the three domains comparing SIA (combined harness+weight updates) against scaffold-iteration-only baseline.
high positive SIA: Self Improving AI with Harness & Weight Updates overall task performance relative to scaffold-only baseline
These results show that per-query configuration of the full retrieval pipeline is a practical alternative to static workload-level tuning.
Authors' conclusion drawn from the reported empirical results on MuSiQue, BrowseComp-Plus, and FinanceBench demonstrating BRANE's performance advantages.
high positive Natural Language Query to Configuration for Retrieval Agents practicality of per-query configuration vs static tuning
BRANE outperforms LLM-routing, rule-based, and fine-tuned Qwen3-4B baselines.
Empirical comparisons against specified baselines (LLM-routing, rule-based approaches, and fine-tuned Qwen3-4B) across the reported benchmark sets. The text does not provide numeric performance metrics or sample sizes in the excerpt.
high positive Natural Language Query to Configuration for Retrieval Agents answer quality (accuracy) and/or cost-quality tradeoff relative to baselines
BRANE matches the best fixed configuration's accuracy at up to 89% lower cost.
Empirical result reported by the authors based on experiments on the named benchmarks (MuSiQue, BrowseComp-Plus, FinanceBench). The provided text states the magnitude ('up to 89% lower cost') but does not give sample sizes or confidence intervals.
high positive Natural Language Query to Configuration for Retrieval Agents inference serving cost while maintaining accuracy
Across MuSiQue, BrowseComp-Plus, and FinanceBench, BRANE consistently pushes the cost-quality Pareto frontier.
Empirical evaluation reported on three benchmark suites: MuSiQue, BrowseComp-Plus, and FinanceBench. The claim is based on experimental comparisons across these datasets; the paper does not state numeric sample sizes in the provided text.
high positive Natural Language Query to Configuration for Retrieval Agents cost-quality tradeoff / Pareto frontier position
Hybrid Fusion significantly accelerated the recovery of smaller Slow AI teams (+6.9% at N=4).
Reported intervention result: Hybrid Fusion produced a +6.9% acceleration in recovery for smaller Slow AI teams, reported at N=4.
high positive The Timing Dependencies of Trust: Speed, Accuracy, and cBCI ... team recovery acceleration (performance improvement) after Hybrid Fusion
Integrating these isolated veridical signals via Hybrid Fusion successfully rescued the Fast AI team (+7.6% at N=8).
Reported intervention result: application of Hybrid Fusion integration produced a +7.6% improvement in Fast AI team performance, reported at N=8.
high positive The Timing Dependencies of Trust: Speed, Accuracy, and cBCI ... team performance improvement after Hybrid Fusion
The Riemannian Oracle adapted to task states by heavily restricting temporal windows (< 0.8s) to intercept fast reflexive compliance and widening windows (> 1.2s) to capture delayed cognitive conflict.
Reported algorithmic behavior of the 2D Adaptive Riemannian Oracle in response to measured spatial covariance: window sizes described as <0.8s for fast states and >1.2s for slow states.
high positive The Timing Dependencies of Trust: Speed, Accuracy, and cBCI ... temporal gating/window size of the Riemannian Oracle
In the Slow AI condition, behavioural teams (N=8) eventually recovered to 100.0%.
Reported team performance metric for behavioural teams in Slow AI condition with N=8; team performance reported to reach 100.0%.
high positive The Timing Dependencies of Trust: Speed, Accuracy, and cBCI ... team accuracy/recovery over time
The proposed policy framework contributes to establishing a foundation for Vietnam to proactively embrace the Agent Economy safely and effectively.
Claim in abstract about the intended contribution/impact of the proposed framework; no empirical evaluation or measured outcomes presented.
high positive Regulatory Policy for the Agent Economy in the Digital Age: ... capacity of Vietnam to embrace Agent Economy safely/effectively (foundation-buil...
The Agent Economy promises substantial gains in productivity and innovation.
Asserted in paper abstract as an anticipated outcome; no empirical measurement, sample size, or quantified effect provided.
high positive Regulatory Policy for the Agent Economy in the Digital Age: ... productivity and innovation gains
Future evaluations should use artifact-level denominators, reproducible parsing rules, correction taxonomies, and independent coding of governance events.
Authors' recommendations based on methodological lessons from this structured self-observed implementation case study and observed parsing/governance challenges.
high positive Persistent AI Agents in Academic Research: A Single-Investig... recommended methodological practices for future evaluations (artifact-level deno...
AI assistance can stabilize an overloaded workflow only when (i) the fraction of tasks handled by AI exceeds a critical threshold, and (ii) the human attention required for review and expected rework is lower than the attention required for manual completion.
Formal analytical conditions derived from the paper's queueing model (model-based theoretical result; no empirical sample reported).
high positive Queue & AI: When Faster Tasks Slow Down the Workflow organizational_efficiency
The paper calls for action by stakeholders to consider human and environmental moderators when adopting AI.
Policy/recommendation statement in the paper's conclusion/abstract; normative recommendation rather than empirical finding.
high positive Position: Adopting AI in Practice Does Not Guarantee the Pro... stakeholder policies and actions regarding AI adoption and moderation
We revise the existing framework to redefine effective organizational determinants and shed light on practical implications including industry and education.
Authors' proposed theoretical revision of an existing framework and discussion of implications; presented as a conceptual contribution within the paper.
high positive Position: Adopting AI in Practice Does Not Guarantee the Pro... organizational determinants and practical implications for industry and educatio...
Most practitioners assume that AI brings productivity boosts owing to enhanced technical capabilities.
Statement of common practitioner belief reported by the authors in the paper's framing; no supporting survey or sample reported in the abstract.
high positive Position: Adopting AI in Practice Does Not Guarantee the Pro... perceived productivity benefits from AI
Adoption of Claude Code increases cumulative lifetime languages used by +0.51.
Panel analysis of 5,838 developers over 28 months using the Callaway & Sant'Anna estimator; treatment = first Claude-co-authored commit.
high positive Coding Beyond Your Training: Claude Code and the Technologic... cumulative lifetime programming languages (count)
Adoption of Claude Code increases the count of newly-used languages by +0.31.
Same dataset and staggered-rollout estimator (Callaway & Sant'Anna), treatment = first Claude-co-authored commit; not-yet-treated controls.
high positive Coding Beyond Your Training: Claude Code and the Technologic... newly-used programming languages (monthly)
Adoption of Claude Code increases Shannon language entropy by +0.14.
Estimated with the doubly robust Callaway & Sant'Anna approach on the 5,838-developer panel over 28 months, using first Claude-co-authored commit as treatment.
high positive Coding Beyond Your Training: Claude Code and the Technologic... Shannon language entropy (diversity of languages used)
Adoption of Claude Code increases the number of distinct programming languages used by a developer by +0.83.
Same panel and staggered-rollout estimation as above (Callaway & Sant'Anna), treatment = first Claude-co-authored commit.
high positive Coding Beyond Your Training: Claude Code and the Technologic... distinct programming languages used (monthly)
Adoption of Claude Code increases the number of repositories a developer contributes to by +1.5 (monthly).
Same panel (5,838 developers, 28 months) and estimator (Callaway & Sant'Anna). Treatment = first Claude-co-authored commit; not-yet-treated controls.
high positive Coding Beyond Your Training: Claude Code and the Technologic... repositories contributed to (monthly)
Adoption of Claude Code is associated with an increase of +41 monthly commits per developer.
Analysis of a panel of 5,838 GitHub developers observed monthly over 28 months, exploiting staggered rollout of Claude Code (May 2025–Jan 2026). Treatment defined by developer's first Claude-co-authored commit; not-yet-treated developers used as controls. Estimates from the doubly robust Callaway and Sant'Anna (2021) staggered-difference-in-differences estimator.
Structured AI-based interventions provide causal evidence that they can transform access to scientific feedback from a largely private advantage into a more widely distributed resource.
Causal inference based on randomized field experiment showing increased revision likelihood and broader uptake of LLM tools across diverse regions and author groups.
high positive Human-AI Collaboration in Science at Scale: A Global Large-s... access and distribution of scientific feedback (measured via treated authors' be...
Effects were strongest among teams with lower h-indexes and earlier career stages.
Heterogeneous treatment effects by team-level metrics (h-index) and career stage reported in the randomized experiment.
high positive Human-AI Collaboration in Science at Scale: A Global Large-s... treatment effect (e.g., revision likelihood) by team h-index and author career s...
Effects were strongest for manuscripts less embedded in the scholarly literature.
Heterogeneous treatment effects reported by manuscript-level embedding in literature (e.g., referencing/citation context) within the randomized experiment.
high positive Human-AI Collaboration in Science at Scale: A Global Large-s... treatment effect (e.g., revision likelihood) by degree of manuscript embeddednes...
Effects of AI feedback were strongest among authors from non-English-dominant research regions.
Heterogeneous treatment effects reported in the randomized experiment stratified by authors' geographic / language-dominance region; sample includes authors from 133 geographic regions.
high positive Human-AI Collaboration in Science at Scale: A Global Large-s... treatment effect on revision likelihood (or other measured outcomes) by region
Exposure to AI feedback increased authors' subsequent use of LLM tools in their future papers, suggesting longer-run shifts in scientific practice.
Follow-up measurements in the randomized field experiment tracking authors' later behavior (use of LLM tools in subsequent papers); comparison between treatment and control authors.
high positive Human-AI Collaboration in Science at Scale: A Global Large-s... subsequent use of LLM tools in future papers