The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests About 🎲 Workforce Futures
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Evidence (8974 claims)

Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.

The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).

Browse by theme

Nine broad, paper-level topics. Click one to filter the claims below.

Adoption
10085 claims
Filter claims →
Productivity
8974 claims
Filtered →
Governance
8062 claims
Filter claims →
Human-AI Collaboration
7749 claims
Filter claims →
Org Design
5057 claims
Filter claims →
Innovation
4896 claims
Filter claims →
Labor Markets
4088 claims
Filter claims →
Skills & Training
3372 claims
Filter claims →
Inequality
2377 claims
Filter claims →

Claims by outcome category

Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.

Outcome Positive Negative Mixed Null Total
Other 882 244 117 1097 2424
Governance & Regulation 1010 469 229 135 1875
Organizational Efficiency 977 235 149 90 1462
Technology Adoption Rate 781 299 143 128 1362
Research Productivity 506 155 74 363 1110
Output Quality 555 219 71 70 915
Decision Quality 395 200 95 54 751
Firm Productivity 523 67 101 27 724
AI Safety & Ethics 262 309 75 36 688
Market Structure 195 201 135 30 566
Task Allocation 248 77 96 38 464
Innovation Output 300 34 55 20 411
Skill Acquisition 207 75 65 21 368
Employment Level 138 67 119 24 350
Fiscal & Macroeconomic 156 80 53 33 329
Task Completion Time 211 38 13 16 280
Firm Revenue 183 52 29 5 270
Consumer Welfare 131 77 48 13 269
Inequality Measures 50 141 54 9 254
Worker Satisfaction 104 85 25 13 227
Error Rate 87 112 11 5 215
Automation Exposure 69 69 37 20 198
Wages & Compensation 102 49 31 11 193
Team Performance 115 30 30 11 187
Regulatory Compliance 88 74 17 7 186
Training Effectiveness 109 22 14 21 168
Developer Productivity 116 21 15 8 161
Job Displacement 12 92 26 1 131
Hiring & Recruitment 57 12 9 5 83
Skill Obsolescence 6 59 10 2 77
Social Protection 43 17 8 2 70
Creative Output 35 21 9 4 70
Labor Share of Income 18 23 17 1 59
Worker Turnover 15 16 4 35
Industry 1 1
Clear
Productivity Remove filter
Meta-analyses show that AI assistance tends to improve human performance compared to working alone.
Reference to existing meta-analyses in the literature reported by the authors (meta-analytic evidence aggregated across studies; no specific meta-analysis names, sample sizes, or quantitative pooled effects provided in the excerpt).
high positive Addressing the Synergy Gap: The Six Elements of the Design S... human performance with AI assistance versus human performance alone
AI is now embedded in healthcare, finance, policy, and many other domains.
Statement in the paper's introduction/abstract summarizing the current deployment of AI across domains (literature observation, no specific empirical study or sample size cited).
high positive Addressing the Synergy Gap: The Six Elements of the Design S... embedding/adoption of AI in multiple domains
Agentic AI does not eliminate engineering discipline; it increases the value of requirements, constraints, traceability, independent verification, and human approval.
Conclusion drawn from synthesis of evidence across multiple domains and argumentation in the paper.
high positive Agentic Agile-V: From Vibe Coding to Verified Engineering in... importance/value of engineering practices (requirements, traceability, verificat...
Agentic Agile-V and the task-level SCOPE-V loop (Specify, Constrain, Orchestrate, Prove, Evolve, Verify) convert conversational intent into structured engineering artifacts and acceptance evidence.
The paper proposes this process framework (the claim is the proposed function of the framework; no empirical evaluation given in the abstract).
high positive Agentic Agile-V: From Vibe Coding to Verified Engineering in... ability to convert conversational intent into structured artifacts and acceptanc...
Controlled studies report productivity gains in some enterprise tasks.
Controlled experimental studies referenced by the paper (specific trials/stats not provided in abstract).
high positive Agentic Agile-V: From Vibe Coding to Verified Engineering in... productivity on enterprise software tasks
These capabilities make software and hardware development faster in some settings.
Aggregated evidence cited in the paper including controlled studies and adoption studies (details not specified in abstract).
high positive Agentic Agile-V: From Vibe Coding to Verified Engineering in... speed of software and hardware development
Agentic AI coding systems can inspect repositories, plan implementation steps, edit files, call tools, run tests, and submit pull requests.
Descriptive synthesis of existing agentic systems and demonstrations referenced in the paper (literature/examples); no single study or sample size given in the abstract.
high positive Agentic Agile-V: From Vibe Coding to Verified Engineering in... agent capabilities (repository inspection, planning, editing, tool use, testing,...
The Claude family leads the benchmark and produces the most professional-looking outputs in our qualitative review.
Empirical result reported from the paper's benchmark and qualitative review of agent outputs (specific metrics, number of agents/tasks, and quantitative scores not provided in the excerpt).
high positive WorkstreamBench: Evaluating LLM Agents on End-to-End Spreads... output professionalism/quality
We develop an evaluation taxonomy comprising three dimensions: Accuracy, Formula, and Format, each comprising fine-grained criteria that reflect professional standards.
Methodological contribution stated in paper; described taxonomy elements (Accuracy, Formula, Format) as part of the evaluation design.
high positive WorkstreamBench: Evaluating LLM Agents on End-to-End Spreads... evaluation criteria/taxonomy
We provide one of the first evaluations of agents on end-to-end spreadsheet tasks, focusing on economically critical financial workflows such as modeling and scenario analysis.
Claim of contribution in the paper; refers to the authors' own evaluation study (details like number of tasks/agents not provided in the excerpt).
high positive WorkstreamBench: Evaluating LLM Agents on End-to-End Spreads... existence of evaluation on end-to-end spreadsheet tasks
Frontier AI labs have developed agents that can construct entire spreadsheets from scratch.
Asserted in paper as background/context; no specific models, numbers, or experimental details provided in the excerpt.
high positive WorkstreamBench: Evaluating LLM Agents on End-to-End Spreads... agent capability to construct spreadsheets
LLM agents are increasingly expected to carry out end-to-end workflows, producing complete artifacts from high-level user instructions.
Framing statement in paper; no empirical data or sample size reported to support the trend claim within the excerpt.
high positive WorkstreamBench: Evaluating LLM Agents on End-to-End Spreads... expectations of agent capabilities (trend)
A six-phase, stepwise implementation framework (ABC-XYZ segmentation, forecast model selection, safety stock calibration, replenishment policy assignment, simulation-based parameter tuning, KPI governance) enables enterprises to achieve 9–16% reductions in inventory costs within existing WMS and ERP architectures.
Practical implications presented in the paper proposing a six-phase implementation framework and asserting expected inventory cost reductions of 9–16% when deployed within existing WMS/ERP.
high positive Equitable railway corridor investment under demand uncertain... expected inventory cost reduction achievable by implementing the proposed framew...
Learning-based control methods deliver up to 16% cost reductions under complex network conditions but require substantial data and governance infrastructure.
Findings from included studies (narrative and/or quantitative results) reporting maximum observed reductions 'up to 16%' and qualitative synthesis noting data/governance requirements.
high positive Equitable railway corridor investment under demand uncertain... inventory cost reduction achieved by learning-based control methods; infrastruct...
The cost reduction from multi-echelon coordination increases significantly with network complexity and lead-time variability.
Pre-specified moderator analyses reported in the paper showing effect size growth with network complexity and lead-time variability.
high positive Equitable railway corridor investment under demand uncertain... magnitude of multi-echelon coordination cost reduction as a function of network ...
Multi-echelon coordination yields a pooled mean cost reduction of 11.4% (95% CI: 6.9–15.9%).
Random-effects meta-analysis pooling percentage cost-reduction effect sizes (reported pooled mean and 95% CI).
high positive Equitable railway corridor investment under demand uncertain... inventory cost reduction from multi-echelon coordination
The advantage of distributional safety stock methods is largest for high-variability SKU segments.
Pre-specified subgroup and moderator analyses reported in the paper indicating greater pooled effects in high-variability SKU segments.
high positive Equitable railway corridor investment under demand uncertain... relative cost-reduction advantage of distributional safety-stock vs normal appro...
Distributional safety stock methods outperform classical normal approximations by a pooled mean of 9.3% (95% CI: 5.8–12.7%) at equivalent service levels.
Random-effects meta-analysis pooling percentage cost-reduction effect sizes (reported pooled mean and 95% CI).
high positive Equitable railway corridor investment under demand uncertain... inventory cost reduction at equivalent service levels
Given the mixed outcomes (some improvements, some new lint/security issues), stronger tool-in-the-loop quality and security gating is motivated for AI-driven development workflows.
Interpretation/recommendation based on observed mix of improvements and introduced issues from the empirical results (PyQu, Pylint, Bandit analyses) and high merge rates.
high positive Quality and Security Signals in AI-Generated Python Refactor... policy/process recommendation (quality/security gating)
73.5% of the analyzed PRs are merged (developer acceptance is high).
Empirical measurement of PR outcomes (merged vs. not merged) in the AIDev dataset of Python refactoring PRs.
high positive Quality and Security Signals in AI-Generated Python Refactor... PR merge rate (acceptance)
Usability is the quality attribute that improves most frequently, improving in 36.5% of the studied changes.
PyQu-based before-and-after analysis of quality attributes on Python refactoring PRs from the AIDev dataset; reported frequency for the 'usability' attribute.
high positive Quality and Security Signals in AI-Generated Python Refactor... usability (one of PyQu's quality attributes)
Agentic commits improve a quality attribute in 22.5% of the studied changes.
Empirical analysis of Python refactoring pull requests from the AIDev dataset using PyQu (an ML-based Python quality assessment tool) to compare quality attributes before and after each change.
high positive Quality and Security Signals in AI-Generated Python Refactor... improvement in any measured code quality attribute (per change)
Current models demonstrate promising spatial grounding, multimodal alignment, and coordinated action execution.
Qualitative and/or quantitative evaluation results in paper indicating strengths in spatial grounding, multimodal alignment, and coordinated action execution.
high positive CutVerse: A Compositional GUI Agents Benchmark for Media Pos... spatial grounding, multimodal alignment, coordinated action execution
We develop a lightweight parser that transforms raw screen recordings and low-level interaction logs into structured, compositional GUI action trajectories with precise grounding.
Methodological contribution described in paper: parser implementation that converts recordings and logs into structured GUI action trajectories.
high positive CutVerse: A Compositional GUI Agents Benchmark for Media Pos... ability to produce structured, grounded GUI action trajectories from recordings/...
The tasks involve dense multimodal interfaces and tightly coupled interaction sequences.
Task descriptions and dataset characteristics in paper stating tasks are complex, long-horizon, multimodal, and tightly coupled.
high positive CutVerse: A Compositional GUI Agents Benchmark for Media Pos... interface complexity and interaction coupling in tasks
We curate expert demonstrations across 7 professional applications (e.g., Premiere Pro, Photoshop), covering 186 complex, long-horizon tasks grounded in authentic editing workflows.
Dataset construction reported in paper: curated expert demonstrations spanning 7 applications and 186 tasks (numbers provided in text).
high positive CutVerse: A Compositional GUI Agents Benchmark for Media Pos... size and scope of demonstration dataset (number of applications and tasks)
We introduce Cutverse, a benchmark designed to systematically evaluate autonomous GUI agents in realistic media post-production environments.
Paper describes the creation of the Cutverse benchmark as a central contribution (design and implementation described in methods).
high positive CutVerse: A Compositional GUI Agents Benchmark for Media Pos... existence and design of a benchmark for GUI agents in media post-production
GUI agents have made significant progress in web navigation and basic operating system tasks.
Background claim stated in paper referencing prior work on GUI agents applied to web navigation and OS tasks (no specific experiments in this paper to support it).
high positive CutVerse: A Compositional GUI Agents Benchmark for Media Pos... capability progress on web navigation and OS tasks
The architecture successfully manages profiles with 14,000+ scientific facts (125k tokens), enabling sustained operation beyond full-context limits.
Reported stress test / capability demonstration in paper: profile size stated as 14,000+ facts and 125k tokens stored and managed by the system.
high positive Episodic-Semantic Memory Architecture for Long-Horizon Scien... number of scientific facts and token footprint the system can manage (profile ca...
The Dual Process system maintains 70-85% accuracy with 1-2 second latency while using 62% fewer tokens (45,434 vs 120,000+ limit) compared to full-context approaches.
Reported empirical results from the large-scale evaluation (1,440 queries / 15,000 messages) comparing Dual Process to full-context models; exact accuracy, latency, and token-count figures provided in the paper.
high positive Episodic-Semantic Memory Architecture for Long-Horizon Scien... accuracy; latency (seconds); token usage
The Dual Process Memory Architecture decouples immediate episodic needs (constant 10-message window) from long-term consolidated knowledge (growing at approximately 3 tokens/message).
System design description and measured consolidation growth rate reported in the paper; empirical observation of growth rate stated.
high positive Episodic-Semantic Memory Architecture for Long-Horizon Scien... episodic window size; long-term memory growth rate (tokens/message)
Agent Skills, structured packages of procedural knowledge loaded into an LLM agent at inference time, are widely reported to improve task pass rates by an average of 16.2 percentage points across diverse domains.
Authors cite prior Skills benchmarks / aggregated reports (benchmark summary referenced in paper); average improvement reported as 16.2 percentage points across tasks in those benchmarks (implied sample of tasks from the referenced benchmark).
high positive When Skills Don't Help: A Negative Result on Procedural Know... task pass rate (task success rate)
Software products and software R&D contributed 50 percent of the 1.2 percentage point acceleration in nonfarm business labor productivity (2017–2024 relative to 2012–2017).
Empirical decomposition comparing productivity growth rates across periods (2017–2024 vs 2012–2017) in the paper; the authors attribute half of the observed 1.2 percentage point acceleration to software products and software R&D.
high positive AI as an Innovation in the Method of Innovation: Implication... acceleration (difference) in nonfarm business labor productivity growth between ...
Software products and software R&D contributed 50 percent of the 2 percent average growth rate in nonfarm business labor productivity from 2017 to 2024.
Empirical decomposition of nonfarm business labor productivity growth in the United States for the period 2017–2024 reported in the paper (the authors attribute shares of the observed 2% average growth to components including software products and software R&D).
high positive AI as an Innovation in the Method of Innovation: Implication... average growth rate in nonfarm business labor productivity (2017–2024)
AI is already materially affecting official productivity measures in the United States.
Empirical decomposition of U.S. productivity data reported in the paper that attributes portions of measured productivity growth to software-related channels linked to AI.
high positive AI as an Innovation in the Method of Innovation: Implication... official productivity measures (U.S. nonfarm business labor productivity)
Using a framework that separates upstream innovation from downstream production suggests that AI boosts both upstream total factor productivity and intangible capital use downstream.
Model/framework decomposition in the paper (theoretical separation of upstream vs downstream, combined with empirical application to productivity data); the paper reports results consistent with increases in upstream TFP and downstream intangible capital use.
high positive AI as an Innovation in the Method of Innovation: Implication... upstream total factor productivity and downstream intangible capital use
Code cleanliness joins model choice, harness, and prompting as a factor that materially affects agent behaviours.
Conclusion drawn from experimental findings that cleanliness materially influenced agent operational metrics (tokens and revisits) even when pass rates were unchanged.
high positive Does Code Cleanliness Affect Coding Agents? A Controlled Min... factors materially affecting agent behaviour (operational footprint/navigation)
Traditional maintainability principles remain highly relevant in the era of AI-driven development, shaping the computational cost and navigational efficiency of coding agents.
Interpretation based on experimental results showing token and navigational efficiency gains on cleaner code (7–8% fewer tokens, 34% fewer revisitations) despite unchanged pass rates.
high positive Does Code Cleanliness Affect Coding Agents? A Controlled Min... relevance of maintainability principles to agent computational cost and navigati...
Agents working on cleaner code reduce file revisitations by 34%.
Empirical measurement across the same experimental trials comparing agent file-revisitation counts between clean and messy repo variants; reported 34% reduction in file revisitations on cleaner code.
high positive Does Code Cleanliness Affect Coding Agents? A Controlled Min... file revisitations (number of times agents revisit files)
Agents working on cleaner code use 7 to 8% fewer tokens.
Empirical measurement across trials (660 trials with Claude Code) comparing token consumption between clean and messy repository variants; reported decrease of 7-8% in tokens when working on cleaner code.
high positive Does Code Cleanliness Affect Coding Agents? A Controlled Min... token usage (number of tokens consumed by agent pipelines)
We author 33 tasks across six such pairs, evaluated through hidden tests at the application's public surface.
Reported experimental design: 33 authored tasks spanning six repository pairs; evaluation used hidden tests executed at the application's public surface.
high positive Does Code Cleanliness Affect Coding Agents? A Controlled Min... number of tasks and pairs used in evaluation
The pairs are constructed in both directions, by agent pipelines that either degrade a clean repository or clean a messy one.
Method description: authors constructed pairs bidirectionally using agent pipelines that modify repositories to create matched clean/messy variants.
high positive Does Code Cleanliness Affect Coding Agents? A Controlled Min... directional construction of repository pairs (degrade or clean)
We introduce an evaluation protocol built around minimal pairs: repositories that match on architecture, dependencies, and external behaviour, but differ on static-analysis rule violations and cognitive complexity.
Methodological description in paper: construction of paired repositories controlling for architecture, dependencies, and external behaviour while varying static-analysis violations and cognitive complexity.
high positive Does Code Cleanliness Affect Coding Agents? A Controlled Min... evaluation protocol (minimal-pair control of repository cleanliness)
A simple prompt checklist can improve LLM responses while reducing unnecessary interaction.
Authors' interpretation/conclusion drawn from the experimental comparisons and rubric scores reported in the paper's results.
high positive Less Back-and-Forth: A Comparative Study of Structured Promp... output_quality and user_interaction
Checklist prompts produced the best quality-effort tradeoff, using fewer average tokens than both raw and clarifying prompts.
Reported comparative statement in the results that checklist prompts used fewer average tokens and produced a better quality-effort tradeoff (no token counts, sample size, or statistical tests reported in the abstract).
high positive Less Back-and-Forth: A Comparative Study of Structured Promp... average_tokens_used (user effort) and output_quality
Checklist-improved prompts achieved the highest mean rubric score, 7.50 out of 8, compared with 5.67 for raw prompts and 6.67 for clarifying-question prompts.
Reported mean rubric scores for each prompt condition in the paper's results (no sample sizes or significance tests provided in the abstract).
high positive Less Back-and-Forth: A Comparative Study of Structured Promp... rubric_score (task completion / correctness / compliance / clarity)
The authors open-source optimize_anything with support for multiple backends as part of the GEPA project at https://github.com/gepa-ai/gepa.
Explicit statement and provided GitHub URL in the paper excerpt.
high positive optimize_anything: A Universal API for Optimizing any Text P... availability of open-source code / tooling
Multi-task search outperforms independent optimization given equivalent per-problem budget through cross-task transfer, with benefits scaling with the number of related tasks.
Reported experiments comparing multi-task search versus independent per-problem optimization under equal per-problem budget; observed cross-task transfer benefits and that benefits increase with more related tasks.
high positive optimize_anything: A Universal API for Optimizing any Text P... optimization performance (e.g., score) under multi-task vs independent optimizat...
Ablations across three domains reveal that actionable side information yields substantially higher final scores than score-only feedback.
Same ablation studies across three domains as above; reported higher final optimization scores when using actionable side information compared to only score feedback.
Ablations across three domains reveal that actionable side information yields faster convergence than score-only feedback.
Paper reports ablation studies in three domains comparing optimization with actionable side information versus score-only feedback and finds faster convergence with side information.
high positive optimize_anything: A Universal API for Optimizing any Text P... convergence speed (time or iterations to converge)