The Commonplace
Home Three-study pilot Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Evidence (323 claims)

Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.

The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).

Browse by theme

Nine broad, paper-level topics. Click one to filter the claims below.

Adoption
21267 claims
Filter claims →
Productivity
17978 claims
Filter claims →
Governance
17038 claims
Filter claims →
Human-AI Collaboration
16914 claims
Filter claims →
Org Design
11104 claims
Filter claims →
Innovation
11087 claims
Filter claims →
Labor Markets
6711 claims
Filter claims →
Skills & Training
5616 claims
Filter claims →
Inequality
4343 claims
Filter claims →

Claims by outcome category

Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.

Outcome Positive Negative Mixed Null Total
Other 1880 496 296 1854 4721
Organizational Efficiency 2906 665 438 180 4210
Governance & Regulation 2162 929 480 247 3866
Technology Adoption Rate 1533 545 278 210 2593
Decision Quality 1391 534 321 173 2429
Output Quality 1298 472 231 145 2153
AI Safety & Ethics 682 821 230 90 1837
Research Productivity 855 253 121 425 1675
Firm Productivity 1105 171 175 73 1531
Task Allocation 735 229 361 99 1433
Market Structure 457 461 251 47 1222
Innovation Output 673 94 108 36 913
Task Completion Time 499 118 43 38 702
Firm Revenue 458 130 61 26 677
Skill Acquisition 381 122 113 34 650
Consumer Welfare 316 176 115 39 648
Employment Level 223 143 177 53 600
Error Rate 246 282 44 19 594
Fiscal & Macroeconomic 283 142 78 52 562
Inequality Measures 103 329 106 13 552
Worker Satisfaction 225 185 63 30 503
Automation Exposure 158 155 72 37 426
Regulatory Compliance 186 126 35 14 362
Team Performance 193 56 51 24 326
Developer Productivity 224 58 27 13 323
Wages & Compensation 148 108 50 17 323
Training Effectiveness 218 44 21 27 313
Job Displacement 23 159 53 5 240
Hiring & Recruitment 109 61 32 11 215
Skill Obsolescence 16 107 26 6 155
Creative Output 71 44 28 6 150
Social Protection 58 31 12 3 104
Labor Share of Income 29 43 25 2 99
Worker Turnover 45 29 6 4 84
Industry 1 1
Google's DORA survey found that individual developer effectiveness increased by 17%, while delivery stability declined by nearly 10%.
Industry-wide developer survey conducted by Google's DORA team.
high mixed AI and the Future of Software Engineering: Expertise, Employ... Individual developer effectiveness and software delivery stability
Bug-fixing performance varied across functional areas, with the reported model rates differing substantially between areas such as Manufacturing and Warehouse.
Table 10 groups task-level mean resolution rates by functional area for GPT-5.2 and Opus-4.6.
high mixed BC-Bench: Evaluating Agentic Engineering in a Domain-Specifi... Bug-fixing mean resolution rate by functional area
In the reported Bug Fixing configurations, claude-opus-4.6 had the highest passˆ5 and mean resolution rate, but it was slower than competing systems.
Comparison of reported model configurations in Table 3; the paper cautions that benchmark versions differ across rows.
high mixed BC-Bench: Evaluating Agentic Engineering in a Domain-Specifi... Bug-fixing resolution reliability and agent execution duration
The study suggests that AI awareness may generate simultaneous engagement-related productivity benefits and insecurity-related risks, so the net productivity effect depends on the relative strength of these channels.
Interpretation of the observed positive AI awareness → work engagement path, negative job insecurity → engagement path, and positive work engagement → individual performance path; based on a cross-sectional sample of 227 workers.
Dijital beceri düzeyi yüksek çalışanlar yapay zekâyı üretkenliği artırıcı bir araç olarak daha etkin kullanırken, düşük dijital becerilere sahip çalışanlar yapay zekânın faydalarından daha sınırlı yararlanabilmektedir.
Georgieff ve Hyee (2022) tarafından bildirilen dijital beceri düzeyine göre farklılaşan uyum ve faydalanma sonuçları.
high mixed Yapay Zekânın Endüstri İlişkilerine Etkileri: Sendika, Toplu... Yapay zekâdan yararlanma ve üretkenlik kazanımı
AI models could approximate entry-level data scientists on routine tasks, but they require verification.
Conclusion drawn from the benchmark evaluations comparing model outputs on routine/structured tasks to expected entry-level data scientist performance (qualitative claim; no quantitative equivalence metrics provided in excerpt).
high mixed Benchmarking AI Performance on End-to-End Data Science Proje... comparative performance between AI models and entry-level data scientists on rou...
The extent to which generative AI models can complete end-to-end data science projects varies considerably by model.
Empirical claim based on the benchmark evaluations of multiple generative AI models using the 40 end-to-end projects and the automated grading pipeline (models compared; exact model count not provided in excerpt).
high mixed Benchmarking AI Performance on End-to-End Data Science Proje... models' ability to complete end-to-end data science projects (performance variat...
Participants who fully delegated coding tasks showed some productivity improvements, but at the cost of learning the library.
Subgroup analysis within the randomized experiment of participants who fully delegated coding to the AI; observed productivity gains for this subgroup alongside reduced learning outcomes (library mastery measures).
high mixed How AI Impacts Skill Formation productivity (task performance) and learning/mastery of the library
The SLMs' limited reasoning ability is the primary bottleneck for task success (i.e., success rates), while framework design is the primary bottleneck for efficiency (energy/resource usage).
Interpretation of empirical results: near-zero task resolution attributed to SLM reasoning limits; differences in energy consumption attributed to architecture differences across frameworks (same SLMs used across frameworks; 150 runs per configuration).
high mixed SWEnergy: An Empirical Study on Energy Efficiency in Agentic... task success (resolution rate) and energy consumption
While a small number of projects exceed an industry-reported estimate of 36 PRs per participant during the three-month observation period, most projects remain below this threshold.
Comparison of projects' PRs-per-participant metric (computed from the dataset of 2,361 repositories) against an industry-reported benchmark of 36 PRs per participant for the three-month window.
high mixed Early Adoption of Agentic Coding Tools by GitHub Projects PRs per participant during three-month period relative to 36-PR benchmark
Meta-analytic evidence shows moderate but heterogeneous effects of agentic/code-generation tools on productivity.
Reference to meta-analytic synthesis across studies reported in the paper (meta-analytic details not provided in abstract).
high mixed Agentic Agile-V: From Vibe Coding to Verified Engineering in... aggregate effect on productivity across studies
High-AIC participants realized outsized gains from GenAI access; low-AIC participants saw limited or even negative marginal returns.
Subgroup analysis of the randomized experiment comparing treatment effects by AIC level; authors report large positive treatment effects for high-AIC subgroup and small or negative effects for low-AIC subgroup.
high mixed Generative AI and the Productivity Divide: Human-AI Compleme... treatment effect on task performance by AIC subgroup
Five other continuous features and three of seven binary patterns from prior SE literature show similar directional disagreement across configurations.
Aggregate empirical finding across the set of features and binary patterns analyzed in the 126-configuration dataset.
high mixed Same Signal, Different Semantics: A Cross-Framework Behavior... directional agreement/disagreement of feature–outcome relations for five continu...
Error rate is the cleanest case: 47 configurations resolve more issues when their error rate is lower, while 48 resolve more when it is higher.
Empirical counts from the paper's analysis of configurations (reported 47 vs 48 configurations showing opposite sign relations between error rate and issue resolution).
high mixed Same Signal, Different Semantics: A Cross-Framework Behavior... issue resolution count/rate as a function of error rate
On most signals, configurations disagree not merely in magnitude but in direction (i.e., the same signal correlates positively with resolution in some configurations and negatively in others).
Across-configuration comparison of behavior–outcome correlations for many signals in the dataset of 126 configurations / 64,380 runs.
high mixed Same Signal, Different Semantics: A Cross-Framework Behavior... direction of correlation between behavioral signals and issue resolution
There is substantial heterogeneity in the productivity effects across settings.
Meta-analytic heterogeneity assessment reported in the paper (subgroup/moderator analyses indicate variability by context). The paper states 'substantial heterogeneity across settings.'
high mixed A meta-analysis of the effect of generative AI on productivi... variation in productivity effect sizes across study contexts
Delegating tasks to genAI can be individually beneficial in the short term even as widespread adoption degrades future model performance (creating a social dilemma).
Result of the paper's behavioral model showing an individual-level incentive to use genAI versus a collective cost from adoption (theoretical/model-based; no empirical sample reported in abstract).
high mixed Generative artificial intelligence reduces social welfare th... individual short-term benefit vs future model performance (collective welfare)
How software developers interact with AI-powered tools, including Large Language Models (LLMs), plays a vital role in how these AI-powered tools impact them.
Based on qualitative analysis of twenty-two interviews with software developers about using LLMs for software development; asserted as a central finding in the paper's analysis.
high mixed Towards an Appropriate Level of Reliance on AI: A Preliminar... impact of AI tools on developers (broadly: productivity, skills, quality)
CLARITI matches GPT-5's resolution rate on underspecified issues while generating 41% fewer questions.
Empirical evaluation comparing CLARITI and GPT-5 on a task set of underspecified software engineering issues; the result reported in the abstract indicates parity in resolution rate and a quantified reduction in questions (41%) but the abstract does not report sample size, test set composition, or statistical significance.
high mixed Asking What Matters: Reward-Driven Clarification for Softwar... resolution rate (task success) and number of clarifying questions generated
The rise of agentic AI development, where LLM-based agents autonomously read, write, navigate, and debug codebases, introduces a new primary consumer with fundamentally different constraints.
Conceptual claim argued in the paper; refers to the emergence of agentic LLM-based tools as new consumers of software artifacts rather than an empirical measurement; no sample size reported.
high mixed Beyond Human-Readable: Rethinking Software Engineering Conve... who/what is the primary consumer of software engineering artifacts (human develo...
Better agents improve the coefficient on human effort but not the exponent (i.e., they reduce the constant factor but do not change the asymptotic scaling class).
Analytic result from the stylized model under the paper's assumptions about task decomposition and novelty fraction ν.
high mixed The Novelty Bottleneck: A Framework for Understanding Human ... human effort (coefficient vs. asymptotic scaling exponent)
The enforcement layer can itself become a primary source of outages by rejecting correct work and causing delegations to exhaust their budgets without accepted writes.
Twelve distinct recorded enforcement incidents; the most expensive involved a repair clamp that blocked both writes and reads to the misdiagnosed target set.
high negative Agent Mesh: Reliability Primitives for Non-Idempotent Agent ... Successful completion and accepted writes under enforcement controls
The dominant delegation failure mode in the incident corpus was a sequence of individually successful actions that stopped changing the outcome rather than an explicit tool error.
Retrospective classification of failures across the incident corpus, supported by the detailed successful-call loop incident.
high negative Agent Mesh: Reliability Primitives for Non-Idempotent Agent ... Delegation convergence and outcome change during agent execution
In a randomized trial of experienced developers working in large, unfamiliar codebases, early-2025 AI tools reduced measured developer speed by approximately 19%, despite developers predicting a 24% speed-up.
METR randomized controlled trial involving 16 experienced developers completing 246 tasks.
high negative AI and the Future of Software Engineering: Expertise, Employ... Developer productivity measured by task completion speed
Fixed implementation spaces can be inefficient because high-level operators may introduce dispatch and intermediate-materialization overhead for simple tasks, while fully custom CUDA implementations can consume substantial search resources on composite workloads.
The paper motivates cross-granularity planning by describing workload-dependent tradeoffs between high-level operators and custom CUDA; this is a methodological argument supported by cited prior work rather than a reported experiment in the supplied text.
high negative HIERA: Workload-Aware Planning Across Implementation Spaces ... Search efficiency and kernel optimization performance
Resolution accuracy was lower for tasks whose gold patches exceeded 10 lines of code than for tasks whose patches changed 1–10 lines.
Table 9 reports mean resolution rates by gold-patch lines of code for GPT-5.2 and Opus-4.6.
high negative BC-Bench: Evaluating Agentic Engineering in a Domain-Specifi... Bug-fixing mean resolution rate by patch size
Agent resolution accuracy declined substantially when the gold patch modified more than one file rather than only one file.
Table 8 compares mean resolution rates by number of files changed in the gold patch for GPT-5.2 and Opus-4.6.
high negative BC-Bench: Evaluating Agentic Engineering in a Domain-Specifi... Bug-fixing mean resolution rate by patch file count
GPT-4.1 performed substantially worse than recent state-of-the-art models on Bug Fixing, with approximately four times lower mean resolution rate and ten times lower passˆ5.
Comparison of GPT-4.1 with newer configurations in Table 3.
high negative BC-Bench: Evaluating Agentic Engineering in a Domain-Specifi... Bug-fixing mean resolution rate and repeated-run reliability
The reported 9.9× ratio should not be interpreted as a controlled measurement of productivity because its denominator is a retrospective estimate of work that was never performed.
Authors’ methodological qualification; the counterfactual was based on retrospective student judgments rather than a measured non-AI baseline, and the team consisted of students rather than the professionals used to price the counterfactual.
high negative Building AI-Intensive Software with AI: Early Results and a ... Validity and generalizability of the estimated productivity/cost ratio
In software development, experienced developers expected a 24% speedup from AI tools but were slowed by 19% -- a 43 percentage-point calibration error.
Controlled trials and independent validations in software engineering reviewed by the authors reporting the stated expected and observed effects for experienced developers.
high negative Quantifying the Expectation-Realisation Gap for Agentic AI S... developer task completion speed / productivity
Agentic AI systems are deployed with expectations of substantial productivity gains, yet rigorous empirical evidence reveals systematic discrepancies between pre-deployment expectations and post-deployment outcomes.
Review of controlled trials and independent validations across multiple domains (software engineering, clinical documentation, clinical decision support) summarized by the authors.
high negative Quantifying the Expectation-Realisation Gap for Agentic AI S... realised productivity gains versus expected productivity gains
API models restore only 27.6% of infeasible formulations (solver interaction gap).
Reported Phase 1 experimental average recovery rate for API models on the 976-problem benchmark.
high negative OptiRepair: Closed-Loop Diagnosis and Repair of Supply Chain... fraction of infeasible formulations restored by API models
Two gaps separate current AI from reliable model repair: solver interaction, as API models restore only 27.6% of infeasible formulations; and operational rationale, as roughly one in four feasible repairs violate supply chain theory.
Authors' analysis of experimental failures attributing them to two distinct failure modes: low solver-interaction success (27.6%) and domain-rationality violations (~25%).
high negative OptiRepair: Closed-Loop Diagnosis and Repair of Supply Chain... solver-interaction repair success; operational rationality compliance
The gap concentrates in Phase 1 repair, where API models average 27.6% recovery rate versus 97.2% for trained models.
Reported breakdown of recovery performance by phase (Phase 1 feasibility repair) from experiments on the benchmark.
high negative OptiRepair: Closed-Loop Diagnosis and Repair of Supply Chain... Phase 1 (feasibility) recovery rate
There is a wealth of rich runtime information that developers routinely access while debugging code, which agents are currently deprived of due to design limitations.
Author argument based on typical developer use of debuggers and the design of current coding agents; qualitative claim within the paper.
high negative Debug2Fix: Can Interactive Debugging Help Coding Agents Fix ... availability/access to runtime debugging information for agents
There is still significant room for improvement in coding agents' bug fixing capabilities.
Author assertion in abstract/introduction describing current limitations of coding agents; no specific empirical measurement provided in the excerpt.
high negative Debug2Fix: Can Interactive Debugging Help Coding Agents Fix ... bug fixing capability of coding agents
Free models produced between 0 and 4 functional outcomes, with DeepSeek and Gemini Free ranking highest among free models.
Results statement summarizing performance range for the five free models (0–4 implemented functions) and naming top performers among them.
high negative Using Natural Language Prompts With AI Models for Low-Cost A... number of successfully implemented functions
Four failure modes (context saturation, memory interference, dependency complexity, reprioritization overhead) cause baseline collaborative ubiquitous agents (CUAs) to degrade from 16.7% to 8.7% completion as load scales from 25% to 100%.
Empirical experiments reported in the paper showing completion rates at different loads; pattern observed across three independent implementations (three CUA implementations). Exact trial/sample counts not provided in excerpt.
high negative CORPGEN: Simulating Corporate Environments with Autonomous D... task completion rate (percent of tasks completed)
Translating natural-language requirements into correct optimization formulations and solver-executable code remains labor-intensive.
Background statement in the paper describing the task and its current workflow; no empirical numbers provided in the abstract.
high negative Constructing Industrial-Scale Optimization Modeling Benchmar... effort to translate NL to optimization formulations/code
AI-provided instructions often result in performance collapse.
Empirical observation from the same set of experiments (20 experiments, 737 participants) comparing AI-led instruction vs. human-led instruction and hybrid configurations.
high negative Why Human Guidance Matters in Collaborative Vibe Coding task performance under AI-provided instructions (measured as performance collaps...
Optimizing large-scale machine learning systems requires navigating a massive hyperparameter search space and designing sophisticated optimizers, architectures, and reward functions; achieving substantial improvements is traditionally a non-trivial task relying on extensive manual iterations.
Background claim in the paper describing the difficulty of engineering improvements for large-scale ML systems; presented as motivation for the proposed approach (conceptual, commonly accepted in ML engineering literature).
high negative Self-Evolving Recommendation System: End-To-End Autonomous M... manual engineering effort required for model optimization
Existing benchmarks rarely calibrate model performance against that of human programmers.
Explicitly stated as a limitation motivating their human-centered evaluation; no counts or proportions of benchmarks are provided in the abstract.
high negative EvoCodeBench: A Human-Performance Benchmark for Self-Evolvin... relative model vs human performance
This approach has encountered a "usability ceiling" manifested as the Intent-Execution Gap (i.e., the fundamental disparity between a creator's high-level intent and the stochastic, black-box nature of current single-shot models).
Conceptual critique presented in the paper; no empirical evaluation, quantitative measurement, or sample described in the provided text.
high negative Vibe AIGC: A New Paradigm for Content Generation via Agentic... Intent-Execution Gap (disparity between user intent and model execution)
Our results challenge the intuition that declarative framework design guarantees AI-assistability.
Aggregate experimental findings across the evaluated frameworks (including DSPy low score despite being declarative) used to argue against the intuition that declarative design alone ensures AI-assistability.
high negative Declarative by Design, Assistable Only by Convention: Benchm... relationship between declarative design and AI-assistability
DSPy -- the most declarative framework by design -- scores lowest (0.07), as its novel abstractions are insufficiently represented in AI training data.
Empirical result: DSPy reported AI-assistability score = 0.07. The causal explanation (insufficient representation in AI training data) is an interpretation provided by the authors.
high negative Declarative by Design, Assistable Only by Convention: Benchm... AI-assistability score for DSPy and proposed cause for low score
Coordinated renaming is a frequent yet challenging task.
Supported by authors' reported formative study and developer survey (paper states study of 609K commits in 100 OSS projects and survey of 205 developers).
high negative Multi-Agent Coordinated Rename Refactoring frequency_and_difficulty_of_coordinated_renaming
Existing systems often rely on a single agent to handle the entire workflow—interpreting issues, navigating large codebases, and implementing fixes—within one reasoning chain, and such monolithic designs force the model to retain irrelevant context, leading to spurious correlations and poor generalization.
Authors' analysis/argument about the architecture of existing single-agent systems and its drawbacks; described qualitatively in the paper (no empirical quantification provided in the excerpt).
high negative BOAD: Discovering Hierarchical Software Engineering Agents v... model generalization / propensity to form spurious correlations due to retained ...
Large language models (LLMs) have shown strong reasoning and coding capabilities, yet they struggle to generalize to real-world software engineering (SWE) problems that are long-horizon and out of distribution.
Statement in paper framing the problem; based on authors' characterization of prior LLM performance on complex/out-of-distribution SWE tasks (no specific dataset or quantitative result provided in the excerpt).
high negative BOAD: Discovering Hierarchical Software Engineering Agents v... generalization performance on long-horizon, out-of-distribution software enginee...
In the domain of TCAD simulation, the scarcity of open-source resources hinders language models from generating valid TCAD code.
Asserted observation in the paper (motivation); no quantitative study or sample size reported in the abstract.
high negative AgenticTCAD: A LLM-based Multi-Agent Framework for Automated... language model ability to generate valid TCAD code (validity of generated TCAD c...
The link from supportive HR practices to performance weakens in a murky middle where algorithmic oversight is present yet hard to interpret.
Reported moderated mediation results from survey of 464 gig workers using Double Machine Learning; described as a nonmonotonic pattern in the abstract.
high negative When Algorithms Manage Humans: A Double Machine Learning App... worker performance (link from HR practices to performance)