Evidence (323 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
21267 claims
Filter claims →
Productivity
17978 claims
Filter claims →
Governance
17038 claims
Filter claims →
Human-AI Collaboration
16914 claims
Filter claims →
Org Design
11104 claims
Filter claims →
Innovation
11087 claims
Filter claims →
Labor Markets
6711 claims
Filter claims →
Skills & Training
5616 claims
Filter claims →
Inequality
4343 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 1880 | 496 | 296 | 1854 | 4721 |
| Organizational Efficiency | 2906 | 665 | 438 | 180 | 4210 |
| Governance & Regulation | 2162 | 929 | 480 | 247 | 3866 |
| Technology Adoption Rate | 1533 | 545 | 278 | 210 | 2593 |
| Decision Quality | 1391 | 534 | 321 | 173 | 2429 |
| Output Quality | 1298 | 472 | 231 | 145 | 2153 |
| AI Safety & Ethics | 682 | 821 | 230 | 90 | 1837 |
| Research Productivity | 855 | 253 | 121 | 425 | 1675 |
| Firm Productivity | 1105 | 171 | 175 | 73 | 1531 |
| Task Allocation | 735 | 229 | 361 | 99 | 1433 |
| Market Structure | 457 | 461 | 251 | 47 | 1222 |
| Innovation Output | 673 | 94 | 108 | 36 | 913 |
| Task Completion Time | 499 | 118 | 43 | 38 | 702 |
| Firm Revenue | 458 | 130 | 61 | 26 | 677 |
| Skill Acquisition | 381 | 122 | 113 | 34 | 650 |
| Consumer Welfare | 316 | 176 | 115 | 39 | 648 |
| Employment Level | 223 | 143 | 177 | 53 | 600 |
| Error Rate | 246 | 282 | 44 | 19 | 594 |
| Fiscal & Macroeconomic | 283 | 142 | 78 | 52 | 562 |
| Inequality Measures | 103 | 329 | 106 | 13 | 552 |
| Worker Satisfaction | 225 | 185 | 63 | 30 | 503 |
| Automation Exposure | 158 | 155 | 72 | 37 | 426 |
| Regulatory Compliance | 186 | 126 | 35 | 14 | 362 |
| Team Performance | 193 | 56 | 51 | 24 | 326 |
| Developer Productivity | 224 | 58 | 27 | 13 | 323 |
| Wages & Compensation | 148 | 108 | 50 | 17 | 323 |
| Training Effectiveness | 218 | 44 | 21 | 27 | 313 |
| Job Displacement | 23 | 159 | 53 | 5 | 240 |
| Hiring & Recruitment | 109 | 61 | 32 | 11 | 215 |
| Skill Obsolescence | 16 | 107 | 26 | 6 | 155 |
| Creative Output | 71 | 44 | 28 | 6 | 150 |
| Social Protection | 58 | 31 | 12 | 3 | 104 |
| Labor Share of Income | 29 | 43 | 25 | 2 | 99 |
| Worker Turnover | 45 | 29 | 6 | 4 | 84 |
| Industry | — | — | — | 1 | 1 |
Google's DORA survey found that individual developer effectiveness increased by 17%, while delivery stability declined by nearly 10%.
Industry-wide developer survey conducted by Google's DORA team.
Bug-fixing performance varied across functional areas, with the reported model rates differing substantially between areas such as Manufacturing and Warehouse.
Table 10 groups task-level mean resolution rates by functional area for GPT-5.2 and Opus-4.6.
In the reported Bug Fixing configurations, claude-opus-4.6 had the highest passˆ5 and mean resolution rate, but it was slower than competing systems.
Comparison of reported model configurations in Table 3; the paper cautions that benchmark versions differ across rows.
The study suggests that AI awareness may generate simultaneous engagement-related productivity benefits and insecurity-related risks, so the net productivity effect depends on the relative strength of these channels.
Interpretation of the observed positive AI awareness → work engagement path, negative job insecurity → engagement path, and positive work engagement → individual performance path; based on a cross-sectional sample of 227 workers.
Dijital beceri düzeyi yüksek çalışanlar yapay zekâyı üretkenliği artırıcı bir araç olarak daha etkin kullanırken, düşük dijital becerilere sahip çalışanlar yapay zekânın faydalarından daha sınırlı yararlanabilmektedir.
Georgieff ve Hyee (2022) tarafından bildirilen dijital beceri düzeyine göre farklılaşan uyum ve faydalanma sonuçları.
AI models could approximate entry-level data scientists on routine tasks, but they require verification.
Conclusion drawn from the benchmark evaluations comparing model outputs on routine/structured tasks to expected entry-level data scientist performance (qualitative claim; no quantitative equivalence metrics provided in excerpt).
The extent to which generative AI models can complete end-to-end data science projects varies considerably by model.
Empirical claim based on the benchmark evaluations of multiple generative AI models using the 40 end-to-end projects and the automated grading pipeline (models compared; exact model count not provided in excerpt).
Participants who fully delegated coding tasks showed some productivity improvements, but at the cost of learning the library.
Subgroup analysis within the randomized experiment of participants who fully delegated coding to the AI; observed productivity gains for this subgroup alongside reduced learning outcomes (library mastery measures).
The SLMs' limited reasoning ability is the primary bottleneck for task success (i.e., success rates), while framework design is the primary bottleneck for efficiency (energy/resource usage).
Interpretation of empirical results: near-zero task resolution attributed to SLM reasoning limits; differences in energy consumption attributed to architecture differences across frameworks (same SLMs used across frameworks; 150 runs per configuration).
While a small number of projects exceed an industry-reported estimate of 36 PRs per participant during the three-month observation period, most projects remain below this threshold.
Comparison of projects' PRs-per-participant metric (computed from the dataset of 2,361 repositories) against an industry-reported benchmark of 36 PRs per participant for the three-month window.
Meta-analytic evidence shows moderate but heterogeneous effects of agentic/code-generation tools on productivity.
Reference to meta-analytic synthesis across studies reported in the paper (meta-analytic details not provided in abstract).
High-AIC participants realized outsized gains from GenAI access; low-AIC participants saw limited or even negative marginal returns.
Subgroup analysis of the randomized experiment comparing treatment effects by AIC level; authors report large positive treatment effects for high-AIC subgroup and small or negative effects for low-AIC subgroup.
Five other continuous features and three of seven binary patterns from prior SE literature show similar directional disagreement across configurations.
Aggregate empirical finding across the set of features and binary patterns analyzed in the 126-configuration dataset.
Error rate is the cleanest case: 47 configurations resolve more issues when their error rate is lower, while 48 resolve more when it is higher.
Empirical counts from the paper's analysis of configurations (reported 47 vs 48 configurations showing opposite sign relations between error rate and issue resolution).
On most signals, configurations disagree not merely in magnitude but in direction (i.e., the same signal correlates positively with resolution in some configurations and negatively in others).
Across-configuration comparison of behavior–outcome correlations for many signals in the dataset of 126 configurations / 64,380 runs.
There is substantial heterogeneity in the productivity effects across settings.
Meta-analytic heterogeneity assessment reported in the paper (subgroup/moderator analyses indicate variability by context). The paper states 'substantial heterogeneity across settings.'
Delegating tasks to genAI can be individually beneficial in the short term even as widespread adoption degrades future model performance (creating a social dilemma).
Result of the paper's behavioral model showing an individual-level incentive to use genAI versus a collective cost from adoption (theoretical/model-based; no empirical sample reported in abstract).
How software developers interact with AI-powered tools, including Large Language Models (LLMs), plays a vital role in how these AI-powered tools impact them.
Based on qualitative analysis of twenty-two interviews with software developers about using LLMs for software development; asserted as a central finding in the paper's analysis.
CLARITI matches GPT-5's resolution rate on underspecified issues while generating 41% fewer questions.
Empirical evaluation comparing CLARITI and GPT-5 on a task set of underspecified software engineering issues; the result reported in the abstract indicates parity in resolution rate and a quantified reduction in questions (41%) but the abstract does not report sample size, test set composition, or statistical significance.
The rise of agentic AI development, where LLM-based agents autonomously read, write, navigate, and debug codebases, introduces a new primary consumer with fundamentally different constraints.
Conceptual claim argued in the paper; refers to the emergence of agentic LLM-based tools as new consumers of software artifacts rather than an empirical measurement; no sample size reported.
Better agents improve the coefficient on human effort but not the exponent (i.e., they reduce the constant factor but do not change the asymptotic scaling class).
Analytic result from the stylized model under the paper's assumptions about task decomposition and novelty fraction ν.
The enforcement layer can itself become a primary source of outages by rejecting correct work and causing delegations to exhaust their budgets without accepted writes.
Twelve distinct recorded enforcement incidents; the most expensive involved a repair clamp that blocked both writes and reads to the misdiagnosed target set.
The dominant delegation failure mode in the incident corpus was a sequence of individually successful actions that stopped changing the outcome rather than an explicit tool error.
Retrospective classification of failures across the incident corpus, supported by the detailed successful-call loop incident.
In a randomized trial of experienced developers working in large, unfamiliar codebases, early-2025 AI tools reduced measured developer speed by approximately 19%, despite developers predicting a 24% speed-up.
METR randomized controlled trial involving 16 experienced developers completing 246 tasks.
Fixed implementation spaces can be inefficient because high-level operators may introduce dispatch and intermediate-materialization overhead for simple tasks, while fully custom CUDA implementations can consume substantial search resources on composite workloads.
The paper motivates cross-granularity planning by describing workload-dependent tradeoffs between high-level operators and custom CUDA; this is a methodological argument supported by cited prior work rather than a reported experiment in the supplied text.
Resolution accuracy was lower for tasks whose gold patches exceeded 10 lines of code than for tasks whose patches changed 1–10 lines.
Table 9 reports mean resolution rates by gold-patch lines of code for GPT-5.2 and Opus-4.6.
Agent resolution accuracy declined substantially when the gold patch modified more than one file rather than only one file.
Table 8 compares mean resolution rates by number of files changed in the gold patch for GPT-5.2 and Opus-4.6.
GPT-4.1 performed substantially worse than recent state-of-the-art models on Bug Fixing, with approximately four times lower mean resolution rate and ten times lower passˆ5.
Comparison of GPT-4.1 with newer configurations in Table 3.
The reported 9.9× ratio should not be interpreted as a controlled measurement of productivity because its denominator is a retrospective estimate of work that was never performed.
Authors’ methodological qualification; the counterfactual was based on retrospective student judgments rather than a measured non-AI baseline, and the team consisted of students rather than the professionals used to price the counterfactual.
In software development, experienced developers expected a 24% speedup from AI tools but were slowed by 19% -- a 43 percentage-point calibration error.
Controlled trials and independent validations in software engineering reviewed by the authors reporting the stated expected and observed effects for experienced developers.
Agentic AI systems are deployed with expectations of substantial productivity gains, yet rigorous empirical evidence reveals systematic discrepancies between pre-deployment expectations and post-deployment outcomes.
Review of controlled trials and independent validations across multiple domains (software engineering, clinical documentation, clinical decision support) summarized by the authors.
API models restore only 27.6% of infeasible formulations (solver interaction gap).
Reported Phase 1 experimental average recovery rate for API models on the 976-problem benchmark.
Two gaps separate current AI from reliable model repair: solver interaction, as API models restore only 27.6% of infeasible formulations; and operational rationale, as roughly one in four feasible repairs violate supply chain theory.
Authors' analysis of experimental failures attributing them to two distinct failure modes: low solver-interaction success (27.6%) and domain-rationality violations (~25%).
The gap concentrates in Phase 1 repair, where API models average 27.6% recovery rate versus 97.2% for trained models.
Reported breakdown of recovery performance by phase (Phase 1 feasibility repair) from experiments on the benchmark.
There is a wealth of rich runtime information that developers routinely access while debugging code, which agents are currently deprived of due to design limitations.
Author argument based on typical developer use of debuggers and the design of current coding agents; qualitative claim within the paper.
There is still significant room for improvement in coding agents' bug fixing capabilities.
Author assertion in abstract/introduction describing current limitations of coding agents; no specific empirical measurement provided in the excerpt.
Free models produced between 0 and 4 functional outcomes, with DeepSeek and Gemini Free ranking highest among free models.
Results statement summarizing performance range for the five free models (0–4 implemented functions) and naming top performers among them.
Four failure modes (context saturation, memory interference, dependency complexity, reprioritization overhead) cause baseline collaborative ubiquitous agents (CUAs) to degrade from 16.7% to 8.7% completion as load scales from 25% to 100%.
Empirical experiments reported in the paper showing completion rates at different loads; pattern observed across three independent implementations (three CUA implementations). Exact trial/sample counts not provided in excerpt.
Translating natural-language requirements into correct optimization formulations and solver-executable code remains labor-intensive.
Background statement in the paper describing the task and its current workflow; no empirical numbers provided in the abstract.
AI-provided instructions often result in performance collapse.
Empirical observation from the same set of experiments (20 experiments, 737 participants) comparing AI-led instruction vs. human-led instruction and hybrid configurations.
Optimizing large-scale machine learning systems requires navigating a massive hyperparameter search space and designing sophisticated optimizers, architectures, and reward functions; achieving substantial improvements is traditionally a non-trivial task relying on extensive manual iterations.
Background claim in the paper describing the difficulty of engineering improvements for large-scale ML systems; presented as motivation for the proposed approach (conceptual, commonly accepted in ML engineering literature).
Existing benchmarks rarely calibrate model performance against that of human programmers.
Explicitly stated as a limitation motivating their human-centered evaluation; no counts or proportions of benchmarks are provided in the abstract.
This approach has encountered a "usability ceiling" manifested as the Intent-Execution Gap (i.e., the fundamental disparity between a creator's high-level intent and the stochastic, black-box nature of current single-shot models).
Conceptual critique presented in the paper; no empirical evaluation, quantitative measurement, or sample described in the provided text.
Our results challenge the intuition that declarative framework design guarantees AI-assistability.
Aggregate experimental findings across the evaluated frameworks (including DSPy low score despite being declarative) used to argue against the intuition that declarative design alone ensures AI-assistability.
DSPy -- the most declarative framework by design -- scores lowest (0.07), as its novel abstractions are insufficiently represented in AI training data.
Empirical result: DSPy reported AI-assistability score = 0.07. The causal explanation (insufficient representation in AI training data) is an interpretation provided by the authors.
Coordinated renaming is a frequent yet challenging task.
Supported by authors' reported formative study and developer survey (paper states study of 609K commits in 100 OSS projects and survey of 205 developers).
Existing systems often rely on a single agent to handle the entire workflow—interpreting issues, navigating large codebases, and implementing fixes—within one reasoning chain, and such monolithic designs force the model to retain irrelevant context, leading to spurious correlations and poor generalization.
Authors' analysis/argument about the architecture of existing single-agent systems and its drawbacks; described qualitatively in the paper (no empirical quantification provided in the excerpt).
Large language models (LLMs) have shown strong reasoning and coding capabilities, yet they struggle to generalize to real-world software engineering (SWE) problems that are long-horizon and out of distribution.
Statement in paper framing the problem; based on authors' characterization of prior LLM performance on complex/out-of-distribution SWE tasks (no specific dataset or quantitative result provided in the excerpt).
In the domain of TCAD simulation, the scarcity of open-source resources hinders language models from generating valid TCAD code.
Asserted observation in the paper (motivation); no quantitative study or sample size reported in the abstract.
The link from supportive HR practices to performance weakens in a murky middle where algorithmic oversight is present yet hard to interpret.
Reported moderated mediation results from survey of 464 gig workers using Double Machine Learning; described as a nonmonotonic pattern in the abstract.