Evidence (672 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
20058 claims
Filter claims →
Productivity
17184 claims
Filter claims →
Governance
16099 claims
Filter claims →
Human-AI Collaboration
16034 claims
Filter claims →
Innovation
10501 claims
Filter claims →
Org Design
10496 claims
Filter claims →
Labor Markets
6444 claims
Filter claims →
Skills & Training
5385 claims
Filter claims →
Inequality
4148 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 1820 | 479 | 278 | 1820 | 4588 |
| Organizational Efficiency | 2711 | 616 | 401 | 173 | 3922 |
| Governance & Regulation | 2075 | 886 | 459 | 246 | 3714 |
| Technology Adoption Rate | 1467 | 530 | 258 | 206 | 2488 |
| Decision Quality | 1281 | 496 | 289 | 152 | 2228 |
| Output Quality | 1227 | 447 | 207 | 138 | 2025 |
| AI Safety & Ethics | 634 | 754 | 207 | 83 | 1688 |
| Research Productivity | 826 | 241 | 114 | 422 | 1624 |
| Firm Productivity | 1052 | 154 | 163 | 66 | 1441 |
| Task Allocation | 685 | 211 | 331 | 99 | 1335 |
| Market Structure | 433 | 423 | 242 | 46 | 1150 |
| Innovation Output | 639 | 91 | 105 | 34 | 871 |
| Task Completion Time | 476 | 113 | 43 | 36 | 672 |
| Firm Revenue | 445 | 126 | 58 | 25 | 656 |
| Skill Acquisition | 364 | 119 | 109 | 34 | 626 |
| Consumer Welfare | 288 | 167 | 104 | 31 | 592 |
| Employment Level | 214 | 140 | 174 | 50 | 582 |
| Error Rate | 230 | 251 | 35 | 16 | 535 |
| Fiscal & Macroeconomic | 268 | 136 | 71 | 50 | 532 |
| Inequality Measures | 100 | 307 | 96 | 12 | 515 |
| Worker Satisfaction | 221 | 173 | 60 | 30 | 484 |
| Automation Exposure | 155 | 138 | 65 | 36 | 398 |
| Regulatory Compliance | 171 | 120 | 30 | 13 | 335 |
| Developer Productivity | 222 | 58 | 27 | 13 | 321 |
| Team Performance | 188 | 56 | 50 | 24 | 320 |
| Wages & Compensation | 146 | 104 | 46 | 16 | 312 |
| Training Effectiveness | 207 | 41 | 21 | 26 | 298 |
| Job Displacement | 23 | 153 | 52 | 4 | 232 |
| Hiring & Recruitment | 102 | 57 | 30 | 11 | 202 |
| Skill Obsolescence | 16 | 102 | 24 | 6 | 148 |
| Creative Output | 71 | 42 | 23 | 6 | 143 |
| Social Protection | 57 | 30 | 11 | 3 | 101 |
| Labor Share of Income | 29 | 42 | 24 | 2 | 97 |
| Worker Turnover | 43 | 29 | 6 | 4 | 82 |
| Industry | — | — | — | 1 | 1 |
An industry survey reported approximately 46% time savings from AI on routine software-development tasks, but less than 10% savings on complex work.
McKinsey industry survey of approximately 4,500 developers.
Mean execution time varied substantially across systems, ranging from 21.8 to 112.8 minutes per task.
Reported execution-time measurements for the evaluated model/scaffold configurations.
Validation work can require substantial expertise and time because fluent AI output may contain substantive errors, yet validation is often treated by the market as undifferentiated work that commands a lower price.
The paper supports the claim with the machine-translation case and cited studies on AI-generated content and overreliance [3, 6, 7, 13, 16]. The paper does not report a sample size or quantitative estimate for this synthesis.
Historical productivity gains have not translated mechanically into equivalent increases in leisure or reductions in work time; institutions influence the conversion of productivity into returned human time.
The paper cites estimates of only about four to five additional hours of U.S. leisure per week over roughly a century, unchanged prime-age market work in the cited measure, and cross-country labor-supply differences associated by Prescott with taxation.
A gross AI time saving does not necessarily become returned or usable human time because institutions may increase workloads and individuals may lack material agency over the time saved.
The paper defines institutional return and material-agency factors and gives a workload-increase example in which gross time savings remain positive while the return fraction can be near zero.
AI can change the human time required per useful outcome, but the sign and magnitude of the change are task-, technology-, and context-dependent.
The paper cites controlled studies showing both time reductions and time increases in specific tasks, including professional writing, customer support, and software development.
With fixed sample sizes, asymptotic experiment conversion is governed by the full spectrum of Rényi divergences, whereas allowing the sample size to depend on realized evidence reduces the relevant constraints to the two directed KL divergence ratios.
The paper contrasts the fixed-sample-size result based on the dominance-ratio result of Mu et al. (2021) with its endogenous-stopping characterization.
MST-5 achieved a slightly lower average weighted deviation than PPO, but its setup time was approximately 60% higher.
Table I reports MST-5 with 1,112 weighted deviation and 50.15 setup time, compared with PPO's 1,144 deviation and 31.71 setup time. The paper explicitly characterizes the setup-time difference as approximately 60%.
DQN achieved the lowest average weighted deviation from scheduled due dates, but it produced substantially more unfinished products than PPO and most other methods.
Table I reports DQN's average weighted deviation as 940, the lowest value, and its unfinished-product count as 144.8, compared with 53.95 for PPO. The text also states that the lower deviation was offset by a considerable increase in unfinished products.
In Era 3, the median PR cycle time was 1.04 days in vLLM and 0.62 days in SGLang, while the 90th-percentile cycle times were 16.8 and 14.3 days, respectively.
Cycle time was calculated from PR creation to merge for the full merged-PR populations in both repositories.
Performance improvements from translation varied substantially by software, with highly optimized C/C++ programs requiring considerably more work to reach speed parity.
The authors describe uneven optimization effort and compare translation behavior across software domains and source languages using qualitative benchmark trends.
Under transparency, the equilibrium algorithm can make prominence convey either good news that deters further search or bad news that encourages further search, depending on the prior probability that the more profitable product is the consumer's best match.
Proposition 2 analytically characterizes three equilibrium algorithm forms across prior beliefs.
Latency optimization is dominated by the model-provider stack and is shaped by task formulation before being affected by preprocessing; the same preprocessing technique may speed up one model-provider stack while slowing down another.
Cross-model and cross-provider latency measurements in the benchmark, including local preprocessing, upload, server processing, and response streaming components.
Skills distilled from non-reasoning trajectories alone were competitive with skills distilled from paired reasoning/non-reasoning trajectories, but the relative performance was domain-dependent.
A source-composition ablation evaluated GPT-5.4-mini on four held-out benchmarks using skills distilled from no-think-only versus paired corpora; each result was based on 3 evaluation seeds.
Experienced open-source developers took 19% longer to complete real tasks when AI assistance was permitted, while estimating afterward that AI assistance had made them 20% faster.
The paper cites a randomized controlled trial by Becker et al. (2025). The supplied text does not report the trial's participant sample size.
Happy buyers have the highest acceptance rate and the lowest rejection rate among the reported buyer-emotion conditions, but 40.82% of their negotiations still reach the maximum number of turns without resolution.
Termination-profile analysis by buyer emotion; deadlocks are negotiations that reach the maximum turn limit.
Emotion conditioning substantially changes negotiation deal rates: angry buyers have an average deal rate of 0.39%, while happy buyers have the highest average deal rate at 28.91%.
Controlled evaluation of six buyer emotions crossed with six seller emotions, two budget conditions, 350 products, three repetitions per product-condition cell, and five language models; results are aggregated across models.
In the paper's controlled testbed, the performance gap associated with myopic planning appears under capability gating, persists for every fixed lookahead horizon up to a chain-length boundary, and disappears in no-gating controls.
The abstract and contribution summary report controlled-testbed comparisons involving gating parameters, fixed lookahead horizons, chain lengths, distractor-build controls, and no-gating controls; no numerical sample size or effect estimate is provided in the supplied text.
Moderate positive temporal steering improves TravelPlanner commonsense constraint performance, while negative steering degrades it; very large positive steering also degrades performance.
The TravelPlanner validation split contains 180 queries covering three difficulty levels and three trip lengths. The primary outcome is Commonsense Constraint Micro Pass Rate, evaluated using the layer-40 temporal direction across steering strengths.
For the two task-level cost case studies, hardening left GPT-5.6 Luna's success rate nearly unchanged on compile-compcert but sharply degraded success on caffe-cifar-10.
Task-level cost sampling with GPT-5.6 Luna using Codex: 100 runs per condition for each of two tasks under control and NIST-derived high; all 200 hardened trajectories were inspected.
Most Terminal-Bench tasks remained solvable under the strictest policy: 82 of 89 tasks had a solvability witness, while 7 tasks were blocked by design.
The authors replayed each task's official reference solution under the strictest policy and, when necessary, authored a policy-compliant reference solution; tasks requiring policy-forbidden actions were classified as blocked by design.
Frontier agent performance increased across all five benchmark groups between 2024 and 2026, but the amount of progress was uneven across domains.
Descriptive frontier analysis using, for each task, the highest score among agents eligible by release quarter, followed by averaging within benchmarks and benchmark groups.
The AMD AITER runtime is required for competitive MLA inference throughput and must be selectively disabled for architectures with incompatible attention head configurations.
Comparative runtime experiments reported in the paper showing MLA throughput improvement with AITER and cases where AITER caused problems on architectures with incompatible head configurations (authors recommend selective disabling).
Pull request description styles are associated with differences in reviewer response timing.
Analysis linking PR description features to reviewer response timing using the AIDev dataset covering PRs generated by five AI coding agents.
Performance, energy use, and time often involve intricate trade-offs under heterogeneous workloads.
Conceptual analysis in the paper describing trade-offs across performance, energy, and latency under varying workloads; no quantitative sample or experiment reported in the excerpt.
The study investigated the impact of text data analytics on audit report lag of listed manufacturing firms in Nigeria using a mixed-method design.
Methodological description provided in the paper's abstract: mixed-method design combining primary firm-level data from external auditors and secondary data from audited annual reports of listed manufacturing firms in Nigeria.
The review found that AI system applicability correlates with occupational task operation completion, wages, employment prospects, and education, influencing business transformation and economic growth.
Main synthesis claim from the systematic literature review of 2024–2025 publications; described as correlations between AI applicability and occupational/economic outcomes.
Learning-curve shape (through its learning speed α) is the primary theoretical determinant of when to stop experimenting; costs determine switching profitability.
Analytical decomposition and theoretical results in the paper linking the learning-curve parameter α to the optimal stopping time and stating role of costs in profitability.
Security Operations Centers (SOCs) face massive, heterogeneous alert streams under minute-level service windows, creating an "Alert Triage Latency Paradox": verbose reasoning chains ensure accuracy and compliance but incur prohibitive latency and token costs, while minimal chains sacrifice transparency and auditability.
Framed as the motivating problem in the paper (abstract). No experimental sample sizes given; presented as conceptual/problem statement supported by the paper's motivating discussion.
Our results show distinct scaling regimes for red- and blue-team tasks.
Empirical results reported in the paper comparing offensive (red-team) and defensive (blue-team) task performance as a function of cost/compute budget.
In coding tasks, low agreeableness leads to large communication shifts that have little effect on milestone completion.
Experimental manipulation of agreeableness in LLMs on structured coding tasks; observed large changes in communication but little change in milestone completion rates. No quantitative effect sizes or sample counts given in the abstract.
Participants' IAT scores were predictive of the time they spent in human-AI collaboration.
Reported predictive relationship between individual IAT scores and measured time spent interacting with/considering resumes during human-AI collaborative screening tasks (likely from regression or correlation analyses); exact statistics and sample size not provided in the excerpt.
Frontier proprietary models achieve near-zero success under GUI-based interaction, whereas COM-based execution yields substantial immediate gains.
Experimental comparison reported in the paper on ComCADBench between GUI-based interaction by proprietary models and COM-based execution (authors report success rates and comparative performance).
The same observation is seen with the amount of changes (e.g., code churn, number of modified files) and with the efforts to merge an agentic PR (e.g., merge time and number of comments).
Reported that pre/post comparison across projects shows mixed/no consistent improvement patterns for code churn, modified files, merge time, and comment counts after instruction-file creation (analysis over 15,549 PRs in 148 projects).
Claude Code completed the pipeline in ~3.4 minutes with silent deviations from the specification, while Codex required ~16 minutes across explicit self-correcting restarts, including an unsolicited performance optimization of the matched filter inner loop.
Reported run-time measurements and qualitative behavior descriptions in paper: timing values (~3.4 min vs ~16 min) and observed behaviors (silent deviations for Claude Code; explicit restarts and an unsolicited optimization by Codex).
AI assistance can generate a deceptive productivity signature: average completion times fall because AI tools typically supply a fast first draft, yet workflow-level performance can deteriorate when a subset of AI errors escapes review and returns as costly downstream rework.
Analytical derivation and discussion based on the paper's queueing model (theoretical/model-based evidence; no empirical sample provided).
Across 78 endpoints, the same model on different endpoints differs in tail latency by an order of magnitude.
Empirical tail-latency measurements across 78 endpoints serving 12 model families.
Provisioned Throughput delivers the lowest latency at low concurrency but saturates its reserved capacity above approximately 20 concurrent users.
Empirical measurements from the instrumented system across concurrency up to 50 users and tier comparisons; the paper reports the observed saturation point near ~20 concurrent users.
Wall-clock time can be reduced to O(√E) through team parallelism, but total human effort remains O(E).
Model-derived result showing parallelism across humans can speed wall-clock completion time while aggregate human effort does not drop asymptotically.
AI-related efficiency improvements can reduce the amount of labour time required while workers remain employed in the short term.
Frontline-worker responses were coded as reduced overtime and partial job substitution, with the paper interpreting these as lower labour requirements without necessarily implying immediate total job loss.
Repair briefs that nominated only unmodifiable files caused a repair window to consume 943 tool turns across 71 minutes, with 48% of the effort spent reading or searching.
Measured tool-turn and duration records from a cross-process failure-repair window; 451 of the 943 turns were read or search operations.
FinVision reduced task completion time by an average of 51 percent for accounting and investment professionals.
A controlled user study involving 48 accounting and investment professionals, as reported in the abstract.
Phase 3 response times were longer in the Low-cost AI condition than in either the High-cost AI or No-AI condition.
Between-condition comparison of Phase 3 response times: Low-cost AI mean=111.81 seconds, High-cost AI mean=86.78 seconds, No-AI mean=92.21 seconds; p<0.01 and p=0.01 for the respective comparisons.
Resilient supply networks should reduce order cancellations, lead-time variance, export interruptions, and time to recovery.
Proposed outcomes and measurement indicators for the international trade continuity layer; no observed results are reported.
In a randomized trial involving experienced open-source developers, access to early-2025 AI tools increased task-completion time by an estimated 19 percent.
The paper reports the findings of Becker, Rush, Barnes, and Rein (2025): 16 experienced developers completed 246 tasks in repositories they knew well.
Distributed and asynchronous work weakens assumptions of co-located, simultaneous labor and creates coordination and latency costs that differ from classical time-per-task metrics.
Conceptual discussion of asynchronous production and coordination, using information-processing and time-cost perspectives.
Any exact simulation of a binary target experiment H using repeated observations from source experiment F must satisfy the statewise information lower bounds I0(H) <= I0(F) E0[N] and I1(H) <= I1(F) E1[N].
Theorem 2.1(a), proved using the likelihood-ratio representation on the stopped sigma-field and conditional Jensen's inequality.
Most AI interactions are brief, typically lasting minutes, whereas evidence-based psychotherapies generally require months of longitudinal engagement.
Conceptual comparison between typical AI interaction patterns and the duration and longitudinal structure of evidence-based psychotherapy.
Models degrade when a task is distributed across several turns rather than stated as a complete single-turn instruction.
The survey cites comparative evidence from prior work examining single-turn versus distributed multi-turn task instructions.
Reducing methodological guidance increased average per-task token consumption and execution time: relative to B1, token use rose by 59% in B2, 25% in B3, and 30% in B4, while execution time rose by 32%, 22%, and 18%, respectively.
Reported average token consumption and execution time per task under B1–B4, shown in Figure 4.