Evidence (2025 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
20058 claims
Filter claims →
Productivity
17184 claims
Filter claims →
Governance
16099 claims
Filter claims →
Human-AI Collaboration
16034 claims
Filter claims →
Innovation
10501 claims
Filter claims →
Org Design
10496 claims
Filter claims →
Labor Markets
6444 claims
Filter claims →
Skills & Training
5385 claims
Filter claims →
Inequality
4148 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 1820 | 479 | 278 | 1820 | 4588 |
| Organizational Efficiency | 2711 | 616 | 401 | 173 | 3922 |
| Governance & Regulation | 2075 | 886 | 459 | 246 | 3714 |
| Technology Adoption Rate | 1467 | 530 | 258 | 206 | 2488 |
| Decision Quality | 1281 | 496 | 289 | 152 | 2228 |
| Output Quality | 1227 | 447 | 207 | 138 | 2025 |
| AI Safety & Ethics | 634 | 754 | 207 | 83 | 1688 |
| Research Productivity | 826 | 241 | 114 | 422 | 1624 |
| Firm Productivity | 1052 | 154 | 163 | 66 | 1441 |
| Task Allocation | 685 | 211 | 331 | 99 | 1335 |
| Market Structure | 433 | 423 | 242 | 46 | 1150 |
| Innovation Output | 639 | 91 | 105 | 34 | 871 |
| Task Completion Time | 476 | 113 | 43 | 36 | 672 |
| Firm Revenue | 445 | 126 | 58 | 25 | 656 |
| Skill Acquisition | 364 | 119 | 109 | 34 | 626 |
| Consumer Welfare | 288 | 167 | 104 | 31 | 592 |
| Employment Level | 214 | 140 | 174 | 50 | 582 |
| Error Rate | 230 | 251 | 35 | 16 | 535 |
| Fiscal & Macroeconomic | 268 | 136 | 71 | 50 | 532 |
| Inequality Measures | 100 | 307 | 96 | 12 | 515 |
| Worker Satisfaction | 221 | 173 | 60 | 30 | 484 |
| Automation Exposure | 155 | 138 | 65 | 36 | 398 |
| Regulatory Compliance | 171 | 120 | 30 | 13 | 335 |
| Developer Productivity | 222 | 58 | 27 | 13 | 321 |
| Team Performance | 188 | 56 | 50 | 24 | 320 |
| Wages & Compensation | 146 | 104 | 46 | 16 | 312 |
| Training Effectiveness | 207 | 41 | 21 | 26 | 298 |
| Job Displacement | 23 | 153 | 52 | 4 | 232 |
| Hiring & Recruitment | 102 | 57 | 30 | 11 | 202 |
| Skill Obsolescence | 16 | 102 | 24 | 6 | 148 |
| Creative Output | 71 | 42 | 23 | 6 | 143 |
| Social Protection | 57 | 30 | 11 | 3 | 101 |
| Labor Share of Income | 29 | 42 | 24 | 2 | 97 |
| Worker Turnover | 43 | 29 | 6 | 4 | 82 |
| Industry | — | — | — | 1 | 1 |
Visibility-enabled metrics such as typing activity, response times, and screen time may improve measurement precision while also encouraging gaming and performative labor that inflate measured productivity without corresponding improvements in output quality.
Conceptual implication for AI economics and productivity measurement; the paper does not report a causal test or quantify the gap between measured productivity and output quality.
Digital transformation has stronger positive associations with substantive carbon disclosure components—carbon-related business practices, carbon governance, and carbon performance—than with the disclosure carrier or reporting format/channel.
Component-level regressions separating substantive content elements from the disclosure-carrier component of the multidimensional CIDQ index.
In the observed Verified comparisons, the treatment-control performance gap was largest at the 20,480-token window and closest to zero at the 262,144-token window.
Three successive Verified comparisons using the same 169 task IDs and fixed 480-second budget; the paper explicitly cautions that the cross-window ordering does not isolate window size from run-era change.
At 128,000 tokens, the primary model retrieved the disclosure for all 12 firms and produced no false retrievals on neutral filings, despite the disclosure having negligible influence on its investment judgment.
Separate retrieval calls were scored against frozen answer keys, with retrieval tested independently from the investment judgment experiment.
For the U.S. FEMA benchmark, PPE outperformed expert baselines on the Socioeconomic and Composite risk indicators, achieving R² of 66.9% versus 61.1%, while performing on par with expert benchmarks across the broader environmental-target suite.
Spatial regression evaluation of 21 FEMA environmental risk scores using approximately 84,000 census-tract observations.
In the in-distribution cat -n regime, Paritok-4B compressed context to 27.8% of its original size and retained 89.3% of uncompressed solve quality.
Evaluation on the same 300 SWE-bench Lite instances using line-numbered cat -n inputs intended to match real agent trajectories.
Across all 300 SWE-bench Lite instances, Paritok-4B compressed agent context to 25.7% of its original size while retaining 86.5% of uncompressed single-shot solve quality.
End-to-end evaluation on all 300 SWE-bench Lite instances comparing compressed and uncompressed context under the paper's single-shot solve harness.
Expert-prompted LLMs achieved the highest accuracy but were more expensive and less consistent than the deterministic PPO policy.
Ablation comparing minimal versus augmented expert prompts and PPO on identical held-out proteins and repeated rollouts; token use, inference cost, and consistency were recorded.
Across 20 BixBench3 tasks containing 138 unique graded artifacts, 13 frontier LLMs achieved overall scores ranging from 0.00 to 0.48.
Benchmark evaluation of 13 models across 20 research-study-scale computational-biology tasks and 1,794 model-artifact evaluations; scores were based on the proportion of artifacts reproduced closely enough to preserve the original biological interpretation.
In GAMEOPT, leading agents often retain requested functionality across six turns, but they do not consistently preserve the core game loop or achieve balanced improvement across gameplay, level design, balance, art, interface, and audio.
The finding comes from 17 six-turn optimization chains totaling 102 requests, evaluated using final-product criteria and regression checks.
In GAMEGEN, agents establish a playable core more reliably than they deliver rich content, robust interfaces, and fully integrated runtime behavior.
The paper summarizes findings from the GAMEGEN evaluation, which combines behavioral rubrics, runtime verification, code inspection, live interaction, and human assessment across 97 generation tasks.
Current coding agents are more reliable at producing playable foundations and implementing explicit requirements than at discovering defects, verifying runtime behavior, and preserving functionality across changes.
This conclusion is drawn from evaluations across GAMEGEN, GAMEFIX, and GAMEOPT, although the supplied text does not report aggregate numerical scores.
Models systematically underrepresent race and religion, which are the strongest real drivers of opinion, while exaggerating the influence of weaker dimensions such as gender and party.
The study compared simulated and human steering magnitudes across 429 dimension-value-by-wave observations.
The study concludes that geospatial transferability depends not only on model design but also on intrinsic geographic differences between source and target regions.
Interpretation of the transfer experiments and regression analysis relating geographic domain shift to transfer performance.
Cross-region transfer performance exhibits substantial spatial heterogeneity and asymmetry across source-target geographic pairs.
Results summarized in the abstract and contribution statement, based on cross-region transfer evaluation and linear mixed-effects regression.
Geospatial transferability varies substantially across target geographic areas.
Cross-region transfer experiments using human mobility generation models trained in one geographic region and evaluated in unseen regions.
In one ablation study, removing the language-model component from several LLM-based forecasters left accuracy unchanged or improved it.
A component ablation study cited by the review, comparing complete LLM-based forecasters with versions lacking the language-model component.
Forecasts that rely on unstructured evidence such as news, filings, or health reports may fail to improve accuracy or calibration over specialized models while requiring substantially more inference compute.
Synthesis of reviewed forecasting-agent studies examining retrieval of unstructured evidence, predictive accuracy, calibration, and computational requirements.
LLM-judge scores varied by approximately 1.5 points on a 0–10 scale across three judge models, although relative ranking within the corpus was largely preserved.
Sensitivity analysis rotating Claude 3.7 Sonnet, Claude 3.5 Haiku, and a 9B in-house judge across the 145-skill corpus.
The default structural gate is permissive: 94.5% of skills cleared the 70-point threshold, but only 48.9% reached 80 points.
Structural scoring of 145 real skills using approximately 50 deterministic rules.
A particular generated sequence is a sample selected under a decoding policy from a learned conditional distribution, so observed output need not represent the model's entire learned distribution.
Technical explanation of greedy decoding, temperature sampling, nucleus/top-p sampling, and the distinction between distributions and samples.
Current AI and large language models have potential to augment psychological assessment and treatment at scale, but they do not yet produce meaningful, sustained clinical change.
Conceptual synthesis of clinical-science literature, AI/LLM capabilities, and ethical and implementation considerations; the paper does not report a primary clinical trial.
The effect of perturbations varies substantially across repositories, with degradation concentrated in a small subset of task instances while other instances are unaffected.
Per-instance degradation analysis compares baseline and perturbed resolve rates for the 54 selected SWE-bench task instances.
Robustness to semantics-preserving perturbations is a joint property of the model, agentic scaffold, and workload rather than a property of the model alone.
The evaluation crosses four models, two scaffolds, and two benchmarks, and observes that scaffold and benchmark changes alter degradation and model rankings.
Model robustness rankings are not consistent across agentic scaffolds: Qwen is among the most robust with mini-SWE agent on SWE-bench Verified but is the most brittle with OpenCode.
The paper compares four models across two scaffolds and two benchmarks and reports scaffold-dependent relative robustness rankings.
GPT-4.1 displays the reverse pattern on counseling: it is a lower-ranked direct solver but the strongest assistant.
Task-level comparison of GPT-4.1's automation and augmentation rankings on the counseling task.
GPT-5-mini ranks first in both the automation and augmentation comparisons, while the unaided GPT-3.5-Turbo worker has the second-best average rank in the augmentation comparison.
Average and task-level model rankings across the two usage regimes, using a fixed GPT-3.5-Turbo worker in augmentation mode.
The model that wins in automation loses in augmentation on five of the seven tasks.
Comparison of the highest-ranked automation and augmentation models at the task level across seven tasks.
The correlation between automation and augmentation rankings varies substantially across tasks, ranging from -0.04 for travel planning to 0.85 for tax preparation.
Task-level comparison of automation and augmentation rankings across the seven benchmark tasks.
Model rankings in automation and augmentation are only modestly correlated, with a model-level rank correlation of about 0.48.
Cross-mode rankings of ten LLMs across seven professional tasks, evaluated using blind pairwise comparisons by four LLM judges and replicated across ten independent runs.
A high grouping score can coexist with imperfect average accuracy: a model may consistently pass some tasks and consistently fail another, producing a tight but partially off-target performance pattern.
A worked example using a six-task Rust suite with five decisive passes and one decisive failure across five runs per task.
The study could not rule out an accuracy gain of up to 4.67 percentage points from explicit high effort.
The upper endpoint of the registered 95% interval for the Sonnet 5 accuracy contrast was +0.0467.
Agent performance at the human frontier is concentrated on GoEmotions: 8 of 12 configurations surpass Human SOTA, whereas CUB200 and Weather each have only one frontier-reaching configuration.
Task-level comparison of the 12 agent configurations per task, consisting of 6 agents evaluated with and without prepared human references.
Current LLM agents reach or surpass human state-of-the-art performance in 10 of 72 experimental configurations, but these successes occur on only 3 of the 6 evaluated tasks.
Evaluation of 6 agents across 6 tasks and 2 reference conditions, yielding 72 configurations; performance was compared with the best collected human result for each task.
ReasoningBank improves benign utility from 0.741 to 0.859, while its aggregate attack success rate decreases from 0.496 to 0.474.
Matched benign and attacked evaluations on the corrected 15-family Banking suite; 135 benign cases and 405 attacked cases per condition across three lineages.
SkillOpt produces both capability gains and regressions: it changes 23 previously wrong instances to correct and 10 previously correct instances to wrong.
Paired transition analysis comparing each evolved instance with the Static run on the identical lineage, family, and variant.
Old upload-earned medals appear more informative in the AI era only when read in isolation; their apparent gains disappear after conditioning on the holder's full credential profile.
Comparisons of badge-count, experience-only, and full-profile regression specifications. Under the badge-count read, slope changes were +0.062 for medals aged 1–2 years and +0.036 for medals aged 2–3 years; the gains vanished under full-profile conditioning and persisted in a balanced panel.
Fresh upload medals lost about one quarter of their predictive slope in the AI era under full-profile conditioning, while fresh code medals did not show a statistically significant decline.
Comparison of pre-AI and AI-era fresh-medal slopes under full-profile conditioning. Fresh upload medals changed by -0.033 from a pre-AI base of 0.122 with P = 0.002; fresh code medals changed by +0.015 with P = 0.12.
Current agents generally make substantial partial progress on realistic professional workflows even when they fail to meet the benchmark's completion threshold.
Across the evaluated models, most average task scores fall between roughly 55 and 75, while success rates remain below one third.
In the authors' layer-locality analysis, the gate projection accounts for 77% of pruning distortion in Qwen3.5, while Qwen2.5 concentrates 67% of distortion on the down projection.
Layer-wise neural-flow graph analysis of pruning-induced signal distortion in two Qwen model families.
For image segmentation, pruning reduces model size by nearly 80% while maintaining approximately constant mean intersection-over-union, whereas default quantization leaves parameter count and MACs unchanged.
Empirical evaluation of compressed segmentation models using pruning and quantization.
The apparent relationship between MMLU and MMLU-Pro scores depends strongly on the chosen scale: for six common systems, the fitted slope is 0.976 under a logit link, 1.043 under a probit link, 1.312 on the linear scale, and 1.984 on the logarithmic scale.
Cross-benchmark linking stress test using six common systems and four score representations.
Across eight armed-model requests, the median realized served total variation was 0.118, while the inherited ℓ∞-based upper bound had a median of 0.996 and was effectively vacuous; Bhattacharyya and sub-Gaussian bounds were 0.159 and 0.230, respectively.
Post hoc comparison of paired exact and compressed logits for eight requests, summarized in Figure 4; the authors explicitly state that this certifies the geometry rather than providing an online witness.
The strongest fixed protocol varied by task: Broadcast was strongest in nine of the ten settings, while PER exceeded Broadcast by 2.7 points for Gemma on SciBench.
Figure 2 reports matched solve/coverage percentages for four protocols across five benchmark conditions and two solver families.
Exact ground-truth skill invocation is neither sufficient nor necessary for successful task execution.
Joint analysis of skill-use parsing and final verifier outcomes during full-pool real execution found that correct skills did not guarantee success and related non-ground-truth skills could still help.
Offline retrieval or selection accuracy does not determine downstream task success in a one-to-one manner.
The study independently evaluated embedding retrieval, explicit agent selection, and full-pool real execution; outputs from the first two procedures were not passed to the execution experiment.
Compared with independent professional estimators, Handoff-H1 produced substantially higher material coverage but lower quantity precision: 86.1% versus 65.5% coverage and 78.8% versus 87.9% Precision@25%.
Direct comparison of Handoff-H1 and independent estimator results under the same benchmark judge and gold standard.
Extraction performance exceeded 90% agreement for document titles and sources, but was weaker for author and reference fields.
Model-extracted metadata fields were compared with verified source records and graded as exact, partial, or absent.
No evaluated LLM-based system led on all four evidence-synthesis tasks.
Comparative evaluation of GPT-5, Claude Sonnet 4, Gemini 2.5 Pro, and NotebookLM across screening, extraction, analysis, and synthesis using a 244-document benchmark and expert reference standards.
Under asymmetric accuracy-improvement costs, an equal-accuracy requirement can substantially lower advantaged-group accuracy while increasing disadvantaged-group accuracy only modestly.
Theoretical firm-design analysis with quadratic development costs and a cost asymmetry in which improving accuracy for the disadvantaged group is more expensive than for the advantaged group.