Evidence (1624 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
20058 claims
Filter claims →
Productivity
17184 claims
Filter claims →
Governance
16099 claims
Filter claims →
Human-AI Collaboration
16034 claims
Filter claims →
Innovation
10501 claims
Filter claims →
Org Design
10496 claims
Filter claims →
Labor Markets
6444 claims
Filter claims →
Skills & Training
5385 claims
Filter claims →
Inequality
4148 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 1820 | 479 | 278 | 1820 | 4588 |
| Organizational Efficiency | 2711 | 616 | 401 | 173 | 3922 |
| Governance & Regulation | 2075 | 886 | 459 | 246 | 3714 |
| Technology Adoption Rate | 1467 | 530 | 258 | 206 | 2488 |
| Decision Quality | 1281 | 496 | 289 | 152 | 2228 |
| Output Quality | 1227 | 447 | 207 | 138 | 2025 |
| AI Safety & Ethics | 634 | 754 | 207 | 83 | 1688 |
| Research Productivity | 826 | 241 | 114 | 422 | 1624 |
| Firm Productivity | 1052 | 154 | 163 | 66 | 1441 |
| Task Allocation | 685 | 211 | 331 | 99 | 1335 |
| Market Structure | 433 | 423 | 242 | 46 | 1150 |
| Innovation Output | 639 | 91 | 105 | 34 | 871 |
| Task Completion Time | 476 | 113 | 43 | 36 | 672 |
| Firm Revenue | 445 | 126 | 58 | 25 | 656 |
| Skill Acquisition | 364 | 119 | 109 | 34 | 626 |
| Consumer Welfare | 288 | 167 | 104 | 31 | 592 |
| Employment Level | 214 | 140 | 174 | 50 | 582 |
| Error Rate | 230 | 251 | 35 | 16 | 535 |
| Fiscal & Macroeconomic | 268 | 136 | 71 | 50 | 532 |
| Inequality Measures | 100 | 307 | 96 | 12 | 515 |
| Worker Satisfaction | 221 | 173 | 60 | 30 | 484 |
| Automation Exposure | 155 | 138 | 65 | 36 | 398 |
| Regulatory Compliance | 171 | 120 | 30 | 13 | 335 |
| Developer Productivity | 222 | 58 | 27 | 13 | 321 |
| Team Performance | 188 | 56 | 50 | 24 | 320 |
| Wages & Compensation | 146 | 104 | 46 | 16 | 312 |
| Training Effectiveness | 207 | 41 | 21 | 26 | 298 |
| Job Displacement | 23 | 153 | 52 | 4 | 232 |
| Hiring & Recruitment | 102 | 57 | 30 | 11 | 202 |
| Skill Obsolescence | 16 | 102 | 24 | 6 | 148 |
| Creative Output | 71 | 42 | 23 | 6 | 143 |
| Social Protection | 57 | 30 | 11 | 3 | 101 |
| Labor Share of Income | 29 | 42 | 24 | 2 | 97 |
| Worker Turnover | 43 | 29 | 6 | 4 | 82 |
| Industry | — | — | — | 1 | 1 |
The AI Publication Footprint is an absolute cumulative Scopus publication count and therefore serves as a proxy for knowledge-production capacity rather than a normalized measure of research intensity or efficiency.
Measurement definition and methodological caveat; publication counts were not normalized per capita or per researcher.
The Station outperformed AlphaEvolve on three of the seven AlphaEvolve problems that did not produce results novel relative to the prior literature, matched it on two, and underperformed it on two.
Portfolio-level comparison across all 12 evaluated problems, categorized by relative performance.
GPU count showed more consistent positive associations with citation and award outcomes than newer hardware generation; GPU count was positively associated with awards, whereas aggregate capability and hardware generation showed no comparably robust award evidence.
Complementary citation analyses and award models, including linear-probability and Firth models; the award analysis used 5,357 papers.
Reported GPU capability is neither necessary nor sufficient for high citation impact: 85.5% of high-capability papers were not highly cited, and most highly cited papers were outside the high-capability group.
Overlap analysis between the top 20% of papers by reported GPU capability and the top 10% by citations within publication-year-by-venue groups.
The annual top 20% of papers ranked by reported GPU capability accounted for 83.9%–89.9% of reported GPU capability during 2020–2023, but only 27%–32% of citations and 20%–33% of paper awards.
Annual concentration analysis comparing the share of total reported GPU capability, citations, and awards attributable to the top 20% of GPU-quantifiable papers.
Research-topic diversification and collaboration with intermediary institutions help explain why sanctions reduce publication quantity while increasing publication quality.
Mediation-style analyses link changes in research topics and coauthorship networks to the observed quantity and quality outcomes, with semi-structured interviews providing contextual validation.
Randomization does not reliably improve performance across different problem-solving contexts; its usefulness remains restricted to decomposable problems.
Simulation of a randomized AI recommendation strategy that samples from top-performing solutions rather than returning a single best recommendation, compared across problem structures and use rates.
The agent harness can substantially change the capability expressed by the same underlying model.
Cross-harness comparisons: MiMo V2.5 Pro scored 16.17 with MiMo Code versus 23.25 with Claude Code, and Kimi K2.7 scored 19.72 with Kimi Code versus 27.34 with Claude Code.
As AI makes proof generation and verification cheaper, mathematical value is shifting toward forms of mathematical meaning-making that current systems cannot yet perform.
Argument based on Tao's account of increasing proof abundance and the paper's taxonomy of concept and conjecture formation; no quantitative estimate of the shift is provided.
The held-out-paper evaluation protocol can test whether the researcher context carries researcher-specific signal, but it cannot establish that a recommended alternative research direction is valuable.
The paper proposes holding out a paper and measuring fidelity and contrast, then explicitly limits what this protocol can validate because unpursued alternatives lack observed ground truth.
The multi-agent pipeline's advantage should not be interpreted as an isolated causal effect of collaboration, because it also selects the best five hypotheses from 20 candidates rather than retaining all five hypotheses from one model, and the conditions use different judge panels.
The paper explicitly identifies inference-time scaling and different evaluation panels as confounds and characterizes the comparison as an observed association.
The agentic RPM reaches or exceeds the performance of the No-RPM baseline, although its early progress is slower because it runs small-scale proxy experiments at each step.
Comparison of performance trajectories over compute time in the end-to-end AIRS-Bench evaluation; the paper attributes the slower early trajectory to per-step sandbox experiments.
AI adoption may increase dispersion in publication or research-output rates across researchers and fields rather than producing uniform productivity gains.
Derived empirical prediction from the theoretical model of plan-relative complementarity and heterogeneous workflows; no direct empirical estimate is reported.
The productive effects of AI in scholarly production are plan-relative and heterogeneous: AI produces the largest gains when it fills specific gaps in a researcher's workflow and when the researcher has strong complementary human capital.
Logical analysis of complementarity and substitution in plan space, supported by conceptual examples and thought experiments.
Models with similar final outcomes can have substantially different execution and feedback-control capabilities.
GPT-5.5 and Gemini-3.1-Pro had similar outcome and Solution Framing scores but different Execution and Feedback Control scores.
Lower-ranked models can reach competitive solutions but do so less consistently across repeated runs.
The paper compares average and best scores across three independent rollouts per model-task pair and notes that Kimi's best@3 is close to higher-ranked models while its avg@3 is substantially lower.
Model differences are larger in typical performance than in peak observed performance: the highest-to-lowest gap is 0.237 for avg@3 versus 0.122 for best@3.
Comparison of avg@3 and best@3 across seven models, each evaluated with three independent rollouts on all 36 tasks.
Current automated research agents operate more like engineering optimizers than fully autonomous researchers.
Systematic evaluation of seven frontier models across 36 long-horizon AI R&D tasks, including process metrics, experience-reuse comparisons, and novelty review.
The review finds that the existing literature is fragmented and calls for causal and microdata studies, particularly on post-2020 developments and cross-country regulatory differences.
Systematic review of 660 documents and the authors' methodological implications regarding limitations of descriptive bibliometric evidence.
The literature exhibits geographic heterogeneity in research emphasis and leadership despite expanding international participation and collaboration.
Geographic distribution analysis combined with author and institutional collaboration-network analysis.
The intellectual structure of virtual influencer research is fragmented, with influence distributed across multiple research teams.
Analysis of publications, citations, and co-authorships in the 116-study bibliometric dataset.
At a fixed flagging rate of p = 0.2%, the top-α and bottom-α flagging rules differed in their estimated productivity changes by 15.3%, with the top-α rule above and the bottom-α rule below the random baseline.
Rank-based comparison of top-(100p)% α papers, bottom-(100p)% α papers, and random Bernoulli(p) flags, with all rules flagging the same fraction of papers.
Open-weight models run locally offer greater control over randomness, but reproducibility still depends on the complete hardware and software stack.
The paper contrasts local execution with proprietary APIs and discusses variation from processors, drivers, numerical libraries, precision settings, batching, and model-serving environments.
More user-centric studies reveal rich contextual detail, yet are fragmented, small-scale and loosely coupled to the mechanisms that shape model behavior.
Author synthesis/critique of user-centric research approaches; stated in the paper but no empirical meta-analysis or sample sizes provided in the excerpt.
Perdomo et al. 2020 formalized this dynamic in their work on performative prediction.
Citation of Perdomo et al. 2020 as the formalization reference (literature claim).
The paper introduces the concept of 'vibe researching' as the AI-era parallel to 'vibe coding.'
Conceptual contribution and definitional introduction in the paper (term coined and elaborated by the author); no empirical validation.
Heuristic approaches (prompt engineering, fine-tuning, repair strategies) are useful for many exploratory tasks but lack the formal statistical guarantees typically required for confirmatory research.
Argument made in the abstract contrasting heuristic approaches' utility for exploratory work versus their lack of formal guarantees; no empirical validation reported in the abstract.
Two strategies are contrasted for obtaining valid estimates of causal effects when using LLMs as synthetic participants: (1) heuristic approaches (prompt engineering, model fine-tuning, repair strategies) and (2) statistical calibration combining auxiliary human data with statistical adjustments.
Paper's methodological framing described in the abstract; classification/contrast of methodological approaches rather than empirical evaluation.
Recognition of the 'staggered adoption' problem has shifted the focus in recent years away from inference towards consistent estimation of treatment effects.
Descriptive claim in the paper about trends in the literature; references to recent methodological work (e.g., Callaway and Sant'Anna (2021)).
The paper concludes by assessing interconnected dimensions (of post-labor futures) to establish a basis for further exploration of potentially transformative economic shifts.
Stated conclusion and research agenda in the review; synthesis of multiple thematic dimensions (qualitative synthesis, no quantitative sample).
The goal of the note is to highlight the fragility of existing forecasts of exponential growth in AI capabilities rather than to establish a rigorous forecast of our own.
Author statement of intent / framing in the paper (method: argumentative / critical appraisal of forecasting methods).
We propose a more complex model that decomposes AI capabilities into base and reasoning capabilities, each exhibiting individual rates of improvement.
Presentation of a theoretical / mathematical model in the note that decomposes capabilities into two components (model specification provided in the paper). This is a modeling/theoretical claim rather than an empirical estimate; sample size not applicable.
The literature on AI-enabled finance transformation is fragmented across FP&A, record-to-report, reporting, audit/assurance, and governance, producing heterogeneous findings.
Systematic literature review (SLR) of peer-reviewed articles with database searches and snowballing; studies mapped across functional silos in the corporate finance function (as reported in the paper).
The magnitude of the increase in paper production among LLM adopters varies substantially by scientific field and author background (range reported 23.7–89.3%).
Stratified analyses across fields and author background in the 2.1M preprint dataset showing heterogeneous percentage increases in production depending on field and author characteristics.
Recovering ground truth causal effects is feasible -- but only with careful modeling choices.
Empirical comparison in the paper between experimental (ground-truth) treatment effects and estimates from observational causal ML approaches applied to the opt-in sample; the provided text indicates such comparisons were made but does not report sample size or exact methods/results.
Literature addressing this gap spans three disconnected streams: (1) AI governance frameworks relying on subjective risk assessment, (2) technical ML bias measurement focused on algorithmic fairness, and (3) organizational implementation approaches failing to translate governance principles into operational practice.
Authors' literature mapping/synthesis identifying three streams (conceptual literature review).
AI labor-impact predictions are based on multimodal data and labor-force productivity modeling relating job and skill creation to economic conditions and workforce development.
Methodological statement from the systematic review describing how predictions in the literature use multimodal data and productivity models (2024–2025 corpus).
Intrinsically interpretable architectures, notably attention-based transformers, are discussed as approaches for time-series explainability.
Survey discussion of model classes and interpretability properties, with attention-based transformers highlighted as a notable interpretable architecture.
Overall, the LLM era coincides with a broader reorganization of scientific exploration, collaboration, and the division of labor.
Aggregate observational findings from linked PubMed Central and OpenAlex datasets covering 775,323 scientists and 137,120 multi-author papers with CRediT statements, comparing patterns before and after 2022.
Through simulations and an empirical application, the paper illustrates the main advantages and limitations of each approach (CBCMs and FBCMs).
Reported simulation studies and at least one empirical application described in the paper (methods section and results).
Although AI working autonomously achieved a 37% reproduction rate, it could be useful for automated screening when human review is cost-prohibitive.
Interpretation in paper: authors note 37% autonomous reproduction rate as potentially useful for large-scale screening where human review is infeasible; based on empirical results of the experiment.
Empirical claims across the reviewed literature vary in methodological rigor and should be viewed with caution before standardized replication.
Meta-level assessment presented in the review of peer‑reviewed literature (2020–2025); no formal quality-assessment statistics provided in the excerpt.
The literature's vocabulary ("self-refine," "self-reward," "self-play," "self-evolve") conflates fundamentally different ambitions.
Qualitative analysis of terminology across the surveyed arXiv papers (2024-2026) reported in the paper's survey and taxonomy section.
The CAD is formalized with a probabilistic model grounded in the fan effect literature in cognitive psychology.
Paper reports a formal probabilistic model drawing on the fan effect literature; model described as the formalization of CAD.
Across four high-stakes domains, assigning different personas is sufficient for AI agents to report divergent, often opposing, conclusions from the same data and question, with findings systematically aligned with those beliefs.
Experimental manipulation across four domains where AI agents were assigned different personas and produced analyses from the same data/question; comparison of resulting conclusions showing divergence and alignment with persona beliefs.
Important gaps remain in the literature and warrant further research.
Paper's abstract statement that the review identifies important gaps that warrant further research (based on review of 194 articles).
The existing literature on AI and economic development remains fragmented, with limited integration across development dimensions.
Conclusion drawn in the abstract from the systematic review of 194 peer-reviewed articles noting fragmentation and limited cross-dimension integration.
The research uses a sequential multi-phase design combining experiments and qualitative fieldwork.
Stated methodology in the abstract (methodological claim about study design). No sample sizes or procedural details provided in the excerpt.
Evidence on the productivity, risk, and resilience implications of AI adoption remains fragmented and dispersed across different fields of research.
Author's assessment of the literature based on the systematic review (PRISMA) of 68 empirical studies published 2015–2025.
Across compression sweeps, real factor archives, and LLM-SRBench tasks, hybrid gains concentrate in weakly represented but target-bearing directions and vanish as the hypothesis space approaches full rank.
Empirical claim based on experiments over compression sweeps, analyses of real factor archives (A-share factor discovery), and LLM-SRBench tasks; no numerical sample sizes or effect magnitudes provided in the abstract.