Evidence (8974 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
10085 claims
Filter claims →
Productivity
8974 claims
Filtered →
Governance
8062 claims
Filter claims →
Human-AI Collaboration
7749 claims
Filter claims →
Org Design
5057 claims
Filter claims →
Innovation
4896 claims
Filter claims →
Labor Markets
4088 claims
Filter claims →
Skills & Training
3372 claims
Filter claims →
Inequality
2377 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 882 | 244 | 117 | 1097 | 2424 |
| Governance & Regulation | 1010 | 469 | 229 | 135 | 1875 |
| Organizational Efficiency | 977 | 235 | 149 | 90 | 1462 |
| Technology Adoption Rate | 781 | 299 | 143 | 128 | 1362 |
| Research Productivity | 506 | 155 | 74 | 363 | 1110 |
| Output Quality | 555 | 219 | 71 | 70 | 915 |
| Decision Quality | 395 | 200 | 95 | 54 | 751 |
| Firm Productivity | 523 | 67 | 101 | 27 | 724 |
| AI Safety & Ethics | 262 | 309 | 75 | 36 | 688 |
| Market Structure | 195 | 201 | 135 | 30 | 566 |
| Task Allocation | 248 | 77 | 96 | 38 | 464 |
| Innovation Output | 300 | 34 | 55 | 20 | 411 |
| Skill Acquisition | 207 | 75 | 65 | 21 | 368 |
| Employment Level | 138 | 67 | 119 | 24 | 350 |
| Fiscal & Macroeconomic | 156 | 80 | 53 | 33 | 329 |
| Task Completion Time | 211 | 38 | 13 | 16 | 280 |
| Firm Revenue | 183 | 52 | 29 | 5 | 270 |
| Consumer Welfare | 131 | 77 | 48 | 13 | 269 |
| Inequality Measures | 50 | 141 | 54 | 9 | 254 |
| Worker Satisfaction | 104 | 85 | 25 | 13 | 227 |
| Error Rate | 87 | 112 | 11 | 5 | 215 |
| Automation Exposure | 69 | 69 | 37 | 20 | 198 |
| Wages & Compensation | 102 | 49 | 31 | 11 | 193 |
| Team Performance | 115 | 30 | 30 | 11 | 187 |
| Regulatory Compliance | 88 | 74 | 17 | 7 | 186 |
| Training Effectiveness | 109 | 22 | 14 | 21 | 168 |
| Developer Productivity | 116 | 21 | 15 | 8 | 161 |
| Job Displacement | 12 | 92 | 26 | 1 | 131 |
| Hiring & Recruitment | 57 | 12 | 9 | 5 | 83 |
| Skill Obsolescence | 6 | 59 | 10 | 2 | 77 |
| Social Protection | 43 | 17 | 8 | 2 | 70 |
| Creative Output | 35 | 21 | 9 | 4 | 70 |
| Labor Share of Income | 18 | 23 | 17 | 1 | 59 |
| Worker Turnover | 15 | 16 | — | 4 | 35 |
| Industry | — | — | — | 1 | 1 |
Productivity
Remove filter
AI-only planning minimizes time and cost.
Controlled, three-condition experiment (AI-only, human-only, hybrid) conducted on a live client deliverable at a mid-sized digital agency; quantitative metrics included time and cost measures (reported alongside estimation accuracy, rework rates, and scope change recovery time).
The bounded-autonomy architecture is a practical, deployed approach for making imperfect language models operationally useful in enterprise systems.
Deployment and reported performance in the described multi-tenant enterprise application evaluation (completion rates, safety interceptions, speedups); the paper synthesizes these empirical results to support the practical claim.
The enterprise application remains the source of truth for business logic and authorization, while the orchestration engine operates over an explicit published actions manifest.
Architectural proposal and implementation details described in the paper; asserted as part of the bounded-autonomy design deployed in the enterprise application.
Several safety properties are structurally enforced by code and intercepted all targeted violations regardless of model output.
Design and deployment of bounded-autonomy architecture with typed action contracts, permission-aware capability exposure, scoped context, validation before side effects, and consumer-side execution boundaries; empirical claim that these code-enforced properties intercepted targeted violations during evaluation.
Both AI conditions delivered 13–18x speedup over manual operation.
Timing/performance comparison across the three experimental conditions (manual operation, unconstrained AI, full bounded autonomy) within the deployed evaluation; reported speedup range 13–18x relative to manual operation.
The bounded-autonomy system completed 23 of 25 tasks with zero unsafe executions.
Evaluation in a deployed multi-tenant enterprise application across 25 scenario trials spanning seven failure families; comparison across three conditions (manual, unconstrained AI with safety layers disabled, full bounded autonomy).
Successful AI implementation in auditing requires an integrated framework that aligns technological readiness, auditor acceptance, and innovation diffusion to sustainably improve audit quality in Indonesia.
Authors' conclusion and recommendation derived from thematic synthesis of reviewed literature and comparative findings.
Comparative analysis indicates Indonesia remains at the early majority stage of AI adoption in auditing.
Authors' comparative synthesis of the reviewed literature and country-specific discussion classifying Indonesia's adoption stage as early majority.
Comparative analysis indicates global audit firms are positioned at the innovators and early adopters’ stage of AI adoption.
Authors' comparative synthesis of the reviewed literature classifying global audit firms' diffusion stage (innovation adoption framework) based on patterns in the articles.
AI implementation has been shown to significantly enhance audit efficiency, accuracy, and overall audit quality.
Synthesis of findings across the reviewed articles (thematic analysis) reporting positive effects of AI on efficiency, accuracy, and audit quality.
Global auditing practices increasingly utilize machine learning, natural language processing, and robotic process automation to support risk-based auditing, fraud detection, and continuous auditing.
Thematic analysis of the 15 selected journal articles identifying dominant AI techniques (ML, NLP, RPA) and common use cases (risk-based auditing, fraud detection, continuous auditing).
Overall, GAI provides a principled and scalable approach to integrating AI-generated information.
Summary claim in the abstract based on the combination of the theoretical properties and empirical results reported in the paper.
Across applications, GAI improves confidence interval coverage without inflating width.
Empirical claim reported across the multiple application studies in the paper (abstract states CI coverage improvement while maintaining or not inflating width); details in main text/appendix presumably contain the quantitative analysis.
In health insurance choice, GAI cuts labeling requirements by over 90% while maintaining decision accuracy.
Reported empirical result from the paper's health insurance choice experiment; abstract gives the >90% reduction claim but does not include sample size or exact metrics in the abstract.
In retail pricing, where all methods access the same auxiliary inputs, GAI consistently outperforms alternative estimators, highlighting the value of its construction rather than differences in information.
Empirical experiment in a retail pricing application comparing multiple estimators given identical auxiliary inputs; stated as consistent outperformance in the abstract (no numerical effect sizes or sample sizes provided there).
In conjoint analysis with weak auxiliary signals, GAI reduces estimation error by about 50% and lowers human labeling requirements by over 75%.
Reported empirical result from the paper's conjoint analysis experiment(s); exact sample size and experimental details are not stated in the abstract.
Empirically, GAI outperforms benchmarks across diverse settings.
Empirical experiments reported across multiple application settings (conjoint analysis, retail pricing, health insurance choice) comparing GAI to alternative estimators/benchmarks.
The authors establish asymptotic normality for the GAI estimator and show a 'safe default' property: relative to human-data-only estimators, GAI weakly improves estimation efficiency under arbitrary auxiliary signals and yields strict gains whenever the auxiliary information is predictive.
The paper claims formal theoretical results (asymptotic normality and efficiency comparisons) — supported by analytic derivations/proofs in the manuscript as referenced in the abstract.
GAI uses an orthogonal moment construction that enables consistent estimation and valid inference with flexible, nonparametric relationship between LLM-generated outputs and human labels.
The paper presents a methodological proposal (Generative Augmented Inference) and states theoretical properties (orthogonal moment construction, consistency, valid inference) — supported by formal asymptotic analysis/proofs in the paper (the abstract references establishing asymptotic normality).
Clear specifications, explicit governance, and ongoing human-AI collaboration are critical for successful scaling of regression automation.
Conclusions and recommendations derived from the case study's lessons and mixed-method evaluation.
The Copilot achieves 30-50% code reuse when generating candidate test scripts.
Quantitative result reported in the paper's evaluation (stated 30-50% code reuse in the abstract/summary).
Mixed-method evaluation shows the AI accelerates script authoring and increases throughput.
Empirical claim based on the paper's mixed-method evaluation (qualitative and quantitative data reported in the case study); specific sample sizes not provided in the summary.
Automated regression testing is essential for maintaining rapid, high-quality delivery in Agile and Scrum organizations.
Introductory/position statement in the paper; general premise motivating the case study (no specific empirical test reported).
AIBuildAI ranks first on MLE-Bench with a medal rate of 63.1%, outperforming all existing baseline methods and matching the capability of highly experienced AI engineers.
Empirical evaluation on MLE-Bench reported in the paper (benchmark ranking, metric = medal rate).
AIBuildAI adopts a hierarchical agent architecture in which a manager agent coordinates three specialized sub-agents: a designer for modeling strategy, a coder for implementation and debugging, and a tuner for training and performance optimization; each sub-agent is itself an LLM-based agent capable of multi-step reasoning and tool use, enabling end-to-end automation of the AI model development process that goes beyond the scope of existing AutoML approaches.
System architecture description in the paper (methods/architecture section).
We introduce AIBuildAI, an AI agent that automatically builds AI models from a task description and training data.
Methodological contribution: system design and implementation described in the paper (introduction/methods).
This tension reveals a pattern we call 'bounded delegation': developers wanted AI to absorb the assembly work surrounding their craft, never the craft itself.
Interpretive result from the paper's qualitative thematic analysis of survey responses (n=860), labeled by the authors as the 'bounded delegation' pattern.
Developers wanted systems enforcing explicit authority scoping, provenance, uncertainty signaling, and least-privilege access throughout.
Reported constraints and desiderata from the thematic analysis of survey responses (n=860).
Developers wanted systems that embed quality signals earlier in their workflow to keep pace with accelerating code generation.
Thematic findings from the paper's human-in-the-loop, multi-model council-based analysis of survey responses (n=860).
Using a human-in-the-loop, multi-model council-based thematic analysis, we identify 22 AI systems that developers want built across five task categories.
Qualitative analysis method described in the paper applied to the survey responses (n=860); result reported as identification of 22 desired AI systems organized into five categories.
The results demonstrate the importance of considering interacting systems of AI agents when doing both capabilities and safety research.
Authors' interpretation/generalization based on experimental findings comparing multi-agent organizations and single agents across tasks and settings.
BTB enables automated evaluation of any LLM or agent, scoring deliverables against 100+ rubric criteria defined by veteran investment bankers to capture stakeholder utility.
Design claim in abstract describing the benchmark's automated scoring system and rubric size (100+ criteria) defined by expert bankers.
For reproducibility all our data and code are provided at https://github.com/scaleapi/scipredict
Explicit reproducibility statement and URL provided in the paper.
SciPredict addresses two critical questions: (a) can LLMs predict the outcome of scientific experiments with sufficient accuracy? and (b) can such predictions be reliably used in the scientific research process?
Statement of research goals and scope in the paper introducing the SciPredict benchmark and accompanying evaluations.
Human experts demonstrate strong calibration: their accuracy increases from ≈5% to ≈80% as they deem outcomes more predictable without conducting the experiment.
Reported stratified accuracy of human experts on SciPredict tasks by self-reported predictability judgments; accuracy rises from ≈5% (when judged not predictable) to ≈80% (when judged predictable).
We introduce SciPredict, a benchmark comprising 405 tasks derived from recent empirical studies in 33 specialized sub-fields of physics, biology, and chemistry.
Construction of the SciPredict benchmark described in the paper; explicitly reports 405 tasks and 33 sub-fields.
The paper documents 14 deliberate conservative assumptions — including frozen base GDP, no AI-on-AI compounding, a permanent friction floor, and conservative capture rates — all of which directionally understate the benefit.
Paper lists 14 conservative modeling assumptions and claims they bias results downward (i.e., understate potential benefits).
Even excluding demand expansion and robotics layers entirely, the direct productivity contribution alone reaches approximately $940 billion per year by 2036.
Model output reported in the paper when removing demand expansion and robotics layers.
In all four scenarios, cumulative net GDP exceeds cumulative AI infrastructure investment before 2036, with the base case achieving payback in 2033.
Model financial calculation comparing cumulative net GDP uplift to cumulative AI infrastructure investment across scenarios; explicit payback year reported for base case.
The base-case scenario yields approximately $1,057 billion in net annual GDP uplift by 2036, equivalent to 3.6 percent of 2024 GDP; the bear case produces $796 billion, the bull case $1,368 billion, and an agentic scenario produces $2,521 billion.
Model scenario outputs presented in the paper (four scenarios differentiated by capture rate and friction assumptions).
Sector-specific productivity gain percentages are anchored to published evidence, including a randomized controlled trial of GitHub Copilot (Kalliamvakou et al., 2023), JPMorgan CEO disclosures, and Cognizant's New Work New World 2026 research.
Paper states productivity percentages are anchored to published evidence and specifically cites Kalliamvakou et al. (2023) RCT, JPMorgan CEO disclosures, and Cognizant (2026).
Organisations should invest in customisation capabilities for AI recruitment tools, implement comprehensive change management strategies, and maintain robust post-hire evaluation procedures.
Authors' recommendations derived from thematic findings and participant perspectives across two firms (qualitative synthesis of n = 22 interviews).
AI functioned optimally as an augmentative technology rather than as a replacement for human decision-makers in recruitment.
Findings: participants across the two case firms described AI being most effective when augmenting human judgment rather than replacing it (interviews n = 22).
AI significantly enhanced efficiency through process standardisation and automation.
Findings based on participant accounts in thematic analysis (interviews n = 22) describing process optimisation and automation benefits.
The Principle of Maximum Heterogeneity reveals a convergence of complex phenomena across fields onto simple underlying design principles with important predictive value for future distributed production systems.
Synthesis claim in the paper arguing cross-field convergence and predictive value based on the theoretical model and conceptual examples; no empirical validation or forecasting trials reported.
The principles derived (including the Principle of Maximum Heterogeneity) can be used as a blueprint for constructing ideal distributed production systems; demonstrated by suggesting specific redesigns for compute systems executing large-scale AI.
Paper includes suggested redesigns for compute systems as demonstrations of the blueprint; these are proposed designs/illustrative applications rather than empirically validated interventions or trials.
The Principle of Maximum Heterogeneity applies recursively across all layers of nested production systems.
Theoretical claim within the paper arguing recursive applicability across nested system layers (e.g., neurons, firms, ecosystems); supported by conceptual reasoning and model exposition rather than empirical multi-layer tests.
The communication topology determines the spatial scale over which heterogeneity spreads in distributed production systems.
Model-based theoretical argument in the paper linking topology to the spatial scale of heterogeneity; illustrated conceptually and via examples but not via empirical sample testing.
Principle of Maximum Heterogeneity: any distributed production system optimising for performance will converge on an increasingly heterogeneous configuration.
Statement of a derived principle from the paper's model (theoretical derivation/argument); demonstration via model reasoning and examples rather than empirical testing; no sample size reported.
A small set of underlying laws generates the complex dynamics observed across fields (biology, economics, neuroscience, computing).
Theoretical argument and synthesis across disciplines within the paper; no empirical or experimental sample size reported.