Evidence (8974 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
10085 claims
Filter claims →
Productivity
8974 claims
Filtered →
Governance
8062 claims
Filter claims →
Human-AI Collaboration
7749 claims
Filter claims →
Org Design
5057 claims
Filter claims →
Innovation
4896 claims
Filter claims →
Labor Markets
4088 claims
Filter claims →
Skills & Training
3372 claims
Filter claims →
Inequality
2377 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 882 | 244 | 117 | 1097 | 2424 |
| Governance & Regulation | 1010 | 469 | 229 | 135 | 1875 |
| Organizational Efficiency | 977 | 235 | 149 | 90 | 1462 |
| Technology Adoption Rate | 781 | 299 | 143 | 128 | 1362 |
| Research Productivity | 506 | 155 | 74 | 363 | 1110 |
| Output Quality | 555 | 219 | 71 | 70 | 915 |
| Decision Quality | 395 | 200 | 95 | 54 | 751 |
| Firm Productivity | 523 | 67 | 101 | 27 | 724 |
| AI Safety & Ethics | 262 | 309 | 75 | 36 | 688 |
| Market Structure | 195 | 201 | 135 | 30 | 566 |
| Task Allocation | 248 | 77 | 96 | 38 | 464 |
| Innovation Output | 300 | 34 | 55 | 20 | 411 |
| Skill Acquisition | 207 | 75 | 65 | 21 | 368 |
| Employment Level | 138 | 67 | 119 | 24 | 350 |
| Fiscal & Macroeconomic | 156 | 80 | 53 | 33 | 329 |
| Task Completion Time | 211 | 38 | 13 | 16 | 280 |
| Firm Revenue | 183 | 52 | 29 | 5 | 270 |
| Consumer Welfare | 131 | 77 | 48 | 13 | 269 |
| Inequality Measures | 50 | 141 | 54 | 9 | 254 |
| Worker Satisfaction | 104 | 85 | 25 | 13 | 227 |
| Error Rate | 87 | 112 | 11 | 5 | 215 |
| Automation Exposure | 69 | 69 | 37 | 20 | 198 |
| Wages & Compensation | 102 | 49 | 31 | 11 | 193 |
| Team Performance | 115 | 30 | 30 | 11 | 187 |
| Regulatory Compliance | 88 | 74 | 17 | 7 | 186 |
| Training Effectiveness | 109 | 22 | 14 | 21 | 168 |
| Developer Productivity | 116 | 21 | 15 | 8 | 161 |
| Job Displacement | 12 | 92 | 26 | 1 | 131 |
| Hiring & Recruitment | 57 | 12 | 9 | 5 | 83 |
| Skill Obsolescence | 6 | 59 | 10 | 2 | 77 |
| Social Protection | 43 | 17 | 8 | 2 | 70 |
| Creative Output | 35 | 21 | 9 | 4 | 70 |
| Labor Share of Income | 18 | 23 | 17 | 1 | 59 |
| Worker Turnover | 15 | 16 | — | 4 | 35 |
| Industry | — | — | — | 1 | 1 |
Productivity
Remove filter
The system outperforms AlphaEvolve's reported circle packing solution (n=26).
Direct comparison reported to AlphaEvolve's circle packing solution with sample size notation n=26 provided in the excerpt; implies evaluation over 26 instances or trials.
The system generates CUDA kernels where 87% match or beat PyTorch.
Reported evaluation of generated CUDA kernels against PyTorch implementations; paper states 87% of generated kernels match or outperform PyTorch.
The system finds scheduling algorithms that cut cloud costs by 40%.
Paper reports that its discovered scheduling algorithms reduce cloud costs by 40%; presumably measured by evaluating cost of scheduled workloads before/after optimization.
The system discovers agent architectures that nearly triple Gemini Flash's ARC-AGI accuracy (32.5% to 89.5%).
Reported comparison to Gemini Flash on the ARC-AGI benchmark with explicit accuracy numbers (32.5% baseline to 89.5% after optimization). Method: discovered agent architectures via LLM-based search; benchmark evaluation on ARC-AGI.
A single AI-based optimization system achieves state-of-the-art results across six diverse tasks.
Paper reports experiments applying a single LLM-based optimization system to six diverse tasks and claims SOTA results across them; no further per-task details provided in the excerpt.
In real-world deployment, GUIDE achieves notable gains: +4.10% ad GMV, +1.40% ad clicks, +1.66% ad cost, and +3.52% ad ROI.
Reported quantitative results from large-scale online deployment on Taobao (real-world A/B/test deployment; exact sample size, duration, and statistical significance are not stated in the excerpt).
Results show GUIDE consistently outperforms state-of-the-art baselines across all scenarios.
Aggregate claim summarizing experimental comparisons versus state-of-the-art baselines across the reported evaluations (public datasets, simulations, and live deployment). Specific baselines, metrics, and statistical details are not provided in the excerpt.
GUIDE employs a Decision Transformer (DT) to jointly model historical bidding actions and environmental state transitions, a Q-value module to guide DT exploration via regularization constraints, and an Inverse Dynamics Module (IDM) that leverages DT-predicted future states to infer robust behaviorally consistent actions as a safe policy fallback.
Detailed methodological description of GUIDE components and their intended roles in the paper (architectural claim).
We propose GUIDE (Generative Auto-Bidding with Unified Modeling and Exploration), a framework that synergistically integrates directed exploration with a safe fallback mechanism.
Methodological contribution described in the paper (design/proposal of the GUIDE framework).
We introduce the concept [of twin agents], distinguish it from digital twins, and outline the research questions this new class of agent demands.
Stated contribution of the paper (conceptual development and research agenda); content claim about what the paper contains rather than an empirical finding.
Cognitive forcing functions and related frameworks address overreliance effectively in contexts where there is a clear boundary between the AI and the human decision-maker.
Claim based on literature and frameworks cited or discussed by the authors (asserted effectiveness in boundary-defined contexts); the abstract does not provide empirical evaluation details or sample sizes.
The next role on that list is more personal: you — digital twins of each individual (twin agents) representing their knowledge, perspective, and communicative style to colleagues when they are unavailable.
Proposed argument supported by the authors' early design work in an ongoing project; conceptual proposal rather than reported empirical validation in the abstract.
Agentic AI has taken on the role of assistant, collaborator, and decision-support tool.
Asserted in the paper's framing/introduction; based on synthesis of prior work and the authors' characterization of current agentic-AI deployments (no empirical sample or quantitative data reported in the abstract).
The evaluation harness records full trajectories and computes auditable partial-credit rewards.
System description in the paper specifying that the evaluation harness captures full action trajectories and implements an auditable partial-credit reward computation.
OpenComputer's hard-coded verifiers align more closely with human adjudication than LLM-as-judge evaluation, especially when success depends on fine-grained application state.
Experimental comparison reported in the paper between hard-coded verifiers and LLM-as-judge evaluations, measured against human adjudication (presumably over the benchmark tasks).
OpenComputer integrates four components: (1) app-specific state verifiers that expose structured inspection endpoints over real applications, (2) a self-evolving verification layer that improves verifier reliability using execution-grounded feedback, (3) a task-generation pipeline that synthesizes realistic and machine-checkable desktop tasks, and (4) an evaluation harness that records full trajectories and computes auditable partial-credit rewards.
System design description presented in the paper (architectural claim listing four components and their intended functions).
In its current form, OpenComputer covers 33 desktop applications and 1,000 finalized tasks spanning browsers, office tools, creative software, development environments, file managers, and communication applications.
System description / reported inventory in the paper (explicit counts of covered applications and finalized tasks).
The paper offers a research agenda for more effective human-AI collaboration in software engineering.
Authors' concluding recommendations and agenda presented in the paper (conceptual / prescriptive contribution).
Humans are retained at key decision points in the workflow to preserve judgment, accountability, and team-level understanding.
Authors' design rationale / argument for human-in-the-loop controls within their proposed workflow (conceptual justification).
The proposed framework spans five stages: PR Creation, PR Augmentation, Reviewer Selection, AI-Assisted Code Review, and PR Retrospective.
Authors' explicit description of their framework stages in the paper (conceptual/design content).
We present a vision for an AI-powered code review workflow combining specialized agents with human-controlled quality gates.
Paper authors' proposed conceptual framework / design contribution (framework description rather than empirical validation).
The rise of Artificial Intelligence (AI) coding assistants has increased code production velocity.
Authors' summary statement about observed effects of AI coding assistants; based on prior literature/observations rather than a reported experiment in this paper's abstract.
Compute expansion increases data-centre electricity pressure.
Public institutional data on compute expansion and data-centre electricity demand analyzed with growth indicators (CAGR, relative growth) showing rising electricity demand associated with compute capacity expansion.
Industrial robots represent persistent cyber-physical action capacity (as evidenced by installations and operational stock).
Use of public data on robot installations and operational stock, summarized via stock-flow ratios and related indicators to characterize persistent robotic action capacity.
AI investment signals broad capital allocation.
Public institutional data on AI investment examined with indicators such as growth multipliers, CAGR and concentration ratios to infer capital allocation patterns.
AI adoption is accelerating.
Analysis of public institutional data on AI adoption using growth indicators (relative growth, CAGR, growth multipliers) within a conceptual-empirical quantitative diagnostic design (no causal econometric model).
The paper recommends staged, governance-aware implementation for responsible AI adoption in SMEs.
Policy and practice recommendation from the reviewer's synthesis and conclusions section.
This review extends the resource-based view to AI-enabled capabilities in SMEs.
Conceptual/theoretical contribution described in the paper based on synthesis of literature and interpretation of AI as a firm capability in SMEs.
AI enhances operational efficiency primarily in recruitment and performance analytics.
Synthesis across the 21 included studies in the review identifying recurring application domains (recruitment, performance analytics) and reported efficiency benefits.
Artificial intelligence (AI) is transforming human resource management (HRM) by automating tasks and enabling data-driven decisions.
Statement synthesized from the systematic literature review (PRISMA-based) of global studies on AI applications in HRM included in the paper; no single empirical estimate reported.
AI excels at structured, retrieval-grounded, and tool-mediated tasks.
Paper's synthesized conclusion from cross-stage analysis; appears to be based on qualitative benchmarking and review rather than a specific randomized trial in the excerpt.
Long-horizon agents can execute experiments, draft manuscripts, and simulate critique with minimal human input.
Qualitative claim based on the paper's end-to-end analysis of AI across the research lifecycle (review of developments through April 2026); no specific trials or sample sizes reported in the excerpt.
Fully automated systems can now generate research papers for as little as $15.
Statement in paper's introduction asserting observed market/practice examples and cost estimates; no specific empirical sample or experiment reported in the excerpt.
The results position value realization as the most informative predictive signal in the dataset and provide an interpretable basis for enterprise-level screening and managerial reflection rather than causal inference.
Interpretation based on model importance diagnostics (ai_iot_advantage_share as key predictor) and explicit statement in the paper emphasizing predictive/interpretive use over causal claims.
Nonlinear diagnostics indicate a threshold-like transition in predicted success around the mid-range of advantage attribution and a saturation pattern at higher values.
Partial dependence plots (PDP) and individual conditional expectation (ICE) analyses reported in the paper showing nonlinear relationships between ai_iot_advantage_share and predicted success.
Across model families, ai_iot_advantage_share emerges as the most stable predictor of reported AI/IoT success.
Feature-importance analyses (permutation importance) and cross-model comparison reported in the paper.
Random Forest achieves the strongest out-of-sample predictive performance and reduces absolute errors relative to Elastic Net for most test observations.
Empirical model comparison using out-of-sample evaluation on the survey dataset (n = 1250); error reduction and relative performance reported in results.
The paper compares a regularized linear baseline (Elastic Net) with nonlinear approaches (Decision Tree and Random Forest) under a consistent out-of-sample evaluation framework.
Methods section: model families listed and out-of-sample evaluation protocol described.
The study uses enterprise survey data from Slovakia and the Czech Republic (n = 1250).
Data description provided in the paper indicating source countries and total sample size.
This study develops and evaluates a firm-level predictive framework for the reported AI/IoT success rate, measured on a bounded 0–100 scale.
Methodological description in the paper: development of a predictive framework and definition of the dependent variable (reported AI/IoT success on 0–100).
Shifting the community's default mindset from optimizing models per task to sampling models from learned weight distributions will accelerate toward an era in which AI systems routinely improve or create other AI systems.
Normative/prognostic statement by the authors outlining the paper's intended impact and vision; not supported by empirical data in the abstract.
Adapter-scale and conditional generation are advancing rapidly.
Authors' assessment of the current research trajectory (statement in abstract); implies multiple recent papers showing progress at adapter and conditional scales but no specific quantification in the abstract.
The authors organize existing methods into a five-stage pipeline and survey applications where weight-space generative approaches are already practical.
Descriptive claim about the content and organization of this position paper (methodology and survey); evidence is the paper itself.
High-performing models occupy low-dimensional, highly structured regions of weight space shaped by symmetry, flatness, modularity, and shared subspaces.
Authors' theoretical/empirical contention synthesizing observations from recent work; presented as an explanatory claim in the paper's abstract rather than a specific experimental result.
Recent advances demonstrate that neural weights can be synthesized on demand, often matching fine-tuning performance while reducing adaptation cost by orders of magnitude.
Claim refers to recent empirical work in the literature showing weight-synthesis methods; no specific papers, sample sizes, or quantified studies are cited in the abstract.
Model checkpoints should be treated as a first-class data modality, and generative modeling in weight space should be standardized as a core machine learning primitive.
Normative argument made by the authors in the position paper (proposal/recommendation); not supported by an empirical study in the abstract.
Neural network checkpoints have quietly become a large-scale data resource: millions of trained weight vectors now exist, each encoding task-, domain-, and architecture-specific knowledge.
Statement in the paper's abstract describing the current state of checkpoints; references to public model zoos and industry practice are implied but not enumerated in the abstract.
Casting customer trajectory prediction as a maximum entropy RL problem balances reward maximization with stochasticity to better reflect customers with bounded rationality.
Methodological proposal and conceptual argument in the paper, supported by empirical comparisons that demonstrate more behaviorally realistic trajectories; direct empirical validation referenced but details not included in excerpt.
Reinforcement learning (maximum entropy RL) generated trajectories align more closely with customer behaviour than Travelling Salesman Problem (TSP) and Probabilistic Nearest Neighbours (PNN) heuristics.
Comparison of RL-generated trajectories to TSP and PNN using real-world trajectory data from a convenience store; alignment metrics reported in the paper (specific metrics and sample size not provided in the excerpt).
Deployment of GrowthGR delivered a non-trivial 0.3% gain in overall search GMV.
Reported result from the same production deployment / online A/B testing on Taobao (overall search GMV improvement claimed); no sample size or experimental details provided in the excerpt.