Evidence (8974 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
10085 claims
Filter claims →
Productivity
8974 claims
Filtered →
Governance
8062 claims
Filter claims →
Human-AI Collaboration
7749 claims
Filter claims →
Org Design
5057 claims
Filter claims →
Innovation
4896 claims
Filter claims →
Labor Markets
4088 claims
Filter claims →
Skills & Training
3372 claims
Filter claims →
Inequality
2377 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 882 | 244 | 117 | 1097 | 2424 |
| Governance & Regulation | 1010 | 469 | 229 | 135 | 1875 |
| Organizational Efficiency | 977 | 235 | 149 | 90 | 1462 |
| Technology Adoption Rate | 781 | 299 | 143 | 128 | 1362 |
| Research Productivity | 506 | 155 | 74 | 363 | 1110 |
| Output Quality | 555 | 219 | 71 | 70 | 915 |
| Decision Quality | 395 | 200 | 95 | 54 | 751 |
| Firm Productivity | 523 | 67 | 101 | 27 | 724 |
| AI Safety & Ethics | 262 | 309 | 75 | 36 | 688 |
| Market Structure | 195 | 201 | 135 | 30 | 566 |
| Task Allocation | 248 | 77 | 96 | 38 | 464 |
| Innovation Output | 300 | 34 | 55 | 20 | 411 |
| Skill Acquisition | 207 | 75 | 65 | 21 | 368 |
| Employment Level | 138 | 67 | 119 | 24 | 350 |
| Fiscal & Macroeconomic | 156 | 80 | 53 | 33 | 329 |
| Task Completion Time | 211 | 38 | 13 | 16 | 280 |
| Firm Revenue | 183 | 52 | 29 | 5 | 270 |
| Consumer Welfare | 131 | 77 | 48 | 13 | 269 |
| Inequality Measures | 50 | 141 | 54 | 9 | 254 |
| Worker Satisfaction | 104 | 85 | 25 | 13 | 227 |
| Error Rate | 87 | 112 | 11 | 5 | 215 |
| Automation Exposure | 69 | 69 | 37 | 20 | 198 |
| Wages & Compensation | 102 | 49 | 31 | 11 | 193 |
| Team Performance | 115 | 30 | 30 | 11 | 187 |
| Regulatory Compliance | 88 | 74 | 17 | 7 | 186 |
| Training Effectiveness | 109 | 22 | 14 | 21 | 168 |
| Developer Productivity | 116 | 21 | 15 | 8 | 161 |
| Job Displacement | 12 | 92 | 26 | 1 | 131 |
| Hiring & Recruitment | 57 | 12 | 9 | 5 | 83 |
| Skill Obsolescence | 6 | 59 | 10 | 2 | 77 |
| Social Protection | 43 | 17 | 8 | 2 | 70 |
| Creative Output | 35 | 21 | 9 | 4 | 70 |
| Labor Share of Income | 18 | 23 | 17 | 1 | 59 |
| Worker Turnover | 15 | 16 | — | 4 | 35 |
| Industry | — | — | — | 1 | 1 |
Productivity
Remove filter
At min-cost, Brick cuts cost 22.15x compared to always using the strongest model.
Empirical evaluation on the 5,504-query benchmark reporting cost reduction at the min-cost operating point.
At a neutral cost-quality profile, Brick achieves 74.11% accuracy at 4.71x lower cost than always using the strongest model.
Empirical evaluation on the 5,504-query benchmark reporting accuracy and relative cost compared to always selecting the strongest model.
On a benchmark of 5,504 queries, Brick at max-quality reaches 76.98% accuracy, beating the best single model (75.02%) and all tested routers.
Empirical evaluation on a benchmark of 5,504 queries; reported accuracy numbers for Brick and the best single model; comparison to other routers.
A continuous preference knob lets operators slide between max-quality and max-saving profiles at deploy time.
Design/feature claim in the paper describing the system's operator control.
We present Brick, a multimodal router that scores each model on six capability dimensions, combines this with a per-query difficulty estimate, and dispatches via a cost-penalized geometric rule.
Methodological description of the proposed system in the paper.
At production scale even small per-request savings become a direct cloud-bill lever.
Argumentative claim in the paper linking per-request savings to total cloud costs; no quantitative backing in the abstract.
Therefore, accounting organizations in the region require targeted training, investment, and institutional support to improve AI adoption and use.
Authors' conclusion/recommendation based on survey and interview findings identifying skills and infrastructure gaps (not empirically tested within the study).
Perceived benefits of employing AI include increased efficiency, improved accuracy, better compliance, and more accurate decisions.
Self-reported perceptions collected via questionnaire and interviews, supported by thematic analysis (sample size not reported).
Accountants in Isabela (Region of Cagayan Valley) demonstrate strong analytical abilities.
Survey questionnaire and interviews analyzed using descriptive statistics and thematic analysis (sample size not reported in text).
There is an urgent necessity for cohesive policy interventions to accelerate the inclusive adoption of digital agriculture in developing economies.
Policy recommendation drawn from the review's synthesis of technological potential and socioeconomic limitations; presented as a conclusion/recommendation rather than a quantified empirical finding.
By transitioning from traditional, intuition-based practices to precision-driven, data-centric methodologies, smart farming facilitates the precise management of crucial inputs such as water, fertilizers, and pesticides, thereby enhancing the yields of staple crops like Zea mays and Glycine max.
Synthesis of agronomic literature presented in the review claiming input-optimization and yield improvements for specific staple crops (maize and soybean); no numeric trial/sample details provided in the abstract.
Emerging technological innovations—including the Internet of Things (IoT), artificial intelligence (AI), unmanned aerial vehicles (UAVs), and blockchain—have a profound impact on optimizing agricultural productivity.
Review article synthesizing literature on multiple technologies (IoT, AI, UAVs, blockchain) and their roles in agriculture; no specific experimental sample sizes provided in the abstract.
The integration of digital agriculture and smart farming technologies represents a transformative evolution in modern agronomy, offering unprecedented solutions to the intertwined crises of global food security, climate change, and resource depletion.
Statement in the review synthesizing existing literature on digital agriculture and its potential impacts; no primary empirical sample or quantitative meta-analysis reported in the abstract.
Extensive experiments show that ComActor achieves state-of-the-art performance on ComCADBench, with strong resilience in long-horizon tasks where baselines collapse, and generalizes to external CAD benchmark.
Empirical results reported by the authors across extensive experiments on ComCADBench and at least one external CAD benchmark (claims of SOTA performance, robustness in long-horizon tasks, and cross-benchmark generalization).
We develop ComActor, a self-correcting agent trained through a progressive three-stage framework, alongside ComForge, a scalable platform for large-scale training in Windows containers.
Authors' methodological and engineering contributions described in the paper (design and implementation of ComActor and ComForge).
We introduce ComCADBench, the first benchmark for agents operating real industrial CAD software.
Paper contribution: creation and release of a new benchmark (authors claim it is the first of its kind).
We identify the Component Object Model (COM) as a unified executable abstraction and propose COM-as-Action: reframing professional software interaction as deterministic program synthesis rather than sequential visual control.
Methodological contribution and framing described by the authors (proposal / conceptual contribution introduced in the paper).
Environment engineering should be treated as a core research direction for developing reliable autonomous research agents.
Authors' recommendation and call to action in the paper (normative claim; no empirical test).
We open-source our code and results.
Statement in the paper that the authors open-sourced code and results (claim about availability; not a performance claim).
EurekAgent discovered new state-of-the-art 26-circle packing results with less than $11 in total API cost.
Specific experimental result reported in the paper: a 26-circle packing result and a stated total API cost under $11 (the excerpt provides the numbers but not the experimental protocol or baselines).
EurekAgent sets new state-of-the-art results on multiple mathematics, kernel engineering, and machine learning tasks.
Empirical claim reported in the paper that EurekAgent achieved state-of-the-art (SOTA) performance on multiple tasks (specific tasks/results and experimental details not included in the excerpt).
EurekAgent engineers the environment along four dimensions: permissions engineering for bounded agent execution and isolated evaluation; artifact engineering for filesystem and Git-based collaboration; budget engineering for budget-aware exploration; and human-in-the-loop engineering for easy human supervision and intervention.
System design specification in the paper describing four engineering dimensions and their intended functions (design description; no empirical validation details in the excerpt).
We present EurekAgent, an environment-engineered agent system for metric-driven autonomous scientific discovery.
Paper reports development and presentation of a system named EurekAgent (system description; design claims in paper).
The bottleneck for autonomous scientific discovery is shifting from prescribing agent workflows to designing agent environments (the resources, constraints, and interfaces that shape agent behavior).
Conceptual argument presented in the paper, motivated by observed improvements in model capabilities (no empirical test or sample size provided in the excerpt).
LLM-based agents have shown increasing potential in automating scientific discovery: given an optimizable metric and an execution environment, they can propose, validate, and iterate scientific solutions, and have produced results that outperform human-designed approaches.
Paper's summary statement citing prior and contemporary LLM-agent results and examples (no specific experiments or sample sizes given in the excerpt).
We interpret HLER as a research harness rather than an autonomous AI scientist: it sharply reduces failures, makes residual weaknesses more visible, and prevents unreliable claims from being advanced as publication-ready outputs.
Interpretation and conclusion drawn from the experimental results (reduced failure rates and qualitative observations about residual weaknesses); presented as the authors' framing of the system.
An 80-run ablation suggests that deterministic computation and human gates contribute independently, with exploratory evidence of complementarity.
Reported ablation experiment of 80 runs isolating the effects of deterministic computation and human decision gates; authors describe independent contributions and exploratory (non-conclusive) evidence of complementarity.
Reliability gains were largest on the least publicly represented dataset, a Qing-dynasty population register.
Reported heterogeneous effect across the four datasets in the experiment, with largest reliability improvement observed on the Qing-dynasty population register dataset (paper links this observation to dataset representation).
Fisher's exact test rejects equality of failure rates between baseline and HLER at p < 0.001.
Statistical test reported in the paper comparing failure rates across conditions (Fisher's exact test result reported as p<0.001).
Using the same underlying model, the same agent decomposition, and identical prompts for the shared reasoning agents, HLER reduced the failure rate to 16% by imposing three architectural commitments: LLMs reason but do not execute data work, data and estimation are handled deterministically, and three human decision gates bind the workflow.
Reported experimental comparison between unconstrained baseline and HLER in the factorial experiment (failure rate under HLER reported as 16%); description of the three architectural commitments provided.
We study this problem through Human-in-the-Loop Economic Research (HLER), a decision architecture based on pre-commitment, decision sequencing, accountability, and attention allocation.
Paper proposes HLER as a specific decision architecture; description and conceptual specification provided in the paper.
Large language models (LLMs) are increasingly used for tasks once reserved for trained researchers, including hypothesis generation, specification choice, and drafting conclusions.
Background/introductory claim in the paper; stated as motivation rather than empirically tested within this study.
projectmem is evaluated through a two-month self-study across 10 projects comprising 207 logged events.
Evaluation description provided in the abstract specifying duration, number of projects, and number of logged events.
projectmem ships as a three-dependency Python package (14 MCP tools, 19 CLI commands, 37 automated tests).
Package composition and counts stated in the abstract; presumably verifiable in the project's repository.
The system runs fully offline with no telemetry; its immutable log also serves as a provenance trail for reproducible, auditable AI-assisted development.
Design and privacy claim made in the abstract (no detailed audit/reproducibility study in the abstract).
We frame this as Memory-as-Governance: memory that does not merely answer the agent but acts on its next action.
Conceptual framing presented in the abstract.
projectmem adds a deterministic pre-action gate that warns an agent before it repeats a previously failed fix or edits a known-fragile file.
Design claim in the abstract describing a pre-action gate feature (no empirical results in abstract).
projectmem records development as an append-only, plain-text event log of typed events (issues, attempts, fixes, decisions, and notes) and deterministically projects that log into compact, AI-readable summaries served through the Model Context Protocol (MCP).
System design and implementation description provided in the abstract (claims about log format, determinism, and that it serves summaries via MCP).
We present projectmem, an open-source, local-first memory and judgment layer for AI coding agents.
Paper describes and links the project's source code repository (https://github.com/riponcm/projectmem); claim of existence and characteristics in abstract.
The complete replication package, including the curated 151-repository monthly panel, is publicly available.
Statement in abstract about data and code availability; implies a replication package and curated panel are released with the paper.
Lines of code grow by +12.8% after adoption (p = 0.003).
Estimated treatment effect on lines of code from the staggered DiD / Borusyak estimator using the same 151-repository panel.
AI coding tools are now used by a majority of developers, and agentic use of these tools has popularized the practice colloquially called "vibe coding".
Statement in paper abstract (background claim); no dataset or analysis in this paper is reported to directly measure share of developers using AI tools or prevalence of "vibe coding".
Selection of predictive quantiles is a critical lever: by calibrating predictive quantiles, one can balance the trade-off between resource efficiency and service reliability.
Systematic analysis and experiments reported in the paper that vary predictive quantile selection and measure resulting resource efficiency and reliability outcomes; authors provide guidelines for calibration.
Foundation (time-series) foundation models demonstrate superior zero-shot forecasting accuracy.
Experimental results reported in the paper comparing zero-shot forecasting accuracy between foundation models and other model classes.
We conduct an extensive evaluation of statistical, deep learning, and foundation models on the CloudCons benchmark.
Experimental study reported in the paper comparing multiple model classes across the datasets and tasks.
We build high-quality datasets that cover diverse workloads from Huawei Cloud, Microsoft Azure, and Google Borg, capturing distinct service characteristics ranging from synchronized diurnal rhythms to stochastic, pulse-like bursts and high-frequency noise.
Dataset construction and description in the paper; datasets sourced from three providers and characterized by different workload patterns.
We propose CloudCons, a comprehensive end-to-end benchmark designed to evaluate forecasting models within the specific context of cloud resource consolidation.
This paper's contribution: development and presentation of the CloudCons benchmark and associated methodology.
The forecast-then-optimize paradigm has emerged to optimize consolidation by anticipating future demands.
Descriptive claim based on prior research and literature cited by the authors; positioned as the methodological approach motivating the study.
These results motivate the need for research to assist practitioners in framing the development of instruction files as a software engineering activity ("Instructions-as-Code").
Conclusion/recommendation drawn from the observed mixed effects of instruction files on agentic PR performance and the exploratory finding about instruction-file characteristics.
Projects that managed to increase their merge rate have substantially longer instruction files, which are also well structured into a higher number of sections and sub-sections.
Exploratory comparison of instruction-file characteristics (length, number of sections/sub-sections) between projects that improved merge rate versus those that did not (described as a 'first exploration' in the paper).