Evidence (8974 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
10085 claims
Filter claims →
Productivity
8974 claims
Filtered →
Governance
8062 claims
Filter claims →
Human-AI Collaboration
7749 claims
Filter claims →
Org Design
5057 claims
Filter claims →
Innovation
4896 claims
Filter claims →
Labor Markets
4088 claims
Filter claims →
Skills & Training
3372 claims
Filter claims →
Inequality
2377 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 882 | 244 | 117 | 1097 | 2424 |
| Governance & Regulation | 1010 | 469 | 229 | 135 | 1875 |
| Organizational Efficiency | 977 | 235 | 149 | 90 | 1462 |
| Technology Adoption Rate | 781 | 299 | 143 | 128 | 1362 |
| Research Productivity | 506 | 155 | 74 | 363 | 1110 |
| Output Quality | 555 | 219 | 71 | 70 | 915 |
| Decision Quality | 395 | 200 | 95 | 54 | 751 |
| Firm Productivity | 523 | 67 | 101 | 27 | 724 |
| AI Safety & Ethics | 262 | 309 | 75 | 36 | 688 |
| Market Structure | 195 | 201 | 135 | 30 | 566 |
| Task Allocation | 248 | 77 | 96 | 38 | 464 |
| Innovation Output | 300 | 34 | 55 | 20 | 411 |
| Skill Acquisition | 207 | 75 | 65 | 21 | 368 |
| Employment Level | 138 | 67 | 119 | 24 | 350 |
| Fiscal & Macroeconomic | 156 | 80 | 53 | 33 | 329 |
| Task Completion Time | 211 | 38 | 13 | 16 | 280 |
| Firm Revenue | 183 | 52 | 29 | 5 | 270 |
| Consumer Welfare | 131 | 77 | 48 | 13 | 269 |
| Inequality Measures | 50 | 141 | 54 | 9 | 254 |
| Worker Satisfaction | 104 | 85 | 25 | 13 | 227 |
| Error Rate | 87 | 112 | 11 | 5 | 215 |
| Automation Exposure | 69 | 69 | 37 | 20 | 198 |
| Wages & Compensation | 102 | 49 | 31 | 11 | 193 |
| Team Performance | 115 | 30 | 30 | 11 | 187 |
| Regulatory Compliance | 88 | 74 | 17 | 7 | 186 |
| Training Effectiveness | 109 | 22 | 14 | 21 | 168 |
| Developer Productivity | 116 | 21 | 15 | 8 | 161 |
| Job Displacement | 12 | 92 | 26 | 1 | 131 |
| Hiring & Recruitment | 57 | 12 | 9 | 5 | 83 |
| Skill Obsolescence | 6 | 59 | 10 | 2 | 77 |
| Social Protection | 43 | 17 | 8 | 2 | 70 |
| Creative Output | 35 | 21 | 9 | 4 | 70 |
| Labor Share of Income | 18 | 23 | 17 | 1 | 59 |
| Worker Turnover | 15 | 16 | — | 4 | 35 |
| Industry | — | — | — | 1 | 1 |
Productivity
Remove filter
AwareLLM reduced mental demand for participants.
Reported results from the user study (comparison to a standard LLM assistant) with 20 participants; abstract reports reductions but gives no quantitative metrics.
AwareLLM led to reductions in cognitive fatigue.
Reported results from the user study comparing AwareLLM to a standard LLM assistant; sample size 20. No quantitative values provided in the abstract.
Compared to a standard LLM assistant, AwareLLM produced statistically significant improvements in task performance.
Results reported from the user study (comparison vs. standard LLM); sample size noted as 20 participants. No numerical effect size provided in the abstract.
AwareLLM dynamically adapts to users' psychophysiological states while analyzing temporal patterns and behavioral tendencies to provide personalized and timely interventions.
Design and claimed operational behavior of the proposed framework as described by authors.
We introduce AwareLLM, a multimodal framework that integrates egocentric vision, pupillometry, eye-gaze tracking, posture detection, heart activity, and large language models to create a proactive and context-aware ecosystem.
System/methods description in paper (architecture/design claim).
Information workers' productivity is significantly influenced by their cognitive states and physiological responses.
Background statement in paper (literature-motivated claim); no study data provided within the abstract to support it.
An AI Workflow Store of hardened and reusable workflows would allow agents to invoke workflows with far greater reliability and security than improvised tool chains.
Vision/proposal in the paper advocating an AI Workflow Store as a solution; presented conceptually without experimental or deployment evidence.
Integrating rigorous software engineering processes into the agentic loop will produce production-grade, hardened, and deterministically-constrained agent workflows that substantially outperform brittle on-the-fly synthesis.
Prescriptive claim / proposed hypothesis in the paper advocating integration of SE practices into agent workflows; offered as a reasoned proposal without empirical results.
Unlike existing datasets, our benchmark utilizes a seed-driven architecture to simulate dynamic environment states and unpredictable API failures, ensuring a deterministic yet diverse evaluation.
Methodological description: seed-driven architecture and simulated API failures; claimed as a distinguishing design feature versus prior datasets.
ComplexMCP provides over 300 meticulously tested tools derived from 7 stateful sandboxes, ranging from office suites to financial systems.
Benchmark construction details reported in the paper: >300 tools, 7 stateful sandboxes (explicit counts provided).
We introduce ComplexMCP, a benchmark designed to evaluate agents in rigorous conditions built on the Model Context Protocol (MCP).
Design and construction of the benchmark reported by authors; methodological description (benchmark/tooling claim).
BenchCAD positions itself as a benchmark for measuring and improving the industrial readiness of multimodal CAD automation.
Authors' stated goal/purpose in the paper/abstract describing BenchCAD as a benchmark intended to measure and guide improvements towards industrial readiness.
Industrial CAD code generation requires models to produce executable parametric programs from visual or textual inputs and to understand 3D structure, infer engineering parameters, and choose CAD operations that reflect design and manufacture.
Problem definition and motivation provided by the authors in the paper/abstract describing the necessary capabilities for industrial CAD code generation.
BenchCAD enables fine-grained analysis across perception, parametric abstraction, and executable program synthesis.
Authors' description of benchmark scope and tasks designed to probe perception (visual understanding), parametric abstraction (inferring engineering parameters), and executable program synthesis (generating runnable CadQuery programs).
BenchCAD evaluates models through visual question answering, code question answering, image-to-code generation, and instruction-guided code editing.
Benchmark design described in the paper/abstract listing four evaluation tasks (VQA, code QA, image-to-code, instruction-guided code editing).
BenchCAD contains 17,900 execution-verified CadQuery programs across 106 industrial part families.
Dataset construction reported in the paper/abstract: explicit statement of 17,900 execution-verified CadQuery programs spanning 106 industrial part families (e.g., bevel gears, compression springs, twist drills).
AI is the most important predictive factor for Lae (based on artificial neural network analysis).
Artificial neural network (ANN) predictive modeling on composite indices for AI and Lae using panel data from 2012–2022 across 30 provincial regions; variable importance ranking from ANN indicates AI as top predictor.
An exogenous shock test using the Big Data Pilot Zone policy further confirms the robustness of the AI–Lae relationship findings.
Policy shock (Big Data Pilot Zone) robustness test performed on the same panel of 30 provincial regions (2012–2022); described as an exogenous shock test corroborating the main results.
Regression results show a positive relationship between firm performance and breadth of AI integration.
Multivariate regression analysis reported in the paper using BTOS AI supplement data (Nov 2025–Jan 2026); association between firm performance (dependent variable) and measures of AI integration breadth (independent variables); sample size and controls not included in excerpt.
Most firms (66%) use AI for task augmentation rather than replacement.
Survey responses about intent/role of AI within firms from the BTOS AI supplement (Nov 2025–Jan 2026); descriptive percent reporting augmentation vs replacement; sample size not provided.
Worker-level AI use appears in 23% of firms (41%, employment-weighted), primarily for writing, document analysis, and information search.
Firm-reported presence of worker-task AI use from the BTOS AI supplement (Nov 2025–Jan 2026); descriptive percentages given, employment-weighted alternative reported; sample size not provided in excerpt.
Among adopter firms, AI is most often used in Sales and Marketing (52%), Strategy (45%), and IT (41%).
Function-specific adoption rates reported from the BTOS AI supplement descriptive statistics (Nov 2025–Jan 2026); sample restricted to adopter firms; sample sizes not stated.
AI use is concentrated in large firms and knowledge-intensive sectors, reaching 50%–60% (60%–70%, employment-weighted) among very large firms in Information, Professional Services, and Finance.
Stratified descriptive statistics by firm size and industry from the BTOS AI supplement (Nov 2025–Jan 2026); employment-weighted estimates reported; exact sample sizes by stratum not provided in excerpt.
Adoption is expected to reach 22% of firms within six months.
Survey question asking firms about expected near-term adoption (BTOS AI supplement, Nov 2025–Jan 2026), producing a stated expected adoption rate; sample size not given.
Employment-weighted adoption rate was 32% (i.e., 32% of employment is in firms using AI in at least one function).
Employment-weighted descriptive statistic from the BTOS AI supplement covering Nov 2025–Jan 2026 (survey-based weighting by employment; sample size not stated).
During Nov 2025–Jan 2026, 18% of firms used AI in at least one function.
Descriptive statistics from the 2026 AI supplement to the U.S. Census Bureau’s Business Trends and Outlook Survey (BTOS), fielded Nov 2025–Jan 2026; nationally representative firm survey (sample size not stated in excerpt).
Regulatory modernisation, secure national data infrastructure and targeted digital training are essential to enable sustainable innovation in valuation practice.
Policy and practitioner recommendations derived from interview data and thematic analysis; synthesis into prescriptive recommendations.
Evaluation indicates improved architectural consistency and deployability compared to general-purpose AI code generation workflows, suggesting that constraint-aware retrieval is essential for aligning AI-assisted service development with production software engineering practices.
Paper reports an evaluation comparing the proposed retrieval-augmented scaffolding approach to general-purpose AI code generation workflows and concludes improvements in architectural consistency and deployability; the excerpt does not provide evaluation design details, metrics, or sample size.
By combining template retrieval with structured interaction, the method embeds production-relevant considerations during service scaffolding.
Paper's description of the mechanism by which the proposed approach operates (template retrieval + structured interaction) to incorporate production concerns; presented as a design claim without detailed empirical quantification in the excerpt.
We propose a retrieval-augmented scaffolding approach that combines platform-based code generation with agentic clarification loops to expose and resolve architectural constraint ambiguities.
Methodological contribution described in the paper: a retrieval-augmented scaffolding method combining template retrieval and agentic clarification loops; this is a proposed approach rather than reported empirical proof in the provided text.
AI-assisted development tools enable rapid prototyping of services.
Stated assertion in paper's introduction/abstract that AI-assisted tools speed up prototyping; no quantitative evaluation or sample size given in the provided text.
The C³ Framework provides implementable design patterns and testable propositions intended to help accounting leaders capture productivity gains from human + AI work while preserving accountability, consistency, and alignment with governance expectations in high-stakes reporting contexts.
Conclusions section stating intended practical utility; presented as intended outcomes of applying the proposed framework, not as empirically demonstrated results in this paper.
The paper proposes a role taxonomy that clarifies review responsibility, escalation thresholds, and evidence retention for human–AI collaboration in accounting.
Results section proposing a role taxonomy as part of the C³ Framework; presented as a design artifact derived from synthesis of research and guidance.
The framework specifies five mandatory control points for high-judgment use cases: source grounding and traceability, independent verification and tie-out, contradiction testing, escalation and approval, and audit-trail logging.
Results section listing five control points as mandatory design elements for high-judgment accounting use cases; conceptual recommendation from synthesis.
The paper develops the C³ Framework—Complementarity, Controls, and Competencies—which maps accounting tasks by task structure and judgment/materiality to recommend collaboration modes.
Results section: conceptual framework developed by the authors based on synthesized literature and guidance; no reported empirical validation in the abstract.
AI accelerates drafting, summarization, and pattern detection in accounting while professionals remain accountable for judgment, materiality, and defensibility in financial reporting and analysis.
Statement in paper summarizing literature and practitioner guidance (2023–2025); conceptual synthesis rather than new empirical data.
AI tools can serve as valuable aids in task splitting, provided there is human oversight to filter out irrelevant tasks.
Paper's conclusion synthesizing experimental results and participant feedback, recommending human-in-the-loop oversight when using AI for task-splitting.
Participants favored a hybrid approach, combining AI tools with conventional methods to maintain high accuracy in planning.
Participant preferences and qualitative feedback reported from the controlled experiment indicating preference for combining AI assistance with human methods; sample size not provided.
AI-assisted approaches can help ensure no important tasks are overlooked during task-splitting.
Reported finding from the experiment indicating AI assistance reduced omissions in task lists (paper statement based on experiment and participant observations); sample size not stated.
AI-assisted approaches can generate more granular task lists than traditional methods.
Experimental comparison reported in the paper showing AI-generated task lists were more granular (based on task lists produced during the controlled experiment); sample size not provided in summary.
Switchcraft saves over $3,600 per million queries.
Cost savings estimate reported in the paper based on the measured 84% reduction applied to a million-query baseline.
Switchcraft reduces inference cost by 84%.
Empirical cost analysis reported in the paper comparing inference cost with and without Switchcraft.
Switchcraft's accuracy matches or exceeds the best individual model.
Empirical comparison reported in the paper between Switchcraft accuracy (82.9%) and accuracies of individual models (details summarized by authors).
Switchcraft achieves 82.9% accuracy.
Empirical evaluation results reported in the paper (accuracy metric measured on the evaluation framework).
Switchcraft operates inline, selecting the lowest-cost model subject to correctness.
Method description in the paper describing Switchcraft's operational design.
We present Switchcraft, the first (to the best of our knowledge) model router optimized for agentic tool calling.
Authors' stated contribution / novelty claim in the paper (method description: Switchcraft).
A lightweight pre-generation router exceeds the best cascade policy on four of five datasets, mainly because it avoids the cheap model's generation cost on queries sent directly to a larger model rather than because of a stronger routing signal.
Empirical experiments across the five benchmarks showing the pre-generation router outperforms best cascade on 4/5 datasets; analysis attributing the advantage primarily to avoided generation cost rather than improved routing accuracy/signal.
The theoretical superiority of SignSGD accurately predicts its faster convergence during the pretraining of a 124M parameter GPT-2 model.
Empirical experiment reported in the paper: pretraining runs of a 124M-parameter GPT-2 model comparing SignSGD (or Muon) vs baseline SGD/variants; details (number of runs, datasets, seeds) are not provided in the abstract.
Extending the sign operator to matrices preserves the optimal scaling with dimensionality: we provide an equivalent optimal lower bound for the Muon optimizer in the matrix domain.
Theoretical extension of the analysis to matrix-valued problems and derivation of a matching optimal lower bound for the Muon optimizer, demonstrating preserved scaling.
SignSGD effectively reduces the complexity by a factor of d under sparse noise, where d is the problem dimension (comparison of SignSGD upper bound with SGD lower bound shows a factor-d improvement).
Theoretical comparison between the derived upper bound for SignSGD and the derived lower bound for SGD within the paper, under the separable/sparse noise model and specified smoothness assumptions.