Evidence (7560 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
9875 claims
Filter claims →
Productivity
8807 claims
Filter claims →
Governance
7870 claims
Filter claims →
Human-AI Collaboration
7560 claims
Filtered →
Org Design
4892 claims
Filter claims →
Innovation
4781 claims
Filter claims →
Labor Markets
4004 claims
Filter claims →
Skills & Training
3308 claims
Filter claims →
Inequality
2332 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 870 | 233 | 116 | 1066 | 2363 |
| Governance & Regulation | 976 | 451 | 218 | 133 | 1809 |
| Organizational Efficiency | 949 | 224 | 144 | 88 | 1416 |
| Technology Adoption Rate | 764 | 287 | 141 | 122 | 1325 |
| Research Productivity | 501 | 152 | 74 | 362 | 1101 |
| Output Quality | 542 | 216 | 69 | 69 | 896 |
| Decision Quality | 387 | 198 | 94 | 54 | 740 |
| Firm Productivity | 513 | 67 | 101 | 27 | 714 |
| AI Safety & Ethics | 249 | 303 | 73 | 36 | 667 |
| Market Structure | 190 | 192 | 134 | 27 | 548 |
| Task Allocation | 243 | 77 | 91 | 36 | 452 |
| Innovation Output | 291 | 33 | 55 | 20 | 401 |
| Skill Acquisition | 206 | 72 | 65 | 21 | 364 |
| Employment Level | 133 | 63 | 115 | 22 | 335 |
| Fiscal & Macroeconomic | 153 | 79 | 52 | 32 | 323 |
| Task Completion Time | 206 | 37 | 12 | 15 | 272 |
| Firm Revenue | 179 | 52 | 29 | 5 | 266 |
| Consumer Welfare | 130 | 76 | 47 | 13 | 266 |
| Inequality Measures | 48 | 137 | 51 | 6 | 242 |
| Worker Satisfaction | 101 | 81 | 25 | 13 | 220 |
| Error Rate | 84 | 110 | 11 | 5 | 210 |
| Wages & Compensation | 98 | 47 | 30 | 10 | 185 |
| Regulatory Compliance | 88 | 73 | 17 | 7 | 185 |
| Automation Exposure | 66 | 64 | 33 | 16 | 182 |
| Team Performance | 105 | 29 | 30 | 11 | 176 |
| Training Effectiveness | 109 | 22 | 14 | 21 | 168 |
| Developer Productivity | 114 | 21 | 14 | 8 | 158 |
| Job Displacement | 12 | 90 | 24 | 1 | 127 |
| Hiring & Recruitment | 57 | 9 | 9 | 5 | 80 |
| Skill Obsolescence | 6 | 56 | 9 | 1 | 72 |
| Social Protection | 43 | 17 | 8 | 2 | 70 |
| Creative Output | 35 | 21 | 9 | 4 | 70 |
| Labor Share of Income | 18 | 21 | 17 | 1 | 57 |
| Worker Turnover | 15 | 16 | — | 4 | 35 |
| Industry | — | — | — | 1 | 1 |
Human Ai Collab
Remove filter
We introduce a new benchmark QuantSightBench to assess prediction-interval forecasting capability and evaluate frontier models under multiple settings, assessing both empirical coverage and interval sharpness.
Methodological contribution reported in the paper: creation of QuantSightBench and its use to evaluate models on empirical coverage and sharpness (paper describes benchmark and evaluation procedure; specific task/sample counts not given in excerpt).
Technology-driven recruitment encompasses Applicant Tracking Systems (ATS), AI-powered screening, video-based interviews, gamified assessments, and data analytics.
Conceptual description in the paper's introduction/background defining the scope of 'technology-driven recruitment'.
The study employed a mixed-methods research design combining a quantitative survey of 150 HR professionals and recruiters across manufacturing, IT, banking, and education sectors with qualitative case study analysis of four organizations in Chhatrapati Sambhajinagar.
Explicit methodological statement in the paper: quantitative survey (N=150) across specified sectors + qualitative case studies of 4 organizations in Chhatrapati Sambhajinagar.
Prior research often treats AI presence as binary, framing it either as a hidden tool or a visible teammate.
Literature-summary claim asserted by the authors (literature review / conceptual critique). No quantitative evidence reported in the abstract.
The LLM fallacy is situated within existing literature on automation bias, cognitive offloading, and human–AI collaboration, but is distinguished as a form of attributional distortion specific to AI-mediated workflows.
Conceptual positioning and literature synthesis in the paper; claim is analytic rather than empirically tested in the abstract.
Less attention has been given to how LLM usage reshapes users' perceptions of their own capabilities.
Literature gap claim from the paper's review of prior research on model reliability, hallucination, and trust calibration; no quantitative synthesis or meta-analysis reported.
The system evaluation was performed in a deployed multi-tenant enterprise application across three conditions: manual operation, unconstrained AI with safety layers disabled, and full bounded autonomy.
Method description in the paper's evaluation section: deployment context (multi-tenant enterprise app), three experimental conditions, and 25 scenario trials spanning seven failure families.
We conducted a year-long longitudinal study of AI use in a high-stakes workplace among cancer specialists.
Methodological statement in the paper indicating a year-long longitudinal empirical study with cancer specialists (no sample size or detailed methods reported in abstract).
We evaluate AIBuildAI on MLE-Bench, a benchmark of realistic Kaggle-style AI development tasks spanning visual, textual, time-series and tabular modalities.
Evaluation methodology described in paper (benchmark selection and task modalities).
We surveyed 860 Microsoft developers to understand where they want AI support, and where they want it to stay out.
Primary empirical method reported in the paper (survey) with sample size explicitly stated as 860 Microsoft developers.
Developers spend roughly one-tenth of their workday writing code.
Statement reported in the paper (abstract). No sample-size or measurement method for this specific statistic provided in the abstract.
We tested 9 frontier models on BTB.
Abstract states that nine frontier models were evaluated using the benchmark.
Completing a BTB task takes bankers up to 21 hours, underscoring the economic stakes of successfully delegating this work to AI.
Reported time-to-complete statistic in abstract (claimed maximum of 21 hours per task); implies measurement of human task completion time by bankers.
BTB requires agents to execute senior banker requests by navigating data rooms, using industry tools (market data platform, SEC filings database), and generating multi-file deliverables including Excel financial models, PowerPoint pitch decks, and PDF/Word reports.
Benchmark design specifications reported in abstract describing the tasks and artifact types agents must produce.
BankerToolBench (BTB) is an open-source benchmark of end-to-end analytical workflows routinely performed by junior investment bankers.
Paper describes BTB design and explicitly states it is open-source and targets end-to-end workflows for junior investment bankers.
We collaborated with 502 investment bankers from leading firms to develop an ecologically valid benchmark grounded in representative work environments.
Reported collaboration/sample size stated in abstract: 502 investment bankers involved in benchmark development.
The global onset of Industry 4.0 and Artificial Intelligence (AI) necessitates a re-evaluation of employment forecasts for Nagpur's medium enterprises.
Interpretive/prescriptive claim based on the paper's framing of technological change (Industry 4.0/AI) and implications for employment forecasting; no empirical sample size or quantitative backing provided in the excerpt.
Medium-scale industries in zones like Butibori and Hingna have traditionally been labor-intensive.
Descriptive statement in the paper about the nature of current industries in Nagpur/MIDC; no sample size or quantitative data reported in the excerpt.
The baselines are implemented as prompts, representing the realistic deployment alternative to a governed framework.
Methodological statement in paper describing how baselines were implemented (as prompts); presented as representing realistic alternative deployment.
We benchmark three systems on an 11-case balanced prior authorization appeal evaluation set.
Methodological statement in paper describing evaluation; sample size explicitly stated as 11 cases.
The two case firms demonstrated contrasting approaches to implementing AI in recruitment.
Findings and case descriptions comparing the two firms' AI recruitment strategies and levels of implementation (n = 2 firms; interviews with 22 participants).
The research contributes by shifting focus to under-researched non-Western workplace settings, particularly technologically advancing Middle Eastern economies like Qatar.
Paper's stated contribution and scope: focus on Qatari organisations and Middle Eastern context.
Four key themes emerged from the data: (1) process optimisation through AI integration, (2) subjectivity in AI-powered recruitment, (3) recruitment strategies in the age of AI, and (4) strategic investments in AI.
Findings: thematic analysis identified these four themes from interview data (n = 22) across the two case firms.
Thematic analysis was used to identify patterns and relationships within the interview data.
Methods: analysis section reporting use of thematic analysis framework.
Data were collected through semi-structured interviews with twenty-two participants across various organisational roles and hierarchical levels.
Methods: semi-structured interviews reported with total participants n = 22 across roles/levels.
The research investigated two prominent Qatari firms with contrasting AI recruitment implementation approaches.
Methods / case selection: two firms were selected and contrasted on their AI recruitment approaches (number of firms = 2).
The study employed an interpretivist philosophy and a case study design.
Methods section: explicitly states interpretivist philosophy and case study design.
All participants had access to the same AI tool; the experiment varied only the structure surrounding its use (behavioral vs cognitive scaffolding vs unstructured).
Experimental design description in the paper: common AI tool provided to all participants; randomization/assignment varied only the scaffolding around AI use.
We found no evidence that information provision drove effects on our behavioural outcomes.
Analysis from the preregistered experiments showing that manipulations of information provision did not produce corresponding changes in measured behaviours (e.g., petition signing, donations).
We observed no evidence of a correlation between AI persuasion effects on attitudes and behaviour.
Analysis reported in the two preregistered experiments comparing AI-induced changes in attitudes with corresponding behavioural outcomes across participants (sample reported in paper).
These results are observational and reflect a single-operator dataset without controlled comparison.
Author statement in the paper describing study limitations.
There is a significant research gap in comparative understanding of generative AI's impact across developed and developing economies; differences in infrastructure, labour markets, and skill distributions may lead to uneven outcomes.
Review observation that the included literature lacks sufficient comparative studies across country-development contexts (explicitly noted as a gap in the paper).
This systematic literature review synthesised findings from 40 empirical and conceptual studies published between 2020 and 2025 using the PRISMA framework (search across Google Scholar and Dimensions.ai), yielding 3,252 database records plus 8 hand-searched studies, of which 40 met the inclusion criteria.
PRISMA-style structured literature search reported in the paper: database search (Google Scholar, Dimensions.ai) returning 3,252 records, 8 hand-searched records, 40 studies meeting inclusion.
The explanatory interface has no significant impact on situational trust.
Trust measured in different forms (situational, learned, cognitive, emotional) in the RCT; authors report no significant effect of explanatory interface on situational trust (N=120).
Under the sequential AI-assisted decision-making paradigm, the explanatory interface has no significant effect on immediate task performance.
Same randomized controlled experiment; authors report no significant effect of explanatory interface on immediate task performance in the sequential paradigm (N=120 total).
The study was a randomized controlled experiment with 120 pre-service teachers.
Randomized controlled experiment reported in the paper; sample described as 120 pre-service teachers.
We instantiate this vision in a controlled study (n=36) comparing the gaze-aware AI assistant to a text-only LLM assistant.
The paper reports running a controlled user study with sample size n=36 directly comparing the gaze-aware assistant against a text-only LLM assistant.
Data were analyzed using partial least squares structural equation modeling (PLS-SEM) implemented in SmartPLS 4.
Methods section statement in paper indicating use of PLS-SEM and SmartPLS 4 for data analysis.
The empirical analysis is based on a questionnaire survey administered to 324 respondents from Romanian organizations operating in IT, services, industry, and public administration.
Questionnaire survey described in paper; sample size explicitly stated as 324 respondents from Romanian organizations across IT, services, industry, and public administration.
All four models converge to similar skill profiles (3.6-point spread), suggesting that text-based automation feasibility may be more skill-dependent than model-dependent.
Comparison across 4 LLMs (LLaMA 3.3 70B, Mistral Large, Qwen 2.5 72B, Gemini 2.5 Flash) with reported 3.6-point spread in skill-profile SAFI scores.
We validate this principle through a controlled experiment on log format token economy across four conditions (human-readable, structured, compressed, and tool-assisted compressed).
Controlled experiment described in the paper comparing four log-format conditions (human-readable, structured, compressed, tool-assisted compressed); exact sample size not reported in the abstract.
For six decades, software engineering principles have been optimized for a single consumer: the human developer.
Historical/position claim asserted in the paper (conceptual/literature-based argument), no empirical sample or quantitative test reported.
Through a causal decomposition that automates one side of agent communication, we separate cooperation failures from competence failures, tracing their origins through agent reasoning analysis.
Method described in the paper: causal decomposition approach that automates one side of communication and analyzes agent reasoning to attribute failures (methodological claim; abstract mentions the approach but gives no sample size or quantitative metrics there).
Capability does not predict cooperation.
Comparative experimental results reported in the paper showing different models with different capability levels achieving substantially different collective cooperation outcomes (specifically comparing OpenAI o3 and o3-mini).
We build a multi-agent setup designed to study cooperative behavior in a frictionless environment, removing all strategic complexity from cooperation.
Methodological description in the paper: design and implementation of a multi-agent experimental setup intended to remove strategic complexity (no sample size or quantitative detail reported in the abstract).
A pre-registered experiment evaluates this thesis in a commons production economy -- where agents share a finite resource pool and collaboratively produce value -- at 50-1,000 agent scale.
Paper states that a pre-registered experiment is planned/described; the experiment context (commons production economy) and planned scale (50-1,000 agents) are specified in the excerpt. No experimental outcomes or effect estimates are reported here.
We instantiate SoP in AgentCity on an EVM-compatible layer-2 blockchain (L2) with a three-tier contract hierarchy (foundational, meta, and operational).
Reported implementation/instantiation described in the paper (system implementation claim). The paper states the platform (AgentCity) and technical details (EVM-compatible L2, three-tier contracts).
In this architecture, smart contracts are the law itself -- the actual legislative output that agents produce and that governs their behavior.
Architectural/design claim in the paper describing conceptual role of smart contracts within SoP; presented as an intended property of the system.
Agents discover, transact with, and delegate to agents owned by other parties without centralized oversight.
Asserted behavior pattern of autonomous agents in the paper's motivation; presented as descriptive claim rather than supported by a reported experiment or dataset in the excerpt.
Autonomous AI agents are beginning to operate across organizational boundaries on the open internet.
Stated as an empirical observation in the paper's introduction/introduction-level motivation; no specific dataset or sample described in the text excerpt.