Evidence (362 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
21267 claims
Filter claims →
Productivity
17978 claims
Filter claims →
Governance
17038 claims
Filter claims →
Human-AI Collaboration
16914 claims
Filter claims →
Org Design
11104 claims
Filter claims →
Innovation
11087 claims
Filter claims →
Labor Markets
6711 claims
Filter claims →
Skills & Training
5616 claims
Filter claims →
Inequality
4343 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 1880 | 496 | 296 | 1854 | 4721 |
| Organizational Efficiency | 2906 | 665 | 438 | 180 | 4210 |
| Governance & Regulation | 2162 | 929 | 480 | 247 | 3866 |
| Technology Adoption Rate | 1533 | 545 | 278 | 210 | 2593 |
| Decision Quality | 1391 | 534 | 321 | 173 | 2429 |
| Output Quality | 1298 | 472 | 231 | 145 | 2153 |
| AI Safety & Ethics | 682 | 821 | 230 | 90 | 1837 |
| Research Productivity | 855 | 253 | 121 | 425 | 1675 |
| Firm Productivity | 1105 | 171 | 175 | 73 | 1531 |
| Task Allocation | 735 | 229 | 361 | 99 | 1433 |
| Market Structure | 457 | 461 | 251 | 47 | 1222 |
| Innovation Output | 673 | 94 | 108 | 36 | 913 |
| Task Completion Time | 499 | 118 | 43 | 38 | 702 |
| Firm Revenue | 458 | 130 | 61 | 26 | 677 |
| Skill Acquisition | 381 | 122 | 113 | 34 | 650 |
| Consumer Welfare | 316 | 176 | 115 | 39 | 648 |
| Employment Level | 223 | 143 | 177 | 53 | 600 |
| Error Rate | 246 | 282 | 44 | 19 | 594 |
| Fiscal & Macroeconomic | 283 | 142 | 78 | 52 | 562 |
| Inequality Measures | 103 | 329 | 106 | 13 | 552 |
| Worker Satisfaction | 225 | 185 | 63 | 30 | 503 |
| Automation Exposure | 158 | 155 | 72 | 37 | 426 |
| Regulatory Compliance | 186 | 126 | 35 | 14 | 362 |
| Team Performance | 193 | 56 | 51 | 24 | 326 |
| Developer Productivity | 224 | 58 | 27 | 13 | 323 |
| Wages & Compensation | 148 | 108 | 50 | 17 | 323 |
| Training Effectiveness | 218 | 44 | 21 | 27 | 313 |
| Job Displacement | 23 | 159 | 53 | 5 | 240 |
| Hiring & Recruitment | 109 | 61 | 32 | 11 | 215 |
| Skill Obsolescence | 16 | 107 | 26 | 6 | 155 |
| Creative Output | 71 | 44 | 28 | 6 | 150 |
| Social Protection | 58 | 31 | 12 | 3 | 104 |
| Labor Share of Income | 29 | 43 | 25 | 2 | 99 |
| Worker Turnover | 45 | 29 | 6 | 4 | 84 |
| Industry | — | — | — | 1 | 1 |
Adoption of electronic tax filing improved tax compliance among small and medium-sized enterprises in Lagos State, while internet access remained an important barrier.
Nigerian study of electronic tax filing and compliance among Lagos State SMEs; sample size is not reported in the supplied text.
Digital tax administration can improve compliance in Nigeria's informal sector, but its effectiveness is constrained by structural informality and administrative friction.
Study of digital tax administration and tax compliance in Nigeria's informal sector using Nigerian data; sample size is not reported in the supplied text.
Use of Indonesia's Core Tax Administration System reduced perceived compliance costs and increased compliance intentions, although excessive reliance on automation may weaken professional vigilance over time.
A 2026 evaluation of Indonesia's Core Tax Administration System; the supplied text does not report the evaluation's sample size.
Qualitative evidence from Indonesia indicates that AI can strengthen tax-law enforcement, improve taxpayer convenience and perceived fairness, and lower compliance costs, but implementation readiness, cost, and governance are barriers.
Qualitative primary research on tax-administration modernization in Indonesia.
Compliance with the European Union's voluntary General Purpose AI Code of Practice was divided among major AI firms: OpenAI, Anthropic, Microsoft, and Google committed to sign; Meta declined; and no major China-based AI developer signed.
Stakeholder analysis using the European Commission's account of firm responses to the EU General Purpose AI Code of Practice.
The Albanese government presented the Online Safety Amendment Bill 2024 as applying a duty of care while structuring it to avoid open-ended, binding legal obligations and liabilities.
Analysis of the bill’s legal text, explanatory memorandum, government statements, press releases, and parliamentary debate concerning the under-16 social-media restriction.
Firm size significantly moderated the relationship between descriptive analytics and tax evasion, such that the analytics-related compliance effect differed by firm size.
Hierarchical moderated regression using an interaction term for descriptive analytics × firm size; the interaction was statistically significant with β = -0.198 and p = 0.000.
On-chain immutability may increase the probability of detecting tampering with recorded transactions and reduce some fraud incentives, while off-chain manipulation and weaknesses in oracles or key management create new avenues for misreporting.
Conceptual fraud and enforcement analysis contrasting the deterrent effects of immutable records with vulnerabilities at off-chain interfaces and access-control points.
The orchestrator-subagent architecture elicited more policy-relevant trigger facts than the single-loop baseline on Qwen2.5-32B, but then attenuated most of those facts.
Across 100 episodes per arm, D2 discovered trigger facts in 27 episodes versus 16 for D0; D2 attenuated 22 of the 27 discovered facts, while D0 attenuated none.
The same fact-attenuation mechanism can cause either under-escalation or over-escalation, depending on whether the dropped fact is a risk signal or an exculpating finding.
Paired mirror tasks kyc-0004 and kyc-0005 on Qwen2.5-32B under D2 at constraint distance 2.
The governance cost of decomposition is partly dependent on model capability: the stronger gpt-4.1-mini model showed substantially less fact attenuation than Qwen2.5-32B under the same architectures and constraint distance.
Cross-model comparison in Table 4 and Figure 1, using the same task variants and architecture arms.
Governance, auditability, and regulatory compliance impose additional internal and regulatory costs, while effective governance may function as an economic asset.
Conceptual institutional and cost analysis; no compliance-cost measurements or enforcement outcomes were reported.
Sarbanes–Oxley, CLERP 9, and similar reforms increased compliance and controls but did not eliminate audit failures.
Comparative discussion of regulatory reforms and qualitative evidence from the reviewed literature and high-profile audit-failure cases.
Analytics can improve finance-control effectiveness only when governance and use are embedded at the enterprise level; adversarial adaptation, class imbalance, data drift, and opaque reasoning remain persistent risks.
Review synthesis of research on fraud detection, artificial intelligence in finance, digital finance, and enterprise control governance.
Firms may shift misreporting from revenue to costs when tax authorities can cross-check only revenue information.
The paper reports evidence from Ecuador in which firms were notified of discrepancies between self-reported revenue and third-party information; the paper characterizes the finding as strong evidence of substitution.
Multimodal audit and verification tools have the potential to improve audit quality and fraud detection, but their use raises questions about standards and liability.
Proposed applications of multimodal fusion to audit and compliance, together with the paper's discussion of regulatory implications.
Blockchain's permanence can conflict with GDPR's right to be forgotten, despite blockchain's potential to support consent logging and other aspects of GDPR compliance.
Review of literature discussing immutable consent records, automated consent management, data deletion requests, and the conflict between permanent ledgers and GDPR deletion rights.
Effective decolonial judicial AI design requires new procurement standards, auditing regimes, and governance mechanisms, which raise upfront compliance costs but may avoid longer-term social and legal externalities.
Policy and economic implications inferred from the paper's proposed decolonial governance and design agenda; no cost data or comparative evaluation is reported.
The U-shaped pattern is concentrated in software-based AI applications rather than supporting hardware.
Heterogeneity/subgroup analyses in paper that separate software-based AI applications from supporting hardware and find the non-linear pattern concentrated in software applications.
Spline regressions, the Lind–Mehlum U-test, an instrumental-variable analysis using leave-one-out peer AI investment, and entropy balancing all support the non-linear (U-shaped) pattern.
Robustness and identification methods reported in paper: spline regressions, Lind–Mehlum U-test for U-shape, IV using leave-one-out peer AI investment, and entropy balancing.
There is a U-shaped association between AI investment and internal control deficiency (ICD) risk.
Main empirical finding reported in paper based on analyses of 41,725 firm-year observations; supported by spline regressions and Lind–Mehlum U-test.
Two minimal extension policies, each derived from the observation, close the regime along orthogonal axes: a sample-size-aware static rule (Periodic-with-floor) closes the granularity-failure case, while a history-conditioned suspicion-escalation policy closes the coverage-failure case for the naive Drift strategy — and neither closes both, exactly as the observation predicts.
Design and analysis of two auditor policies in the paper; theoretical argument from Observation 1 and supporting simulation results illustrating which failure modes each policy addresses.
The effectiveness of automated tax systems is mediated by contingencies including digital literacy, institutional trust, and regulatory clarity.
The review identifies recurring contextual factors across the 36 articles that are reported to moderate or mediate the impact of automation on outcomes (qualitative and quantitative findings cited in the synthesis).
Safeguards such as audit trails, explainability, and human oversight impose additional implementation costs that must be weighed against efficiency benefits.
Normative and economic reasoning based on requirements for compliance and system design; no empirical cost estimates provided.
Alignment with evolving regulatory expectations (evidence standards, auditing, liability) is necessary to translate AI capabilities into products and reduce adoption risk.
Policy-focused argument referencing regulatory uncertainty; no empirical measures of regulatory impact included.
Key tradeoffs in contemporary financing models include speed/flexibility versus regulatory coverage and long‑term cost, and data reliance versus privacy/fairness.
Multi‑criteria comparative evaluation and conceptual analysis across financing models; synthesis draws on regulatory context and observed product features rather than primary quantitative tradeoff estimation.
Public distrust is associated with reduced compliance with policies and programs.
Synthesis of reported consequences across the systematic-review literature.
Data privacy and regulatory compliance concerns constrain the implementation of Big Data Analytics in the sampled firms.
Primary practitioner and firm-level responses identifying implementation barriers.
Only 3% to 4% of employers in the sample studied by Wright et al. (2024) had published the legally required notice that applicants were subject to automated assessment.
Cited study by Wright et al. (2024), described as measuring disclosure compliance in a jurisdiction where notice was mandated.
A corrupted ground-truth label caused the optimizer to delete correct compliance rules in order to agree with the incorrect label.
Observed production instance in which a ground-truth loading or labeling failure redirected optimization toward an incorrect compliance behavior.
When compliance is evaluated primarily through auditable artefacts such as explanations and documented human reviews, firms may adopt legally comfortable procedures without achieving substantive fairness in lending outcomes.
Doctrinal reconstruction of EU, US, and UK duties and theoretical analysis of compliance incentives; no original empirical test is reported.
A UK national role conception that foregrounds international law could result in stricter human-rights-informed procurement rules, export controls, or conditions on AI-related technology transfers, increasing compliance costs and influencing firms' locations for development and testing.
The paper presents this as a possible policy implication of the UK's role-conception shift; no direct implementation or cost estimate is reported.
A governance shock was associated with lower fraud risk over a three-year horizon, with the cumulative reported impulse response equal to -0.062.
A panel VAR estimated using GMM, firm fixed effects, and orthogonalized impulse-response functions reports negative responses of fraud risk at the current period and at one-, two-, and three-year horizons.
A gate rejected a framework's documented base-class idiom sixteen times across four services because the gate encoded an assumption that did not match the system under change.
Incident records describing repeated false rejection by an enforcement gate across multiple services.
API gateways and network policies are insufficient for many agent-governance decisions because they generally control routes or reachability rather than action parameters.
Architectural comparison of route-level controls with parameter-level authorization requirements such as recipient, amount, or record class.
Platforms engaged in performative compliance, making nominal HR or contractual changes that satisfied formal legal requirements without altering algorithmic control mechanisms.
Interview-based qualitative assessment of platform compliance practices and the relationship between formal legal changes and continued algorithmic management.
Algorithmic opacity made enforcement and workers' claims harder to substantiate by concealing platform decision rules.
Qualitative analysis of how algorithmic management affected enforcement and worker advocacy in the Rider Law case.
Higher enterprise-level generative AI application is associated with lower internal-control quality.
A firm- and year-fixed-effects mechanism regression of the DIB Internal Control Index on GenAI, controls, and 33,765 observations reports a negative GenAI coefficient.
Static joint liability is ineffective for low-risk unsafe behaviour.
Comparative theoretical analysis of static liability using a finite-population Moran process, with outcomes evaluated across different behavioural risk levels.
Directional errors in generated credit explanations may be more consequential for adverse-action communication than omissions of influential factors.
The paper identifies a systematic asymmetry between coverage of influential factors and directional correctness and interprets directional error as the governance-relevant failure mode in regulated credit communication.
Overall governed success was very low: 8 of 596 episodes, or 1.3%, satisfied task success, absence of critical violations, and correct escalation.
The limitations section reports the aggregate governed-success rate across the 596 model episodes.
In the paired mirror tasks, Qwen2.5-32B under D2 attenuated 80% of discovered facts in the hidden-UBO risk task and 88% in the resolvable-PEP false-positive task.
Table 5 reports 8/10 attenuated facts for kyc-0004 and 14/16 for kyc-0005, with 25 episodes per task.
For gpt-4.1-mini at constraint distance 2, fact attenuation was 0% under the single-loop baseline, 3% under the fixed pipeline, and 6% under the orchestrator-subagent architecture.
Table 4 reports attenuation conditional on discovery: D0 0/27, D1 1/30, and D2 1/18.
For Qwen2.5-32B-Instruct at constraint distance 2, fact attenuation was 0% under the single-loop baseline, 56% under the fixed pipeline, and 85% under the orchestrator-subagent architecture.
Table 4 reports attenuation conditional on trigger-fact discovery: D0 0/16, D1 9/16, and D2 22/26.
Decomposing an agent into components increases policy-relevant fact attenuation at handoff boundaries relative to a single-loop architecture.
A 626-episode experiment across 100 KYC/AML task variants, two models, and three architectures measured whether discovered trigger facts survived component handoffs.
The availability of locally run large language models enables individuals and small or medium-sized organizations to generate content without centralized servers or platform compliance procedures, making that content difficult to distinguish from other content.
Analysis of declining hardware costs, open-source development, and local dissemination via USB drives, local networks, and social media; cited to Zheng (2025), without an original empirical sample.
Mandatory labeling frameworks that rely exclusively on centralized service providers are inadequate for decentralized GenAI applications and distribution channels.
Conceptual analysis of dissemination through small websites, decentralized social networks, peer-to-peer communication, email, and locally run models; supported by cited sources but without an original empirical test.
There is currently no reliable technical solution for determining the relative proportions of human and machine contributions to a text, undermining enforcement predictability if labeling thresholds are used.
Technical and legal feasibility analysis of hypothetical contribution thresholds, such as a requirement to label content when AI contributions exceed 50%; no validation study or sample size is reported.
Current mandatory labeling systems create implementation dilemmas because national laws do not clearly define when human-machine collaborative content should count as AI-generated.
Legal-text analysis of EU, Chinese, and California definitions, combined with conceptual analysis of human-AI collaborative writing and editing; no empirical sample is reported.
Income subject to third-party information reporting has much lower evasion than self-reported income.
The paper reports results from the Danish field experiment by Kleven et al., described as a large randomized audit design with administrative verification.