Evidence (648 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
21267 claims
Filter claims →
Productivity
17978 claims
Filter claims →
Governance
17038 claims
Filter claims →
Human-AI Collaboration
16914 claims
Filter claims →
Org Design
11104 claims
Filter claims →
Innovation
11087 claims
Filter claims →
Labor Markets
6711 claims
Filter claims →
Skills & Training
5616 claims
Filter claims →
Inequality
4343 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 1880 | 496 | 296 | 1854 | 4721 |
| Organizational Efficiency | 2906 | 665 | 438 | 180 | 4210 |
| Governance & Regulation | 2162 | 929 | 480 | 247 | 3866 |
| Technology Adoption Rate | 1533 | 545 | 278 | 210 | 2593 |
| Decision Quality | 1391 | 534 | 321 | 173 | 2429 |
| Output Quality | 1298 | 472 | 231 | 145 | 2153 |
| AI Safety & Ethics | 682 | 821 | 230 | 90 | 1837 |
| Research Productivity | 855 | 253 | 121 | 425 | 1675 |
| Firm Productivity | 1105 | 171 | 175 | 73 | 1531 |
| Task Allocation | 735 | 229 | 361 | 99 | 1433 |
| Market Structure | 457 | 461 | 251 | 47 | 1222 |
| Innovation Output | 673 | 94 | 108 | 36 | 913 |
| Task Completion Time | 499 | 118 | 43 | 38 | 702 |
| Firm Revenue | 458 | 130 | 61 | 26 | 677 |
| Skill Acquisition | 381 | 122 | 113 | 34 | 650 |
| Consumer Welfare | 316 | 176 | 115 | 39 | 648 |
| Employment Level | 223 | 143 | 177 | 53 | 600 |
| Error Rate | 246 | 282 | 44 | 19 | 594 |
| Fiscal & Macroeconomic | 283 | 142 | 78 | 52 | 562 |
| Inequality Measures | 103 | 329 | 106 | 13 | 552 |
| Worker Satisfaction | 225 | 185 | 63 | 30 | 503 |
| Automation Exposure | 158 | 155 | 72 | 37 | 426 |
| Regulatory Compliance | 186 | 126 | 35 | 14 | 362 |
| Team Performance | 193 | 56 | 51 | 24 | 326 |
| Developer Productivity | 224 | 58 | 27 | 13 | 323 |
| Wages & Compensation | 148 | 108 | 50 | 17 | 323 |
| Training Effectiveness | 218 | 44 | 21 | 27 | 313 |
| Job Displacement | 23 | 159 | 53 | 5 | 240 |
| Hiring & Recruitment | 109 | 61 | 32 | 11 | 215 |
| Skill Obsolescence | 16 | 107 | 26 | 6 | 155 |
| Creative Output | 71 | 44 | 28 | 6 | 150 |
| Social Protection | 58 | 31 | 12 | 3 | 104 |
| Labor Share of Income | 29 | 43 | 25 | 2 | 99 |
| Worker Turnover | 45 | 29 | 6 | 4 | 84 |
| Industry | — | — | — | 1 | 1 |
Technology-based information services can improve passenger satisfaction, while human-provided services remain important when passengers need reassurance or exception handling.
The paper cites aviation self-service and passenger-experience studies, including Choi et al. (2024), and summarizes their findings.
Passenger outcomes from digital aviation services vary according to digital readiness, perceived control, privacy sensitivity, language proficiency, travel familiarity, journey complexity, and accessibility needs.
The paper synthesizes passenger-segmentation and smart-airport research, including Halpern, Mwesiumo, Budd, et al. (2021) and Rubio-Andrada et al. (2023).
The null behavioral findings do not establish that AI assistance has no downstream spending effects; the observed behavioral estimates were too imprecise to rule out medium-to-large effects.
The authors compared behavioral point estimates with the study's minimum detectable effect and explicitly characterized the results as quantitative bounds rather than evidence of absence.
AI-assisted books increased from near zero in 2022 to more than half of new releases in 2025; new Amazon releases approximately tripled, average quality fell, and estimated consumer surplus increased by about 7%.
The paper reports market-level observations and estimates concerning AI-assisted book prevalence, Amazon release volume, average quality, and reader consumer surplus.
Rule-based pricing remains highly competitive with reinforcement-learning pricing when the two families are compared directly.
Direct comparison of bill-sharing, MMR, and SDR benchmarks with RL pricing policies in the PV-only configuration.
The interaction between algorithmic transparency and perceived control was significant under cultural-context moderation, while the main structural paths did not differ by cultural context.
Multi-group analysis comparing consumers from Brazil and Indonesia.
Reasoning models increased token consumption faster than token prices declined, causing seller prices per token and buyer prices per completed task to diverge.
The paper compares token-denominated seller prices with task-denominated buyer prices using observed per-task token consumption.
When the guarded and unguarded conditions use a unified offer schema and buyer chooser, the Both-minus-None welfare contrasts change to +7.2 for 1.5B, −13.9 for 3B, and +23.8 for 14B.
Controlled E1 rerun with one schema and one chooser in all treatment cells; 30 profiles per model, with profile-bootstrap confidence intervals.
A year-long viewer ablation produces 1.74% more video views but 2.13% less view time relative to disabling exploration.
A year-long viewer-side ablation experiment comparing the exploration mechanism with disablement.
Algorithmic dynamic pricing can preserve affordability for mass segments while extracting surplus from buyers with higher willingness to pay, creating distributional and consumer-surplus effects.
Conceptual analysis of dynamic pricing, versioning, and time-limited offers; the paper reports no pricing data, welfare estimates, or causal identification.
Technology may lower access costs and democratize investment through robo-advice, while also increasing informational asymmetries between sophisticated algorithmic traders and retail investors.
Conceptual implications drawn by the review; no causal welfare estimate is reported.
High-quality journalism has public-good characteristics, while algorithmic amplification of low-quality content can generate negative externalities.
Normative and economic framing in the discussion of externalities and public goods; no welfare estimate or empirical comparison is provided.
Editorial choices and algorithmic amplification jointly shape what audiences see, while editorial standards, algorithmic signals, and regulatory legitimacy jointly shape what audiences trust.
Conceptual decomposition of the framework into visibility and credibility mechanisms; not supported by an empirical sample.
Country-level perceived corruption functions as a contextual moderator of the relationship between experience–expectation disconfirmation and review sentiment.
The study integrates expectancy–disconfirmation and institutional trust theories and applies statistical moderation tests to Booking.com review sentiment.
The broader directional relationship between home-country corruption, experience–expectation discrepancies, and review sentiment is reproduced in a separate Dubai sample, supporting generalizability across cities.
Robustness analysis using 242,541 Booking.com hotel reviews from Dubai and comparing the findings with the primary New York City sample.
Tourists from low-corruption countries show stronger emotional responses to discrepancies between their hotel experience and expectations, with sentiment amplified in both positive and negative directions.
Moderation tests examining whether home-country perceived corruption conditions the relationship between experience–expectation disconfirmation and textual sentiment in the New York City sample of 56,260 reviews.
Scaling a model parameter in a given direction increases welfare for agents whose preferences align with that direction across deployment queries and decreases welfare for agents whose preferences covary negatively with it.
The paper derives the directional derivative of individual welfare with respect to parameter scaling and interprets the sign of the resulting covariance-like expression.
Changing reasoning effort affects search depth in opposite directions across the two tested models: inspections fell from 3.12 to 1.91 per session for Gemini Flash Lite but rose from 0.39 to 5.83 for Gemini Pro.
Follow-up reasoning-effort experiment covering four effort levels for Gemini Flash Lite and three for Gemini Pro, with 500 replications per cell.
The effect of listing position on final hotel choice is heterogeneous across models: it is statistically indistinguishable from zero for Gemini Flash and Gemini Pro, but negative and statistically significant for Claude Sonnet and Gemini Flash Lite.
Linear probability models of final choice using 500 sessions per model. Reported choice-position coefficients were −0.000014 (p=0.338) for Flash, −0.000008 (p=0.606) for Pro, −0.000035 (p=0.018) for Sonnet, and −0.000080 (p<0.001) for Flash Lite.
For most tested LLMs, inspection probability is non-monotonic in position: inspection declines from the top of the page to a minimum around ranks 68–74 and then rises toward the bottom.
Models including both linear and quadratic position terms. The quadratic term was positive and significant for three of four LLMs, producing a U-shaped inspection curve.
Avoidance of binding duties trades off lower short-term political and administrative costs for platforms against potential long-term welfare losses from unmitigated harms to children.
Normative welfare analysis of platform compliance costs, enforcement design, and child-harm externalities; the paper provides no quantified welfare estimate.
Improved fraud detection and triage can increase market transparency and investor confidence, while uneven adoption may create cross-jurisdictional arbitrage opportunities for fraudsters.
Discussion of market-level effects of AI-supported financial-fraud detection and differences in adoption across jurisdictions.
Personalization may increase consumer value through more efficient matching, but tailored price extraction, reduced autonomy and privacy, and over-personalization can reduce consumer surplus or create negative externalities.
Conceptual welfare analysis identifying both matching benefits and harms from price extraction, privacy loss, reduced autonomy, ad fatigue, and reduced discovery; no consumer experiment or welfare estimate is reported.
Finer-grained personalization can enable more precise menu and individualized pricing, potentially increasing firm markups while creating distributional and consumer-welfare concerns.
Conceptual economic analysis of the pricing implications of personalized targeting; no pricing data, market sample, or causal estimate is reported.
In the model, an arbitrarily large GDP is welfare-relevant to humans only through their financial ownership share of the corporate machine network.
Theoretical welfare analysis sets human production and consumption to zero and summarizes the human stake using the ownership-share state variable ε_t.
Equalizing algorithmic accuracy does not necessarily equalize who benefits from AI, because deployment-facing regulation can narrow measured performance gaps while inducing unequal equilibrium use.
The paper's integrated model of upstream algorithm design and downstream physician adoption, in which liability changes both algorithmic accuracy choices and group-specific physician reliance.
The paper identifies a potential welfare tradeoff in AI physiotherapy: lower unit costs and broader access may be accompanied by losses in empathy, individualized judgment, and complex reasoning.
Normative welfare analysis; the paper explicitly notes that clinical outcomes, cost-effectiveness, and distributional effects require future empirical evaluation.
At low prior beliefs, transparency reduces consumer welfare whenever search costs are prohibitive; when search is viable, its welfare effect may be positive or negative depending on search and information frictions.
Proposition 3 analytically characterizes welfare effects across prior beliefs, search costs, and information frictions.
On the typical heating scenario, SAC reduced thermal discomfort by 33.1% relative to the baseline, but increased operating cost from €0.413 to €0.631.
Closed-loop BOPTEST evaluation over a 14-day typical heating scenario. Table 2 reports baseline discomfort of 9.446 K h and cost of €0.413, compared with SAC discomfort of 6.322 K h and cost of €0.631.
On the peak heating scenario, SAC reduced thermal discomfort by 90.7% relative to the baseline while increasing operating cost by 11.5%.
Closed-loop evaluation on the BOPTEST emulator over a 14-day peak heating scenario. Baseline discomfort was 8.382 K h and cost was €0.909; SAC discomfort was 0.777 K h and cost was €1.013.
Total surplus declines persistently during the learning transition only in the high-price-stickiness regime.
Numerical comparison of total surplus across low- and high-price-stickiness dynamic Cournot markets.
Consumers tend to share more data for personalized services in health and financial contexts than for retail promotions, which many perceive as intrusive.
Review synthesis citing Gao et al. (2023); the supplied text does not report the underlying study's sample size or quantitative effect.
The reviewed literature reports that younger consumers generally prioritize convenience and share data for personalized services, whereas older consumers tend to be more cautious about data sharing.
Review discussion citing Mu and Zhang (2025) and related consumer-behavior literature; no pooled estimate or sample size for this age comparison is reported in the supplied text.
Consumers face a trade-off between the convenience and benefits of personalized marketing and concerns about privacy and possible misuse of personal data.
Review synthesis of literature on the personalization-privacy paradox, including studies identifying trust and perceived sacrifice as mediating factors.
Personalized targeting and automation may increase short-run firm profitability and consumer surplus from relevance while also creating manipulation and autonomy-related externalities that reduce welfare in unmeasured ways.
Economic interpretation in the implications section; no quantitative estimate or causal design is reported.
The benefits of generative AI in information environments coexist with harms: it can improve content access, creativity, and potential civic tools while fragmenting shared facts, reducing trust, and impairing democratic discourse.
Interpretivist synthesis of qualitative literature, policy documents, governance reports, and case studies.
Cultural and regulatory contexts significantly influence consumer responses to virtual influencers, implying that virtual influencer strategies should be adapted to different market environments.
Findings summarized from Khalfallah and Keller's (2025) systematic review of 51 journal articles.
In the product-turnover dynamic environment, the model permits an exact period-three cycle in footprint, fee, and welfare and also permits chaotic equilibrium paths.
Dynamic theoretical extension in which active product categories can turn over and the footprint can contract; the paper states that it characterizes a period-three window and establishes existence of chaotic paths.
AI-driven underwriting may expand startup financing, but data-driven credit scoring can embed bias and opacity.
Qualitative synthesis of literature on fintech, data-driven credit scoring, AI underwriting, and startup financing.
The marginal effect of AI accuracy on the value of information is inverted-U-shaped: moderately accurate AI amplifies information frictions, whereas sufficiently accurate AI helps mitigate them.
Comparative-statics analysis of the welfare value of revealing physician quality across AI-accuracy levels.
Asymmetric information about physician quality does not universally reduce social welfare; welfare loss occurs only when standard care is unreliable and AI accuracy is sufficiently low.
Comparison of social welfare under full information and asymmetric information across parameter regimes in the principal-agent model.
Prompts mentioning macroeconomic uncertainty cause the LLM to recommend more saving and lower equity holdings.
The authors compare advice responses to prompts that do and do not mention macroeconomic uncertainty within their survey-and-simulation framework.
For p < x, optimal persuasive signaling increases chatbot engagement from zero under no persuasion to a positive equilibrium level, while the user's expected payoff remains equal to the no-persuasion payoff.
Proposition 1 compares the no-persuasion and optimal-signaling equilibria in the baseline analytical model.
In post-AGI conditions, market price and financial return should be treated as partial signals of value rather than as definitions of value.
Conceptual argument that prices may reflect residual scarcity, property rights, and market power without capturing total social value, while output and financial returns may diverge from agency, health, trust, ecological stability, and distributional outcomes.
The evaluation uses a fixed $50 budget over 10,000 rounds, and each policy is evaluated over 30 paired seeds.
This is the stated experimental design in the benchmark setup.
PA-DCT maintains competitive service quality while spending only 39–43% of the available wallet across the three evaluated market scenarios.
The claim is based on the 402Pilot-Bench evaluation with a fixed $50 budget, 10,000 rounds, three scenarios, and 30 paired seeds per policy.
The paper states that increases in income beyond moderate thresholds yield diminishing returns to subjective well-being.
Cited empirical literature, specifically Easterlin (1974) and Kahneman and Deaton (2010); the paper does not present a new estimate.
Temporal advertising intrusiveness moderates the relationship between advertising-supported monetization and search-query engagement, whereas visual intrusiveness does not have a significant moderating effect.
The paper reports experimental tests of visual ad format and temporal ad length across four experiments with N = 1063.
SEO strategies that prioritize attention capture over relevance or exploit ranking vulnerabilities can reduce consumer welfare through lower information quality, whereas semantic and intent-focused SEO can improve matching between consumers and information.
Conceptual welfare analysis in the review contrasting manipulative SEO practices with relevance- and intent-oriented optimization; no quantified consumer-welfare estimate is reported.
Social-media platforms achieved substantial functional gains, including expanded access, lower barriers to speech and organization, and greater distribution capacity, while also experiencing a trust and legitimacy crisis that generated prospects of audience withdrawal and increased regulation.
The paper's synthesis of prior analysis of platforms, cited as Abiri and Guidi (2023), together with media-law scholarship treating platforms as governance institutions.