Evidence (2228 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
20058 claims
Filter claims →
Productivity
17184 claims
Filter claims →
Governance
16099 claims
Filter claims →
Human-AI Collaboration
16034 claims
Filter claims →
Innovation
10501 claims
Filter claims →
Org Design
10496 claims
Filter claims →
Labor Markets
6444 claims
Filter claims →
Skills & Training
5385 claims
Filter claims →
Inequality
4148 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 1820 | 479 | 278 | 1820 | 4588 |
| Organizational Efficiency | 2711 | 616 | 401 | 173 | 3922 |
| Governance & Regulation | 2075 | 886 | 459 | 246 | 3714 |
| Technology Adoption Rate | 1467 | 530 | 258 | 206 | 2488 |
| Decision Quality | 1281 | 496 | 289 | 152 | 2228 |
| Output Quality | 1227 | 447 | 207 | 138 | 2025 |
| AI Safety & Ethics | 634 | 754 | 207 | 83 | 1688 |
| Research Productivity | 826 | 241 | 114 | 422 | 1624 |
| Firm Productivity | 1052 | 154 | 163 | 66 | 1441 |
| Task Allocation | 685 | 211 | 331 | 99 | 1335 |
| Market Structure | 433 | 423 | 242 | 46 | 1150 |
| Innovation Output | 639 | 91 | 105 | 34 | 871 |
| Task Completion Time | 476 | 113 | 43 | 36 | 672 |
| Firm Revenue | 445 | 126 | 58 | 25 | 656 |
| Skill Acquisition | 364 | 119 | 109 | 34 | 626 |
| Consumer Welfare | 288 | 167 | 104 | 31 | 592 |
| Employment Level | 214 | 140 | 174 | 50 | 582 |
| Error Rate | 230 | 251 | 35 | 16 | 535 |
| Fiscal & Macroeconomic | 268 | 136 | 71 | 50 | 532 |
| Inequality Measures | 100 | 307 | 96 | 12 | 515 |
| Worker Satisfaction | 221 | 173 | 60 | 30 | 484 |
| Automation Exposure | 155 | 138 | 65 | 36 | 398 |
| Regulatory Compliance | 171 | 120 | 30 | 13 | 335 |
| Developer Productivity | 222 | 58 | 27 | 13 | 321 |
| Team Performance | 188 | 56 | 50 | 24 | 320 |
| Wages & Compensation | 146 | 104 | 46 | 16 | 312 |
| Training Effectiveness | 207 | 41 | 21 | 26 | 298 |
| Job Displacement | 23 | 153 | 52 | 4 | 232 |
| Hiring & Recruitment | 102 | 57 | 30 | 11 | 202 |
| Skill Obsolescence | 16 | 102 | 24 | 6 | 148 |
| Creative Output | 71 | 42 | 23 | 6 | 143 |
| Social Protection | 57 | 30 | 11 | 3 | 101 |
| Labor Share of Income | 29 | 42 | 24 | 2 | 97 |
| Worker Turnover | 43 | 29 | 6 | 4 | 82 |
| Industry | — | — | — | 1 | 1 |
AI-enabled recruitment is associated with a shift toward data-driven decision-making in which machine intelligence complements human expertise.
The paper's literature synthesis and qualitative examination of AI-powered recruitment practices, including chatbots, video interviews, targeted job advertisements, and predictive analytics.
Ichnology is a useful testbed for formalizing inference because trace evidence is highly ambiguous: one organism can produce diverse traces, while similar traces can result from different organisms or abiotic processes.
Conceptual analysis of ambiguity in trace-based interpretation and comparison of biotic and abiotic explanations.
For middle managers, AI has both positive and negative effects: it supports data analysis and managerial decision-making while creating concerns about automation of some managerial responsibilities.
Cross-study synthesis of findings differentiated by organizational level.
Static models linking stable managerial traits to stable decisions are insufficient because the effects of GM characteristics depend on dynamic competence, situational expression of values, leadership adaptability, and recognition of gendered traits.
Thematic synthesis of studies on GM characteristics and leadership styles, used to qualify the dispositional assumption.
The effect of citations on false-information detection depends on both users' digital literacy and the type of falsehood being evaluated.
The experiment estimated heterogeneous treatment effects using interaction terms for digital literacy and separately analyzed misinformation and disinformation outcomes.
The commitment effect is concentrated among models rather than universal: three of the 12 models were strongly seduced by the panels, four never committed under any panel, three committed regardless of the panel, and two responded weakly.
Disaggregated analysis of commitment behavior across the 12-model frontier-model roster.
nDCG@K also produced a borderline finding in the example audit, despite the score-delta, top-K-retention, and merit-aware rate-gap metrics remaining within tolerance.
The paper compares ranking-quality results with score, retention, and merit-aware fairness metrics.
Mean absolute rank change produced borderline audit findings in the example corpus, including a finding on the neutral baseline configuration.
The paper reports the results of its rank-stability metric in the illustrative audit.
Higher model capability postpones but does not eliminate the retrieval–integration gap: the largest open-weight model retained a 3.4-percentage-point effect at 128,000 tokens, while a lower-capability commercial system eventually lost retrieval as well as decision influence.
Comparative experiments across models with different capability levels and context lengths.
The retrieval–integration gap replicated across three independently trained model families, although the context length at which the gap became binding differed across models.
Replication experiments conducted across three independently trained model families.
The system's bullish prior was associated with BUY being the clean prediction on 40.0% of days, compared with 26.2% for SELL.
Reported clean-system decision frequencies used to explain why some target directions were more difficult to induce.
In an agentic retrieval setting, changing the dial affects the information searched for, the evidence selected, and the evidence reflected in the final analysis.
Additional agentic retrieval evaluation examining information acquisition, evidence selection, and final generated analyses under different dial settings.
Under identical balanced bullish and bearish evidence, changing Qwen3-8B to its calibrated neutral setting changed both JPMorgan Chase's and NVIDIA's decisions from buy to sell and shifted the rationales toward greater emphasis on downside risks.
Paired response-level comparison for two securities under the same evidence, before and after applying the calibrated neuron intervention.
The direction of the dial's effect is model-specific: increasing the intervention coefficient shifts the investment-bias score in the opposite direction for DeepSeek-R1-14B and Mistral-24B compared with the other evaluated models.
Cross-model response curves relating intervention strength to the aggregate investment-bias score.
The reachable range of investment-bias scores is model-dependent: four models approach the full interval from -1 to 1, whereas DeepSeek-R1-14B has a narrower response range.
Evaluation of the investment-bias score over intervention-strength values for each of the five models.
AI lowers the cost of generating and comparing alternative operationalizations of a concept, but it cannot determine which operationalization answers the research question; that judgment requires theory and substantive expertise.
Review's discussion of construct definition, competing rubrics, nomological networks, and the role of domain expertise.
Different measurement functions applied to the same underlying concept can preserve different features of the data and therefore support different empirical conclusions.
Formal discussion of measurement as a projection from high-dimensional reality into a lower-dimensional representation, combined with the existence of multiple plausible operationalizations.
The availability of AI shifts the bottleneck in empirical measurement from finding any scalable measure of a phenomenon to choosing among many plausible measures.
Authors' synthesis of the reduced cost of applying AI measurement functions and the increased number of choices involving rubrics, models, prompts, training data, and tuning strategies.
Predictive accuracy alone does not necessarily produce business impact or supply-chain resilience; predictions must also be timely, calibrated, interpretable, and connected to decision rights and feasible response options.
Conceptual argument supported by distinctions among predictive intelligence, decision actionability, and resilience orchestration; the paper cites prior analytics literature but reports no original test.
The observed pattern of lower accuracy on modern statutory updates is consistent with the paper’s hypothesis of ‘precedent overfitting,’ but the study does not establish that precedent overfitting is the causal mechanism.
Comparison of model performance across historical doctrinal control categories and cases involving the 2018 amendments; the models were evaluated as black boxes without access to weights, attention mechanisms, or retrieval rankings.
Fact payloads carried the gold evidence in 98–99% of cases, while chunk reading accuracy declined from 81% to 73% as more text was supplied.
Decomposition of the held-out question-answering results into evidence coverage and reading accuracy across token budgets.
The same systematic review found that human judgment remained decisive under high uncertainty, while several included studies identified bias and diminished trust as unresolved concerns even when efficiency gains were present.
López-Solís et al. (2025) systematic review of 30 studies, as summarized by the paper.
A technically accurate forecast can have less business value when delivered after planning decisions are fixed than a somewhat less accurate forecast embedded in a responsive decision process.
Conceptual inference in the review based on the relationship between decision latency, process integration, and business value; no quantitative test is reported in the paper.
Algorithmic recruitment tools are not inherently more biased than the human-mediated processes they replace; their effects on equity depend substantially on system design, training data, and auditing practices.
Review of countervailing evidence comparing structured or algorithmic screening with unstructured human interviews, citing Bogen and Rieke (2018) and Stone et al. (2024).
Managers in the OECD research often associated algorithmic management with improved decision quality and efficiency, but also reported concerns about unclear accountability, opaque algorithmic logic and effects on workers.
Findings from the OECD survey of more than 6,000 firms in six countries; the paper reports both perceived benefits and governance or worker-related concerns.
Persona information was most beneficial when habitual travel information was unavailable.
Factorial comparison of prompting configurations with and without habitual travel information and with persona information.
For Qwen2.5-7B, judge scores showed above-chance per-step sign agreement with replay contribution, but did not identify or concentrate on pivotal steps better than the shuffled-control benchmark.
Judge sign agreement was 84/139=60.4%, with a 95% interval excluding 50%; however, judge precision-at-pivotal lift was 1.000 with an interval containing the chance value of 1.0, and rank fidelity was indistinguishable from its own shuffle.
Calibration regret and behavioral overreliance can diverge: high-quality AI on hard tasks produced the highest regret despite not producing the highest overreliance.
Table 3 reports final-state outcomes for four combinations of AI quality and task difficulty, averaged over 20 seeds.
A naive decision rule that classified a 1/5 pass rate as scattered and non-zeroable was false; inspecting failure causes showed that a shared cause could still be corrected completely by one rule.
The reported replication contradicted the earlier pass-rate-only classification; manual inspection found homogeneous failures, and the rule raised the rate to 5/5.
In solution-selection tasks, models often select incorrect options even when valid solutions are present; NOTA improves average accuracy but models rarely use it, including when no valid option is present.
Verification-only selection experiments with a 'none of the above' option across solution-concept tasks.
Multiple-choice prompting can improve indeterminacy detection for some models while inducing false indeterminacy judgments for others.
Partial-order query comparison across prompting formats and models.
The study evaluates three explanation pipelines: a tabular pipeline using XGBoost with SHAP, a network pipeline using a graph neural network with GNNExplainer, and a bimodal pipeline combining tabular and network evidence.
Experimental design using Freddie Mac single-family loan-level data and separate prediction, post-hoc explanation, and LLM narrative-generation stages.
LLM-generated credit-risk explanations reliably identify influential factors but are less reliable at correctly stating the direction of those factors' influence.
Automated checks of generated narratives against the underlying SHAP and GNNExplainer evidence across the tabular, network, and bimodal pipelines.
A LoRA-recovered compressed variant can remain fully parseable and achieve 71% strict BoolQ accuracy while sending 97 of 100 predictions to a single class, resulting in 52.6% balanced accuracy.
BoolQ evaluation of compressed and LoRA-recovered language-model variants, including prediction-distribution analysis over 100 predictions.
Among the two evaluated theory-of-mind skills, next-action prediction predicts negotiation outcomes, whereas preference inference alone does not.
Analysis relating the two ToM components—preference inference and next-action prediction—to negotiation outcomes.
Path-based random-utility models provide interpretable marginal utilities, elasticities, values of time, and welfare analysis, but their use can be limited by enormous path spaces and unknown or misspecified choice-set generation.
Review of path-based random-utility models, their estimands, and path-enumeration and choice-set limitations.
Machine-learning capabilities such as occupancy-ratio methods, graph and sequence models, constrained learning, transfer, data fusion, and multi-agent learning do not by themselves establish preference recovery or valid policy counterfactuals.
Critical synthesis of method families and their evidentiary limits in transportation applications.
Offline reinforcement learning based on logged behavior is prescriptive rather than automatically descriptive of the behavior-generating process.
Conceptual distinction in the review between improving a policy from logged data and recovering descriptive behavioral preferences.
Imitation learning can reproduce observed behavior without explaining the underlying preferences or behavioral mechanism.
The review distinguishes policy or occupancy matching from structural recovery of utility or reward.
The equivalence between recursive choice, dynamic discrete choice, and maximum-entropy IRL does not by itself establish that the models identify the same behavioral quantities.
The review distinguishes utility, reward, policy, occupancy, constraints, and observation error as different estimands with different interpretations and counterfactual implications.
Centaur, a model fine-tuned on more than ten million human choices across hundreds of experiments, predicts human behavior on held-out tasks better than bespoke models but remains only weakly equivalent to human cognition.
The paper summarizes the Centaur model's training and evaluation results and argues that predictive success does not establish shared underlying mechanisms.
GPT-5 achieved the highest screening recall, at 91.8%, but had lower specificity.
Comparative screening evaluation against expert-annotated inclusion and exclusion decisions for the 244-document benchmark subset.
The PaperFindingBench evaluation combines exact-match scoring for approximately 27% of queries with an LLM judge scoring agent-supplied evidence for the remaining approximately 73%.
Benchmark scoring protocol described in Table 1 and the PaperFindingBench setup.
Across the cited research, the effect of human-in-the-loop oversight on decision quality is uneven.
The author synthesizes studies reporting limited effects on discrimination, mixed results in child welfare, and cases where humans improve technologically mediated processes.
Human intervention against malfunctioning algorithmic risk-prediction tools in child-welfare decisions has limited effectiveness, with results that are mixed.
The chapter's summary of De-Arteaga, Fogliato, and Chouldechova's study of human oversight in child-welfare decision-making.
AI-generated predictions do not eliminate the need for human judgment, especially for economic decisions involving ethical considerations, strategic priorities, and institutional constraints.
Conceptual synthesis drawing on the distinction between prediction and judgment and on decision-theory literature; no empirical sample is reported.
The effects of AI use depend on system design, use intensity and frequency, user autonomy, relational context, and the time horizon over which outcomes are evaluated.
The paper identifies these variables as determinants in its dynamic framework.
For finance forecasting and treasury, a modest improvement in forecast accuracy can create substantial value when it occurs near a liquidity threshold, whereas a larger statistical improvement may be irrelevant if it does not change a decision.
Decision-oriented synthesis of forecasting and treasury literature, citing decision-focused evaluation of predictive information.
Counterfactual evaluation is necessary in collections analytics because observed payments may have occurred naturally and should not automatically be attributed to the predictive intervention.
Methodological recommendation in the order-to-cash discussion concerning intervention evaluation and suitable baselines.
Decision quality and process value should be evaluated separately from predictive accuracy because a technically accurate model may fail to change actions or improve financial, operational, control, or service outcomes.
Conceptual framework distinguishing predictive quality, decision quality, and process value; supported by examples involving collections and journal-anomaly detection.