Evidence (1688 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
20058 claims
Filter claims →
Productivity
17184 claims
Filter claims →
Governance
16099 claims
Filter claims →
Human-AI Collaboration
16034 claims
Filter claims →
Innovation
10501 claims
Filter claims →
Org Design
10496 claims
Filter claims →
Labor Markets
6444 claims
Filter claims →
Skills & Training
5385 claims
Filter claims →
Inequality
4148 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 1820 | 479 | 278 | 1820 | 4588 |
| Organizational Efficiency | 2711 | 616 | 401 | 173 | 3922 |
| Governance & Regulation | 2075 | 886 | 459 | 246 | 3714 |
| Technology Adoption Rate | 1467 | 530 | 258 | 206 | 2488 |
| Decision Quality | 1281 | 496 | 289 | 152 | 2228 |
| Output Quality | 1227 | 447 | 207 | 138 | 2025 |
| AI Safety & Ethics | 634 | 754 | 207 | 83 | 1688 |
| Research Productivity | 826 | 241 | 114 | 422 | 1624 |
| Firm Productivity | 1052 | 154 | 163 | 66 | 1441 |
| Task Allocation | 685 | 211 | 331 | 99 | 1335 |
| Market Structure | 433 | 423 | 242 | 46 | 1150 |
| Innovation Output | 639 | 91 | 105 | 34 | 871 |
| Task Completion Time | 476 | 113 | 43 | 36 | 672 |
| Firm Revenue | 445 | 126 | 58 | 25 | 656 |
| Skill Acquisition | 364 | 119 | 109 | 34 | 626 |
| Consumer Welfare | 288 | 167 | 104 | 31 | 592 |
| Employment Level | 214 | 140 | 174 | 50 | 582 |
| Error Rate | 230 | 251 | 35 | 16 | 535 |
| Fiscal & Macroeconomic | 268 | 136 | 71 | 50 | 532 |
| Inequality Measures | 100 | 307 | 96 | 12 | 515 |
| Worker Satisfaction | 221 | 173 | 60 | 30 | 484 |
| Automation Exposure | 155 | 138 | 65 | 36 | 398 |
| Regulatory Compliance | 171 | 120 | 30 | 13 | 335 |
| Developer Productivity | 222 | 58 | 27 | 13 | 321 |
| Team Performance | 188 | 56 | 50 | 24 | 320 |
| Wages & Compensation | 146 | 104 | 46 | 16 | 312 |
| Training Effectiveness | 207 | 41 | 21 | 26 | 298 |
| Job Displacement | 23 | 153 | 52 | 4 | 232 |
| Hiring & Recruitment | 102 | 57 | 30 | 11 | 202 |
| Skill Obsolescence | 16 | 102 | 24 | 6 | 148 |
| Creative Output | 71 | 42 | 23 | 6 | 143 |
| Social Protection | 57 | 30 | 11 | 3 | 101 |
| Labor Share of Income | 29 | 42 | 24 | 2 | 97 |
| Worker Turnover | 43 | 29 | 6 | 4 | 82 |
| Industry | — | — | — | 1 | 1 |
Educational data should be treated as an economic asset with potential negative externalities, including privacy harms and surveillance, requiring data-governance and property-rights analysis.
Normative political-economy implication drawn from the essay's analysis of platformisation and datafication; this is a proposed analytical framework rather than an empirically estimated result.
Realizing a required distinction may require revealing additional information, creating a tradeoff between privacy and the ability to make the required judgment.
Conceptual analysis of realization and informational repair; illustrated across institutional and automated decision contexts.
Sales experts showed only partial agreement with the features identified by the model as important: some features with strong predictive power were regarded by experts as unintuitive or unimportant.
Five sales experts reviewed SHAP explanations for four correctly classified regional sales cases through a structured questionnaire, including ratings of feature importance and agreement with explanations.
Companies that deploy AI agents increase societal vulnerability to agentic-AI risks, while the same companies can also build societal resilience through design decisions that improve responses to failure.
The paper's explicit claim that deployment increases exposure, paired with its framework for coping and adaptive contributions.
The quality differences associated with RAMP maturity should not be interpreted as causal because maturity is observational and correlated engineering discipline or model capability may explain part of the observed gap.
The study uses observational repository maturity labels and a reanalysis of an existing panel rather than randomized assignment to maturity levels.
The persuasive attack showed the opposite directional pattern: it had a 53.3% ASR for BUY-targeting versus 41.9% for SELL-targeting.
Pooled micro-average attack success rates across five assets, reported in Table 1.
For non-persuasive attacks, BUY-targeted attacks were less effective than SELL-targeted attacks: data poisoning was lower by 2.7 percentage points, objective hijacking by 5.8 points, and indirect injection by 4.2 points.
Comparison of SELL- and BUY-targeted ASRs in Table 1, with the authors attributing the pattern to the system's bullish prior and differences in the set of attackable days.
Attack success did not increase monotonically with an attack's position or depth in the trading pipeline.
Comparison of attack success rates by targeted role: the Researcher attack outperformed the Trader attack despite occurring earlier in the pipeline, while analyst-level attacks had intermediate or lower success rates.
World models have achieved functional substitution for interaction and controllability in specific scenarios, but they remain short of traditional simulators in formal physical-law guarantees, structured state feedback, and reproducibility of long-horizon evolution.
Comparative capability analysis of 200 papers spanning latent-dynamics, autoregressive, diffusion, JEPA, explicit 3D/4D, and occupancy-centric methods.
The paper identifies reinforcement learning as highly capable for dynamic scheduling environments but particularly difficult to interpret and govern.
Narrative synthesis of the workforce-management literature, including cited applications of Q-learning, Deep Q-Networks, and policy-gradient methods.
The hardware verification chain is non-uniform: Lean kernel checking is used for specification-to-emitted-artifact correspondence, while Yosys SAT-based equivalence checks are used for Verilog correspondence and netlist equivalence, with a synthesis miter as an intermediate checker.
Description of four verification links and their respective checkers; the paper explicitly identifies the absence of a general public Verilog-to-Lean importer.
For Claude-Haiku-4.5 on 5G-Faults FT, GPT-5.5 disagreed with both Gemini judges more often, producing pairwise agreement rates of 0.860 with each, while the two Gemini judges agreed at 0.980.
Pairwise binary-grade agreement over 100 fault-diagnosis samples; Table V reports 0.860, 0.860, and 0.980.
LLMs support substantial task competence, abstraction, transfer, and in-context learning, while lacking human-like understanding, unified beliefs or intentions, and phenomenal consciousness.
Conceptual synthesis of technical and cognitive-science literature on LLM representations, task performance, and cognition.
LLMs can produce some verbatim regurgitation of training data, but treating all output as copied text ignores generative recombination and targeted memorization risks.
Diagnostic analysis of the training-data-memorization misconception and discussion of parametric memory as lossy compression.
Instruction tuning and RLHF directly update the model parameters and therefore reshape the same weights that encode knowledge and competence, rather than adding only a detachable safety filter.
Technical account of supervised instruction tuning and reinforcement learning from human feedback, supported by cited technical literature.
Deflationary and anthropomorphic folk theories each capture genuine features of LLMs but become misleading when treated as complete accounts of what LLM-based systems are.
Conceptual analysis of recurring slogans and framings across scientific, public, and institutional discourse, supported by cited literature.
Human capital analytics can improve retention and matching but also creates risks related to surveillance and biased decisions.
The paper's discussion of potential distributional effects and ethical risks associated with workplace analytics; no empirical risk estimate is reported.
Trust formation depends on characteristics of the trustor, attributes of the trustee, and contextual features.
Review synthesis of multifactorial explanations of trust formation across the included literature.
Credibility judgments about information are distinct from trust in people or institutions, and should be measured separately.
Conceptual distinction synthesized across the systematic review and its discussion of trust-related credibility measures.
RLHF, constitutional AI, system instructions, and content moderation primarily change model outputs and do not restructure the epistemic foundation established during pretraining.
Description of alignment and deployment techniques and the essay's comparison of their functional effects.
The trained policy retained its performance under prompt ablations: safe success was 97.63% with the full prompt, 97.63% without least-privilege wording, and 97.39% with a short one-line prompt.
Prompt-ablation study on a fixed 206-task routine evaluation set with 1,648 episodes per prompt-policy pair.
Conditional independence must hold after conditioning not only on the latent true label but also on observed features of the annotation task, because this allows shared sources of measurement difficulty to be accounted for.
The framework assumes proxy measurements are independent conditional on the true label and observed features such as text complexity, difficulty, or embeddings; the paper discusses this as the core identification assumption.
Natural-language instructions do not guarantee compliance; they shift probability toward compliant behavior, whereas grammar-expressible constraints can be enforced deterministically through constrained decoding.
This is presented as an architectural distinction between constrained decoding and natural-language configuration in the paper's axioms.
The evaluation demonstrates monitoring and mitigation of the studied private-channel attacks under controlled auction conditions, but it does not establish spontaneous emergence of a latent communication protocol.
The latent-collusion benchmark uses a fixed code optimized offline and explicitly primes the receiver to infer strategic intent; the authors describe the setting as a controlled, receiver-primed attack.
Rater heterogeneity in the human evaluation dominates differences between explanation systems.
Human-rating analysis comparing professional and non-professional evaluators of bimodal explanations.
Across three independently evolved lineages, SkillOpt's utility gain, exposure increase, and increase in unauthorized-state changes occur in all three lineages, but its total ASR increase occurs in only two of three lineages.
Lineage-level deltas relative to Static, with three complete end-to-end evolution repetitions.
For SkillOpt, conditional attack success after exposure decreases from 0.605 to 0.562, while aggregate attack success increases from 0.496 to 0.530.
Decomposition of attacked episodes into exposure and conditional attack success, using the frozen injection goals; 405 attacked cases per condition, with 400 SkillOpt cases after five provider-side exclusions.
The measured value of a prompt depends substantively on how the artifact is defined; semantically inert random padding can receive nearly the full description-length value even when it contributes nothing to the artifact's substantive content.
The paper gives constructions in which a random suffix is appended to a proof or other artifact and supplied by the prompt, causing the prompt to capture nearly the suffix's description length.
Faithful debugging context can sometimes mislead even strong models when the observed symptom is far from the underlying fault.
Per-case qualitative analysis of context-augmented outcomes, including faithful error outputs from a test case whose observation was distant from the fault.
Frontier LLM agents primarily operate at the C0 level of consciousness while exhibiting emerging C1-like capabilities.
The paper applies the C0-C1-C2 taxonomy of consciousness to the functional properties of frontier LLM agents, including global access to learned information.
The Opportunity Gate is designed to trigger ads for purchase planning or concrete commercial options, while abstaining in sensitive, purely informational, or disruptive contexts.
Description of the gate’s production rules and learned triggering model; this is a system-design claim rather than a measured outcome.
Safety concerns are a defining constraint on human–AI collaboration in manufacturing and shape the deployment, oversight, and acceptability of AI tools.
Experts identified hazardous machinery, complex shop-floor interactions, and regulatory requirements as factors shaping AI use; the study did not provide a quantitative safety estimate.
Perceived unfairness risk reduces managers' trust in AI, but trust does not necessarily improve managers' decision quality.
The paper summarizes a survey of 161 organizational decision-makers examining unfairness risk, trust in AI, and decision-making quality.
Generative AI applications in retail knowledge interfaces remain at an early stage of development and require governance because of hallucination, privacy leakage, and brand-inconsistency risks.
Qualitative review and synthesis matrix of the selected literature; no measured incidence or causal effect is reported.
Algorithms and creators jointly co-produce creators' digital identities, making digital traces strategically curatable but unstable and contingent on platform governance.
Qualitative and observational case examples from platformed creative work, synthesized with scholarship on digital identity and algorithmic governance.
DPGs with open or accessible datasets and standardized data pipelines may accelerate AI model training and local innovation, but they require governance to address privacy, bias, and data-reuse externalities.
Conceptual discussion of the relationship between DPG data infrastructure, AI innovation, and governance risks; no empirical effect estimate is reported.
Basic security failures still account for most successful intrusions, and current AI systems primarily scale the exploitation of existing weaknesses rather than creating access where cyber hygiene is sound.
The paper synthesizes observations from reported campaigns and cited cybersecurity analyses, including the continued role of unpatched systems, default credentials, and exposed services.
The study's validation exercise reproduced the error rates reported in the COMPAS fairness dispute, with the results depending on a reference population that the original analysis did not state.
A validation exercise replicated published COMPAS error rates and traced them to an unstated reference-population choice.
The identity of the group classified as disadvantaged reversed across defensible fairness specifications.
Study 1 compared fairness outcomes across alternative metric families and other defensible audit specifications while holding the underlying systems and data fixed.
Across 91,572 defensible specifications of four AI decision systems, the regulatory four-fifths fairness verdict flipped in all seven system–protected-attribute pairings examined.
Study 1 enumerated the full multiverse of defensible metric, threshold, reference-population, and comparison-group specifications for four decision systems.
The modeled probability of a common-mode failure increases with edge-case pressure and does not depend on the character-shaping allocation α.
Common-mode failure specification q(M) = (1 − e−βq A(M)) · e0, with M determined by edge-case pressure, plus the stated independence from α.
In the model, character fragility is independent of deployment scale at the per-interaction level; increasing deployment scale raises the aggregate number of fragility manifestations rather than the per-interaction fragility rate.
Model specification pfrag(α) = p(0)frag αn and the paper's distinction between per-interaction probability and aggregate event count.
The baseline character fragility rate p(0)frag is the dominant determinant of the optimal safety-design allocation, shifting α* by 0.50 across its examined range and having a substantially larger effect than tail severity, filter quality, or common-mode failure probability.
Parameter-sensitivity analysis within the analytical model and simulation scenarios.
Causal identification, interpretability, robustness checks, privacy safeguards, and defenses against adversarial manipulation are necessary for multimodal accounting research to produce economically valid findings.
Surveyed methodological challenges and recommendations, including causal designs, human-in-the-loop validation, privacy-preserving methods, and robustness analysis.
The review introduces Algorithmic Reciprocity, Conditional Reciprocity, and Multi-Dimensional Algorithmic Justice as constructs for analyzing platform-based employment relationships.
Conceptual constructs developed through thematic synthesis of 39 empirical studies.
The paper cannot determine from available evidence whether the absence of Indian results on agentic coding benchmarks reflects a genuine capability gap or only a reporting gap.
The study uses only publicly reported benchmark results, does not run evaluations, and treats unreported scores as ambiguous between non-evaluation, nondisclosure, strategic exclusion, or inapplicability.
A trait can improve individual survival while simultaneously increasing vulnerability to extinction at the species level.
The paper contrasts evidence that larger brains can improve individual survival with evidence that they impose reproductive and developmental costs that increase species-level vulnerability.
The persistence of long-lived non-intelligent lineages demonstrates that general intelligence is not necessary for extraordinary evolutionary success, although it does not by itself prove that general intelligence is maladaptive.
Cross-taxon comparison of lineage durations and a stated qualification that the comparison establishes non-necessity rather than direct maladaptiveness.
AI regulations requiring validation, transparency, and audits could improve model quality and social welfare while increasing model-development costs and time to market.
Policy trade-off analysis; the paper does not provide empirical estimates of quality gains, welfare effects, costs, or delays.
The paper's LLM-usage estimates are subject to uncertainty because they assume that linear extrapolation from 2018–2022 provides a faithful estimate of human word frequencies in 2025.
Authors' stated limitation: post-2022 human vocabulary may change through imitation of, or deliberate avoidance of, LLM-associated vocabulary.