Evidence (7560 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
9875 claims
Filter claims →
Productivity
8807 claims
Filter claims →
Governance
7870 claims
Filter claims →
Human-AI Collaboration
7560 claims
Filtered →
Org Design
4892 claims
Filter claims →
Innovation
4781 claims
Filter claims →
Labor Markets
4004 claims
Filter claims →
Skills & Training
3308 claims
Filter claims →
Inequality
2332 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 870 | 233 | 116 | 1066 | 2363 |
| Governance & Regulation | 976 | 451 | 218 | 133 | 1809 |
| Organizational Efficiency | 949 | 224 | 144 | 88 | 1416 |
| Technology Adoption Rate | 764 | 287 | 141 | 122 | 1325 |
| Research Productivity | 501 | 152 | 74 | 362 | 1101 |
| Output Quality | 542 | 216 | 69 | 69 | 896 |
| Decision Quality | 387 | 198 | 94 | 54 | 740 |
| Firm Productivity | 513 | 67 | 101 | 27 | 714 |
| AI Safety & Ethics | 249 | 303 | 73 | 36 | 667 |
| Market Structure | 190 | 192 | 134 | 27 | 548 |
| Task Allocation | 243 | 77 | 91 | 36 | 452 |
| Innovation Output | 291 | 33 | 55 | 20 | 401 |
| Skill Acquisition | 206 | 72 | 65 | 21 | 364 |
| Employment Level | 133 | 63 | 115 | 22 | 335 |
| Fiscal & Macroeconomic | 153 | 79 | 52 | 32 | 323 |
| Task Completion Time | 206 | 37 | 12 | 15 | 272 |
| Firm Revenue | 179 | 52 | 29 | 5 | 266 |
| Consumer Welfare | 130 | 76 | 47 | 13 | 266 |
| Inequality Measures | 48 | 137 | 51 | 6 | 242 |
| Worker Satisfaction | 101 | 81 | 25 | 13 | 220 |
| Error Rate | 84 | 110 | 11 | 5 | 210 |
| Wages & Compensation | 98 | 47 | 30 | 10 | 185 |
| Regulatory Compliance | 88 | 73 | 17 | 7 | 185 |
| Automation Exposure | 66 | 64 | 33 | 16 | 182 |
| Team Performance | 105 | 29 | 30 | 11 | 176 |
| Training Effectiveness | 109 | 22 | 14 | 21 | 168 |
| Developer Productivity | 114 | 21 | 14 | 8 | 158 |
| Job Displacement | 12 | 90 | 24 | 1 | 127 |
| Hiring & Recruitment | 57 | 9 | 9 | 5 | 80 |
| Skill Obsolescence | 6 | 56 | 9 | 1 | 72 |
| Social Protection | 43 | 17 | 8 | 2 | 70 |
| Creative Output | 35 | 21 | 9 | 4 | 70 |
| Labor Share of Income | 18 | 21 | 17 | 1 | 57 |
| Worker Turnover | 15 | 16 | — | 4 | 35 |
| Industry | — | — | — | 1 | 1 |
Human Ai Collab
Remove filter
We developed Pocket MonstARs, a controlled gamified abstraction of HRC warehouse inventory picking in which virtual monsters serve as proxies for pick targets, while labeled and object-marked boxes preserve the real-world identification demands of the picking task.
Methods section / system description in the paper describing the experimental testbed used for the user study (Pocket MonstARs).
We conducted a small-scale randomized experiment to measure uplift from human-agent collaboration on real-world computational reproducibility tasks.
Randomized experiment described in the paper (authors report it was small-scale; details in methods section).
SEC S-1 filings combine historical financial statements, governance structures, pro forma and common-control accounting treatments, capital-formation narratives, and underwriting-sensitive risk disclosures within substantially longer documents than typical periodic filings.
Descriptive claim in paper about the composition and length of S-1 filings relative to periodic filings.
Finance Agent v2 narrowly deals with periodic reporting from publicly traded companies (SEC 10-K and 10-Q filings).
Direct statement in paper describing Finance Agent v2's task scope as focused on periodic reporting (SEC 10-K and 10-Q).
FacProcessTwin was evaluated in a real-world case study of an Australian food manufacturer covering 16 production process flows spanning chilled, frozen, and aseptic shelf-stable product categories and including process variations within the same product.
Case study description in paper; sample explicitly stated as 16 production process flows.
AI's effects on writing from the reader side are distinct from those on the production (writer) side.
Empirical results showing readers' accusatory behavior and in-group gatekeeping dynamics operate separately from producers' use of generative AI; argued by contrasting detection/production literature with current reader-side observations.
This research extends signaling theory by showing that substitute signals used socially can grow even when inaccurate if the underlying detection problem cannot be solved at the non-expert level.
Theoretical inference supported by empirical pattern: proliferation of inaccurate yet socially useful accusations (e.g., 'AI slop') in the absence of reliable non-expert detection.
We close by sketching the shape of a harmonised tiered framework and the empirical evaluation needed to calibrate it.
Statement of the paper's concluding contribution in the abstract: a proposed harmonised framework and recommended empirical evaluation; presented as a proposal rather than tested result.
From this we derive a six-dimensional taxonomy (disclosure, responsibility, human oversight, licensing, enforcement, maintainer workload), an ordinal Policy Maturity Score, and a mapping of documented agent incidents onto the dimensions each policy fails to govern.
Stated as study outputs in abstract; indicates development of taxonomy, maturity score, and incident-to-dimension mapping based on the comparative analysis and process tracing.
We compare policies across six organisations (SymPy, LLVM, matplotlib, OpenInfra, the Apache Software Foundation, and the Linux Foundation) using Most-Similar Systems Design with indicator-based coding and process tracing for SymPy and LLVM.
Explicit methods statement in abstract describing sample (six organisations) and methods (Most-Similar Systems Design, indicator-based coding, process tracing for two organisations).
Open-source software, however, evolves through a process designed for humans: contributor agreements, codes of conduct, and review norms all assume a legally accountable person who can attest to provenance and answer reviewer questions.
Descriptive claim about standard open-source governance and contribution processes stated in the abstract; no empirical measurement or coded sample cited in the abstract.
We run a pre-specified 2*4 factorial experiment with 280 complete research runs across four datasets.
Experimental design and sample size reported in paper: 2x4 factorial experiment, 280 complete research runs, four datasets.
We evaluate the system in two controlled human-subject experiments comparing AI-based pre-mediation with professional human mediators in a multi-issue negotiation scenario.
Reported experimental design: two controlled human-subject studies conducted by the authors (details and comparisons described in paper).
The pipeline's components are not autonomous and do not interact peer-to-peer; outputs are passed forward in a fixed sequence (single-party pipeline).
Implementation detail described in the paper (system architecture specification).
We discuss design tradeoffs, failure modes, and lessons learned from operating autonomous AI agents at scale.
Paper statement indicating inclusion of discussion sections on tradeoffs, failure modes, and operational lessons; descriptive/meta claim about paper content.
The core problem is not the absence of explanation but the absence of structured reasoning in the first place.
Conceptual argument/proposed reframing presented in the paper; no empirical test reported.
We ran a controlled three-arm ablation on a production valuation agent: A = plain web-only LLM analyst; B = adds public structured tools + a 14-dimension valuation playbook, verifier, objectivity policy and red-team; C = adds the proprietary Noah AI corpus of curated pipeline, trial and deal intelligence.
Description of experimental arms and setup used in the study (methodological statement).
The evidence base was concentrated in system-facing applications that detect or shape inequities within recruitment, evaluation and exposure systems.
Synthesis result from the scoping review indicating thematic concentration across included studies (as reported in abstract).
ALE is organized around a task taxonomy with 55 subfields grouped into 13 industry clusters covering 1K+ tasks.
Author-provided counts describing the benchmark taxonomy and task pool.
ALE covers non-physical industries defined with reference to O*NET / SOC 2018 (the U.S. federal occupational taxonomy).
Design specification described in the paper referencing O*NET / SOC 2018.
We evaluated seven models (including Gemini, Claude, and GPT families) by comparing their zero-shot estimates against self-reported skill ratings from 27 participants.
Method description: evaluation of seven LLMs comparing zero-shot model estimates to self-reported skill ratings; 27 participants provided self-reports.
AI deployment should be evaluated not only by average task speed, but by its overall effects on congestion, rework, and the robustness of human oversight under load.
Policy/recommendation based on the paper's theoretical results and derived implications from the queueing model (conceptual/prescriptive conclusion; no empirical testing reported).
The divergence between mean task speed and system-level delay caused by AI assistance is labeled the 'variance wedge'.
Definition/terminology introduced in the paper as part of its conceptual framing; supported by the analytic model description.
The benchmark probes 18 mainstream LLMs across four prompting strategies.
Benchmark experiments described in the paper evaluate 18 mainstream LLMs using four different prompting strategies applied to the collected dataset.
GENSTRAT generates a distribution of two-player zero-sum imperfect-information card games.
Design specification in paper; reported generated pool size of 2,000 games (abstract).
The paper provides a taxonomy of minimum input artifacts for agentic software, firmware, and hardware work; a conversation-to-contract gate; risk-adaptive workflows; and an evidence-bundle acceptance model for agent-generated artifacts.
Declared contributions in the paper (deliverables/artefacts produced by the research; no empirical validation provided in the abstract).
The central problem for agentic engineering is no longer prompt engineering; it is engineering process control.
Argument and synthesis presented by the paper (conceptual claim based on reviewed evidence).
The results define three operating regimes.
Summary claim in results/conclusions indicating categorization of outcomes into three regimes.
We performed a large-scale evaluation spanning 15,000 messages with cross-model validation across six LLMs from three families (OpenAI, Anthropic, Google), totaling 1,440 queries.
Study design and reported sample sizes and model counts provided in the paper.
Experiments are run with and without access to Causely under two scenarios: an active incident and a healthy baseline.
Methodological description in the paper describing the two experimental conditions (with/without Causely) and two scenarios (active incident, healthy baseline).
Experiments compare four agent configurations (Claude Code, OpenAI Codex, HolmesGPT with Sonnet and Gemini backends).
Methodological description listing the four agent configurations used in experiments.
We evaluate this value proposition through a benchmark study conducted in a controlled setting with injected faults in a 24-microservice OpenTelemetry demo application.
Methodological description in the paper specifying a controlled benchmark with an OpenTelemetry demo application composed of 24 microservices.
We introduce a taxonomy organized by influence tier, corresponding to interventions on progressively more latent variables: product mentions, information framing, behavioral redirection, and long-term preference shaping.
Paper contribution: authors present a four-tier taxonomy as a conceptual framework; this is a descriptive/constructive claim about the content of the paper itself.
This study presents a sociotechnical audit of six commercial LLMs by comparing their reasoning with a Delphi-derived rubric constructed from the responses of twenty infrastructure professionals.
Method: sociotechnical audit comparing six commercial LLMs to a rubric created via a Delphi process with 20 infrastructure professionals (Delphi-derived rubric).
The framework reframes the central question of autonomous software engineering from whether a foundation model can produce a patch to whether the model-harness-environment system can produce a verifiably correct, attributed, and maintainable change.
Conceptual reframing and argument presented in the abstract as a conclusion of the proposed framework and evaluation approach.
We formalize this substrate as 'AI Harness Engineering' and identify eleven component responsibilities: task specification, context selection, tool access, project memory, task state, observability, failure attribution, verification, permissions, entropy auditing, and intervention recording.
Methodological/conceptual contribution described in the paper (abstract) that lists eleven component responsibilities as part of the formalization.
A symmetric six-gate producer audit separates LLM-engineering failures (template collapse, refusal, internal-ID leakage) from genuine commercial steering.
Methodological claim describing a six-gate producer audit procedure in the paper to diagnose engineering failures vs. commercial steering.
Holding the traveler and bundle fixed, the steering delta is read off between a commission-aware prompt and a minimum-disclosure factual template (paired counterfactual).
Method description of the paired counterfactual experimental design used by TourMart.
We propose TourMart, an applied intelligent-system audit instrument for LLM-OTA commission governance, driven by two governance levers — lambda (gain on message-induced perception) and kappa (budget-normalized cap on how far the message can shift perceived welfare).
Methodological proposal described in paper: design of an audit instrument and two formal levers (lambda, kappa).
Online travel agents (Booking, Trip.com, Expedia) have replaced ranked-list interfaces with conversational LLM agents that compress many options into one sentence of advice.
Descriptive assertion in paper about product/industry UI change; no empirical sample or formal measurement reported in excerpt.
We position DAO-governed decentralized physical infrastructure networks (DePIN) within a vertically integrated stack that links energy and sensing to connectivity, storage/compute, models, and robots.
Architectural/framework description in the paper that maps DePIN elements into a vertically integrated stack; conceptual/mapping method without empirical measurement.
We evaluate 4 popular agent harnesses and 7 foundation models on Workspace-Bench.
Experimental setup reported in the paper listing 4 agent harnesses and 7 foundation models used in evaluations.
Weight-based memory generalizes by applying abstract rules to inputs never seen before.
Conceptual claim grounded in the paper's theoretical distinction between weight-based learning and retrieval; references Complementary Learning Systems theory; no empirical sample in abstract.
Retrieval generalizes by similarity to stored cases.
Conceptual claim stated in paper (distinction between retrieval-based and weight-based generalization); supported by theoretical characterization, not empirical data in abstract.
The process of synthesizing information is inherently iterative: users explore content, identify relationships between concepts, and continuously reorganize their mental models.
Conceptual description of the cognitive/process characteristics in the paper's background/motivation (no empirical measurement reported).
Many practical machine learning applications are online and sequential, meaning prior decisions inform future ones — a setting in which fairness challenges differ from standard supervised learning.
Background claim in the paper motivating the work; literature context and conceptual discussion rather than new empirical data.
The paper establishes a taxonomy of forgetting mechanisms: passive decay-based, active deletion-based, safety-triggered, and adaptive reinforcement-based.
Explicit taxonomy presented in paper (listed in abstract).
We evaluate Aether over synthetic network change scenarios covering main classes of network changes and on past incidents from a major ISP operational network.
Evaluation methodology stated in paper abstract: tested on synthetic scenarios and historical incidents from one major ISP (no numeric sample size provided in abstract).
Expert assessment involved three senior academics producing reports and appointment-level syntheses.
Paper states that three senior academics produced assessment reports and synthesised appointment-level recommendations; n=3 assessors.
The distillation pipeline used an eight-layer extraction method and a nine-module skill architecture grounded in local, closed-corpus analysis.
Methods description in paper specifying an eight-layer extraction approach and nine-module skill architecture; presented as the technical design of the distillation pipeline.