Evidence (535 claims)
Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.
The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).
Browse by theme
Nine broad, paper-level topics. Click one to filter the claims below.
Adoption
20058 claims
Filter claims →
Productivity
17184 claims
Filter claims →
Governance
16099 claims
Filter claims →
Human-AI Collaboration
16034 claims
Filter claims →
Innovation
10501 claims
Filter claims →
Org Design
10496 claims
Filter claims →
Labor Markets
6444 claims
Filter claims →
Skills & Training
5385 claims
Filter claims →
Inequality
4148 claims
Filter claims →
Claims by outcome category
Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.
| Outcome | Positive | Negative | Mixed | Null | Total |
|---|---|---|---|---|---|
| Other | 1820 | 479 | 278 | 1820 | 4588 |
| Organizational Efficiency | 2711 | 616 | 401 | 173 | 3922 |
| Governance & Regulation | 2075 | 886 | 459 | 246 | 3714 |
| Technology Adoption Rate | 1467 | 530 | 258 | 206 | 2488 |
| Decision Quality | 1281 | 496 | 289 | 152 | 2228 |
| Output Quality | 1227 | 447 | 207 | 138 | 2025 |
| AI Safety & Ethics | 634 | 754 | 207 | 83 | 1688 |
| Research Productivity | 826 | 241 | 114 | 422 | 1624 |
| Firm Productivity | 1052 | 154 | 163 | 66 | 1441 |
| Task Allocation | 685 | 211 | 331 | 99 | 1335 |
| Market Structure | 433 | 423 | 242 | 46 | 1150 |
| Innovation Output | 639 | 91 | 105 | 34 | 871 |
| Task Completion Time | 476 | 113 | 43 | 36 | 672 |
| Firm Revenue | 445 | 126 | 58 | 25 | 656 |
| Skill Acquisition | 364 | 119 | 109 | 34 | 626 |
| Consumer Welfare | 288 | 167 | 104 | 31 | 592 |
| Employment Level | 214 | 140 | 174 | 50 | 582 |
| Error Rate | 230 | 251 | 35 | 16 | 535 |
| Fiscal & Macroeconomic | 268 | 136 | 71 | 50 | 532 |
| Inequality Measures | 100 | 307 | 96 | 12 | 515 |
| Worker Satisfaction | 221 | 173 | 60 | 30 | 484 |
| Automation Exposure | 155 | 138 | 65 | 36 | 398 |
| Regulatory Compliance | 171 | 120 | 30 | 13 | 335 |
| Developer Productivity | 222 | 58 | 27 | 13 | 321 |
| Team Performance | 188 | 56 | 50 | 24 | 320 |
| Wages & Compensation | 146 | 104 | 46 | 16 | 312 |
| Training Effectiveness | 207 | 41 | 21 | 26 | 298 |
| Job Displacement | 23 | 153 | 52 | 4 | 232 |
| Hiring & Recruitment | 102 | 57 | 30 | 11 | 202 |
| Skill Obsolescence | 16 | 102 | 24 | 6 | 148 |
| Creative Output | 71 | 42 | 23 | 6 | 143 |
| Social Protection | 57 | 30 | 11 | 3 | 101 |
| Labor Share of Income | 29 | 42 | 24 | 2 | 97 |
| Worker Turnover | 43 | 29 | 6 | 4 | 82 |
| Industry | — | — | — | 1 | 1 |
Hallucination-detection approaches based on response dispersion can provide sample-level evidence of hallucination without access to model internals, but they cannot detect confident, consistent errors.
Formal review of reference-free, internal-state, and retrieval-alignment detection methods, including the worked example in the paper.
Adding GDB and Git MCP servers left ECC largely unchanged for the three tested models but increased extraneous valid calls.
The Easy-Noise and Hard-Noise suites interleaved external-server tasks with design tasks and were compared with clean matched suites for three representative models.
The CRDT substrate guarantees preservation of concurrent edits at the byte level but does not guarantee semantic compatibility or compilation correctness.
Formal discussion of strong eventual consistency and an example in which incompatible function signatures both survive the merge and break downstream compilation.
In the cited ERP study, simple coding-agent tasks succeeded reliably without ERP-specific tooling, while increasing task complexity exposed failures involving lazy heuristics, hallucinated system state, dropped constraints, and over-confidence.
Summary of arXiv:2604.13107, described as testing a coding agent against an open-core ERP system; the paper does not report the study's sample size or success rates.
In the binary JudgeAgent policy, 87.8% of disagreements with human reviewers were false rejections, while 12.2% were false acceptances.
Comparison between early binary JudgeAgent PASS/FAIL decisions and human judgments on the Grocery and Alcohol dataset.
SENTRY has 0.55 precision for the positive high/medium-risk class on the held-out test set.
Per-class evaluation on 35 high/medium-risk cases; Table 3 reports positive-class precision of 0.55.
Unsupervised anomaly-detection methods can identify novel or rare fraud patterns, but they often produce high false-positive rates by flagging benign uncommon behavior.
Survey of clustering, density estimation, isolation forests, autoencoders, and change-point detection methods and their reported operational limitations.
Models performed substantially better on issue categories with explicit lexical signals than on categories requiring contextual inference of intent.
The paper reports mean category performance of 0.835 for the defined-term category and 0.427 for Incorrect Capitalization in Context.
Allowing models to return null when a structural solution is infeasible improves infeasibility detection but creates a bias toward incorrectly declaring feasible instances infeasible.
Comparison of structural-reasoning performance with and without an explicit indeterminacy or null-response option.
The residual damage remained stochastic for the most capable model in the exploratory frontier pass: its one damaging task had an estimated per-run damage probability of 0.16.
The frontier Opus model was evaluated on the CAB-gated task for 32 runs and damaged it in 5/32 runs, yielding an estimated probability of approximately 0.16.
Improved AI-based fraud detection may induce fraudsters to adopt adaptive or adversarial strategies, creating an arms-race dynamic between detection technology and fraud innovation.
Conceptual strategic analysis; the paper proposes game-theoretic or agent-based research but reports no direct empirical test of adaptive fraud behavior.
Audit effectiveness is shaped by regulatory settings, auditor incentives and competencies, behavioural biases, and the increasing complexity of fraud schemes.
Thematic coding of 32 studies identified four recurring drivers: regulatory environment, auditor independence, professional competence and scepticism, and fraud complexity and sophistication.
Static-analysis tools such as CCC and tracehash improved the translation workflow but were insufficient to eliminate all translation errors.
The authors developed call-graph, structural-comparison, instrumentation, and debugger-based tools, then observed that translation errors remained and were revealed by testing on real data.
Failure-mode profiles are idiosyncratic at the trace level but generalize substantially better at the cell level.
MAST/MAD failure-mode analysis using 1,242 LLM-judged traces with 14 binary failure-mode labels; comparison of mean absolute error and correlation for trace-level versus cell-level profiles.
Final transaction state alone can miss evaluated errors: 99 of 861 capability-coverage runs without full credit reached a state also produced by a full-credit execution of the same task.
Comparison of incomplete or erroneous execution traces with full-credit executions using reconstructed final World state and process-level scoring.
The SCOPE routing component is a feasible approximate executor rather than an exact optimizer of the downstream routing-completion value.
The routing decoder uses capacity masks and multi-start decoding, retains the lowest-cost feasible rollout under the implemented metric, and does not use OR-Tools at inference time.
Model performance is evaluated using mean absolute error, exact-match accuracy, and accuracy within one Likert-scale point of the observed response.
The evaluation section formally defines MAE, exact-match accuracy, and relaxed accuracy (Acc±1).
A behavioral failure analysis reveals architecture-distinctive error signatures that aggregate scoring cannot discriminate.
Authors performed a behavioral failure analysis on agent traces and report distinctive error patterns associated with different agent architectures, arguing aggregate scores mask these differences.
Agent-authored code shows modestly elevated corrective rates (26.3% vs. 23.0%), while human code shows higher adaptive rates.
Comparison of modification-type profiles (corrective vs. adaptive) between agent-authored and human-authored code; reported percentages for corrective edits and a qualitative statement about adaptive edits.
Teams of independent AI agents with strict role boundaries (planners, executors, critics, experts), organized as a 'team of rivals' with opposing incentives, can catch and minimize errors within the final product at a small cost to the velocity of actions.
Architectural design and reported system demonstration described in the paper (implementation of specialized agent teams and orchestration); no explicit numeric sample size provided for evaluation of this specific tradeoff in the excerpt.
Current frontier models can demonstrate coherent multi-step behavior, but substantial capability gaps remain before achieving human-level task completion in realistic workplace settings.
Summary conclusion in abstract combining observed multi-step behaviors and the reported ≈40% failure rate, plus failure-mode analysis across the 150 tasks.
Automatiseringen kan minska återkommande manuella fel, men noggrannheten formas i praktiken genom samspel mellan automatiserad tolkning, inbyggda kontroller och mänsklig granskning.
Kvalitativa observationer och respondentuttalanden i studien som beskriver att felminskning uppstår men att slutlig noggrannhet beror på tekniska kontroller och mänsklig inblandning.
Prompting- and embedding-based automated methods for surfacing known failure patterns produce mixed results across methods, i.e., they do not consistently surface the known failures.
Authors tested prompting and embedding-based approaches against the known meta-label failure groups and report heterogeneous / mixed performance across methods (no numerical effect sizes provided in the summary).
This parametric inflation is driven by the leverage structure of the regressor matrix rather than by the error variance: heteroskedasticity-robust standard errors do not directly address the leverage-driven finite-sample distortion, whereas randomization-based inference is insulated from both error-distributional and variance-structural departures by construction.
Analytical argument and simulation/robustness experiments under non-Gaussian and heteroskedastic errors presented in the paper.
Architectural smell density (ASD) declines by 6.7% (p = 0.004), but this decline is a denominator effect resulting from lines-of-code growth rather than an actual architectural improvement.
Observed ASD change computed from estimated smell counts and LOC changes in the 151-repository panel and interpreted by decomposing density into numerator (smells) and denominator (LOC).
A safety monitor condition reduces sabotage success, but 56% of participants still accept the malicious code, ignoring its warnings.
Experimental manipulation: one condition included a safety monitor. Authors report that the monitor reduced sabotage success (no absolute reduction magnitude reported here) and that 56% of participants in that context accepted malicious code despite warnings.
Specialized detectors generally perform better but remain inconsistent across generators and can produce false positives on real-damaged samples.
Experimental comparison showing specialized AI-generated image detectors outperform MLLMs on some generator subsets, yet show variability across generators and some false positives on genuine damaged images.
Failures are structured by task family and execution surface, with HR, management, and multi-system business workflows as persistent bottlenecks and local workspace repair comparatively easier but unsaturated.
Error-mode analysis across the 105 tasks and evaluated models reported in experiments; authors identify task-family-level patterns (HR, management, multi-system workflows) and relative ease of local workspace repair.
Study 1 quantifies confirmation bias through controlled experiments on 250 CVE vulnerability/patch pairs evaluated across four state-of-the-art models under five framing conditions for the review prompt.
Controlled experiment described in the paper: 250 CVE vulnerability/patch pairs evaluated across four state-of-the-art LLMs under five prompt framing conditions.
Helicoid dynamics is a specific failure regime: a system engages competently, drifts into error, accurately names what went wrong, then reproduces the same pattern at a higher level of sophistication, recognizing it is looping and continuing nonetheless.
Definition introduced in the paper and illustrated by the reported case series; the claim is conceptual/phenomenological rather than a statistical result.
The review characterizes hallucination as a systemic property of probabilistic generation rather than merely a transient defect of particular model releases.
Discussion of formal lower-bound and computability results, supplemented by empirical examples of hallucination-related harms.
Effects committed by previous invocations of a delegation can remain visible and make a correct, idempotent component unwinnable.
An event-log incident in which the log was outside the transactional workspace; 21 events accumulated across six invocations, while each invocation correctly published exactly three events.
A progress signal based on identifiers that are constant by construction can produce a guaranteed false stall decision on the third repair round.
One incident used the failing check name as the progress identifier; the identifier remained constant even while failing tests decreased and new tests passed. Two additional recovery paths were found to have the same constant-signal problem by replay.
Agents can enter non-converging loops composed entirely of successful tool calls, making error-rate circuit breakers unable to detect the failure.
A recorded verifier incident in which the same tool call with one distinct payload was issued repeatedly; every call returned success and the run stopped only after human intervention.
With a fixed finite truncation order, proxy collisions can affect entire rows and columns of the panel, and the fraction of affected cells is of order δα(R) + δγ(R).
The paper defines the sets of conflated individual and time types and explains that a unit or time period in these sets contaminates a whole row or column of the N × T panel.
Across testing of more than 100 large language models, only approximately 55% of AI-generated code samples were free of known security issues.
Cross-model code-security audit covering more than 100 large language models.
Early evaluations found that approximately 40% of GitHub Copilot-generated code contained security vulnerabilities, and developers frequently judged their own insecure AI-assisted code to be safe.
Security evaluations of Copilot-generated code, including testing across programming languages and studies of developers' security judgments.
Among agent-first repositories, the increase in static-analysis warnings was approximately 1.7 times larger in repositories without committed AI configuration than in repositories with committed configuration.
The association was estimated from the maturity-stratified reanalysis of an existing agent-adoption panel; the authors caution that maturity is observational and may be confounded by engineering discipline or model capability.
AI-generated code can contain substantive defects, including insecure coding patterns, that developers must detect before the code is deployed.
The paper cites empirical assessment of GitHub Copilot code contributions [13]. The paper does not report the underlying study's sample size or quantitative defect rate.
Premature termination and repetitive retry loops were approximately 2.0-fold and 2.2-fold more common, respectively, in the lowest task-score quantile than across all attempts.
Failure-mode annotations of 260 model-task attempts, with enrichment assessed using two-sided Fisher's exact tests and Benjamini–Hochberg correction.
The number of failure-mode tags was strongly negatively correlated with mean task score across models.
An LLM judge assigned up to 10 failure-mode tags to each model-task attempt; the correlation between total tags and mean task score was Spearman ρ = −0.92 with p = 9.9 × 10−6.
In GAMEFIX, agents perform substantially worse when defects are hidden, and near-complete repair is uncommon for tasks containing multiple bugs.
The GAMEFIX protocol evaluates both explicitly reported defects and defects that agents must discover themselves, using deterministic Fail-to-Pass repair tests and Pass-to-Pass regression checks across 100 repair tasks per run.
The abundance of AI-generated measures creates risks of noisy, poorly validated, or difficult-to-interpret variables, referred to by the authors as “AI slop.”
Review's discussion of the many degrees of freedom in AI measurement and the resulting challenges of variable selection, validation, interpretation, and post-selection inference.
On the dollar scale, the model's median absolute percent error was approximately 51.3%.
Operationally interpretable evaluation using median absolute percent error after converting predictions from the log scale to dollar prices.
Meta AI had the highest High-Confidence Error Rate, with 31.7% of its tested verdicts being incorrect while receiving a confidence score of at least 9 out of 10; Perplexity had a rate of 15.0% and ChatGPT 6.7%.
The paper’s HCER metric was calculated from model correctness and self-reported confidence scores across the 60-case battery.
Surfacing TRACE-enriched attributes on product detail pages reduced the missing-or-incorrect-item rate by 1.08% relative to the control group.
Outcome from the five-week randomized online A/B test; the reported 95% confidence interval was [-2.04%, -0.13%] with p=0.026.
In the same cited study, organizations that invested in digital transformation without concurrent workforce digital-literacy investment experienced a 19% increase in digital-tool-related operational errors compared with organizations that paired technology investment with structured digital-skills development.
The paper summarizes Amadi-Echendu and Ikechukwu (2022); the study design and sample size are not reported in the supplied text.
ContractScrub is designed as a recall-sensitive task because missed contract defects are considered more costly than false-positive flags.
The evaluation focuses primarily on recall, arguing that reviewers can more easily verify whether a flagged issue is real than discover previously unidentified issues.
Current LLMs perform worse on end-to-end contract scrubbing than their performance on seemingly related general-purpose capabilities would suggest.
The authors compare ContractScrub performance with the capabilities represented by general benchmarks and argue that models fall short when those capabilities must be jointly applied to full-document legal review.
All evaluated models had F1 scores below 0.650 on the contract-scrubbing benchmark.
Model predictions were deterministically compared with gold annotations using precision, recall, and F1 across nine issue categories.