The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Evidence (535 claims)

Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.

The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).

Browse by theme

Nine broad, paper-level topics. Click one to filter the claims below.

Adoption
20058 claims
Filter claims →
Productivity
17184 claims
Filter claims →
Governance
16099 claims
Filter claims →
Human-AI Collaboration
16034 claims
Filter claims →
Innovation
10501 claims
Filter claims →
Org Design
10496 claims
Filter claims →
Labor Markets
6444 claims
Filter claims →
Skills & Training
5385 claims
Filter claims →
Inequality
4148 claims
Filter claims →

Claims by outcome category

Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.

Outcome Positive Negative Mixed Null Total
Other 1820 479 278 1820 4588
Organizational Efficiency 2711 616 401 173 3922
Governance & Regulation 2075 886 459 246 3714
Technology Adoption Rate 1467 530 258 206 2488
Decision Quality 1281 496 289 152 2228
Output Quality 1227 447 207 138 2025
AI Safety & Ethics 634 754 207 83 1688
Research Productivity 826 241 114 422 1624
Firm Productivity 1052 154 163 66 1441
Task Allocation 685 211 331 99 1335
Market Structure 433 423 242 46 1150
Innovation Output 639 91 105 34 871
Task Completion Time 476 113 43 36 672
Firm Revenue 445 126 58 25 656
Skill Acquisition 364 119 109 34 626
Consumer Welfare 288 167 104 31 592
Employment Level 214 140 174 50 582
Error Rate 230 251 35 16 535
Fiscal & Macroeconomic 268 136 71 50 532
Inequality Measures 100 307 96 12 515
Worker Satisfaction 221 173 60 30 484
Automation Exposure 155 138 65 36 398
Regulatory Compliance 171 120 30 13 335
Developer Productivity 222 58 27 13 321
Team Performance 188 56 50 24 320
Wages & Compensation 146 104 46 16 312
Training Effectiveness 207 41 21 26 298
Job Displacement 23 153 52 4 232
Hiring & Recruitment 102 57 30 11 202
Skill Obsolescence 16 102 24 6 148
Creative Output 71 42 23 6 143
Social Protection 57 30 11 3 101
Labor Share of Income 29 42 24 2 97
Worker Turnover 43 29 6 4 82
Industry 1 1
Hallucination-detection approaches based on response dispersion can provide sample-level evidence of hallucination without access to model internals, but they cannot detect confident, consistent errors.
Formal review of reference-free, internal-state, and retrieval-alignment detection methods, including the worked example in the paper.
high mixed LAAF: A Layered Accountability Architecture Framework for LL... Hallucination detection capability
Adding GDB and Git MCP servers left ECC largely unchanged for the three tested models but increased extraneous valid calls.
The Easy-Noise and Hard-Noise suites interleaved external-server tasks with design tasks and were compared with clean matched suites for three representative models.
high mixed Benchmarking AI Agents for Hardware Design Automation via MC... Expected Call Coverage and Extraneous Valid Call Ratio
The CRDT substrate guarantees preservation of concurrent edits at the byte level but does not guarantee semantic compatibility or compilation correctness.
Formal discussion of strong eventual consistency and an example in which incompatible function signatures both survive the merge and break downstream compilation.
high mixed AgentRoom: Concurrent Multi-Agent Coding in a CRDT-Backed Sh... Semantic compatibility and downstream compilation after concurrent edits
In the cited ERP study, simple coding-agent tasks succeeded reliably without ERP-specific tooling, while increasing task complexity exposed failures involving lazy heuristics, hallucinated system state, dropped constraints, and over-confidence.
Summary of arXiv:2604.13107, described as testing a coding agent against an open-core ERP system; the paper does not report the study's sample size or success rates.
high mixed Applying Anthropic Primitives at Large Enterprises: Harness ... Reliability and failure modes of coding-agent task execution in an ERP environme...
In the binary JudgeAgent policy, 87.8% of disagreements with human reviewers were false rejections, while 12.2% were false acceptances.
Comparison between early binary JudgeAgent PASS/FAIL decisions and human judgments on the Grocery and Alcohol dataset.
high mixed TRACE: Agentic Catalog Enrichment with Multi-source Evidence... Types of disagreement between automated judgments and human judgments
SENTRY has 0.55 precision for the positive high/medium-risk class on the held-out test set.
Per-class evaluation on 35 high/medium-risk cases; Table 3 reports positive-class precision of 0.55.
high mixed SENTRY: Deterministic, Intelligent Risk Assessment for IT Ch... Proportion of changes flagged high/medium risk that subsequently belong to the h...
Unsupervised anomaly-detection methods can identify novel or rare fraud patterns, but they often produce high false-positive rates by flagging benign uncommon behavior.
Survey of clustering, density estimation, isolation forests, autoencoders, and change-point detection methods and their reported operational limitations.
high mixed From Sampling to Surveillance: Evaluating the Effectiveness ... Detection of novel fraud patterns and false-positive rate
Models performed substantially better on issue categories with explicit lexical signals than on categories requiring contextual inference of intent.
The paper reports mean category performance of 0.835 for the defined-term category and 0.427 for Incorrect Capitalization in Context.
high mixed ContractScrub: A benchmark for final review of legal contrac... Category-level recall in contract-error detection
Allowing models to return null when a structural solution is infeasible improves infeasibility detection but creates a bias toward incorrectly declaring feasible instances infeasible.
Comparison of structural-reasoning performance with and without an explicit indeterminacy or null-response option.
high mixed Preference Reasoning under Indeterminacy in Large Language M... Detection of infeasible instances and false declarations of non-existence on fea...
The residual damage remained stochastic for the most capable model in the exploratory frontier pass: its one damaging task had an estimated per-run damage probability of 0.16.
The frontier Opus model was evaluated on the CAB-gated task for 32 runs and damaged it in 5/32 runs, yielding an estimated probability of approximately 0.16.
high mixed No Task Fails Every Time: Why One-Shot Audits Are Structural... Per-run probability of damage on the residual damaging task
Improved AI-based fraud detection may induce fraudsters to adopt adaptive or adversarial strategies, creating an arms-race dynamic between detection technology and fraud innovation.
Conceptual strategic analysis; the paper proposes game-theoretic or agent-based research but reports no direct empirical test of adaptive fraud behavior.
high mixed Bridging the expectation performance gap in fraud detection:... Fraud sophistication and effectiveness of fraud detection
Audit effectiveness is shaped by regulatory settings, auditor incentives and competencies, behavioural biases, and the increasing complexity of fraud schemes.
Thematic coding of 32 studies identified four recurring drivers: regulatory environment, auditor independence, professional competence and scepticism, and fraud complexity and sophistication.
high mixed Bridging the expectation performance gap in fraud detection:... Audit effectiveness and fraud-detection performance
Static-analysis tools such as CCC and tracehash improved the translation workflow but were insufficient to eliminate all translation errors.
The authors developed call-graph, structural-comparison, instrumentation, and debugger-based tools, then observed that translation errors remained and were revealed by testing on real data.
high mixed Static analysis-guided agentic AI translation enables Rust a... Translation error detection and residual error rate
Failure-mode profiles are idiosyncratic at the trace level but generalize substantially better at the cell level.
MAST/MAD failure-mode analysis using 1,242 LLM-judged traces with 14 binary failure-mode labels; comparison of mean absolute error and correlation for trace-level versus cell-level profiles.
high mixed Deployment Decision Reliability: A Generalizability-Theory F... Generalizability of multi-agent failure-mode profiles
Final transaction state alone can miss evaluated errors: 99 of 861 capability-coverage runs without full credit reached a state also produced by a full-credit execution of the same task.
Comparison of incomplete or erroneous execution traces with full-credit executions using reconstructed final World state and process-level scoring.
high mixed Agentic Commerce World: An Auditable and Verifiable Environm... Ability of final-state evaluation to detect process errors
The SCOPE routing component is a feasible approximate executor rather than an exact optimizer of the downstream routing-completion value.
The routing decoder uses capacity masks and multi-start decoding, retains the lowest-cost feasible rollout under the implemented metric, and does not use OR-Tools at inference time.
high mixed SCOPE: Supply-Chain Operations through Coupled Policies for ... Routing feasibility and route cost.
Model performance is evaluated using mean absolute error, exact-match accuracy, and accuracy within one Likert-scale point of the observed response.
The evaluation section formally defines MAE, exact-match accuracy, and relaxed accuracy (Acc±1).
high mixed Simulating Tenant Responses to Energy Policy Interventions w... Ordinal prediction error and classification agreement with observed survey label...
A behavioral failure analysis reveals architecture-distinctive error signatures that aggregate scoring cannot discriminate.
Authors performed a behavioral failure analysis on agent traces and report distinctive error patterns associated with different agent architectures, arguing aggregate scores mask these differences.
high mixed General Agent Evaluation error signature patterns vs. aggregate scoring
Agent-authored code shows modestly elevated corrective rates (26.3% vs. 23.0%), while human code shows higher adaptive rates.
Comparison of modification-type profiles (corrective vs. adaptive) between agent-authored and human-authored code; reported percentages for corrective edits and a qualitative statement about adaptive edits.
high mixed Will It Survive? Deciphering the Fate of AI-Generated Code i... proportion of modifications classified as corrective (versus adaptive)
Teams of independent AI agents with strict role boundaries (planners, executors, critics, experts), organized as a 'team of rivals' with opposing incentives, can catch and minimize errors within the final product at a small cost to the velocity of actions.
Architectural design and reported system demonstration described in the paper (implementation of specialized agent teams and orchestration); no explicit numeric sample size provided for evaluation of this specific tradeoff in the excerpt.
high mixed If You Want Coherence, Orchestrate a Team of Rivals: Multi-A... error interception/minimization and action velocity
Current frontier models can demonstrate coherent multi-step behavior, but substantial capability gaps remain before achieving human-level task completion in realistic workplace settings.
Summary conclusion in abstract combining observed multi-step behaviors and the reported ≈40% failure rate, plus failure-mode analysis across the 150 tasks.
high mixed The Hierarchy of Agentic Capabilities: Evaluating Frontier M... multi-step task completion ability vs gap to human-level performance
Automatiseringen kan minska återkommande manuella fel, men noggrannheten formas i praktiken genom samspel mellan automatiserad tolkning, inbyggda kontroller och mänsklig granskning.
Kvalitativa observationer och respondentuttalanden i studien som beskriver att felminskning uppstår men att slutlig noggrannhet beror på tekniska kontroller och mänsklig inblandning.
high mixed AI-baserad automatisering i leverantorsfakturahantering : En... felränta / noggrannhet (error rate / accuracy)
Prompting- and embedding-based automated methods for surfacing known failure patterns produce mixed results across methods, i.e., they do not consistently surface the known failures.
Authors tested prompting and embedding-based approaches against the known meta-label failure groups and report heterogeneous / mixed performance across methods (no numerical effect sizes provided in the summary).
This parametric inflation is driven by the leverage structure of the regressor matrix rather than by the error variance: heteroskedasticity-robust standard errors do not directly address the leverage-driven finite-sample distortion, whereas randomization-based inference is insulated from both error-distributional and variance-structural departures by construction.
Analytical argument and simulation/robustness experiments under non-Gaussian and heteroskedastic errors presented in the paper.
high mixed Testing the Significance of the Difference-in-Differences Co... source of finite-sample size distortion (leverage) and comparative robustness of...
Architectural smell density (ASD) declines by 6.7% (p = 0.004), but this decline is a denominator effect resulting from lines-of-code growth rather than an actual architectural improvement.
Observed ASD change computed from estimated smell counts and LOC changes in the 151-repository panel and interpreted by decomposing density into numerator (smells) and denominator (LOC).
high mixed Mining Architectural Quality Under Agentic AI Adoption: A Ca... architectural smell density (ASD)
A safety monitor condition reduces sabotage success, but 56% of participants still accept the malicious code, ignoring its warnings.
Experimental manipulation: one condition included a safety monitor. Authors report that the monitor reduced sabotage success (no absolute reduction magnitude reported here) and that 56% of participants in that context accepted malicious code despite warnings.
high mixed Coding with "Enemy": Can Human Developers Detect AI Agent Sa... acceptance of malicious code / sabotage success under safety monitor
Specialized detectors generally perform better but remain inconsistent across generators and can produce false positives on real-damaged samples.
Experimental comparison showing specialized AI-generated image detectors outperform MLLMs on some generator subsets, yet show variability across generators and some false positives on genuine damaged images.
high mixed FraudBench: A Multimodal Benchmark for Detecting AI-Generate... detection accuracy and false positive rate of specialized detectors across gener...
Failures are structured by task family and execution surface, with HR, management, and multi-system business workflows as persistent bottlenecks and local workspace repair comparatively easier but unsaturated.
Error-mode analysis across the 105 tasks and evaluated models reported in experiments; authors identify task-family-level patterns (HR, management, multi-system workflows) and relative ease of local workspace repair.
high mixed Claw-Eval-Live: A Live Agent Benchmark for Evolving Real-Wor... failure distribution by task family / execution surface
Study 1 quantifies confirmation bias through controlled experiments on 250 CVE vulnerability/patch pairs evaluated across four state-of-the-art models under five framing conditions for the review prompt.
Controlled experiment described in the paper: 250 CVE vulnerability/patch pairs evaluated across four state-of-the-art LLMs under five prompt framing conditions.
high mixed Measuring and Exploiting Confirmation Bias in LLM-Assisted S... confirmation bias as measured by vulnerability detection performance
Helicoid dynamics is a specific failure regime: a system engages competently, drifts into error, accurately names what went wrong, then reproduces the same pattern at a higher level of sophistication, recognizing it is looping and continuing nonetheless.
Definition introduced in the paper and illustrated by the reported case series; the claim is conceptual/phenomenological rather than a statistical result.
high mixed AI Knows What's Wrong But Cannot Fix It: Helicoid Dynamics i... incidence and qualitative characterization of the helicoid pattern in LLM intera...
The review characterizes hallucination as a systemic property of probabilistic generation rather than merely a transient defect of particular model releases.
Discussion of formal lower-bound and computability results, supplemented by empirical examples of hallucination-related harms.
high negative LAAF: A Layered Accountability Architecture Framework for LL... Reliability and factuality of LLM outputs
Effects committed by previous invocations of a delegation can remain visible and make a correct, idempotent component unwinnable.
An event-log incident in which the log was outside the transactional workspace; 21 events accumulated across six invocations, while each invocation correctly published exactly three events.
high negative Agent Mesh: Reliability Primitives for Non-Idempotent Agent ... Correctness and recoverability of repeated component executions
A progress signal based on identifiers that are constant by construction can produce a guaranteed false stall decision on the third repair round.
One incident used the failing check name as the progress identifier; the identifier remained constant even while failing tests decreased and new tests passed. Two additional recovery paths were found to have the same constant-signal problem by replay.
high negative Agent Mesh: Reliability Primitives for Non-Idempotent Agent ... Accuracy of the progress/stall detector during repair
Agents can enter non-converging loops composed entirely of successful tool calls, making error-rate circuit breakers unable to detect the failure.
A recorded verifier incident in which the same tool call with one distinct payload was issued repeatedly; every call returned success and the run stopped only after human intervention.
high negative Agent Mesh: Reliability Primitives for Non-Idempotent Agent ... Detection of agent non-convergence by error-rate breakers
With a fixed finite truncation order, proxy collisions can affect entire rows and columns of the panel, and the fraction of affected cells is of order δα(R) + δγ(R).
The paper defines the sets of conflated individual and time types and explains that a unit or time period in these sets contaminates a whole row or column of the N × T panel.
high negative Nonparametric Identification of Two-Way Unobserved Heterogen... Fraction of panel cells affected by proxy misclassification
Across testing of more than 100 large language models, only approximately 55% of AI-generated code samples were free of known security issues.
Cross-model code-security audit covering more than 100 large language models.
high negative AI and the Future of Software Engineering: Expertise, Employ... Share of AI-generated code samples free of known security flaws
Early evaluations found that approximately 40% of GitHub Copilot-generated code contained security vulnerabilities, and developers frequently judged their own insecure AI-assisted code to be safe.
Security evaluations of Copilot-generated code, including testing across programming languages and studies of developers' security judgments.
high negative AI and the Future of Software Engineering: Expertise, Employ... Security vulnerability rate in AI-generated code and developers' security assess...
Among agent-first repositories, the increase in static-analysis warnings was approximately 1.7 times larger in repositories without committed AI configuration than in repositories with committed configuration.
The association was estimated from the maturity-stratified reanalysis of an existing agent-adoption panel; the authors caution that maturity is observational and may be confounded by engineering discipline or model capability.
high negative A Few Pages of Markdown: Committed AI Configuration and Lowe... Change in static-analysis warnings after coding-agent adoption
AI-generated code can contain substantive defects, including insecure coding patterns, that developers must detect before the code is deployed.
The paper cites empirical assessment of GitHub Copilot code contributions [13]. The paper does not report the underlying study's sample size or quantitative defect rate.
high negative From Producing to Validating: How AI Is Deskilling Freelance... Defects and security vulnerabilities in generated code
Premature termination and repetitive retry loops were approximately 2.0-fold and 2.2-fold more common, respectively, in the lowest task-score quantile than across all attempts.
Failure-mode annotations of 260 model-task attempts, with enrichment assessed using two-sided Fisher's exact tests and Benjamini–Hochberg correction.
high negative BixBench3: Benchmarking AI agents on research-study-scale co... Frequency of agent failure modes during task execution
The number of failure-mode tags was strongly negatively correlated with mean task score across models.
An LLM judge assigned up to 10 failure-mode tags to each model-task attempt; the correlation between total tags and mean task score was Spearman ρ = −0.92 with p = 9.9 × 10−6.
high negative BixBench3: Benchmarking AI agents on research-study-scale co... Failure-mode frequency and benchmark task score
In GAMEFIX, agents perform substantially worse when defects are hidden, and near-complete repair is uncommon for tasks containing multiple bugs.
The GAMEFIX protocol evaluates both explicitly reported defects and defects that agents must discover themselves, using deterministic Fail-to-Pass repair tests and Pass-to-Pass regression checks across 100 repair tasks per run.
high negative GameXpert-Bench: How Far Are Coding Agents from Expert Game ... Bug discovery and successful multi-bug repair without regressions
The abundance of AI-generated measures creates risks of noisy, poorly validated, or difficult-to-interpret variables, referred to by the authors as “AI slop.”
Review's discussion of the many degrees of freedom in AI measurement and the resulting challenges of variable selection, validation, interpretation, and post-selection inference.
high negative The Measurement Revolution? Credible Measurement and Inferen... Quality and interpretability of AI-generated variables
On the dollar scale, the model's median absolute percent error was approximately 51.3%.
Operationally interpretable evaluation using median absolute percent error after converting predictions from the log scale to dollar prices.
high negative A Machine Learning Framework for Price Estimation in Air For... Dollar-scale acquisition price estimation error
Meta AI had the highest High-Confidence Error Rate, with 31.7% of its tested verdicts being incorrect while receiving a confidence score of at least 9 out of 10; Perplexity had a rate of 15.0% and ChatGPT 6.7%.
The paper’s HCER metric was calculated from model correctness and self-reported confidence scores across the 60-case battery.
high negative Can Legal AI Know When It Is Wrong? And Do Students Know Whe... Rate of incorrect legal verdicts delivered with high self-reported confidence
Surfacing TRACE-enriched attributes on product detail pages reduced the missing-or-incorrect-item rate by 1.08% relative to the control group.
Outcome from the five-week randomized online A/B test; the reported 95% confidence interval was [-2.04%, -0.13%] with p=0.026.
high negative TRACE: Agentic Catalog Enrichment with Multi-source Evidence... Rate of missing or incorrect items
In the same cited study, organizations that invested in digital transformation without concurrent workforce digital-literacy investment experienced a 19% increase in digital-tool-related operational errors compared with organizations that paired technology investment with structured digital-skills development.
The paper summarizes Amadi-Echendu and Ikechukwu (2022); the study design and sample size are not reported in the supplied text.
high negative Reskilling and Upskilling in the Age of Automation: Continuo... Operational errors related to digital tools.
ContractScrub is designed as a recall-sensitive task because missed contract defects are considered more costly than false-positive flags.
The evaluation focuses primarily on recall, arguing that reviewers can more easily verify whether a flagged issue is real than discover previously unidentified issues.
high negative ContractScrub: A benchmark for final review of legal contrac... Relative cost of false negatives versus false positives in contract review
Current LLMs perform worse on end-to-end contract scrubbing than their performance on seemingly related general-purpose capabilities would suggest.
The authors compare ContractScrub performance with the capabilities represented by general benchmarks and argue that models fall short when those capabilities must be jointly applied to full-document legal review.
high negative ContractScrub: A benchmark for final review of legal contrac... End-to-end contract-scrubbing performance
All evaluated models had F1 scores below 0.650 on the contract-scrubbing benchmark.
Model predictions were deterministically compared with gold annotations using precision, recall, and F1 across nine issue categories.
high negative ContractScrub: A benchmark for final review of legal contrac... F1 score for contract-defect identification