The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests About 🎲 Workforce Futures
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Evidence (7560 claims)

Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.

The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).

Browse by theme

Nine broad, paper-level topics. Click one to filter the claims below.

Adoption
9875 claims
Filter claims →
Productivity
8807 claims
Filter claims →
Governance
7870 claims
Filter claims →
Human-AI Collaboration
7560 claims
Filtered →
Org Design
4892 claims
Filter claims →
Innovation
4781 claims
Filter claims →
Labor Markets
4004 claims
Filter claims →
Skills & Training
3308 claims
Filter claims →
Inequality
2332 claims
Filter claims →

Claims by outcome category

Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.

Outcome Positive Negative Mixed Null Total
Other 870 233 116 1066 2363
Governance & Regulation 976 451 218 133 1809
Organizational Efficiency 949 224 144 88 1416
Technology Adoption Rate 764 287 141 122 1325
Research Productivity 501 152 74 362 1101
Output Quality 542 216 69 69 896
Decision Quality 387 198 94 54 740
Firm Productivity 513 67 101 27 714
AI Safety & Ethics 249 303 73 36 667
Market Structure 190 192 134 27 548
Task Allocation 243 77 91 36 452
Innovation Output 291 33 55 20 401
Skill Acquisition 206 72 65 21 364
Employment Level 133 63 115 22 335
Fiscal & Macroeconomic 153 79 52 32 323
Task Completion Time 206 37 12 15 272
Firm Revenue 179 52 29 5 266
Consumer Welfare 130 76 47 13 266
Inequality Measures 48 137 51 6 242
Worker Satisfaction 101 81 25 13 220
Error Rate 84 110 11 5 210
Wages & Compensation 98 47 30 10 185
Regulatory Compliance 88 73 17 7 185
Automation Exposure 66 64 33 16 182
Team Performance 105 29 30 11 176
Training Effectiveness 109 22 14 21 168
Developer Productivity 114 21 14 8 158
Job Displacement 12 90 24 1 127
Hiring & Recruitment 57 9 9 5 80
Skill Obsolescence 6 56 9 1 72
Social Protection 43 17 8 2 70
Creative Output 35 21 9 4 70
Labor Share of Income 18 21 17 1 57
Worker Turnover 15 16 4 35
Industry 1 1
Clear
Human Ai Collab Remove filter
Only 12% of AI market value is used in physical activities.
Descriptive aggregate: authors categorize and report that 12% of estimated AI market value maps to physical activities.
high negative Where can AI be used? Insights from a deep ontology of work ... share of AI market value by activity type (physical)
Applying them to hardware-in-the-loop (HIL) embedded and Internet-of-Things (IoT) systems remains challenging due to the tight coupling between software logic and physical hardware behavior; code that compiles successfully may still fail when deployed on real devices because of timing constraints, peripheral initialization requirements, or hardware-specific behaviors.
Conceptual/engineering reasoning stated in the paper describing known HIL/IoT failure modes (no experimental quantification provided in this excerpt).
high negative Skilled AI Agents for Embedded and IoT Systems Development code failure / runtime correctness when deployed to hardware
Across heterogeneous learners, a common broadcast curriculum can be slower than personalized instruction by a factor linear in the number of learner types.
Theoretical comparative result in the model (analysis of broadcast vs personalized curricula across heterogeneous learner types; abstract states factor linear in number of types).
high negative A Mathematical Theory of Understanding speed of instruction / time to learn under broadcast curriculum vs personalized ...
The findings provide evidence against cue-based accounts of lie detection more generally.
Authors' interpretation: because lie-detection accuracy did not decrease despite changes to visual cues (retouching, backgrounds, avatars), the results challenge theories that rely on superficial cues for lie detection.
high negative Through the Looking-Glass: AI-Mediated Video Communication R... validity of cue-based accounts of lie detection
Participants' confidence in their judgments declined in AI-mediated videos, particularly when some participants used avatars while others did not.
Experimental comparisons across conditions with varying levels of AI mediation; subgroup/condition contrast highlighting larger declines in mixed-avatar settings.
high negative Through the Looking-Glass: AI-Mediated Video Communication R... participants' confidence in their lie-detection judgments
Perceived trust in speakers declined in AI-mediated videos.
Experimental results from the two preregistered online experiments comparing perceived trust across varying levels of AI mediation (retouching, background replacement, avatars).
high negative Through the Looking-Glass: AI-Mediated Video Communication R... perceived trust in speakers
AI-based tools that mediate, enhance or generate parts of video communication may interfere with how people evaluate trustworthiness and credibility.
Motivating claim stated in the paper's introduction/abstract; not an empirical finding but a hypothesis motivating the experiments.
high negative Through the Looking-Glass: AI-Mediated Video Communication R... evaluation of trustworthiness and credibility (general)
Compositional spatial reasoning remains a formidable challenge for state-of-the-art VLMs (as revealed by our evaluation).
Empirical results from the evaluation of the 37 VLMs on the MultihopSpatial benchmark showing poor performance on multi-hop/compositional queries.
high negative MultihopSpatial: Multi-hop Compositional Spatial Reasoning B... performance on compositional/multi-hop spatial reasoning tasks
Existing benchmarks predominantly focus on elementary, single-hop relations and neglect multi-hop compositional spatial reasoning and precise visual grounding needed for real-world scenarios.
Literature/benchmark survey and motivation presented by the authors comparing characteristics of prior benchmarks vs. the proposed needs.
high negative MultihopSpatial: Multi-hop Compositional Spatial Reasoning B... scope/complexity of spatial reasoning tasks in existing benchmarks
Significant limitations emerged in case law citations, with most cited cases being non-existent or incorrectly referenced.
Authors' review of the case citations produced by the four AI engines for the single transcript, finding many citations were fabricated or misreferenced.
high negative Robot Wingman: Using AI to Assess an Employment Termination accuracy of case law citations (error rate / hallucination rate)
Initial adaptation challenges to AI integration were identified among employees.
Participants in semi-structured interviews (n=12) reported initial difficulties adapting to AI tools; themes relating to early adaptation challenges were coded.
high negative AI-AUGMENTED WORKFORCE: THE IMPACT OF ARTIFICIAL INTELLIGENC... initial adaptation challenges to AI
There is a measurement asymmetry in standard LLM evaluation: unconstrained prompts can inflate constraint-adherence scores and mask the practical value of structured prompting.
Analysis of evaluation results from the controlled study showing that unconstrained (simple) prompts sometimes achieve high constraint-adherence scores, leading to misleading evaluation of structured prompts' benefits.
high negative Evaluating 5W3H Structured Prompting for Intent Alignment in... constraint_adherence_scores / evaluation_bias
Traditional paradigms, specifically the resource-based view and the dynamic capabilities framework, operate under closed-system, first-order cybernetic assumptions that fail to capture the dissipative nature of algorithmic agents.
Conceptual critique presented in the paper's theoretical argumentation (literature critique and re-framing); no empirical sample reported.
high negative Governing Human–AI Co-Evolution: Intelligentization Capabili... explanatory_power_of_management_theory (ability to account for AI-driven organiz...
AI usage predicts work disengagement behavior via emotional exhaustion elicited by AI-associated technostressors.
Four-stage longitudinal study (survey) of finance professionals (N=285); mediation analysis testing AI usage -> technostressors -> emotional exhaustion -> work disengagement, based on SOR framework.
high negative Autonomous enhancement or emotional depletion? The dual-path... work disengagement behavior (mediated by emotional exhaustion from technostresso...
These findings highlight fundamental challenges in the numerical and time-series reasoning for current LLMs and motivate future research in financial intelligence.
Interpretation of experimental results in the paper: authors conclude that the observed limited gains (particularly on trading-signal/time-series aspects) indicate shortcomings in LLM numerical and time-series reasoning.
high negative FinTradeBench: A Financial Reasoning Benchmark for LLMs LLMs' numerical and time-series reasoning capability (qualitative conclusion fro...
There is a central design tension in human-AI systems: maximizing short-term hybrid capability does not necessarily preserve long-term human cognitive competence.
Conceptual/theoretical claim derived from the framework and discussion in the paper (argument and mathematical framing), no empirical sample or longitudinal data presented in the excerpt.
high negative Cognitive Amplification vs Cognitive Delegation in Human-AI ... long-term human cognitive competence
Rather than broad job losses, evidence points to a reallocation at the entry level: AI automates tasks typically assigned to junior staff, shifting the nature of entry-level roles.
Synthesis of firm- and task-level empirical studies reported in the brief documenting automation of routine/junior tasks and changes in job-task composition; specific sample sizes vary by cited study and are not provided in the brief.
high negative AI, Productivity, and Labor Markets: A Review of the Empiric... automation of entry-level/junior tasks and changes to entry-level job content
Confirmation bias poses a weakness in LLM-based code review, with implications on how AI-assisted development tools are deployed.
Synthesis of findings from Study 1 (framing-induced detection failures) and Study 2 (practical exploitability and partial mitigation via debiasing).
high negative Measuring and Exploiting Confirmation Bias in LLM-Assisted S... reliability/security of LLM-based code review
Adversarial framing succeeds in 88% of cases against Claude Code (autonomous agent) in real project configurations where adversaries can iteratively refine their framing to increase attack success.
Study 2 experiments in real project configurations with iterative adversary refinement evaluated against Claude Code (autonomous agent); reported 88% success rate.
high negative Measuring and Exploiting Confirmation Bias in LLM-Assisted S... attack success rate (vulnerability reintroduction accepted/not detected)
Adversarial pull request framing (e.g., labeled as security improvements or urgent functionality fixes) succeeds in reintroducing known vulnerabilities in 35% of cases against GitHub Copilot under one-shot attacks.
Study 2 experiments simulating adversarial pull requests evaluated against GitHub Copilot (interactive assistant); reported success rate 35% for one-shot attacks.
high negative Measuring and Exploiting Confirmation Bias in LLM-Assisted S... attack success rate (vulnerability reintroduction accepted/not detected)
The framing effect is strongly asymmetric: false negatives increase sharply while false positive rates change little.
Comparison of false negative and false positive rates across framing conditions in Study 1 experiments (250 CVE pairs across models).
high negative Measuring and Exploiting Confirmation Bias in LLM-Assisted S... false negative rate and false positive rate
Framing a change as bug-free reduces vulnerability detection rates by 16-93%.
Result reported from Study 1 controlled experiments across models and framing conditions (250 CVE pairs).
high negative Measuring and Exploiting Confirmation Bias in LLM-Assisted S... vulnerability detection rate
AI-only baselines perform near or below the median of competition participants.
Comparison of AI-only baseline performance to the distribution of competition participant results reported in the paper (competition with 29 teams / 80 participants).
high negative AgentDS Technical Report: Benchmarking the Future of Human-A... relative performance rank of AI-only baselines vs participants
Our results show that current AI agents struggle with domain-specific reasoning.
Outcome of the competition reported in the paper comparing AI-only baselines to participant submissions across the AgentDS tasks (competition data from 29 teams / 80 participants); reported aggregate performance indicating AI weakness on domain-specific tasks.
high negative AgentDS Technical Report: Benchmarking the Future of Human-A... domain-specific reasoning performance
LLM-generated peer reviews place significantly less weight on clarity and significance of the research.
Comparative analysis between LLM-generated reviews and human reviews from the conference dataset; reported as a statistically significant difference but exact statistics and sample size not provided in the excerpt.
high negative How LLMs Distort Our Written Language importance/weight given to clarity and significance in peer review content
Significantly more heavy LLM users reported that the writing was less creative and not in their voice.
Self-reported measures from participants in the human user study comparing heavy LLM users to others; no sample size or exact statistics provided in the excerpt.
high negative How LLMs Distort Our Written Language self-reported creativity and 'in-your-voice' authenticity of writing
The gap between informal natural language requirements and precise program behavior (the 'intent gap') has always plagued software engineering, but AI-generated code amplifies it to an unprecedented scale.
Conceptual claim and argumentation in the paper; presented as an observed escalation in the scale of the existing 'intent gap' due to AI code generation. No quantitative evidence or sample size given in the excerpt.
high negative Intent Formalization: A Grand Challenge for Reliable Coding ... mismatch between intended and actual program behavior (intent gap) / resulting c...
These dynamics amplify initial disparities and produce persistent performance gaps across the population.
Main theoretical conclusion of the paper: analysis of the proposed dynamical system showing amplification and persistence of gaps (authors' demonstrated result).
high negative Actionable Recourse in Competitive Environments: A Dynamic G... magnitude and persistence of performance disparities across population over time
Exclusion-based cohesion can produce state-contingent illusory precision together with effective input concentration and dynamic lock-in simultaneously—i.e., these phenomena co-occur under the model's parameter regimes.
Analytical model results showing co-occurrence of multiple adverse phenomena (bias that grows in tails, illusory precision, input concentration, lock-in) under the same exclusion mechanisms; derived within the paper's theoretical framework.
high negative Cohesion as Concentration: Exclusion-Driven Fragility in Fin... co-occurrence of multiple adverse outcomes: tail bias, observed disagreement, ef...
When the anchor belief is updated from internally filtered aggregates, the system can exhibit dynamic lock-in: delayed recognition of regime shifts followed by abrupt correction.
Analytical dynamics studied in the model when anchor updates depend on filtered (excluded) aggregates; derivations demonstrate delayed detection and abrupt adjustments. This is a theoretical/dynamical model result, no empirical data.
high negative Cohesion as Concentration: Exclusion-Driven Fragility in Fin... delay in regime recognition and magnitude/timing of corrective update
Exclusion leads to effective concentration of decision inputs: the effective number of independent inputs falls below the nominal participant count.
Model-derived analytic result showing that report shrinkage and discarding reduce effective information contributions, quantified relative to nominal participation in the theoretical framework. No empirical sample.
high negative Cohesion as Concentration: Exclusion-Driven Fragility in Fin... effective number of independent decision inputs (information concentration)
Exclusion-based cohesion induces 'illusory precision': observed disagreement can fall while actual estimation error in tail regimes rises (i.e., lower recorded variance despite higher true error).
Theoretical result derived from the signal-aggregation model showing a regime in which filtered reports reduce observed variance even as tail-regime estimation error increases. No empirical validation provided.
high negative Cohesion as Concentration: Exclusion-Driven Fragility in Fin... observed disagreement (reported variance) versus true estimation error in tail r...
Relative to a full-inclusion benchmark, exclusion-based cohesion produces state-contingent bias that is small in normal regimes but grows sharply under regime displacement (tail events).
Analytical comparisons between the exclusion model and a full-inclusion benchmark within the theoretical model; derivations showing bias as a function of regime and exclusion parameters. The result is from model analysis, not empirical data.
high negative Cohesion as Concentration: Exclusion-Driven Fragility in Fin... estimation bias (especially under regime displacement/tail events)
Limitations include possible limited organizational generalizability due to a single Fortune 500 lab context; ABS results depend on model specification/calibration; and operational definitions of 'resilience' and 'planning cycle' require careful reading.
Authors' reported limitations based on study design: single lab context (n = 23), dependence of ABS on model choices, and nontrivial operational definitions.
high negative The Algorithmic Canvas: On the Autopoietic Redefinition of S... generalizability and robustness of study findings
Some declines (in self-efficacy and meaningfulness) from passive AI use persist after participants return to manual work.
Within-experiment assessment of outcomes after participants returned to manual (no-AI) tasks following the AI-use manipulation in the pre-registered experiment (N = 269); reported persistent reductions in self-efficacy and meaningfulness for the passive condition.
high negative Relying on AI at work reduces self-efficacy, ownership, and ... self-efficacy; perceived meaningfulness (measured post-return to manual work)
Passive use of AI reduces perceived meaningfulness of work.
Pre-registered experiment (N = 269) with self-reported measure of work meaningfulness; passive-copy condition showed lower meaningfulness ratings than No-AI and Active-collaboration conditions.
high negative Relying on AI at work reduces self-efficacy, ownership, and ... perceived meaningfulness of work
Passive use of AI reduces psychological ownership of the produced outputs.
Same pre-registered experiment (N = 269). Participants in the passive-copy AI condition reported lower psychological ownership of their outputs (self-report scales) relative to No-AI and Active-collaboration conditions.
high negative Relying on AI at work reduces self-efficacy, ownership, and ... psychological ownership of outputs
Passive use of AI (copying AI-generated output) reduces workers' self-efficacy.
Pre-registered between-subjects experiment (N = 269) using occupation-specific writing tasks. Participants assigned to a passive-copy AI condition reported lower self-efficacy (self-reported confidence to complete tasks without AI) compared to the No-AI (manual) and Active-collaboration conditions.
high negative Relying on AI at work reduces self-efficacy, ownership, and ... self-efficacy (confidence to complete tasks without AI)
Problem C is the practical difficulty of attributing responsibility and agency across distributed socio-technical systems (robots, algorithms, institutions, humans).
Conceptual diagnosis developed in the paper and exemplified with vignettes from three application domains; defined as an analytic concept rather than empirically measured.
high negative Examining ethical challenges in human–robot interaction usin... ability to attribute responsibility/agency in distributed socio-technical system...
Provider incentives may be misaligned (e.g., optimizing for engagement or test performance instead of durable learning), requiring contracts, regulation, or purchaser design to align incentives.
Consensus from interdisciplinary workshop (50 scholars) highlighting incentive risks and market-design considerations; descriptive, not empirical.
high negative The Future of Feedback: How Can AI Help Transform Feedback t... provider optimization metrics (engagement/test performance) vs. durable learning...
Extensive learner data needed to personalize AI feedback raises privacy and data-governance concerns (consent, storage, usage).
Qualitative consensus from workshop participants (50 scholars) noting data-collection requirements and governance risks; no empirical governance studies included.
high negative The Future of Feedback: How Can AI Help Transform Feedback t... volume/type of learner data collected; privacy risk indicators; compliance with ...
Automated feedback may not capture pedagogical nuances expert teachers use (motivation, socio-emotional cues, complex reasoning), limiting pedagogical fit.
Expert syntheses from the workshop of 50 scholars highlighting limits of automation relative to expert teacher judgment; no empirical comparisons presented.
high negative The Future of Feedback: How Can AI Help Transform Feedback t... coverage of socio-emotional and complex-reasoning cues in feedback; corresponden...
AI-generated feedback can be incorrect, misleading, or misaligned with learning objectives; assessing feedback quality is nontrivial.
Repeated concern raised across workshop participants (50 scholars) in qualitative synthesis; noted as a substantive risk and open challenge rather than empirically quantified here.
high negative The Future of Feedback: How Can AI Help Transform Feedback t... feedback factual correctness; alignment with stated learning objectives; rate of...
Exposure to top-rated exemplar papers produced large reductions in interquartile range (IQR) of estimates—within converging measure families, IQR fell by roughly 80–99%.
Stage 3 of the protocol: after agents were shown top-rated exemplar papers, measured within-measure-family IQRs of agents' estimates decreased substantially; reported quantitative reduction range of 80%–99% within measure families that converged.
high negative Nonstandard Errors in AI Agents percentage reduction in interquartile range (IQR) of effect estimates within mea...
Frontier language models and human editors do not reliably reproduce the evaluative signal contained in institutional publication records.
Comparison of zero-shot frontier-model average accuracy (31%) and human-panel majority-vote accuracy (42%) versus fine-tuned models (up to 59% and higher in economics), indicating that neither zero-shot frontier models nor the human panels matched fine-tuned performance on the held-out benchmarks.
high negative Machines acquire scientific taste from institutional traces Relative prediction accuracy on held-out benchmark(s) of research-pitch quality
Eleven frontier language models (proprietary and open) averaged 31% accuracy on a held-out four-tier benchmark of management research pitches (chance ≈25%); this is only marginally above chance.
Zero-shot (or as-provided) evaluation of eleven state-of-the-art language models on the held-out four-tier management pitches benchmark, yielding an average accuracy of 31% versus chance ≈25%. (Exact list of models and number of benchmark examples not provided in the supplied text.)
high negative Machines acquire scientific taste from institutional traces Accuracy on the four-tier management research-pitch benchmark
Generalization across domains and long-term robustness to adversarial adaptation require further validation.
Authors explicitly note the need for further validation; the paper's reported experiments do not (in the provided summary) disclose broad domain coverage, longitudinal tests, or adversarial evolution studies.
high negative CoMAI: A Collaborative Multi-Agent Framework for Robust and ... generalization across domains; long-term robustness to adaptive adversaries
A modular system may increase engineering complexity and compute overhead compared to a single LLM endpoint.
Authors' caveat in the paper noting higher engineering and compute costs as a trade-off for modularity; the summary does not provide quantitative cost or latency measurements.
high negative CoMAI: A Collaborative Multi-Agent Framework for Robust and ... engineering complexity and compute/resource overhead
Quality of CoMAI depends on rubric design and on how the finite-state machine and agent prompts are specified.
Authors' noted limitation/caveat in the paper that system performance hinges on rubric and prompt/FSM design choices; this is a qualitative dependency rather than an empirically quantified effect in the summary.
high negative CoMAI: A Collaborative Multi-Agent Framework for Robust and ... assessment quality as a function of rubric/FSM/agent prompt design
Using C.A.P. entails trade-offs: potential increases in latency and compute cost and a risk of over-correction (unnecessary clarification).
Paper explicitly notes these trade-offs as part of the design discussion and proposes measuring latency, compute cost, and unnecessary clarification rate in evaluations; this is an acknowledged design risk rather than an empirically quantified result.
high negative A Context Alignment Pre-processor for Enhancing the Coherenc... response latency, compute cost per session, rate of unnecessary clarifications