The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests About 🎲 Workforce Futures
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Evidence (7560 claims)

Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.

The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).

Browse by theme

Nine broad, paper-level topics. Click one to filter the claims below.

Adoption
9875 claims
Filter claims →
Productivity
8807 claims
Filter claims →
Governance
7870 claims
Filter claims →
Human-AI Collaboration
7560 claims
Filtered →
Org Design
4892 claims
Filter claims →
Innovation
4781 claims
Filter claims →
Labor Markets
4004 claims
Filter claims →
Skills & Training
3308 claims
Filter claims →
Inequality
2332 claims
Filter claims →

Claims by outcome category

Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.

Outcome Positive Negative Mixed Null Total
Other 870 233 116 1066 2363
Governance & Regulation 976 451 218 133 1809
Organizational Efficiency 949 224 144 88 1416
Technology Adoption Rate 764 287 141 122 1325
Research Productivity 501 152 74 362 1101
Output Quality 542 216 69 69 896
Decision Quality 387 198 94 54 740
Firm Productivity 513 67 101 27 714
AI Safety & Ethics 249 303 73 36 667
Market Structure 190 192 134 27 548
Task Allocation 243 77 91 36 452
Innovation Output 291 33 55 20 401
Skill Acquisition 206 72 65 21 364
Employment Level 133 63 115 22 335
Fiscal & Macroeconomic 153 79 52 32 323
Task Completion Time 206 37 12 15 272
Firm Revenue 179 52 29 5 266
Consumer Welfare 130 76 47 13 266
Inequality Measures 48 137 51 6 242
Worker Satisfaction 101 81 25 13 220
Error Rate 84 110 11 5 210
Wages & Compensation 98 47 30 10 185
Regulatory Compliance 88 73 17 7 185
Automation Exposure 66 64 33 16 182
Team Performance 105 29 30 11 176
Training Effectiveness 109 22 14 21 168
Developer Productivity 114 21 14 8 158
Job Displacement 12 90 24 1 127
Hiring & Recruitment 57 9 9 5 80
Skill Obsolescence 6 56 9 1 72
Social Protection 43 17 8 2 70
Creative Output 35 21 9 4 70
Labor Share of Income 18 21 17 1 57
Worker Turnover 15 16 4 35
Industry 1 1
Clear
Human Ai Collab Remove filter
We developed Pocket MonstARs, a controlled gamified abstraction of HRC warehouse inventory picking in which virtual monsters serve as proxies for pick targets, while labeled and object-marked boxes preserve the real-world identification demands of the picking task.
Methods section / system description in the paper describing the experimental testbed used for the user study (Pocket MonstARs).
high neutral ARTOO-DARTU: Studying AR-HRC With AR Obstruction Mitigation ... experimental testbed design (game abstraction fidelity to picking task constrain...
We conducted a small-scale randomized experiment to measure uplift from human-agent collaboration on real-world computational reproducibility tasks.
Randomized experiment described in the paper (authors report it was small-scale; details in methods section).
high neutral Life After Benchmark Saturation: A Case Study of CORE-Bench human-agent collaboration uplift (measured via task completion time and success)
SEC S-1 filings combine historical financial statements, governance structures, pro forma and common-control accounting treatments, capital-formation narratives, and underwriting-sensitive risk disclosures within substantially longer documents than typical periodic filings.
Descriptive claim in paper about the composition and length of S-1 filings relative to periodic filings.
high neutral IPO Finance Agent: Evaluation of LLM Financial Analysts beyo... document complexity and length
Finance Agent v2 narrowly deals with periodic reporting from publicly traded companies (SEC 10-K and 10-Q filings).
Direct statement in paper describing Finance Agent v2's task scope as focused on periodic reporting (SEC 10-K and 10-Q).
high neutral IPO Finance Agent: Evaluation of LLM Financial Analysts beyo... benchmark task scope (periodic filings)
FacProcessTwin was evaluated in a real-world case study of an Australian food manufacturer covering 16 production process flows spanning chilled, frozen, and aseptic shelf-stable product categories and including process variations within the same product.
Case study description in paper; sample explicitly stated as 16 production process flows.
high neutral FacProcessTwin: An LLM-Based System for Process Twin Develop... scope and diversity of evaluation (number and types of process flows covered)
AI's effects on writing from the reader side are distinct from those on the production (writer) side.
Empirical results showing readers' accusatory behavior and in-group gatekeeping dynamics operate separately from producers' use of generative AI; argued by contrasting detection/production literature with current reader-side observations.
high neutral "That's AI Slop, You Bot!" Studying Accusations, Evidence, a... differences between reader-side responses and writer-side production effects
This research extends signaling theory by showing that substitute signals used socially can grow even when inaccurate if the underlying detection problem cannot be solved at the non-expert level.
Theoretical inference supported by empirical pattern: proliferation of inaccurate yet socially useful accusations (e.g., 'AI slop') in the absence of reliable non-expert detection.
high neutral "That's AI Slop, You Bot!" Studying Accusations, Evidence, a... social signalling dynamics / theoretical extension
We close by sketching the shape of a harmonised tiered framework and the empirical evaluation needed to calibrate it.
Statement of the paper's concluding contribution in the abstract: a proposed harmonised framework and recommended empirical evaluation; presented as a proposal rather than tested result.
high neutral Regulating the Machine Contributor: Governance and Policy Al... proposed harmonised framework and specification of needed empirical evaluation
From this we derive a six-dimensional taxonomy (disclosure, responsibility, human oversight, licensing, enforcement, maintainer workload), an ordinal Policy Maturity Score, and a mapping of documented agent incidents onto the dimensions each policy fails to govern.
Stated as study outputs in abstract; indicates development of taxonomy, maturity score, and incident-to-dimension mapping based on the comparative analysis and process tracing.
high neutral Regulating the Machine Contributor: Governance and Policy Al... policy taxonomy completeness and policy maturity (ordinal score); mapping of inc...
We compare policies across six organisations (SymPy, LLVM, matplotlib, OpenInfra, the Apache Software Foundation, and the Linux Foundation) using Most-Similar Systems Design with indicator-based coding and process tracing for SymPy and LLVM.
Explicit methods statement in abstract describing sample (six organisations) and methods (Most-Similar Systems Design, indicator-based coding, process tracing for two organisations).
high neutral Regulating the Machine Contributor: Governance and Policy Al... policy characteristics across six open-source organisations
Open-source software, however, evolves through a process designed for humans: contributor agreements, codes of conduct, and review norms all assume a legally accountable person who can attest to provenance and answer reviewer questions.
Descriptive claim about standard open-source governance and contribution processes stated in the abstract; no empirical measurement or coded sample cited in the abstract.
high neutral Regulating the Machine Contributor: Governance and Policy Al... design assumptions of open-source contribution processes (legal accountability/p...
We run a pre-specified 2*4 factorial experiment with 280 complete research runs across four datasets.
Experimental design and sample size reported in paper: 2x4 factorial experiment, 280 complete research runs, four datasets.
high neutral (Human) Attention Is (Still) All You Need: Human oversight m... number of complete research runs (experimental sample)
We evaluate the system in two controlled human-subject experiments comparing AI-based pre-mediation with professional human mediators in a multi-issue negotiation scenario.
Reported experimental design: two controlled human-subject studies conducted by the authors (details and comparisons described in paper).
high neutral Automated Mediator for Human Negotiation: Pre-Mediation via ... comparative evaluation between AI-mediated and human-mediated pre-mediation
The pipeline's components are not autonomous and do not interact peer-to-peer; outputs are passed forward in a fixed sequence (single-party pipeline).
Implementation detail described in the paper (system architecture specification).
high neutral Automated Mediator for Human Negotiation: Pre-Mediation via ... pipeline execution model (fixed-sequence, non-autonomous modules)
We discuss design tradeoffs, failure modes, and lessons learned from operating autonomous AI agents at scale.
Paper statement indicating inclusion of discussion sections on tradeoffs, failure modes, and operational lessons; descriptive/meta claim about paper content.
high neutral Autonomous Incident Resolution at Hyperscale: An Agentic AI ... discussion of design tradeoffs, failure modes, and lessons learned
The core problem is not the absence of explanation but the absence of structured reasoning in the first place.
Conceptual argument/proposed reframing presented in the paper; no empirical test reported.
high neutral Beyond Post-hoc Explanation: Toward Glassbox AI via Probabil... presence of structured reasoning vs. post-hoc explanation
We ran a controlled three-arm ablation on a production valuation agent: A = plain web-only LLM analyst; B = adds public structured tools + a 14-dimension valuation playbook, verifier, objectivity policy and red-team; C = adds the proprietary Noah AI corpus of curated pipeline, trial and deal intelligence.
Description of experimental arms and setup used in the study (methodological statement).
high neutral AI Scientists Are Only as Good as Their Evidence: A Stratifi... experimental treatment definitions (method)
The evidence base was concentrated in system-facing applications that detect or shape inequities within recruitment, evaluation and exposure systems.
Synthesis result from the scoping review indicating thematic concentration across included studies (as reported in abstract).
high neutral Artificial intelligence applications supporting women’s care... focus of existing empirical studies (system-facing vs individual-facing applicat...
ALE is organized around a task taxonomy with 55 subfields grouped into 13 industry clusters covering 1K+ tasks.
Author-provided counts describing the benchmark taxonomy and task pool.
high neutral Agents' Last Exam taxonomy breadth (subfields, clusters, number of tasks)
ALE covers non-physical industries defined with reference to O*NET / SOC 2018 (the U.S. federal occupational taxonomy).
Design specification described in the paper referencing O*NET / SOC 2018.
high neutral Agents' Last Exam scope of industries covered by the benchmark
We evaluated seven models (including Gemini, Claude, and GPT families) by comparing their zero-shot estimates against self-reported skill ratings from 27 participants.
Method description: evaluation of seven LLMs comparing zero-shot model estimates to self-reported skill ratings; 27 participants provided self-reports.
high neutral Can AI Guess What You Know? Performance Comparison of Large ... comparison between model zero-shot skill estimates and self-reported skill ratin...
AI deployment should be evaluated not only by average task speed, but by its overall effects on congestion, rework, and the robustness of human oversight under load.
Policy/recommendation based on the paper's theoretical results and derived implications from the queueing model (conceptual/prescriptive conclusion; no empirical testing reported).
high neutral Queue & AI: When Faster Tasks Slow Down the Workflow organizational_efficiency
The divergence between mean task speed and system-level delay caused by AI assistance is labeled the 'variance wedge'.
Definition/terminology introduced in the paper as part of its conceptual framing; supported by the analytic model description.
high neutral Queue & AI: When Faster Tasks Slow Down the Workflow task_completion_time
The benchmark probes 18 mainstream LLMs across four prompting strategies.
Benchmark experiments described in the paper evaluate 18 mainstream LLMs using four different prompting strategies applied to the collected dataset.
high neutral Benchmarking LLMs for Community Governance Simulation with L... coverage of models and prompting strategies in benchmark (number of LLMs and pro...
GENSTRAT generates a distribution of two-player zero-sum imperfect-information card games.
Design specification in paper; reported generated pool size of 2,000 games (abstract).
high neutral GENSTRAT: Toward a Science of Strategic Reasoning in Large L... game distribution (two-player zero-sum imperfect-information card games)
The paper provides a taxonomy of minimum input artifacts for agentic software, firmware, and hardware work; a conversation-to-contract gate; risk-adaptive workflows; and an evidence-bundle acceptance model for agent-generated artifacts.
Declared contributions in the paper (deliverables/artefacts produced by the research; no empirical validation provided in the abstract).
high neutral Agentic Agile-V: From Vibe Coding to Verified Engineering in... availability of process artifacts and workflow models for agentic engineering
The central problem for agentic engineering is no longer prompt engineering; it is engineering process control.
Argument and synthesis presented by the paper (conceptual claim based on reviewed evidence).
high neutral Agentic Agile-V: From Vibe Coding to Verified Engineering in... primary bottleneck affecting agentic engineering effectiveness (process control ...
The results define three operating regimes.
Summary claim in results/conclusions indicating categorization of outcomes into three regimes.
high neutral Cross-domain benchmarks reveal when coordinated AI agents im... classification into operating regimes
We performed a large-scale evaluation spanning 15,000 messages with cross-model validation across six LLMs from three families (OpenAI, Anthropic, Google), totaling 1,440 queries.
Study design and reported sample sizes and model counts provided in the paper.
high neutral Episodic-Semantic Memory Architecture for Long-Horizon Scien... evaluation sample size and cross-model coverage
Experiments are run with and without access to Causely under two scenarios: an active incident and a healthy baseline.
Methodological description in the paper describing the two experimental conditions (with/without Causely) and two scenarios (active incident, healthy baseline).
high neutral Causely: A Causal Intelligence Layer for Enterprise AI A Ben... experimental condition (Causely vs. no Causely) across two scenarios
Experiments compare four agent configurations (Claude Code, OpenAI Codex, HolmesGPT with Sonnet and Gemini backends).
Methodological description listing the four agent configurations used in experiments.
high neutral Causely: A Causal Intelligence Layer for Enterprise AI A Ben... agent configuration comparisons
We evaluate this value proposition through a benchmark study conducted in a controlled setting with injected faults in a 24-microservice OpenTelemetry demo application.
Methodological description in the paper specifying a controlled benchmark with an OpenTelemetry demo application composed of 24 microservices.
high neutral Causely: A Causal Intelligence Layer for Enterprise AI A Ben... benchmark evaluation setup (24-microservice demo with injected faults)
We introduce a taxonomy organized by influence tier, corresponding to interventions on progressively more latent variables: product mentions, information framing, behavioral redirection, and long-term preference shaping.
Paper contribution: authors present a four-tier taxonomy as a conceptual framework; this is a descriptive/constructive claim about the content of the paper itself.
high neutral Generative AI Advertising as a Problem of Trustworthy Commer... categorization of types of commercial influence in generative systems
This study presents a sociotechnical audit of six commercial LLMs by comparing their reasoning with a Delphi-derived rubric constructed from the responses of twenty infrastructure professionals.
Method: sociotechnical audit comparing six commercial LLMs to a rubric created via a Delphi process with 20 infrastructure professionals (Delphi-derived rubric).
high neutral Governance risks of AI reasoning in urban infrastructure thr... comparison of LLM reasoning to expert-derived rubric
The framework reframes the central question of autonomous software engineering from whether a foundation model can produce a patch to whether the model-harness-environment system can produce a verifiably correct, attributed, and maintainable change.
Conceptual reframing and argument presented in the abstract as a conclusion of the proposed framework and evaluation approach.
high neutral AI Harness Engineering: A Runtime Substrate for Foundation-M... ability of the overall system (model+harness+environment) to produce verifiably ...
We formalize this substrate as 'AI Harness Engineering' and identify eleven component responsibilities: task specification, context selection, tool access, project memory, task state, observability, failure attribution, verification, permissions, entropy auditing, and intervention recording.
Methodological/conceptual contribution described in the paper (abstract) that lists eleven component responsibilities as part of the formalization.
high neutral AI Harness Engineering: A Runtime Substrate for Foundation-M... completeness and scope of responsibilities required for a runtime harness
A symmetric six-gate producer audit separates LLM-engineering failures (template collapse, refusal, internal-ID leakage) from genuine commercial steering.
Methodological claim describing a six-gate producer audit procedure in the paper to diagnose engineering failures vs. commercial steering.
high neutral TourMart: A Parametric Audit Instrument for Commission Steer... ability to distinguish engineering failures from commercial steering
Holding the traveler and bundle fixed, the steering delta is read off between a commission-aware prompt and a minimum-disclosure factual template (paired counterfactual).
Method description of the paired counterfactual experimental design used by TourMart.
high neutral TourMart: A Parametric Audit Instrument for Commission Steer... steering delta (difference in acceptance between commission-aware and minimum-di...
We propose TourMart, an applied intelligent-system audit instrument for LLM-OTA commission governance, driven by two governance levers — lambda (gain on message-induced perception) and kappa (budget-normalized cap on how far the message can shift perceived welfare).
Methodological proposal described in paper: design of an audit instrument and two formal levers (lambda, kappa).
high neutral TourMart: A Parametric Audit Instrument for Commission Steer... audit instrument capability for measuring message-induced perception shifts unde...
Online travel agents (Booking, Trip.com, Expedia) have replaced ranked-list interfaces with conversational LLM agents that compress many options into one sentence of advice.
Descriptive assertion in paper about product/industry UI change; no empirical sample or formal measurement reported in excerpt.
high neutral TourMart: A Parametric Audit Instrument for Commission Steer... interface format (ranked-list → single-sentence conversational recommendation)
We position DAO-governed decentralized physical infrastructure networks (DePIN) within a vertically integrated stack that links energy and sensing to connectivity, storage/compute, models, and robots.
Architectural/framework description in the paper that maps DePIN elements into a vertically integrated stack; conceptual/mapping method without empirical measurement.
high neutral DAO-enabled decentralized physical AI: A new paradigm for hu... conceptual integration of DePIN components into a vertical infrastructure stack
We evaluate 4 popular agent harnesses and 7 foundation models on Workspace-Bench.
Experimental setup reported in the paper listing 4 agent harnesses and 7 foundation models used in evaluations.
high neutral Workspace-Bench 1.0: Benchmarking AI Agents on Workspace Tas... number of agent harnesses and foundation models evaluated
Weight-based memory generalizes by applying abstract rules to inputs never seen before.
Conceptual claim grounded in the paper's theoretical distinction between weight-based learning and retrieval; references Complementary Learning Systems theory; no empirical sample in abstract.
high neutral Contextual Agentic Memory is a Memo, Not True Memory type of generalization performed by weight-based memory
Retrieval generalizes by similarity to stored cases.
Conceptual claim stated in paper (distinction between retrieval-based and weight-based generalization); supported by theoretical characterization, not empirical data in abstract.
high neutral Contextual Agentic Memory is a Memo, Not True Memory type of generalization performed by retrieval systems
The process of synthesizing information is inherently iterative: users explore content, identify relationships between concepts, and continuously reorganize their mental models.
Conceptual description of the cognitive/process characteristics in the paper's background/motivation (no empirical measurement reported).
high neutral MindTrellis: Co-Creating Knowledge Structures with AI throug... iterative nature of knowledge synthesis (exploration, relation identification, r...
Many practical machine learning applications are online and sequential, meaning prior decisions inform future ones — a setting in which fairness challenges differ from standard supervised learning.
Background claim in the paper motivating the work; literature context and conceptual discussion rather than new empirical data.
high neutral Fairness under uncertainty in sequential decisions characterization of ML application setting (online/sequential)
The paper establishes a taxonomy of forgetting mechanisms: passive decay-based, active deletion-based, safety-triggered, and adaptive reinforcement-based.
Explicit taxonomy presented in paper (listed in abstract).
high neutral FSFM: A Biologically-Inspired Framework for Selective Forget... classification of forgetting mechanisms
We evaluate Aether over synthetic network change scenarios covering main classes of network changes and on past incidents from a major ISP operational network.
Evaluation methodology stated in paper abstract: tested on synthetic scenarios and historical incidents from one major ISP (no numeric sample size provided in abstract).
high neutral Aether: Network Validation Using Agentic AI and Digital Twin evaluation dataset composition (synthetic scenarios + past ISP incidents)
Expert assessment involved three senior academics producing reports and appointment-level syntheses.
Paper states that three senior academics produced assessment reports and synthesised appointment-level recommendations; n=3 assessors.
high neutral The Relic Condition: When Published Scholarship Becomes Mate... expert assessment procedure (number and type of assessors)
The distillation pipeline used an eight-layer extraction method and a nine-module skill architecture grounded in local, closed-corpus analysis.
Methods description in paper specifying an eight-layer extraction approach and nine-module skill architecture; presented as the technical design of the distillation pipeline.
high neutral The Relic Condition: When Published Scholarship Becomes Mate... pipeline architecture (layers/modules)