The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests About 🎲 Workforce Futures
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Evidence (8974 claims)

Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.

The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).

Browse by theme

Nine broad, paper-level topics. Click one to filter the claims below.

Adoption
10085 claims
Filter claims →
Productivity
8974 claims
Filtered →
Governance
8062 claims
Filter claims →
Human-AI Collaboration
7749 claims
Filter claims →
Org Design
5057 claims
Filter claims →
Innovation
4896 claims
Filter claims →
Labor Markets
4088 claims
Filter claims →
Skills & Training
3372 claims
Filter claims →
Inequality
2377 claims
Filter claims →

Claims by outcome category

Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.

Outcome Positive Negative Mixed Null Total
Other 882 244 117 1097 2424
Governance & Regulation 1010 469 229 135 1875
Organizational Efficiency 977 235 149 90 1462
Technology Adoption Rate 781 299 143 128 1362
Research Productivity 506 155 74 363 1110
Output Quality 555 219 71 70 915
Decision Quality 395 200 95 54 751
Firm Productivity 523 67 101 27 724
AI Safety & Ethics 262 309 75 36 688
Market Structure 195 201 135 30 566
Task Allocation 248 77 96 38 464
Innovation Output 300 34 55 20 411
Skill Acquisition 207 75 65 21 368
Employment Level 138 67 119 24 350
Fiscal & Macroeconomic 156 80 53 33 329
Task Completion Time 211 38 13 16 280
Firm Revenue 183 52 29 5 270
Consumer Welfare 131 77 48 13 269
Inequality Measures 50 141 54 9 254
Worker Satisfaction 104 85 25 13 227
Error Rate 87 112 11 5 215
Automation Exposure 69 69 37 20 198
Wages & Compensation 102 49 31 11 193
Team Performance 115 30 30 11 187
Regulatory Compliance 88 74 17 7 186
Training Effectiveness 109 22 14 21 168
Developer Productivity 116 21 15 8 161
Job Displacement 12 92 26 1 131
Hiring & Recruitment 57 12 9 5 83
Skill Obsolescence 6 59 10 2 77
Social Protection 43 17 8 2 70
Creative Output 35 21 9 4 70
Labor Share of Income 18 23 17 1 59
Worker Turnover 15 16 4 35
Industry 1 1
Clear
Productivity Remove filter
A formative study (N = 8) and a within-subjects summative evaluation (N = 16) comparing Pista to a baseline agent demonstrated that active participation in execution influenced not only task outcomes but also users' comprehension of the task, their perception of the agent, and their sense of role within the workflow.
Empirical evaluation consisting of a formative study with N=8 and a within-subjects summative evaluation with N=16 comparing Pista to a baseline agent (authors report influence on task outcomes, comprehension, perception, and role).
high positive Auditing and Controlling AI Agent Actions in Spreadsheets task outcomes (primary claim), plus user comprehension, perception, and role sen...
We introduce Pista, a spreadsheet AI agent that decomposes execution into auditable, controllable actions, providing users with visibility into the agent's decision-making process and the capacity to intervene at each step.
System description / design contribution presented by the authors (implementation description rather than empirical evidence).
high positive Auditing and Controlling AI Agent Actions in Spreadsheets availability of auditable, controllable actions and ability to intervene
Selective forgetting should be considered a fundamental capability for next-generation LLM agents operating in real-world, resource-constrained scenarios.
Conclusion/argument in paper based on conceptual analysis and reported empirical benefits.
high positive FSFM: A Biologically-Inspired Framework for Selective Forget... necessity of selective forgetting for future LLM agents
The work bridges cognitive neuroscience (hippocampal indexing/consolidation theory and Ebbinghaus forgetting curve) and AI systems to inform forgetting mechanisms.
Claimed theoretical grounding and cross-disciplinary framing in paper (stated in abstract).
high positive FSFM: A Biologically-Inspired Framework for Selective Forget... theoretical alignment between neuroscience and AI forgetting mechanisms
Empirical results show security performance with 100% elimination of security risks.
Reported experimental result in abstract claiming full elimination of security risks.
high positive FSFM: A Biologically-Inspired Framework for Selective Forget... security risk elimination
Empirical results show content quality improved by +29.2% signal-to-noise ratio.
Reported experimental result in abstract (signal-to-noise ratio improvement).
high positive FSFM: A Biologically-Inspired Framework for Selective Forget... content quality (signal-to-noise ratio)
Empirical results show access efficiency improved by +8.49%.
Reported experimental result in abstract.
Building on advances in LLM agent architectures and vector databases, the paper presents detailed specifications, implementation strategies, and empirical validation from controlled experiments.
Methodological claim in abstract indicating implementation and controlled experiments (no experimental details in abstract).
high positive FSFM: A Biologically-Inspired Framework for Selective Forget... presence of implementation details and experimental validation
Selective forgetting improves security through active forgetting of malicious inputs, sensitive data, and privacy-compromising content.
Authors' taxonomy and safety-triggered forgetting mechanism; abstract reports empirical security performance (100% elimination of security risks).
high positive FSFM: A Biologically-Inspired Framework for Selective Forget... security performance (elimination of security risks)
Selective forgetting improves content quality by dynamically updating outdated preferences and context.
Conceptual claim supported by authors' implementation and empirical validation; abstract reports content quality improvement (signal-to-noise ratio).
high positive FSFM: A Biologically-Inspired Framework for Selective Forget... content quality (signal-to-noise ratio)
A well-designed forgetting mechanism improves efficiency via intelligent memory pruning.
Claim supported by authors' framework and controlled experiments reported in the paper (abstract references empirical results for access efficiency).
In resource-constrained environments, a well-designed forgetting mechanism is as crucial as remembering.
Argument and conceptual analysis in paper; motivated by theoretical considerations and (claimed) empirical validation.
high positive FSFM: A Biologically-Inspired Framework for Selective Forget... relative importance of forgetting vs remembering for system performance
The findings point to a staged progression of AI utility from low-consequence assistance toward higher-order automation, as trust, infrastructure, and verification mature.
Synthesis of interview responses (over 30) indicating current use cases are lower-risk assistance and that stakeholders expect (or prefer) gradual progression toward automation contingent on trust/infrastructure/verification improvements.
high positive Agentic AI in Engineering and Manufacturing: Industry Perspe... trajectory of AI deployment (from assistance to automation) conditional on matur...
Reliability, verification, and auditability are central requirements for adoption, driving human-in-the-loop frameworks and governance aligned with existing engineering reviews.
Consistent themes from interviews (over 30) indicating stakeholders prioritize reliability, verifiability, and audit trails, leading to preference for human-in-the-loop designs integrated with current review processes.
high positive Agentic AI in Engineering and Manufacturing: Industry Perspe... requirements driving adoption decisions (reliability, verification, auditability...
Higher-value agentic gains come from orchestrating multi-step workflows across tools.
Observed and reported in interviews (over 30) with stakeholders in engineering and manufacturing workflows describing value from agentic orchestration across tools.
high positive Agentic AI in Engineering and Manufacturing: Industry Perspe... value generated by agentic AI when coordinating multi-step toolchains
Near-term AI gains cluster around structured, repetitive work and data-intensive synthesis.
Qualitative findings from an exploratory state-of-practice study based on over 30 semi-structured interviews across four stakeholder groups (large enterprises, small/medium firms, AI developers, and CAD/CAM/CAE vendors).
high positive Agentic AI in Engineering and Manufacturing: Industry Perspe... locations/types of tasks where AI provides near-term value (structured/repetitiv...
SWE-chat is a living dataset; our collection pipeline automatically and continually discovers and processes sessions from public repositories.
Description of the dataset collection infrastructure and pipeline provided in the paper; operational behavior asserted by authors.
high positive SWE-chat: Coding Agent Interactions From Real Users in the W... dataset collection process (automated, continual discovery from public repositor...
The dataset currently contains 6,000 sessions, comprising more than 63,000 user prompts and 355,000 agent tool calls.
Descriptive statistics reported by the authors based on their dataset collection pipeline (dataset metadata).
high positive SWE-chat: Coding Agent Interactions From Real Users in the W... dataset size (sessions, prompts, agent tool calls)
We present SWE-chat, the first large-scale dataset of real coding agent sessions collected from open-source developers in the wild.
Paper authorship / dataset description; dataset curated and presented by the paper as a contribution. No external validation provided in excerpt.
high positive SWE-chat: Coding Agent Interactions From Real Users in the W... existence and scale of the SWE-chat dataset (novel dataset release)
Statelessness is the load-bearing property explaining enterprises' preference for weaker but replayable retrieval pipelines, and DPM demonstrates this property is attainable without the decisioning penalty retrieval pays.
Synthesis/conclusion based on theoretical argument and empirical results presented (architectural analysis + experiments showing DPM performance and auditability).
high positive Stateless Decision Memory for Enterprise AI Agents trade-off between stateless architectures and decisioning performance / auditabi...
The audit surface follows the same one-versus-N pattern: DPM logs two LLM calls per decision while summarization logs 83-97 on LongHorizon-Bench.
Empirical measurement on LongHorizon-Bench reported in the paper: logged LLM calls per decision are 2 for DPM vs 83-97 for summarization.
high positive Stateless Decision Memory for Enterprise AI Agents number of LLM calls logged per decision (audit surface)
DPM is additionally 7-15x faster at binding budgets, making one LLM call at decision time instead of N.
Empirical runtime/efficiency measurement reported in the paper (range 7-15x speedup) comparing number of LLM calls and latency under tight memory budgets.
high positive Stateless Decision Memory for Enterprise AI Agents decision-time latency / number of LLM calls
At a 20x compression ratio, DPM improves reasoning coherence by +0.53 (Cohen's h=1.13, p=0.0034) compared to summarization-based memory (paired permutation, n=10).
Paired permutation test over 10 cases at a 20x compression ratio; reported effect +0.53 with Cohen's h=1.13 and p=0.0034.
high positive Stateless Decision Memory for Enterprise AI Agents reasoning coherence
At a 20x compression ratio, DPM improves factual precision by +0.52 (Cohen's h=1.17, p=0.0014) compared to summarization-based memory (paired permutation, n=10).
Paired permutation test over 10 cases at a 20x compression ratio; reported effect +0.52 with Cohen's h=1.17 and p=0.0014.
high positive Stateless Decision Memory for Enterprise AI Agents factual precision
On ten regulated decisioning cases at three memory budgets, DPM matches summarization-based memory at generous budgets and substantially outperforms it when the budget binds.
Empirical evaluation on 10 decisioning cases across three memory budgets; comparison between DPM and summarization-based memory as reported in the paper (n=10).
high positive Stateless Decision Memory for Enterprise AI Agents relative performance (match/outperform) of DPM vs summarization-based memory acr...
We propose Deterministic Projection Memory (DPM): an append-only event log plus one task-conditioned projection at decision time.
Method/architectural proposal described in the paper.
high positive Stateless Decision Memory for Enterprise AI Agents architecture design (DPM specification)
Long-term prospects of agentic AI include catalyzing accelerated innovation in physical design via autonomous algorithm discovery, continuous tool improvement, and closed-loop learning from large design corpora.
Forward-looking conclusion in the paper; framed as the authors' projection based on survey synthesis rather than as an empirically demonstrated outcome in the abstract.
high positive Invited: Agentic AI for Physical Design R&D: Status and Pros... autonomous algorithm discovery, continuous tool improvement, closed-loop learnin...
Interfaces between agentic systems and traditional EDA frameworks are a key area of focus and enable tighter integration of agent capabilities into existing design workflows.
Survey highlights interfaces between agents and EDA frameworks as a focus area; claim is descriptive of research direction rather than reporting empirical outcomes.
high positive Invited: Agentic AI for Physical Design R&D: Status and Pros... development and importance of interfaces between agents and EDA frameworks
Autonomous agents can explore heuristic spaces for placement, routing, and partitioning, enabling autonomous exploration of design heuristics.
Presented as an emphasized capability/area of research in the survey; the abstract asserts this possibility but does not report empirical benchmarks or sample sizes.
high positive Invited: Agentic AI for Physical Design R&D: Status and Pros... autonomous exploration of heuristic spaces (placement, routing, partitioning)
Tool-integrated agents can be used for algorithm evolution, debugging, and workflow automation in physical design R&D.
Paper emphasizes this as a primary area of application in the survey; rationale and examples are discussed but no quantitative trial sizes are given in the abstract.
high positive Invited: Agentic AI for Physical Design R&D: Status and Pros... use of agents for algorithm evolution, debugging, and workflow automation
Agentic AI systems can comprehend user specifications, modify code, run EDA tools, analyze results, perform multi-step reasoning, and iteratively refine design heuristics—unlike earlier ML uses that focused narrowly on prediction or optimization subroutines.
Descriptive claim in the paper contrasting agentic AI capabilities with earlier ML approaches; presented as an overview of functional capabilities rather than empirical measurement.
high positive Invited: Agentic AI for Physical Design R&D: Status and Pros... breadth of tasks agentic AI systems can perform (spec comprehension, code modifi...
Recent advances in large language models (LLMs) and tool-using autonomous agents present new opportunities for accelerating research and development in physical design.
Stated as a central thesis in the paper's abstract/survey; based on the authors' synthesis of recent advances and emerging applications (no empirical sample or quantified evaluation reported in the abstract).
high positive Invited: Agentic AI for Physical Design R&D: Status and Pros... acceleration of research and development in physical design
The same user study (n=32) reports improvements in subjective measures including fluency and user preference for RAPIDDS over non-adaptive systems.
User study (n=32) reporting subjective questionnaire/ratings (fluency, preference) comparing RAPIDDS vs non-adaptive baselines.
high positive Multi-Cycle Spatio-Temporal Adaptation in Human-Robot Teamin... subjective fluency and user preference
A user study (n=32) shows significant plan improvement compared to non-adaptive systems across objective metrics such as efficiency and proximity.
User study reported in paper with sample size n=32 comparing RAPIDDS to non-adaptive systems on objective metrics (efficiency, proximity); significance claimed.
high positive Multi-Cycle Spatio-Temporal Adaptation in Human-Robot Teamin... efficiency and proximity (objective plan metrics)
An ablation study in simulation and a physical robot scenario demonstrates the importance of dual (task + motion) adaptation.
Ablation experiments reported in paper (simulation and physical robot experiments comparing full RAPIDDS to ablated variants).
high positive Multi-Cycle Spatio-Temporal Adaptation in Human-Robot Teamin... plan performance when removing components (effect of dual adaptation)
RAPIDDS jointly adapts task schedules and steers diffusion models of robot motions to maximize efficiency and minimize proximity accounting for individualized models.
Algorithmic method described in paper combining schedule optimization with motion steering (method section).
high positive Multi-Cycle Spatio-Temporal Adaptation in Human-Robot Teamin... efficiency and proximity of joint plans
In the ICT industry, Tobin's Q significantly increased following AI adoption (heterogeneous positive effect).
Subgroup/heterogeneity analysis within the main sample (KOSDAQ firms 2018–2025), estimating the post-adoption effect of AI on Tobin's Q in firms classified as ICT.
high positive The Dynamic Causal Effects of Corporate AI Adoption on Profi... Tobin's Q (market value) in ICT-industry firms
Our baseline model finds evidence that AI is productivity enhancing.
Results from the paper's stated baseline empirical model using BEA industry-account-based measures; model specification described by authors.
ClawNet enables multiple users to collaborate securely through their respective agents.
Capability claim about the instantiated system (authors assert that ClawNet enables secure multi-user collaboration; excerpt contains no empirical security evaluation or user study).
high positive ClawNet: Human-Symbiotic Agent Network for Cross-User Autono... secure multi-user collaboration enabled by agent-mediated interactions
We instantiate this paradigm in ClawNet, an identity-governed agent collaboration framework that enforces identity binding and authorization verification through a central orchestrator.
Implementation claim: authors state they built ClawNet as an instantiation of their paradigm (paper describes framework/architecture; no experimental evaluation included in excerpt).
high positive ClawNet: Human-Symbiotic Agent Network for Cross-User Autono... existence of an implemented framework (ClawNet) enforcing identity binding and a...
Action-level accountability logs every operation against its owner's identity and authorization, ensuring full auditability.
Design claim describing an accountability primitive (paper asserts logging and auditability as a property; no audit or verification evidence shown in excerpt).
high positive ClawNet: Human-Symbiotic Agent Network for Cross-User Autono... auditability of agent actions (logging tied to owner identity/authorization)
Scoped authorization enforces per-identity access control and escalates boundary violations to the owner.
Design/specification claim describing the scoped authorization governance primitive in the proposed paradigm (no empirical or security evaluation provided in excerpt).
high positive ClawNet: Human-Symbiotic Agent Network for Cross-User Autono... access control enforcement and escalation behavior
The paradigm rests on three governance primitives: (1) a layered identity architecture that separates a Manager Agent from multiple context-specific Identity Agents; the Manager Agent holds global knowledge but is architecturally isolated from external communication.
Architectural/design claim describing the proposed layered identity primitive (presentation of design; no empirical validation in excerpt).
high positive ClawNet: Human-Symbiotic Agent Network for Cross-User Autono... identity architecture and information flow constraints
We propose a human-symbiotic agent paradigm in which each user owns a permanently bound agent system that collaborates on the owner's behalf, forming a network whose nodes are humans rather than agents.
Design proposal / conceptual architecture presented in the paper (no large-scale deployment or empirical evaluation described in excerpt).
high positive ClawNet: Human-Symbiotic Agent Network for Cross-User Autono... structure of agent networks (human-centric vs agent-centric) and delegation mode...
The next frontier for AI agents lies not in stronger individual capability, but in the digitization of human collaborative relationships.
Normative/strategic claim advanced by the authors as the central thesis (conceptual argument, no empirical test reported).
high positive ClawNet: Human-Symbiotic Agent Network for Cross-User Autono... focus of AI-agent development (individual capability vs collaboration digitizati...
Human productivity rests on the social and organizational relationships through which people coordinate, negotiate, and delegate.
Theoretical/argumentative claim presented as background motivation (conceptual reasoning, citation not provided in excerpt).
high positive ClawNet: Human-Symbiotic Agent Network for Cross-User Autono... human productivity as mediated by social/organizational relationships
Time Series Augmented Generation (TSAG) enables LLM agents to delegate quantitative tasks to verifiable external tools.
Description of TSAG framework in paper stating delegation mechanism to external verifiable tools for quantitative computations.
high positive Time Series Augmented Generation for Financial Applications delegation capability to external tools
We publicly release the evaluation framework and empirical insights to foster standardized research on reliable financial AI.
Paper states that the framework, benchmark, and empirical results are released publicly by the authors.
high positive Time Series Augmented Generation for Financial Applications public release of resources
The results demonstrate that capable agents can achieve near-perfect tool-use accuracy with minimal hallucination, validating the tool-augmented paradigm.
Empirical results from the authors' experiments on the 100-question benchmark across multiple agents; paper states agents achieve 'near-perfect' tool-use accuracy and 'minimal' hallucination.
high positive Time Series Augmented Generation for Financial Applications tool-use accuracy; hallucination rate
We apply this methodology in a large-scale empirical study using our framework, Time Series Augmented Generation (TSAG), where an LLM agent delegates quantitative tasks to verifiable, external tools.
Paper reports applying the TSAG framework in an empirical study in which agents call external tools to perform quantitative computations; described as 'large-scale' and implemented by the authors.
high positive Time Series Augmented Generation for Financial Applications use of external/verifiable tools by LLM agents