The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests About 🎲 Workforce Futures
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Evidence (8974 claims)

Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.

The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).

Browse by theme

Nine broad, paper-level topics. Click one to filter the claims below.

Adoption
10085 claims
Filter claims →
Productivity
8974 claims
Filtered →
Governance
8062 claims
Filter claims →
Human-AI Collaboration
7749 claims
Filter claims →
Org Design
5057 claims
Filter claims →
Innovation
4896 claims
Filter claims →
Labor Markets
4088 claims
Filter claims →
Skills & Training
3372 claims
Filter claims →
Inequality
2377 claims
Filter claims →

Claims by outcome category

Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.

Outcome Positive Negative Mixed Null Total
Other 882 244 117 1097 2424
Governance & Regulation 1010 469 229 135 1875
Organizational Efficiency 977 235 149 90 1462
Technology Adoption Rate 781 299 143 128 1362
Research Productivity 506 155 74 363 1110
Output Quality 555 219 71 70 915
Decision Quality 395 200 95 54 751
Firm Productivity 523 67 101 27 724
AI Safety & Ethics 262 309 75 36 688
Market Structure 195 201 135 30 566
Task Allocation 248 77 96 38 464
Innovation Output 300 34 55 20 411
Skill Acquisition 207 75 65 21 368
Employment Level 138 67 119 24 350
Fiscal & Macroeconomic 156 80 53 33 329
Task Completion Time 211 38 13 16 280
Firm Revenue 183 52 29 5 270
Consumer Welfare 131 77 48 13 269
Inequality Measures 50 141 54 9 254
Worker Satisfaction 104 85 25 13 227
Error Rate 87 112 11 5 215
Automation Exposure 69 69 37 20 198
Wages & Compensation 102 49 31 11 193
Team Performance 115 30 30 11 187
Regulatory Compliance 88 74 17 7 186
Training Effectiveness 109 22 14 21 168
Developer Productivity 116 21 15 8 161
Job Displacement 12 92 26 1 131
Hiring & Recruitment 57 12 9 5 83
Skill Obsolescence 6 59 10 2 77
Social Protection 43 17 8 2 70
Creative Output 35 21 9 4 70
Labor Share of Income 18 23 17 1 59
Worker Turnover 15 16 4 35
Industry 1 1
Clear
Productivity Remove filter
Small language models (SLMs) are more cost-efficient and amenable to on-device inference.
Background claim in the paper's introduction/abstract; motivated by literature and the paper's focus on on-device SLMs.
high positive When Cloud Agents Meet Device Agents: Lessons from Hybrid Mu... monetary cost and feasibility of on-device inference
Frontier large language models (LLMs), typically hosted in the cloud, offer strong performance across a wide range of tasks at substantially high cost.
Background claim in the paper's introduction/abstract; supported by literature context and framing rather than a specific experiment in this paper.
high positive When Cloud Agents Meet Device Agents: Lessons from Hybrid Mu... task performance (accuracy/quality) of LLMs and associated monetary cost
Managers can make both (Agentic Technical Debt and Stochastic Tax) visible through lightweight dashboards and governance controls.
Prescriptive/recommendation in the paper; authors state they 'outline' approaches for managers to surface these concepts using dashboards and governance controls. No empirical evaluation or case study evidence reported in the provided excerpt.
high positive Governing Technical Debt in Agentic AI Systems visibility/monitoring of agentic technical debt and operating burden via dashboa...
Using the three metrics (data product adoption, time-to-find, time-to-insight) ties platform success to measurable business value rather than internal activity.
Argument in the paper about metric selection and their role in assessing platform success (methodological rationale).
high positive Beyond the Data Mesh Illusion: Designing Modern AI-augmented... alignment of platform success metrics with business value
A staged framework that shifts ownership from hub to spokes avoids both centralized bottlenecks and uncoordinated decentralization.
Organizational/process recommendation presented in the paper as a way to manage decentralization (design rationale).
high positive Beyond the Data Mesh Illusion: Designing Modern AI-augmented... avoidance of centralized bottlenecks and uncoordinated decentralization (organiz...
Natural-language conversational interfaces democratize access for business users and expose historically underutilized enterprise data.
Proposed UX/interaction benefit asserted in the paper (design claim; no empirical measurement reported in the excerpt).
high positive Beyond the Data Mesh Illusion: Designing Modern AI-augmented... data access and usage by business users (adoption of previously underutilized da...
Large language models (LLMs) that automate governance tasks also lower the barrier for domain practitioners to develop genuine cross-functional expertise spanning business and data engineering, enabling spoke teams to take on greater end-to-end ownership without proportionally increasing their dependence on the hub.
Argument in the paper linking AI/LLM capabilities to skill enablement and reduced hub dependence (conceptual claim; no empirical results in the excerpt).
high positive Beyond the Data Mesh Illusion: Designing Modern AI-augmented... skill acquisition / reduction in dependence on central hub
Domain spokes own business semantics, product backlogs, and local iteration cadence, progressively assuming greater responsibility as they mature (shifting operational ownership outward over time).
Architectural/organizational design element described in the paper (procedural proposal for staged ownership transfer).
high positive Beyond the Data Mesh Illusion: Designing Modern AI-augmented... task allocation and ownership over data product lifecycle
A central hub (Center of Excellence) can provide shared platform services, policy automation, and AI-enabled governance that automatically standardizes data products, generates quality rules, drafts data contracts, and reviews changes for regressions.
Functional capabilities described in the proposed architecture; presented as what the hub component will provide (design/specification).
high positive Beyond the Data Mesh Illusion: Designing Modern AI-augmented... automation and standardization of governance tasks (e.g., quality rules, contrac...
An AI-augmented hub-and-spoke model layered on a modern lakehouse architecture can relax the flexibility-versus-control trade-off inherent in enterprise data platforms.
Proposed architectural solution and theoretical argument in the paper (design proposal; no reported experimental/field results provided in the text excerpt).
high positive Beyond the Data Mesh Illusion: Designing Modern AI-augmented... balance between flexibility (domain self-service) and centralized control (gover...
Ongoing efforts of the initiative aim to incorporate benchmarks that address concerns about bias by considering alternative perspectives and human centered use cases.
Statement of planned/ongoing work in the paper regarding future benchmark inclusion to address bias and human-centered use cases; no empirical results provided.
high positive BEAMS: Benchmarking and Evaluating AI for Modeling and Simul... planned incorporation of bias-aware benchmarks and human-centered use case consi...
Implemented tests include causal translation, model iteration, causal reasoning, conformance, model behavior explanation, suggested model building steps, and suggested model fixes.
Specific list of implemented test categories provided in the paper; descriptive/reporting evidence from the initiative's work.
high positive BEAMS: Benchmarking and Evaluating AI for Modeling and Simul... types/categories of tests implemented
Tests for several distinct categories of evaluation have been implemented and applied to AI tools that support qualitative model building, quantitative model building, and model discussion.
Paper reports that a set of tests have been implemented and applied to AI tools across qualitative and quantitative modeling and discussion; no sample sizes or numeric evaluation results provided in the excerpt.
high positive BEAMS: Benchmarking and Evaluating AI for Modeling and Simul... existence and application of implemented evaluation tests across types of modeli...
A steering group focuses on prioritizing potential benchmarks, while a technical group focuses on implementing the benchmarks in the form of automated tests.
Organizational description in the paper specifying roles (steering group and technical group); no quantitative evaluation reported.
high positive BEAMS: Benchmarking and Evaluating AI for Modeling and Simul... organizational roles for benchmark prioritization and implementation
The open source sd ai project hosted by the initiative establishes transparency and enables contributions to be shared broadly.
Descriptive statement about the open-source project hosted by the initiative; no empirical measures of transparency or contribution sharing provided.
high positive BEAMS: Benchmarking and Evaluating AI for Modeling and Simul... transparency and breadth of contributions enabled by the open source sd ai proje...
The initiative uses open digital and organizational infrastructure to collaboratively evaluate AI tools for modeling and simulation.
Descriptive claim in the paper about organizational approach (open infrastructure and collaborative evaluation); no empirical testing or sample size reported.
high positive BEAMS: Benchmarking and Evaluating AI for Modeling and Simul... use of open infrastructure for collaborative evaluation
The BEAMS Initiative aims to guide the development of AI tools for modeling and simulation toward forms that are responsible and ethical by establishing benchmarks for human centered modeling and simulation practices.
Descriptive statement about the Initiative's stated aims and purpose in the paper; organizational description rather than empirical evidence.
high positive BEAMS: Benchmarking and Evaluating AI for Modeling and Simul... existence and purpose of the BEAMS Initiative (benchmarking for responsible/ethi...
Tools that can automate aspects of modeling practice must complement human expertise, not replace it.
Normative claim made in the paper (argument about human-centered design); no empirical evidence or sample size reported.
high positive BEAMS: Benchmarking and Evaluating AI for Modeling and Simul... relationship between automated modeling tools and human expertise (complementari...
AI tools to support real world decision making must be able to build simulation models that inform their recommendations and render them interpretable.
Normative assertion in the paper (position statement / requirement); no empirical study or sample size reported.
high positive BEAMS: Benchmarking and Evaluating AI for Modeling and Simul... ability of AI tools to build interpretable simulation models that inform recomme...
The agentic future is not predetermined; leaders must both skate to where the puck is going and actively steer it toward a good place, ensuring innovation delivers welfare gains felt by businesses and consumers around the world.
Normative recommendation offered by the authors; based on conceptual argument and interpretation of the framework rather than empirical testing in the excerpt.
high positive From Augmentation to Reconstruction: Guiding the AI Disrupti... policy/leadership influence on welfare distribution of AI-driven innovation
These complementary investments produce the familiar 'productivity J-curve' of general-purpose technologies.
Stated as an economic analogy/claim drawing on general-purpose technology literature; presented as an asserted mechanism rather than shown with new empirical estimates in the excerpt.
high positive From Augmentation to Reconstruction: Guiding the AI Disrupti... productivity trajectory (J-curve) following complementary investments
The most consequential disruption resides in the third stage (Reconstruction) where workflows and markets are rebuilt around delegation, machine-to-machine interaction, continuous monitoring, and auditable constraints.
Theoretical claim in the paper backed by conceptual reasoning and illustrative sector examples; no quantitative evidence provided in the excerpt.
high positive From Augmentation to Reconstruction: Guiding the AI Disrupti... magnitude/importance of disruption arising from Reconstruction-stage changes
The system preserves human agency via override mechanisms.
Design description of the collaborative forecasting system that explicitly includes override controls for human users.
high positive Schnitzel-Prediction: Designing Human-Ai Collaboration For C... preservation of human agency (ability to override algorithmic forecasts)
The paper provides a rigorous blueprint for designing synergistic, trustworthy, and diagnostic operational planning tools, contributing to the discourse on human-AI collaboration and sustainable information systems (IS).
Stated contribution in the paper's conclusions: presentation of a blueprint and implications for human-AI collaboration and sustainable IS.
high positive Schnitzel-Prediction: Designing Human-Ai Collaboration For C... guidance/blueprint for operational planning tool design
Two think-aloud sessions show that human judgment remains critical for high-uncertainty events.
Qualitative evaluation consisting of two think-aloud sessions reported in the paper.
high positive Schnitzel-Prediction: Designing Human-Ai Collaboration For C... importance/role of human judgment in handling high-uncertainty forecasting event...
Algorithmic benchmarking reduced forecast errors by 30% over naive baselines.
Quantitative algorithmic benchmarking reported in the evaluation section of the paper (comparison vs. naive baselines).
Teams interacting with more embodied agents display conversational patterns that more closely resemble human–human dialogue.
Conversational analysis comparing dialogue patterns across teams interacting with different embodiment levels; the abstract reports greater similarity to human–human dialogue for teams with higher embodiment agents, but does not provide the similarity metric values or sample sizes.
high positive Teaming Up with Artificial Agents in Non-routine Analytical ... conversational pattern similarity to human–human dialogue
Human-only teams are more likely to complete all tasks successfully (higher task completion success) than mixed human–AI teams.
Comparison of task completion success between human-only teams and mixed teams in the escape room experiment as reported in the paper; no numerical completion rates provided in the abstract.
high positive Teaming Up with Artificial Agents in Non-routine Analytical ... task completion / success rate
Risk-aware layered automation can materially reduce review bottlenecks created by AI-driven code growth without compromising production safety.
Synthesis conclusion based on RADAR deployment results, telemetry (535K+ reviewed diffs, 331K+ landed), and comparative analyses (before-after and difference-in-differences) reported in the paper.
high positive Automating Low-Risk Code Review at Meta: RADAR, Risk Calibra... reduction in review bottlenecks and preservation of production safety
RADAR reduces median diff review wall time by 35%.
Efficiency outcomes reported via telemetry and difference-in-differences analysis stated in the paper; median diff review wall time reduction reported as 35%. Sample likely drawn from RADAR telemetry (535K+ diffs) though not explicitly stated for this metric in the excerpt.
high positive Automating Low-Risk Code Review at Meta: RADAR, Risk Calibra... median diff review wall time
RADAR reduces median time to close by over 330%.
Efficiency outcomes reported via telemetry and difference-in-differences analysis stated in the paper; median time-to-close reduction reported as 'over 330%'. Underlying sample for efficiency analysis likely from the RADAR telemetry (535K+ diffs), though the excerpt does not give the precise sample for this metric.
high positive Automating Low-Risk Code Review at Meta: RADAR, Risk Calibra... median time to close for diffs
The Production Incident rate for RADAR-reviewed diffs is 1/50 that of non-RADAR diffs.
Comparative observational analysis reported in the paper; production incident rate for RADAR-reviewed diffs compared to non-RADAR diffs, with the relative rate given as 1/50. Exact absolute counts not provided in the excerpt; overall RADAR telemetry covers 535K+ diffs.
high positive Automating Low-Risk Code Review at Meta: RADAR, Risk Calibra... production incident rate (RADAR vs non-RADAR)
The revert rate for RADAR-reviewed diffs is 1/3 that of non-RADAR diffs.
Comparative observational analysis reported in the paper contrasting RADAR-reviewed diffs with non-RADAR diffs. Underlying counts and exact sample split not provided in the excerpt; overall RADAR telemetry covers 535K+ diffs.
high positive Automating Low-Risk Code Review at Meta: RADAR, Risk Calibra... diff revert rate (RADAR vs non-RADAR)
Relaxing the Diff Risk Score threshold from the 25th to the 50th percentile increased the approve rate to 60.31%.
Policy threshold comparison reported in the paper using observational before-after comparisons and system telemetry; approval rate reported as 60.31% after threshold change.
high positive Automating Low-Risk Code Review at Meta: RADAR, Risk Calibra... approve rate of diffs under RADAR as a function of Diff Risk Score threshold
RADAR has reviewed 535K+ diffs and landed 331K+ changes.
System deployment telemetry reported in the paper: 'RADAR has reviewed 535K+ diffs and landed 331K+.'
high positive Automating Low-Risk Code Review at Meta: RADAR, Risk Calibra... number of diffs reviewed and diffs landed by RADAR
Agentic AI was responsible for over 80% of that growth in code volume.
Attribution analysis reported in the paper linking growth in code/diff volume to agentic AI sources; described as 'over 80% of that growth.' The underlying attribution method is not detailed in the excerpt.
high positive Automating Low-Risk Code Review at Meta: RADAR, Risk Calibra... share of growth in code/diff volume attributable to agentic AI
Per-developer diff volume rose 51% (year over year) at Meta.
Internal telemetry/observational analysis reported in the paper; stated as a 51% increase in per-developer diff volume. No explicit sample size for this specific measure provided in the excerpt.
high positive Automating Low-Risk Code Review at Meta: RADAR, Risk Calibra... per-developer diff volume (year-over-year change)
At Meta, significant lines of code per human-landed diff grew by 105.9% year over year.
Internal telemetry/observational analysis reported in the paper; stated as a year-over-year percentage growth for Meta. No sample size for this specific measure provided in the excerpt.
high positive Automating Low-Risk Code Review at Meta: RADAR, Risk Calibra... lines of code per human-landed diff (year-over-year growth)
We discuss implications for Information Systems (IS) design and propose future field evaluations.
Paper includes a discussion section outlining IS design implications and suggestions for future empirical/field work.
high positive Multi Agent Systems In The Lean Startup Cycle: Operationalis... proposed implications and future research directions
The approach preserves statistical rigour, traceability, and nuanced Persevere/Iterate decisions when accelerating experimentation.
Reported outcomes of controlled simulations and description of system design that enforces statistical procedures and logging; stated in manuscript as findings.
high positive Multi Agent Systems In The Lean Startup Cycle: Operationalis... statistical rigour, traceability, and decision quality in experimentation (Perse...
Logs render capabilities observable at the feature level, turning 'agentic AI' into a disciplined experimentation infrastructure rather than a generic assistant.
Implementation logs and descriptions from the Node.js instantiation reported in the paper; qualitative claim about observability and traceability at the feature level.
high positive Multi Agent Systems In The Lean Startup Cycle: Operationalis... feature-level observability/traceability of experimentation activities
The Multi Agent System reduces time-to-validated-learning by roughly an order of magnitude while preserving statistical rigour, traceability, and nuanced Persevere/Iterate decisions.
Results from the controlled simulations reported in the paper (comparison between agentic multi-agent system and manual B-M-L cycles).
high positive Multi Agent Systems In The Lean Startup Cycle: Operationalis... time-to-validated-learning (and preservation of statistical rigour, traceability...
Controlled simulations compare agentic and manual B-M-L cycles on feature ideas.
Reported controlled simulation experiments in the paper comparing agentic (multi-agent) and manual B-M-L cycles; methodological description present in manuscript.
high positive Multi Agent Systems In The Lean Startup Cycle: Operationalis... comparison of agentic vs manual B-M-L cycles (experimentation performance metric...
We instantiate them in a Node.js package instrumenting a production-grade SaaS codebase.
Implementation artifact reported in the paper (Node.js package) and description of instrumentation on a production-grade SaaS codebase.
high positive Multi Agent Systems In The Lean Startup Cycle: Operationalis... existence and instantiation of a Node.js package that instruments a SaaS codebas...
Drawing on the Dynamic Capabilities View, we derive fifteen meta-requirements and thirty-three design principles (consolidated into seven goal-directed groups) for sensing, seizing, reconfiguring, orchestration, and governance.
Design-theory derivation reported in the paper (counts of meta-requirements and design principles are stated in the manuscript).
high positive Multi Agent Systems In The Lean Startup Cycle: Operationalis... number and organization of derived meta-requirements and design principles
We propose a multi-agent artefact that operationalises the Build–Measure–Learn (B-M-L) cycle as a closed-loop control system.
Design science study described in the paper; conceptual derivation and artifact instantiation (Node.js package) reported in the manuscript.
high positive Multi Agent Systems In The Lean Startup Cycle: Operationalis... operationalisation of the Build–Measure–Learn cycle as a closed-loop control sys...
AI adoption raises real output.
Panel local projections linking establishment-level AI adoption (share of job postings requiring AI skills) to real output across 13 industries over 2017-2025.
AI adoption raises labor productivity.
Panel local projections estimating the effect of establishment-level AI-skill posting share on labor productivity across 13 industries (2017-2025).
The positive impact of AI application on enterprise innovation efficiency is stronger in labor-intensive firms.
Heterogeneity/subsample analysis using the 2012–2023 A-share panel indicating larger AI effects for labor-intensive firms.
high positive Research on the Influence Mechanism of Artificial Intelligen... enterprise innovation efficiency (heterogeneous by labor intensity)
The positive impact of AI application on enterprise innovation efficiency is stronger in asset-intensive firms.
Heterogeneity/subsample analysis on 2012–2023 A-share firm panel showing larger estimated AI effects for firms characterized as asset-intensive.
high positive Research on the Influence Mechanism of Artificial Intelligen... enterprise innovation efficiency (heterogeneous by asset intensity)