The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests About 🎲 Workforce Futures
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Evidence (8974 claims)

Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.

The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).

Browse by theme

Nine broad, paper-level topics. Click one to filter the claims below.

Adoption
10085 claims
Filter claims →
Productivity
8974 claims
Filtered →
Governance
8062 claims
Filter claims →
Human-AI Collaboration
7749 claims
Filter claims →
Org Design
5057 claims
Filter claims →
Innovation
4896 claims
Filter claims →
Labor Markets
4088 claims
Filter claims →
Skills & Training
3372 claims
Filter claims →
Inequality
2377 claims
Filter claims →

Claims by outcome category

Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.

Outcome Positive Negative Mixed Null Total
Other 882 244 117 1097 2424
Governance & Regulation 1010 469 229 135 1875
Organizational Efficiency 977 235 149 90 1462
Technology Adoption Rate 781 299 143 128 1362
Research Productivity 506 155 74 363 1110
Output Quality 555 219 71 70 915
Decision Quality 395 200 95 54 751
Firm Productivity 523 67 101 27 724
AI Safety & Ethics 262 309 75 36 688
Market Structure 195 201 135 30 566
Task Allocation 248 77 96 38 464
Innovation Output 300 34 55 20 411
Skill Acquisition 207 75 65 21 368
Employment Level 138 67 119 24 350
Fiscal & Macroeconomic 156 80 53 33 329
Task Completion Time 211 38 13 16 280
Firm Revenue 183 52 29 5 270
Consumer Welfare 131 77 48 13 269
Inequality Measures 50 141 54 9 254
Worker Satisfaction 104 85 25 13 227
Error Rate 87 112 11 5 215
Automation Exposure 69 69 37 20 198
Wages & Compensation 102 49 31 11 193
Team Performance 115 30 30 11 187
Regulatory Compliance 88 74 17 7 186
Training Effectiveness 109 22 14 21 168
Developer Productivity 116 21 15 8 161
Job Displacement 12 92 26 1 131
Hiring & Recruitment 57 12 9 5 83
Skill Obsolescence 6 59 10 2 77
Social Protection 43 17 8 2 70
Creative Output 35 21 9 4 70
Labor Share of Income 18 23 17 1 59
Worker Turnover 15 16 4 35
Industry 1 1
Clear
Productivity Remove filter
We establish a comprehensive readability model that synthesizes textual, structural, program, and visual features of code.
Description in paper of a newly constructed readability model combining textual, structural, program, and visual features; model development is presented as a methodological contribution (no numeric effect size).
high positive The Readability Spectrum: Patterns, Issues, and Prompt Effec... code_readability (measured via the proposed readability model)
The study demonstrates that recent archival case evidence can be used rigorously to analyze an emerging strategic phenomenon without reducing the study to a purely descriptive literature review.
Methodological claim supported by the paper's demonstration of within-case coding and cross-case pattern matching applied to recent archival documents for the four firms.
high positive Artificial Intelligence Enabled Competitive Intelligence as ... validity and rigor of archival case methods for studying emerging strategic phen...
The paper develops a process view of AIECI built on sensing, interpretation, and orchestration as the sequence through which AI inputs are transformed into competitive intelligence capability, intelligence-informed decisions, and economic outcomes.
Theoretical contribution synthesized from cross-case analysis and conceptual development within the paper.
high positive Artificial Intelligence Enabled Competitive Intelligence as ... conceptual/process model of how AI inputs are transformed into economic outcomes...
Competitive intelligence (the process of sensing, interpreting, and orchestrating responses) rather than AI as a standalone automation tool is the strategic mechanism through which value is created.
Theoretical argument supported by within-case coding and cross-case synthesis of archival materials from four firms demonstrating how AI functions as part of an intelligence infrastructure rather than as isolated automation.
high positive Artificial Intelligence Enabled Competitive Intelligence as ... role of competitive intelligence as the mechanism linking AI inputs to economic ...
Across the four cases, AIECI delivered strategic speed under uncertainty (faster, better-timed decisions in uncertain environments).
Archival case evidence (public disclosures and corporate materials) showing firms using AI-enabled intelligence to accelerate decision cycles and respond more quickly to market signals.
high positive Artificial Intelligence Enabled Competitive Intelligence as ... strategic speed under uncertainty (reduced time-to-decision and faster strategic...
Across the four cases, AIECI improved allocation quality (better targeting and resource allocation decisions).
Within- and cross-case coding of corporate materials from the four sampled firms reporting improvements in campaign targeting, budget allocation, and resource deployment linked to AI-driven intelligence.
high positive Artificial Intelligence Enabled Competitive Intelligence as ... improved allocation quality (better targeting/allocating marketing and operation...
Across the four cases, AIECI produced efficiency gains and cost relief for firms.
Cross-case evidence from archival corporate disclosures and reports for Walmart, Unilever, Sprinklr, and DoubleVerify showing operational/marketing efficiencies and cost savings linked to AI-enabled competitive intelligence.
high positive Artificial Intelligence Enabled Competitive Intelligence as ... efficiency improvements and cost relief (reduced costs or improved resource use ...
Across the four cases, AIECI generated value through revenue acceleration.
Cross-case findings from a qualitative comparative multiple-case design using public archival evidence (annual reports, 10-Ks, earnings releases, corporate materials) for four firms (Walmart, Unilever, Sprinklr, DoubleVerify).
high positive Artificial Intelligence Enabled Competitive Intelligence as ... revenue acceleration (increased sales or faster revenue growth attributed to AIE...
Policy options should centre on building institutional capacity for AGI situational awareness, strengthening Europe's position in the AI value chain, and developing frameworks for international stability in an era of increasingly capable AI systems.
Paper's recommended policy agenda derived from its assessment of risks and gaps (as stated in abstract); the abstract does not report empirical testing of these options or quantified expected effects.
high positive Europe and the Geopolitics of AGI: The Need for a Preparedne... governance_and_regulation
These findings point to a need for a coordinated European preparedness agenda.
Paper's synthesis and policy recommendation based on the identified capability and governance gaps (as stated in abstract); recommendation not supported by quantified impact estimates in the abstract.
high positive Europe and the Geopolitics of AGI: The Need for a Preparedne... governance_and_regulation
A plausible window for AGI emergence falls between 2030 and 2040, or potentially earlier, though substantial uncertainty remains.
Paper's synthesis of empirical trends in AI capabilities, expert forecasting surveys, and policy analysis (as stated in abstract). No specific sample size or survey details provided in the abstract.
Visualizing spatial (localization) uncertainty in the annotation interface improves human-in-the-loop annotation (i.e., localization uncertainty is a lever to improve annotation quality/efficiency).
Synthesis/interpretation in the paper based on the controlled study results (120 participants) and box-level analysis showing improved label quality and reduced time when uncertainty cues were shown.
high positive From Model Uncertainty to Human Attention: Localization-Awar... human-in-the-loop annotation quality and efficiency
A box-level analysis confirms that the uncertainty cues redirect annotator effort toward high-uncertainty predictions and away from well-localized boxes.
Box-level analysis reported in paper comparing annotator behavior across predicted boxes with differing localization uncertainty; analysis shows effort reallocation toward boxes labeled as high-uncertainty.
high positive From Model Uncertainty to Human Attention: Localization-Awar... annotator effort allocation across predicted boxes
In the same controlled study, participants who received uncertainty cues were faster overall (reduced annotation time).
Same controlled user study with 120 participants comparing interfaces with and without spatial-uncertainty visualizations; paper reports that participants with cues were faster overall.
In a controlled study with 120 participants, those receiving uncertainty cues achieve higher label quality.
Controlled user study reported in the paper; 120 participants; comparison between annotators who received visualized spatial-uncertainty cues via a purpose-built interface and those who did not; paper reports label quality outcomes.
The model identifies simple measures/conditions that characterize when productivity paradoxes and skill polarization arise.
Theoretical derivations and analytical characterizations within the model yielding threshold conditions and measures parameterizing when paradoxical outcomes occur (model-based; no empirical validation).
high positive Human-AI Productivity Paradoxes: Modeling the Interplay of S... predictive conditions/thresholds for productivity paradoxes and skill polarizati...
Sustainable progress requires collaborative integration of humans and machines, rather than replacement.
Normative conclusion/recommendation stated in the paper based on study findings (argument for augmented intelligence over replacement).
high positive Augmented Intelligence: Resolving the AI integration-obsoles... approach to AI-human integration
This research presents the innovative Marketing Intelligence Operations (MIO) Framework and a practical AI Adoption Readiness Scorecard, enabling leaders to manage the operational balance between transformative efficiency improvements and human capital vulnerability.
Paper states that it introduces a new framework and a practical scorecard as deliverables of the research (descriptive claim about the paper's contributions).
high positive Augmented Intelligence: Resolving the AI integration-obsoles... AI adoption readiness / operational management capability
AI-integrated Marketing Intelligence Operations (MIO) quantitatively improves campaign Return on Investment (ROI) by 47%.
Reported as an empirical result from the paper's mixed-methods study (the paper states use of audits, surveys, and NLP analysis to evaluate MIO outcomes).
high positive Augmented Intelligence: Resolving the AI integration-obsoles... campaign Return on Investment (ROI)
Deploying LegalCheck in the Municipality of Amsterdam demonstrated substantial efficiency gains, improved legal consistency, and positive user acceptance.
Summary claim based on the real-world deployment outcomes described in the paper (timing improvements, consistency/factual accuracy statements, and reported positive reception by professionals); specific quantitative metrics and sample sizes are not fully reported in the excerpt.
high positive LegalCheck: Retrieval- and Context-Augmented Generation for ... efficiency (time), legal consistency, user acceptance
The system produced explainable outputs based on actual regulations and prior cases, providing citations/explainability that support legal reasoning.
Paper describes retrieval from curated legal knowledge bases and generation of outputs grounded in regulations and prior cases during the Amsterdam deployment; presented as a feature of the system and supported by expert review.
high positive LegalCheck: Retrieval- and Context-Augmented Generation for ... explainability / traceability of generated legal reasoning to source regulations...
LegalCheck uses a combination of Retrieval-Augmented Generation (RAG) and Context-Augmented Generation (CAG) with curated legal knowledge bases and controlled prompting to retrieve relevant laws and precedents and incorporate case-specific details into coherent drafts.
System architecture and methodology described in the paper (design/implementation claim).
high positive LegalCheck: Retrieval- and Context-Augmented Generation for ... n/a (system design / method description)
Legal professionals found that the system ensured a consistent application of legal standards without replacing human judgment.
Reported qualitative feedback from professionals in the Municipality of Amsterdam deployment and the system design that includes an expert-in-the-loop review; no formal measurement of 'replacement' was reported.
high positive LegalCheck: Retrieval- and Context-Augmented Generation for ... consistency in application of legal standards and preservation of human oversigh...
Legal professionals found that the system reduced their workload.
Reported user feedback from legal professionals during the Municipality of Amsterdam deployment; qualitative statements that professionals experienced workload reduction (no numeric workload metrics or sample size reported).
high positive LegalCheck: Retrieval- and Context-Augmented Generation for ... perceived workload of legal professionals
The system's output captured the vast majority of required legal reasoning—often 80% to 100% of essential content.
Reported coverage statistic from the deployment/evaluation described in the paper (phrased as 'often 80% to 100% of essential content'); exact evaluation method, sample size, and measurement protocol are not provided in the excerpt.
high positive LegalCheck: Retrieval- and Context-Augmented Generation for ... proportion of essential legal reasoning/content captured in generated drafts
LegalCheck maintained high legal consistency and factual accuracy when generating draft letters.
Evaluation during real-world deployment with expert-in-the-loop review and feedback from legal professionals in the Municipality of Amsterdam; claims of high consistency and factual accuracy are reported but no formal numeric accuracy metric or sample size is provided in the text.
high positive LegalCheck: Retrieval- and Context-Augmented Generation for ... legal consistency and factual accuracy of generated letters
LegalCheck produced near-final advice letters in minutes rather than hours.
Reported results from a real-world deployment within the Municipality of Amsterdam; system logs / timing comparisons between human drafting time (hours) and LegalCheck-assisted drafting time (minutes) are described in the paper (no explicit numeric sample size reported).
high positive LegalCheck: Retrieval- and Context-Augmented Generation for ... time to produce advice/objection response letters
We outline a research program for the runtime systems that foundation-model software agents will require.
Paper claims to present a forward-looking research agenda or program (stated in abstract); this is a conceptual contribution rather than an empirical finding.
high positive AI Harness Engineering: A Runtime Substrate for Foundation-M... research directions needed for runtime systems for foundation-model software age...
Applied to a controlled validation task, the framework yields episode packages whose evidence structure varies systematically with harness level: lower levels produce only a final patch, while higher levels produce reproduction logs, failure attributions, deterministic requirement checks, and structured verification reports.
Empirical application described in the abstract: framework applied to a controlled validation task showing systematic variation in episode-package evidence structure across harness levels. The abstract does not report sample size or statistical measures.
high positive AI Harness Engineering: A Runtime Substrate for Foundation-M... evidence structure of episode packages produced (types of artifacts: final patch...
We propose a trace-based evaluation protocol that converts each agent run into an auditable episode package.
Methodological proposal described in the abstract proposing a trace-based protocol and an auditable episode package format; no quantitative evaluation details provided in the abstract.
high positive AI Harness Engineering: A Runtime Substrate for Foundation-M... auditability of agent runs (availability of trace-based episode packages)
We operationalize the harness through a four-level ladder (H0–H3) that progressively exposes runtime support to the agent.
Design contribution described in the paper (abstract) introducing a four-level ladder (H0–H3) as an operationalization of the harness concept.
high positive AI Harness Engineering: A Runtime Substrate for Foundation-M... degree of runtime support exposed to an agent across harness levels
Foundation models have transformed automated code generation.
Statement in paper's abstract referring to broad impact of foundation models on automated code generation; likely supported by citations and literature overview within the paper (no sample size or quantitative study reported in the abstract).
high positive AI Harness Engineering: A Runtime Substrate for Foundation-M... ability of foundation models to generate code (automation of coding tasks)
Authorship preservation should be a design priority for AI tools deployed in identity-relevant, behavior-dependent tasks.
Authors' recommendation based on experimental results showing negative motivational and behavioral consequences of delegating authorship to LLMs despite improved objective goal quality.
high positive Optimized but Unowned: How AI-Authored Goals Undermine the M... design recommendation (no empirical outcome measured)
Mediation analyses identified psychological ownership as the mechanism: it mediated the authorship effect on every downstream motivational and behavioral outcome, while objective goal quality did not.
Mediation analyses reported in the preregistered experiment (authors tested psychological ownership and objective goal quality as mediators of authorship effects on multiple downstream outcomes); preregistered N = 470.
high positive Optimized but Unowned: How AI-Authored Goals Undermine the M... mediating effect of psychological ownership on authorship => motivational and be...
At two-week follow-up, 72.8% of self-authored participants had acted on two or more of their goals, compared to 46.6% in the LLM condition.
Behavioral follow-up measure collected two weeks after the intervention in the preregistered experiment; percentages reported in the paper/abstract. (Follow-up completion N not specified in the abstract.)
high positive Optimized but Unowned: How AI-Authored Goals Undermine the M... proportion of participants who acted on two or more goals within two weeks (beha...
LLM-generated goals scored higher on SMART criteria (specificity, measurability, achievability, relevance, and time-boundedness).
Preregistered randomized experiment comparing self-authored vs LLM-authored goals derived from a personal reflection; reported effect size d = 2.26; total preregistered N = 470.
high positive Optimized but Unowned: How AI-Authored Goals Undermine the M... SMART criteria score (objective goal quality)
The Agent-First paradigm is orthogonal and complementary to transport-layer standards such as MCP, operating as the semantic application layer above existing tool discovery and invocation protocols.
Conceptual argument and mapping presented in the paper asserting interoperability/orthogonality with transport-layer standards (e.g., MCP).
high positive Agent-First Tool API: A Semantic Interface Paradigm for Ente... compatibility_with_transport_layer_standards
Agent-First APIs improve autonomous error recovery by 5.8x (compared to optimized CRUD baselines).
Reported comparative experiments on 50 real operational tasks measuring autonomous error recovery capability.
high positive Agent-First Tool API: A Semantic Interface Paradigm for Ente... autonomous_error_recovery
Agent-First APIs reduce required human interventions by 72.7% (compared to optimized CRUD baselines).
Same set of comparative experiments on 50 real operational tasks reported in the paper.
high positive Agent-First Tool API: A Semantic Interface Paradigm for Ente... required_human_interventions
Comparative experiments on 50 real operational tasks demonstrate that Agent-First APIs achieve 88% end-to-end task success rate versus 64% for optimized CRUD baselines (+37.5%).
Empirical comparative experiments reported in the paper on 50 real operational tasks, comparing Agent-First APIs to optimized CRUD baselines.
high positive Agent-First Tool API: A Semantic Interface Paradigm for Ente... end-to-end_task_success_rate
The paradigm is implemented and validated in a production multi-tenant SaaS platform serving 85 registered tools across 6 business domains.
Reported production implementation and deployment statistics (platform with 85 registered tools spanning 6 business domains).
high positive Agent-First Tool API: A Semantic Interface Paradigm for Ente... deployment_of_paradigm_on_production_SaaS_platform
We propose the Agent-First Tool API paradigm, comprising three integrated mechanisms: (1) a Six-Verb Semantic Protocol that decomposes tool interactions into search, resolve, preview, execute, verify, and recover phases; (2) a Normalized Tool Contract (NTC) providing structured decision-support metadata including confidence scores, evidence chains, and suggested next actions; and (3) a dual-layer governance pipeline combining static capability policies with dynamic risk escalation.
Design and specification presented in the paper (proposed architecture and components).
high positive Agent-First Tool API: A Semantic Interface Paradigm for Ente... proposed_API_paradigm_and_components
LLMs can help generate more correct and functional code compared to participant-generated solutions.
Comparative analysis of generated solutions reported in the paper (no sample-size for solutions explicitly stated in the abstract). The paper states LLM-assisted solutions were more correct/functional.
high positive "Like Taking the Path of Least Resistance": Exploring the Im... correctness and functionality of generated code
Qualitative analysis of participants' interactions and interviews revealed four different human-LLM collaboration modes supporting various problem-solving strategies.
Qualitative analysis of interaction logs and retrospective interviews from the study participants (N=20) reported in the paper; identification of four collaboration modes described.
high positive "Like Taking the Path of Least Resistance": Exploring the Im... types of collaboration modes
We conducted a within-subject study followed by retrospective interviews with programmers (N=20).
Stated methods in the paper: within-subject experimental design plus retrospective interviews; sample size explicitly given as N=20.
Organizations classified as 'Proactive Integrators' can reduce the risk of obsolescence by up to 53%.
Subgroup finding reported in the study (reduction estimate for organizations labeled 'Proactive Integrators'); specific subgroup sample not provided in abstract.
high positive The AI-engineering imperative - Navigating synergy and obsol... reduction in risk of skills obsolescence
AI-assisted engineering teams can achieve a 24% increase in productivity.
Empirical finding reported by the study, derived from the mixed-methods analysis (survey of 320 orgs, Delphi with 40 experts, and case studies of 5 industries as described in abstract).
high positive The AI-engineering imperative - Navigating synergy and obsol... increase in productivity of AI-assisted engineering teams
Entities that strategically implement AI can enhance their innovation cycles by up to 30%.
Statement in paper (presented as a forecast/estimate; no specific study or sample detailed in abstract).
high positive The AI-engineering imperative - Navigating synergy and obsol... improvement in innovation cycle speed/efficiency
AwareLLM opens new avenues for Human-AI collaboration where technology adapts to users' needs rather than users adhering to technological constraints.
Authorial/conceptual claim based on the proposed framework and study results; presented as a broader implication rather than a direct empirical finding.
high positive AwareLLM: A Proactive Multimodal Ecosystem for Personalized ... human-AI collaboration potential
Participants described AwareLLM's personalized interventions as timely and relevant, helping them boost their confidence and deepen engagement with their work.
Qualitative user feedback reported in the study (participant descriptions); sample size 20. No coding details or counts provided in the abstract.
high positive AwareLLM: A Proactive Multimodal Ecosystem for Personalized ... confidence and engagement (subjective reports)