The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests About 🎲 Workforce Futures
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Evidence (8974 claims)

Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.

The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).

Browse by theme

Nine broad, paper-level topics. Click one to filter the claims below.

Adoption
10085 claims
Filter claims →
Productivity
8974 claims
Filtered →
Governance
8062 claims
Filter claims →
Human-AI Collaboration
7749 claims
Filter claims →
Org Design
5057 claims
Filter claims →
Innovation
4896 claims
Filter claims →
Labor Markets
4088 claims
Filter claims →
Skills & Training
3372 claims
Filter claims →
Inequality
2377 claims
Filter claims →

Claims by outcome category

Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.

Outcome Positive Negative Mixed Null Total
Other 882 244 117 1097 2424
Governance & Regulation 1010 469 229 135 1875
Organizational Efficiency 977 235 149 90 1462
Technology Adoption Rate 781 299 143 128 1362
Research Productivity 506 155 74 363 1110
Output Quality 555 219 71 70 915
Decision Quality 395 200 95 54 751
Firm Productivity 523 67 101 27 724
AI Safety & Ethics 262 309 75 36 688
Market Structure 195 201 135 30 566
Task Allocation 248 77 96 38 464
Innovation Output 300 34 55 20 411
Skill Acquisition 207 75 65 21 368
Employment Level 138 67 119 24 350
Fiscal & Macroeconomic 156 80 53 33 329
Task Completion Time 211 38 13 16 280
Firm Revenue 183 52 29 5 270
Consumer Welfare 131 77 48 13 269
Inequality Measures 50 141 54 9 254
Worker Satisfaction 104 85 25 13 227
Error Rate 87 112 11 5 215
Automation Exposure 69 69 37 20 198
Wages & Compensation 102 49 31 11 193
Team Performance 115 30 30 11 187
Regulatory Compliance 88 74 17 7 186
Training Effectiveness 109 22 14 21 168
Developer Productivity 116 21 15 8 161
Job Displacement 12 92 26 1 131
Hiring & Recruitment 57 12 9 5 83
Skill Obsolescence 6 59 10 2 77
Social Protection 43 17 8 2 70
Creative Output 35 21 9 4 70
Labor Share of Income 18 23 17 1 59
Worker Turnover 15 16 4 35
Industry 1 1
Clear
Productivity Remove filter
PRISM automatically generates test cases from plain-language agent requirements.
Methodological description in abstract stating PRISM takes plain-language requirements and automatically generates test cases.
high positive PRISM: Prompt Reliability via Iterative Simulation and Monit... test-case generation from requirements
PRISM successfully identifies and repairs production regressions caused by LLM behavioral drift within a 24-hour detection window.
Reported result in abstract claiming detection and repair of production regressions within 24 hours during deployment.
high positive PRISM: Prompt Reliability via Iterative Simulation and Monit... time-to-detection/repair of production regressions
PRISM achieves 99% production reliability across all evaluated agents.
Reported quantitative outcome in abstract from the three-week evaluation over 35 agents.
PRISM reduces median prompt authoring time from 2 days to under 30 minutes.
Reported quantitative outcome in abstract from the three-week evaluation across the 35 agents.
high positive PRISM: Prompt Reliability via Iterative Simulation and Monit... median prompt authoring time
Policymakers should combine support for technological development with strategic investments in finance, trade integration, and public infrastructure to maximize AI's economic benefits and transform its potential into sustainable and inclusive growth.
Policy recommendation derived from the empirical findings (positive AI effects and positive interactions with financial innovation, trade openness, and government consumption) reported for 19 G20 countries (2005–2023) using GMM.
high positive Artificial intelligence and economic growth in G20 economies... economic growth (implied)
The interaction between AI and government final consumption expenditure helps strengthen economic growth by improving public infrastructure, institutional quality, and capacity to leverage new technologies.
GMM interaction specifications using panel data for 19 G20 countries (2005–2023); reported AI × government final consumption expenditure interaction coefficient is positive and statistically significant, with interpretation linking it to public infrastructure and institutional capacity.
The interaction between AI and trade openness is positive and significant, underscoring the role of international trade in technological diffusion and competitiveness to boost growth.
GMM interaction models on panel data (19 G20 countries, 2005–2023); reported AI × trade openness interaction coefficient is positive and statistically significant.
The interaction between AI and financial innovation has a positive and significant impact on economic growth, indicating that innovative finance mediates AI's technological potential into tangible economic gains.
GMM models with interaction terms using panel data of 19 G20 countries (2005–2023); reported AI × financial innovation interaction coefficient is positive and statistically significant.
AI-related innovation has a positive and significant effect on economic growth (linear model, GMM).
Panel analysis of 19 G20 countries (2005–2023) using the Generalized Method of Moments (GMM) linear model; reported positive and statistically significant coefficient for AI-related innovation.
Synthetic scenarios in the paper illustrate that the revised metric distinguishes between frequent low-leverage use, semantically repetitive prompting, and more autonomous, higher-consequence AI-assisted work.
Paper includes synthetic scenario simulations/illustrations demonstrating metric behavior across different usage patterns (synthetic examples; no real-world sample reported).
high positive Intelligence Impact Quotient (IIQ): A Framework for Measurin... ability of the metric to discriminate types of AI use
The authors derive sub-daily update rules and a bounded interpretation layer for estimated efficiency and financial impact from the IIQ metric.
Analytic derivation in the methods: paper presents update rules (sub-daily) and an interpretation layer mapping IIQ to estimated efficiency and financial impact (theoretical derivation / worked examples). No empirical validation sample reported.
high positive Intelligence Impact Quotient (IIQ): A Framework for Measurin... estimated efficiency and financial impact
The formulation produces a raw Intelligence Adoption Index (IAI) and a normalized 0-1000 IIQ index for comparison between heterogeneous users and units.
Methodological description: authors define a raw IAI and describe normalization to a 0–1000 IIQ scale for comparability (model/specification). No empirical sample reported.
high positive Intelligence Impact Quotient (IIQ): A Framework for Measurin... normalized adoption/index score across users/units
IIQ combines a novelty-weighted, time-decayed token stock with usage frequency, a grace-period recency gate, organizational leverage, task complexity, and autonomy to form its measurement.
Methodological formulation in the paper: component-level specification of the IIQ metric (theoretical specification / algorithmic description). No empirical validation sample reported.
high positive Intelligence Impact Quotient (IIQ): A Framework for Measurin... operationalization of AI usage (components driving the metric)
The Intelligence Impact Quotient (IIQ) is a composite metric intended to quantify the depth to which AI systems are integrated into organizational work and their impact.
Paper framing and definition: the authors introduce IIQ as a composite metric and describe its purpose as quantifying AI integration depth and impact (conceptual/methodological description). No empirical sample reported.
high positive Intelligence Impact Quotient (IIQ): A Framework for Measurin... depth of AI integration into work / AI impact
Das Dokument leistet einen Beitrag zu den laufenden Bemühungen der G7 und der OECD, die Verbreitung innovativer, vertrauenswürdiger und produktivitätssteigernder KI im Einklang mit den KI-Grundsätzen der OECD zu fördern.
Descriptive claim about the paper's intended contribution to G7/OECD efforts and alignment with OECD AI Principles (self-declared by the paper).
high positive Einführung von KI in kleinen und mittleren Unternehmen Policy-aligned Förderung vertrauenswürdiger, produktivitätssteigernder KI
Die Erkenntnisse unterstreichen, dass die Regierungen Strategien unterstützen sollten, die die Einführung von KI in KMU beschleunigen und eine digitale Transformation fördern, die allen zugutekommt.
Policy recommendation based on the paper's synthesis of data analysis and case studies; presented as the paper's conclusion (no causal estimate provided in excerpt).
high positive Einführung von KI in kleinen und mittleren Unternehmen Wirksamkeit staatlicher Strategien zur Beschleunigung der KI-Einführung in KMU /...
Das Dokument führt eine Taxonomie der KI-nutzenden KMU auf Basis des digitalen Reifegrads, der Komplexität der Nutzung und des Umfangs der Anwendung ein, die darauf abzielt, die Politikgestaltung zu unterstützen.
Descriptive statement that the paper develops a taxonomy (method: taxonomy construction based on those three dimensions); presented as part of the paper's contributions (no empirical validation details in excerpt).
high positive Einführung von KI in kleinen und mittleren Unternehmen Kategorisierung/Taxonomie von KI-nutzenden KMU
Dieses Diskussionspapier wurde auf Ersuchen der G7-Präsidentschaft vom OECD-Sekretariat erstellt, um Hintergrundmaterial für die Diskussionen der G7 über einen Blueprint zur Einführung von KI in KMU bereitzustellen.
Descriptive statement of the paper's provenance and purpose (administrative/factual about document preparation).
high positive Einführung von KI in kleinen und mittleren Unternehmen Erstellung und Zweck des Diskussionspapiers
Im Rahmen der G7-Präsidentschaft Kanadas 2025 wurde die beschleunigte Einführung von KI in KMU zu einer Hauptpriorität erklärt.
Factual statement about G7 policy priorities as reported by the paper (administrative/policy fact reported by OECD secretariat).
high positive Einführung von KI in kleinen und mittleren Unternehmen political/policy prioritization of AI adoption in KMU
Künstliche Intelligenz (KI) ist ein vielversprechender Ansatz, um Produktivität und Innovation in Unternehmen, insbesondere kleinen und mittleren Unternehmen (KMU), zu steigern.
Authoritative statement in the paper's summary; based on literature review and general argumentation rather than a specific empirical test reported in this excerpt.
high positive Einführung von KI in kleinen und mittleren Unternehmen Produktivität und Innovation in Unternehmen (insbesondere KMU)
The research contributes to the literature on technology adoption in developing economies and offers policymakers and business leaders in sub-Saharan Africa valuable insights.
Paper's stated contribution in the abstract; a general claim about the study's scholarly and policy relevance rather than a quantifiable empirical result.
high positive Estimation of Firm Labour Productivity and Sales Growth from... contribution to literature and policy relevance
Targeted policy interventions — such as upskilling initiatives and supportive regulatory frameworks — are important to harness AI’s benefits while mitigating adverse impacts on workers.
Paper conclusion/recommendation drawn from empirical findings (positive association of AI with productivity and sales, plus observed cross-country variation). This is presented as a policy implication; no empirical evaluation of specific policies is reported in the excerpt.
high positive Estimation of Firm Labour Productivity and Sales Growth from... policy effectiveness in harnessing AI benefits and mitigating worker impacts (re...
AI adoption has a significant positive relationship with firm sales growth in the selected sub-Saharan African countries.
Same firm-level World Bank Enterprise Surveys (2007–2024) and regression methods (FGLS, robust OLS, HDFE) as above. Paper statement: "AI has a significant positive relationship with ... sales growth." Exact sample size and numeric effect not provided in excerpt.
AI adoption has a significant positive relationship with firm labour productivity in the selected sub-Saharan African countries.
Firm-level dataset from the World Bank Enterprise Surveys covering 2007–2024; empirical analysis using feasible generalized least squares (FGLS), robust OLS, and high-dimensional fixed effects (HDFE) linear regressions. Paper statement: "AI has a significant positive relationship with firm labour productivity." Exact firm sample size not reported in the provided excerpt.
There is a positive spillover effect on AI-ineligible chats: treated workers adapted their multitasking workflow to devote greater attention to these chats.
Experiment-level observations comparing worker behavior on AI-ineligible chats between treatment and control; treated workers reallocated attention/effort (multitasking workflow changes) leading to improved attention on AI-ineligible chats.
high positive Agentic AI and Human-in-the-Loop Interventions: Field Experi... attention/effort devoted to AI-ineligible chats (spillover effect)
Early intervention is essential for sustaining high post-escalation intervention effort.
Temporal analysis of intervention timing within the randomized experiment showing an association between earlier human intervention after escalation and higher subsequent intervention effort.
high positive Agentic AI and Human-in-the-Loop Interventions: Field Experi... post-escalation intervention effort as a function of intervention timing
Human intervention preserves service quality in algorithm-triggered technical escalations (unresolved customer issues beyond the AI's capability).
Experimental subgroup analysis of escalations categorized as algorithm-triggered technical escalations; post-escalation human interventions were observed to maintain service quality in these cases.
high positive Agentic AI and Human-in-the-Loop Interventions: Field Experi... service quality after technical escalations
PRIF yielded an average ROI of 83%.
Reported financial evaluation/ROI estimate following PRIF adoption in the paper (derived from pilot/case study cost-benefit or sample analysis).
high positive Enhancing Forensic Accounting Practice: A Proactive Risk Man... return on investment (ROI %)
PRIF adoption reduced financial misstatements by 47%.
Reported change in financial misstatement incidence after PRIF implementation in the paper's evaluation (case studies/forensic report analysis).
high positive Enhancing Forensic Accounting Practice: A Proactive Risk Man... financial misstatements (reduction %)
PRIF adoption reduced compliance resolution time by 58%.
Reported performance metric after PRIF adoption in pilot/case studies described in the paper.
high positive Enhancing Forensic Accounting Practice: A Proactive Risk Man... compliance resolution time (reduction %)
Client retention was 91% for high SCI versus 54% for low SCI.
Reported retention rates stratified by SCI levels in paper (presumably derived from the sample used for SCI analysis).
The Stakeholder Communication Index (SCI) revealed a strong correlation (r = 0.83) between report quality and client retention.
Statistical analysis reported in paper linking SCI-derived report quality scores to client retention; correlation coefficient r = 0.83 provided.
high positive Enhancing Forensic Accounting Practice: A Proactive Risk Man... correlation between report quality and client retention (r)
Accuracy increased from 62% to 89–94% after integration of AI and blockchain.
Reported accuracy figures in results section based on PRIF evaluation (presumably from analyzed forensic reports/case studies).
high positive Enhancing Forensic Accounting Practice: A Proactive Risk Man... detection/analysis accuracy (%)
Integration of AI and blockchain reduced the risk detection time from 47 days post-event to 9–22 days pre-event.
Reported results from PRIF implementation/pilot using case studies and forensic report analysis (paper cites these temporal comparisons).
high positive Enhancing Forensic Accounting Practice: A Proactive Risk Man... risk detection time (days)
This study pioneers a Proactive Risk Intelligence Framework (PRIF) for Chartered Accountant (CA) firms, targeting gaps in risk anticipation, stakeholder communication, and compliance.
Paper description of study objective and framework development (mixed-method design, interviews, case studies, forensic report analysis).
high positive Enhancing Forensic Accounting Practice: A Proactive Risk Man... creation and introduction of PRIF (framework development)
We distill practical design principles for selecting logging policies when operational constraints prevent implementing the theoretical optimum.
Abstract states the paper provides practical design principles derived from the theoretical work; basis is methodological/theoretical synthesis (no empirical sample size provided in abstract).
high positive Logging Policy Design for Off-Policy Evaluation practical guidance / design principles for logging policy selection under constr...
We demonstrate the importance of treatment selection when gathering data for OPE, and describe theoretically optimal approaches when this is a firm's primary objective.
Abstract claims demonstration and description of theoretically optimal approaches; evidence likely consists of analytical results and/or illustrative demonstrations in the paper (no sample size reported in abstract).
high positive Logging Policy Design for Off-Policy Evaluation impact of treatment (logging) selection on OPE performance and derivation of opt...
We propose a unifying framework for logging policy design and derive optimal policies in canonical informational regimes where the target policy and reward distribution are (i) known, (ii) unknown, and (iii) partially known through priors or noisy estimates at logging time.
Abstract states the development of a framework and derivations of optimal policies across specified informational regimes; evidence is theoretical derivations (no empirical details in abstract).
high positive Logging Policy Design for Off-Policy Evaluation optimal logging policies for minimizing OPE error under different informational ...
In practice OPE accuracy depends heavily on the logging policy used to collect data for computing the estimate.
Assertion in abstract motivated by the authors' study of logging policy design; implies analytical results in the paper relating logging policy to OPE error (no sample size given in abstract).
high positive Logging Policy Design for Off-Policy Evaluation OPE accuracy / OPE error
Off-policy evaluation (OPE) estimates the value of a target treatment policy (e.g., a recommender system) using data collected by a different logging policy, enabling high-stakes experimentation without live deployment.
Statement in abstract describing OPE and its role; conceptual/theoretical description (no sample size or empirical study reported in the abstract).
high positive Logging Policy Design for Off-Policy Evaluation ability to estimate target policy value without live deployment (OPE capability)
These verified assertions improve users' performance on code-comprehension tasks in a user study with more than 400 participants.
User study reported in the paper: a study involving more than 400 participants measured performance on code-comprehension tasks with and without the verified assertions (sample size reported as >400 participants).
high positive Viverra: Text-to-Code with Guarantees users' performance on code-comprehension tasks
Evaluation on 18 diverse programming tasks suggests that Viverra can efficiently generate code with verified assertions.
Empirical evaluation reported in the paper: a test set of 18 programming tasks was used to evaluate Viverra's ability to generate code with verified assertions (sample size = 18 tasks).
high positive Viverra: Text-to-Code with Guarantees rate/ability to generate code with verified assertions
Viverra verifies those assertions in a compositional and best-effort manner via a portfolio of bounded model checkers.
Method description: the paper states that verification is done compositionally and in a best-effort way using a portfolio of bounded model checkers (implementation/algorithmic claim).
high positive Viverra: Text-to-Code with Guarantees verification of assertions using bounded model checkers
Given a natural-language task description, Viverra prompts an LLM to synthesize a C program together with candidate assertions expressing safety and correctness properties.
Method section description: the workflow described in the paper explicitly states LLM prompting to produce C programs and candidate assertions (methodological claim, illustrated with examples).
high positive Viverra: Text-to-Code with Guarantees generation of C program plus candidate assertions
Viverra automatically produces formally verified annotations alongside generated code to aid users' understanding of the generated program.
System description in the paper: Viverra is presented as a system that generates code together with formally verified annotations; implementation details and demonstration are described (no precise external benchmark cited here).
high positive Viverra: Text-to-Code with Guarantees availability of formally verified annotations alongside generated code
Adversarial inputs evolved using a small proxy model retain high effectiveness against large commercial LRMs (strong transferability).
Reported transfer experiments in the abstract showing that evolved adversarial inputs from a small proxy model remain effective against larger commercial models; no numeric transfer success rates provided in abstract.
high positive Inducing Overthink: Hierarchical Genetic Algorithm-based DoS... transfer effectiveness of adversarial inputs (ability to induce overthinking / i...
Across four state-of-the-art reasoning models, the proposed method substantially amplifies output length, achieving up to a 26.1x increase on the MATH benchmark and consistently outperforming benign and manually crafted missing-premise baselines.
Experimental results reported in the abstract: evaluations on the MATH benchmark and comparisons against benign and manually crafted missing-premise baselines across four SOTA models.
high positive Inducing Overthink: Hierarchical Genetic Algorithm-based DoS... output length (response length) on MATH benchmark
We propose an automated black-box framework that induces overthinking in LRMs by systematically perturbing the logical structure of input problems using a hierarchical genetic algorithm (HGA) operating on structured problem decompositions and optimizing a composite fitness function to maximize response length and reflective overthinking markers.
Methodological description of the proposed approach (HGA and composite fitness) as presented by the authors in the abstract.
high positive Inducing Overthink: Hierarchical Genetic Algorithm-based DoS... ability to induce overthinking / increase in response length (method capability)
Function signatures, constraints and style descriptions emerge as the most influential prompt dimensions affecting the readability of LLM-generated code.
Systematic examination of multiple prompt dimensions in the paper, reporting that function signatures, constraints, and style descriptions had the largest measured influence on readability scores.
high positive The Readability Spectrum: Patterns, Issues, and Prompt Effec... impact_of_prompt_dimensions_on_readability
We evaluate the readability of code generated by mainstream LLMs under 5,869 scenarios extracted from large code bases including World of Code (WoC) and LeetCode.
Empirical evaluation reported in paper using 5,869 scenarios drawn from WoC and LeetCode; LLM-generated code samples were produced and scored with the readability model.
high positive The Readability Spectrum: Patterns, Issues, and Prompt Effec... coverage of evaluation / dataset size for readability assessment