The Commonplace
Home Dashboard Papers Evidence Digests 🎲

Evidence (7448 claims)

Adoption
5267 claims
Productivity
4560 claims
Governance
4137 claims
Human-AI Collaboration
3103 claims
Labor Markets
2506 claims
Innovation
2354 claims
Org Design
2340 claims
Skills & Training
1945 claims
Inequality
1322 claims

Evidence Matrix

Claim counts by outcome category and direction of finding.

Outcome Positive Negative Mixed Null Total
Other 378 106 59 455 1007
Governance & Regulation 379 176 116 58 739
Research Productivity 240 96 34 294 668
Organizational Efficiency 370 82 63 35 553
Technology Adoption Rate 296 118 66 29 513
Firm Productivity 277 34 68 10 394
AI Safety & Ethics 117 177 44 24 364
Output Quality 244 61 23 26 354
Market Structure 107 123 85 14 334
Decision Quality 168 74 37 19 301
Fiscal & Macroeconomic 75 52 32 21 187
Employment Level 70 32 74 8 186
Skill Acquisition 89 32 39 9 169
Firm Revenue 96 34 22 152
Innovation Output 106 12 21 11 151
Consumer Welfare 70 30 37 7 144
Regulatory Compliance 52 61 13 3 129
Inequality Measures 24 68 31 4 127
Task Allocation 75 11 29 6 121
Training Effectiveness 55 12 12 16 96
Error Rate 42 48 6 96
Worker Satisfaction 45 32 11 6 94
Task Completion Time 78 5 4 2 89
Wages & Compensation 46 13 19 5 83
Team Performance 44 9 15 7 76
Hiring & Recruitment 39 4 6 3 52
Automation Exposure 18 17 9 5 50
Job Displacement 5 31 12 48
Social Protection 21 10 6 2 39
Developer Productivity 29 3 3 1 36
Worker Turnover 10 12 3 25
Skill Obsolescence 3 19 2 24
Creative Output 15 5 3 1 24
Labor Share of Income 10 4 9 23
Four primary application areas were identified: (1) behavioural monitoring and feedback, (2) predictive risk modelling, (3) decision support and AI classifiers, and (4) limit‑setting and self‑exclusion tools.
Thematic synthesis of included studies categorizing described applications into four main areas (review taxonomy).
high null result Deep technologies and safer gambling: A systematic review. application area classification (categorical counts / thematic presence)
Searches were performed in Web of Science, PubMed, Scopus, EBSCO and IEEE, plus manual searches, following PRISMA guidelines.
Methods section of the review specifying databases searched and PRISMA-guided review process.
high null result Deep technologies and safer gambling: A systematic review. search strategy / databases searched (qualitative)
The review included 68 empirical and methodological studies on deep technologies in online gambling.
Systematic review following PRISMA; searches of Web of Science, PubMed, Scopus, EBSCO, IEEE and manual searching produced 68 included studies (count reported in paper).
high null result Deep technologies and safer gambling: A systematic review. number of included studies (study count = 68)
The collection includes a mix of methodological papers, empirical applications demonstrating ecological insight, and translational work focused on policy or conservation practice.
Study-types categorization provided in the paper (descriptive tally/characterization of the kinds of contributions in the collection).
high null result Towards ‘digital ecology’: Advances in integrating artificia... types of studies present in the collection
Methods in the collection span from automated image and signal processing for routine tasks to integrated modelling that couples ecological theory with data‑driven methods.
Methods-scope summary in the paper describing the range of AI/ML approaches used across the collection (descriptive across studies).
high null result Towards ‘digital ecology’: Advances in integrating artificia... range of methodological approaches used
The collection uses large ecological observational datasets such as camera‑trap imagery, sensor streams, biodiversity surveys, and other high‑volume ecological monitoring data.
Data & methods section listing the data types represented across the reviewed papers (descriptive inventory of dataset types used in the collection).
high null result Towards ‘digital ecology’: Advances in integrating artificia... types of data used in ecological AI research
Recommendation (research): Future research should link AI adoption to objective performance metrics (profitability, default rates, processing times) and use longitudinal or quasi-experimental designs to identify causal effects.
Authors' suggested research directions noted in the summary, motivated by limitations of cross-sectional, self-reported data.
high null result From Data to Decisions: Harnessing Artificial Intelligence f... research design and outcome measurement (recommendation)
The summary omits important reporting details: p-values, standard errors, model control variables, and exact variable operationalizations are not provided.
Explicit reporting gap noted in the paper summary (absence of p-values, SEs, controls, and operationalization details).
high null result From Data to Decisions: Harnessing Artificial Intelligence f... statistical reporting completeness
Because the data are cross-sectional and self-reported, the design limits causal inference about AI adoption causing the observed outcomes.
Study design (cross-sectional survey, self-reported measures) and explicit limitation noted in the paper summary.
high null result From Data to Decisions: Harnessing Artificial Intelligence f... ability to infer causality
Key measures are self-reported Likert scales for AI adoption/usage and the dependent outcomes (financial decision-making efficiency, operational efficiency, financial resilience, and AI-based analytics effectiveness).
Measurement description in Methods: independent and dependent variables reported as self-reported Likert measures collected in the cross-sectional survey.
high null result From Data to Decisions: Harnessing Artificial Intelligence f... measurement type (self-reported Likert scales)
The study is a cross-sectional quantitative survey of 312 professionals in banks, fintechs, and financial service firms.
Study design and sample description reported in Data & Methods; sample size explicitly given as N = 312 and composition described as professionals across financial institutions, fintech organizations, and financial service companies.
The SKILL.md used in the with-skill condition encodes workflow logic, API patterns, and business rules as portable domain guidance for agents.
Paper description of the with-skill intervention specifying the content and intended role of SKILL.md.
high null result SKILLS: Structured Knowledge Injection for LLM-Driven Teleco... presence and content type of injected domain guidance (workflow logic, API patte...
We evaluated open-weight models under two conditions: baseline (generic agent with tool access but no domain guidance) and with-skill (agent augmented with a portable SKILL.md document encoding workflow logic, API patterns, and business rules).
Experimental design in paper describing the two agent conditions; SKILL.md described as the injected domain guidance artifact.
high null result SKILLS: Structured Knowledge Injection for LLM-Driven Teleco... experimental condition (baseline vs with-skill)
Each scenario is grounded in live mock API servers with seeded production-representative data, MCP tool interfaces, and deterministic evaluation rubrics combining response content checks, tool-call verification, and database state assertions.
Methods/benchmark design described in paper specifying environment: live mock APIs, seeded data, MCP tool interfaces, and deterministic evaluation combining content checks, tool-call verification, and DB assertions.
high null result SKILLS: Structured Knowledge Injection for LLM-Driven Teleco... evaluation environment fidelity and evaluation criteria (content checks, tool-ca...
SKILLS comprises 37 telecom operations scenarios spanning 8 TM Forum Open API domains (TMF620, TMF621, TMF622, TMF628, TMF629, TMF637, TMF639, TMF724).
Framework specification in the paper; explicit statement of scenario count (37) and list of 8 TMF Open API domains.
high null result SKILLS: Structured Knowledge Injection for LLM-Driven Teleco... coverage: number of scenarios (37) and number of API domains (8) included
We introduce SKILLS (Structured Knowledge Injection for LLM-driven Service Lifecycle operations), a benchmark framework for telecom operations.
Paper describes the design and release of the SKILLS benchmark framework as the contribution; methods section outlines framework components and usage.
high null result SKILLS: Structured Knowledge Injection for LLM-Driven Teleco... existence and definition of the SKILLS benchmark framework
The paper identifies three core mechanisms underlying calibrated trust and complementarity: (1) calibrated trust balancing reliance and oversight, (2) complementarity–trust interaction for optimal performance, and (3) dynamic feedback loops producing reinforcing learning cycles.
Explicit identification of mechanisms claimed in the paper's synthesis; this is a descriptive claim about the paper's content rather than an empirical finding—no sample or empirical test reported in the abstract.
high null result Optimising Human– AI Decision Performance: A Trust and Cap... n/a (identification of theoretical mechanisms)
AI-adopting firms do not increase capital expenditures following adoption.
Firm-level capex analysis showing no significant change in capital expenditures for adopters versus nonadopters post-adoption in the paper's empirical framework.
high null result AI and Productivity: The Role of Innovation capital expenditures (capex)
It remains unclear how developers' general programming and security-specific experience, and the type of AI tool used (free vs. paid), affect the security of the resulting software — motivating this study.
Paper's stated research gap/motivation: the authors identify uncertainty in the literature regarding interactions between developer experience, AI tool tier (free vs. paid), and resulting code security.
high null result The Impact of AI-Assisted Development on Software Security: ... the combined effect of developer experience and AI tool type on code security (i...
Participants were assigned a security-related programming task using either no AI tools, the free version, or the paid version of Gemini.
Experimental design described in the paper: random/conditional assignment of participants into three groups (no AI, free Gemini, paid Gemini) performing the same security-related programming task.
high null result The Impact of AI-Assisted Development on Software Security: ... experimental condition (tool used) as it relates to subsequent code security out...
We conducted a quantitative programming study with software developers (n = 159) exploring the impact of Google's AI tool Gemini on code security.
Explicit methodological statement in the paper: a quantitative study with 159 participating software developers assigned to experimental conditions to evaluate Gemini's impact on security-related programming tasks.
high null result The Impact of AI-Assisted Development on Software Security: ... impact of Gemini on code security (security of code produced in the study)
The authors surveyed workers and developers on a representative sample of 171 tasks and used language models (LMs) to scale ratings to 10,131 computer-assisted tasks across all U.S. occupations.
Study methodology reported in the paper: surveys of 'workers and developers' on 171 tasks, plus LM-based scaling to 10,131 tasks (coverage claims across U.S. occupations).
high null result Are We Automating the Joy Out of Work? Designing AI to Augme... coverage and scaling of task-level ratings (number of tasks surveyed and number ...
SWE-Skills-Bench is available at https://github.com/GeniusHTX/SWE-Skills-Bench.
Repository URL provided in the paper for the benchmark's code/data.
high null result SWE-Skills-Bench: Do Agent Skills Actually Help in Real-Worl... public availability (URL) of the benchmark
SWE-Skills-Bench provides a testbed for evaluating the design, selection, and deployment of skills in software engineering agents.
Benchmark design pairs skills, repositories, and deterministic verification tests; intended use stated by authors as a testbed for evaluation of skills.
high null result SWE-Skills-Bench: Do Agent Skills Actually Help in Real-Worl... availability of a benchmarking testbed for evaluating agent skills
39 of 49 skills yield zero pass-rate improvement.
Empirical evaluation over 49 skills and ~565 task instances reporting that 39 skills produced no improvement in test pass rate when injected.
high null result SWE-Skills-Bench: Do Agent Skills Actually Help in Real-Worl... change in task acceptance-test pass rate (zero improvement)
The authors introduce a deterministic verification framework that maps each task's acceptance criteria to execution-based tests, enabling controlled paired evaluation with and without the skill.
Method: creation of a deterministic verification framework that converts acceptance criteria into executable tests; used to perform paired evaluations (with skill vs. without skill).
high null result SWE-Skills-Bench: Do Agent Skills Actually Help in Real-Worl... ability to deterministically verify task acceptance criteria via execution-based...
SWE-Skills-Bench pairs 49 public SWE skills with authentic GitHub repositories pinned at fixed commits and requirement documents with explicit acceptance criteria, yielding approximately 565 task instances across six SWE subdomains.
Benchmark construction: 49 public skills, repositories pinned to fixed commits, requirement documents with acceptance criteria, producing ~565 task instances spanning six SWE subdomains (as reported by the paper).
high null result SWE-Skills-Bench: Do Agent Skills Actually Help in Real-Worl... number of skill-repo-task instances (~565) and coverage across six subdomains
The article introduces a novel Bayesian Item Response Theory framework that quantifies human–AI synergy by separately estimating individual ability, collaborative ability, and AI model capability while controlling for task difficulty.
Methodological contribution described in the paper: development and application of a Bayesian Item Response Theory model that includes separate parameters for individual ability, collaborative ability, AI model capability, and task difficulty (method section of the paper).
high null result Quantifying and Optimizing Human-AI Synergy: Evidence-Based ... estimated parameters for individual ability, collaborative ability, AI model cap...
The Planner is trained via Supervised Fine-Tuning (SFT) to internalize diagnostic capabilities and then aligned with business outcomes (conversion rate) via Reinforcement Learning (RL).
Method description in the paper specifying SFT initialization followed by RL alignment targeting conversion rate (UCVR) as reward signal.
high null result Probe-then-Plan: Environment-Aware Planning for Industrial E... Planner diagnostic behavior and policy alignment with conversion rate (model tra...
EASP's Offline Data Synthesis stage: a Teacher Agent synthesizes diverse, execution-validated plans by diagnosing the probed environment.
Method description in the paper detailing the Teacher Agent's role in synthesizing execution-validated plans during offline data synthesis.
high null result Probe-then-Plan: Environment-Aware Planning for Industrial E... synthesized execution-validated search plans (data generation outcome)
The Probe-then-Plan mechanism uses a lightweight Retrieval Probe to expose the retrieval snapshot, enabling the Planner to diagnose execution gaps and generate grounded search plans.
Methodological description in the paper: design and implementation of Retrieval Probe and Planner; validated through synthesized data and downstream evaluations (offline and online).
high null result Probe-then-Plan: Environment-Aware Planning for Industrial E... retrieval snapshot exposure and Planner diagnostic output (implementation/functi...
Descriptive statistics, reliability tests, regression analysis, and structural equation modelling (SEM) were employed to analyse the relationships between AI adoption and entrepreneurial outcomes.
Methods section reporting use of descriptive statistics, reliability tests, regression analysis, and SEM to evaluate relationships between AI adoption and measured outcomes.
high null result Entrepreneurship in the Era of Artificial Intelligence: Rede... not applicable (methodological detail)
The study used a quantitative research design and collected data from 350 entrepreneurs and managers of small and medium-sized enterprises (SMEs) who had adopted AI in their business operations.
Methods section of the paper specifying a quantitative design and a sample size of 350 AI-adopting SME entrepreneurs/managers.
high null result Entrepreneurship in the Era of Artificial Intelligence: Rede... not applicable (methodological detail)
The study used portfolio-level analysis to compare the financial outcomes of portfolios constructed using AI-driven ESG indicators with those based on conventional ESG ratings.
Methodological statement in the paper: portfolio-level analysis and comparative design. The summary does not specify the number of portfolios, asset universes, time frame, or construction rules.
high null result Green Intelligence in Finance: Artificial Intelligence-Drive... Study methodology (portfolio-level comparative analysis)
A quantitative methodology was employed, utilizing a structured questionnaire administered to 400 small business owners.
Explicit methodological statement in the paper: structured questionnaire survey with sample size N=400 small business owners.
high null result The role of artificial intelligence in enhancing financial l... method / sample (use of structured questionnaire; sample size = 400)
The study uses a game-theoretic model involving a foundation model provider and two competing downstream firms to analyze how policy interventions affect consumer surplus in the AI supply chain.
Methodological description in the paper: a formal game-theoretic model with one upstream provider and two downstream competing firms; equilibrium analysis and comparative statics are performed on model outcomes (prices, qualities, profits, consumer surplus).
high null result The Economics of AI Supply Chain Regulation model equilibrium outcomes (prices, qualities, provider profit, downstream profi...
Foi realizada etnografia organizacional orientada ao SCF, com roteiro e triangulação de evidências.
Método qualitativo divulgado no resumo: etnografia organizacional com roteiro e triangulação; o resumo não fornece número de organizações, duração ou amostragem.
high null result A FRICÇÃO PSICOANTROPOLÓGICA (SCF - Symbolic-Cognitive Frict... evidências qualitativas da existência e manifestação da fricção psicoantropológi...
Foi construído e validado um instrumento psicométrico (escala SCF-30) e calculado um índice 0–100, com modelagem por Equações Estruturais (SEM) e testes de confiabilidade/validade.
Descrição metodológica explícita no resumo: construção e validação da escala SCF-30, uso de SEM e testes de confiabilidade e validade. O resumo não detalha estatísticas, amostra ou resultados numéricos.
high null result A FRICÇÃO PSICOANTROPOLÓGICA (SCF - Symbolic-Cognitive Frict... pontuação SCF (índice 0–100) e propriedades psicométricas da escala SCF-30 (conf...
O SCF é operacionalizado por três vetores centrais: Percepção de Complexidade (PC), Aversão ao Risco Institucional (AR) e Inércia Cultural (IC).
Estrutura conceitual e operacional apresentada no artigo; especificação explícita dos três vetores como componentes do construto SCF.
high null result A FRICÇÃO PSICOANTROPOLÓGICA (SCF - Symbolic-Cognitive Frict... componentes constituintes do construto SCF (PC, AR, IC)
This research conducts a critical analysis of the ethical implications of artificial intelligence in terms of job displacement during the fifth industrial revolution.
Author-declared methodology: a literature-based critical analysis drawing on novel studies and the existing body of literature; no further methodological details (e.g., inclusion criteria, databases searched) provided in the excerpt.
high null result A Study on Work-Life Balance of Women Employees in the IT Se... ethical implications of AI-related job displacement
This study uses panel data on agricultural firms listed on the Shanghai and Shenzhen A-share markets from 2007 to 2023 and applies a multidimensional fixed-effects model to estimate the impact of AI on firms’ total factor productivity (TFP).
Methodological statement in the paper: dataset = panel of listed agricultural firms (Shanghai and Shenzhen A-share markets), time period 2007–2023; empirical approach = multidimensional fixed-effects model.
high null result Artificial intelligence and the sustainable development of a... study design / estimation of AI impact on total factor productivity (TFP)
Degree, betweenness, and eigenvector centrality metrics were used to identify structural vulnerabilities and leverage points in the construction supply chain network.
Paper reports calculation of degree, betweenness, and eigenvector centrality to outline vulnerabilities; specific metrics and interpretations are reported (e.g., degree centrality value for brokers).
high null result Social-Network Analytics of Construction Supply Chain network centrality measures (degree, betweenness, eigenvector) as indicators of ...
Thematic coding translated reported interactions into nodes and edges of a complex network and grouped challenges into thematic categories.
Methods described: thematic coding applied to interview data to create network structure and to generate challenge categories (six main categories, 16 open codes reported).
high null result Social-Network Analytics of Construction Supply Chain conversion of qualitative interactions into network structure and thematic categ...
This study combines empirical, semi-structured interviews with social network analytics to map construction supply chain relationships and vulnerabilities.
Methods reported in the paper: use of semi-structured interviews plus social network analysis (thematic coding to create nodes/edges, calculation of network metrics). Sample size not specified in the abstract.
high null result Social-Network Analytics of Construction Supply Chain research method integration (interviews + social network analytics)
Distinguishing between base models and fine-tuned systems is important for researchers using LLMs to study cultural patterns, because fine-tuning and alignment can change the behaviors relevant to behavioral research.
Analytical distinction and methodological guidance in the paper; claim grounded in conceptual reasoning about model development workflows rather than a specific experimental demonstration in the excerpt.
high null result The Third Ambition: Artificial Intelligence and the Science ... impact of model provenance (base vs fine-tuned) on suitability for behavioral/cu...
Contemporary artificial intelligence research has been organized around two dominant ambitions: productivity (treating AI systems as tools for accelerating work and economic output) and alignment (ensuring increasingly capable systems behave safely and in accordance with human values).
Literature synthesis and conceptual framing within the paper (review of prevailing research agendas and priorities in AI literature). No original empirical sample or experiment reported for this claim in the provided text.
high null result The Third Ambition: Artificial Intelligence and the Science ... categorization of dominant research ambitions in contemporary AI (productivity v...
This study analyzes comments and statements from party members in OECD countries from 2016 to 2025 through content analysis, examining media interviews, speeches, and debates.
Description of the study's data and method: content analysis of party member comments and statements drawn from media interviews, speeches, and debates across OECD countries over the 2016–2025 period (sample size and selection details not reported in the excerpt).
high null result Political Ideology, Artificial Intelligence (AI), and Labor ... dataset composition and methodological approach (sources and timeframe of analyz...
The study contributes to the literature by integrating evidence across higher education, vocational training, and lifelong learning to emphasize the need for balanced policy approaches to skill formation.
Stated contribution in the paper: cross-pathway synthesis of existing empirical evidence and secondary data (methods described as comparative synthesis; no primary empirical contribution reported in the summary).
high null result Balancing Higher Education, Vocational Training, and Lifelon... scholarly contribution / integrative synthesis
The study uses secondary data and comparative evidence from prior empirical studies to analyze relationships between higher education, vocational education, and lifelong learning.
Stated methodology in the paper: analysis of secondary data and synthesis of prior empirical/comparative studies (no primary data collection; no sample sizes reported).
high null result Balancing Higher Education, Vocational Training, and Lifelon... methodological approach / data sources
This study analyzed survey data from 466 Chinese food delivery riders using structural equation modeling and bootstrapping procedures, modeling work pressure as a mediator and perceived autonomy as a moderator.
Statement in abstract describing sample size (466 Chinese food delivery riders) and analytic approach (SEM and bootstrapping) and modeled variables (work pressure mediator, perceived autonomy moderator).
high null result Not all algorithmic controls are equal: the double-edged imp... methodology / analysis approach