The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests About 🎲 Workforce Futures
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Evidence (7560 claims)

Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.

The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).

Browse by theme

Nine broad, paper-level topics. Click one to filter the claims below.

Adoption
9875 claims
Filter claims →
Productivity
8807 claims
Filter claims →
Governance
7870 claims
Filter claims →
Human-AI Collaboration
7560 claims
Filtered →
Org Design
4892 claims
Filter claims →
Innovation
4781 claims
Filter claims →
Labor Markets
4004 claims
Filter claims →
Skills & Training
3308 claims
Filter claims →
Inequality
2332 claims
Filter claims →

Claims by outcome category

Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.

Outcome Positive Negative Mixed Null Total
Other 870 233 116 1066 2363
Governance & Regulation 976 451 218 133 1809
Organizational Efficiency 949 224 144 88 1416
Technology Adoption Rate 764 287 141 122 1325
Research Productivity 501 152 74 362 1101
Output Quality 542 216 69 69 896
Decision Quality 387 198 94 54 740
Firm Productivity 513 67 101 27 714
AI Safety & Ethics 249 303 73 36 667
Market Structure 190 192 134 27 548
Task Allocation 243 77 91 36 452
Innovation Output 291 33 55 20 401
Skill Acquisition 206 72 65 21 364
Employment Level 133 63 115 22 335
Fiscal & Macroeconomic 153 79 52 32 323
Task Completion Time 206 37 12 15 272
Firm Revenue 179 52 29 5 266
Consumer Welfare 130 76 47 13 266
Inequality Measures 48 137 51 6 242
Worker Satisfaction 101 81 25 13 220
Error Rate 84 110 11 5 210
Wages & Compensation 98 47 30 10 185
Regulatory Compliance 88 73 17 7 185
Automation Exposure 66 64 33 16 182
Team Performance 105 29 30 11 176
Training Effectiveness 109 22 14 21 168
Developer Productivity 114 21 14 8 158
Job Displacement 12 90 24 1 127
Hiring & Recruitment 57 9 9 5 80
Skill Obsolescence 6 56 9 1 72
Social Protection 43 17 8 2 70
Creative Output 35 21 9 4 70
Labor Share of Income 18 21 17 1 57
Worker Turnover 15 16 4 35
Industry 1 1
Clear
Human Ai Collab Remove filter
We evaluate SIA across three contrasting domains: Chinese legal charge classification (LawBench), low-level GPU kernel optimisation, and single-cell RNA denoising.
Experimental design described in the paper (three benchmark domains used for evaluation).
high null result SIA: Self Improving AI with Harness & Weight Updates domains/tasks used for evaluation
We propose SIA, a self-improving loop in which a language-model agent (the Feedback-Agent) updates both the harness and the weights of a task-specific agent.
Methodological contribution described in the paper (proposal of a new combined approach; implementation details presumably in methods).
high null result SIA: Self Improving AI with Harness & Weight Updates capability of an agent to update both harness and weights
These two silos (harness-update and test-time training) operate in isolation.
Authors' characterization of the research landscape presented in the paper (conceptual claim/literature observation).
high null result SIA: Self Improving AI with Harness & Weight Updates degree of integration between research lines
Two largely disjoint research lines attack this bottleneck: the harness-update school (a meta-agent rewrites the scaffold while model weights are fixed) and the test-time training school (hand-written RL pipelines update model weights while the harness is fixed).
Paper's literature/positioning claim classifying prior work into two categories (conceptual/literature summary).
high null result SIA: Self Improving AI with Harness & Weight Updates classification of prior research approaches
Seventeen operators completed continuous search tasks under high cognitive workload while their spatial covariance was mapped using a 2D Adaptive Riemannian Oracle.
Methodological description in the paper: 17 human operators performed continuous search tasks in a Virtual Reality drone task; spatial covariance recorded using a 2D Adaptive Riemannian Oracle.
high null result The Timing Dependencies of Trust: Speed, Accuracy, and cBCI ... experiment sample and measurement modality (operators; spatial covariance mappin...
Outputs are graded by a fact-anchored chain of rubrics, averaging 35.6 binary criteria per task.
Benchmark grading methodology reported by the authors, with a reported average of 35.6 binary criteria per task (presumably calculated across the benchmark tasks).
high null result JobBench: Aligning Agent Work With Human Will granularity of evaluation (number of binary rubric criteria per task)
JobBench covers 130 agentic tasks across 35 occupations.
Dataset/benchmark composition reported by the authors (explicit counts provided in the paper).
high null result JobBench: Aligning Agent Work With Human Will scope/coverage of the benchmark (number of tasks and occupations)
The study contributes a taxonomy of AI workforce impact, a Workforce Resilience Readiness Score (WRRS), an AI Workforce Trust Index (AWTI), an Ethical Automation Boundary concept, and a pilot empirical validation design.
Declared methodological and conceptual contributions in the paper (these are presented as deliverables of the study; no validated results reported in the excerpt).
high null result From Automation Panic to Workforce Resilience: A Governance ... new measurement/conceptual tools (taxonomy, WRRS, AWTI, Ethical Automation Bound...
The International Labour Organization's 2025 update highlights the need to assess the exposure of generative AI at the task level using task data, expert input, and AI model predictions.
Reference to ILO 2025 update recommendation described in the paper (policy/technical guidance rather than primary empirical data in the excerpt).
high null result From Automation Panic to Workforce Resilience: A Governance ... recommended assessment methods for AI exposure (task-level approach)
A path analysis was used to trace structural relationships between HR quality, effectiveness perceptions, and AI readiness.
Paper reports a path analysis linking composite HR quality indices, perceived HR effectiveness, and AI readiness measures; uses same survey sample.
high null result Determinants of Artificial Intelligence Adoption in Public S... AI readiness and perceived HR effectiveness
A binary logistic regression modelling active AI adoption was estimated with McFadden R² = 0.032.
Reported logistic regression model fit (McFadden R² = 0.032) for AI adoption outcome using the survey data.
high null result Determinants of Artificial Intelligence Adoption in Public S... active AI adoption (binary)
An OLS regression was estimated explaining perceived HR effectiveness with R² = 0.446.
Reported OLS model fit statistics in the paper (R-squared = 0.446); model explains perceived HR effectiveness using survey data.
high null result Determinants of Artificial Intelligence Adoption in Public S... perceived HR effectiveness
Constructed and validated a composite index of external HR quality factors with Cronbach's α = 0.959.
Measurement validation reported in the paper; Cronbach's alpha reported for external HR factors.
high null result Determinants of Artificial Intelligence Adoption in Public S... external HR quality index reliability
Constructed and validated a composite index of internal HR quality factors with Cronbach's α = 0.924.
Measurement validation reported in the paper; Cronbach's alpha reported for internal HR factors.
high null result Determinants of Artificial Intelligence Adoption in Public S... internal HR quality index reliability
A large-scale empirical survey of 12,562 public servants was conducted in June 2025 in Kazakhstan.
Statement in paper specifying survey sample and date; sample of public servants N = 12,562, June 2025.
high null result Determinants of Artificial Intelligence Adoption in Public S... AI adoption determinants (survey data collection)
We compare and benchmark strategy profiles adopted by open and proprietary state-of-the-art language models deployed in AgentSociety against best response.
Empirical benchmarking experiments comparing multiple language models' strategy profiles to best-response strategies (experimental evaluation / benchmarking).
high null result AgentSociety: Incentivizing Agentic Social Intelligence strategy profiles of open and proprietary language models versus best-response
Identification limits prevent a strict causal claim; the paper outlines an agenda for cleaner tests.
Authors' explicit caveat in the abstract noting limits to identification and stating they outline future cleaner tests.
high null result Coding Beyond Your Training: Claude Code and the Technologic... causal identification credibility / limitations
The analysis exploits the staggered rollout of Claude Code across GitHub between May 2025 and January 2026, using a panel of 5,838 developers observed monthly over 28 months, with treatment defined by a developer's first Claude-co-authored commit and not-yet-treated developers as controls, and estimates obtained via the doubly robust Callaway and Sant'Anna (2021) estimator.
Methods and data description as stated in the abstract: staggered rollout timing, sample size (5,838), observation window (28 months), treatment definition (first Claude-co-authored commit), estimator (Callaway & Sant'Anna 2021).
high null result Coding Beyond Your Training: Claude Code and the Technologic... study design / identification strategy
Results are robust to two stricter activity filters.
Robustness checks reported in the paper applying two stricter activity filters to the sample; claim refers to consistency of estimated effects under these alternate sample definitions.
high null result Coding Beyond Your Training: Claude Code and the Technologic... sensitivity/robustness of estimated treatment effects to stricter activity filte...
We conducted a global large-scale randomized field experiment, delivering customized LLM-generated feedback for over 31,000 arXiv preprints across 150 fields and more than 45,000 researchers from 133 geographic regions.
Statement in paper describing experimental design and scale: randomized field experiment; sample described as >31,000 preprints, >45,000 researchers, 150 fields, 133 regions.
high null result Human-AI Collaboration in Science at Scale: A Global Large-s... n/a (description of experimental sample and coverage)
Decision-makers (DMs) are similarly ambiguity-seeking and ambiguity-generated insensitive (a-insensitive) regardless of whether the analyst is human or a machine learning (ML) model.
Incentivized laboratory experiment in which participants' ambiguity attitudes were measured for forecasts attributed to human and ML analysts; comparison of ambiguity-seeking and a-insensitivity across analyst type reported in the paper (sample size not reported in abstract).
high null result Trusting human versus machine predictions as a decision unde... ambiguity attitude (ambiguity-seeking and a-insensitivity)
There is a significant deficiency in India-centric qualitative investigations on human-AI collaboration in the IT sector.
Authors' review of peer-reviewed literature and secondary data concluding a gap in India-focused qualitative studies (literature gap analysis). No numeric count provided.
high null result Human–AI Collaboration in the Indian IT Industry: A Qualitat... quantity/coverage of India-centric qualitative research
The same bias was not observed when imagining help from another human participant.
Empirical comparison reported in the abstract: predictions about receiving help from another human did not show the same faster-than-reality bias as predictions about AI assistance (from the same preregistered study, N = 1237).
high null result Cognitive offloading and the speedup illusion in human-AI in... predicted completion time when imagining help from another human
Actual completion times between independent completion and AI-assisted completion did not differ.
Empirical result reported in the abstract comparing measured completion times for independent vs. AI-assisted task completion in the preregistered study (N = 1237).
high null result Cognitive offloading and the speedup illusion in human-AI in... actual completion time
We conducted a preregistered large-scale behavioral study (N = 1237) to characterize mismatches between expectations and reality, with a focus on simple cognitive tasks.
Authors report study design and sample size in the abstract: preregistered behavioral experiment with N = 1237 participants.
high null result Cognitive offloading and the speedup illusion in human-AI in... study design / sample size (methodological claim)
The degree of persuasiveness for LLM-based narrative explanations did not meaningfully impact decision accuracy over a simple AI prediction alone.
Large-scale human behavioral experiment comparing decision accuracy with AI prediction alone versus AI prediction plus narrative explanations of varying persuasiveness (method described in paper).
We sample 50 benchmark games from a 2,000-game generated pool and evaluate nine frontier and open-weight LLMs in a head-to-head tournament with over 36,000 matches.
Empirical setup reported in the paper's abstract: 50 sampled games, 2,000-game pool, nine LLMs, >36,000 head-to-head matches.
high null result GENSTRAT: Toward a Science of Strategic Reasoning in Large L... evaluation sample size / tournament scale (matches run)
We interviewed 24 product-focused individuals at a large technology firm about how AI has impacted their own work, their work within their product team, and their professional interactions.
Qualitative semi-structured interviews with 24 product-focused employees at a single large technology firm; sample size = 24.
high null result Beyond the Org Chart: AI and the Transformation of Invisible... description of sample and data collection
This study is a systematic literature review conducted following PRISMA 2020 guidelines synthesizing peer-reviewed studies published between 2019 and 2025 identified via searches in Scopus, Web of Science and Google Scholar.
Author-stated methodology in the paper: PRISMA 2020 systematic literature review covering 2019–2025 with database searches in Scopus, Web of Science, and Google Scholar.
high null result Yapay Zeka Sistemleri ve İnsan İşbirliğinin Psikolojik, Sosy... scope and coverage of literature search / methodological transparency
This scoping review adhered to the PRISMA-ScR guidelines and encompassed 29 peer-reviewed empirical studies published from 2020 to 2025.
Methods statement in the paper (explicit methodological description).
high null result The influence of AI-Driven Employee Performance Management (... scope and methodological adherence of the review (PRISMA-ScR; n=29 studies)
Large language models are routinely used as automated evaluators (to review code, moderate content, or score outputs), often with many items passing through one conversation.
Background/introductory claim in the paper describing common practice; not an experimental result but contextual motivation.
high null result AMEL: Accumulated Message Effects on LLM Judgments prevalence of LLM use as automated evaluators
Position of biased turns does not matter: five biased turns placed anywhere in a 50-turn history produce the same shift.
Follow-up experiment manipulating the positions of biased turns within 50-turn histories and observing equivalent bias magnitudes.
high null result AMEL: Accumulated Message Effects on LLM Judgments dependence of AMEL on the position of biased messages in conversation history
Bias does not grow with context length: 5 prior turns and 50 produce the same shift (Spearman |r| < 0.01; OLS slope p = 0.80).
Correlation and OLS analysis of bias magnitude versus context-length (number of prior turns) reported in the experiments.
high null result AMEL: Accumulated Message Effects on LLM Judgments relationship between context length and magnitude of AMEL
We conducted 75,898 API calls to 11 models from 4 providers (OpenAI, Anthropic, Google, and four open-source models).
Descriptive statement of the experimental scope reported in the paper: total number of API calls and models/providers tested.
high null result AMEL: Accumulated Message Effects on LLM Judgments experimental sample size / scope (number of API calls and models)
When execution is standardized on a cheaper Gemini Flash scaffold (separating planning from execution), a pooled 32-game planner bakeoff is consistent with near-equality (p approx 0.821).
Empirical experiment: 32-game planner-only comparison where execution was standardized; reported p-value ≈ 0.821 indicating no significant difference among planners.
high null result Evaluating Large Language Models as Live Strategic Agents: P... planner performance equality (pooled test)
We study this setting in a timed multi-phase Risk environment with explicit victory targets and repeated planning and execution cycles.
Methodological description of the experimental environment used in the paper (timed multi-phase Risk environment with explicit victory targets and repeated cycles).
high null result Evaluating Large Language Models as Live Strategic Agents: P... experimental_environment_description
Identification of effects uses within-firm variation with firm and city-by-year fixed effects.
Identification strategy reported in abstract: within-firm variation under firm and city-by-year fixed effects.
high null result Toward Sustainable Workforce Development: How AI Reshapes Sk... identification approach / econometric controls
The study measures four skill-category demand shares and their within-category importance from job-description text.
Methodological statement in abstract: measurement of four skill-category demand shares and within-category importance via job-description text.
high null result Toward Sustainable Workforce Development: How AI Reshapes Sk... skill-category demand shares and within-category importance
AI exposure is decomposed into displacement and augmentation components based on task routineness.
Methodological claim in abstract: decomposition of exposure into displacement and augmentation using a routineness criterion for tasks.
high null result Toward Sustainable Workforce Development: How AI Reshapes Sk... decomposed AI exposure measures (displacement vs augmentation)
The authors construct firm-by-year potential AI exposure via semantic matching between AI patent texts and detailed occupation task descriptions.
Method description in abstract: semantic matching of AI patent texts to occupation task descriptions to build firm-by-year exposure.
high null result Toward Sustainable Workforce Development: How AI Reshapes Sk... firm-by-year potential AI exposure (constructed measure)
The study uses approximately 67 million online job postings from two major Chinese recruitment platforms (2019–2024).
Statement in paper abstract describing dataset size and source (job postings from two major Chinese recruitment platforms over 2019–2024).
high null result Toward Sustainable Workforce Development: How AI Reshapes Sk... dataset size and coverage (number of job postings, platforms, years)
The study extends the Technology Acceptance Model (TAM), Dynamic Capabilities Theory, and the Technology-Organisation-Environment (TOE) framework into the qualitative, emerging-economy entrepreneurial context.
Authors' stated theoretical contribution based on mapping thematic results to TAM, Dynamic Capabilities, and TOE frameworks within analysis and discussion sections.
high null result Navigating the Intelligence Frontier: AI Adoption as a Succe... theoretical contribution / framework extension
This study employed an interpretivist, qualitative research design using sixteen in-depth semi-structured interviews with entrepreneurs across fintech, edtech, health-tech, logistics, retail, and SaaS in Delhi/NCR, India, and used Braun & Clarke's (2006) six-phase thematic analysis framework.
Explicit methodological description in the paper: interpretivist qualitative design; n=16 in-depth semi-structured interviews across specified sectors in Delhi/NCR; thematic analysis following Braun & Clarke (2006).
high null result Navigating the Intelligence Frontier: AI Adoption as a Succe... research design / data collection (qualitative interviews)
Using a qualitative approach with 17 expert interviews from employees at startups.
Methods statement in paper specifying qualitative study design and sample size of 17 interviews.
high null result From Prompt To Process: Qualitative Insights On How Genai Us... study methodology and sample
Process-related insights into how GenAI transforms startups are limited.
Authors' literature positioning / gap statement in paper (no empirical metric provided).
high null result From Prompt To Process: Qualitative Insights On How Genai Us... availability of process-related insights in literature
The paper's findings are based on three pre-registered user studies with a combined sample size of N = 2691.
Statement in the paper's abstract reporting three pre-registered user studies and combined N = 2691.
high null result The efficiency-gain illusion: People underestimate the rate ... study_sample_description
Light AI users perform similarly to matched users who do not use AI.
Same controlled logical reasoning experiment with on-demand AI assistance comparing light AI users to matched non-users (sample size not stated in abstract).
high null result The Impact of AI Usage and Informativeness on Skill Developm... post-AI performance / skill development
We map that space through six interconnected elements: sociotechnical context, decision-making frameworks, human decision participants, AI capabilities, interaction, and holistic evaluation.
The paper's proposed analytical/framework contribution listing six elements (descriptive of the authors' mapping work).
high null result Addressing the Synergy Gap: The Six Elements of the Design S... n/a (framework description)
Most current work treats human-AI combination as an engineering problem and concentrates on interpretability, trust calibration, or interface design.
Authors' characterization of the existing literature and dominant research foci (qualitative literature assessment; no quantitative breakdown provided).
high null result Addressing the Synergy Gap: The Six Elements of the Design S... research focus/themes in human-AI combination literature
We call this persistent shortfall the 'synergy gap.'
Terminology/definition introduced by the authors in the paper (conceptual claim, not an empirical finding).
high null result Addressing the Synergy Gap: The Six Elements of the Design S... n/a (terminology defining a phenomenon)