The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Evidence (1624 claims)

Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.

The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).

Browse by theme

Nine broad, paper-level topics. Click one to filter the claims below.

Adoption
20058 claims
Filter claims →
Productivity
17184 claims
Filter claims →
Governance
16099 claims
Filter claims →
Human-AI Collaboration
16034 claims
Filter claims →
Innovation
10501 claims
Filter claims →
Org Design
10496 claims
Filter claims →
Labor Markets
6444 claims
Filter claims →
Skills & Training
5385 claims
Filter claims →
Inequality
4148 claims
Filter claims →

Claims by outcome category

Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.

Outcome Positive Negative Mixed Null Total
Other 1820 479 278 1820 4588
Organizational Efficiency 2711 616 401 173 3922
Governance & Regulation 2075 886 459 246 3714
Technology Adoption Rate 1467 530 258 206 2488
Decision Quality 1281 496 289 152 2228
Output Quality 1227 447 207 138 2025
AI Safety & Ethics 634 754 207 83 1688
Research Productivity 826 241 114 422 1624
Firm Productivity 1052 154 163 66 1441
Task Allocation 685 211 331 99 1335
Market Structure 433 423 242 46 1150
Innovation Output 639 91 105 34 871
Task Completion Time 476 113 43 36 672
Firm Revenue 445 126 58 25 656
Skill Acquisition 364 119 109 34 626
Consumer Welfare 288 167 104 31 592
Employment Level 214 140 174 50 582
Error Rate 230 251 35 16 535
Fiscal & Macroeconomic 268 136 71 50 532
Inequality Measures 100 307 96 12 515
Worker Satisfaction 221 173 60 30 484
Automation Exposure 155 138 65 36 398
Regulatory Compliance 171 120 30 13 335
Developer Productivity 222 58 27 13 321
Team Performance 188 56 50 24 320
Wages & Compensation 146 104 46 16 312
Training Effectiveness 207 41 21 26 298
Job Displacement 23 153 52 4 232
Hiring & Recruitment 102 57 30 11 202
Skill Obsolescence 16 102 24 6 148
Creative Output 71 42 23 6 143
Social Protection 57 30 11 3 101
Labor Share of Income 29 42 24 2 97
Worker Turnover 43 29 6 4 82
Industry 1 1
The AI Publication Footprint is an absolute cumulative Scopus publication count and therefore serves as a proxy for knowledge-production capacity rather than a normalized measure of research intensity or efficiency.
Measurement definition and methodological caveat; publication counts were not normalized per capita or per researcher.
high mixed AI Publication Footprint and National AI Readiness: Global G... National AI knowledge-production capacity as measured by cumulative publication ...
The Station outperformed AlphaEvolve on three of the seven AlphaEvolve problems that did not produce results novel relative to the prior literature, matched it on two, and underperformed it on two.
Portfolio-level comparison across all 12 evaluated problems, categorized by relative performance.
high mixed Autonomous Mathematical Discovery in an Open-World Multi-Age... Relative benchmark performance against AlphaEvolve
GPU count showed more consistent positive associations with citation and award outcomes than newer hardware generation; GPU count was positively associated with awards, whereas aggregate capability and hardware generation showed no comparably robust award evidence.
Complementary citation analyses and award models, including linear-probability and Firth models; the award analysis used 5,357 papers.
high mixed More Computational Resources Do Not Ensure Higher Scholarly ... Citation impact and paper-award status
Reported GPU capability is neither necessary nor sufficient for high citation impact: 85.5% of high-capability papers were not highly cited, and most highly cited papers were outside the high-capability group.
Overlap analysis between the top 20% of papers by reported GPU capability and the top 10% by citations within publication-year-by-venue groups.
high mixed More Computational Resources Do Not Ensure Higher Scholarly ... Overlap between high reported GPU capability and high citation impact
The annual top 20% of papers ranked by reported GPU capability accounted for 83.9%–89.9% of reported GPU capability during 2020–2023, but only 27%–32% of citations and 20%–33% of paper awards.
Annual concentration analysis comparing the share of total reported GPU capability, citations, and awards attributable to the top 20% of GPU-quantifiable papers.
high mixed More Computational Resources Do Not Ensure Higher Scholarly ... Concentration of reported GPU capability versus citations and paper awards
Research-topic diversification and collaboration with intermediary institutions help explain why sanctions reduce publication quantity while increasing publication quality.
Mediation-style analyses link changes in research topics and coauthorship networks to the observed quantity and quality outcomes, with semi-structured interviews providing contextual validation.
high mixed Science under sanctions: The impact of the entity list on Ch... Publication quantity and bibliometric quality
Randomization does not reliably improve performance across different problem-solving contexts; its usefulness remains restricted to decomposable problems.
Simulation of a randomized AI recommendation strategy that samples from top-performing solutions rather than returning a single best recommendation, compared across problem structures and use rates.
high mixed Navigating Epistemic Monocultures in AI-Driven Science: A Si... Collective problem-solving performance
The agent harness can substantially change the capability expressed by the same underlying model.
Cross-harness comparisons: MiMo V2.5 Pro scored 16.17 with MiMo Code versus 23.25 with Claude Code, and Kimi K2.7 scored 19.72 with Kimi Code versus 27.34 with Claude Code.
high mixed ASI-Bench: At the Dawn of Artificial Superintelligence Scientific research score as a function of agent harness
As AI makes proof generation and verification cheaper, mathematical value is shifting toward forms of mathematical meaning-making that current systems cannot yet perform.
Argument based on Tao's account of increasing proof abundance and the paper's taxonomy of concept and conjecture formation; no quantitative estimate of the shift is provided.
high mixed Assessing LLMs' mathematical abilities requires understandin... relative scarcity and value of mathematical activities
The held-out-paper evaluation protocol can test whether the researcher context carries researcher-specific signal, but it cannot establish that a recommended alternative research direction is valuable.
The paper proposes holding out a paper and measuring fidelity and contrast, then explicitly limits what this protocol can validate because unpursued alternatives lack observed ground truth.
high mixed Personalized Auto-Research: Towards a True AI Co-Scientist Validity of held-out evaluation as a measure of personalization signal versus re...
The multi-agent pipeline's advantage should not be interpreted as an isolated causal effect of collaboration, because it also selects the best five hypotheses from 20 candidates rather than retaining all five hypotheses from one model, and the conditions use different judge panels.
The paper explicitly identifies inference-time scaling and different evaluation panels as confounds and characterizes the comparison as an observed association.
high mixed Reconstruction: A Blind Benchmark for Recovering Research Id... Interpretation of the multi-agent versus single-model Match-rate difference
The agentic RPM reaches or exceeds the performance of the No-RPM baseline, although its early progress is slower because it runs small-scale proxy experiments at each step.
Comparison of performance trajectories over compute time in the end-to-end AIRS-Bench evaluation; the paper attributes the slower early trajectory to per-step sandbox experiments.
high mixed AI Research Preference Models Performance trajectory and average normalized score over compute time
AI adoption may increase dispersion in publication or research-output rates across researchers and fields rather than producing uniform productivity gains.
Derived empirical prediction from the theoretical model of plan-relative complementarity and heterogeneous workflows; no direct empirical estimate is reported.
high mixed How to Do Research With <scp>AI</scp> : An Austrian Capital ... Dispersion of publication and research-output rates
The productive effects of AI in scholarly production are plan-relative and heterogeneous: AI produces the largest gains when it fills specific gaps in a researcher's workflow and when the researcher has strong complementary human capital.
Logical analysis of complementarity and substitution in plan space, supported by conceptual examples and thought experiments.
Models with similar final outcomes can have substantially different execution and feedback-control capabilities.
GPT-5.5 and Gemini-3.1-Pro had similar outcome and Solution Framing scores but different Execution and Feedback Control scores.
high mixed Beyond Final Scores: A Systematic Evaluation of Agents for L... Process capability scores: execution reliability and feedback control
Lower-ranked models can reach competitive solutions but do so less consistently across repeated runs.
The paper compares average and best scores across three independent rollouts per model-task pair and notes that Kimi's best@3 is close to higher-ranked models while its avg@3 is substantially lower.
high mixed Beyond Final Scores: A Systematic Evaluation of Agents for L... Consistency and peak performance across repeated agent runs
Model differences are larger in typical performance than in peak observed performance: the highest-to-lowest gap is 0.237 for avg@3 versus 0.122 for best@3.
Comparison of avg@3 and best@3 across seven models, each evaluated with three independent rollouts on all 36 tasks.
high mixed Beyond Final Scores: A Systematic Evaluation of Agents for L... Between-model variation in average and best-observed task scores
Current automated research agents operate more like engineering optimizers than fully autonomous researchers.
Systematic evaluation of seven frontier models across 36 long-horizon AI R&D tasks, including process metrics, experience-reuse comparisons, and novelty review.
high mixed Beyond Final Scores: A Systematic Evaluation of Agents for L... Overall autonomous research capability, including solution formulation, implemen...
The review finds that the existing literature is fragmented and calls for causal and microdata studies, particularly on post-2020 developments and cross-country regulatory differences.
Systematic review of 660 documents and the authors' methodological implications regarding limitations of descriptive bibliometric evidence.
high mixed Mapping Data-Driven Governance in Sharing Economy Platforms:... Methodological maturity and evidence quality of the research field
The literature exhibits geographic heterogeneity in research emphasis and leadership despite expanding international participation and collaboration.
Geographic distribution analysis combined with author and institutional collaboration-network analysis.
high mixed Scientific Mapping of Green Finance: A Bibliometric Analysis... Geographic distribution of research emphasis, participation, and leadership
The intellectual structure of virtual influencer research is fragmented, with influence distributed across multiple research teams.
Analysis of publications, citations, and co-authorships in the 116-study bibliometric dataset.
high mixed Virtual influencers in advertising and marketing communicati... Concentration and structure of scholarly influence and collaboration
At a fixed flagging rate of p = 0.2%, the top-α and bottom-α flagging rules differed in their estimated productivity changes by 15.3%, with the top-α rule above and the bottom-α rule below the random baseline.
Rank-based comparison of top-(100p)% α papers, bottom-(100p)% α papers, and random Bernoulli(p) flags, with all rules flagging the same fraction of papers.
high mixed A robust association between LLM use and scientific producti... Change in author productivity under alternative paper-flagging rules
Open-weight models run locally offer greater control over randomness, but reproducibility still depends on the complete hardware and software stack.
The paper contrasts local execution with proprietary APIs and discusses variation from processors, drivers, numerical libraries, precision settings, batching, and model-serving environments.
high mixed Randomness in large language models: What researchers need t... Control and reproducibility of model inference
More user-centric studies reveal rich contextual detail, yet are fragmented, small-scale and loosely coupled to the mechanisms that shape model behavior.
Author synthesis/critique of user-centric research approaches; stated in the paper but no empirical meta-analysis or sample sizes provided in the excerpt.
high mixed Real-World AI Evaluation: How FRAME Generates Systematic Evi... research_scope_and_linkage_to_mechanisms
Perdomo et al. 2020 formalized this dynamic in their work on performative prediction.
Citation of Perdomo et al. 2020 as the formalization reference (literature claim).
high mixed The Stability of Online Algorithms in Performative Predictio... existence_of_a_formal_framework_for_performative_prediction
The paper introduces the concept of 'vibe researching' as the AI-era parallel to 'vibe coding.'
Conceptual contribution and definitional introduction in the paper (term coined and elaborated by the author); no empirical validation.
high mixed Vibe Researching as Wolf Coming: Can AI Agents with Skills R... conceptual framing of research practice
Heuristic approaches (prompt engineering, fine-tuning, repair strategies) are useful for many exploratory tasks but lack the formal statistical guarantees typically required for confirmatory research.
Argument made in the abstract contrasting heuristic approaches' utility for exploratory work versus their lack of formal guarantees; no empirical validation reported in the abstract.
high mixed This human study did not involve human subjects: Validating ... suitability of heuristic approaches for exploratory vs. confirmatory research (v...
Two strategies are contrasted for obtaining valid estimates of causal effects when using LLMs as synthetic participants: (1) heuristic approaches (prompt engineering, model fine-tuning, repair strategies) and (2) statistical calibration combining auxiliary human data with statistical adjustments.
Paper's methodological framing described in the abstract; classification/contrast of methodological approaches rather than empirical evaluation.
high mixed This human study did not involve human subjects: Validating ... methods for obtaining valid causal estimates from LLM-based simulations
Recognition of the 'staggered adoption' problem has shifted the focus in recent years away from inference towards consistent estimation of treatment effects.
Descriptive claim in the paper about trends in the literature; references to recent methodological work (e.g., Callaway and Sant'Anna (2021)).
high mixed Improved Inference for CSDID Using the Cluster Jackknife research focus/priorities in the econometrics literature (estimation vs inferenc...
The paper concludes by assessing interconnected dimensions (of post-labor futures) to establish a basis for further exploration of potentially transformative economic shifts.
Stated conclusion and research agenda in the review; synthesis of multiple thematic dimensions (qualitative synthesis, no quantitative sample).
high mixed Post-Labor Economics: A Systematic Review research agenda / preparedness for transformative shifts
The goal of the note is to highlight the fragility of existing forecasts of exponential growth in AI capabilities rather than to establish a rigorous forecast of our own.
Author statement of intent / framing in the paper (method: argumentative / critical appraisal of forecasting methods).
high mixed Are AI Capabilities Increasing Exponentially? A Competing Hy... robustness/fragility of exponential-growth forecasts
We propose a more complex model that decomposes AI capabilities into base and reasoning capabilities, each exhibiting individual rates of improvement.
Presentation of a theoretical / mathematical model in the note that decomposes capabilities into two components (model specification provided in the paper). This is a modeling/theoretical claim rather than an empirical estimate; sample size not applicable.
high mixed Are AI Capabilities Increasing Exponentially? A Competing Hy... structure of capability growth (base vs reasoning components and their rates)
The literature on AI-enabled finance transformation is fragmented across FP&A, record-to-report, reporting, audit/assurance, and governance, producing heterogeneous findings.
Systematic literature review (SLR) of peer-reviewed articles with database searches and snowballing; studies mapped across functional silos in the corporate finance function (as reported in the paper).
high mixed PEMETAAN DOMAIN PROSES, OUTCOME, DAN TATA KELOLA AI-ENABLED ... fragmentation of literature and heterogeneity of findings
The magnitude of the increase in paper production among LLM adopters varies substantially by scientific field and author background (range reported 23.7–89.3%).
Stratified analyses across fields and author background in the 2.1M preprint dataset showing heterogeneous percentage increases in production depending on field and author characteristics.
high mixed Scientific production in the era of Large Language Models heterogeneity in change in paper production across fields and author backgrounds
Recovering ground truth causal effects is feasible -- but only with careful modeling choices.
Empirical comparison in the paper between experimental (ground-truth) treatment effects and estimates from observational causal ML approaches applied to the opt-in sample; the provided text indicates such comparisons were made but does not report sample size or exact methods/results.
high mixed Reevaluating Causal Estimation Methods with Data from a Prod... accuracy/ability of observational methods to recover experimental (ground-truth)...
Literature addressing this gap spans three disconnected streams: (1) AI governance frameworks relying on subjective risk assessment, (2) technical ML bias measurement focused on algorithmic fairness, and (3) organizational implementation approaches failing to translate governance principles into operational practice.
Authors' literature mapping/synthesis identifying three streams (conceptual literature review).
high mixed Making AI Risk Assessment More Objective: Addressing Sociote... structure and connectedness of relevant literature streams
AI labor-impact predictions are based on multimodal data and labor-force productivity modeling relating job and skill creation to economic conditions and workforce development.
Methodological statement from the systematic review describing how predictions in the literature use multimodal data and productivity models (2024–2025 corpus).
high mixed The algorithmic management of job loss and creation in the e... predictive modeling of labor impacts
Intrinsically interpretable architectures, notably attention-based transformers, are discussed as approaches for time-series explainability.
Survey discussion of model classes and interpretability properties, with attention-based transformers highlighted as a notable interpretable architecture.
high mixed Explainable Artificial Intelligence for Economic Time Series... intrinsic interpretability of model architectures
Overall, the LLM era coincides with a broader reorganization of scientific exploration, collaboration, and the division of labor.
Aggregate observational findings from linked PubMed Central and OpenAlex datasets covering 775,323 scientists and 137,120 multi-author papers with CRediT statements, comparing patterns before and after 2022.
high mixed Scientific exploration, collaboration and labor division in ... aggregate changes in exploration, collaboration, and division of labor
Through simulations and an empirical application, the paper illustrates the main advantages and limitations of each approach (CBCMs and FBCMs).
Reported simulation studies and at least one empirical application described in the paper (methods section and results).
high mixed Identifying Treatment and Spillover Effects with Control-Bas... comparative performance (advantages and limitations) of CBCMs and FBCMs
Although AI working autonomously achieved a 37% reproduction rate, it could be useful for automated screening when human review is cost-prohibitive.
Interpretation in paper: authors note 37% autonomous reproduction rate as potentially useful for large-scale screening where human review is infeasible; based on empirical results of the experiment.
high mixed AI-assisted teams outperform AI-led teams but not human-only... potential_value_for_screening
Empirical claims across the reviewed literature vary in methodological rigor and should be viewed with caution before standardized replication.
Meta-level assessment presented in the review of peer‑reviewed literature (2020–2025); no formal quality-assessment statistics provided in the excerpt.
high mixed From data to decisions: A narrative review of business intel... methodological rigor / reproducibility of empirical studies
The literature's vocabulary ("self-refine," "self-reward," "self-play," "self-evolve") conflates fundamentally different ambitions.
Qualitative analysis of terminology across the surveyed arXiv papers (2024-2026) reported in the paper's survey and taxonomy section.
high mixed Recursive Self-Improvement in AI: From Bounded Self-Refineme... terminology/conceptual clarity in literature
The CAD is formalized with a probabilistic model grounded in the fan effect literature in cognitive psychology.
Paper reports a formal probabilistic model drawing on the fan effect literature; model described as the formalization of CAD.
high mixed The Context Access Divide: Interaction-Level Architecture as... formal modeling of context-access effects (theoretical task-success dynamics)
Across four high-stakes domains, assigning different personas is sufficient for AI agents to report divergent, often opposing, conclusions from the same data and question, with findings systematically aligned with those beliefs.
Experimental manipulation across four domains where AI agents were assigned different personas and produced analyses from the same data/question; comparison of resulting conclusions showing divergence and alignment with persona beliefs.
high mixed The Agentic Garden of Forking Paths direction and content of reported conclusions by AI agents given persona assignm...
Important gaps remain in the literature and warrant further research.
Paper's abstract statement that the review identifies important gaps that warrant further research (based on review of 194 articles).
The existing literature on AI and economic development remains fragmented, with limited integration across development dimensions.
Conclusion drawn in the abstract from the systematic review of 194 peer-reviewed articles noting fragmentation and limited cross-dimension integration.
high mixed Artificial Intelligence and Economic Development: A Systemat... literature_integration / interdisciplinarity
The research uses a sequential multi-phase design combining experiments and qualitative fieldwork.
Stated methodology in the abstract (methodological claim about study design). No sample sizes or procedural details provided in the excerpt.
high mixed Strategic Adoption of AI-Enabled Decision-Making Systems: De... methodological approach to studying managerial agency
Evidence on the productivity, risk, and resilience implications of AI adoption remains fragmented and dispersed across different fields of research.
Author's assessment of the literature based on the systematic review (PRISMA) of 68 empirical studies published 2015–2025.
high mixed AI Adoption in Local Government: Productivity, Systemic Risk... state of evidence (fragmentation across fields)
Across compression sweeps, real factor archives, and LLM-SRBench tasks, hybrid gains concentrate in weakly represented but target-bearing directions and vanish as the hypothesis space approaches full rank.
Empirical claim based on experiments over compression sweeps, analyses of real factor archives (A-share factor discovery), and LLM-SRBench tasks; no numerical sample sizes or effect magnitudes provided in the abstract.