The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests About 🎲 Workforce Futures
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Evidence (8974 claims)

Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.

The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).

Browse by theme

Nine broad, paper-level topics. Click one to filter the claims below.

Adoption
10085 claims
Filter claims →
Productivity
8974 claims
Filtered →
Governance
8062 claims
Filter claims →
Human-AI Collaboration
7749 claims
Filter claims →
Org Design
5057 claims
Filter claims →
Innovation
4896 claims
Filter claims →
Labor Markets
4088 claims
Filter claims →
Skills & Training
3372 claims
Filter claims →
Inequality
2377 claims
Filter claims →

Claims by outcome category

Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.

Outcome Positive Negative Mixed Null Total
Other 882 244 117 1097 2424
Governance & Regulation 1010 469 229 135 1875
Organizational Efficiency 977 235 149 90 1462
Technology Adoption Rate 781 299 143 128 1362
Research Productivity 506 155 74 363 1110
Output Quality 555 219 71 70 915
Decision Quality 395 200 95 54 751
Firm Productivity 523 67 101 27 724
AI Safety & Ethics 262 309 75 36 688
Market Structure 195 201 135 30 566
Task Allocation 248 77 96 38 464
Innovation Output 300 34 55 20 411
Skill Acquisition 207 75 65 21 368
Employment Level 138 67 119 24 350
Fiscal & Macroeconomic 156 80 53 33 329
Task Completion Time 211 38 13 16 280
Firm Revenue 183 52 29 5 270
Consumer Welfare 131 77 48 13 269
Inequality Measures 50 141 54 9 254
Worker Satisfaction 104 85 25 13 227
Error Rate 87 112 11 5 215
Automation Exposure 69 69 37 20 198
Wages & Compensation 102 49 31 11 193
Team Performance 115 30 30 11 187
Regulatory Compliance 88 74 17 7 186
Training Effectiveness 109 22 14 21 168
Developer Productivity 116 21 15 8 161
Job Displacement 12 92 26 1 131
Hiring & Recruitment 57 12 9 5 83
Skill Obsolescence 6 59 10 2 77
Social Protection 43 17 8 2 70
Creative Output 35 21 9 4 70
Labor Share of Income 18 23 17 1 59
Worker Turnover 15 16 4 35
Industry 1 1
Clear
Productivity Remove filter
The performance gain of human-AI collaboration over the model-only baselines is significant, irrespective of model size.
Reported results from the experimental comparison across conditions and three model sizes (3B, 8B, 70B) with N=112 participants; paper states the performance gain is significant across sizes (no numeric effect sizes or p-values provided in the excerpt).
high positive Seeking Information with RAG-Assistants: Does Model Size Mat... task accuracy / performance
The framework addresses AI-specific challenges including model versioning, human-AI interaction dynamics, contamination and spillover effects, and equitable impact assessment.
Paper lists and provides guidance on AI-specific methodological issues (model versioning, interaction dynamics, contamination/spillover, equity). This is a descriptive claim about topics the framework covers, not an empirical evaluation of solutions.
high positive Principles and Guidelines for Randomized Controlled Trials i... coverage of AI-specific methodological challenges in evaluation guidelines
The framework implements a graded transparency and repeatability framework.
Paper extends TOP-guideline-derived transparency principle into a graded scheme for transparency and repeatability; described as an operational feature of the proposed framework.
high positive Principles and Guidelines for Randomized Controlled Trials i... graded transparency and repeatability practices for AI RCTs
The framework integrates heterogeneity analysis and practical significance assessment.
Paper reports inclusion of guidance on analyzing heterogenous treatment effects and assessing practical significance; presented as part of guidelines rather than tested across datasets.
high positive Principles and Guidelines for Randomized Controlled Trials i... inclusion of heterogeneity and practical significance analysis in evaluation pra...
The framework formalizes causal inference through RCT methodology for AI contexts.
Paper states adoption of randomized controlled trial methods and causal inference framing for AI impact evaluation; described as methodological proposition rather than validated application.
high positive Principles and Guidelines for Randomized Controlled Trials i... use of RCTs to support causal inference in AI evaluations
Our framework extends prior work by centering evaluation on human performance rather than model output alone.
Paper claims a conceptual shift: focus on human performance metrics; supported by argumentative rationale and literature references rather than empirical demonstration.
high positive Principles and Guidelines for Randomized Controlled Trials i... focus of evaluation metrics (human performance vs. model output)
The principles and guidelines serve three key roles for AI evaluation RCTs: a design tool for planning studies, an evaluation rubric for assessing existing work, and a blueprint for standard setting as the field converges on norms.
Paper's stated intended uses/positioning of the framework; presented as roles in the discussion/positioning section rather than empirically validated roles.
high positive Principles and Guidelines for Randomized Controlled Trials i... utility of the framework in planning, evaluating, and standard-setting
We operationalize all five principles into 33 guidelines adapted for AI evaluation RCT contexts, expressed as requirements with rationales, implementation instructions, and evidence bases.
Paper reports a concrete output: 33 guidelines derived from the five principles, with each guideline presented as requirement + rationale + implementation instructions + evidence base (documented in paper content).
high positive Principles and Guidelines for Randomized Controlled Trials i... availability of operational guidelines for AI RCTs
The paper adopts the (Shadish et al., 2002) four-validity framework and extends it with a fifth principle on transparency, repeatability, and verification adapted from the Transparency and Openness Promotion (TOP) Guidelines (Center for Open Science, 2025).
Explicit methodological choice described in the paper: adoption of Shadish et al. four-validity framework and addition of a transparency/repeatability principle based on TOP Guidelines; documented in the text as design decision.
high positive Principles and Guidelines for Randomized Controlled Trials i... methodological framework / validity criteria
The framework draws on established experimental practices from disciplines with established RCT traditions, including software engineering, economics, clinical and health sciences, and psychology.
Paper reports literature review and cross-disciplinary synthesis as the methodological foundation for the framework (references to those disciplines). No empirical cross-disciplinary experiment reported.
high positive Principles and Guidelines for Randomized Controlled Trials i... methodological comprehensiveness / interdisciplinary grounding
This work establishes a foundational framework for standardizing AI evaluation RCTs (sometimes called human uplift studies).
Paper's stated contribution: development of a conceptual framework integrating RCT design principles for AI evaluation. Based on literature synthesis and methodological argumentation rather than empirical testing.
high positive Principles and Guidelines for Randomized Controlled Trials i... standardization of AI evaluation RCTs / evaluation methodology
The paper introduces a Specification Governance Model (SGM), grounded in Transaction Cost Economics, and provides a practical governance decision guide.
Conceptual/modeling contribution described in the paper: SGM grounded in TCE with an applied decision guide (theoretical plus prescriptive).
high positive The Productivity-Reliability Paradox: Specification-Driven G... governance decision-making for specification practices
The paper proposes the AI-Augmented Methodology Taxonomy (AAMT), classifying six methodologies under three AI integration tiers.
Conceptual contribution: taxonomy introduced and described in the paper (six methodologies, three tiers).
high positive The Productivity-Reliability Paradox: Specification-Driven G... existence and classification of methodologies (taxonomic contribution)
Telemetry across 10,000+ developers shows a 98% increase in pull requests.
Observational telemetry data aggregated across >10,000 developers reported in the paper; metric reported is percent increase in pull request count.
high positive The Productivity-Reliability Paradox: Specification-Driven G... number of pull requests (pull_request_count)
Controlled studies report 20-56% productivity gains on well-scoped tasks.
Aggregate of multiple controlled experimental studies cited in the paper (2022–2026); reported as observed productivity improvements on well-scoped tasks in those studies. Specific study-level sample sizes not reported in the claim text.
The platform was used to support compound AI use cases at Salesforce, specifically Agentforce (autonomous AI agents) and ApexGuru (AI-powered code analysis).
Paper states the deployment was developed at Salesforce and lists Agentforce and ApexGuru as supported use cases; this is an implementation/adoption claim rather than a quantitative result.
high positive Scalable Inference Architectures for Compound AI Systems: A ... support/adoption by named applications (Agentforce, ApexGuru)
The architecture enables compound AI systems to: (a) scale model invocations in parallel, (b) handle bursty multi-agent workloads, and (c) support rapid model iteration — capabilities essential for operationalizing agentic AI at enterprise scale.
Paper provides case studies (Agentforce, ApexGuru) and operational lessons from production deployment to support these functional claims; the provided text does not include numerical benchmarks for each capability individually nor sample sizes.
high positive Scalable Inference Architectures for Compound AI Systems: A ... scalability of model invocations, ability to handle bursty workloads, support fo...
The modular, platform-agnostic inference architecture integrates serverless execution, dynamic autoscaling, and MLOps pipelines to deliver consistent low-latency inference across multi-component agent workflows.
System design and production deployment description in the paper; claim supported by implementation details and reported production performance (qualitative and operational evidence), but no detailed experimental protocol or sample sizes are given in the provided text.
high positive Scalable Inference Architectures for Compound AI Systems: A ... consistency of low-latency inference (multi-component agent workflows)
The platform delivered 30 to 40% cost savings relative to prior static deployments.
Reported production cost comparisons between the new modular inference architecture and prior static deployments (paper states "30 to 40% cost savings"); the provided text does not include details on cost components, time period, or sample size.
high positive Scalable Inference Architectures for Compound AI Systems: A ... infrastructure / inference cost
The deployment produced up to 3.9x throughput improvement compared to prior static deployments.
Reported production results comparing throughput of the modular inference architecture to prior static deployments (statement in the paper: "up to 3.9x throughput improvement"); no sample size or confidence intervals provided in the provided text.
The production deployment achieved over 50% reduction in tail latency (P95) compared to prior static deployments.
Reported production results comparing the modular inference architecture to prior static deployments (production measurements of P95 tail latency); paper states this was observed in production but does not report sample size or detailed statistical tests in the provided text.
We release the benchmark, harness, sweep configurations, and full run corpus.
Statement of artifact release in the paper; verifiable by checking the project's repository or supplementary materials.
high positive AgentFloor: How Far Up the tool use Ladder Can Small Open-We... availability of released materials (benchmark and run corpus)
These findings suggest a practical design principle for agentic systems: use smaller open-weight models for the broad base of routine actions, and reserve large frontier models for the narrower class of tasks that truly demand deeper planning and control.
Synthesis/recommendation drawn from the empirical results on AgentFloor showing where small/mid models suffice and where frontier models have advantage; prescriptive claim rather than a direct empirical measurement.
high positive AgentFloor: How Far Up the tool use Ladder Can Small Open-We... recommended task routing strategy for agentic systems (model assignment to task ...
The gap appears most clearly on long-horizon planning tasks that require sustained coordination and reliable constraint tracking over many steps, where frontier models still hold an advantage, though neither side reaches strong reliability.
Performance breakdown by capability tier on AgentFloor showing frontier (GPT-5) advantage on long-horizon planning/constraint-tracking tasks; both model groups have low absolute reliability on these tasks according to reported results.
high positive AgentFloor: How Far Up the tool use Ladder Can Small Open-We... performance on long-horizon planning tasks (ability to sustain coordination and ...
We evaluate 16 open-weight models, from 0.27B to 32B parameters, alongside GPT-5 across 16,542 scored runs.
Empirical evaluation reported in the paper: 16 open-weight models spanning specified parameter sizes, inclusion of GPT-5, and a total of 16,542 scored runs (reported counts).
high positive AgentFloor: How Far Up the tool use Ladder Can Small Open-We... evaluation runs (model-by-task performance across 16,542 scored runs)
We introduce AgentFloor, a deterministic 30-task benchmark organized as a six-tier capability ladder, spanning instruction following, tool use, multi-step coordination, and long-horizon planning under persistent constraints.
Paper describes the design of the benchmark: deterministic, 30 tasks, organized into six tiers covering specified capabilities. This is a descriptive claim about the artifact introduced in the work.
high positive AgentFloor: How Far Up the tool use Ladder Can Small Open-We... benchmark construction (30 tasks, six-tier capability ladder)
TokenArena is a methodology, not a single ranking; we publish full provenance and limitations and welcome external replication.
Author statement about the intended use of the benchmark and transparency practices (publication of provenance and limitations).
high positive Token Arena: A Continuous Benchmark Unifying Energy and Cogn... positioning of TokenArena as a methodological framework with published provenanc...
We release the framework, schema, probe and eval harness, and a v1.0 leaderboard snapshot under CC BY 4.0.
Author statement of artifact release (license explicitly CC BY 4.0).
high positive Token Arena: A Continuous Benchmark Unifying Energy and Cogn... availability of TokenArena artifacts and leaderboard under CC BY 4.0
We introduce TokenArena, a continuous benchmark that measures inference at endpoint granularity along five core axes (output speed, time to first token, workload-blended price, effective context, and quality on the live endpoint) and synthesizes them, together with a modeled energy estimate, into three headline composites: joules per correct answer, dollars per correct answer, and endpoint fidelity (output-distribution similarity to a first-party reference).
Methodological contribution described by the authors; framework specification and composite metrics defined in the paper.
high positive Token Arena: A Continuous Benchmark Unifying Energy and Cogn... five core axes (output speed, time to first token, workload-blended price, effec...
Qiushi Engine performed thousands of LLM-mediated reasoning, measurement and revision actions during its investigations (e.g., 3,242 LLM calls, 1,242 tool calls).
Operational logs and activity counts reported in the paper: 145.9 million tokens, 3,242 LLM calls, 1,242 tool calls, 163 research notes, 44 scripts.
high positive End-to-end autonomous scientific discovery on a real optical... scale of automated research activity (counts of LLM calls, tool calls, notes, sc...
Qiushi Engine combines nonlinear research phases, Meta-Trace memory and a dual-layer architecture to maintain adaptive and stable research trajectories across long-horizon investigations.
System architecture and methods section describing nonlinear research phases, Meta-Trace memory, and dual-layer architecture; demonstrated operation across long-horizon tasks in experiments (thousands of LLM and tool calls).
high positive End-to-end autonomous scientific discovery on a real optical... ability to maintain adaptive and stable research trajectories over long-horizon ...
The AI-discovered optical bilinear mechanism suggests a route towards high-speed, energy-efficient optical hardware for pairwise computation.
Interpretive claim based on the structural analogy between the discovered optical bilinear interaction and Transformer attention; conceptual argument provided in the paper rather than measured hardware speed or energy benchmarks.
high positive End-to-end autonomous scientific discovery on a real optical... potential for high-speed, energy-efficient optical hardware (conceptual implicat...
In an open-ended study (145.9 million tokens, 3,242 LLM calls, 1,242 tool calls, 163 research notes and 44 scripts), Qiushi Engine proposes and experimentally validates an optical bilinear interaction, a physical mechanism structurally analogous to a core operation in Transformer attention.
Open-ended experimental study reported in the paper with the listed activity metrics (145.9M tokens, 3,242 LLM calls, etc.); experimental investigation and measurements presented claiming validation of optical bilinear interaction and drawing structural analogy to Transformer attention's pairwise operation.
high positive End-to-end autonomous scientific discovery on a real optical... experimental validation of an optical bilinear interaction mechanism
Qiushi Engine autonomously reproduces a published transmission-matrix experiment on a non-original platform.
Experimental reproduction reported in the paper; description of executing the published transmission-matrix experiment using the Qiushi Engine on a different (non-original) optical platform and presenting measured results comparing to published experiment.
high positive End-to-end autonomous scientific discovery on a real optical... successful reproduction of a published transmission-matrix experiment (experimen...
Qiushi Discovery Engine is an LLM-based agentic system for end-to-end autonomous scientific discovery on a real optical platform.
Description and implementation of the Qiushi Engine combining LLM-based agentic control with an optical experimental platform; system design and end-to-end experiments reported in the paper (no randomized trial; system demonstration).
high positive End-to-end autonomous scientific discovery on a real optical... existence and operation of an end-to-end autonomous LLM-driven discovery system ...
The paper formalizes these limitations, addresses four alternative views, and proposes a co-existence solution plus a call to action for system builders, benchmark designers, and the memory community.
Meta-claim about the paper's content: formalization, rebuttals, and recommendations stated in the abstract; no empirical sample reported in abstract.
high positive Contextual Agentic Memory is a Memo, Not True Memory proposed research and design agenda (co-existence of lookup and weight-based mem...
Complementary Learning Systems (CLS) theory shows biological intelligence solved this problem by pairing fast hippocampal exemplar storage with slow neocortical weight consolidation.
Appeal to established neuroscience theory (CLS); the paper draws on CLS literature to justify the two-system solution in biology; no new empirical sample reported in abstract.
high positive Contextual Agentic Memory is a Memo, Not True Memory memory architecture in biological intelligence (hippocampus + neocortex)
Models across all three families acquire interpretable mechanical reasoning strategies without fine-tuning.
Observation reported for the three open-source models used in experiments (Llama 3.3 70B, Qwen3 4B, Qwen3 MoE 30B-A3B) showing emergent, interpretable mechanical reasoning during the iterative design process without any model fine-tuning.
high positive Language Models Refine Mechanical Linkage Designs Through Sy... acquisition of interpretable mechanical reasoning strategies
The system correctly diagnoses underconstraint failure modes 35.6% of the time.
Reported diagnostic accuracy for underconstraint failure mode in the experimental results (35.6%).
high positive Language Models Refine Mechanical Linkage Designs Through Sy... accuracy in diagnosing underconstraint failure mode
The system correctly diagnoses overconstraint failure modes 56.3% of the time.
Reported diagnostic accuracy for overconstraint failure mode in the experimental results (56.3%).
high positive Language Models Refine Mechanical Linkage Designs Through Sy... accuracy in diagnosing overconstraint failure mode
78.6% of iterative refinement trajectories show measurable improvement.
Reported aggregate statistic from the experimental evaluation of iterative refinement trajectories (percentage improvement across trajectories).
high positive Language Models Refine Mechanical Linkage Designs Through Sy... presence of measurable improvement across iterative refinement trajectories
The modular architecture improves structural validity by up to 134% over monolithic baselines.
Empirical results reported across six motion targets and three models comparing modular architecture to monolithic baselines; the paper reports an improvement in structural validity up to 134%.
high positive Language Models Refine Mechanical Linkage Designs Through Sy... structural validity of linkage designs
The modular architecture reduces geometric error by up to 68% over monolithic baselines.
Empirical results reported across six engineering-relevant motion targets and three open-source models comparing the modular architecture to monolithic baselines; the paper states a maximum reduction of geometric error of 68%.
Language models can systematically improve linkage designs through symbolic representations.
Reported experiments using a modular architecture combining language-model agents and numerical optimisers across six engineering-relevant motion targets and three open-source models (Llama 3.3 70B, Qwen3 4B, Qwen3 MoE 30B-A3B); comparisons reported versus monolithic baselines.
high positive Language Models Refine Mechanical Linkage Designs Through Sy... quality of linkage designs (geometric error, structural validity)
Scalable synthetic computer creation, together with at-scale simulations, is highly promising as a foundational substrate for agent self-improvement and agentic reinforcement learning in long-horizon productivity scenarios.
Authors' conclusion/argument based on the methods and preliminary experimental results presented in the paper (interpretive claim rather than a quantified empirical result).
high positive Synthetic Computers at Scale for Long-Horizon Productivity S... suitability as a substrate for agent self-improvement and agentic RL
Given that personas are abundant at billion scale, this methodology can in principle scale to millions or even billions of synthetic user worlds with sufficient compute, enabling broader coverage of diverse professions, roles, contexts, environments, and productivity needs.
Argumentative/theoretical scalability claim based on the abundance of personas and the scalable design of the methodology (no empirical demonstration at millions/billions scale reported).
high positive Synthetic Computers at Scale for Long-Horizon Productivity S... scalability potential (number of synthetic user worlds producible)
Each run requires over 8 hours of agent runtime and spans more than 2,000 turns on average.
Reported runtime and turn-count metrics from the preliminary experiments (per-run runtime >8 hours; per-run average >2,000 turns).
high positive Synthetic Computers at Scale for Long-Horizon Productivity S... agent runtime per simulation run; number of turns per run
In preliminary experiments, we create 1,000 synthetic computers and run long-horizon simulations on them.
Reported preliminary experiment count in the paper (explicit statement: 1,000 synthetic computers were created and simulated).
high positive Synthetic Computers at Scale for Long-Horizon Productivity S... number of synthetic computers created and simulated
Conditioned on each synthetic computer, we run long-horizon simulations: one agent creates productivity objectives that are specific to the computer's user and require multiple professional deliverables and about a month of human work; another agent then acts as that user and keeps working across the computer ... until these objectives are completed.
Description of the two-agent simulation procedure in the paper (simulation design: objective-creating agent and user-acting agent executing tasks across the synthetic computer).
high positive Synthetic Computers at Scale for Long-Horizon Productivity S... ability to simulate long-horizon, user-conditioned productivity workflows
We introduce Synthetic Computers at Scale, a scalable methodology for creating such environments with realistic folder hierarchies and content-rich artifacts (e.g., documents, spreadsheets, and presentations).
Methodological description and implementation presented in the paper (design and procedures for generating synthetic computers and artifact types).
high positive Synthetic Computers at Scale for Long-Horizon Productivity S... creation of synthetic computer environments with realistic folder hierarchies an...