The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests About 🎲 Workforce Futures
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Evidence (8807 claims)

Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.

The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).

Browse by theme

Nine broad, paper-level topics. Click one to filter the claims below.

Adoption
9875 claims
Filter claims →
Productivity
8807 claims
Filtered →
Governance
7870 claims
Filter claims →
Human-AI Collaboration
7560 claims
Filter claims →
Org Design
4892 claims
Filter claims →
Innovation
4781 claims
Filter claims →
Labor Markets
4004 claims
Filter claims →
Skills & Training
3308 claims
Filter claims →
Inequality
2332 claims
Filter claims →

Claims by outcome category

Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.

Outcome Positive Negative Mixed Null Total
Other 870 233 116 1066 2363
Governance & Regulation 976 451 218 133 1809
Organizational Efficiency 949 224 144 88 1416
Technology Adoption Rate 764 287 141 122 1325
Research Productivity 501 152 74 362 1101
Output Quality 542 216 69 69 896
Decision Quality 387 198 94 54 740
Firm Productivity 513 67 101 27 714
AI Safety & Ethics 249 303 73 36 667
Market Structure 190 192 134 27 548
Task Allocation 243 77 91 36 452
Innovation Output 291 33 55 20 401
Skill Acquisition 206 72 65 21 364
Employment Level 133 63 115 22 335
Fiscal & Macroeconomic 153 79 52 32 323
Task Completion Time 206 37 12 15 272
Firm Revenue 179 52 29 5 266
Consumer Welfare 130 76 47 13 266
Inequality Measures 48 137 51 6 242
Worker Satisfaction 101 81 25 13 220
Error Rate 84 110 11 5 210
Wages & Compensation 98 47 30 10 185
Regulatory Compliance 88 73 17 7 185
Automation Exposure 66 64 33 16 182
Team Performance 105 29 30 11 176
Training Effectiveness 109 22 14 21 168
Developer Productivity 114 21 14 8 158
Job Displacement 12 90 24 1 127
Hiring & Recruitment 57 9 9 5 80
Skill Obsolescence 6 56 9 1 72
Social Protection 43 17 8 2 70
Creative Output 35 21 9 4 70
Labor Share of Income 18 21 17 1 57
Worker Turnover 15 16 4 35
Industry 1 1
Clear
Productivity Remove filter
FacProcessTwin was evaluated in a real-world case study of an Australian food manufacturer covering 16 production process flows spanning chilled, frozen, and aseptic shelf-stable product categories and including process variations within the same product.
Case study description in paper; sample explicitly stated as 16 production process flows.
high neutral FacProcessTwin: An LLM-Based System for Process Twin Develop... scope and diversity of evaluation (number and types of process flows covered)
We run a pre-specified 2*4 factorial experiment with 280 complete research runs across four datasets.
Experimental design and sample size reported in paper: 2x4 factorial experiment, 280 complete research runs, four datasets.
high neutral (Human) Attention Is (Still) All You Need: Human oversight m... number of complete research runs (experimental sample)
The ABC-Bench tasks require a combination of biology and software expertise.
Authors' description of task design highlighting that tasks (robot control code, DNA design, screening evasion) combine biological domain knowledge with software/coding skills.
high neutral ABC-Bench: An Agentic Bio-Capabilities Benchmark for Biosecu... skill mix required for benchmark tasks
ABC-Bench evaluates LLM agents on both benign and dual-use biology tasks, including: writing code to operate liquid handling robots, designing DNA fragments for in vitro assembly, and evading DNA synthesis screening.
Description of benchmark tasks provided in the paper; task list enumerated by the authors as components of ABC-Bench.
high neutral ABC-Bench: An Agentic Bio-Capabilities Benchmark for Biosecu... types of tasks included in ABC-Bench
We introduce the Agentic Bio-Capabilities Benchmark (ABC-Bench), a suite of tasks to measure agentic biosecurity-relevant capabilities.
The paper describes the design and composition of ABC-Bench and presents evaluation results using it; this is a methodological contribution asserted by the authors.
high neutral ABC-Bench: An Agentic Bio-Capabilities Benchmark for Biosecu... existence and composition of the ABC-Bench benchmark
We discuss design tradeoffs, failure modes, and lessons learned from operating autonomous AI agents at scale.
Paper statement indicating inclusion of discussion sections on tradeoffs, failure modes, and operational lessons; descriptive/meta claim about paper content.
high neutral Autonomous Incident Resolution at Hyperscale: An Agentic AI ... discussion of design tradeoffs, failure modes, and lessons learned
We ran a controlled three-arm ablation on a production valuation agent: A = plain web-only LLM analyst; B = adds public structured tools + a 14-dimension valuation playbook, verifier, objectivity policy and red-team; C = adds the proprietary Noah AI corpus of curated pipeline, trial and deal intelligence.
Description of experimental arms and setup used in the study (methodological statement).
high neutral AI Scientists Are Only as Good as Their Evidence: A Stratifi... experimental treatment definitions (method)
ALE is organized around a task taxonomy with 55 subfields grouped into 13 industry clusters covering 1K+ tasks.
Author-provided counts describing the benchmark taxonomy and task pool.
high neutral Agents' Last Exam taxonomy breadth (subfields, clusters, number of tasks)
ALE covers non-physical industries defined with reference to O*NET / SOC 2018 (the U.S. federal occupational taxonomy).
Design specification described in the paper referencing O*NET / SOC 2018.
high neutral Agents' Last Exam scope of industries covered by the benchmark
We evaluated seven models (including Gemini, Claude, and GPT families) by comparing their zero-shot estimates against self-reported skill ratings from 27 participants.
Method description: evaluation of seven LLMs comparing zero-shot model estimates to self-reported skill ratings; 27 participants provided self-reports.
high neutral Can AI Guess What You Know? Performance Comparison of Large ... comparison between model zero-shot skill estimates and self-reported skill ratin...
At inference time, BRANE selects the configuration that maximizes predicted correctness penalized by cost, exposing a tunable cost-quality tradeoff without retraining.
Method description and algorithmic claim in the paper (selection rule maximizing predicted correctness with cost penalty). No empirical sample size required for algorithmic description.
high neutral Natural Language Query to Configuration for Retrieval Agents cost-quality tradeoff exposed by selection strategy
We propose BRANE, which uses an LLM to convert each query into workload-specific characteristics, then trains a lightweight per-configuration predictor that estimates whether the pipeline will answer the query correctly.
Method description in the paper: BRANE architecture and training procedure (LLM-based feature extraction + per-configuration correctness predictor). No numeric sample size reported for method description.
high neutral Natural Language Query to Configuration for Retrieval Agents method (feature extraction and predictor training)
AI deployment should be evaluated not only by average task speed, but by its overall effects on congestion, rework, and the robustness of human oversight under load.
Policy/recommendation based on the paper's theoretical results and derived implications from the queueing model (conceptual/prescriptive conclusion; no empirical testing reported).
high neutral Queue & AI: When Faster Tasks Slow Down the Workflow organizational_efficiency
The divergence between mean task speed and system-level delay caused by AI assistance is labeled the 'variance wedge'.
Definition/terminology introduced in the paper as part of its conceptual framing; supported by the analytic model description.
high neutral Queue & AI: When Faster Tasks Slow Down the Workflow task_completion_time
The study used standard scientific methods, employing a comparative approach and inductive and deductive methods to identify patterns of interaction between legal regulation and technological development.
Methodology section of the paper explicitly states the use of comparative, inductive and deductive methods and theoretical synthesis.
high neutral ECONOMIC SYSTEMS IN THE CONTEXT OF DIGITALISATION AND AI: TH... methodological approach used in the study
The paper develops a theoretical and legal model that treats law as an integral part of the economic system influencing income distribution, labour relations, market structure and productivity dynamics.
Model construction through synthesis of theoretical perspectives using inductive and deductive methods and comparative legal analysis (methodology described in the paper).
high neutral ECONOMIC SYSTEMS IN THE CONTEXT OF DIGITALISATION AND AI: TH... role of legal frameworks in shaping economic institutional conditions (income di...
The paper provides a taxonomy of minimum input artifacts for agentic software, firmware, and hardware work; a conversation-to-contract gate; risk-adaptive workflows; and an evidence-bundle acceptance model for agent-generated artifacts.
Declared contributions in the paper (deliverables/artefacts produced by the research; no empirical validation provided in the abstract).
high neutral Agentic Agile-V: From Vibe Coding to Verified Engineering in... availability of process artifacts and workflow models for agentic engineering
The central problem for agentic engineering is no longer prompt engineering; it is engineering process control.
Argument and synthesis presented by the paper (conceptual claim based on reviewed evidence).
high neutral Agentic Agile-V: From Vibe Coding to Verified Engineering in... primary bottleneck affecting agentic engineering effectiveness (process control ...
We performed a large-scale evaluation spanning 15,000 messages with cross-model validation across six LLMs from three families (OpenAI, Anthropic, Google), totaling 1,440 queries.
Study design and reported sample sizes and model counts provided in the paper.
high neutral Episodic-Semantic Memory Architecture for Long-Horizon Scien... evaluation sample size and cross-model coverage
Experiments are run with and without access to Causely under two scenarios: an active incident and a healthy baseline.
Methodological description in the paper describing the two experimental conditions (with/without Causely) and two scenarios (active incident, healthy baseline).
high neutral Causely: A Causal Intelligence Layer for Enterprise AI A Ben... experimental condition (Causely vs. no Causely) across two scenarios
Experiments compare four agent configurations (Claude Code, OpenAI Codex, HolmesGPT with Sonnet and Gemini backends).
Methodological description listing the four agent configurations used in experiments.
high neutral Causely: A Causal Intelligence Layer for Enterprise AI A Ben... agent configuration comparisons
We evaluate this value proposition through a benchmark study conducted in a controlled setting with injected faults in a 24-microservice OpenTelemetry demo application.
Methodological description in the paper specifying a controlled benchmark with an OpenTelemetry demo application composed of 24 microservices.
high neutral Causely: A Causal Intelligence Layer for Enterprise AI A Ben... benchmark evaluation setup (24-microservice demo with injected faults)
The framework reframes the central question of autonomous software engineering from whether a foundation model can produce a patch to whether the model-harness-environment system can produce a verifiably correct, attributed, and maintainable change.
Conceptual reframing and argument presented in the abstract as a conclusion of the proposed framework and evaluation approach.
high neutral AI Harness Engineering: A Runtime Substrate for Foundation-M... ability of the overall system (model+harness+environment) to produce verifiably ...
We formalize this substrate as 'AI Harness Engineering' and identify eleven component responsibilities: task specification, context selection, tool access, project memory, task state, observability, failure attribution, verification, permissions, entropy auditing, and intervention recording.
Methodological/conceptual contribution described in the paper (abstract) that lists eleven component responsibilities as part of the formalization.
high neutral AI Harness Engineering: A Runtime Substrate for Foundation-M... completeness and scope of responsibilities required for a runtime harness
We construct an evaluation framework on five function-calling benchmarks and train a DistilBERT-based classifier, deployed under a latency budget.
Methods / experimental setup reported in the paper: five function-calling benchmarks and a DistilBERT classifier trained and deployed under latency constraints.
high neutral Switchcraft: AI Model Router for Agentic Tool Calling evaluation framework and classifier training/deployment
We evaluate 4 popular agent harnesses and 7 foundation models on Workspace-Bench.
Experimental setup reported in the paper listing 4 agent harnesses and 7 foundation models used in evaluations.
high neutral Workspace-Bench 1.0: Benchmarking AI Agents on Workspace Tas... number of agent harnesses and foundation models evaluated
The same curated index is exposed as a Gosset MCP server that any frontier model can call as a tool.
System description in paper noting that the curated index is available via a Gosset MCP server for external models to call.
high neutral Curated AI beats frontier LLMs at pharma asset discovery availability of curated index as callable MCP server
All five systems receive the same natural-language query and the same JSON output schema.
Methodological detail reported in paper describing controlled inputs across systems.
high neutral Curated AI beats frontier LLMs at pharma asset discovery consistency of input/query and output schema across systems
We benchmark Gosset ... against four frontier systems with web access (Claude Opus 4.7, GPT 5.5, Gemini 3.1 Pro, Perplexity sonar-pro) on ten niche oncology/immunology targets.
Experimental benchmark described in paper: direct comparison of Gosset versus four named models on 10 targets; methodological statement.
high neutral Curated AI beats frontier LLMs at pharma asset discovery comparative retrieval performance on 10 niche oncology/immunology targets
Weight-based memory generalizes by applying abstract rules to inputs never seen before.
Conceptual claim grounded in the paper's theoretical distinction between weight-based learning and retrieval; references Complementary Learning Systems theory; no empirical sample in abstract.
high neutral Contextual Agentic Memory is a Memo, Not True Memory type of generalization performed by weight-based memory
Retrieval generalizes by similarity to stored cases.
Conceptual claim stated in paper (distinction between retrieval-based and weight-based generalization); supported by theoretical characterization, not empirical data in abstract.
high neutral Contextual Agentic Memory is a Memo, Not True Memory type of generalization performed by retrieval systems
The process of synthesizing information is inherently iterative: users explore content, identify relationships between concepts, and continuously reorganize their mental models.
Conceptual description of the cognitive/process characteristics in the paper's background/motivation (no empirical measurement reported).
high neutral MindTrellis: Co-Creating Knowledge Structures with AI throug... iterative nature of knowledge synthesis (exploration, relation identification, r...
The paper establishes a taxonomy of forgetting mechanisms: passive decay-based, active deletion-based, safety-triggered, and adaptive reinforcement-based.
Explicit taxonomy presented in paper (listed in abstract).
high neutral FSFM: A Biologically-Inspired Framework for Selective Forget... classification of forgetting mechanisms
We evaluate Aether over synthetic network change scenarios covering main classes of network changes and on past incidents from a major ISP operational network.
Evaluation methodology stated in paper abstract: tested on synthetic scenarios and historical incidents from one major ISP (no numeric sample size provided in abstract).
high neutral Aether: Network Validation Using Agentic AI and Digital Twin evaluation dataset composition (synthetic scenarios + past ISP incidents)
Generally speaking, these systems place an agent in a feedback loop in which it can write code, compile that code to an assembly of CAD model(s), visualize the model, and then iteratively refine its code based on visual and other feedback.
Descriptive claim about the general architecture of Agent-Aided Design systems as asserted by the authors (methodological description), not an empirical test; no quantitative evaluation provided here.
high neutral Agent-Aided Design for Dynamic CAD Models system architecture / iterative design loop (agent writes code, compiles, visual...
The paper provides lessons for scaling regression automation and enabling effective human-AI teaming in Agile settings.
Stated contribution of the paper (synthesis of lessons from the industrial case study).
high neutral Human-AI Collaboration for Scaling Agile Regression Testing:... availability of lessons and guidance
The Copilot was integrated with Hacon's CI pipelines and operates asynchronously as a 'silent AI teammate', producing candidate scripts for human review.
System integration and deployment description within the case study (implementation detail reported in the paper).
high neutral Human-AI Collaboration for Scaling Agile Regression Testing:... operational mode and integration with CI (asynchronous candidate generation for ...
We conducted an exploratory industrial case study of the Hacon Test Automation Copilot, an agentic AI system that generates system-level regression test scripts from validated specifications using retrieval-augmented generation and a multi-agent workflow.
Methodological claim: description of the study design and the system; the paper reports a single industrial case study at Hacon (a Siemens company).
high neutral Human-AI Collaboration for Scaling Agile Regression Testing:... capability to generate system-level regression test scripts
A variance decomposition indicates that most expert disagreement about long-run macroeconomic outcomes is driven by differing beliefs about the economic effects of highly capable AI, rather than disagreement about the pace of AI capability progress.
Authors' variance-decomposition analysis of survey responses separating components due to beliefs about AI capabilities vs. beliefs about economic effects given capabilities (methodological details referenced but not provided in excerpt).
high neutral Forecasting the Economic Effects of AI sources of expert disagreement (capabilities vs. economic effects)
A life insurance system integrated into an industry partner mobile app was tested in two experiments.
Paper reports two experiments running the ARQuest-enabled life insurance system inside a partner mobile app; experimental setup is stated though sample sizes are not provided in the excerpt.
high neutral AI in Insurance: Adaptive Questionnaires for Improved Risk P... experimental evaluation of system in partner app
The paper addresses three institutional audiences: enterprise finance and operations teams; government and regulatory bodies developing AI labor displacement frameworks; and financial markets requiring a machine labor index as a long-duration economic signal.
Stated intended audiences in the paper (descriptive statement).
high neutral HEWU: A Standardized Framework for Measuring Machine-Generat... intended institutional audiences
BCR is a minimalist, single-stage training paradigm that trains the model to solve N problems simultaneously within a shared context window, rewarded purely by per-instance accuracy.
Methodological description presented in the paper describing the training procedure and objective (single-stage, per-instance accuracy reward, N-problem batching in shared context).
high neutral Batched Contextual Reinforcement: A Task-Scaling Law for Eff... training paradigm characteristics (simplicity, stage count, reward structure)
The framework is calibrated with O*NET task data, a survey of 3,778 domain experts, and GPT-4o-derived task decompositions, and implemented in computer vision.
Calibration and empirical implementation using O*NET, a domain expert survey (n=3,778), and GPT-4o task decompositions; applied to computer vision tasks.
high neutral Economics of Human and AI Collaboration: When is Partial Aut... validity of calibration / empirical grounding of the framework
We introduce an entropy-based measure of task complexity that maps model accuracy into a labor substitution ratio, quantifying human labor displacement at each accuracy level.
New metric proposed in the paper (entropy-based task complexity) and mapping procedure from accuracy to substitution ratio; implemented in the framework.
high neutral Economics of Human and AI Collaboration: When is Partial Aut... labor substitution ratio (human labor displaced per unit accuracy)
Costinot and Werning (2023) develop a sufficient-statistic approach and find optimal technology taxes of 1–3.7% on robots.
Citation reported in the paper summarizing Costinot and Werning (2023)'s quantitative sufficient-statistic estimate.
high neutral NBER WORKING PAPER SERIES optimal robot tax rate
Guerreiro et al. (2022) characterize optimal Mirrleesian tax system with automation and find that robot taxes should be transitional—high when incumbent workers cannot retrain, converging to zero as new cohorts adjust skill investments.
Citation reported in the paper summarizing Guerreiro et al. (2022)'s theoretical result on transitional robot taxes.
high neutral NBER WORKING PAPER SERIES optimal robot tax path over time
If labor becomes economically redundant, the policy focus shifts from steering innovation to redesigning public finance and redistribution (e.g., new tax instruments, redistribution mechanisms).
Theoretical scenario analysis in the paper with references to related works (Korinek and Juelfs 2024; Korinek and Lockwood 2026).
high neutral NBER WORKING PAPER SERIES policy priority shift (steering -> public finance/redistribution)
Evaluation is carried out under three frozen context configurations (diff only: config_A; diff with file content: config_B; full context: config_C) enabling systematic ablation of context provision strategies.
Methodological description: three fixed context configurations defined and used for ablation experiments.
high neutral SWE-PRBench: Benchmarking AI Code Review Quality Against Pul... effect of context-provision design on model performance
Traffic performance is evaluated using the Fundamental Diagram (FD) under varying driver heterogeneity, heterogeneous time-gap penetration levels, and different shares of RL-controlled vehicles.
Description of experimental/evaluation setup in the paper: macroscopic evaluation via Fundamental Diagram across varied scenario parameters. No numeric sample size provided in the claim text.
high neutral Macroscopic Characteristics of Mixed Traffic Flow with Deep ... traffic performance (via Fundamental Diagram) under varied heterogeneity and RL ...
CriQ is a sister app to Dream11, India's largest fantasy sports platform with over 250 million users.
Descriptive statement in the paper providing context about the application domain and user base.