The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Evidence (2025 claims)

Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.

The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).

Browse by theme

Nine broad, paper-level topics. Click one to filter the claims below.

Adoption
20058 claims
Filter claims →
Productivity
17184 claims
Filter claims →
Governance
16099 claims
Filter claims →
Human-AI Collaboration
16034 claims
Filter claims →
Innovation
10501 claims
Filter claims →
Org Design
10496 claims
Filter claims →
Labor Markets
6444 claims
Filter claims →
Skills & Training
5385 claims
Filter claims →
Inequality
4148 claims
Filter claims →

Claims by outcome category

Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.

Outcome Positive Negative Mixed Null Total
Other 1820 479 278 1820 4588
Organizational Efficiency 2711 616 401 173 3922
Governance & Regulation 2075 886 459 246 3714
Technology Adoption Rate 1467 530 258 206 2488
Decision Quality 1281 496 289 152 2228
Output Quality 1227 447 207 138 2025
AI Safety & Ethics 634 754 207 83 1688
Research Productivity 826 241 114 422 1624
Firm Productivity 1052 154 163 66 1441
Task Allocation 685 211 331 99 1335
Market Structure 433 423 242 46 1150
Innovation Output 639 91 105 34 871
Task Completion Time 476 113 43 36 672
Firm Revenue 445 126 58 25 656
Skill Acquisition 364 119 109 34 626
Consumer Welfare 288 167 104 31 592
Employment Level 214 140 174 50 582
Error Rate 230 251 35 16 535
Fiscal & Macroeconomic 268 136 71 50 532
Inequality Measures 100 307 96 12 515
Worker Satisfaction 221 173 60 30 484
Automation Exposure 155 138 65 36 398
Regulatory Compliance 171 120 30 13 335
Developer Productivity 222 58 27 13 321
Team Performance 188 56 50 24 320
Wages & Compensation 146 104 46 16 312
Training Effectiveness 207 41 21 26 298
Job Displacement 23 153 52 4 232
Hiring & Recruitment 102 57 30 11 202
Skill Obsolescence 16 102 24 6 148
Creative Output 71 42 23 6 143
Social Protection 57 30 11 3 101
Labor Share of Income 29 42 24 2 97
Worker Turnover 43 29 6 4 82
Industry 1 1
Visibility-enabled metrics such as typing activity, response times, and screen time may improve measurement precision while also encouraging gaming and performative labor that inflate measured productivity without corresponding improvements in output quality.
Conceptual implication for AI economics and productivity measurement; the paper does not report a causal test or quantify the gap between measured productivity and output quality.
high mixed The Digital Visibility‐Belonging Paradox: Digital Cohesion a... Measured productivity and substantive output quality
Digital transformation has stronger positive associations with substantive carbon disclosure components—carbon-related business practices, carbon governance, and carbon performance—than with the disclosure carrier or reporting format/channel.
Component-level regressions separating substantive content elements from the disclosure-carrier component of the multidimensional CIDQ index.
high mixed Environmental Transparency Through Digital Transformation: E... Component-level carbon information disclosure quality
In the observed Verified comparisons, the treatment-control performance gap was largest at the 20,480-token window and closest to zero at the 262,144-token window.
Three successive Verified comparisons using the same 169 task IDs and fixed 480-second budget; the paper explicitly cautions that the cross-window ordering does not isolate window size from run-era change.
high mixed Same Model, Different Harness: Different Coding-Agent Result... Treatment-control gap in mean per-task F2PF and complete solutions across contex...
At 128,000 tokens, the primary model retrieved the disclosure for all 12 firms and produced no false retrievals on neutral filings, despite the disclosure having negligible influence on its investment judgment.
Separate retrieval calls were scored against frozen answer keys, with retrieval tested independently from the investment judgment experiment.
high mixed Reading Is Not Using: Retrieval, Judgment, and the Design of... Disclosure retrieval accuracy and downstream decision influence
For the U.S. FEMA benchmark, PPE outperformed expert baselines on the Socioeconomic and Composite risk indicators, achieving R² of 66.9% versus 61.1%, while performing on par with expert benchmarks across the broader environmental-target suite.
Spatial regression evaluation of 21 FEMA environmental risk scores using approximately 84,000 census-tract observations.
high mixed Planetary Prediction Engine: Autonomous Geospatial Predictio... R² for FEMA Socioeconomic, Composite, and broader environmental risk indicators
In the in-distribution cat -n regime, Paritok-4B compressed context to 27.8% of its original size and retained 89.3% of uncompressed solve quality.
Evaluation on the same 300 SWE-bench Lite instances using line-numbered cat -n inputs intended to match real agent trajectories.
high mixed Paritok-4B: Intent-Conditioned Context Compression for Codin... Context compression ratio and retained solve quality under line-numbered agent i...
Across all 300 SWE-bench Lite instances, Paritok-4B compressed agent context to 25.7% of its original size while retaining 86.5% of uncompressed single-shot solve quality.
End-to-end evaluation on all 300 SWE-bench Lite instances comparing compressed and uncompressed context under the paper's single-shot solve harness.
high mixed Paritok-4B: Intent-Conditioned Context Compression for Codin... Context compression ratio and single-shot software issue solve quality
Expert-prompted LLMs achieved the highest accuracy but were more expensive and less consistent than the deterministic PPO policy.
Ablation comparing minimal versus augmented expert prompts and PPO on identical held-out proteins and repeated rollouts; token use, inference cost, and consistency were recorded.
high mixed Federation Is Nearly Free, Reasoning Is Not: Tradeoffs for A... Prediction accuracy, inference cost, and cross-rollout consistency
Across 20 BixBench3 tasks containing 138 unique graded artifacts, 13 frontier LLMs achieved overall scores ranging from 0.00 to 0.48.
Benchmark evaluation of 13 models across 20 research-study-scale computational-biology tasks and 1,794 model-artifact evaluations; scores were based on the proportion of artifacts reproduced closely enough to preserve the original biological interpretation.
high mixed BixBench3: Benchmarking AI agents on research-study-scale co... Overall artifact-reproduction performance on computational-biology research task...
In GAMEOPT, leading agents often retain requested functionality across six turns, but they do not consistently preserve the core game loop or achieve balanced improvement across gameplay, level design, balance, art, interface, and audio.
The finding comes from 17 six-turn optimization chains totaling 102 requests, evaluated using final-product criteria and regression checks.
high mixed GameXpert-Bench: How Far Are Coding Agents from Expert Game ... Functionality preservation and quality improvement during multi-turn game optimi...
In GAMEGEN, agents establish a playable core more reliably than they deliver rich content, robust interfaces, and fully integrated runtime behavior.
The paper summarizes findings from the GAMEGEN evaluation, which combines behavioral rubrics, runtime verification, code inspection, live interaction, and human assessment across 97 generation tasks.
high mixed GameXpert-Bench: How Far Are Coding Agents from Expert Game ... Completeness, richness, interface robustness, and runtime integration of generat...
Current coding agents are more reliable at producing playable foundations and implementing explicit requirements than at discovering defects, verifying runtime behavior, and preserving functionality across changes.
This conclusion is drawn from evaluations across GAMEGEN, GAMEFIX, and GAMEOPT, although the supplied text does not report aggregate numerical scores.
high mixed GameXpert-Bench: How Far Are Coding Agents from Expert Game ... Relative reliability of coding agents across game-generation, repair, and optimi...
Models systematically underrepresent race and religion, which are the strongest real drivers of opinion, while exaggerating the influence of weaker dimensions such as gender and party.
The study compared simulated and human steering magnitudes across 429 dimension-value-by-wave observations.
high mixed Large language models simulate intersectional synthetic iden... Magnitude of demographic effects on simulated opinion distributions
The study concludes that geospatial transferability depends not only on model design but also on intrinsic geographic differences between source and target regions.
Interpretation of the transfer experiments and regression analysis relating geographic domain shift to transfer performance.
high mixed Quantifying geographic domain shift to decouple the geospati... Geospatial transferability of human mobility generation models
Cross-region transfer performance exhibits substantial spatial heterogeneity and asymmetry across source-target geographic pairs.
Results summarized in the abstract and contribution statement, based on cross-region transfer evaluation and linear mixed-effects regression.
high mixed Quantifying geographic domain shift to decouple the geospati... Geospatial transfer performance of human mobility generation models
Geospatial transferability varies substantially across target geographic areas.
Cross-region transfer experiments using human mobility generation models trained in one geographic region and evaluated in unseen regions.
high mixed Quantifying geographic domain shift to decouple the geospati... Model transfer performance across geographic target areas
In one ablation study, removing the language-model component from several LLM-based forecasters left accuracy unchanged or improved it.
A component ablation study cited by the review, comparing complete LLM-based forecasters with versions lacking the language-model component.
Forecasts that rely on unstructured evidence such as news, filings, or health reports may fail to improve accuracy or calibration over specialized models while requiring substantially more inference compute.
Synthesis of reviewed forecasting-agent studies examining retrieval of unstructured evidence, predictive accuracy, calibration, and computational requirements.
high mixed LLM-based Agents for Forecasting and Prediction: Methods, Tr... Forecast accuracy, forecast calibration, and inference compute
LLM-judge scores varied by approximately 1.5 points on a 0–10 scale across three judge models, although relative ranking within the corpus was largely preserved.
Sensitivity analysis rotating Claude 3.7 Sonnet, Claude 3.5 Haiku, and a 9B in-house judge across the 145-skill corpus.
high mixed Evaluating Skills, Not Just Agents: Agentic Continuous Evalu... Absolute LLM-judge scores and relative ranking of skills
The default structural gate is permissive: 94.5% of skills cleared the 70-point threshold, but only 48.9% reached 80 points.
Structural scoring of 145 real skills using approximately 50 deterministic rules.
high mixed Evaluating Skills, Not Just Agents: Agentic Continuous Evalu... Distribution of structural skill-quality scores
A particular generated sequence is a sample selected under a decoding policy from a learned conditional distribution, so observed output need not represent the model's entire learned distribution.
Technical explanation of greedy decoding, temperature sampling, nucleus/top-p sampling, and the distinction between distributions and samples.
high mixed Six misconceptions about large language models: A minimal mo... Output diversity and predictability
Current AI and large language models have potential to augment psychological assessment and treatment at scale, but they do not yet produce meaningful, sustained clinical change.
Conceptual synthesis of clinical-science literature, AI/LLM capabilities, and ethical and implementation considerations; the paper does not report a primary clinical trial.
high mixed A framework for evidence-based psychotherapy with AI (EBP-AI... Meaningful, sustained clinical change from AI-augmented psychological assessment...
The effect of perturbations varies substantially across repositories, with degradation concentrated in a small subset of task instances while other instances are unaffected.
Per-instance degradation analysis compares baseline and perturbed resolve rates for the 54 selected SWE-bench task instances.
high mixed A Jagged Frontier: Evaluating Robustness of Code Agents to S... Per-instance change in issue resolve rate after repository perturbation.
Robustness to semantics-preserving perturbations is a joint property of the model, agentic scaffold, and workload rather than a property of the model alone.
The evaluation crosses four models, two scaffolds, and two benchmarks, and observes that scaffold and benchmark changes alter degradation and model rankings.
high mixed A Jagged Frontier: Evaluating Robustness of Code Agents to S... Variation in issue-resolution robustness across model, scaffold, benchmark, and ...
Model robustness rankings are not consistent across agentic scaffolds: Qwen is among the most robust with mini-SWE agent on SWE-bench Verified but is the most brittle with OpenCode.
The paper compares four models across two scaffolds and two benchmarks and reports scaffold-dependent relative robustness rankings.
high mixed A Jagged Frontier: Evaluating Robustness of Code Agents to S... Relative resolve-rate degradation under perturbation across models and scaffolds...
GPT-4.1 displays the reverse pattern on counseling: it is a lower-ranked direct solver but the strongest assistant.
Task-level comparison of GPT-4.1's automation and augmentation rankings on the counseling task.
high mixed CentaurBench: Benchmarking LLM Capabilities on Augmenting vs... Relative rank of GPT-4.1 as a direct counseling-response generator versus an ass...
GPT-5-mini ranks first in both the automation and augmentation comparisons, while the unaided GPT-3.5-Turbo worker has the second-best average rank in the augmentation comparison.
Average and task-level model rankings across the two usage regimes, using a fixed GPT-3.5-Turbo worker in augmentation mode.
high mixed CentaurBench: Benchmarking LLM Capabilities on Augmenting vs... Average rank of models and baseline conditions in automation and augmentation
The model that wins in automation loses in augmentation on five of the seven tasks.
Comparison of the highest-ranked automation and augmentation models at the task level across seven tasks.
high mixed CentaurBench: Benchmarking LLM Capabilities on Augmenting vs... Whether the top automation model is also the top augmentation model on each task
The correlation between automation and augmentation rankings varies substantially across tasks, ranging from -0.04 for travel planning to 0.85 for tax preparation.
Task-level comparison of automation and augmentation rankings across the seven benchmark tasks.
high mixed CentaurBench: Benchmarking LLM Capabilities on Augmenting vs... Task-specific rank correlation between automation and augmentation performance
Model rankings in automation and augmentation are only modestly correlated, with a model-level rank correlation of about 0.48.
Cross-mode rankings of ten LLMs across seven professional tasks, evaluated using blind pairwise comparisons by four LLM judges and replicated across ten independent runs.
high mixed CentaurBench: Benchmarking LLM Capabilities on Augmenting vs... Rank correlation between model performance as autonomous task solvers and as ass...
A high grouping score can coexist with imperfect average accuracy: a model may consistently pass some tasks and consistently fail another, producing a tight but partially off-target performance pattern.
A worked example using a six-task Rust suite with five decisive passes and one decisive failure across five runs per task.
high mixed Grouping the Stochastic Machine: Precision, Not Capability, ... Task pass rate and outcome grouping across repeated engineering tasks
The study could not rule out an accuracy gain of up to 4.67 percentage points from explicit high effort.
The upper endpoint of the registered 95% interval for the Sonnet 5 accuracy contrast was +0.0467.
high mixed The Price of Thinking: Reasoning Effort as a Model-Specific ... Difference in answer accuracy between explicit high effort and omitted effort
Agent performance at the human frontier is concentrated on GoEmotions: 8 of 12 configurations surpass Human SOTA, whereas CUB200 and Weather each have only one frontier-reaching configuration.
Task-level comparison of the 12 agent configurations per task, consisting of 6 agents evaluated with and without prepared human references.
high mixed When AI Designs AI: Innovation or Imitation? Number of agent configurations matching or exceeding human SOTA by task
Current LLM agents reach or surpass human state-of-the-art performance in 10 of 72 experimental configurations, but these successes occur on only 3 of the 6 evaluated tasks.
Evaluation of 6 agents across 6 tasks and 2 reference conditions, yielding 72 configurations; performance was compared with the best collected human result for each task.
high mixed When AI Designs AI: Innovation or Imitation? Whether agent-designed methods reach or exceed the best collected human performa...
ReasoningBank improves benign utility from 0.741 to 0.859, while its aggregate attack success rate decreases from 0.496 to 0.474.
Matched benign and attacked evaluations on the corrected 15-family Banking suite; 135 benign cases and 405 attacked cases per condition across three lineages.
high mixed Auditing Self-Evolution in Financial Agents: Capability Gain... Benign task utility and aggregate prompt-injection attack success
SkillOpt produces both capability gains and regressions: it changes 23 previously wrong instances to correct and 10 previously correct instances to wrong.
Paired transition analysis comparing each evolved instance with the Static run on the identical lineage, family, and variant.
high mixed Auditing Self-Evolution in Financial Agents: Capability Gain... Transitions in benign task correctness
Old upload-earned medals appear more informative in the AI era only when read in isolation; their apparent gains disappear after conditioning on the holder's full credential profile.
Comparisons of badge-count, experience-only, and full-profile regression specifications. Under the badge-count read, slope changes were +0.062 for medals aged 1–2 years and +0.036 for medals aged 2–3 years; the gains vanished under full-profile conditioning and persisted in a balanced panel.
high mixed Stranded credentials: how a skill-signaling market absorbed ... Predictive informativeness of older upload-earned medals
Fresh upload medals lost about one quarter of their predictive slope in the AI era under full-profile conditioning, while fresh code medals did not show a statistically significant decline.
Comparison of pre-AI and AI-era fresh-medal slopes under full-profile conditioning. Fresh upload medals changed by -0.033 from a pre-AI base of 0.122 with P = 0.002; fresh code medals changed by +0.015 with P = 0.12.
high mixed Stranded credentials: how a skill-signaling market absorbed ... Predictive informativeness of fresh upload- and code-earned medals for subsequen...
Current agents generally make substantial partial progress on realistic professional workflows even when they fail to meet the benchmark's completion threshold.
Across the evaluated models, most average task scores fall between roughly 55 and 75, while success rates remain below one third.
high mixed StartupBench: Benchmarking General-Purpose Agents on Market-... Average task score versus strict end-to-end completion rate
In the authors' layer-locality analysis, the gate projection accounts for 77% of pruning distortion in Qwen3.5, while Qwen2.5 concentrates 67% of distortion on the down projection.
Layer-wise neural-flow graph analysis of pruning-induced signal distortion in two Qwen model families.
high mixed Large Models for Small Devices: Recent Advances and Empirica... Share of pruning-induced neural-flow distortion by projection layer
For image segmentation, pruning reduces model size by nearly 80% while maintaining approximately constant mean intersection-over-union, whereas default quantization leaves parameter count and MACs unchanged.
Empirical evaluation of compressed segmentation models using pruning and quantization.
high mixed Large Models for Small Devices: Recent Advances and Empirica... Model size, parameter count, MACs, and segmentation mIoU
The apparent relationship between MMLU and MMLU-Pro scores depends strongly on the chosen scale: for six common systems, the fitted slope is 0.976 under a logit link, 1.043 under a probit link, 1.312 on the linear scale, and 1.984 on the logarithmic scale.
Cross-benchmark linking stress test using six common systems and four score representations.
high mixed Frontier AI Forecasting Has a Measurement Problem: An Audit ... Cross-version or cross-benchmark score comparability
Across eight armed-model requests, the median realized served total variation was 0.118, while the inherited ℓ∞-based upper bound had a median of 0.996 and was effectively vacuous; Bhattacharyya and sub-Gaussian bounds were 0.159 and 0.230, respectively.
Post hoc comparison of paired exact and compressed logits for eight requests, summarized in Figure 4; the authors explicitly state that this certifies the geometry rather than providing an online witness.
high mixed Pricing the Risk of Runtime Compression: Anytime-Valid Admis... Per-request median served total variation and corresponding upper bounds
The strongest fixed protocol varied by task: Broadcast was strongest in nine of the ten settings, while PER exceeded Broadcast by 2.7 points for Gemma on SciBench.
Figure 2 reports matched solve/coverage percentages for four protocols across five benchmark conditions and two solver families.
high mixed LLMs Can Predict Failure Risk, But Struggle to Predict Which... Protocol-specific solve coverage
Exact ground-truth skill invocation is neither sufficient nor necessary for successful task execution.
Joint analysis of skill-use parsing and final verifier outcomes during full-pool real execution found that correct skills did not guarantee success and related non-ground-truth skills could still help.
high mixed Demystifying Agent Skills: Why They Work-Until They Don't Final task success conditional on skill invocation
Offline retrieval or selection accuracy does not determine downstream task success in a one-to-one manner.
The study independently evaluated embedding retrieval, explicit agent selection, and full-pool real execution; outputs from the first two procedures were not passed to the execution experiment.
high mixed Demystifying Agent Skills: Why They Work-Until They Don't Relationship between skill-retrieval accuracy and final task success
Compared with independent professional estimators, Handoff-H1 produced substantially higher material coverage but lower quantity precision: 86.1% versus 65.5% coverage and 78.8% versus 87.9% Precision@25%.
Direct comparison of Handoff-H1 and independent estimator results under the same benchmark judge and gold standard.
high mixed Handoff-H1: An Orchestrated Vision-Agent System for Material... Tradeoff between takeoff completeness and quantity accuracy
Extraction performance exceeded 90% agreement for document titles and sources, but was weaker for author and reference fields.
Model-extracted metadata fields were compared with verified source records and graded as exact, partial, or absent.
high mixed Knowledge Synthesis Review Framework: Task-Level Benchmarkin... Metadata extraction agreement by field
No evaluated LLM-based system led on all four evidence-synthesis tasks.
Comparative evaluation of GPT-5, Claude Sonnet 4, Gemini 2.5 Pro, and NotebookLM across screening, extraction, analysis, and synthesis using a 244-document benchmark and expert reference standards.
high mixed Knowledge Synthesis Review Framework: Task-Level Benchmarkin... Relative task-level performance across evidence-synthesis tasks
Under asymmetric accuracy-improvement costs, an equal-accuracy requirement can substantially lower advantaged-group accuracy while increasing disadvantaged-group accuracy only modestly.
Theoretical firm-design analysis with quadratic development costs and a cost asymmetry in which improving accuracy for the disadvantaged group is more expensive than for the advantaged group.
high mixed Algorithm Design and Physician Liability Group-specific algorithmic accuracy