The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Evidence (672 claims)

Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.

The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).

Browse by theme

Nine broad, paper-level topics. Click one to filter the claims below.

Adoption
20058 claims
Filter claims →
Productivity
17184 claims
Filter claims →
Governance
16099 claims
Filter claims →
Human-AI Collaboration
16034 claims
Filter claims →
Innovation
10501 claims
Filter claims →
Org Design
10496 claims
Filter claims →
Labor Markets
6444 claims
Filter claims →
Skills & Training
5385 claims
Filter claims →
Inequality
4148 claims
Filter claims →

Claims by outcome category

Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.

Outcome Positive Negative Mixed Null Total
Other 1820 479 278 1820 4588
Organizational Efficiency 2711 616 401 173 3922
Governance & Regulation 2075 886 459 246 3714
Technology Adoption Rate 1467 530 258 206 2488
Decision Quality 1281 496 289 152 2228
Output Quality 1227 447 207 138 2025
AI Safety & Ethics 634 754 207 83 1688
Research Productivity 826 241 114 422 1624
Firm Productivity 1052 154 163 66 1441
Task Allocation 685 211 331 99 1335
Market Structure 433 423 242 46 1150
Innovation Output 639 91 105 34 871
Task Completion Time 476 113 43 36 672
Firm Revenue 445 126 58 25 656
Skill Acquisition 364 119 109 34 626
Consumer Welfare 288 167 104 31 592
Employment Level 214 140 174 50 582
Error Rate 230 251 35 16 535
Fiscal & Macroeconomic 268 136 71 50 532
Inequality Measures 100 307 96 12 515
Worker Satisfaction 221 173 60 30 484
Automation Exposure 155 138 65 36 398
Regulatory Compliance 171 120 30 13 335
Developer Productivity 222 58 27 13 321
Team Performance 188 56 50 24 320
Wages & Compensation 146 104 46 16 312
Training Effectiveness 207 41 21 26 298
Job Displacement 23 153 52 4 232
Hiring & Recruitment 102 57 30 11 202
Skill Obsolescence 16 102 24 6 148
Creative Output 71 42 23 6 143
Social Protection 57 30 11 3 101
Labor Share of Income 29 42 24 2 97
Worker Turnover 43 29 6 4 82
Industry 1 1
An industry survey reported approximately 46% time savings from AI on routine software-development tasks, but less than 10% savings on complex work.
McKinsey industry survey of approximately 4,500 developers.
high mixed AI and the Future of Software Engineering: Expertise, Employ... Time savings from AI-assisted development by task complexity
Mean execution time varied substantially across systems, ranging from 21.8 to 112.8 minutes per task.
Reported execution-time measurements for the evaluated model/scaffold configurations.
high mixed FrontierChallenge: Evaluating Scientific Workflow Completion Mean scientific workflow execution time
Validation work can require substantial expertise and time because fluent AI output may contain substantive errors, yet validation is often treated by the market as undifferentiated work that commands a lower price.
The paper supports the claim with the machine-translation case and cited studies on AI-generated content and overreliance [3, 6, 7, 13, 16]. The paper does not report a sample size or quantitative estimate for this synthesis.
high mixed From Producing to Validating: How AI Is Deskilling Freelance... Effort and expertise required to detect and correct AI errors, relative to compe...
Historical productivity gains have not translated mechanically into equivalent increases in leisure or reductions in work time; institutions influence the conversion of productivity into returned human time.
The paper cites estimates of only about four to five additional hours of U.S. leisure per week over roughly a century, unchanged prime-age market work in the cited measure, and cross-country labor-supply differences associated by Prescott with taxation.
high mixed Human Computational Capital and the AI Time Dividend: A Cond... Leisure time, market work time, and institutional conversion of productivity int...
A gross AI time saving does not necessarily become returned or usable human time because institutions may increase workloads and individuals may lack material agency over the time saved.
The paper defines institutional return and material-agency factors and gives a workload-increase example in which gross time savings remain positive while the return fraction can be near zero.
high mixed Human Computational Capital and the AI Time Dividend: A Cond... Conversion of technical AI time savings into usable human time
AI can change the human time required per useful outcome, but the sign and magnitude of the change are task-, technology-, and context-dependent.
The paper cites controlled studies showing both time reductions and time increases in specific tasks, including professional writing, customer support, and software development.
high mixed Human Computational Capital and the AI Time Dividend: A Cond... Human time required per useful task outcome
With fixed sample sizes, asymptotic experiment conversion is governed by the full spectrum of Rényi divergences, whereas allowing the sample size to depend on realized evidence reduces the relevant constraints to the two directed KL divergence ratios.
The paper contrasts the fixed-sample-size result based on the dominance-ratio result of Mu et al. (2021) with its endogenous-stopping characterization.
high mixed The Order of Binary Experiments under Endogenous Stopping Asymptotic number of source observations required to reproduce another experimen...
MST-5 achieved a slightly lower average weighted deviation than PPO, but its setup time was approximately 60% higher.
Table I reports MST-5 with 1,112 weighted deviation and 50.15 setup time, compared with PPO's 1,144 deviation and 31.71 setup time. The paper explicitly characterizes the setup-time difference as approximately 60%.
high mixed Reinforcement Learning-Based Production Scheduling in an Ind... Weighted due-date deviation and setup time
DQN achieved the lowest average weighted deviation from scheduled due dates, but it produced substantially more unfinished products than PPO and most other methods.
Table I reports DQN's average weighted deviation as 940, the lowest value, and its unfinished-product count as 144.8, compared with 53.95 for PPO. The text also states that the lower deviation was offset by a considerable increase in unfinished products.
high mixed Reinforcement Learning-Based Production Scheduling in an Ind... Weighted deviation from due dates and number of unfinished products
In Era 3, the median PR cycle time was 1.04 days in vLLM and 0.62 days in SGLang, while the 90th-percentile cycle times were 16.8 and 14.3 days, respectively.
Cycle time was calculated from PR creation to merge for the full merged-PR populations in both repositories.
high mixed Engineering Signals of Human-AI Collaboration in the Agentic... Time from pull-request creation to merge
Performance improvements from translation varied substantially by software, with highly optimized C/C++ programs requiring considerably more work to reach speed parity.
The authors describe uneven optimization effort and compare translation behavior across software domains and source languages using qualitative benchmark trends.
high mixed Static analysis-guided agentic AI translation enables Rust a... Execution speed relative to the original implementation
Under transparency, the equilibrium algorithm can make prominence convey either good news that deters further search or bad news that encourages further search, depending on the prior probability that the more profitable product is the consumer's best match.
Proposition 2 analytically characterizes three equilibrium algorithm forms across prior beliefs.
high mixed Algorithm Transparency and Search Manipulation: Steering vs.... Consumer search continuation and interpretation of product prominence
Latency optimization is dominated by the model-provider stack and is shaped by task formulation before being affected by preprocessing; the same preprocessing technique may speed up one model-provider stack while slowing down another.
Cross-model and cross-provider latency measurements in the benchmark, including local preprocessing, upload, server processing, and response streaming components.
Skills distilled from non-reasoning trajectories alone were competitive with skills distilled from paired reasoning/non-reasoning trajectories, but the relative performance was domain-dependent.
A source-composition ablation evaluated GPT-5.4-mini on four held-out benchmarks using skills distilled from no-think-only versus paired corpora; each result was based on 3 evaluation seeds.
high mixed Reason Wide, Not Deep: Amortizing the Reasoning Premium into... Held-out benchmark pass rate
Experienced open-source developers took 19% longer to complete real tasks when AI assistance was permitted, while estimating afterward that AI assistance had made them 20% faster.
The paper cites a randomized controlled trial by Becker et al. (2025). The supplied text does not report the trial's participant sample size.
high mixed AI Evaluation Should Measure Verification Cost, Not Correctn... Actual task completion time and developers' perceived speed under AI assistance
Happy buyers have the highest acceptance rate and the lowest rejection rate among the reported buyer-emotion conditions, but 40.82% of their negotiations still reach the maximum number of turns without resolution.
Termination-profile analysis by buyer emotion; deadlocks are negotiations that reach the maximum turn limit.
high mixed Deal Me Maybe: The Role of Emotions in Multi-Agent Negotiati... Acceptance, rejection, and deadlock termination rates
Emotion conditioning substantially changes negotiation deal rates: angry buyers have an average deal rate of 0.39%, while happy buyers have the highest average deal rate at 28.91%.
Controlled evaluation of six buyer emotions crossed with six seller emotions, two budget conditions, 350 products, three repetitions per product-condition cell, and five language models; results are aggregated across models.
high mixed Deal Me Maybe: The Role of Emotions in Multi-Agent Negotiati... Negotiation deal rate, defined as the proportion of negotiations ending in agree...
In the paper's controlled testbed, the performance gap associated with myopic planning appears under capability gating, persists for every fixed lookahead horizon up to a chain-length boundary, and disappears in no-gating controls.
The abstract and contribution summary report controlled-testbed comparisons involving gating parameters, fixed lookahead horizons, chain lengths, distractor-build controls, and no-gating controls; no numerical sample size or effect estimate is provided in the supplied text.
high mixed Capability-Gated Planning: Cost-to-Goal Discovery and the Li... Planner performance and expected discovery cost across gating and lookahead cond...
Moderate positive temporal steering improves TravelPlanner commonsense constraint performance, while negative steering degrades it; very large positive steering also degrades performance.
The TravelPlanner validation split contains 180 queries covering three difficulty levels and three trip lengths. The primary outcome is Commonsense Constraint Micro Pass Rate, evaluated using the layer-40 temporal direction across steering strengths.
high mixed Intertemporal Preference Steering in Qwen3 via Contrastive A... Commonsense Constraint Micro Pass Rate on multi-day travel itineraries
For the two task-level cost case studies, hardening left GPT-5.6 Luna's success rate nearly unchanged on compile-compcert but sharply degraded success on caffe-cifar-10.
Task-level cost sampling with GPT-5.6 Luna using Codex: 100 runs per condition for each of two tasks under control and NIST-derived high; all 200 hardened trajectories were inspected.
high mixed Permission Denied: Policy-Graded Evaluation of Coding Agents... Task success rate under policy hardening
Most Terminal-Bench tasks remained solvable under the strictest policy: 82 of 89 tasks had a solvability witness, while 7 tasks were blocked by design.
The authors replayed each task's official reference solution under the strictest policy and, when necessary, authored a policy-compliant reference solution; tasks requiring policy-forbidden actions were classified as blocked by design.
high mixed Permission Denied: Policy-Graded Evaluation of Coding Agents... Task solvability under policy enforcement
Frontier agent performance increased across all five benchmark groups between 2024 and 2026, but the amount of progress was uneven across domains.
Descriptive frontier analysis using, for each task, the highest score among agents eligible by release quarter, followed by averaging within benchmarks and benchmark groups.
high mixed Messier: A High-Resolution Corpus for Cross-Benchmark Agent ... Frontier pass rate by benchmark group over time
The AMD AITER runtime is required for competitive MLA inference throughput and must be selectively disabled for architectures with incompatible attention head configurations.
Comparative runtime experiments reported in the paper showing MLA throughput improvement with AITER and cases where AITER caused problems on architectures with incompatible head configurations (authors recommend selective disabling).
high mixed Architecture-Aware LLM Inference Optimization on AMD Instinc... inference throughput and runtime compatibility with attention head configuration...
Pull request description styles are associated with differences in reviewer response timing.
Analysis linking PR description features to reviewer response timing using the AIDev dataset covering PRs generated by five AI coding agents.
high mixed How AI Coding Agents Communicate: A Study of Pull Request De... reviewer response time (timing)
Performance, energy use, and time often involve intricate trade-offs under heterogeneous workloads.
Conceptual analysis in the paper describing trade-offs across performance, energy, and latency under varying workloads; no quantitative sample or experiment reported in the excerpt.
high mixed AI Chips and the Economics of Computer trade-offs among performance, energy use, and time
The study investigated the impact of text data analytics on audit report lag of listed manufacturing firms in Nigeria using a mixed-method design.
Methodological description provided in the paper's abstract: mixed-method design combining primary firm-level data from external auditors and secondary data from audited annual reports of listed manufacturing firms in Nigeria.
high mixed Audit Report Lag Responsiveness of Text Data Analytics: Time... impact of text data analytics on audit report lag
The review found that AI system applicability correlates with occupational task operation completion, wages, employment prospects, and education, influencing business transformation and economic growth.
Main synthesis claim from the systematic literature review of 2024–2025 publications; described as correlations between AI applicability and occupational/economic outcomes.
high mixed The algorithmic management of job loss and creation in the e... task operation completion, wages, employment prospects, education
Learning-curve shape (through its learning speed α) is the primary theoretical determinant of when to stop experimenting; costs determine switching profitability.
Analytical decomposition and theoretical results in the paper linking the learning-curve parameter α to the optimal stopping time and stating role of costs in profitability.
high mixed The Challenger: When Do New Data Sources Justify Switching M... determinants of optimal stopping time and switching profitability
Security Operations Centers (SOCs) face massive, heterogeneous alert streams under minute-level service windows, creating an "Alert Triage Latency Paradox": verbose reasoning chains ensure accuracy and compliance but incur prohibitive latency and token costs, while minimal chains sacrifice transparency and auditability.
Framed as the motivating problem in the paper (abstract). No experimental sample sizes given; presented as conceptual/problem statement supported by the paper's motivating discussion.
Our results show distinct scaling regimes for red- and blue-team tasks.
Empirical results reported in the paper comparing offensive (red-team) and defensive (blue-team) task performance as a function of cost/compute budget.
high mixed Beyond Success Rate: Cost-Aware Evaluation of Offensive and ... scaling_behavior_of_task_performance
In coding tasks, low agreeableness leads to large communication shifts that have little effect on milestone completion.
Experimental manipulation of agreeableness in LLMs on structured coding tasks; observed large changes in communication but little change in milestone completion rates. No quantitative effect sizes or sample counts given in the abstract.
high mixed When Does Personality Composition Matter for Multi-Agent LLM... milestone completion (task completion success)
Participants' IAT scores were predictive of the time they spent in human-AI collaboration.
Reported predictive relationship between individual IAT scores and measured time spent interacting with/considering resumes during human-AI collaborative screening tasks (likely from regression or correlation analyses); exact statistics and sample size not provided in the excerpt.
high mixed Resume Screening, Fast and Slow: (Biased) AI Recommendations... time spent in human-AI collaboration (resume viewing / interaction time)
Frontier proprietary models achieve near-zero success under GUI-based interaction, whereas COM-based execution yields substantial immediate gains.
Experimental comparison reported in the paper on ComCADBench between GUI-based interaction by proprietary models and COM-based execution (authors report success rates and comparative performance).
high mixed ComAct: Reframing Professional Software Manipulation via COM... success rate on CAD tasks under GUI-based interaction vs COM-based execution
The same observation is seen with the amount of changes (e.g., code churn, number of modified files) and with the efforts to merge an agentic PR (e.g., merge time and number of comments).
Reported that pre/post comparison across projects shows mixed/no consistent improvement patterns for code churn, modified files, merge time, and comment counts after instruction-file creation (analysis over 15,549 PRs in 148 projects).
high mixed Toward Instructions-as-Code: Understanding the Impact of Ins... amount of changes (code churn, number of modified files) and effort to merge (ti...
Claude Code completed the pipeline in ~3.4 minutes with silent deviations from the specification, while Codex required ~16 minutes across explicit self-correcting restarts, including an unsolicited performance optimization of the matched filter inner loop.
Reported run-time measurements and qualitative behavior descriptions in paper: timing values (~3.4 min vs ~16 min) and observed behaviors (silent deviations for Claude Code; explicit restarts and an unsolicited optimization by Codex).
AI assistance can generate a deceptive productivity signature: average completion times fall because AI tools typically supply a fast first draft, yet workflow-level performance can deteriorate when a subset of AI errors escapes review and returns as costly downstream rework.
Analytical derivation and discussion based on the paper's queueing model (theoretical/model-based evidence; no empirical sample provided).
Across 78 endpoints, the same model on different endpoints differs in tail latency by an order of magnitude.
Empirical tail-latency measurements across 78 endpoints serving 12 model families.
Provisioned Throughput delivers the lowest latency at low concurrency but saturates its reserved capacity above approximately 20 concurrent users.
Empirical measurements from the instrumented system across concurrency up to 50 users and tier comparisons; the paper reports the observed saturation point near ~20 concurrent users.
high mixed Latency and Cost of Multi-Agent Intelligent Tutoring at Scal... response time (latency) and saturation threshold (concurrency where reserved cap...
Wall-clock time can be reduced to O(√E) through team parallelism, but total human effort remains O(E).
Model-derived result showing parallelism across humans can speed wall-clock completion time while aggregate human effort does not drop asymptotically.
high mixed The Novelty Bottleneck: A Framework for Understanding Human ... wall-clock task completion time and total human effort
AI-related efficiency improvements can reduce the amount of labour time required while workers remain employed in the short term.
Frontline-worker responses were coded as reduced overtime and partial job substitution, with the paper interpreting these as lower labour requirements without necessarily implying immediate total job loss.
high negative The Impact of Artificial Intelligence on Low-Skilled Employm... Working time and quantity of tasks performed
Repair briefs that nominated only unmodifiable files caused a repair window to consume 943 tool turns across 71 minutes, with 48% of the effort spent reading or searching.
Measured tool-turn and duration records from a cross-process failure-repair window; 451 of the 943 turns were read or search operations.
high negative Agent Mesh: Reliability Primitives for Non-Idempotent Agent ... Agent effort and duration spent diagnosing a misrouted failure
FinVision reduced task completion time by an average of 51 percent for accounting and investment professionals.
A controlled user study involving 48 accounting and investment professionals, as reported in the abstract.
high negative Frontiers in FinTech: Multimodal Foundation Models for Finan... Task completion time for financial analysis tasks
Phase 3 response times were longer in the Low-cost AI condition than in either the High-cost AI or No-AI condition.
Between-condition comparison of Phase 3 response times: Low-cost AI mean=111.81 seconds, High-cost AI mean=86.78 seconds, No-AI mean=92.21 seconds; p<0.01 and p=0.01 for the respective comparisons.
high negative How AI Assistance Affects Human Skill Development: A Study o... Response time during the post-AI unassisted assessment
Resilient supply networks should reduce order cancellations, lead-time variance, export interruptions, and time to recovery.
Proposed outcomes and measurement indicators for the international trade continuity layer; no observed results are reported.
high negative An Integrated Big Data and Predictive Analytics Framework fo... Order cancellations, lead-time variance, export interruptions, and recovery time
In a randomized trial involving experienced open-source developers, access to early-2025 AI tools increased task-completion time by an estimated 19 percent.
The paper reports the findings of Becker, Rush, Barnes, and Rein (2025): 16 experienced developers completed 246 tasks in repositories they knew well.
high negative Human Computational Capital and the AI Time Dividend: A Cond... Software-development task completion time
Distributed and asynchronous work weakens assumptions of co-located, simultaneous labor and creates coordination and latency costs that differ from classical time-per-task metrics.
Conceptual discussion of asynchronous production and coordination, using information-processing and time-cost perspectives.
high negative Time as a Factor of Competitiveness: From Taylorist Time Stu... Coordination latency and process efficiency in asynchronous work
Any exact simulation of a binary target experiment H using repeated observations from source experiment F must satisfy the statewise information lower bounds I0(H) <= I0(F) E0[N] and I1(H) <= I1(F) E1[N].
Theorem 2.1(a), proved using the likelihood-ratio representation on the stopped sigma-field and conditional Jensen's inequality.
high negative The Order of Binary Experiments under Endogenous Stopping Expected number of observations required for exact sequential simulation
Most AI interactions are brief, typically lasting minutes, whereas evidence-based psychotherapies generally require months of longitudinal engagement.
Conceptual comparison between typical AI interaction patterns and the duration and longitudinal structure of evidence-based psychotherapy.
high negative A framework for evidence-based psychotherapy with AI (EBP-AI... Longitudinal treatment engagement and duration
Models degrade when a task is distributed across several turns rather than stated as a complete single-turn instruction.
The survey cites comparative evidence from prior work examining single-turn versus distributed multi-turn task instructions.
high negative Multi-turn Conversational AI from Text to Multimodal Interac... Task performance under multi-turn versus single-turn instruction presentation
Reducing methodological guidance increased average per-task token consumption and execution time: relative to B1, token use rose by 59% in B2, 25% in B3, and 30% in B4, while execution time rose by 32%, 22%, and 18%, respectively.
Reported average token consumption and execution time per task under B1–B4, shown in Figure 4.
high negative ASI-Bench: At the Dawn of Artificial Superintelligence Computational cost and execution time per research task