The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests About 🎲 Workforce Futures
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Evidence (8974 claims)

Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.

The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).

Browse by theme

Nine broad, paper-level topics. Click one to filter the claims below.

Adoption
10085 claims
Filter claims →
Productivity
8974 claims
Filtered →
Governance
8062 claims
Filter claims →
Human-AI Collaboration
7749 claims
Filter claims →
Org Design
5057 claims
Filter claims →
Innovation
4896 claims
Filter claims →
Labor Markets
4088 claims
Filter claims →
Skills & Training
3372 claims
Filter claims →
Inequality
2377 claims
Filter claims →

Claims by outcome category

Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.

Outcome Positive Negative Mixed Null Total
Other 882 244 117 1097 2424
Governance & Regulation 1010 469 229 135 1875
Organizational Efficiency 977 235 149 90 1462
Technology Adoption Rate 781 299 143 128 1362
Research Productivity 506 155 74 363 1110
Output Quality 555 219 71 70 915
Decision Quality 395 200 95 54 751
Firm Productivity 523 67 101 27 724
AI Safety & Ethics 262 309 75 36 688
Market Structure 195 201 135 30 566
Task Allocation 248 77 96 38 464
Innovation Output 300 34 55 20 411
Skill Acquisition 207 75 65 21 368
Employment Level 138 67 119 24 350
Fiscal & Macroeconomic 156 80 53 33 329
Task Completion Time 211 38 13 16 280
Firm Revenue 183 52 29 5 270
Consumer Welfare 131 77 48 13 269
Inequality Measures 50 141 54 9 254
Worker Satisfaction 104 85 25 13 227
Error Rate 87 112 11 5 215
Automation Exposure 69 69 37 20 198
Wages & Compensation 102 49 31 11 193
Team Performance 115 30 30 11 187
Regulatory Compliance 88 74 17 7 186
Training Effectiveness 109 22 14 21 168
Developer Productivity 116 21 15 8 161
Job Displacement 12 92 26 1 131
Hiring & Recruitment 57 12 9 5 83
Skill Obsolescence 6 59 10 2 77
Social Protection 43 17 8 2 70
Creative Output 35 21 9 4 70
Labor Share of Income 18 23 17 1 59
Worker Turnover 15 16 4 35
Industry 1 1
Clear
Productivity Remove filter
AI-only planning minimizes time and cost.
Controlled, three-condition experiment (AI-only, human-only, hybrid) conducted on a live client deliverable at a mid-sized digital agency; quantitative metrics included time and cost measures (reported alongside estimation accuracy, rework rates, and scope change recovery time).
The bounded-autonomy architecture is a practical, deployed approach for making imperfect language models operationally useful in enterprise systems.
Deployment and reported performance in the described multi-tenant enterprise application evaluation (completion rates, safety interceptions, speedups); the paper synthesizes these empirical results to support the practical claim.
high positive Bounded Autonomy for Enterprise AI: Typed Action Contracts a... operational usefulness of LLMs in enterprise context
The enterprise application remains the source of truth for business logic and authorization, while the orchestration engine operates over an explicit published actions manifest.
Architectural proposal and implementation details described in the paper; asserted as part of the bounded-autonomy design deployed in the enterprise application.
high positive Bounded Autonomy for Enterprise AI: Typed Action Contracts a... system design property (source-of-truth and orchestration behavior)
Several safety properties are structurally enforced by code and intercepted all targeted violations regardless of model output.
Design and deployment of bounded-autonomy architecture with typed action contracts, permission-aware capability exposure, scoped context, validation before side effects, and consumer-side execution boundaries; empirical claim that these code-enforced properties intercepted targeted violations during evaluation.
high positive Bounded Autonomy for Enterprise AI: Typed Action Contracts a... interception of targeted violations / enforcement of safety properties
Both AI conditions delivered 13–18x speedup over manual operation.
Timing/performance comparison across the three experimental conditions (manual operation, unconstrained AI, full bounded autonomy) within the deployed evaluation; reported speedup range 13–18x relative to manual operation.
high positive Bounded Autonomy for Enterprise AI: Typed Action Contracts a... task completion time (speedup vs. manual)
The bounded-autonomy system completed 23 of 25 tasks with zero unsafe executions.
Evaluation in a deployed multi-tenant enterprise application across 25 scenario trials spanning seven failure families; comparison across three conditions (manual, unconstrained AI with safety layers disabled, full bounded autonomy).
high positive Bounded Autonomy for Enterprise AI: Typed Action Contracts a... tasks completed / unsafe executions
Successful AI implementation in auditing requires an integrated framework that aligns technological readiness, auditor acceptance, and innovation diffusion to sustainably improve audit quality in Indonesia.
Authors' conclusion and recommendation derived from thematic synthesis of reviewed literature and comparative findings.
high positive Implementing Artificial Intelligence in Auditing: A Systemat... requirements for successful AI implementation and resulting audit quality
Comparative analysis indicates Indonesia remains at the early majority stage of AI adoption in auditing.
Authors' comparative synthesis of the reviewed literature and country-specific discussion classifying Indonesia's adoption stage as early majority.
high positive Implementing Artificial Intelligence in Auditing: A Systemat... innovation diffusion stage of Indonesia's auditing sector
Comparative analysis indicates global audit firms are positioned at the innovators and early adopters’ stage of AI adoption.
Authors' comparative synthesis of the reviewed literature classifying global audit firms' diffusion stage (innovation adoption framework) based on patterns in the articles.
high positive Implementing Artificial Intelligence in Auditing: A Systemat... innovation diffusion stage of global audit firms
AI implementation has been shown to significantly enhance audit efficiency, accuracy, and overall audit quality.
Synthesis of findings across the reviewed articles (thematic analysis) reporting positive effects of AI on efficiency, accuracy, and audit quality.
high positive Implementing Artificial Intelligence in Auditing: A Systemat... audit efficiency, audit accuracy, audit quality
Global auditing practices increasingly utilize machine learning, natural language processing, and robotic process automation to support risk-based auditing, fraud detection, and continuous auditing.
Thematic analysis of the 15 selected journal articles identifying dominant AI techniques (ML, NLP, RPA) and common use cases (risk-based auditing, fraud detection, continuous auditing).
high positive Implementing Artificial Intelligence in Auditing: A Systemat... use of specific AI techniques and application areas in auditing
Overall, GAI provides a principled and scalable approach to integrating AI-generated information.
Summary claim in the abstract based on the combination of the theoretical properties and empirical results reported in the paper.
high positive Generative Augmented Inference scalability and principled integration of AI-generated information
Across applications, GAI improves confidence interval coverage without inflating width.
Empirical claim reported across the multiple application studies in the paper (abstract states CI coverage improvement while maintaining or not inflating width); details in main text/appendix presumably contain the quantitative analysis.
high positive Generative Augmented Inference confidence interval coverage and width (statistical inference quality)
In health insurance choice, GAI cuts labeling requirements by over 90% while maintaining decision accuracy.
Reported empirical result from the paper's health insurance choice experiment; abstract gives the >90% reduction claim but does not include sample size or exact metrics in the abstract.
high positive Generative Augmented Inference human labeling requirements; decision accuracy
In retail pricing, where all methods access the same auxiliary inputs, GAI consistently outperforms alternative estimators, highlighting the value of its construction rather than differences in information.
Empirical experiment in a retail pricing application comparing multiple estimators given identical auxiliary inputs; stated as consistent outperformance in the abstract (no numerical effect sizes or sample sizes provided there).
high positive Generative Augmented Inference estimator performance in retail pricing (e.g., predictive or decision accuracy /...
In conjoint analysis with weak auxiliary signals, GAI reduces estimation error by about 50% and lowers human labeling requirements by over 75%.
Reported empirical result from the paper's conjoint analysis experiment(s); exact sample size and experimental details are not stated in the abstract.
high positive Generative Augmented Inference estimation error; human labeling requirements
Empirically, GAI outperforms benchmarks across diverse settings.
Empirical experiments reported across multiple application settings (conjoint analysis, retail pricing, health insurance choice) comparing GAI to alternative estimators/benchmarks.
high positive Generative Augmented Inference overall performance relative to benchmarks (estimation error / predictive perfor...
The authors establish asymptotic normality for the GAI estimator and show a 'safe default' property: relative to human-data-only estimators, GAI weakly improves estimation efficiency under arbitrary auxiliary signals and yields strict gains whenever the auxiliary information is predictive.
The paper claims formal theoretical results (asymptotic normality and efficiency comparisons) — supported by analytic derivations/proofs in the manuscript as referenced in the abstract.
high positive Generative Augmented Inference estimation efficiency (asymptotic variance / efficiency relative to baseline)
GAI uses an orthogonal moment construction that enables consistent estimation and valid inference with flexible, nonparametric relationship between LLM-generated outputs and human labels.
The paper presents a methodological proposal (Generative Augmented Inference) and states theoretical properties (orthogonal moment construction, consistency, valid inference) — supported by formal asymptotic analysis/proofs in the paper (the abstract references establishing asymptotic normality).
high positive Generative Augmented Inference consistent estimation and valid inference (statistical estimation properties)
Clear specifications, explicit governance, and ongoing human-AI collaboration are critical for successful scaling of regression automation.
Conclusions and recommendations derived from the case study's lessons and mixed-method evaluation.
high positive Human-AI Collaboration for Scaling Agile Regression Testing:... success of scaling regression automation / effectiveness of human-AI teaming
The Copilot achieves 30-50% code reuse when generating candidate test scripts.
Quantitative result reported in the paper's evaluation (stated 30-50% code reuse in the abstract/summary).
high positive Human-AI Collaboration for Scaling Agile Regression Testing:... code reuse in generated test scripts
Mixed-method evaluation shows the AI accelerates script authoring and increases throughput.
Empirical claim based on the paper's mixed-method evaluation (qualitative and quantitative data reported in the case study); specific sample sizes not provided in the summary.
high positive Human-AI Collaboration for Scaling Agile Regression Testing:... script authoring speed and throughput
Automated regression testing is essential for maintaining rapid, high-quality delivery in Agile and Scrum organizations.
Introductory/position statement in the paper; general premise motivating the case study (no specific empirical test reported).
high positive Human-AI Collaboration for Scaling Agile Regression Testing:... ability to maintain rapid, high-quality delivery
AIBuildAI ranks first on MLE-Bench with a medal rate of 63.1%, outperforming all existing baseline methods and matching the capability of highly experienced AI engineers.
Empirical evaluation on MLE-Bench reported in the paper (benchmark ranking, metric = medal rate).
high positive AIBuildAI: An AI Agent for Automatically Building AI Models medal rate (task success rate) on MLE-Bench
AIBuildAI adopts a hierarchical agent architecture in which a manager agent coordinates three specialized sub-agents: a designer for modeling strategy, a coder for implementation and debugging, and a tuner for training and performance optimization; each sub-agent is itself an LLM-based agent capable of multi-step reasoning and tool use, enabling end-to-end automation of the AI model development process that goes beyond the scope of existing AutoML approaches.
System architecture description in the paper (methods/architecture section).
high positive AIBuildAI: An AI Agent for Automatically Building AI Models system architecture and claimed capabilities (multistep reasoning, tool use, end...
We introduce AIBuildAI, an AI agent that automatically builds AI models from a task description and training data.
Methodological contribution: system design and implementation described in the paper (introduction/methods).
high positive AIBuildAI: An AI Agent for Automatically Building AI Models ability to produce AI models from task descriptions and training data
This tension reveals a pattern we call 'bounded delegation': developers wanted AI to absorb the assembly work surrounding their craft, never the craft itself.
Interpretive result from the paper's qualitative thematic analysis of survey responses (n=860), labeled by the authors as the 'bounded delegation' pattern.
high positive To Copilot and Beyond: 22 AI Systems Developers Want Built preferred boundary of automation / delegation
Developers wanted systems enforcing explicit authority scoping, provenance, uncertainty signaling, and least-privilege access throughout.
Reported constraints and desiderata from the thematic analysis of survey responses (n=860).
high positive To Copilot and Beyond: 22 AI Systems Developers Want Built desired governance/security features for AI tools (authority scoping, provenance...
Developers wanted systems that embed quality signals earlier in their workflow to keep pace with accelerating code generation.
Thematic findings from the paper's human-in-the-loop, multi-model council-based analysis of survey responses (n=860).
high positive To Copilot and Beyond: 22 AI Systems Developers Want Built requested placement/timing of quality signals in developer workflow
Using a human-in-the-loop, multi-model council-based thematic analysis, we identify 22 AI systems that developers want built across five task categories.
Qualitative analysis method described in the paper applied to the survey responses (n=860); result reported as identification of 22 desired AI systems organized into five categories.
high positive To Copilot and Beyond: 22 AI Systems Developers Want Built catalog of desired AI systems and task categories
The results demonstrate the importance of considering interacting systems of AI agents when doing both capabilities and safety research.
Authors' interpretation/generalization based on experimental findings comparing multi-agent organizations and single agents across tasks and settings.
high positive AI Organizations are More Effective but Less Aligned than In... research priorities/considerations for capabilities and safety research (implica...
BTB enables automated evaluation of any LLM or agent, scoring deliverables against 100+ rubric criteria defined by veteran investment bankers to capture stakeholder utility.
Design claim in abstract describing the benchmark's automated scoring system and rubric size (100+ criteria) defined by expert bankers.
high positive BankerToolBench: Evaluating AI Agents in End-to-End Investme... number of rubric criteria for automated evaluation
For reproducibility all our data and code are provided at https://github.com/scaleapi/scipredict
Explicit reproducibility statement and URL provided in the paper.
high positive SciPredict: Can LLMs Predict the Outcomes of Scientific Expe... data_and_code_availability
SciPredict addresses two critical questions: (a) can LLMs predict the outcome of scientific experiments with sufficient accuracy? and (b) can such predictions be reliably used in the scientific research process?
Statement of research goals and scope in the paper introducing the SciPredict benchmark and accompanying evaluations.
high positive SciPredict: Can LLMs Predict the Outcomes of Scientific Expe... research_questions_addressed
Human experts demonstrate strong calibration: their accuracy increases from ≈5% to ≈80% as they deem outcomes more predictable without conducting the experiment.
Reported stratified accuracy of human experts on SciPredict tasks by self-reported predictability judgments; accuracy rises from ≈5% (when judged not predictable) to ≈80% (when judged predictable).
high positive SciPredict: Can LLMs Predict the Outcomes of Scientific Expe... calibration_of_human_confidence_vs_accuracy
We introduce SciPredict, a benchmark comprising 405 tasks derived from recent empirical studies in 33 specialized sub-fields of physics, biology, and chemistry.
Construction of the SciPredict benchmark described in the paper; explicitly reports 405 tasks and 33 sub-fields.
The paper documents 14 deliberate conservative assumptions — including frozen base GDP, no AI-on-AI compounding, a permanent friction floor, and conservative capture rates — all of which directionally understate the benefit.
Paper lists 14 conservative modeling assumptions and claims they bias results downward (i.e., understate potential benefits).
high positive AI Capex Is Justified: A Bottom-Up Sectoral Estimate of Arti... directional bias of model assumptions relative to potential benefits
Even excluding demand expansion and robotics layers entirely, the direct productivity contribution alone reaches approximately $940 billion per year by 2036.
Model output reported in the paper when removing demand expansion and robotics layers.
high positive AI Capex Is Justified: A Bottom-Up Sectoral Estimate of Arti... direct productivity contribution to annual GDP by 2036 excluding demand expansio...
In all four scenarios, cumulative net GDP exceeds cumulative AI infrastructure investment before 2036, with the base case achieving payback in 2033.
Model financial calculation comparing cumulative net GDP uplift to cumulative AI infrastructure investment across scenarios; explicit payback year reported for base case.
high positive AI Capex Is Justified: A Bottom-Up Sectoral Estimate of Arti... year when cumulative net GDP exceeds cumulative AI infrastructure investment (pa...
The base-case scenario yields approximately $1,057 billion in net annual GDP uplift by 2036, equivalent to 3.6 percent of 2024 GDP; the bear case produces $796 billion, the bull case $1,368 billion, and an agentic scenario produces $2,521 billion.
Model scenario outputs presented in the paper (four scenarios differentiated by capture rate and friction assumptions).
high positive AI Capex Is Justified: A Bottom-Up Sectoral Estimate of Arti... net annual GDP uplift by 2036 (US, scenario-specific)
Sector-specific productivity gain percentages are anchored to published evidence, including a randomized controlled trial of GitHub Copilot (Kalliamvakou et al., 2023), JPMorgan CEO disclosures, and Cognizant's New Work New World 2026 research.
Paper states productivity percentages are anchored to published evidence and specifically cites Kalliamvakou et al. (2023) RCT, JPMorgan CEO disclosures, and Cognizant (2026).
high positive AI Capex Is Justified: A Bottom-Up Sectoral Estimate of Arti... sector-specific productivity gain percentages used in the model
Organisations should invest in customisation capabilities for AI recruitment tools, implement comprehensive change management strategies, and maintain robust post-hire evaluation procedures.
Authors' recommendations derived from thematic findings and participant perspectives across two firms (qualitative synthesis of n = 22 interviews).
high positive The augmented recruiter: examining AI integration and decisi... recommended_organisational_practices_for_AI_recruitment
AI functioned optimally as an augmentative technology rather than as a replacement for human decision-makers in recruitment.
Findings: participants across the two case firms described AI being most effective when augmenting human judgment rather than replacing it (interviews n = 22).
high positive The augmented recruiter: examining AI integration and decisi... role_of_AI (augmentation vs replacement)
AI significantly enhanced efficiency through process standardisation and automation.
Findings based on participant accounts in thematic analysis (interviews n = 22) describing process optimisation and automation benefits.
high positive The augmented recruiter: examining AI integration and decisi... efficiency (process standardisation and automation)
The Principle of Maximum Heterogeneity reveals a convergence of complex phenomena across fields onto simple underlying design principles with important predictive value for future distributed production systems.
Synthesis claim in the paper arguing cross-field convergence and predictive value based on the theoretical model and conceptual examples; no empirical validation or forecasting trials reported.
high positive The Principle of Maximum Heterogeneity Optimises Productivit... predictive value of the model/principles for future distributed production syste...
The principles derived (including the Principle of Maximum Heterogeneity) can be used as a blueprint for constructing ideal distributed production systems; demonstrated by suggesting specific redesigns for compute systems executing large-scale AI.
Paper includes suggested redesigns for compute systems as demonstrations of the blueprint; these are proposed designs/illustrative applications rather than empirically validated interventions or trials.
high positive The Principle of Maximum Heterogeneity Optimises Productivit... design-guided performance improvements in compute systems for large-scale AI (pr...
The Principle of Maximum Heterogeneity applies recursively across all layers of nested production systems.
Theoretical claim within the paper arguing recursive applicability across nested system layers (e.g., neurons, firms, ecosystems); supported by conceptual reasoning and model exposition rather than empirical multi-layer tests.
high positive The Principle of Maximum Heterogeneity Optimises Productivit... emergence/spread of heterogeneity across nested layers
The communication topology determines the spatial scale over which heterogeneity spreads in distributed production systems.
Model-based theoretical argument in the paper linking topology to the spatial scale of heterogeneity; illustrated conceptually and via examples but not via empirical sample testing.
high positive The Principle of Maximum Heterogeneity Optimises Productivit... spatial scale/spread of heterogeneity as a function of communication topology
Principle of Maximum Heterogeneity: any distributed production system optimising for performance will converge on an increasingly heterogeneous configuration.
Statement of a derived principle from the paper's model (theoretical derivation/argument); demonstration via model reasoning and examples rather than empirical testing; no sample size reported.
high positive The Principle of Maximum Heterogeneity Optimises Productivit... degree of heterogeneity in agent/configuration space
A small set of underlying laws generates the complex dynamics observed across fields (biology, economics, neuroscience, computing).
Theoretical argument and synthesis across disciplines within the paper; no empirical or experimental sample size reported.
high positive The Principle of Maximum Heterogeneity Optimises Productivit... explanatory coverage of complex system dynamics