The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Evidence (2228 claims)

Search and filter individual claims pulled from the papers. Looking for a specific finding ("what's the effect on wages?"), you're in the right place. Want to compare whole outcome categories against each other instead? Use the Evidence Explorer.

The board below groups claims two ways: by broad theme (nine paper-level topics) and by outcome category (the 34 claim-level outcomes that the Explorer and Syntheses also use).

Browse by theme

Nine broad, paper-level topics. Click one to filter the claims below.

Adoption
20058 claims
Filter claims →
Productivity
17184 claims
Filter claims →
Governance
16099 claims
Filter claims →
Human-AI Collaboration
16034 claims
Filter claims →
Innovation
10501 claims
Filter claims →
Org Design
10496 claims
Filter claims →
Labor Markets
6444 claims
Filter claims →
Skills & Training
5385 claims
Filter claims →
Inequality
4148 claims
Filter claims →

Claims by outcome category

Counts by direction of finding. These are the same 34 outcome categories the Explorer compares and the Syntheses are written for. A linked row has a published synthesis.

Outcome Positive Negative Mixed Null Total
Other 1820 479 278 1820 4588
Organizational Efficiency 2711 616 401 173 3922
Governance & Regulation 2075 886 459 246 3714
Technology Adoption Rate 1467 530 258 206 2488
Decision Quality 1281 496 289 152 2228
Output Quality 1227 447 207 138 2025
AI Safety & Ethics 634 754 207 83 1688
Research Productivity 826 241 114 422 1624
Firm Productivity 1052 154 163 66 1441
Task Allocation 685 211 331 99 1335
Market Structure 433 423 242 46 1150
Innovation Output 639 91 105 34 871
Task Completion Time 476 113 43 36 672
Firm Revenue 445 126 58 25 656
Skill Acquisition 364 119 109 34 626
Consumer Welfare 288 167 104 31 592
Employment Level 214 140 174 50 582
Error Rate 230 251 35 16 535
Fiscal & Macroeconomic 268 136 71 50 532
Inequality Measures 100 307 96 12 515
Worker Satisfaction 221 173 60 30 484
Automation Exposure 155 138 65 36 398
Regulatory Compliance 171 120 30 13 335
Developer Productivity 222 58 27 13 321
Team Performance 188 56 50 24 320
Wages & Compensation 146 104 46 16 312
Training Effectiveness 207 41 21 26 298
Job Displacement 23 153 52 4 232
Hiring & Recruitment 102 57 30 11 202
Skill Obsolescence 16 102 24 6 148
Creative Output 71 42 23 6 143
Social Protection 57 30 11 3 101
Labor Share of Income 29 42 24 2 97
Worker Turnover 43 29 6 4 82
Industry 1 1
AI-enabled recruitment is associated with a shift toward data-driven decision-making in which machine intelligence complements human expertise.
The paper's literature synthesis and qualitative examination of AI-powered recruitment practices, including chatbots, video interviews, targeted job advertisements, and predictive analytics.
high mixed AI-driven talent acquisition in Indian IT firms: A qualitati... Allocation of recruitment decision-making between human expertise and AI-based s...
Ichnology is a useful testbed for formalizing inference because trace evidence is highly ambiguous: one organism can produce diverse traces, while similar traces can result from different organisms or abiotic processes.
Conceptual analysis of ambiguity in trace-based interpretation and comparison of biotic and abiotic explanations.
high mixed The epistemic architecture of palaeontological reasoning: tr... Reliability of tracemaker and process attribution
For middle managers, AI has both positive and negative effects: it supports data analysis and managerial decision-making while creating concerns about automation of some managerial responsibilities.
Cross-study synthesis of findings differentiated by organizational level.
high mixed Artificial Intelligence and Its Influences on Enterprise Emp... Managerial decision support and concerns about managerial-task automation
Static models linking stable managerial traits to stable decisions are insufficient because the effects of GM characteristics depend on dynamic competence, situational expression of values, leadership adaptability, and recognition of gendered traits.
Thematic synthesis of studies on GM characteristics and leadership styles, used to qualify the dispositional assumption.
high mixed EXPRESS: A Review of Upper Echelons Theory in Hospitality: G... Managerial decision-making and leadership effectiveness
The effect of citations on false-information detection depends on both users' digital literacy and the type of falsehood being evaluated.
The experiment estimated heterogeneous treatment effects using interaction terms for digital literacy and separately analyzed misinformation and disinformation outcomes.
high mixed The Roles of Source Citations in Identifying Misinformation ... False-information detection, separated into misinformation and disinformation de...
The commitment effect is concentrated among models rather than universal: three of the 12 models were strongly seduced by the panels, four never committed under any panel, three committed regardless of the panel, and two responded weakly.
Disaggregated analysis of commitment behavior across the 12-model frontier-model roster.
high mixed Calibrated Enough to Know, Not Calibrated to Act: Fabricated... Variation across models in commitment behavior under evidence panels
nDCG@K also produced a borderline finding in the example audit, despite the score-delta, top-K-retention, and merit-aware rate-gap metrics remaining within tolerance.
The paper compares ranking-quality results with score, retention, and merit-aware fairness metrics.
high mixed Counterfactual Bias Testing for Application Tracking System Ranking quality measured by nDCG@K
Mean absolute rank change produced borderline audit findings in the example corpus, including a finding on the neutral baseline configuration.
The paper reports the results of its rank-stability metric in the illustrative audit.
Higher model capability postpones but does not eliminate the retrieval–integration gap: the largest open-weight model retained a 3.4-percentage-point effect at 128,000 tokens, while a lower-capability commercial system eventually lost retrieval as well as decision influence.
Comparative experiments across models with different capability levels and context lengths.
high mixed Reading Is Not Using: Retrieval, Judgment, and the Design of... Disclosure influence on investment judgment and disclosure retrieval across mode...
The retrieval–integration gap replicated across three independently trained model families, although the context length at which the gap became binding differed across models.
Replication experiments conducted across three independently trained model families.
high mixed Reading Is Not Using: Retrieval, Judgment, and the Design of... Relationship between disclosure retrieval and influence on investment judgment a...
The system's bullish prior was associated with BUY being the clean prediction on 40.0% of days, compared with 26.2% for SELL.
Reported clean-system decision frequencies used to explain why some target directions were more difficult to induce.
high mixed Poisoning Agentic Alpha: Adversarial Vulnerabilities Across ... Baseline distribution of clean BUY and SELL trading decisions
In an agentic retrieval setting, changing the dial affects the information searched for, the evidence selected, and the evidence reflected in the final analysis.
Additional agentic retrieval evaluation examining information acquisition, evidence selection, and final generated analyses under different dial settings.
high mixed Your AI, On a Dial: Controlling Investment Bias in LLMs with... Information-search behavior, evidence selection, and evidence use in investment ...
Under identical balanced bullish and bearish evidence, changing Qwen3-8B to its calibrated neutral setting changed both JPMorgan Chase's and NVIDIA's decisions from buy to sell and shifted the rationales toward greater emphasis on downside risks.
Paired response-level comparison for two securities under the same evidence, before and after applying the calibrated neuron intervention.
high mixed Your AI, On a Dial: Controlling Investment Bias in LLMs with... Security-level buy/sell decision and rationale evidence emphasis
The direction of the dial's effect is model-specific: increasing the intervention coefficient shifts the investment-bias score in the opposite direction for DeepSeek-R1-14B and Mistral-24B compared with the other evaluated models.
Cross-model response curves relating intervention strength to the aggregate investment-bias score.
high mixed Your AI, On a Dial: Controlling Investment Bias in LLMs with... Change in aggregate buy-versus-sell investment stance as intervention strength i...
The reachable range of investment-bias scores is model-dependent: four models approach the full interval from -1 to 1, whereas DeepSeek-R1-14B has a narrower response range.
Evaluation of the investment-bias score over intervention-strength values for each of the five models.
high mixed Your AI, On a Dial: Controlling Investment Bias in LLMs with... Range of aggregate investment-bias scores reachable through neuron intervention
AI lowers the cost of generating and comparing alternative operationalizations of a concept, but it cannot determine which operationalization answers the research question; that judgment requires theory and substantive expertise.
Review's discussion of construct definition, competing rubrics, nomological networks, and the role of domain expertise.
high mixed The Measurement Revolution? Credible Measurement and Inferen... Quality and appropriateness of construct definitions
Different measurement functions applied to the same underlying concept can preserve different features of the data and therefore support different empirical conclusions.
Formal discussion of measurement as a projection from high-dimensional reality into a lower-dimensional representation, combined with the existence of multiple plausible operationalizations.
high mixed The Measurement Revolution? Credible Measurement and Inferen... Empirical conclusions produced by alternative measurements
The availability of AI shifts the bottleneck in empirical measurement from finding any scalable measure of a phenomenon to choosing among many plausible measures.
Authors' synthesis of the reduced cost of applying AI measurement functions and the increased number of choices involving rubrics, models, prompts, training data, and tuning strategies.
high mixed The Measurement Revolution? Credible Measurement and Inferen... Choice and credibility of measurement variables
Predictive accuracy alone does not necessarily produce business impact or supply-chain resilience; predictions must also be timely, calibrated, interpretable, and connected to decision rights and feasible response options.
Conceptual argument supported by distinctions among predictive intelligence, decision actionability, and resilience orchestration; the paper cites prior analytics literature but reports no original test.
high mixed An Integrated Big Data and Predictive Analytics Framework fo... Business impact and supply-chain resilience from predictive analytics
The observed pattern of lower accuracy on modern statutory updates is consistent with the paper’s hypothesis of ‘precedent overfitting,’ but the study does not establish that precedent overfitting is the causal mechanism.
Comparison of model performance across historical doctrinal control categories and cases involving the 2018 amendments; the models were evaluated as black boxes without access to weights, attention mechanisms, or retrieval rankings.
high mixed Can Legal AI Know When It Is Wrong? And Do Students Know Whe... Pattern of legal reasoning errors across historical principles versus modern sta...
Fact payloads carried the gold evidence in 98–99% of cases, while chunk reading accuracy declined from 81% to 73% as more text was supplied.
Decomposition of the held-out question-answering results into evidence coverage and reading accuracy across token budgets.
high mixed RAG Deserves an Index: Why Ingest-Time Compilation Beats Que... Evidence coverage and reading accuracy as context increases
The same systematic review found that human judgment remained decisive under high uncertainty, while several included studies identified bias and diminished trust as unresolved concerns even when efficiency gains were present.
López-Solís et al. (2025) systematic review of 30 studies, as summarized by the paper.
high mixed Generative AI in Enterprise Decision Intelligence: Opportuni... Role of human judgment, perceived trust, and bias in AI-assisted strategic decis...
A technically accurate forecast can have less business value when delivered after planning decisions are fixed than a somewhat less accurate forecast embedded in a responsive decision process.
Conceptual inference in the review based on the relationship between decision latency, process integration, and business value; no quantitative test is reported in the paper.
high mixed Big Data, Artificial Intelligence, and Machine Learning Inte... Business value of predictive analytics as a function of forecast timing and inte...
Algorithmic recruitment tools are not inherently more biased than the human-mediated processes they replace; their effects on equity depend substantially on system design, training data, and auditing practices.
Review of countervailing evidence comparing structured or algorithmic screening with unstructured human interviews, citing Bogen and Rieke (2018) and Stone et al. (2024).
high mixed Artificial intelligence, multilingualism, and career sustain... Equity and bias in recruitment decisions
Managers in the OECD research often associated algorithmic management with improved decision quality and efficiency, but also reported concerns about unclear accountability, opaque algorithmic logic and effects on workers.
Findings from the OECD survey of more than 6,000 firms in six countries; the paper reports both perceived benefits and governance or worker-related concerns.
high mixed CHAPTER 6. STRATEGIC MANAGEMENT OF HUMAN–AI COLLABORATION: N... Perceived decision quality, efficiency, accountability and worker impacts
Persona information was most beneficial when habitual travel information was unavailable.
Factorial comparison of prompting configurations with and without habitual travel information and with persona information.
high mixed An Agentic Approach for Active Data Collection, Travel Behav... LLM travel-mode prediction accuracy as a function of persona and habitual-travel...
For Qwen2.5-7B, judge scores showed above-chance per-step sign agreement with replay contribution, but did not identify or concentrate on pivotal steps better than the shuffled-control benchmark.
Judge sign agreement was 84/139=60.4%, with a 95% interval excluding 50%; however, judge precision-at-pivotal lift was 1.000 with an interval containing the chance value of 1.0, and rank fidelity was indistinguishable from its own shuffle.
high mixed Credit Without Ground Truth: Auditing Step-Level Credit Assi... Per-step sign agreement and concentration of judge credit on causally pivotal st...
Calibration regret and behavioral overreliance can diverge: high-quality AI on hard tasks produced the highest regret despite not producing the highest overreliance.
Table 3 reports final-state outcomes for four combinations of AI quality and task difficulty, averaged over 20 seeds.
high mixed Modeling AI Overreliance as a Complex Adaptive System Calibration regret and behavioral overreliance
A naive decision rule that classified a 1/5 pass rate as scattered and non-zeroable was false; inspecting failure causes showed that a shared cause could still be corrected completely by one rule.
The reported replication contradicted the earlier pass-rate-only classification; manual inspection found homogeneous failures, and the rule raised the rate to 5/5.
high mixed Grouping the Stochastic Machine: Precision, Not Capability, ... Validity of pass-rate-only diagnosis of failure recoverability
In solution-selection tasks, models often select incorrect options even when valid solutions are present; NOTA improves average accuracy but models rarely use it, including when no valid option is present.
Verification-only selection experiments with a 'none of the above' option across solution-concept tasks.
high mixed Preference Reasoning under Indeterminacy in Large Language M... Selection accuracy, abstention frequency, and calibration regarding the existenc...
Multiple-choice prompting can improve indeterminacy detection for some models while inducing false indeterminacy judgments for others.
Partial-order query comparison across prompting formats and models.
high mixed Preference Reasoning under Indeterminacy in Large Language M... Accuracy and calibration in identifying whether pairwise preference queries are ...
The study evaluates three explanation pipelines: a tabular pipeline using XGBoost with SHAP, a network pipeline using a graph neural network with GNNExplainer, and a bimodal pipeline combining tabular and network evidence.
Experimental design using Freddie Mac single-family loan-level data and separate prediction, post-hoc explanation, and LLM narrative-generation stages.
high mixed Communicating Credit Risk with Large Language Models: Evalua... Explanation quality across evidence modalities
LLM-generated credit-risk explanations reliably identify influential factors but are less reliable at correctly stating the direction of those factors' influence.
Automated checks of generated narratives against the underlying SHAP and GNNExplainer evidence across the tabular, network, and bimodal pipelines.
high mixed Communicating Credit Risk with Large Language Models: Evalua... Factor identification and directional correctness of explanation narratives
A LoRA-recovered compressed variant can remain fully parseable and achieve 71% strict BoolQ accuracy while sending 97 of 100 predictions to a single class, resulting in 52.6% balanced accuracy.
BoolQ evaluation of compressed and LoRA-recovered language-model variants, including prediction-distribution analysis over 100 predictions.
high mixed Large Models for Small Devices: Recent Advances and Empirica... Strict accuracy, balanced accuracy, output parseability, and prediction-class di...
Among the two evaluated theory-of-mind skills, next-action prediction predicts negotiation outcomes, whereas preference inference alone does not.
Analysis relating the two ToM components—preference inference and next-action prediction—to negotiation outcomes.
Path-based random-utility models provide interpretable marginal utilities, elasticities, values of time, and welfare analysis, but their use can be limited by enormous path spaces and unknown or misspecified choice-set generation.
Review of path-based random-utility models, their estimands, and path-enumeration and choice-set limitations.
high mixed Learning Sequential Mobility Choice: A Review of Route and A... Interpretability and validity of route-choice estimation
Machine-learning capabilities such as occupancy-ratio methods, graph and sequence models, constrained learning, transfer, data fusion, and multi-agent learning do not by themselves establish preference recovery or valid policy counterfactuals.
Critical synthesis of method families and their evidentiary limits in transportation applications.
high mixed Learning Sequential Mobility Choice: A Review of Route and A... Preference recovery and policy-counterfactual validity
Offline reinforcement learning based on logged behavior is prescriptive rather than automatically descriptive of the behavior-generating process.
Conceptual distinction in the review between improving a policy from logged data and recovering descriptive behavioral preferences.
high mixed Learning Sequential Mobility Choice: A Review of Route and A... Descriptive validity versus policy improvement from logged mobility data
Imitation learning can reproduce observed behavior without explaining the underlying preferences or behavioral mechanism.
The review distinguishes policy or occupancy matching from structural recovery of utility or reward.
high mixed Learning Sequential Mobility Choice: A Review of Route and A... Interpretability and structural explanation of observed mobility behavior
The equivalence between recursive choice, dynamic discrete choice, and maximum-entropy IRL does not by itself establish that the models identify the same behavioral quantities.
The review distinguishes utility, reward, policy, occupancy, constraints, and observation error as different estimands with different interpretations and counterfactual implications.
high mixed Learning Sequential Mobility Choice: A Review of Route and A... Behavioral identification and validity of policy counterfactuals
Centaur, a model fine-tuned on more than ten million human choices across hundreds of experiments, predicts human behavior on held-out tasks better than bespoke models but remains only weakly equivalent to human cognition.
The paper summarizes the Centaur model's training and evaluation results and argues that predictive success does not establish shared underlying mechanisms.
high mixed Process-Constituted Intelligence: A Shared Criterion for Hum... Prediction of human behavior on held-out tasks and process-level equivalence
GPT-5 achieved the highest screening recall, at 91.8%, but had lower specificity.
Comparative screening evaluation against expert-annotated inclusion and exclusion decisions for the 244-document benchmark subset.
high mixed Knowledge Synthesis Review Framework: Task-Level Benchmarkin... Screening recall and specificity
The PaperFindingBench evaluation combines exact-match scoring for approximately 27% of queries with an LLM judge scoring agent-supplied evidence for the remaining approximately 73%.
Benchmark scoring protocol described in Table 1 and the PaperFindingBench setup.
high mixed Competing at Every Price Point with Agentic Evolution over a... Literature-retrieval adjusted micro-F1 under exact-match and LLM-judged scoring
Across the cited research, the effect of human-in-the-loop oversight on decision quality is uneven.
The author synthesizes studies reporting limited effects on discrimination, mixed results in child welfare, and cases where humans improve technologically mediated processes.
high mixed Lost in the Loop: Who Is the ‘Human’ of the Human in the Loo... Decision quality under human oversight of automated systems
Human intervention against malfunctioning algorithmic risk-prediction tools in child-welfare decisions has limited effectiveness, with results that are mixed.
The chapter's summary of De-Arteaga, Fogliato, and Chouldechova's study of human oversight in child-welfare decision-making.
high mixed Lost in the Loop: Who Is the ‘Human’ of the Human in the Loo... Human capacity to detect or correct erroneous algorithmic risk scores
AI-generated predictions do not eliminate the need for human judgment, especially for economic decisions involving ethical considerations, strategic priorities, and institutional constraints.
Conceptual synthesis drawing on the distinction between prediction and judgment and on decision-theory literature; no empirical sample is reported.
high mixed Artificial Intelligence-Driven Economic Decision Making: A C... Human involvement in value-based and institutionally constrained decision-making
The effects of AI use depend on system design, use intensity and frequency, user autonomy, relational context, and the time horizon over which outcomes are evaluated.
The paper identifies these variables as determinants in its dynamic framework.
high mixed Can Artificial Intelligence Make Us Happier? Dynamic Capabil... AI-mediated well-being and accumulated capability, autonomy, relatedness, and de...
For finance forecasting and treasury, a modest improvement in forecast accuracy can create substantial value when it occurs near a liquidity threshold, whereas a larger statistical improvement may be irrelevant if it does not change a decision.
Decision-oriented synthesis of forecasting and treasury literature, citing decision-focused evaluation of predictive information.
high mixed Business Process Improvement in Finance Operations through P... Economic value of forecasts and resulting treasury or planning decisions
Counterfactual evaluation is necessary in collections analytics because observed payments may have occurred naturally and should not automatically be attributed to the predictive intervention.
Methodological recommendation in the order-to-cash discussion concerning intervention evaluation and suitable baselines.
high mixed Business Process Improvement in Finance Operations through P... Incremental effect of collections interventions on payment timing
Decision quality and process value should be evaluated separately from predictive accuracy because a technically accurate model may fail to change actions or improve financial, operational, control, or service outcomes.
Conceptual framework distinguishing predictive quality, decision quality, and process value; supported by examples involving collections and journal-anomaly detection.
high mixed Business Process Improvement in Finance Operations through P... Decision quality and downstream process outcomes