The Commonplace
Home Three-study pilot Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

AI has improved forecasting and investment workflows, but public evidence to 31 Aug 2026 shows no general AI method that reliably produces sustained, capacity-aware net returns across listed equities and crypto; stronger point-in-time, execution-aware, and prospective evaluations are required before profitability claims can be accepted.

Artificial Intelligence in Equity and Crypto Markets: Progress, Profitability Evidence, and the Limits of Automated Investing
Linsen Zhu, Mengqing Cai · September 04, 2026
arxiv review_meta medium evidence 8/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Linsen Zhu unresolved corpus identity
  2. Mengqing Cai unresolved corpus identity
This critical review finds that AI has driven meaningful upstream advances in prediction, representation and workflow integration for equities and crypto, but the public evidence does not show any general AI architecture producing persistent, capacity-aware net alpha after realistic implementation costs and across regimes.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Artificial intelligence (AI) now supports investment workflows from data and prediction through research, portfolios, execution, and tool use. Technical capability, however, is not evidence of investment profitability. This critical state-of-the-art review examines public research available through 31 August 2026 on listed equities, exchange-traded funds, centralized crypto spot, perpetual futures, and on-chain markets. We organize evidence with an alpha-translation chain: point-in-time information must yield a stable signal, feasible positions, executable orders, and risk-adjusted returns after costs. Across machine learning, time-series foundation models, financial language models, reinforcement learning, and agents, the examined record shows real but mainly upstream progress in prediction, text processing, portfolio design, and workflow integration. Evidence is thinner for durable net performance. Temporal contamination, repeated selection, survivorship, weak benchmarks, implementation costs, venue mechanics, and capacity can break translation to net alpha. Strong historical results coexist with predictor decay, corrected look-ahead failures, mixed prospective evidence, and few audited live-capital records. Crypto adds informative state but requires separate treatment of spot, perpetual, and decentralized cash flows and execution. Within the public evidence examined here, no general AI architecture is shown to deliver persistent, cross-regime, capacity-aware net alpha. More credible claims require point-in-time data and models, decision-aligned objectives, joint portfolio--execution evaluation, controlled adaptation, prospective tests, and authority-matched governance. These conditions can improve evidence and implementation; they do not guarantee profit.

Summary

Main Finding

Public research through 31 August 2026 shows clear technical progress in AI components of investment workflows (prediction, representation, NLP, agents, portfolio design, execution tooling), but that progress has not translated into convincing public evidence of persistent, cross‑regime, capacity‑aware net alpha. Many promising historical results weaken or vanish once realistic temporality, selection, implementation costs, venue mechanics and capacity are accounted for. The literature supports meaningful upstream gains but not a general, portable claim that any AI architecture reliably produces sustainable net alpha.

Key Points

  • Scope and aim

    • Critical state-of-the-art review (not a formal systematic review) of public work on listed equities/ETFs, centralized crypto spot, perpetual futures, and on‑chain/DEX markets up to 2026‑08‑31.
    • Focus is economic validation: whether model outputs transduce into implementable, persistent, risk‑adjusted excess returns after realistic costs.
  • Alpha‑translation chain (conceptual core)

    • Investment value requires a chain: point‑in‑time information → learned representation/signal → decision rule/portfolio → executable orders/fills → realized returns after fees, impact, funding, venue losses → persistence at scale.
    • Failures can occur at any stage (e.g., look‑ahead leakage, selection/multiple‑testing, infeasible portfolios, latency/impact, venue‑specific frictions, crowding).
  • Evidence framework

    • Unit of evidence is a precisely bounded claim; seven evidentiary dimensions: Temporality (T), Selection control (S), Portfolio mapping (P), Implementation realism (I), Risk & benchmark (R), External validity (X), Operational provenance (O).
    • Common study archetypes: capability descriptions (CAP), task benchmarks (TASK), time‑valid statistical tests (OOS), gross simulations (GROSS), net/risk‑adjusted simulations (NET), precommitted paper/prospective portfolios (PROSP), audited live‑capital records (LIVE), independent cross‑regime/capacity studies (EXT).
  • What the literature shows

    • Real technical advances: cross‑sectional ML, time‑series foundation models, financial LLMs and NLP for disclosure and news, RL and cost‑aware portfolio learning, multimodal agents integrating analysis and execution.
    • Strong historical or task‑metric results exist in many studies (nonlinear predictors, decision‑focused policies), but important fragilities:
    • Temporal contamination / look‑ahead leaks (e.g., pretraining memory in LLMs can masquerade as forecasting).
    • Selection bias and multiple‑testing (search over many features/architectures).
    • Survivorship and data‑revision issues.
    • Weak or inappropriate benchmarks / risk adjustments.
    • Unrealistic or omitted implementation costs (impact, funding, gas, MEV, borrow constraints).
    • Capacity and crowding: high nominal returns concentrated in low‑capacity securities or short windows.
    • Predictor decay and lack of prospective/live evidence; few audited live‑capital records.
    • Crypto requires separate treatment by instrument:
    • Centralized spot: fragmented venues, custody and data‑integrity risks.
    • Perpetual futures: funding payments, mark price dynamics, liquidation risk, leverage constraints.
    • On‑chain/DEX: auditability of state but unique frictions (gas auctions, failed tx, MEV, AMM impermanent loss).
  • Recommendations distilled from the review

    • Use point‑in‑time data and model checkpoints.
    • Align objectives to economic outcomes (decision‑aligned loss functions, capacity‑aware objectives).
    • Evaluate signals jointly with portfolio mapping and execution (cost‑aware learning).
    • Conduct controlled, precommitted prospective tests; publish operational provenance where possible.
    • Implement governance matched to trading authority (monitoring, rollback, documentation).
    • Treat productivity gains (automation, faster research) separately from alpha claims.

Data & Methods

  • Literature search and inclusion

    • Iterative search across finance, economics, ML, NLP and market‑microstructure venues, following citations and using publisher / vendor materials for operational claims.
    • Prioritize peer‑reviewed work and dated preprints (arXiv) for fast‑moving topics. Exclude unrelated AI uses (credit, fraud) unless directly tied to traded‑asset decisions.
    • Not database‑complete; this is a critical synthesis, not a registered systematic review or meta‑analysis.
  • Evidence coding and interpretation

    • Each claim assessed on the seven evidence dimensions (T,S,P,I,R,X,O). Absence = unknown (not assumed favorable).
    • Evaluation archetypes (CAP, TASK, OOS, GROSS, NET, PROSP, LIVE, EXT) used to identify what a study actually supports and what remains unresolved.
    • Emphasis on temporality: true out‑of‑sample requires that decision‑time information exclude any post‑decision data (explicit condition It ∩ Ft+1 = ∅).
  • Examples of methodological pitfalls documented

    • Look‑ahead / data leakage: models trained with future‑revised inputs or contaminated pretraining.
    • Cost modeling as a simplification: linear per‑share costs understate nonlinear impact and capacity limits.
    • Multiple‑testing: many hyperparameters/architectures/features raise the discovery threshold.
    • LLM memorization: pretraining can encode past facts that resemble forecasts unless controlled.

Implications for AI Economics

  • For researchers

    • Move beyond task metrics and historical backtests: build evaluation designs that report all seven evidence dimensions and include prospective or live tests where feasible.
    • Incorporate costs, venue mechanics and capacity into model training (decision‑aligned objectives, cost‑aware RL or portfolio learning).
    • Provide point‑in‑time checkpoints, reproducible pipelines, and explicit selection budgets to reduce p‑hacking and leakage.
    • Study the dynamics of diffusion and crowding: how AI adoption changes signal returns and market impact over time.
  • For asset managers and practitioners

    • Do not equate technical capability with deployable alpha. Demand transparent evidence: point‑in‑time data, costed net simulations, prospective or audited live records, and clearly stated capacity.
    • Treat AI as an organizing layer that can increase productivity and risk‑control even if incremental alpha is limited.
    • Prioritize governance, monitoring, and rollback mechanisms for agentic systems that have delegated trading authority.
  • For policy makers and market operators

    • Encourage or require disclosures that improve operational provenance (timestamped model checkpoints, execution records) for public claims about trading performance.
    • Recognize instrument‑specific frictions: regulation and surveillance should account for crypto‑native risks (MEV, failed tx, custody).
    • Consider how diffusion of AI trading affects market stability, liquidity provision and concentration of risk.
  • For theories of market efficiency and asset pricing

    • Evidence is mixed: AI has improved signal extraction but many signals are fragile once economic translation is fully modeled. This supports a nuanced view where markets remain contestable — some predictability is exploitable, but implementation frictions, capacity limits and adaptive competition often attenuate net alpha.
    • Research should explicitly model the equilibrium feedback between AI adoption and signal returns (endogenous capacity and impact).
  • Practical checklist suggested by the review (for credible AI‑alpha claims)

    • Provide point‑in‑time data and model artifacts.
    • Report selection/search budget and sensitivity to alternative choices.
    • Map signals to implementable portfolios with explicit leverage/borrow/turnover constraints.
    • Simulate or measure realistic implementation costs (impact, funding, gas, MEV).
    • Benchmark with stated risk model and report uncertainty/serial correlation.
    • Show external validity (different regimes, scales) or a precommitted prospective test.
    • Document operational provenance (paper/live portfolio, audit trail).

Caveat: the review finds durable public evidence lacking for a general AI architecture that consistently yields sustainable net alpha up to the literature cutoff. This conclusion does not deny profitable proprietary systems may exist; it highlights the gap between upstream technical gains and the higher evidentiary bar needed to claim deployable, persistent, scaleable alpha.

Assessment

Paper Typereview_meta Evidence Strengthmedium — This is a careful, critical synthesis of public empirical and methodological literature through 31 Aug 2026 and it anchors conclusions to primary studies, but it is not a new causal identification study nor a systematic, bias-assessed meta-analysis; conclusions therefore rest on heterogeneous published evidence and selective, though transparent, choice of sources. Methods Rigormedium — The paper proposes a clear, rigorous evaluative framework (alpha-translation chain; seven-dimension evidence profile), applies it across asset strata, and documents inclusion/exclusion rules; however it explicitly forgoes a registered protocol, exhaustive systematic retrieval, duplicate screening, and formal risk-of-bias scoring, which reduces methodological completeness relative to a full systematic review. SampleA critical state-of-the-art review of public research and official material available through 31 August 2026, covering studies and documentation on listed equities, ETFs, centralized crypto spot markets, crypto perpetual futures, and on-chain (DEX) markets; prioritized peer-reviewed articles and official proceedings, plus dated working papers and arXiv versions and vendor documentation. Themesinnovation adoption GeneralizabilityLimited to public, documented evidence available by 31 Aug 2026 (may omit proprietary, confidential, or later results)., Excludes many asset classes (fixed income, FX, options, commodities, private markets) and so findings do not generalize to those markets., Heterogeneity in study designs, universes, benchmarks, and cost assumptions limits pooled inference., Conclusions about 'no general AI architecture delivering persistent net alpha' apply to the public record examined and do not preclude specific proprietary implementations., Time-bounded: rapid methodological or market changes after the cut-off could alter conclusions.

Claims (12)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Within the public evidence examined through 31 August 2026, no general AI architecture, foundation model, or agent architecture is shown to deliver persistent, cross-regime, capacity-aware net alpha. Other null_result Persistent, cross-regime, capacity-aware risk-adjusted investment performance after implementation costs
Reading fidelity high
Study strength medium
not reported
0.24
The reviewed literature shows meaningful technical progress in representation, prediction, textual information processing, portfolio design, and operational integration, but evidence for durable net investment performance is thinner. Other mixed Technical investment-workflow capabilities and durable net investment performance
Reading fidelity high
Study strength medium
not reported
0.24
A look-ahead correction can eliminate a previously reported machine-learning alpha. Other negative Reported machine-learning return advantage or alpha after correcting information-timing leakage
Reading fidelity high
Study strength medium
not reported
0.24
Large language models can recall financial facts from before their pretraining cutoff, making apparent forecasts potentially reflect memorization rather than genuine forecasting. Ai Safety And Ethics negative Validity of LLM-based financial forecasting under point-in-time information constraints
Reading fidelity high
Study strength medium
not reported
0.24
Cost-unaware return predictors can favor fleeting, low-scale opportunities and produce unattractive implementable portfolios, whereas incorporating trading costs into the learning objective changes which information is valuable. Task Allocation negative Implementability and portfolio frontier quality of learned investment strategies after trading costs
Reading fidelity high
Study strength medium
not reported
0.24
Prospective AI-managed household portfolios can become concentrated without generating statistically significant abnormal returns. Other mixed Portfolio concentration and statistically significant abnormal returns in prospective AI-managed household portfolios
Reading fidelity high
Study strength medium
statistically significant abnormal returns: not detected
0.24
Early relative performance of AI-labelled hedge funds can decay as the technology diffuses. Other negative Relative performance persistence of AI-labelled hedge funds
Reading fidelity high
Study strength medium
not reported
0.24
In Hou et al. (2020), a majority of 452 anomalies failed conventional statistical significance in the authors' reconstruction, and an even larger share failed a stricter multiple-testing-aware threshold. Other negative Statistical significance and robustness of published return anomalies
Reading fidelity high
Study strength high
n=452
majority of 452 anomalies failed conventional significance; an even larger share failed the higher threshold
0.4
Chen and Zimmermann (2022) reproduced almost all of 319 published characteristics under a transparent common code base. Other positive Replication of published asset-pricing characteristics
Reading fidelity high
Study strength high
n=319
almost all of 319 characteristics reproduced
0.4
High-turnover strategies are disproportionately eroded by trading costs, and trading rules can materially change implementability. Other negative Net strategy performance and implementability after trading costs
Reading fidelity high
Study strength high
not reported
0.4
For some investment styles, institutional transaction data indicate lower trading costs and higher capacity than several pessimistic trading-cost models imply. Other positive Trading costs and strategy capacity
Reading fidelity high
Study strength medium
costs lower and capacity higher than several pessimistic estimates
0.24
The AI-specific crypto evidence base is thinner than the equity evidence base. Other negative Depth and strength of publicly available evidence for AI-enabled investment performance
Reading fidelity high
Study strength medium
not reported
0.24

Notes