The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Large language models pick sector stocks that outperform in calm markets but stumble in volatility; integrating LLM stock selection with classical portfolio optimization yields more reliable returns.

Generative AI-enhanced Sector-based Investment Portfolio Construction
Alina Voronina, Oleksandr Romanko, Ruiwen Cao, Roy H. Kwon, Rafael Mendoza-Arriaga · December 31, 2025
arxiv descriptive medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Alina Voronina unresolved corpus identity
  2. Oleksandr Romanko unresolved corpus identity
  3. Ruiwen Cao unresolved corpus identity
  4. Roy H. Kwon unresolved corpus identity
  5. Rafael Mendoza-Arriaga unresolved corpus identity

Semantic Scholar

Latest observation:

  1. A. Voronina provider ID
  2. O. Romanko provider ID
  3. Rui Cao provider ID
  4. R. Kwon provider ID
  5. Rafael Mendoza-Arriaga provider ID
LLMs can enhance sector equity selection and beat sector indices in stable markets, but many models underperform during volatile regimes, while combining LLM selection with traditional optimization improves consistency and performance.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

This paper investigates how Large Language Models (LLMs) from leading providers (OpenAI, Google, Anthropic, DeepSeek, and xAI) can be applied to quantitative sector-based portfolio construction. We use LLMs to identify investable universes of stocks within S&P 500 sector indices and evaluate how their selections perform when combined with classical portfolio optimization methods. Each model was prompted to select and weight 20 stocks per sector, and the resulting portfolios were compared with their respective sector indices across two distinct out-of-sample periods: a stable market phase (January-March 2025) and a volatile phase (April-June 2025). Our results reveal a strong temporal dependence in LLM portfolio performance. During stable market conditions, LLM-weighted portfolios frequently outperformed sector indices on both cumulative return and risk-adjusted (Sharpe ratio) measures. However, during the volatile period, many LLM portfolios underperformed, suggesting that current models may struggle to adapt to regime shifts or high-volatility environments underrepresented in their training data. Importantly, when LLM-based stock selection is combined with traditional optimization techniques, portfolio outcomes improve in both performance and consistency. This study contributes one of the first multi-model, cross-provider evaluations of generative AI algorithms in investment management. It highlights that while LLMs can effectively complement quantitative finance by enhancing stock selection and interpretability, their reliability remains market-dependent. The findings underscore the potential of hybrid AI-quantitative frameworks, integrating LLM reasoning with established optimization techniques, to produce more robust and adaptive investment strategies.

Summary

Main Finding

LLMs from multiple providers can meaningfully enhance sector-level stock selection and — when combined with classical optimization — produce portfolios that outperform sector benchmarks in stable market conditions. However, their performance is temporally dependent: many LLM-driven portfolios underperform during high-volatility regime shifts. Hybrid approaches that use LLMs for selection and quantitative optimization for weighting produce the most consistent and robust results.

Key Points

  • Multi-model, cross-provider evaluation: 12 LLMs from OpenAI, Google, Anthropic, DeepSeek, and xAI were tested on the same sector-selection task.
  • Sector focus: The study operates at the S&P 500 sector level (11 sectors) instead of market-wide stock selection.
  • Prompting protocol: Each model was prompted to select and (initially) weight at least 20 tickers per sector intended to outperform the respective sector index.
  • Portfolio types compared (per model × sector): LLM-weighted, equally weighted, minimum-variance, maximum-expected-return, and maximum-Sharpe portfolios built from the LLM-selected universe.
  • Out-of-sample evaluation: Two distinct OOS windows — a stable market (Jan–Mar 2025) and a volatile market (Apr–Jun 2025) — to test regime sensitivity.
  • Performance metrics: Cumulative return and risk-adjusted performance (Sharpe ratio); statistical analysis linked sector outcomes mainly to volatility rather than diversification.
  • Results summary:
    • Stable period: Many LLM-weighted portfolios outperformed their sector indices in returns and Sharpe.
    • Volatile period: Many LLM portfolios underperformed, suggesting limited adaptability to sudden regime changes or underrepresented volatility in training data.
    • Hybridization: Applying classical mean–variance optimization to the LLM-selected universes improved both performance and consistency across regimes.
  • Sector heterogeneity: Energy, Financials, and Information Technology sectors tended to yield better LLM-driven results; Consumer Staples tended to lag.

Data & Methods

  • Stock universe construction:
    • Reconstructed the 11 S&P 500 sector constituent lists using GICS classification and public sources (S&P methodology, S&P Global counts, full S&P 500 lists from Wikipedia, TradingView checks).
    • Adjusted invalid LLM outputs to ensure full allocations matched legitimate sector constituents.
  • Market data:
    • Adjusted closing prices from Yahoo Finance for individual stocks and sector indices (note: Energy listed as GSPE on Yahoo).
  • Models evaluated (examples and versions reported):
    • OpenAI: GPT-4o, GPT-4.1 (gpt-4.1-2025-04-14), o4-mini, GPT-5 (knowledge cutoff Sept 30, 2024).
    • Google: Gemini 2.5 Pro (cutoff Jan 2025).
    • Anthropic: Claude Sonnet 3.7, Claude Sonnet 4, Claude Opus 4.
    • xAI: Grok 3, Grok 3 Mini.
    • DeepSeek: DeepSeek-V3 (deepseek-chat), DeepSeek-R1 (deepseek-reasoner).
  • Prompting and selection:
    • Standardized, sector-specific prompts requesting at least 20 tickers and weights aimed to outperform the sector index.
    • Repeated prompting and verification to mitigate variance/hallucination; most-frequent picks concept noted in prior related work.
  • Optimization:
    • Mean–variance optimization used to compute minimum-variance, max-return, and max-Sharpe portfolios using the LLM-selected universes.
  • Evaluation:
    • In-sample: 5 years prior to Jan 2025 for model inputs/selection context.
    • Out-of-sample: Jan–Mar 2025 (stable) and Apr–Jun 2025 (volatile).
    • Performance comparisons versus each sector index and across portfolio construction methods.
  • Limitations noted in methods:
    • Some model training windows overlap with early 2025 data (e.g., Claude 4), potentially giving limited exposure to OOS conditions.
    • No detailed discussion in the excerpt about transaction costs, turnover, liquidity constraints, or execution slippage.

Implications for AI Economics

  • Practical asset management:
    • LLMs can be useful idea generators for stock selection and for producing human-interpretable rationale, helping research teams expand coverage at low marginal cost.
    • Best practice appears to be hybridization: use generative models for candidate generation and qualitative reasoning, then apply quantitative optimization, risk controls, and execution-aware adjustments.
  • Model risk & regime sensitivity:
    • Temporal dependence implies model risk is substantial: LLMs trained on historical or pre-2025 data may underperform when market structure or volatility regimes shift.
    • Firms should stress-test LLM outputs across regimes, incorporate monitoring for regime shifts, and avoid overreliance on LLM recommendations without quantitative overlays.
  • Need for domain specialization and updating:
    • Domain-specific LLMs (e.g., BloombergGPT) or models regularly fine-tuned on up-to-date market data are likely to be more robust, especially for high-volatility regimes.
    • Real-time data access and continual fine-tuning or ensembling can reduce lag between model knowledge and market reality.
  • Explainability and governance:
    • The black-box nature and inconsistency of LLMs raise governance, compliance, and auditability concerns — important for regulated financial institutions.
    • Hybrid designs increase interpretability by coupling LLM rationale with transparent optimization math, improving acceptability to stakeholders and regulators.
  • Product and market design:
    • Potential for new AI-enhanced sector products (research services, ETF selection aids, bespoke sector baskets) but deployment must consider operational costs, transaction frictions, and reputational risk.
  • Research directions:
    • Future work should quantify turnover, transaction costs, and implementation shortfall, test robustness across more regimes and time-horizons, compare domain-specific vs general LLMs, and explore ensembles/consensus approaches.
    • Investigate mechanisms to reduce hallucination (structured outputs, retrieval-augmented generation, fact-checking layers) and formalize model-confidence signals for portfolio automation decisions.

Overall, this paper provides evidence that LLMs have practical value in sector-level portfolio construction but are best used within hybrid frameworks that mitigate temporal and model risks.

Assessment

Paper Typedescriptive Evidence Strengthmedium — The study evaluates multiple LLMs out-of-sample across two distinct market regimes and compares model-based portfolios to benchmark sector indices, which provides useful empirical evidence; however, the evaluation is limited to two short 3-month windows, a fixed stock count per sector, likely lacks comprehensive statistical inference and robustness checks (e.g., sensitivity to prompts, multiple testing adjustments), and may omit realistic frictions like transaction costs and market impact, reducing confidence in broader causal or predictive claims. Methods Rigormedium — Rigor is strengthened by a multi-provider, cross-model design and the use of classical portfolio optimization to combine LLM selection with quantitative methods, but is weakened by short horizons, potential sample and selection biases, likely omission of turnover/transaction costs and significance testing, and insufficient reported robustness analyses (prompt engineering sensitivity, alternative weighting schemes, longer/rolling out-of-sample tests). SampleFor each S&P 500 sector index, five leading LLM providers (OpenAI, Google, Anthropic, DeepSeek, xAI) were prompted to select and weight 20 stocks per sector; resulting LLM-weighted portfolios were evaluated against their sector indices over two out-of-sample periods: Jan–Mar 2025 (stable) and Apr–Jun 2025 (volatile); classical portfolio optimization methods were also applied to combine LLM selection with quantitative weighting. Themesadoption innovation GeneralizabilityEvaluation limited to two short (3-month) out-of-sample periods, so results may not hold over longer horizons or other market cycles., Restricted to S&P 500 sectors and a fixed 20-stock-per-sector design—findings may not generalize to other universes, smaller-cap stocks, or other asset classes., Only a handful of LLM providers tested; model updates or other vendors could behave differently., Potential sensitivity to prompt wording and prompt-engineering choices not fully explored., Likely omission or limited treatment of transaction costs, turnover, slippage, and market impact which affect real-world implementability., Results may depend on the specific optimization methods used and hyperparameters, limiting portability to different portfolio construction pipelines.

Claims (7)

ClaimDirectionOutcomeConfidence & EvidenceDetails
We applied LLMs from OpenAI, Google, Anthropic, DeepSeek, and xAI to quantitative sector-based portfolio construction by prompting each model to select and weight 20 stocks per S&P 500 sector. Other positive LLM-based stock selection and weighting procedure (selection coverage across sectors)
Reading fidelity high
Study strength medium
not reported
0.18
Portfolios constructed from LLM-weighted stock selections frequently outperformed their respective sector indices on cumulative return and risk-adjusted (Sharpe ratio) measures during a stable market period (January–March 2025). Other positive cumulative return; Sharpe ratio (risk-adjusted return)
Reading fidelity high
Study strength medium
not reported
0.18
During a volatile market period (April–June 2025), many LLM-based portfolios underperformed their sector indices, indicating temporal dependence of LLM portfolio performance. Other negative relative performance vs. sector indices (cumulative return and Sharpe ratio implied)
Reading fidelity high
Study strength medium
not reported
0.18
The observed underperformance in volatile markets suggests current LLMs may struggle to adapt to regime shifts or high-volatility environments that are underrepresented in their training data. Other negative model adaptability to market regime shifts / robustness under volatility
Reading fidelity high
Study strength medium
not reported
0.18
Combining LLM-based stock selection with traditional portfolio optimization techniques improves portfolio outcomes in both performance and consistency. Other positive portfolio performance and consistency (risk/return metrics implied)
Reading fidelity high
Study strength medium
not reported
0.18
This study provides one of the first multi-model, cross-provider evaluations of generative AI algorithms in investment management. Other positive novelty / breadth of multi-provider evaluation
Reading fidelity high
Study strength speculative
not reported
0.03
LLMs can effectively complement quantitative finance by enhancing stock selection and interpretability, but their reliability remains market-dependent. Other mixed effectiveness of LLMs as complements to quantitative methods (stock selection quality and interpretability) and market-dependent reliability
Reading fidelity high
Study strength medium
not reported
0.18

Notes