0 cumulative citations
View corpus contextLarge language models pick sector stocks that outperform in calm markets but stumble in volatility; integrating LLM stock selection with classical portfolio optimization yields more reliable returns.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
This paper investigates how Large Language Models (LLMs) from leading providers (OpenAI, Google, Anthropic, DeepSeek, and xAI) can be applied to quantitative sector-based portfolio construction. We use LLMs to identify investable universes of stocks within S&P 500 sector indices and evaluate how their selections perform when combined with classical portfolio optimization methods. Each model was prompted to select and weight 20 stocks per sector, and the resulting portfolios were compared with their respective sector indices across two distinct out-of-sample periods: a stable market phase (January-March 2025) and a volatile phase (April-June 2025). Our results reveal a strong temporal dependence in LLM portfolio performance. During stable market conditions, LLM-weighted portfolios frequently outperformed sector indices on both cumulative return and risk-adjusted (Sharpe ratio) measures. However, during the volatile period, many LLM portfolios underperformed, suggesting that current models may struggle to adapt to regime shifts or high-volatility environments underrepresented in their training data. Importantly, when LLM-based stock selection is combined with traditional optimization techniques, portfolio outcomes improve in both performance and consistency. This study contributes one of the first multi-model, cross-provider evaluations of generative AI algorithms in investment management. It highlights that while LLMs can effectively complement quantitative finance by enhancing stock selection and interpretability, their reliability remains market-dependent. The findings underscore the potential of hybrid AI-quantitative frameworks, integrating LLM reasoning with established optimization techniques, to produce more robust and adaptive investment strategies.
Summary
Main Finding
LLMs from multiple providers can meaningfully enhance sector-level stock selection and — when combined with classical optimization — produce portfolios that outperform sector benchmarks in stable market conditions. However, their performance is temporally dependent: many LLM-driven portfolios underperform during high-volatility regime shifts. Hybrid approaches that use LLMs for selection and quantitative optimization for weighting produce the most consistent and robust results.
Key Points
- Multi-model, cross-provider evaluation: 12 LLMs from OpenAI, Google, Anthropic, DeepSeek, and xAI were tested on the same sector-selection task.
- Sector focus: The study operates at the S&P 500 sector level (11 sectors) instead of market-wide stock selection.
- Prompting protocol: Each model was prompted to select and (initially) weight at least 20 tickers per sector intended to outperform the respective sector index.
- Portfolio types compared (per model × sector): LLM-weighted, equally weighted, minimum-variance, maximum-expected-return, and maximum-Sharpe portfolios built from the LLM-selected universe.
- Out-of-sample evaluation: Two distinct OOS windows — a stable market (Jan–Mar 2025) and a volatile market (Apr–Jun 2025) — to test regime sensitivity.
- Performance metrics: Cumulative return and risk-adjusted performance (Sharpe ratio); statistical analysis linked sector outcomes mainly to volatility rather than diversification.
- Results summary:
- Stable period: Many LLM-weighted portfolios outperformed their sector indices in returns and Sharpe.
- Volatile period: Many LLM portfolios underperformed, suggesting limited adaptability to sudden regime changes or underrepresented volatility in training data.
- Hybridization: Applying classical mean–variance optimization to the LLM-selected universes improved both performance and consistency across regimes.
- Sector heterogeneity: Energy, Financials, and Information Technology sectors tended to yield better LLM-driven results; Consumer Staples tended to lag.
Data & Methods
- Stock universe construction:
- Reconstructed the 11 S&P 500 sector constituent lists using GICS classification and public sources (S&P methodology, S&P Global counts, full S&P 500 lists from Wikipedia, TradingView checks).
- Adjusted invalid LLM outputs to ensure full allocations matched legitimate sector constituents.
- Market data:
- Adjusted closing prices from Yahoo Finance for individual stocks and sector indices (note: Energy listed as GSPE on Yahoo).
- Models evaluated (examples and versions reported):
- OpenAI: GPT-4o, GPT-4.1 (gpt-4.1-2025-04-14), o4-mini, GPT-5 (knowledge cutoff Sept 30, 2024).
- Google: Gemini 2.5 Pro (cutoff Jan 2025).
- Anthropic: Claude Sonnet 3.7, Claude Sonnet 4, Claude Opus 4.
- xAI: Grok 3, Grok 3 Mini.
- DeepSeek: DeepSeek-V3 (deepseek-chat), DeepSeek-R1 (deepseek-reasoner).
- Prompting and selection:
- Standardized, sector-specific prompts requesting at least 20 tickers and weights aimed to outperform the sector index.
- Repeated prompting and verification to mitigate variance/hallucination; most-frequent picks concept noted in prior related work.
- Optimization:
- Mean–variance optimization used to compute minimum-variance, max-return, and max-Sharpe portfolios using the LLM-selected universes.
- Evaluation:
- In-sample: 5 years prior to Jan 2025 for model inputs/selection context.
- Out-of-sample: Jan–Mar 2025 (stable) and Apr–Jun 2025 (volatile).
- Performance comparisons versus each sector index and across portfolio construction methods.
- Limitations noted in methods:
- Some model training windows overlap with early 2025 data (e.g., Claude 4), potentially giving limited exposure to OOS conditions.
- No detailed discussion in the excerpt about transaction costs, turnover, liquidity constraints, or execution slippage.
Implications for AI Economics
- Practical asset management:
- LLMs can be useful idea generators for stock selection and for producing human-interpretable rationale, helping research teams expand coverage at low marginal cost.
- Best practice appears to be hybridization: use generative models for candidate generation and qualitative reasoning, then apply quantitative optimization, risk controls, and execution-aware adjustments.
- Model risk & regime sensitivity:
- Temporal dependence implies model risk is substantial: LLMs trained on historical or pre-2025 data may underperform when market structure or volatility regimes shift.
- Firms should stress-test LLM outputs across regimes, incorporate monitoring for regime shifts, and avoid overreliance on LLM recommendations without quantitative overlays.
- Need for domain specialization and updating:
- Domain-specific LLMs (e.g., BloombergGPT) or models regularly fine-tuned on up-to-date market data are likely to be more robust, especially for high-volatility regimes.
- Real-time data access and continual fine-tuning or ensembling can reduce lag between model knowledge and market reality.
- Explainability and governance:
- The black-box nature and inconsistency of LLMs raise governance, compliance, and auditability concerns — important for regulated financial institutions.
- Hybrid designs increase interpretability by coupling LLM rationale with transparent optimization math, improving acceptability to stakeholders and regulators.
- Product and market design:
- Potential for new AI-enhanced sector products (research services, ETF selection aids, bespoke sector baskets) but deployment must consider operational costs, transaction frictions, and reputational risk.
- Research directions:
- Future work should quantify turnover, transaction costs, and implementation shortfall, test robustness across more regimes and time-horizons, compare domain-specific vs general LLMs, and explore ensembles/consensus approaches.
- Investigate mechanisms to reduce hallucination (structured outputs, retrieval-augmented generation, fact-checking layers) and formalize model-confidence signals for portfolio automation decisions.
Overall, this paper provides evidence that LLMs have practical value in sector-level portfolio construction but are best used within hybrid frameworks that mitigate temporal and model risks.
Assessment
Claims (7)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| We applied LLMs from OpenAI, Google, Anthropic, DeepSeek, and xAI to quantitative sector-based portfolio construction by prompting each model to select and weight 20 stocks per S&P 500 sector. Other | positive | LLM-based stock selection and weighting procedure (selection coverage across sectors) |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Portfolios constructed from LLM-weighted stock selections frequently outperformed their respective sector indices on cumulative return and risk-adjusted (Sharpe ratio) measures during a stable market period (January–March 2025). Other | positive | cumulative return; Sharpe ratio (risk-adjusted return) |
Reading fidelity
high
Study strength
medium
|
not reported
|
| During a volatile market period (April–June 2025), many LLM-based portfolios underperformed their sector indices, indicating temporal dependence of LLM portfolio performance. Other | negative | relative performance vs. sector indices (cumulative return and Sharpe ratio implied) |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The observed underperformance in volatile markets suggests current LLMs may struggle to adapt to regime shifts or high-volatility environments that are underrepresented in their training data. Other | negative | model adaptability to market regime shifts / robustness under volatility |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Combining LLM-based stock selection with traditional portfolio optimization techniques improves portfolio outcomes in both performance and consistency. Other | positive | portfolio performance and consistency (risk/return metrics implied) |
Reading fidelity
high
Study strength
medium
|
not reported
|
| This study provides one of the first multi-model, cross-provider evaluations of generative AI algorithms in investment management. Other | positive | novelty / breadth of multi-provider evaluation |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| LLMs can effectively complement quantitative finance by enhancing stock selection and interpretability, but their reliability remains market-dependent. Other | mixed | effectiveness of LLMs as complements to quantitative methods (stock selection quality and interpretability) and market-dependent reliability |
Reading fidelity
high
Study strength
medium
|
not reported
|