The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Two large language models diverge: DeepSeek-V3 is more accurate and stable at predicting firms' earnings direction, while GPT-4 gives superior explanatory insight and an edge during unstable recovery periods; combining both approaches appears to offer the best decision‑support trade-off.

COMPARING CHATGPT AND DEEPSEEK IN FINANCIAL FORECASTING: CROSS-REGIONAL EVIDENCE FROM TURKEY, EUROPE, AND THE UNITED STATES
Milad Stanikzai, Ali Bayrakdaroğlu · January 01, 2026 · EMC Review - Časopis za ekonomiju - APEIRON
openalex descriptive low evidence 7/10 relevance Full text usable extracted full text DOI Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

OpenAlex

Latest observation:

  1. Milad Stanikzai provider ID
  2. Ali Bayrakdaroğlu provider ID

Semantic Scholar

Latest observation:

  1. Milad Stanikzai provider ID
  2. Ali Bayrakdaroğlu provider ID
In a 2022–2024 panel of 15 firms, DeepSeek-V3 produced more accurate and stable earnings-direction forecasts while GPT-4 delivered richer qualitative explanations and stronger performance in unstable, recovery-driven scenarios, suggesting hybrid systems may best trade off reliability and interpretability.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

This research compares the forecasting ability of two latest large language models ChatGPT (GPT-4) and DeepSeek-V3 on predicting corporate direction of earnings based on historical financial reports alone. The panel data of 15 Turkish, European, and American companies operating in the banking, technology, and industrial sectors for 2022-2024 have been considered. Projections were made with a consistent prompt style and evaluated for direction accuracy, confusion matrices, interpretive insight scores, and sector-region performance.Findings suggest that while ChatGPT offers greater qualitative explanation and performs better in unstable, recovery-based scenarios, DeepSeek offers greater accuracy, stability, and performance in stable, regulation-based environments. The paper brings into focus critical trade-offs between interpretability and reliability and suggests that the use of hybrid models can offer the best solution to financial forecasting.These findings oppose Efficient Market Hypothesis assumptions and highlight the need for decision-support systems that integrate AI interpretability, domain expertise, and responsible model selection. The research offers prescriptive advice to analysts, institutions, and developers who aim to incorporate LLMs into financial decisions.

Summary

Main Finding

When given five years of historical financial statements (2017–2021) to predict earnings direction for 2022–2024, GPT‑4 (ChatGPT) produced richer, more interpretable justifications and performed relatively better in unstable or recovery scenarios, while DeepSeek‑V3 produced more stable and more accurate directional forecasts in stable, regulation‑driven environments. The paper emphasizes a trade‑off between interpretability (ChatGPT) and reliability/stability (DeepSeek) and recommends hybrid approaches for practical decision support. These results challenge strict forms of the Efficient Market Hypothesis by showing systematic, model-dependent predictive differences from identical inputs.

Key Points

  • Models compared: OpenAI GPT‑4 (ChatGPT) vs DeepSeek‑V3‑Chat.
  • Task: directional earnings forecast (Increase / Decrease / Stable) plus approximate % change and a short justification, using only audited financial statements (2017–2021).
  • Sample: 15 publicly listed firms across three regions (Turkey, Europe, USA) and three sectors (banking, technology, industrial). Examples: Apple, Amazon, JPMorgan, BNP Paribas, Nestlé, Volkswagen, Arçelik, Tüpraş, Garanti BBVA, Ziraat, Halkbank.
  • Input features: Total Revenue, EPS, Net Income, Total Assets, Total Liabilities, Shareholders’ Equity, EBIT, Current Ratio (provided in Excel).
  • Currency handling: analyses done in firms’ local reporting currencies (no cross‑firm currency conversion); forecasts used intra‑firm year‑over‑year percentage changes to avoid nominal distortion.
  • Prompt engineering: standardized template, chain‑of‑thought (CoT) prompting, constrained to financial statement data only, structured tabular output required. Separate sessions per model-company-year.
  • Evaluation metrics: direction accuracy, confusion matrices, bias/consistency metrics, and an interpretive “insight quality” score (qualitative scoring of explanations).
  • Principal empirical finding: ChatGPT excels in generating plausible reasoning and handles volatile/recovery cases better; DeepSeek is more accurate and consistent in well‑regulated, stable environments (notably some banking contexts).
  • Interpretation: results highlight model‑specific heuristic choices and training biases; neither model dominates across all sectors/regions.
  • Recommendation: hybrid systems that combine an interpretable LLM for rationale and a sturdier/statistically calibrated model for final directional decision provide improved decision support.

Data & Methods

  • Design: comparative longitudinal forecasting exercise. Models used only firm financial statements (2017–2021) to predict direction of earnings for 2022–2024.
  • Sample selection criteria: asset size, exchange listing, data availability; 15 firms across Turkey, Europe, USA in banking, technology, industrial/energy.
  • Financial statements used: Income Statement, Balance Sheet, Cash Flow, Statement of Changes in Equity. Data sources: KAP (Turkey), official annual reports (Europe, IFRS), investor relations/AnnualReports.com (USA, US GAAP).
  • Feature engineering / derived metrics:
    • EPS (IAS 33 / TFRS 33 methods when not directly reported).
    • EBIT: computed differently for corporates (operating income + interest expense) and banks (approximated as net interest income + non‑interest income − operating expenses or operating profit before provisions & taxes).
    • Current ratio and standard balance‑sheet aggregates derived per Penman/Damodaran conventions.
  • Prompting protocol:
    • Unified “forget previous instructions” template making the model act as a financial expert.
    • Explicit constraints: use only provided financials; produce year‑by‑year direction, % estimate, and ≤3 sentence justification.
    • Chain‑of‑thought encouraged to increase reasoning traceability.
  • Model execution: identical inputs and prompt format supplied in separate sessions to GPT‑4 and DeepSeek‑V3.
  • Evaluation:
    • Directional predictions compared to realized outcomes for 2022–2024.
    • Metrics: accuracy, class confusion matrices (false positives/negatives per class), consistency over years, bias tendencies, and a qualitative interpretive score for explanations.
  • Limitations acknowledged by authors:
    • Small sample (15 firms) limiting generalizability.
    • Exclusion of macro/market sentiment and non‑financial textual sources.
    • No currency normalization across firms limits cross‑firm magnitude comparison (authors argue intra‑firm % changes mitigate this).
    • Dependence on prompt design, model versions, and possible hallucinations or table‑reading errors.
    • Tests are retrospective; not a live trading or real‑time decision experiment.

Implications for AI Economics

  • For theory: systematic, model‑dependent predictive differences using identical inputs present an empirical challenge to strict EMH claims that all useful information is already fully and equivalently priced; model processing matters.
  • For practice (analysts / institutions):
    • Match model choice to environment: use more interpretable LLMs for volatile/recovery situations where nuanced reasoning aids judgment; prefer more stable/accurate models (or calibrated ensembling) in regulated, stable sectors (some banking contexts).
    • Adopt hybrid decision‑support workflows: combine LLM rationale (for auditability and human review) with a calibrated predictive engine for final directional calls.
    • Insist on prompt standardization, CoT-style explanations, and version control for reproducibility and audit trails.
  • For model developers:
    • Improve sectoral calibration (accounting standard diversity, bank vs corporate accounting), table understanding, and robustness to input formatting.
    • Provide tools for quantitative calibration of LLM probability outputs (to reduce over/underconfidence) and for structured explainability metrics to support auditability in finance.
  • For regulators and auditors:
    • Require transparency around prompt and model versions used in financial decision contexts; require backtesting and documentation of model performance across regions/sectors.
    • Encourage standards for interpretability scoring and stress‑testing LLMs on accounting standards heterogeneity (IFRS vs US GAAP vs local TFRS).
  • For research:
    • Scale up sample sizes, include macro/market sentiment and textual disclosures, and benchmark against human analysts and specialized predictive ML models.
    • Explore formal ensemble/hybrid strategies and calibration methods that reconcile interpretability and directional accuracy.
    • Test real‑time deployment, economic value (e.g., trading or risk management performance), and fairness/ethical considerations in automated forecasting.

Overall, the paper provides a focused, controlled comparison that surfaces meaningful trade‑offs (interpretability vs reliability) when using general‑purpose LLMs for financial forecasting and makes a strong case for hybrid, audited decision‑support systems tailored by sector and regime.

Assessment

Paper Typedescriptive Evidence Strengthlow — The study uses a very small convenience panel (15 firms) over a short timeframe (2022–2024), lacks causal identification, has limited out-of-sample validation, and may be sensitive to prompt wording, model versions, and selection of firms/sectors, so results are suggestive but not robustly generalizable. Methods Rigormedium — The authors use a consistent prompting protocol and multiple evaluation metrics (direction accuracy, confusion matrices, interpretive scores, sector-region breakdowns), which are appropriate for model comparison; however, the small sample, lack of statistical testing details, no pre-registration or benchmark baselines (e.g., analyst/market forecasts), and limited robustness checks weaken methodological rigor. SamplePanel of 15 publicly listed firms across Turkey, Europe, and the U.S. in banking, technology, and industrial sectors, using only historical financial reports from 2022–2024 as input; models (GPT-4/ChatGPT and DeepSeek-V3) produced earnings-direction forecasts that were evaluated on direction accuracy, confusion matrices, interpretive insight scores, and sector-region performance. Themeshuman_ai_collab adoption productivity governance GeneralizabilitySmall sample (15 firms) limits representativeness, Only three sectors studied (banking, technology, industrial); excludes other industries, Geographic coverage limited to Turkey, Europe, and the U.S.; results may not hold in other markets, Short time window (2022–2024) that includes unusual macro events, Findings may be sensitive to prompt design and proprietary model versions, Only direction of earnings forecasted (not magnitude) and only from financial reports—omits market and alternative data, No assessment of economic impact (e.g., trading performance or firm decisions)

Claims (10)

ClaimDirectionOutcomeConfidence & EvidenceDetails
The study uses panel data of 15 Turkish, European, and American companies operating in the banking, technology, and industrial sectors for 2022-2024. Other null_result dataset composition (firms, sectors, regions, years)
Reading fidelity high
Study strength high
n=15
0.3
Projections from both models were produced using a consistent prompt style and evaluated using direction accuracy, confusion matrices, interpretive insight scores, and sector-region performance comparisons. Decision Quality null_result forecast evaluation metrics (direction accuracy, confusion matrices, interpretive insight scores, sector-region performance)
Reading fidelity high
Study strength high
n=15
0.3
ChatGPT (GPT-4) provides greater qualitative explanation (stronger interpretive insight scores) than DeepSeek-V3. Ai Safety And Ethics positive quality of qualitative explanations / interpretive insight scores
Reading fidelity medium
Study strength medium
n=15
0.11
ChatGPT performs better in unstable, recovery-based forecasting scenarios. Decision Quality positive forecast direction accuracy in unstable/recovery scenarios
Reading fidelity medium
Study strength medium
n=15
0.11
DeepSeek-V3 delivers greater accuracy, stability, and overall forecasting performance in stable, regulation-based environments. Decision Quality positive forecast accuracy and stability in stable/regulation-based scenarios
Reading fidelity medium
Study strength medium
n=15
0.11
The results highlight a critical trade-off between interpretability (qualitative explanation) and reliability (forecasting accuracy/stability). Ai Safety And Ethics mixed trade-off between interpretability and reliability of model outputs
Reading fidelity high
Study strength medium
n=15
0.18
Hybrid models (combining strengths of models like ChatGPT and DeepSeek) can offer the best solution to financial forecasting. Organizational Efficiency positive proposed performance improvements from model hybridization
Reading fidelity high
Study strength speculative
not reported
0.03
These findings challenge (oppose) the assumptions of the Efficient Market Hypothesis by demonstrating that LLMs can extract predictive signal from historical financial reports. Market Structure negative ability of LLMs to predict earnings direction (implication for market efficiency)
Reading fidelity medium
Study strength low
n=15
0.05
The paper recommends decision-support systems that integrate AI interpretability, domain expertise, and responsible model selection when incorporating LLMs into financial decision-making. Governance And Regulation positive recommended system features (interpretability, domain expertise, model selection)
Reading fidelity high
Study strength speculative
not reported
0.03
The research offers prescriptive advice to analysts, institutions, and developers aiming to incorporate LLMs into financial decisions. Governance And Regulation positive practical recommendations/advice uptake
Reading fidelity high
Study strength speculative
not reported
0.03

Notes