0 cumulative citations
View corpus contextTwo large language models diverge: DeepSeek-V3 is more accurate and stable at predicting firms' earnings direction, while GPT-4 gives superior explanatory insight and an edge during unstable recovery periods; combining both approaches appears to offer the best decision‑support trade-off.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
0 cumulative citations
View corpus contextThis research compares the forecasting ability of two latest large language models ChatGPT (GPT-4) and DeepSeek-V3 on predicting corporate direction of earnings based on historical financial reports alone. The panel data of 15 Turkish, European, and American companies operating in the banking, technology, and industrial sectors for 2022-2024 have been considered. Projections were made with a consistent prompt style and evaluated for direction accuracy, confusion matrices, interpretive insight scores, and sector-region performance.Findings suggest that while ChatGPT offers greater qualitative explanation and performs better in unstable, recovery-based scenarios, DeepSeek offers greater accuracy, stability, and performance in stable, regulation-based environments. The paper brings into focus critical trade-offs between interpretability and reliability and suggests that the use of hybrid models can offer the best solution to financial forecasting.These findings oppose Efficient Market Hypothesis assumptions and highlight the need for decision-support systems that integrate AI interpretability, domain expertise, and responsible model selection. The research offers prescriptive advice to analysts, institutions, and developers who aim to incorporate LLMs into financial decisions.
Summary
Main Finding
When given five years of historical financial statements (2017–2021) to predict earnings direction for 2022–2024, GPT‑4 (ChatGPT) produced richer, more interpretable justifications and performed relatively better in unstable or recovery scenarios, while DeepSeek‑V3 produced more stable and more accurate directional forecasts in stable, regulation‑driven environments. The paper emphasizes a trade‑off between interpretability (ChatGPT) and reliability/stability (DeepSeek) and recommends hybrid approaches for practical decision support. These results challenge strict forms of the Efficient Market Hypothesis by showing systematic, model-dependent predictive differences from identical inputs.
Key Points
- Models compared: OpenAI GPT‑4 (ChatGPT) vs DeepSeek‑V3‑Chat.
- Task: directional earnings forecast (Increase / Decrease / Stable) plus approximate % change and a short justification, using only audited financial statements (2017–2021).
- Sample: 15 publicly listed firms across three regions (Turkey, Europe, USA) and three sectors (banking, technology, industrial). Examples: Apple, Amazon, JPMorgan, BNP Paribas, Nestlé, Volkswagen, Arçelik, Tüpraş, Garanti BBVA, Ziraat, Halkbank.
- Input features: Total Revenue, EPS, Net Income, Total Assets, Total Liabilities, Shareholders’ Equity, EBIT, Current Ratio (provided in Excel).
- Currency handling: analyses done in firms’ local reporting currencies (no cross‑firm currency conversion); forecasts used intra‑firm year‑over‑year percentage changes to avoid nominal distortion.
- Prompt engineering: standardized template, chain‑of‑thought (CoT) prompting, constrained to financial statement data only, structured tabular output required. Separate sessions per model-company-year.
- Evaluation metrics: direction accuracy, confusion matrices, bias/consistency metrics, and an interpretive “insight quality” score (qualitative scoring of explanations).
- Principal empirical finding: ChatGPT excels in generating plausible reasoning and handles volatile/recovery cases better; DeepSeek is more accurate and consistent in well‑regulated, stable environments (notably some banking contexts).
- Interpretation: results highlight model‑specific heuristic choices and training biases; neither model dominates across all sectors/regions.
- Recommendation: hybrid systems that combine an interpretable LLM for rationale and a sturdier/statistically calibrated model for final directional decision provide improved decision support.
Data & Methods
- Design: comparative longitudinal forecasting exercise. Models used only firm financial statements (2017–2021) to predict direction of earnings for 2022–2024.
- Sample selection criteria: asset size, exchange listing, data availability; 15 firms across Turkey, Europe, USA in banking, technology, industrial/energy.
- Financial statements used: Income Statement, Balance Sheet, Cash Flow, Statement of Changes in Equity. Data sources: KAP (Turkey), official annual reports (Europe, IFRS), investor relations/AnnualReports.com (USA, US GAAP).
- Feature engineering / derived metrics:
- EPS (IAS 33 / TFRS 33 methods when not directly reported).
- EBIT: computed differently for corporates (operating income + interest expense) and banks (approximated as net interest income + non‑interest income − operating expenses or operating profit before provisions & taxes).
- Current ratio and standard balance‑sheet aggregates derived per Penman/Damodaran conventions.
- Prompting protocol:
- Unified “forget previous instructions” template making the model act as a financial expert.
- Explicit constraints: use only provided financials; produce year‑by‑year direction, % estimate, and ≤3 sentence justification.
- Chain‑of‑thought encouraged to increase reasoning traceability.
- Model execution: identical inputs and prompt format supplied in separate sessions to GPT‑4 and DeepSeek‑V3.
- Evaluation:
- Directional predictions compared to realized outcomes for 2022–2024.
- Metrics: accuracy, class confusion matrices (false positives/negatives per class), consistency over years, bias tendencies, and a qualitative interpretive score for explanations.
- Limitations acknowledged by authors:
- Small sample (15 firms) limiting generalizability.
- Exclusion of macro/market sentiment and non‑financial textual sources.
- No currency normalization across firms limits cross‑firm magnitude comparison (authors argue intra‑firm % changes mitigate this).
- Dependence on prompt design, model versions, and possible hallucinations or table‑reading errors.
- Tests are retrospective; not a live trading or real‑time decision experiment.
Implications for AI Economics
- For theory: systematic, model‑dependent predictive differences using identical inputs present an empirical challenge to strict EMH claims that all useful information is already fully and equivalently priced; model processing matters.
- For practice (analysts / institutions):
- Match model choice to environment: use more interpretable LLMs for volatile/recovery situations where nuanced reasoning aids judgment; prefer more stable/accurate models (or calibrated ensembling) in regulated, stable sectors (some banking contexts).
- Adopt hybrid decision‑support workflows: combine LLM rationale (for auditability and human review) with a calibrated predictive engine for final directional calls.
- Insist on prompt standardization, CoT-style explanations, and version control for reproducibility and audit trails.
- For model developers:
- Improve sectoral calibration (accounting standard diversity, bank vs corporate accounting), table understanding, and robustness to input formatting.
- Provide tools for quantitative calibration of LLM probability outputs (to reduce over/underconfidence) and for structured explainability metrics to support auditability in finance.
- For regulators and auditors:
- Require transparency around prompt and model versions used in financial decision contexts; require backtesting and documentation of model performance across regions/sectors.
- Encourage standards for interpretability scoring and stress‑testing LLMs on accounting standards heterogeneity (IFRS vs US GAAP vs local TFRS).
- For research:
- Scale up sample sizes, include macro/market sentiment and textual disclosures, and benchmark against human analysts and specialized predictive ML models.
- Explore formal ensemble/hybrid strategies and calibration methods that reconcile interpretability and directional accuracy.
- Test real‑time deployment, economic value (e.g., trading or risk management performance), and fairness/ethical considerations in automated forecasting.
Overall, the paper provides a focused, controlled comparison that surfaces meaningful trade‑offs (interpretability vs reliability) when using general‑purpose LLMs for financial forecasting and makes a strong case for hybrid, audited decision‑support systems tailored by sector and regime.
Assessment
Claims (10)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| The study uses panel data of 15 Turkish, European, and American companies operating in the banking, technology, and industrial sectors for 2022-2024. Other | null_result | dataset composition (firms, sectors, regions, years) |
Reading fidelity
high
Study strength
high
|
n=15
|
| Projections from both models were produced using a consistent prompt style and evaluated using direction accuracy, confusion matrices, interpretive insight scores, and sector-region performance comparisons. Decision Quality | null_result | forecast evaluation metrics (direction accuracy, confusion matrices, interpretive insight scores, sector-region performance) |
Reading fidelity
high
Study strength
high
|
n=15
|
| ChatGPT (GPT-4) provides greater qualitative explanation (stronger interpretive insight scores) than DeepSeek-V3. Ai Safety And Ethics | positive | quality of qualitative explanations / interpretive insight scores |
Reading fidelity
medium
Study strength
medium
|
n=15
|
| ChatGPT performs better in unstable, recovery-based forecasting scenarios. Decision Quality | positive | forecast direction accuracy in unstable/recovery scenarios |
Reading fidelity
medium
Study strength
medium
|
n=15
|
| DeepSeek-V3 delivers greater accuracy, stability, and overall forecasting performance in stable, regulation-based environments. Decision Quality | positive | forecast accuracy and stability in stable/regulation-based scenarios |
Reading fidelity
medium
Study strength
medium
|
n=15
|
| The results highlight a critical trade-off between interpretability (qualitative explanation) and reliability (forecasting accuracy/stability). Ai Safety And Ethics | mixed | trade-off between interpretability and reliability of model outputs |
Reading fidelity
high
Study strength
medium
|
n=15
|
| Hybrid models (combining strengths of models like ChatGPT and DeepSeek) can offer the best solution to financial forecasting. Organizational Efficiency | positive | proposed performance improvements from model hybridization |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| These findings challenge (oppose) the assumptions of the Efficient Market Hypothesis by demonstrating that LLMs can extract predictive signal from historical financial reports. Market Structure | negative | ability of LLMs to predict earnings direction (implication for market efficiency) |
Reading fidelity
medium
Study strength
low
|
n=15
|
| The paper recommends decision-support systems that integrate AI interpretability, domain expertise, and responsible model selection when incorporating LLMs into financial decision-making. Governance And Regulation | positive | recommended system features (interpretability, domain expertise, model selection) |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| The research offers prescriptive advice to analysts, institutions, and developers aiming to incorporate LLMs into financial decisions. Governance And Regulation | positive | practical recommendations/advice uptake |
Reading fidelity
high
Study strength
speculative
|
not reported
|