0 cumulative citations
View corpus contextGenerative AI parsing macro releases uncovers a currency-strength signal: a simple long–short strategy based on the AI-derived AIFX index yields a Sharpe ratio above 0.7 and retains significant alpha after standard FX factor controls, and several robustness checks argue performance is not solely due to LLM memorization.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
0 cumulative citations
View corpus contextI revisit the exchange rate disconnect puzzle, first documented by Meese and Rogoff (1983), using generative artificial intelligence (AI) to forecast currency returns based on economic fundamentals. Using ChatGPT and DeepSeek, I analyze a comprehensive dataset of economic data releases for major currency pairs and measure the fundamental strength of each currency. These AI-powered fundamentals exhibit significant cross-sectional predictive power. A simple trading strategy that goes long currencies with strong fundamentals and short currencies with weak fundamentals generates a Sharpe ratio exceeding 0.7 per annum. The excess returns of this strategy remain significant after controlling for traditional currency factors. To mitigate concerns of look-ahead bias, I run multiple exercises to ensure that predictability stems from AI reasoning rather than memorization. Finally, I explore the potential sources of predictability and find evidence that the Taylor rule framework, generally used by central banks to set interest rates, is a key mechanism connecting exchange rates to economic fundamentals.
Summary
Main Finding
AI language models (GPT-4o and DeepSeek‑V3) can convert structured macroeconomic data releases into a simple, high‑signal currency fundamentals index (AIFX) that materially predicts future exchange‑rate returns. A cross‑sectional trading strategy based on AIFX (long currencies with strong AI‑implied fundamentals, short those with weak fundamentals) delivers economically significant returns (annualized Sharpe > 0.7) and a robust alpha after controlling for standard currency factors. Evidence points to Taylor‑rule related variables (inflation, employment, broad activity) as the key economic channels.
Key Points
-
Data and scope
- Covers 544 unique economic indicators across G‑10 economies (USD, EUR, JPY, GBP, CHF, CAD, AUD, NZD, SEK, NOK).
- Economic calendar from Investing.com: realized, previous, consensus forecast; sample Jan 1996–Oct 2024; 174,820 data points.
- FX: nine USD/foreign exchange rates (end‑of‑day London) and 1‑month forward rates from Bloomberg.
-
AI processing and signal construction
- Structured prompt to LLMs (GPT‑4o baseline; DeepSeek‑V3 replication): provide indicator name, actual, previous, forecast, and currency (explicitly exclude date/time).
- LLM returns: short analysis + directional label ∈ {STRENGTHEN, WEAKEN, INSIGNIFICANT OR UNCERTAIN}.
- For each currency and lookback window L = (t − τ, t]:
- Strength = (# STRENGTHEN in L) / (total # releases in L)
- Weakness = (# WEAKEN in L) / (total # releases in L)
- AIFX = Strength − Weakness (net AI‑implied fundamental strength)
-
Trading strategy and main performance
- Cross‑sectional strategy (AIFX strategy): monthly sort by AIFX; long top two currencies, short bottom two (equal weights).
- Lookback windows tested 1–60 months; strongest predictive performance for 36–60 month lookbacks.
- Reported annualized Sharpe ratio > 0.7; cumulative performance shows persistent gains.
- After controlling for conventional currency factors (Dollar, Dollar Carry, Carry, Momentum, Value), a statistically significant alpha remains, accounting for ~74% of the strategy’s average return.
-
Robustness and validation
- Replication with DeepSeek‑V3 yields similar results.
- Alternative Weighted AIFX (weights by release importance) and a time‑series variant both produce economically meaningful returns.
- Turnover moderate — authors argue transaction costs unlikely to erase returns.
- Four exercises to address look‑ahead/memorization concerns:
- Ask LLM to guess release year — guesses poorly (only ~5.6% correct per year), suggesting LLMs are not implicitly using timing.
- Difference‑in‑differences between GPT‑3.5 (training cutoff Sep 2021) and GPT‑4o (Oct 2023): no significant relative performance drop for GPT‑3.5 post‑cutoff window.
- Test whether LLMs simply memorize realized correlations between macro variables and next‑month returns — find no evidence of such memory.
- Construct a “pure hindsight” portfolio from what the model could remember about past monthly FX returns; AIFX returns are orthogonal to this control.
- Conclusion: results are unlikely to be driven purely by look‑ahead bias and are consistent with LLM reasoning over structured inputs.
-
Economic mechanism
- Top predictive indicator categories: Inflation, Employment, Broad economic activity — consistent with Taylor‑rule / monetary policy channel.
- Predictability largely driven by positive news (STRENGTHEN signals); negative news elicits stronger immediate market reaction (less delayed predictability).
- Interpretation: central bank policy responses (and asymmetric responses to negative vs positive news) are a plausible mechanism connecting fundamentals to exchange rates.
Data & Methods
- Data sources
- Investing.com economic calendar (actual, previous, forecast) for 544 indicators, Jan 1996–Oct 2024.
- Bloomberg end‑of‑day FX rates (USD price of 1 unit of foreign currency) and 1‑month forward rates.
- LLM setup and prompt
- Prompt includes only indicator name, realized, previous, forecast values, and currency label; instructs model to behave as financial analyst and output (ANALYSIS, DIRECTION).
- Directional labels mapped to positive/negative/neutral news.
- Baseline model: GPT‑4o via API. Robustness model: DeepSeek‑V3 (and comparisons to GPT‑3.5 for cutoff tests).
- Variable construction
- For each currency c and time t, for lookback τ compute Strengthc,t,τ, Weaknessc,t,τ, and AIFXc,t,τ = Strength − Weakness.
- Also construct a Weighted AIFX variant where each release is weighted by estimated importance.
- Strategies and evaluation
- Cross‑sectional sort: monthly rebalancing; long top 2 / short bottom 2 by AIFX (tested for τ = 1..60 months).
- Time‑series alternative strategy also tested.
- Control regressions: regress strategy excess returns on Dollar, Dollar Carry, Carry, Momentum, Value factors (monthly panel regressions and factor regressions).
- Statistical tests: performance metrics (mean, sd, Sharpe), panel regressions predicting next‑month returns, replication across models, and four look‑ahead robustness exercises described above.
Implications for AI Economics
-
Methodological implications
- LLMs can be used as structured quantitative analysts: when fed cleaned numerical inputs (actual, forecast, previous) and a clear instruction, they yield economically interpretable signals that aggregate heterogeneous releases into a succinct fundamentals index.
- AI opens a scalable route to convert large economic calendars into high‑frequency fundamentals signals without bespoke econometric feature engineering.
- The approach complements — rather than replaces — standard factor models: AIFX captures incremental information orthogonal to common currency factors.
-
Economic and markets implications
- Reopens parts of the “exchange rate disconnect” debate: richer, AI‑synthesized fundamentals (especially Taylor‑rule related indicators) can deliver practical predictability in FX markets.
- Highlights the central role of monetary policy expectations (inflation, employment, activity) and asymmetric market responses to good vs bad news in exchange‑rate dynamics.
- Practical potential for asset managers, FX desks, and macro hedge funds to incorporate AI‑derived fundamentals into allocation, hedging, and alpha strategies — subject to replication and transaction‑cost analysis.
-
Cautions and open questions
- Replicability and model dependence: results hinge on prompt design, LLM version, and pretraining data; thorough out‑of‑sample, realtime paper‑trading, and multi‑model replication are important next steps.
- Remaining leakage risk: although the paper implements multiple controls for look‑ahead bias, absolute exclusion of subtle data‑leakage via pretraining is challenging; independent realtime implementation would be the definitive test.
- Economic interpretation limits: while evidence supports the Taylor‑rule channel, causality is not definitively established — alternative channels (risk premia, global liquidity, sovereign risk) deserve further testing.
- Governance and operational risks: using LLMs in production trading raises issues around model updates, prompt stability, and regulatory/compliance transparency.
Summary takeaway: Carefully prompted LLMs can synthesize large structured macro release datasets into a compact fundamentals index (AIFX) that delivers robust, economically meaningful exchange‑rate predictability, with Taylor‑rule related variables as the likely economic transmission channel — but practical deployment requires rigorous replication and real‑time validation.
Assessment
Claims (12)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| AI-derived fundamentals, measured using the AIFX index, have significant cross-sectional predictive power for future exchange-rate returns. Other | positive | Future exchange-rate returns |
Reading fidelity
high
Study strength
medium
|
n=174820
|
| A trading strategy that goes long currencies with strong AIFX signals and short currencies with weak AIFX signals produces an annualized Sharpe ratio exceeding 0.7. Other | positive | Risk-adjusted trading-strategy returns |
Reading fidelity
high
Study strength
medium
|
Sharpe ratio exceeding 0.7 per annum
|
| The AIFX strategy has statistically significant alpha after controlling for traditional currency factors, and the alpha accounts for 74% of the strategy's average return. Other | positive | Factor-adjusted excess returns of the AIFX strategy |
Reading fidelity
high
Study strength
medium
|
74% of the AIFX strategy's average return
|
| The AIFX index predicts next-month exchange-rate returns in panel regressions. Other | positive | Next-month exchange-rate returns |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Inflation, employment, and broad economic activity indicators are the most important categories for forecasting exchange rates in the AI-based analysis. Other | positive | Exchange-rate return predictability by economic-indicator category |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The predictive signal is largely driven by positive news implying currency appreciation rather than negative news implying depreciation. Other | positive | Long-horizon exchange-rate return predictability |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Negative economic news generates a stronger immediate market reaction than positive economic news, leaving less delayed predictive content at longer horizons. Other | mixed | Immediate and delayed exchange-rate reactions to economic news |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The results remain consistent when the analysis is replicated using DeepSeek-V3 instead of GPT-4o. Other | positive | Trading-strategy performance and exchange-rate predictability |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The paper finds no significant difference in the relative performance of GPT-3.5 and GPT-4o between periods covered by both models' training data and the later period covered only by GPT-4o. Other | null_result | Relative predictive performance of GPT-3.5 versus GPT-4o |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Only 5.6% of the AI model's guesses about the year of an economic-data release are correct on average within each year. Ai Safety And Ethics | null_result | Accuracy of inferred economic-data release year |
Reading fidelity
high
Study strength
medium
|
5.6% correct guesses
|
| The return of the AIFX strategy is orthogonal to the return of a pure hindsight portfolio constructed from what the AI model can remember about monthly currency returns. Other | null_result | Correlation or return comovement between AIFX and hindsight portfolios |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The dataset contains 544 unique economic indicators and 174,820 data points spanning January 1996 to October 2024. Other | positive | Scope of the economic-data sample |
Reading fidelity
high
Study strength
high
|
n=174820
544 unique indicators
|