The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Large language models nudge users toward textbook financial planning: simulated adherence to GPT-5.2 recommendations raises stock-market participation, increases savings buffers and produces age‑declining equity shares. But advice differs by user — men, prior AI users and the financially literate receive more equity-heavy and higher-saving guidance, with most of the gender gap driven by what people ask and a smaller portion by how the model responds.

AI Financial Advice: Supply, Demand, and Life Cycle Implications
Taha Choukhmane, Tim de Silva, Weidong Lin, Matthew Akuzawa · August 03, 2026
arxiv quasi_experimental medium evidence 8/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Taha Choukhmane unresolved corpus identity
  2. Tim de Silva unresolved corpus identity
  3. Weidong Lin unresolved corpus identity
  4. Matthew Akuzawa unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Taha Choukhmane provider ID
  2. Tim de Silva provider ID
  3. Weidong Lin provider ID
  4. Matthew Akuzawa provider ID
Using user-written prompts and a calibrated life-cycle model, the paper shows that following GPT-5.2 advice would, in simulation, move individuals closer to life-cycle prescriptions (broad stock participation, age-declining equity shares, larger savings buffers), while advice varies systematically by gender, financial literacy and prior AI experience—two-thirds of the gender equity-share gap stems from different prompts and one-third from model responses to gender labels.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

We ask a representative sample to write prompts seeking spending and investing advice from LLMs, then simulate the lifetime effects of following the advice under realistic asset and labor market conditions. Applying this method to GPT-5.2, we find following the advice would move respondents toward life cycle theory: broader participation in diversified equity funds, age-declining equity shares, and larger savings buffers. Recommendations vary systematically by gender, prior AI experience, and financial literacy. For gender, two-thirds of recommended equity-share differences arise from men and women writing different prompts (demand), while one-third arise from gender labels attached to otherwise identical prompts (supply).

Summary

Main Finding

LLM-generated financial advice (tested on GPT-5.2, with checks on Gemini 3 Flash and GPT-5.6 Terra) would, on average, move users closer to standard life-cycle recommendations: near-universal participation in diversified equity funds, equity shares that decline with age after about 45, and larger precautionary savings buffers. However, the advice also departs from optimal life-cycle prescriptions in systematic ways (e.g., strong reliance on heuristics, weak active rebalancing, undersmoothing of consumption). Advice varies meaningfully across user groups (gender, financial literacy, prior AI use), and these differences accumulate into modest lifetime wealth gaps (~4–6% at age 60). Much of the gender gap in equity recommendations is demand-driven (two-thirds), with the remainder arising from supply-side differences in model responses.

Key Points

  • Aggregate effect:
    • Following GPT-5.2 advice increases stock-market participation, concentrates risky holdings in diversified equity funds, and produces age-declining equity shares after ~45.
    • Nearly all simulated agents build savings buffers > $10k by age 30 under LLM advice (relative to observed low baseline wealth for many respondents).
  • Systematic departures from normative life-cycle prescriptions:
    • Advice implies a high discount factor (high patience).
    • Frequent use of simple heuristics (round-number targets, fixed withdrawal rules).
    • Poor consumption smoothing: too little decumulation in retirement and large consumption drops after job loss, even with liquid wealth.
    • Portfolios tend to drift with realized returns instead of active rebalancing.
  • Heterogeneity and lifecycle accumulation:
    • Advice differs by gender, financial literacy, and prior AI experience: sampling prompts from men, high-literacy respondents, or prior-AI users yields ~4–6% higher wealth at age 60.
    • For gender, ~2/3 of the recommended equity-share gap arises because men and women write different prompts (demand); ~1/3 arises because identical prompts elicit different model responses when labeled male vs female (supply).
    • Race labels produced smaller differences and no detectable supply-side effect.
  • Robustness:
    • Main quantitative patterns are similar across GPT-5.2, Gemini 3 Flash, and GPT-5.6 Terra.
  • Methodological contribution:
    • Introduces a survey-of-prompts approach combined with a calibrated life-cycle simulation and a two-step LLM pipeline (textual advice → LLM translation into quantitative choices).

Data & Methods

  • Survey:
    • Demographically balanced U.S. sample recruited via Prolific.
    • 1,000 complete responses collected; 952 passed authenticity and specificity screens.
    • Each respondent wrote three free-text prompts: (1) description of financial situation, (2) spending vs saving advice, (3) investing advice.
    • Collected respondent demographics, financial literacy, and prior AI use for stratified analysis.
  • LLM usage and experimental design:
    • Primary model: GPT-5.2. Also tested Gemini 3 Flash and GPT-5.6 Terra for robustness.
    • For gender supply-demand decomposition: used prompts that did not state gender, then randomly attached gender labels (e.g., “I am a man”) to isolate supply-side effects.
    • Each simulated-period LLM query was independent (no persistent memory); only the simulated state variables linked periods.
  • Life-cycle model and simulation:
    • Quantitative, stochastic life-cycle model in Gourinchas–Parker tradition, calibrated to U.S. data (SIPP, SSA, CRSP referenced for calibration inputs).
    • Features: stochastic income, job-to-job and unemployment transitions, stochastic asset returns, mortality risk, and U.S. tax/social-insurance system.
    • Asset classes: fixed income, diversified equity funds, individual stocks, and other risky assets (crypto, gold, commodities, collectibles).
    • Matching procedure: simulated agents (ages 22–90) were randomly matched to survey prompts from respondents with similar age, income, and employment. The prompt’s state variables were replaced by the simulated agent’s current state and submitted to the LLM.
    • Two-step LLM pipeline per period: (1) obtain textual advice from the LLM; (2) use a second LLM to translate the textual advice into quantitative recommendations for saving, consumption, and asset allocation.
    • Benchmarks: (i) optimal policy functions from dynamic optimization (life-cycle benchmark), (ii) “academic prompt” benchmark (structured, researcher-designed prompt that references life-cycle planning and best interest).
  • Outcome measures:
    • Consumption, saving, asset allocation paths over the life cycle and resulting wealth at age 60 (and other ages).
    • Diagnostics tied to normative life-cycle criteria: participation, diversification, age profile of equity share, consumption smoothing, active rebalancing.

Implications for AI Economics

  • Demand matters as much as supply. Heterogeneity in prompts (what users ask and how they present information) is a major driver of heterogeneity in advice and cumulative wealth outcomes. Studying or regulating AI financial advice requires attention to user behavior, prompt design, and interface choices—not only model architecture.
  • LLMs can potentially raise average financial quality at scale. General-purpose LLMs already nudge many users toward life-cycle-consistent behaviors (diversified equity exposure, savings buffers). This suggests a role for low-cost, conversational AI in reducing naïve or suboptimal financial behaviors at population scale.
  • Risks of uneven benefits and bias. Because advice varies systematically with user characteristics and because some supply-side demographic sensitivities exist (the gender label effect), LLMs may exacerbate or reproduce existing inequalities unless addressed. Diagnostics and debiasing interventions (prompt-standardization, model fine-tuning, evaluation benchmarks) are needed.
  • Need for lifecycle benchmarks and diagnostics. Embedding LLM advice in calibrated life-cycle simulations provides a useful, theory-grounded set of diagnostics (consumption smoothing, diversification, rebalancing) for evaluating model performance and progress over time. Policymakers and platform designers should adopt such longitudinal benchmarks rather than rely solely on one-shot or static evaluations.
  • Design and policy levers:
    • Prompt engineering and structured elicitation (e.g., “academic” or standardized prompts) can reduce heuristic-driven errors (better smoothing, higher-quality saving guidance) but may not solve all issues (e.g., active rebalancing).
    • Product design (templates, required disclosures, constrained action recommendations) could reduce harmful heterogeneity and supply-side biases.
    • Regulation and consumer protection should monitor demographic sensitivities and the real-world accumulation of advice-driven outcomes.
  • Future research directions:
    • Reapply the survey-and-simulation framework across other populations, countries, financial domains (insurance, housing, credit), and evolving LLMs.
    • Test interventions: prompt templates, model fine-tuning for fiduciary behavior, persistent memory/stateful advice, or hybrid human–AI advisory setups.
    • Empirically track whether simulated advice translates into behavioral change in field settings and quantify welfare impacts beyond wealth (utility, consumption volatility, bequests).
  • Caveats:
    • Results describe advice from specific-generation general-purpose LLMs and may change as models are optimized for financial advice or as user behavior evolves.
    • The simulation assumes users follow LLM advice exactly and that the LLM has no memory across periods—both simplifications relative to real-world use.

Assessment

Paper Typequasi_experimental Evidence Strengthmedium — The paper provides credible, internally consistent evidence on what LLMs recommend and how those recommendations vary by user characteristics, including an experimental randomized label to identify supply-side differences; however, the paper does not observe actual behavioral responses to advice in the field — lifetime economic effects are inferred from simulations that assume compliance and from translations of textual advice into quantitative choices, which weakens external causal claims about realized economic impacts. Methods Rigorhigh — The authors combine a reasonably large, demographically balanced survey (N≈952 validated prompts) with a carefully calibrated stochastic life-cycle model (SIPP/SSA/CRSP calibration), translate textual advice into quantitative recommendations, randomize gender labels to identify supply-side effects, and conduct cross-model robustness checks; main limitations include reliance on Prolific sample representativeness, measurement error in translating text to numeric advice (performed via LLM), the strong assumption that agents follow advice perfectly, and the modelling choices in the life-cycle calibration. SampleSurvey of ~1,000 U.S. adults recruited on Prolific (1,000 complete responses, 952 passing authenticity and specificity checks), demographically balanced; respondents wrote three free-text prompts (financial situation, spending vs saving, investing); survey collected demographics, financial literacy, prior AI experience. Primary LLM: GPT-5.2; robustness using Gemini 3 Flash and GPT-5.6 Terra. Life-cycle model calibrated to U.S. data (SIPP, SSA, CRSP) and simulated ages ~22–90; simulations match agents to prompts by similar age/income/employment and iterate annual independent LLM queries (text advice) plus an LLM-based translation into quantitative choices. Themeshuman_ai_collab inequality IdentificationCombines a demographically balanced survey of user-written LLM prompts with randomized label experiments and calibrated life-cycle simulation: (1) random assignment of gender labels to otherwise identical prompts isolates supply-side (model) differences in recommendations; (2) comparing simulated life-cycle outcomes when sampling prompts from different respondent subgroups (e.g., men vs women; high vs low financial literacy; prior AI users vs non-users) identifies how demand-side heterogeneity in prompts maps into cumulative wealth differences; (3) robustness checks use multiple LLMs (GPT-5.2, Gemini 3 Flash, GPT-5.6 Terra) and an ‘academic’ benchmark prompt; causal claims about model behavior and supply-side effects rely on the randomized label; claims about real-world welfare effects are simulation-based (conditional on following LLM advice) rather than direct causal evidence from observed behavior. GeneralizabilityProlific sample may not perfectly represent the broader U.S. population (selection on platform participation and survey-taking behavior)., Simulated life-cycle effects assume full or rule-following adherence to LLM advice; real-world compliance likely lower and heterogeneous., Translation of free-text LLM advice into quantitative savings/portfolio choices introduces measurement error and depends on the translation procedure (also LLM-based)., Findings are calibrated to historical U.S. income, return, and policy parameters; results may not generalize to other countries or different macro/market regimes., LLM behavior evolves rapidly; results for GPT-5.2/5.6 and Gemini 3 Flash may differ for future or provider-specific models., Prompts were constrained to spending/saving/investing; other financial domains (insurance, credit, housing) are not covered., Queries are independent across periods (no memory), which may not reflect real multi-turn advisory interactions.

Claims (12)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Following GPT-5.2's financial recommendations would move most respondents closer to broad life-cycle theory, including greater participation in diversified equity funds, equity shares that decline with age after 45, and larger savings buffers. Consumer Welfare positive Simulated household saving and portfolio-allocation behavior relative to observed behavior and life-cycle-theory benchmarks
Reading fidelity high
Study strength medium
n=952
0.48
Following the advice would lead virtually all simulated individuals to accumulate more than $10,000 in savings by age 30. Consumer Welfare positive Financial wealth or savings-buffer accumulation by age 30
Reading fidelity high
Study strength medium
above $10,000 by age 30
0.48
GPT-5.2 advice would induce near-universal stock-market participation, with risky holdings concentrated in diversified equity funds rather than individual stocks, crypto, gold, commodities, or collectibles. Adoption Rate positive Stock-market participation and composition of risky-asset holdings
Reading fidelity high
Study strength medium
near-universal participation
0.48
Prompts mentioning macroeconomic uncertainty cause the LLM to recommend more saving and lower equity holdings. Consumer Welfare mixed Recommended saving rates and equity holdings conditional on macroeconomic-uncertainty prompts
Reading fidelity high
Study strength medium
not reported
0.48
The LLM's advice insufficiently smooths consumption, recommending too little retirement decumulation and sharp consumption declines after job loss even when simulated individuals have sufficient liquid wealth. Consumer Welfare negative Consumption smoothing after job loss and during retirement
Reading fidelity high
Study strength medium
not reported
0.48
A more structured academic prompt improves consumption smoothing and reduces reliance on saving and withdrawal heuristics, but does not increase active portfolio rebalancing. Decision Quality mixed Consumption smoothing, reliance on heuristics, and active portfolio rebalancing
Reading fidelity high
Study strength medium
not reported
0.48
Simulated wealth at age 60 is approximately 5% higher when using prompts written by men, respondents with high financial-literacy scores, or respondents who previously used AI for financial guidance. Consumer Welfare positive Average simulated wealth at age 60
Reading fidelity high
Study strength medium
around 5% higher
0.48
The LLM recommends lower equity shares for prompts written by women and individuals with lower financial literacy, and lower saving rates for prompts written by individuals without prior AI use. Inequality negative Recommended equity shares and saving rates across respondent groups
Reading fidelity high
Study strength medium
not reported
0.48
Approximately two-thirds of the gender gap in recommended diversified-equity shares is demand-driven, while approximately one-third is supply-driven by the gender label attached to otherwise identical prompts. Inequality negative Gender gap in recommended diversified-equity shares and its demand-versus-supply decomposition
Reading fidelity high
Study strength medium
two-thirds demand-driven; one-third supply-driven
0.48
Gender differences in recommended holdings of non-diversified assets, including individual stocks or crypto, are entirely demand-driven. Inequality negative Gender differences in recommended holdings of individual stocks, crypto, and other non-diversified assets
Reading fidelity high
Study strength medium
entirely demand-driven
0.48
Explicit race labels leave recommendations essentially unchanged and produce no detectable supply-side effect. Inequality null_result Changes in financial recommendations caused by explicit race labels
Reading fidelity high
Study strength medium
no supply-side effect
0.48
The main results are quantitatively similar when the analysis uses Gemini 3 Flash and GPT-5.6 Terra instead of GPT-5.2. Decision Quality positive Robustness of simulated financial-advice outcomes across LLM models
Reading fidelity high
Study strength low
not reported
0.24

Notes