The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

AI investment advice meaningfully shifts investor risk exposure but only partially: in a randomized experiment with 400 Korean pension holders, aggressive versus conservative GPT recommendations transmit roughly 37% of their difference into final portfolios, increasing expected returns and volatility while producing no detectable gain in risk-adjusted performance.

Do People Follow AI Advice? Evidence from a Pension Portfolio Choice Experiment
Hongseok Choi, Jeongbin Kim, Matthew Kovach, Kyu-Min Lee, Euncheol Shin, Hector Tzavellas · August 11, 2026
arxiv rct high evidence 9/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Hongseok Choi unresolved corpus identity
  2. Jeongbin Kim unresolved corpus identity
  3. Matthew Kovach unresolved corpus identity
  4. Kyu-Min Lee unresolved corpus identity
  5. Euncheol Shin unresolved corpus identity
  6. Hector Tzavellas unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Hongseok Choi provider ID
  2. Jeongbin Kim provider ID
  3. Matthew Kovach provider ID
  4. Kyu-Min Lee provider ID
  5. Euncheol Shin provider ID
  6. Hector Tzavellas provider ID
In an incentivized RCT with 400 Korean DC pension participants, assignment to aggressive versus conservative AI portfolio recommendations causally shifts final allocations—about 37% of the recommendation gap transmits into choices—raising expected return and volatility but not improving Sharpe ratios.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

We study how differences in AI-generated financial recommendations are transmitted into individual portfolio choices. In an experiment with 400 employed adults enrolled in workplace defined contribution pension plans in South Korea, participants allocate a hypothetical pension balance across eleven products and may revise it after receiving one of two fixed AI-generated recommendations. A $2 \times 2$ design randomizes recommendation content and whether the recommendation includes a short rationale. Approximately 37$\%$ of the experimentally induced difference between the aggressive and conservative recommendations passes through to final portfolios. This causal contrast changes expected portfolio return, volatility, allocations across risk grades, and the number of products held, but produces no detectable difference in computed Sharpe ratios. 81$\%$ of participants revise. Among revisers, 95$\%$ move toward the assigned recommendation and implement about half of the suggested adjustment. Rationales do not detectably alter pass-through. These results show that users partially and selectively transmit recommendation content into economically meaningful differences in risk exposure while retaining substantial weight on their initial choices.

Summary

Main Finding

A single, nonbinding AI portfolio recommendation causally shifts individuals’ pension allocations substantially but incompletely. Random assignment to an aggressive versus a conservative GPT-generated recommendation produces a 37% pass-through of the experimentally induced difference into final portfolios: participants move meaningfully toward the recommended portfolios (changes in expected return, volatility, risk-grade allocations, and number of products held) but retain substantial weight on their initial choices. A short verbal rationale accompanying the numeric recommendation does not detectably change pass-through.

Key Points

  • Design and sample

    • n = 400 employed adults in South Korea, ages 35–55, all covered by workplace defined-contribution (DC) pension plans.
    • Between-subjects 2×2: recommendation content (GPT-4 Turbo = relatively aggressive vs GPT-4o = relatively conservative) × presence/absence of a short rationale. Participants were told the recommendation came from AI but not which model.
    • Quotas ensured 100 participants per treatment cell; only desktop/laptop respondents were eligible.
  • Main quantitative results

    • Aggregate causal pass-through: ≈ 37% of the difference between the aggressive and conservative recommendations appears in participants’ final portfolios (robust across specifications; rejects both 0% and 100% pass-through).
    • 81% of participants revised their portfolio after seeing the recommendation.
    • Among revisers: 95% moved in the direction of the assigned recommendation and implemented roughly half (~50%) of the suggested adjustment (intensive margin).
    • Incomplete pass-through also reflects an extensive margin: ~19% did not revise at all.
    • Pass-through affected expected return and portfolio risk (volatility), reduced allocations to principal-protected/low-risk products, increased allocations to high/very-high-risk products, and reduced number of products held (concentration), with roughly 28–40% of recommendation differences passing through across these dimensions.
    • No detectable pass-through in measures of risk-adjusted performance (Sharpe ratios) or in concentration measures intended to capture improvements in risk-adjusted outcomes.
    • The short verbal rationale did not produce a detectable shift in average portfolio position nor change the pass-through magnitude.
  • Behavioral heterogeneity (descriptive)

    • Participants who believed their baseline dominated the recommendation revised as often as others but directed less of their revision toward the AI.
    • Participants whose baseline was less efficient than the recommendation revised more, but did not end up significantly closer to the recommendation.
    • These relationships are correlational (beliefs elicited post-decision), not causal.

Data & Methods

  • Task

    • Two-stage incentivized hypothetical pension allocation: participants first choose an initial allocation over 11 pension products (table of multi-horizon gross returns, discrete risk grade, product-level standard deviation, etc.), then see an AI recommendation and may revise. Choices were payoff-relevant via a performance-based bonus linking expected return and risk to payoffs.
    • Eleven-product menu matched Ha et al. (2019) to create a realistic, controlled DC choice environment for Korean subjects.
  • Treatment generation

    • Both recommendations were produced from identical product information via prompts (in Korean) to two GPT versions used in spring 2024: GPT-4 Turbo and GPT-4o.
    • Each model was queried 100 times at temperature 0.5; the modal allocation across draws was used as the recommendation.
      • GPT-4 Turbo modal allocation (32/100 draws): (10, 10, 5, 15, 0, 0, 10, 20, 0, 0, 30) — relatively aggressive, concentrated into higher-return/risk products (esp. products 8 and 11).
      • GPT-4o modal allocation (16/100 draws): (20, 20, 10, 10, 10, 10, 5, 5, 0, 5, 5) — relatively conservative and diversified.
    • In the rationale conditions, a short model-specific explanatory paragraph (drawn from generations that produced the modal allocation) accompanied the numeric recommendation.
  • Measurement of pass-through

    • Each recommendation was normalized along the line connecting the conservative recommendation (= 0) and the aggressive recommendation (= 1). Participants’ final portfolios were projected onto this line; the randomized difference in average final positions between the two recommendation assignments equals the share of the experimentally induced difference that passed through.
    • Primary robustness: change-score, baseline-adjusted, and covariate-adjusted specifications all produced similar pass-through estimates.
  • Outcomes analyzed

    • Portfolio position on the conservative → aggressive line (primary pass-through metric).
    • Expected portfolio return and volatility.
    • Allocations by discrete risk grade (principal-protected, low, medium, high, very high).
    • Number of products held (breadth/concentration).
    • Sharpe ratios and concentration measures (risk-adjusted performance).
  • Pre-analysis

    • Analysis plans preregistered (AsPredicted). The design isolates causal effect of recommendation content (two specific GPT-produced portfolios), not a general causal statement about the GPT models themselves.

Implications for AI Economics

  • Recommendation content matters: Different general-purpose LLMs exposed to identical product information can produce materially different financial recommendations, and a substantial fraction of those differences can transmit into users’ economic decisions. Thus model choice and prompt/output design can affect population-level risk exposures even without ongoing delegated management.
  • Partial and selective transmission: AI does not simply replace human judgment; it shifts it. Most users move toward AI recommendations but only partially, and many make additional orthogonal adjustments. Predicting behavioral impacts of AI advisory systems therefore requires modeling both content generation and users’ advice-integration process (extensive and intensive margins).
  • Risk exposure vs. risk-adjusted performance: The observed pass-through changes levels and composition of risk assumed by individuals but does not detectably improve risk-adjusted outcomes (Sharpe ratio). Deploying AI advice can therefore alter aggregate risk-taking without clear welfare improvement unless model recommendations are explicitly calibrated to improve risk-adjusted performance.
  • Explanations may not help: Short verbal rationales did not measurably increase pass-through in this setting, suggesting that simple textual explanations may be insufficient to change how users integrate AI advice in multidimensional allocation tasks.
  • Policy and product design relevance: Regulators and firms should note that accessible, general-purpose AI advisors can change users’ portfolio risk exposures in economically meaningful ways even when users do not fully accept recommendations. Transparency about model behavior, testing for systematic differences across models/prompts, and attention to how advice is presented (beyond short rationales) matter for consumer protection and product design.
  • External validity and limits: Results are specific to a controlled pension-allocation exercise with Korean DC participants and to two GPT variants circa 2024; findings speak to likely patterns (partial compliance, directional adjustments) rather than universal magnitudes. Future work should test other populations, continuous/delegated advisory products, repeated interactions, and welfare consequences of different recommendation-generation strategies.

Assessment

Paper Typerct Evidence Strengthhigh — Randomized assignment to two substantively different recommendations provides clean causal identification for the effect of those specific recommendations on portfolio choices; preregistration, balanced quotas, robustness across specifications, and detailed behavioral decomposition strengthen internal validity. Limitations include a one-time hypothetical (but incentivized) task, a single-country sample, and inference limited to the particular AI outputs used. Methods Rigorhigh — Design strengths: preregistered RCT, balanced demographic quotas, clear 2×2 manipulation separating content vs rationale, well-defined outcome measures (portfolio weights, expected return, volatility, Sharpe), and robustness checks; analysis connects aggregate experimental contrast to individual-level revision margins. Limitations: hypothetical/incentivized allocation rather than observed real-world trading, use of modal GPT outputs (reliance on specific model draws), and external validity constraints (single country, age band, desktop users). Sample400 employed adults in South Korea aged 35–55 covered by workplace defined contribution (DC) pension plans, recruited via an online panel (Macromill Embrain) with quotas balancing age and gender cells; participants completed a desktop/laptop-based, incentivized hypothetical allocation of a pension balance across 11 pension products and were randomized into four treatment arms (GPT-4 Turbo vs GPT-4o recommendation × rationale vs no rationale). Themeshuman_ai_collab adoption IdentificationRandomized 2×2 between-subjects experiment: participants (N=400) were randomly assigned to receive one of two AI-generated portfolio recommendations (an aggressive portfolio from GPT-4 Turbo vs a conservative portfolio from GPT-4o) and independently to see the recommendation with or without a short rationale; causal pass-through is identified by comparing final portfolio positions across randomized recommendation assignments (normalized on the line between the two recommendations). Analysis is preregistered and uses change-score, baseline-adjusted, and covariate-adjusted specifications. GeneralizabilitySample limited to South Korean, employed DC participants aged 35–55—may not generalize to other countries, age groups, retirees, unemployed, or self-employed individuals., One-time, hypothetical but incentivized allocation may not reflect real-world long-run portfolio behavior or actual financial commitments., Recommendations come from two specific GPT model outputs (modal draws) in Korean; results may vary with different prompts, models, languages, or repeated/adaptive advice., Study restricted to desktop users and an online panel—selection bias vs broader retail investor population., Preregistered behavioural heterogeneity findings (e.g., belief correlates) are descriptive rather than causal because those characteristics were not randomized.

Claims (9)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Approximately 37% of the experimentally induced difference between the aggressive and conservative AI recommendations passes through to participants' final pension portfolios. Task Allocation positive The share of the difference in assigned AI recommendation content reflected in final portfolio choices
Reading fidelity high
Study strength high
n=400
approximately 37%
1.0
Assignment to the aggressive AI recommendation increases expected portfolio return. Consumer Welfare positive Expected return of the final pension portfolio
Reading fidelity high
Study strength high
n=400
approximately 28 to 40 percent of the difference between recommendations
1.0
Assignment to the aggressive AI recommendation increases portfolio risk and shifts allocations toward higher-risk products. Consumer Welfare positive Portfolio volatility and allocation to high- and very-high-risk products
Reading fidelity high
Study strength high
n=400
approximately 28 to 40 percent of the difference between recommendations
1.0
Assignment to the aggressive AI recommendation reduces the number of products held in participants' final portfolios. Task Allocation negative Number of investment products held in the final pension portfolio
Reading fidelity high
Study strength high
n=400
approximately 28 to 40 percent of the difference between recommendations
1.0
The AI recommendation treatment produces no detectable difference in computed Sharpe ratios. Consumer Welfare null_result Sharpe ratio of the final pension portfolio
Reading fidelity high
Study strength high
n=400
1.0
Short rationales accompanying the numerical AI recommendations do not detectably alter average portfolio positions or the degree of pass-through. Task Allocation null_result Average final portfolio position and recommendation-content pass-through
Reading fidelity high
Study strength high
n=400
1.0
81% of participants revise their initial portfolios after receiving an AI recommendation. Task Allocation positive Whether participants revised their initial pension allocation
Reading fidelity high
Study strength medium
n=400
81%
0.6
Among participants who revise, 95% move in a direction aligned with the assigned AI recommendation. Task Allocation positive Whether a revising participant's portfolio movement was toward the assigned recommendation
Reading fidelity high
Study strength medium
n=324
95%
0.6
Among participants who revise, they implement approximately half of the adjustment suggested by the AI recommendation. Task Allocation positive Fraction of the AI-suggested portfolio adjustment implemented by revisers
Reading fidelity high
Study strength medium
n=324
approximately half of the suggested adjustment
0.6

Notes