0 cumulative citations
View corpus contextLanguage models steer money unevenly: while fund picks are broadly consistent across investor types, LLMs recommend different investment sizes and allocate more capital to non-Black and male managers—biases that are often stronger when demographics are signaled implicitly by names.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Abstract Do large language models (LLMs) generate unbiased financial advice across investor and fund manager demographics? We develop a two-sided audit framework to evaluate demographic bias in LLM-generated investment advice and apply it to multiple large language models, with GPT-4 Turbo as the primary baseline. On the investor side, fund selections are similar across demographic groups and rely on financial criteria, but recommended investment amounts vary when investor names signal race or gender, despite identical age and income. On the fund manager side, capital allocations favor non-Black and male managers: racial disparities persist even under explicit disclosure, while gender-related differences are more pronounced under name-based cues. Bias patterns are qualitatively similar across models, with differences in magnitude between implicit and explicit demographic signaling. These results suggest that, even when LLMs incorporate core financial reasoning, demographic signals can affect allocation decisions, with effects that tend to be stronger under implicit signaling, potentially replicating existing market inequalities and raising concerns about impartiality in financial advising. The proposed audit framework provides a generalizable approach for identifying and evaluating demographic bias in AI-driven financial advisory systems.
Summary
Main Finding
LLMs can reproduce demographic disparities in investment advice even when financial fundamentals are held constant. Using a two-sided audit (investor-side and fund-manager-side), the authors find that fund selection is largely driven by financial criteria, but recommended investment amounts and capital allocations vary by demographic signals: investor names that imply race or gender affect recommended amounts, and fund managers who are Black or implicitly signaled as female receive systematically lower allocations. Racial disparities persist even with explicit disclosure; gender disparities are stronger under implicit (name-based) signaling. Results generalize qualitatively across multiple LLMs but vary in magnitude and sign.
Key Points
- Two-sided audit framework: tests both (a) investor-side effects (fund choice and recommended investment amount) and (b) fund-manager-side effects (capital allocated to managers).
- Signal manipulation: demographic cues introduced either implicitly (names) or explicitly (stated race/gender) while holding age, income, risk tolerance, and other financial attributes constant.
- Investor-side: fund selection (which funds are recommended) is similar across demographic groups and follows financial metrics; however, recommended investment amounts vary by inferred race/gender from names.
- Fund-manager-side:
- Racial bias: Black fund managers receive lower recommended allocations than White counterparts; effect persists under explicit disclosure of race.
- Gender bias: Female managers receive lower allocations when gender is implied via name, but explicit disclosure of gender often removes statistically significant differences.
- Implicit signals tend to produce stronger bias effects than explicit disclosures.
- Model heterogeneity: GPT-4 Turbo used as baseline; additional tests with GPT-4.1, GPT-4o, Claude 3.5 Sonnet, and Llama 3.1 8B show similar qualitative patterns but differ in magnitude and sometimes sign.
- Robustness: Enriching prompts (adding horizon, risk tolerance, long-term criteria) and varying age/income shows many disparities remain—especially for younger or lower-income profiles.
- Ethical and fiduciary concerns: demographic-driven allocation changes raise questions about impartiality, duty of care, and the legitimacy of deploying LLMs in financial advice.
Data & Methods
- Experimental design: controlled audit study that systematically manipulates demographic signals while keeping financial profiles identical. Two complementary experiments:
- Investor-side: present identical investor financial profiles with varied names (and sometimes explicit demographic labels) and record recommended funds and investment amounts.
- Fund-manager-side: present fund descriptions with manager identity signaled via name or explicit race/gender and record capital allocation recommendations.
- Models tested: primary baseline GPT-4 Turbo (chosen for performance/efficiency reproducibility) and replication on GPT-4.1, GPT-4o, Claude 3.5 Sonnet, Llama 3.1 8B.
- Signal types: implicit (names chosen to signal race/gender) vs explicit (directly stated race/gender).
- Robustness checks: enriched prompts (risk tolerance, horizon, objectives), varying investor age and income, long-term evaluation criteria for managers.
- Statistical analysis: comparisons of selections and allocation amounts across demographic treatments; significance testing reported for observed disparities (paper reports statistically significant racial effects and context-dependent gender effects).
- Conceptual grounding: links to audit-study methodology from labor, lending, housing; draws on behavioral finance insights (structured vs open-ended decisions) and literature on algorithmic fairness, RLHF, and test-time control.
Implications for AI Economics
- Distributional effects on capital: If deployed at scale, LLM-driven advisory tools could systematically redirect capital away from Black managers and certain demographic investor groups, reinforcing existing market inequalities in fund flows, fundraising, and wealth accumulation.
- Market efficiency and allocative justice: Demographic-driven allocation distortions can reduce allocative efficiency and block potentially valuable managers from capital, with long-run impacts on competition, innovation, and returns.
- Fiduciary and regulatory risks: Financial institutions using LLMs may face fiduciary duty breaches (duty of care/loyalty) if advice is influenced by irrelevant demographic cues. Regulators may demand audits, transparency, or constraints on model use in advisory roles.
- Model-specific risk: Heterogeneity across models implies institutions could inadvertently introduce model-specific biases; audits must be model- and prompt-specific rather than relying on generic claims of neutrality.
- Design and governance responses:
- Operational: implement regular two-sided audits, sanitize demographic cues in inputs where appropriate, enforce standardized structured decision protocols that reduce open-ended allocations, and require human-in-the-loop review for allocation decisions.
- Technical: incorporate fairness objectives into RLHF, apply counterfactual or constraint-based post-processing to allocations, and develop test-time controls to neutralize implicit demographic associations.
- Policy: mandate disclosure of model use in advisory contexts, require provenance/ audit logs for allocation recommendations, and develop regulatory standards for fairness testing in financial AI.
- Research agenda: evaluate real-world impact of LLM recommendations on fund flows, study mechanisms causing persistent racial vs gender disparities, and develop domain-specific fairness metrics and correction methods for financial allocation tasks.
Limitations noted by the authors: synthetic audit profiles may differ from complex real-world advisor–client interactions; results depend on prompt framing and model versions (LLMs evolve); observed effects measure associations in model outputs, not deployed institutional outcomes.
Assessment
Claims (10)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| We develop a two-sided audit framework to evaluate demographic bias in LLM-generated investment advice. Other | positive | ability_to_detect_bias_via_framework |
Reading fidelity
high
Study strength
medium
|
not reported
|
| We apply the audit framework to multiple large language models, with GPT-4 Turbo as the primary baseline. Other | positive | model_comparison_and_baseline_evaluation |
Reading fidelity
high
Study strength
medium
|
not reported
|
| On the investor side, fund selections are similar across demographic groups and rely on financial criteria. Task Allocation | null_result | fund_selection |
Reading fidelity
high
Study strength
medium
|
not reported
|
| On the investor side, recommended investment amounts vary when investor names signal race or gender, despite identical age and income. Task Allocation | negative | recommended_investment_amount |
Reading fidelity
high
Study strength
medium
|
not reported
|
| On the fund manager side, capital allocations favor non-Black and male managers. Task Allocation | negative | capital_allocation_to_managers |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Racial disparities in capital allocations persist even under explicit disclosure of manager demographics. Task Allocation | negative | capital_allocation_by_race_with_explicit_disclosure |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Gender-related differences in capital allocation are more pronounced under name-based (implicit) cues than under explicit disclosure. Task Allocation | negative | capital_allocation_by_gender_under_implicit_vs_explicit_signaling |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Bias patterns are qualitatively similar across models, with differences in magnitude between implicit and explicit demographic signaling. Other | mixed | bias_pattern_across_models_and_signaling_modes |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Demographic signals can affect allocation decisions, with effects that tend to be stronger under implicit signaling, potentially replicating existing market inequalities and raising concerns about impartiality in financial advising. Inequality | negative | potential_replication_of_market_inequalities_and_impartiality_concerns |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| The proposed audit framework provides a generalizable approach for identifying and evaluating demographic bias in AI-driven financial advisory systems. Other | positive | generalizability_of_audit_framework |
Reading fidelity
high
Study strength
medium
|
not reported
|