The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Language models steer money unevenly: while fund picks are broadly consistent across investor types, LLMs recommend different investment sizes and allocate more capital to non-Black and male managers—biases that are often stronger when demographics are signaled implicitly by names.

Who Invests, Who Gets Funded: Gender and Racial Bias in LLM-Generated Investment Advice
Ye Emma Wang, Kexin Gu · February 06, 2026 · Journal of Business Ethics
openalex quasi_experimental medium evidence 8/10 relevance Full text usable extracted full text DOI Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

OpenAlex

Latest observation:

  1. Ye Emma Wang provider ID
  2. Kexin Gu provider ID
Controlled audits of multiple LLMs show similar fund choices across investor demographics but systematically different recommended investment amounts and capital allocations that favor non-Black and male managers, with stronger effects when demographics are signaled implicitly by names.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Abstract Do large language models (LLMs) generate unbiased financial advice across investor and fund manager demographics? We develop a two-sided audit framework to evaluate demographic bias in LLM-generated investment advice and apply it to multiple large language models, with GPT-4 Turbo as the primary baseline. On the investor side, fund selections are similar across demographic groups and rely on financial criteria, but recommended investment amounts vary when investor names signal race or gender, despite identical age and income. On the fund manager side, capital allocations favor non-Black and male managers: racial disparities persist even under explicit disclosure, while gender-related differences are more pronounced under name-based cues. Bias patterns are qualitatively similar across models, with differences in magnitude between implicit and explicit demographic signaling. These results suggest that, even when LLMs incorporate core financial reasoning, demographic signals can affect allocation decisions, with effects that tend to be stronger under implicit signaling, potentially replicating existing market inequalities and raising concerns about impartiality in financial advising. The proposed audit framework provides a generalizable approach for identifying and evaluating demographic bias in AI-driven financial advisory systems.

Summary

Main Finding

LLMs can reproduce demographic disparities in investment advice even when financial fundamentals are held constant. Using a two-sided audit (investor-side and fund-manager-side), the authors find that fund selection is largely driven by financial criteria, but recommended investment amounts and capital allocations vary by demographic signals: investor names that imply race or gender affect recommended amounts, and fund managers who are Black or implicitly signaled as female receive systematically lower allocations. Racial disparities persist even with explicit disclosure; gender disparities are stronger under implicit (name-based) signaling. Results generalize qualitatively across multiple LLMs but vary in magnitude and sign.

Key Points

  • Two-sided audit framework: tests both (a) investor-side effects (fund choice and recommended investment amount) and (b) fund-manager-side effects (capital allocated to managers).
  • Signal manipulation: demographic cues introduced either implicitly (names) or explicitly (stated race/gender) while holding age, income, risk tolerance, and other financial attributes constant.
  • Investor-side: fund selection (which funds are recommended) is similar across demographic groups and follows financial metrics; however, recommended investment amounts vary by inferred race/gender from names.
  • Fund-manager-side:
    • Racial bias: Black fund managers receive lower recommended allocations than White counterparts; effect persists under explicit disclosure of race.
    • Gender bias: Female managers receive lower allocations when gender is implied via name, but explicit disclosure of gender often removes statistically significant differences.
    • Implicit signals tend to produce stronger bias effects than explicit disclosures.
  • Model heterogeneity: GPT-4 Turbo used as baseline; additional tests with GPT-4.1, GPT-4o, Claude 3.5 Sonnet, and Llama 3.1 8B show similar qualitative patterns but differ in magnitude and sometimes sign.
  • Robustness: Enriching prompts (adding horizon, risk tolerance, long-term criteria) and varying age/income shows many disparities remain—especially for younger or lower-income profiles.
  • Ethical and fiduciary concerns: demographic-driven allocation changes raise questions about impartiality, duty of care, and the legitimacy of deploying LLMs in financial advice.

Data & Methods

  • Experimental design: controlled audit study that systematically manipulates demographic signals while keeping financial profiles identical. Two complementary experiments:
    • Investor-side: present identical investor financial profiles with varied names (and sometimes explicit demographic labels) and record recommended funds and investment amounts.
    • Fund-manager-side: present fund descriptions with manager identity signaled via name or explicit race/gender and record capital allocation recommendations.
  • Models tested: primary baseline GPT-4 Turbo (chosen for performance/efficiency reproducibility) and replication on GPT-4.1, GPT-4o, Claude 3.5 Sonnet, Llama 3.1 8B.
  • Signal types: implicit (names chosen to signal race/gender) vs explicit (directly stated race/gender).
  • Robustness checks: enriched prompts (risk tolerance, horizon, objectives), varying investor age and income, long-term evaluation criteria for managers.
  • Statistical analysis: comparisons of selections and allocation amounts across demographic treatments; significance testing reported for observed disparities (paper reports statistically significant racial effects and context-dependent gender effects).
  • Conceptual grounding: links to audit-study methodology from labor, lending, housing; draws on behavioral finance insights (structured vs open-ended decisions) and literature on algorithmic fairness, RLHF, and test-time control.

Implications for AI Economics

  • Distributional effects on capital: If deployed at scale, LLM-driven advisory tools could systematically redirect capital away from Black managers and certain demographic investor groups, reinforcing existing market inequalities in fund flows, fundraising, and wealth accumulation.
  • Market efficiency and allocative justice: Demographic-driven allocation distortions can reduce allocative efficiency and block potentially valuable managers from capital, with long-run impacts on competition, innovation, and returns.
  • Fiduciary and regulatory risks: Financial institutions using LLMs may face fiduciary duty breaches (duty of care/loyalty) if advice is influenced by irrelevant demographic cues. Regulators may demand audits, transparency, or constraints on model use in advisory roles.
  • Model-specific risk: Heterogeneity across models implies institutions could inadvertently introduce model-specific biases; audits must be model- and prompt-specific rather than relying on generic claims of neutrality.
  • Design and governance responses:
    • Operational: implement regular two-sided audits, sanitize demographic cues in inputs where appropriate, enforce standardized structured decision protocols that reduce open-ended allocations, and require human-in-the-loop review for allocation decisions.
    • Technical: incorporate fairness objectives into RLHF, apply counterfactual or constraint-based post-processing to allocations, and develop test-time controls to neutralize implicit demographic associations.
    • Policy: mandate disclosure of model use in advisory contexts, require provenance/ audit logs for allocation recommendations, and develop regulatory standards for fairness testing in financial AI.
  • Research agenda: evaluate real-world impact of LLM recommendations on fund flows, study mechanisms causing persistent racial vs gender disparities, and develop domain-specific fairness metrics and correction methods for financial allocation tasks.

Limitations noted by the authors: synthetic audit profiles may differ from complex real-world advisor–client interactions; results depend on prompt framing and model versions (LLMs evolve); observed effects measure associations in model outputs, not deployed institutional outcomes.

Assessment

Paper Typequasi_experimental Evidence Strengthmedium — The controlled manipulations provide credible causal evidence about how LLM outputs respond to demographic signals within the tested models, and results are consistent across models and implicit/explicit cue conditions; however, strength is limited by uncertainty about prompt sensitivity, the representativeness of the scenarios and funds tested, possible multiple-testing/selection issues, and absence of evidence linking model outputs to real-world investor behavior or market outcomes. Methods Rigormedium — The study employs a systematic, two-sided audit framework with controlled variation and multi-model comparisons—strong design features for probing model behavior—but the paper (based on the abstract) does not report details on sample size, robustness checks for prompt phrasing, the scope of funds/managers tested, or corrections for multiple comparisons, and external validation (e.g., human-in-the-loop or market data) appears absent. SampleSynthetic investor profiles (identical age and income) with varied demographic signals (name-based implicit cues and explicit demographic disclosure), synthetic fund manager profiles varied by race and gender signals, a set of candidate funds and manager attributes used for allocation tasks, and outputs collected from multiple LLMs with GPT-4 Turbo as the primary baseline; outcomes measured include fund selection, recommended investment amounts, and capital allocation recommendations. Themesinequality governance IdentificationTwo-sided audit using controlled prompt interventions: the authors hold investors' financial attributes constant while varying demographic signals (implicit name-based cues and explicit disclosures) for both investors and fund managers, then compare LLM outputs (fund choice, recommended investment amount, capital allocation) across demographic conditions; they repeat this across multiple LLMs (GPT-4 Turbo primary) and use statistical comparisons to attribute differences in recommendations to the demographic cues. GeneralizabilityResults are specific to the particular LLM versions and may change as models are updated., Synthetic prompt scenarios may not capture full complexity of real advisor-client interactions or platform UX., Name-based cues are an imperfect proxy for race/gender in real-world settings and cultural contexts., Limited set of funds/managers and tested scenarios may not represent the broader investment universe., Findings on model outputs do not automatically translate to real-world investment flows or market-level impacts.

Claims (10)

ClaimDirectionOutcomeConfidence & EvidenceDetails
We develop a two-sided audit framework to evaluate demographic bias in LLM-generated investment advice. Other positive ability_to_detect_bias_via_framework
Reading fidelity high
Study strength medium
not reported
0.48
We apply the audit framework to multiple large language models, with GPT-4 Turbo as the primary baseline. Other positive model_comparison_and_baseline_evaluation
Reading fidelity high
Study strength medium
not reported
0.48
On the investor side, fund selections are similar across demographic groups and rely on financial criteria. Task Allocation null_result fund_selection
Reading fidelity high
Study strength medium
not reported
0.48
On the investor side, recommended investment amounts vary when investor names signal race or gender, despite identical age and income. Task Allocation negative recommended_investment_amount
Reading fidelity high
Study strength medium
not reported
0.48
On the fund manager side, capital allocations favor non-Black and male managers. Task Allocation negative capital_allocation_to_managers
Reading fidelity high
Study strength medium
not reported
0.48
Racial disparities in capital allocations persist even under explicit disclosure of manager demographics. Task Allocation negative capital_allocation_by_race_with_explicit_disclosure
Reading fidelity high
Study strength medium
not reported
0.48
Gender-related differences in capital allocation are more pronounced under name-based (implicit) cues than under explicit disclosure. Task Allocation negative capital_allocation_by_gender_under_implicit_vs_explicit_signaling
Reading fidelity high
Study strength medium
not reported
0.48
Bias patterns are qualitatively similar across models, with differences in magnitude between implicit and explicit demographic signaling. Other mixed bias_pattern_across_models_and_signaling_modes
Reading fidelity high
Study strength medium
not reported
0.48
Demographic signals can affect allocation decisions, with effects that tend to be stronger under implicit signaling, potentially replicating existing market inequalities and raising concerns about impartiality in financial advising. Inequality negative potential_replication_of_market_inequalities_and_impartiality_concerns
Reading fidelity high
Study strength speculative
not reported
0.08
The proposed audit framework provides a generalizable approach for identifying and evaluating demographic bias in AI-driven financial advisory systems. Other positive generalizability_of_audit_framework
Reading fidelity high
Study strength medium
not reported
0.48

Notes