0 cumulative citations
View corpus contextPeople rate certified-financial-planner–style advice higher than AI-style advice across readability, trust and willingness to rely, but mislabeling AI as expert raises perceived fit and overall quality; correct disclosure adds little beyond message cues while inaccurate attribution distorts users' trust calibration.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
As generative AI increasingly becomes a common source of daily decision-making, including financial choices, it is critical to understand how people evaluate AI-generated financial advice. We conducted a preregistered vignette experiment (N = 285) in which substantive financial content---including facts, numerical values, recommendation direction, and core reasoning---was held constant while communication style varied across AI Financial Assistant (AI), Certified Financial Planner (Expert), and Online Community Forum (OC) advice. Displayed source attribution was independently manipulated through correctly labeled, unlabeled, and mislabeled conditions, allowing us to separate attribution effects from source-specific communication cues. Expert advice was rated more favorably than AI advice on 9 of 10 outcomes (|d|=0.20--0.47), and this advantage remained visible without source labels, where Expert advice outperformed AI advice on 8 of 10 outcomes (up to d=0.60). Correct labels added limited differentiation, whereas mislabeling increased ratings of AI advice for situational fit and overall quality (d=0.42 for each) and attenuated the Expert advantage in situational fit (d=-0.36). Descriptive analyses further showed that AI advice was most responsive to displayed attribution and, conversely, that advice-style differences were most visible under an AI label. These findings show that financial-advice evaluations are shaped jointly by displayed attribution and message-level communication cues. We position disclosure not as a neutral transparency mechanism, but as an interpretive frame whose accuracy and interaction with message cues can shape trust and reliance.
Summary
Main Finding
When the substantive financial content is held constant, people rate expert-style financial advice substantially higher than AI-style advice across most dimensions (presentation, perceived safety, authority, trust, and willingness to rely). These differences largely reflect message-level communication cues (tone, structure, reasoning format) rather than labels alone. Correct attribution adds little beyond those cues, but incorrect attribution (mislabeling) can meaningfully distort evaluations — especially boosting AI-style content when it is mislabeled as expert — and AI-labeled advice tends to elicit greater scrutiny.
Key Points
- Experimental result summary:
- Expert-style advice outperformed AI-style advice on 9 of 10 outcome measures (Cohen’s d ≈ 0.20–0.47).
- Even when no source label was shown (unlabeled), Expert > AI on 8 of 10 outcomes; the largest unlabeled effect was d = 0.60 for situational fit.
- Correct source labels provided limited additional differentiation beyond message cues.
- Mislabeling produced selective shifts: AI advice shown with a non-AI label increased situational fit and overall quality ratings (d = 0.42 for each).
- Presenting AI-style advice with an Expert label (vs. unlabeled) improved perceived situational fit and overall quality by d ≈ 0.47.
- Dimension-specific findings:
- Readability: Expert > AI (d ≈ 0.47); Online Community (OC) > AI (d ≈ 0.53).
- Risk/harm perception: Expert seen as safer than AI (harm risk d ≈ −0.27); AI judged more likely to mislead novices than Expert (d ≈ −0.30).
- Source authority & behavioral intentions: Expert > AI on perceived knowledge, overall quality, trust intention, and reliance intention (ds ≈ 0.20–0.27).
- OC advice was rated less knowledgeable than AI (d ≈ −0.19) despite higher readability on some measures.
- Interaction patterns:
- AI-style content is most sensitive to the displayed source label.
- Advice-style differences are most salient when advice is labeled as AI — suggesting AI labels trigger increased scrutiny.
- Interpretation:
- The “AI evaluation gap” largely reflects how messages are written/formatted; disclosure affects interpretation but is not purely neutral — inaccurate labels can mislead evaluative calibration.
Data & Methods
- Design: Preregistered vignette experiment (N = 285 U.S. adults recruited on Prolific).
- Stimuli:
- Eight realistic personal-finance scenarios varying stakes, external uncertainty, and verifiability (2×2×2 design; participants saw 4 scenarios).
- Three advice-style treatments (same substantive content across treatments): AI-style (neutral, analytical), Expert-style (structured, principle-based), Online Community (OC)-style (informal, experience-based). Message-level communication cues varied while facts, numbers, recommendation direction, and core reasoning were held constant.
- Source-labeling manipulation with three arms: correctly labeled, unlabeled, and mislabeled (label swaps).
- Measures: Ten 7-point Likert items after each vignette assessing readability, reasoning clarity, situational fit, risk acknowledgment, financial harm risk, misleadingness, perceived source knowledge, overall quality, trust intention, and reliance intention.
- Analysis: Mixed-effects regressions with participant random intercepts, scenario fixed effects, and scenario familiarity as a covariate; descriptive nonparametric sensitivity analyses (Kruskal–Wallis) for label and style sensitivity.
- Limitations noted by authors: bundled manipulation of style (cannot isolate single cue), vignette-based self-reports (not real financial behavior), constructed stimuli (not necessarily representative of all real-world sources), U.S. Prolific sample, and exploratory/descriptive sensitivity analyses not confirmatory.
Implications for AI Economics
- Market demand and competition:
- Communication style functions as a quality signal. Providers (human advisors and AI products) may compete on message design (structure, explicit reasoning, readability) as much as on provenance. AI firms that adopt expert-like communication patterns may capture more trust and market share without changing underlying model accuracy.
- Mislabeling (intentional or accidental) can distort consumer demand and reduce effective competition: labeling AI advice as human expert could artificially inflate demand for lower-cost AI solutions, while presenting human advice as AI could depress perceived quality.
- Information asymmetries and signaling:
- Labels are imperfect signals and disclosure accuracy is crucial. Inaccurate provenance can lead to miscalibration of trust and reliance, increasing the risk of consumer harm in financial markets.
- Designers and regulators should treat source attribution as an interpretive frame rather than a neutral tag; poor attribution interacts with message cues to change perceived risk and reliance.
- Consumer welfare and financial stability:
- Over- or under-reliance driven by label-driven framing or stylistic cues (rather than substantive quality) can worsen consumer outcomes, particularly for less financially literate users who rely on perceived authority and readability.
- Because AI labels appear to heighten scrutiny, accurate AI labelling combined with transparent reasoning could improve calibration (reduce unwarranted deference and inappropriate trust), potentially lowering misinformed financial actions.
- Policy and regulatory implications:
- Disclosure rules should emphasize accuracy of provenance and require more than a source label — e.g., standardized reasoning disclosures, assumptions, and verifiable calculations for financial recommendations.
- Anti-mislabeling enforcement matters: the asymmetric harm from incorrect labels (larger distortions than any benefit from correct labels) argues for regulatory attention to provenance fidelity and penalties for deceptive attribution.
- Standards for “explainability” in financial advice should prioritize actionable transparency (explicit assumptions, verifiable claims) because message cues materially affect evaluations and reliance.
- Pricing and contracting:
- Human advisors might justify price premia partly via communication-style signaling (principled, structured explanations). If AI providers can credibly mimic those styles and pair them with transparent, auditable reasoning, price competition could intensify.
- Contracts and certification (e.g., third-party verification of provenance and reasoning) could become valuable services to reduce information frictions and support proper trust calibration.
- Research and monitoring priorities:
- Empirical work should measure downstream behavioral impacts (actual financial choices and welfare), heterogeneity by financial literacy, and long-term market effects of stylistic imitation and labeling practices.
- Monitoring for strategic mislabeling and its effects on consumer finance outcomes should be a priority for regulators and platform governance.
Overall, the paper shows that AI’s competitive position in financial advice markets depends critically on communication style and correct attribution; policies that require accurate provenance plus richer, standardized transparency about reasoning will better align trust and reliance with true advice quality.
Assessment
Claims (11)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Expert-style financial advice was rated more favorably than AI-style advice on 9 of 10 evaluation outcomes. Decision Quality | positive | Relative evaluations of expert- versus AI-style advice across readability, reasoning clarity, situational fit, risk acknowledgment, financial harm risk, misleadingness, perceived source knowledge, overall quality, trust intention, and reliance intention |
Reading fidelity
high
Study strength
medium
|
n=285
|d| = 0.20–0.47
|
| Expert-style advice was rated as more readable than AI-style advice. Output Quality | positive | Perceived readability of financial advice |
Reading fidelity
high
Study strength
medium
|
n=285
d = 0.47
|
| Expert-style advice was perceived as a better fit for the protagonist's situation than AI-style advice. Decision Quality | positive | Perceived situational fit of the advice |
Reading fidelity
high
Study strength
medium
|
n=285
d = 0.39
|
| Expert-style advice was perceived as posing less risk of financial harm than AI-style advice. Ai Safety And Ethics | negative | Perceived risk that the advice would cause financial harm |
Reading fidelity
high
Study strength
medium
|
n=285
d = −0.27
|
| AI-style advice was rated as more likely than expert-style advice to mislead a person with limited financial knowledge. Ai Safety And Ethics | negative | Perceived likelihood that advice would mislead a financially inexperienced person |
Reading fidelity
high
Study strength
medium
|
n=285
d = −0.30
|
| Compared with AI-style advice, expert-style advice received higher ratings for perceived source knowledge, overall quality, trust intention, and reliance intention. Worker Satisfaction | positive | Perceived source knowledge, overall advice quality, intention to trust, and intention to rely on the advice |
Reading fidelity
high
Study strength
medium
|
n=285
d = 0.20–0.27
|
| Even without source labels, expert-style advice was evaluated more favorably than AI-style advice on 8 of 10 dimensions. Decision Quality | positive | Differences in advice evaluations based on message-level communication cues without explicit source attribution |
Reading fidelity
high
Study strength
medium
|
n=285
|d| = 0.27–0.60
|
| Correct source labels added limited explanatory value beyond the communication cues in the advice messages. Governance And Regulation | null_result | Incremental effect of accurate source labels on advice evaluations |
Reading fidelity
high
Study strength
medium
|
n=285
d = 0.40 for situational fit
|
| Mislabeling AI-style advice with a non-AI source label increased ratings of its situational fit and overall quality. Output Quality | positive | Perceived situational fit and overall quality of AI-style financial advice |
Reading fidelity
high
Study strength
medium
|
n=285
d = 0.42 for each
|
| Mislabeling attenuated the expert-style advantage over AI-style advice in situational fit. Decision Quality | negative | Difference between expert- and AI-style advice in perceived situational fit under incorrect labeling |
Reading fidelity
high
Study strength
medium
|
n=285
d = −0.36
|
| AI-style advice was more responsive to displayed source attribution than expert-style advice in the descriptive label-sensitivity analysis. Decision Quality | positive | Changes in situational-fit and overall-quality ratings for the same AI-style advice under different displayed labels |
Reading fidelity
high
Study strength
low
|
n=285
d = 0.47 for situational fit and d = 0.47 for overall quality
|