The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References About 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

People rate certified-financial-planner–style advice higher than AI-style advice across readability, trust and willingness to rely, but mislabeling AI as expert raises perceived fit and overall quality; correct disclosure adds little beyond message cues while inaccurate attribution distorts users' trust calibration.

How People Evaluate AI-, Expert-, and Peer-Style Financial Advice
Aryan Ramchandra Kapadia, Eshwar Chandrasekharan, Koustuv Saha · August 10, 2026
arxiv rct medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Aryan Ramchandra Kapadia unresolved corpus identity
  2. Eshwar Chandrasekharan unresolved corpus identity
  3. Koustuv Saha unresolved corpus identity

Semantic Scholar

Latest observation:

  1. A. Kapadia provider ID
  2. Eshwar Chandrasekharan provider ID
  3. Koustuv Saha provider ID
In a preregistered randomized vignette experiment (N=285), expert-style financial advice was rated higher than AI-style on most dimensions, and displayed source attribution (especially mislabeling) selectively shifted evaluations, with AI-style content most sensitive to labels.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

As generative AI increasingly becomes a common source of daily decision-making, including financial choices, it is critical to understand how people evaluate AI-generated financial advice. We conducted a preregistered vignette experiment (N = 285) in which substantive financial content---including facts, numerical values, recommendation direction, and core reasoning---was held constant while communication style varied across AI Financial Assistant (AI), Certified Financial Planner (Expert), and Online Community Forum (OC) advice. Displayed source attribution was independently manipulated through correctly labeled, unlabeled, and mislabeled conditions, allowing us to separate attribution effects from source-specific communication cues. Expert advice was rated more favorably than AI advice on 9 of 10 outcomes (|d|=0.20--0.47), and this advantage remained visible without source labels, where Expert advice outperformed AI advice on 8 of 10 outcomes (up to d=0.60). Correct labels added limited differentiation, whereas mislabeling increased ratings of AI advice for situational fit and overall quality (d=0.42 for each) and attenuated the Expert advantage in situational fit (d=-0.36). Descriptive analyses further showed that AI advice was most responsive to displayed attribution and, conversely, that advice-style differences were most visible under an AI label. These findings show that financial-advice evaluations are shaped jointly by displayed attribution and message-level communication cues. We position disclosure not as a neutral transparency mechanism, but as an interpretive frame whose accuracy and interaction with message cues can shape trust and reliance.

Summary

Main Finding

When the substantive financial content is held constant, people rate expert-style financial advice substantially higher than AI-style advice across most dimensions (presentation, perceived safety, authority, trust, and willingness to rely). These differences largely reflect message-level communication cues (tone, structure, reasoning format) rather than labels alone. Correct attribution adds little beyond those cues, but incorrect attribution (mislabeling) can meaningfully distort evaluations — especially boosting AI-style content when it is mislabeled as expert — and AI-labeled advice tends to elicit greater scrutiny.

Key Points

  • Experimental result summary:
    • Expert-style advice outperformed AI-style advice on 9 of 10 outcome measures (Cohen’s d ≈ 0.20–0.47).
    • Even when no source label was shown (unlabeled), Expert > AI on 8 of 10 outcomes; the largest unlabeled effect was d = 0.60 for situational fit.
    • Correct source labels provided limited additional differentiation beyond message cues.
    • Mislabeling produced selective shifts: AI advice shown with a non-AI label increased situational fit and overall quality ratings (d = 0.42 for each).
    • Presenting AI-style advice with an Expert label (vs. unlabeled) improved perceived situational fit and overall quality by d ≈ 0.47.
  • Dimension-specific findings:
    • Readability: Expert > AI (d ≈ 0.47); Online Community (OC) > AI (d ≈ 0.53).
    • Risk/harm perception: Expert seen as safer than AI (harm risk d ≈ −0.27); AI judged more likely to mislead novices than Expert (d ≈ −0.30).
    • Source authority & behavioral intentions: Expert > AI on perceived knowledge, overall quality, trust intention, and reliance intention (ds ≈ 0.20–0.27).
    • OC advice was rated less knowledgeable than AI (d ≈ −0.19) despite higher readability on some measures.
  • Interaction patterns:
    • AI-style content is most sensitive to the displayed source label.
    • Advice-style differences are most salient when advice is labeled as AI — suggesting AI labels trigger increased scrutiny.
  • Interpretation:
    • The “AI evaluation gap” largely reflects how messages are written/formatted; disclosure affects interpretation but is not purely neutral — inaccurate labels can mislead evaluative calibration.

Data & Methods

  • Design: Preregistered vignette experiment (N = 285 U.S. adults recruited on Prolific).
  • Stimuli:
    • Eight realistic personal-finance scenarios varying stakes, external uncertainty, and verifiability (2×2×2 design; participants saw 4 scenarios).
    • Three advice-style treatments (same substantive content across treatments): AI-style (neutral, analytical), Expert-style (structured, principle-based), Online Community (OC)-style (informal, experience-based). Message-level communication cues varied while facts, numbers, recommendation direction, and core reasoning were held constant.
    • Source-labeling manipulation with three arms: correctly labeled, unlabeled, and mislabeled (label swaps).
  • Measures: Ten 7-point Likert items after each vignette assessing readability, reasoning clarity, situational fit, risk acknowledgment, financial harm risk, misleadingness, perceived source knowledge, overall quality, trust intention, and reliance intention.
  • Analysis: Mixed-effects regressions with participant random intercepts, scenario fixed effects, and scenario familiarity as a covariate; descriptive nonparametric sensitivity analyses (Kruskal–Wallis) for label and style sensitivity.
  • Limitations noted by authors: bundled manipulation of style (cannot isolate single cue), vignette-based self-reports (not real financial behavior), constructed stimuli (not necessarily representative of all real-world sources), U.S. Prolific sample, and exploratory/descriptive sensitivity analyses not confirmatory.

Implications for AI Economics

  • Market demand and competition:
    • Communication style functions as a quality signal. Providers (human advisors and AI products) may compete on message design (structure, explicit reasoning, readability) as much as on provenance. AI firms that adopt expert-like communication patterns may capture more trust and market share without changing underlying model accuracy.
    • Mislabeling (intentional or accidental) can distort consumer demand and reduce effective competition: labeling AI advice as human expert could artificially inflate demand for lower-cost AI solutions, while presenting human advice as AI could depress perceived quality.
  • Information asymmetries and signaling:
    • Labels are imperfect signals and disclosure accuracy is crucial. Inaccurate provenance can lead to miscalibration of trust and reliance, increasing the risk of consumer harm in financial markets.
    • Designers and regulators should treat source attribution as an interpretive frame rather than a neutral tag; poor attribution interacts with message cues to change perceived risk and reliance.
  • Consumer welfare and financial stability:
    • Over- or under-reliance driven by label-driven framing or stylistic cues (rather than substantive quality) can worsen consumer outcomes, particularly for less financially literate users who rely on perceived authority and readability.
    • Because AI labels appear to heighten scrutiny, accurate AI labelling combined with transparent reasoning could improve calibration (reduce unwarranted deference and inappropriate trust), potentially lowering misinformed financial actions.
  • Policy and regulatory implications:
    • Disclosure rules should emphasize accuracy of provenance and require more than a source label — e.g., standardized reasoning disclosures, assumptions, and verifiable calculations for financial recommendations.
    • Anti-mislabeling enforcement matters: the asymmetric harm from incorrect labels (larger distortions than any benefit from correct labels) argues for regulatory attention to provenance fidelity and penalties for deceptive attribution.
    • Standards for “explainability” in financial advice should prioritize actionable transparency (explicit assumptions, verifiable claims) because message cues materially affect evaluations and reliance.
  • Pricing and contracting:
    • Human advisors might justify price premia partly via communication-style signaling (principled, structured explanations). If AI providers can credibly mimic those styles and pair them with transparent, auditable reasoning, price competition could intensify.
    • Contracts and certification (e.g., third-party verification of provenance and reasoning) could become valuable services to reduce information frictions and support proper trust calibration.
  • Research and monitoring priorities:
    • Empirical work should measure downstream behavioral impacts (actual financial choices and welfare), heterogeneity by financial literacy, and long-term market effects of stylistic imitation and labeling practices.
    • Monitoring for strategic mislabeling and its effects on consumer finance outcomes should be a priority for regulators and platform governance.

Overall, the paper shows that AI’s competitive position in financial advice markets depends critically on communication style and correct attribution; policies that require accurate provenance plus richer, standardized transparency about reasoning will better align trust and reliance with true advice quality.

Assessment

Paper Typerct Evidence Strengthmedium — Random assignment of labels and counterbalancing of advice versions provides credible causal identification for labeling and style effects within the experiment, and the study was preregistered; however, outcomes are self-reported vignette ratings (not real behavior), stimuli are experimentally constructed rather than observed real-world AI outputs, the sample is a U.S. Prolific convenience sample, and ecological/generalizability limitations reduce external validity. Methods Rigormedium — Strong internal design choices: preregistration, careful parity of substantive content across styles, randomization, attention checks, and appropriate mixed-effects modeling; weaknesses include bundled manipulation of many message features (so individual drivers are unidentified), reliance on hypothetical text vignettes and intentions rather than behavioral outcomes, modest sample size, and a single online convenience sample. SampleN = 285 U.S. adults recruited via Prolific after exclusions (incomplete responses, failed attention checks, very fast completions); each participant evaluated four vignettes drawn from eight personal-finance scenarios; participants were assigned to one of three labeling arms and saw all three advice styles at least once (one style twice); demographic breakdown reported in paper Table 1. Themeshuman_ai_collab adoption IdentificationRandomized vignette experiment: participants were randomly assigned to labeling arms (correctly labeled, unlabeled, mislabeled) and to counterbalanced advice-version orders; substantive financial content was held constant while communication style varied; analyses use mixed-effects regressions with participant random intercepts and scenario fixed effects (and covariates) to estimate causal effects of displayed attribution and message-style on ratings. GeneralizabilityProlific U.S. convenience sample limits representativeness across countries, demographics, and real-world clients, Vignette/text-only stimuli may not capture richer multimodal or interactive AI assistant behavior, Constructed advice preserved content parity but may not reflect diversity/quality of real AI or human professional outputs, Outcomes are self-reported evaluations and intentions, not observed financial decisions or downstream outcomes, Bundled manipulation of multiple message features prevents isolating which aspects of communication drive effects

Claims (11)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Expert-style financial advice was rated more favorably than AI-style advice on 9 of 10 evaluation outcomes. Decision Quality positive Relative evaluations of expert- versus AI-style advice across readability, reasoning clarity, situational fit, risk acknowledgment, financial harm risk, misleadingness, perceived source knowledge, overall quality, trust intention, and reliance intention
Reading fidelity high
Study strength medium
n=285
|d| = 0.20–0.47
0.6
Expert-style advice was rated as more readable than AI-style advice. Output Quality positive Perceived readability of financial advice
Reading fidelity high
Study strength medium
n=285
d = 0.47
0.6
Expert-style advice was perceived as a better fit for the protagonist's situation than AI-style advice. Decision Quality positive Perceived situational fit of the advice
Reading fidelity high
Study strength medium
n=285
d = 0.39
0.6
Expert-style advice was perceived as posing less risk of financial harm than AI-style advice. Ai Safety And Ethics negative Perceived risk that the advice would cause financial harm
Reading fidelity high
Study strength medium
n=285
d = −0.27
0.6
AI-style advice was rated as more likely than expert-style advice to mislead a person with limited financial knowledge. Ai Safety And Ethics negative Perceived likelihood that advice would mislead a financially inexperienced person
Reading fidelity high
Study strength medium
n=285
d = −0.30
0.6
Compared with AI-style advice, expert-style advice received higher ratings for perceived source knowledge, overall quality, trust intention, and reliance intention. Worker Satisfaction positive Perceived source knowledge, overall advice quality, intention to trust, and intention to rely on the advice
Reading fidelity high
Study strength medium
n=285
d = 0.20–0.27
0.6
Even without source labels, expert-style advice was evaluated more favorably than AI-style advice on 8 of 10 dimensions. Decision Quality positive Differences in advice evaluations based on message-level communication cues without explicit source attribution
Reading fidelity high
Study strength medium
n=285
|d| = 0.27–0.60
0.6
Correct source labels added limited explanatory value beyond the communication cues in the advice messages. Governance And Regulation null_result Incremental effect of accurate source labels on advice evaluations
Reading fidelity high
Study strength medium
n=285
d = 0.40 for situational fit
0.6
Mislabeling AI-style advice with a non-AI source label increased ratings of its situational fit and overall quality. Output Quality positive Perceived situational fit and overall quality of AI-style financial advice
Reading fidelity high
Study strength medium
n=285
d = 0.42 for each
0.6
Mislabeling attenuated the expert-style advantage over AI-style advice in situational fit. Decision Quality negative Difference between expert- and AI-style advice in perceived situational fit under incorrect labeling
Reading fidelity high
Study strength medium
n=285
d = −0.36
0.6
AI-style advice was more responsive to displayed source attribution than expert-style advice in the descriptive label-sensitivity analysis. Decision Quality positive Changes in situational-fit and overall-quality ratings for the same AI-style advice under different displayed labels
Reading fidelity high
Study strength low
n=285
d = 0.47 for situational fit and d = 0.47 for overall quality
0.3

Notes