0 cumulative citations
View corpus contextConversational AI is already a major channel for personal finance queries, but users largely retain control: people ask ChatGPT and Gemini for information and advice rather than handing over execution, with true delegation limited mainly to budgeting and tracking.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
As AI increasingly participates in human decision making, understanding how decision-making authority is distributed between humans and AI has become a fundamental behavioural question. We introduce a behavioural measurement framework combining intent and delegated decision authority to quantify what consumers seek from AI and how much decision-making authority they assign to it. Applied to 1.5 million real-world ChatGPT and Gemini interactions from 6,304 users in the United States and India, we find that financial services are already a substantial AI use case. Consumers overwhelmingly use AI to retrieve information and shape financial judgement, while delegation of financial execution remains rare. By shifting attention from conversation topics to delegated decision authority, this work establishes a behavioural baseline for measuring the transition to increasingly agentic AI.
Summary
Main Finding
Conversational AI is already a major channel for personal finance interactions, but users overwhelmingly use it for information and advice (inform & shape decisions). True delegation of financial execution to AI (agentic, transactional actions) is rare in real-world ChatGPT and Gemini use — largely limited to budgeting/monitoring tasks — establishing a behavioural baseline before the wider adoption of agentic financial assistants.
Key Points
- Scale and scope
- Dataset: 1.53M user prompts from 6,304 opt-in users (2,499 US; 3,805 India), covering ChatGPT and Gemini (Aug–Oct 2025).
- After topical segmentation: ~292k US subchats and ~213k India subchats.
- About half of users engaged in finance-related conversations during the study window.
- Behavioural lens: intent × decision authority
- Introduced a two-dimensional framework: (a) behavioural intent (12 financial intents) and (b) decision authority (DA) levels:
- Level 1 Inform: provide information/explanations.
- Level 2 Shape: advisory/analytical influence on choices.
- Level 3 Act: autonomous/transactional execution or automation.
- Most real conversations fall into Inform (Level 1) and Shape (Level 2). Level 3 (Act) conversations are uncommon (<~1% in raw snapshot; higher-agentic intents under 1%).
- Introduced a two-dimensional framework: (a) behavioural intent (12 financial intents) and (b) decision authority (DA) levels:
- Empirical patterns
- Financial services are among the largest topical domains for conversational AI use.
- Common intents: simple retrieval, complex research, product/strategy comparison and optimisation, personal financial analysis and planning — i.e., information and advice roles dominate.
- Delegation/execution (transferring decision-making authority to the model) is rare, primarily observed in budgeting, monitoring, and a few instruction-led execution examples.
- Taxonomy & tagging
- Developed a MECE financial services taxonomy (Investments, Retail Banking & Credit, Tax, Payments, Benefits, Insurance, Business Finance, Finance Infrastructure, etc.).
- Entity extraction yielded 2.7k correctly matched financial-service keywords (plus many unmatched general terms).
Data & Methods
- Data source and preprocessing
- Opt-in sample of ChatGPT and Gemini chat histories from US and India; metadata includes user messages, AI responses, timestamps.
- Non-English India messages (≈144k) translated to English (Google Translate) to standardise modelling.
- Chat segmentation heuristics split long sessions into topical subchats (thresholds and token-overlap rules), producing ~505k subchats total.
- Finance detection
- Task: binary classifier to label subchats as finance vs non-finance.
- Training set: 5.8k annotated subchats (sampled across markets and platforms) seeded with a 2.4k financial keyword dictionary.
- Model: LongFormer-based fine-tuned transformer (chosen for long context; avg subchat ≈2k tokens).
- Performance: Accuracy 96.5%, F1 97.3% on test split.
- Financial-services tagging
- Two-stage approach: LLM-assisted entity extraction (GPT‑4o‑mini) → manual mapping to FS taxonomy.
- Result: 2.7k matched FS keywords; deterministic tagging of finance subchats that mention FS keywords.
- Intent and decision-authority classification
- Intent taxonomy: 12 labels (e.g., Delegated Financial Decision Execution; Instruction-led Execution; Automation & Monitoring; Product/Strategy Optimisation; Comparison; Problem Resolution; Personal Analysis; Planning; Complex Research; Simple Retrieval; Learning; Creation).
- Mapping of intents to DA levels (Inform / Shape / Act).
- Training data: 3.4k manually labelled finance subchats (2.8k organic + 600 synthetic examples from GPT‑5.5 to enrich higher-agentic cases).
- Model: BigBird (sparse-attention Transformer for long contexts), fine-tuned for multi-label classification.
- Note: synthetic augmentation used to partially offset the rarity of high-DA examples in the snapshot.
- Limitations noted by authors
- Sample is not a general-population random sample: skewed demographically (US sample younger, more female, lower income; India sample heavily male and young).
- Translation errors may affect Indian non-English messages.
- Low prevalence of agentic intents in the observed period limits inference about future delegation adoption.
- Use of synthetic examples to train rare intents may bias models toward detecting agentic intents that are rare in real data.
Implications for AI Economics
- Measurement baseline for delegation
- The paper provides an operationalizable behavioural metric (decision authority levels) to measure the gradual transfer of cognitive and decision-making responsibilities from humans to AI — a crucial variable for empirical work on task automation and substitution in services.
- Labor and demand-side effects
- Current dominant use-cases (inform + shape) suggest AI complements human judgement (advisory role) rather than substituting execution-heavy financial roles. This points to demand shifts toward higher-order advisory and monitoring tasks rather than immediate large-scale job displacement in transactional financial roles.
- Product design and market structure
- Low delegation today implies firms can focus on hybrid human-in-the-loop product designs, verification layers, and interfaces that preserve user control while improving decision support. As agentic capabilities increase, product-market fit and value-capture depend on trusted, auditable delegation mechanisms.
- Regulation, liability, and consumer protection
- If delegation increases, regulatory attention must shift from content/accuracy (advice quality) to authorization, execution safety, liability allocation, and disclosure of delegated authority. The DA taxonomy can help regulators define thresholds for when stricter oversight is needed.
- Financial markets and consumer welfare
- Widespread use of AI for research and product optimisation can change consumer search costs, price competition among providers, and the market for paid financial advice. Monitoring whether information/shape uses improve outcomes or introduce correlated errors (e.g., herd behaviour) is essential.
- Research agenda for AI economists
- Track temporal dynamics of DA in conversations to measure adoption of agentic features.
- Quantify effect of AI advice (shape) on actual financial behaviour and outcomes (e.g., savings, investment choices, fee minimisation).
- Model complementarities between algorithmic advice and financial intermediaries; study pricing and liability when AI executes transactions.
- Use the framework to study distributional effects (who delegates? demographic/wealth gradients in DA assignment).
- Policy takeaways
- Policymakers should monitor not just topical AI use but the delegated authority dimension to anticipate when governance, licensing, and consumer-protection frameworks need to be tightened.
If you want, I can extract the paper’s quantitative breakdowns of intents and DA-level shares (per country / per platform), summarize the taxonomy mapping, or produce example conversation archetypes illustrating Inform vs Shape vs Act.
Assessment
Claims (8)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| The study analyzes 1.53 million user prompts from 6,304 users in the United States and India, covering ChatGPT and Gemini histories collected during August–October 2025. Other | positive | Scale of the human–AI financial decision-making corpus |
Reading fidelity
high
Study strength
medium
|
n=6304
1.53M user prompts
|
| Approximately half of the users in the study engaged in finance-related conversations with conversational AI during the study period. Adoption Rate | positive | Share of users engaging with AI about financial services or money-related issues |
Reading fidelity
high
Study strength
medium
|
n=6304
approximately half of users
|
| Consumers predominantly use conversational AI for financial information retrieval and for shaping financial judgment, rather than for executing financial transactions. Decision Quality | mixed | Distribution of financial AI interactions across information, advisory, and execution roles |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Delegation of financial execution to AI is rare and is largely confined to budgeting and financial tracking. Task Allocation | negative | Prevalence and domain concentration of delegated financial execution |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The dataset contained fewer than 1% of conversations with intents involving delegated financial decision execution, instruction-led financial execution, or financial automation and monitoring. Automation Exposure | negative | Prevalence of highly agentic financial intents |
Reading fidelity
high
Study strength
medium
|
under 1%
|
| The finance-conversation classifier achieved 96.5 accuracy and an F1 score of 97.3 on its held-out test set. Other | positive | Finance-subchat classification performance |
Reading fidelity
high
Study strength
high
|
n=5800
Accuracy = 96.5; F1 = 97.3
|
| The US and India samples differ substantially from their respective general-population benchmarks in age, gender, income, and employment characteristics. Other | mixed | Sample representativeness and demographic composition |
Reading fidelity
high
Study strength
medium
|
n=6304
US: 58.2% aged 18–34 versus 29.0% GenPop; India: 89.4% aged 18–34 versus 41.2% GenPop
|
| The financial-intent training dataset contained 3,400 manually annotated subchats, including 2,800 real conversations and 600 synthetic examples. Other | positive | Size and composition of the financial-intent classification training dataset |
Reading fidelity
high
Study strength
medium
|
n=3400
2.8k real conversations plus 600 synthetic examples
|