The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Conversational AI is already a major channel for personal finance queries, but users largely retain control: people ask ChatGPT and Gemini for information and advice rather than handing over execution, with true delegation limited mainly to budgeting and tracking.

From Information to Delegation: Mapping Human-AI Financial Decision Making
Iman Munire Bilal, Yingcan Carol Wang, Ajan Raj, Filippo Giovagnini, Pranav Tewari, Yuwei Zhang, Mei-Chen Zoe Liou, Qamar Zaman · August 03, 2026
arxiv descriptive medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Iman Munire Bilal unresolved corpus identity
  2. Yingcan Carol Wang unresolved corpus identity
  3. Ajan Raj unresolved corpus identity
  4. Filippo Giovagnini unresolved corpus identity
  5. Pranav Tewari unresolved corpus identity
  6. Yuwei Zhang unresolved corpus identity
  7. Mei-Chen Zoe Liou unresolved corpus identity
  8. Qamar Zaman unresolved corpus identity

Semantic Scholar

Latest observation:

  1. I. Bilal provider ID
  2. Yingcan Wang provider ID
  3. Ajan Raj provider ID
  4. Filippo Giovagnini provider ID
  5. Pranav Tewari provider ID
  6. Yuwei Zhang provider ID
  7. Mei-Chen Zoe Liou provider ID
  8. Qamar Zaman provider ID
Using 1.5 million ChatGPT and Gemini interactions from US and India users, the paper finds conversational AI is commonly used to retrieve information and shape financial judgments, while delegation of transactional execution is rare and concentrated in budgeting/monitoring.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

As AI increasingly participates in human decision making, understanding how decision-making authority is distributed between humans and AI has become a fundamental behavioural question. We introduce a behavioural measurement framework combining intent and delegated decision authority to quantify what consumers seek from AI and how much decision-making authority they assign to it. Applied to 1.5 million real-world ChatGPT and Gemini interactions from 6,304 users in the United States and India, we find that financial services are already a substantial AI use case. Consumers overwhelmingly use AI to retrieve information and shape financial judgement, while delegation of financial execution remains rare. By shifting attention from conversation topics to delegated decision authority, this work establishes a behavioural baseline for measuring the transition to increasingly agentic AI.

Summary

Main Finding

Conversational AI is already a major channel for personal finance interactions, but users overwhelmingly use it for information and advice (inform & shape decisions). True delegation of financial execution to AI (agentic, transactional actions) is rare in real-world ChatGPT and Gemini use — largely limited to budgeting/monitoring tasks — establishing a behavioural baseline before the wider adoption of agentic financial assistants.

Key Points

  • Scale and scope
    • Dataset: 1.53M user prompts from 6,304 opt-in users (2,499 US; 3,805 India), covering ChatGPT and Gemini (Aug–Oct 2025).
    • After topical segmentation: ~292k US subchats and ~213k India subchats.
    • About half of users engaged in finance-related conversations during the study window.
  • Behavioural lens: intent × decision authority
    • Introduced a two-dimensional framework: (a) behavioural intent (12 financial intents) and (b) decision authority (DA) levels:
      • Level 1 Inform: provide information/explanations.
      • Level 2 Shape: advisory/analytical influence on choices.
      • Level 3 Act: autonomous/transactional execution or automation.
    • Most real conversations fall into Inform (Level 1) and Shape (Level 2). Level 3 (Act) conversations are uncommon (<~1% in raw snapshot; higher-agentic intents under 1%).
  • Empirical patterns
    • Financial services are among the largest topical domains for conversational AI use.
    • Common intents: simple retrieval, complex research, product/strategy comparison and optimisation, personal financial analysis and planning — i.e., information and advice roles dominate.
    • Delegation/execution (transferring decision-making authority to the model) is rare, primarily observed in budgeting, monitoring, and a few instruction-led execution examples.
  • Taxonomy & tagging
    • Developed a MECE financial services taxonomy (Investments, Retail Banking & Credit, Tax, Payments, Benefits, Insurance, Business Finance, Finance Infrastructure, etc.).
    • Entity extraction yielded 2.7k correctly matched financial-service keywords (plus many unmatched general terms).

Data & Methods

  • Data source and preprocessing
    • Opt-in sample of ChatGPT and Gemini chat histories from US and India; metadata includes user messages, AI responses, timestamps.
    • Non-English India messages (≈144k) translated to English (Google Translate) to standardise modelling.
    • Chat segmentation heuristics split long sessions into topical subchats (thresholds and token-overlap rules), producing ~505k subchats total.
  • Finance detection
    • Task: binary classifier to label subchats as finance vs non-finance.
    • Training set: 5.8k annotated subchats (sampled across markets and platforms) seeded with a 2.4k financial keyword dictionary.
    • Model: LongFormer-based fine-tuned transformer (chosen for long context; avg subchat ≈2k tokens).
    • Performance: Accuracy 96.5%, F1 97.3% on test split.
  • Financial-services tagging
    • Two-stage approach: LLM-assisted entity extraction (GPT‑4o‑mini) → manual mapping to FS taxonomy.
    • Result: 2.7k matched FS keywords; deterministic tagging of finance subchats that mention FS keywords.
  • Intent and decision-authority classification
    • Intent taxonomy: 12 labels (e.g., Delegated Financial Decision Execution; Instruction-led Execution; Automation & Monitoring; Product/Strategy Optimisation; Comparison; Problem Resolution; Personal Analysis; Planning; Complex Research; Simple Retrieval; Learning; Creation).
    • Mapping of intents to DA levels (Inform / Shape / Act).
    • Training data: 3.4k manually labelled finance subchats (2.8k organic + 600 synthetic examples from GPT‑5.5 to enrich higher-agentic cases).
    • Model: BigBird (sparse-attention Transformer for long contexts), fine-tuned for multi-label classification.
    • Note: synthetic augmentation used to partially offset the rarity of high-DA examples in the snapshot.
  • Limitations noted by authors
    • Sample is not a general-population random sample: skewed demographically (US sample younger, more female, lower income; India sample heavily male and young).
    • Translation errors may affect Indian non-English messages.
    • Low prevalence of agentic intents in the observed period limits inference about future delegation adoption.
    • Use of synthetic examples to train rare intents may bias models toward detecting agentic intents that are rare in real data.

Implications for AI Economics

  • Measurement baseline for delegation
    • The paper provides an operationalizable behavioural metric (decision authority levels) to measure the gradual transfer of cognitive and decision-making responsibilities from humans to AI — a crucial variable for empirical work on task automation and substitution in services.
  • Labor and demand-side effects
    • Current dominant use-cases (inform + shape) suggest AI complements human judgement (advisory role) rather than substituting execution-heavy financial roles. This points to demand shifts toward higher-order advisory and monitoring tasks rather than immediate large-scale job displacement in transactional financial roles.
  • Product design and market structure
    • Low delegation today implies firms can focus on hybrid human-in-the-loop product designs, verification layers, and interfaces that preserve user control while improving decision support. As agentic capabilities increase, product-market fit and value-capture depend on trusted, auditable delegation mechanisms.
  • Regulation, liability, and consumer protection
    • If delegation increases, regulatory attention must shift from content/accuracy (advice quality) to authorization, execution safety, liability allocation, and disclosure of delegated authority. The DA taxonomy can help regulators define thresholds for when stricter oversight is needed.
  • Financial markets and consumer welfare
    • Widespread use of AI for research and product optimisation can change consumer search costs, price competition among providers, and the market for paid financial advice. Monitoring whether information/shape uses improve outcomes or introduce correlated errors (e.g., herd behaviour) is essential.
  • Research agenda for AI economists
    • Track temporal dynamics of DA in conversations to measure adoption of agentic features.
    • Quantify effect of AI advice (shape) on actual financial behaviour and outcomes (e.g., savings, investment choices, fee minimisation).
    • Model complementarities between algorithmic advice and financial intermediaries; study pricing and liability when AI executes transactions.
    • Use the framework to study distributional effects (who delegates? demographic/wealth gradients in DA assignment).
  • Policy takeaways
    • Policymakers should monitor not just topical AI use but the delegated authority dimension to anticipate when governance, licensing, and consumer-protection frameworks need to be tightened.

If you want, I can extract the paper’s quantitative breakdowns of intents and DA-level shares (per country / per platform), summarize the taxonomy mapping, or produce example conversation archetypes illustrating Inform vs Shape vs Act.

Assessment

Paper Typedescriptive Evidence Strengthmedium — Large-scale real-world dataset (1.5M prompts from 6,304 users) provides strong descriptive evidence on usage patterns, but findings are observational and susceptible to sampling bias (opt-in panel), measurement error from automated classification, heavy synthetic augmentation for rare/high-agentic intents, and platform/time specificity which limit causal inference and external validity. Methods Rigormedium — The authors use reasonable preprocessing (translation, topic segmentation), modern long-context transformer classifiers (LongFormer/BigBird), and report strong performance for the finance detector; however, key concerns include limited annotated training data relative to scale, extensive use of synthetic examples (very high for delegated-execution intents), only partial reporting of classifier evaluation metrics (intent classifier metrics are not shown in the supplied text), and heuristic segmentation rules that were only lightly validated. SampleConversation histories from 6,304 opt-in users (2,499 US; 3,805 India) collected Aug–Oct 2025 via a MeasureProtocol panel, containing 1.53M user prompts across ChatGPT and Gemini (1.31M ChatGPT prompts, 214.6k Gemini prompts); non-English India messages (≈144k) were machine-translated; conversations were segmented into ≈505k subchats and labeled via a mix of manual annotation and fine-tuned transformer models. Themeshuman_ai_collab adoption GeneralizabilityOpt-in panel; not a representative sample of general population or all conversational-AI users, Demographics skew (US sample younger/female/lower income; India sample heavily male and very young), Only two AI platforms (ChatGPT, Gemini) and two countries (US, India) during a specific three-month window (Aug–Oct 2025), Machine translation of non-English content may introduce errors and cultural nuance loss, Automated classification and extensive synthetic augmentation (especially for rare high-delegation intents) may bias intent prevalence estimates, Segmentation heuristics may imperfectly split/merge topical units, affecting counts

Claims (8)

ClaimDirectionOutcomeConfidence & EvidenceDetails
The study analyzes 1.53 million user prompts from 6,304 users in the United States and India, covering ChatGPT and Gemini histories collected during August–October 2025. Other positive Scale of the human–AI financial decision-making corpus
Reading fidelity high
Study strength medium
n=6304
1.53M user prompts
0.18
Approximately half of the users in the study engaged in finance-related conversations with conversational AI during the study period. Adoption Rate positive Share of users engaging with AI about financial services or money-related issues
Reading fidelity high
Study strength medium
n=6304
approximately half of users
0.18
Consumers predominantly use conversational AI for financial information retrieval and for shaping financial judgment, rather than for executing financial transactions. Decision Quality mixed Distribution of financial AI interactions across information, advisory, and execution roles
Reading fidelity high
Study strength medium
not reported
0.18
Delegation of financial execution to AI is rare and is largely confined to budgeting and financial tracking. Task Allocation negative Prevalence and domain concentration of delegated financial execution
Reading fidelity high
Study strength medium
not reported
0.18
The dataset contained fewer than 1% of conversations with intents involving delegated financial decision execution, instruction-led financial execution, or financial automation and monitoring. Automation Exposure negative Prevalence of highly agentic financial intents
Reading fidelity high
Study strength medium
under 1%
0.18
The finance-conversation classifier achieved 96.5 accuracy and an F1 score of 97.3 on its held-out test set. Other positive Finance-subchat classification performance
Reading fidelity high
Study strength high
n=5800
Accuracy = 96.5; F1 = 97.3
0.3
The US and India samples differ substantially from their respective general-population benchmarks in age, gender, income, and employment characteristics. Other mixed Sample representativeness and demographic composition
Reading fidelity high
Study strength medium
n=6304
US: 58.2% aged 18–34 versus 29.0% GenPop; India: 89.4% aged 18–34 versus 41.2% GenPop
0.18
The financial-intent training dataset contained 3,400 manually annotated subchats, including 2,800 real conversations and 600 synthetic examples. Other positive Size and composition of the financial-intent classification training dataset
Reading fidelity high
Study strength medium
n=3400
2.8k real conversations plus 600 synthetic examples
0.18

Notes