7 cumulative citations
View corpus contextPeople carry expectations about a multipurpose AI across tasks and update them cautiously: belief adjustments run at roughly half the Bayesian rate and a 10-point higher posterior in one task raises priors in the next by 3–4 points. Delegation hinges more on users' subjective accuracy beliefs than on self-confidence, implying persistent, path-dependent reliance patterns.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Large language models (LLMs) increasingly support heterogeneous tasks within a single interface, requiring users to form, update, and act upon beliefs about one system across domains with different reliability profiles. Understanding how such beliefs transfer across tasks and shape delegation is therefore critical for the design of multipurpose AI systems. We report a preregistered experiment (N=240; 7,200 trials) in which participants interacted with a controlled AI simulation across grammar checking, travel planning, and visual question answering, each with fixed, domain-typical accuracy levels. Delegation was operationalized as a binary reliance decision: accepting the AI's output versus acting independently, and belief dynamics were evaluated against Bayesian benchmarks. We find three main results. First, participants do not reset beliefs between tasks: priors in a new task depend on posteriors from the previous task, with a 10-point increase predicting a 3-4 point higher subsequent prior. Second, within tasks, belief updating follows the Bayesian direction but is substantially conservative, proceeding at roughly half the normative Bayesian rate. Third, delegation is driven primarily by subjective beliefs about AI accuracy rather than self-confidence, though confidence independently reduces reliance when beliefs are held constant. Together, these findings show that users form global, path-dependent expectations about multipurpose AI systems, update them conservatively, and rely on AI primarily based on subjective beliefs rather than objective performance. We discuss implications for expectation calibration, reliance design, and the risks of belief spillovers in deployed LLM-based interfaces.
Summary
Main Finding
Users interacting with a single multipurpose AI do not treat tasks independently. Beliefs about the AI’s accuracy carry over across tasks (priors are anchored to previous posteriors), within-task belief updating is directionally Bayesian but substantially conservative (about half the normative Bayesian updating rate), and delegation (accepting AI output vs. acting independently) is driven mainly by subjective beliefs about AI accuracy rather than by users’ self-confidence. Dispositional trust and AI literacy raise initial priors.
Key Points
- Experiment: preregistered, N = 240 participants, 7,200 trials (30 trials per participant). Tasks: grammar checking, travel planning, visual question answering. AI outputs were pre-scripted to enforce fixed, domain-typical accuracy levels.
- Cross-task spillover: participants do not reset priors when starting a new task. A 10-point increase in posterior belief from the previous task predicts a ~3–4 point higher prior in the next task.
- Conservative updating: within tasks, belief updates move in the Bayesian-predicted direction but at roughly ~50% of the Bayesian (Beta–Binomial) benchmark — i.e., systematic under-reaction (conservatism bias).
- Delegation operationalized as a binary accept/reject decision. Delegation is primarily explained by subjective beliefs about the AI’s accuracy; self-confidence has an independent (negative) effect on reliance only when beliefs are held constant.
- Individual differences: dispositional trust in automation predicts higher initial priors; AI literacy also independently increases initial priors. Need for cognition and related traits were measured as moderators.
- The study intentionally isolates a one-shot-style delegation decision (accept AI output vs act independently) to provide a clean mechanism-level account; richer mixed-initiative interactions were not modeled here.
Data & Methods
- Design: Sequential, multi-task behavioral experiment with controlled AI simulation (pre-scripted outputs) so ground-truth task accuracies were known and stable by task.
- Participants: 240 (preregistered). Total trials = 7,200 (30 trials per participant across three tasks).
- Tasks: grammar correction, travel planning, visual question answering — chosen for cross-domain heterogeneity and domain-typical accuracy differences.
- Measures:
- Trial-level subjective belief ratings about AI accuracy (priors and posteriors).
- Binary delegation choice: accept AI output vs choose own answer.
- Self-confidence (task-level), dispositional trust in automation (TiA), AI literacy (MAILS), need for cognition (NCS-6).
- Benchmarks & analysis:
- Normative benchmark: Bayesian updating modeled with a Beta–Binomial framework (pseudocounts).
- Main tests: (a) regression of next-task priors on lagged posteriors to quantify carryover; (b) comparison of observed within-task update magnitudes to Bayesian predictions to estimate conservatism; (c) logistic regressions predicting delegation from lagged beliefs and confidence; (d) tests linking dispositional measures to initial priors.
- Key quantitative results reported: 10-point posterior rise → ~3–4 point higher next-task prior; within-task updating roughly half the Bayesian rate.
Implications for AI Economics
- Reputation and spillovers across product lines:
- Multipurpose AI providers face strong cross-task reputation externalities. Good performance in one domain will raise user expectations in other domains (increasing initial adoption), while failures in any domain can depress demand broadly. Firms should anticipate and manage these cross-domain reputation effects when deploying bundled or multi-capability products.
- Pricing, bundling, and product strategy:
- Because users generalize beliefs across tasks, bundling multiple capabilities into one interface can create value beyond direct functionality (positive spillovers) but also increases downside risk (one domain’s failure harms others). This affects optimal bundling and pricing strategies: firms may charge a premium for bundled convenience but must invest more heavily in per-task guarantees or monitoring to avoid systemic reputation losses.
- Platform design and signaling:
- Per-task reliability signals, explicit per-task performance metrics, and segmented reputational systems (separate accuracy indicators or certifications by task) can reduce harmful belief spillovers and miscalibration. Platforms that fail to provide per-task signalization may induce over- or under-utilization of capabilities, reducing social welfare.
- Adoption dynamics and market competition:
- Conservative updating implies slower convergence of user beliefs to true per-task performance. Entrants with superior performance may face slow demand growth (users underweight new evidence), favoring incumbents with existing positive priors. Conversely, incumbents suffering observable errors could lose cross-domain trust rapidly. These dynamics influence diffusion models, investment in quality improvement, and marketing strategies.
- Labor market and productivity externalities:
- Since delegation is driven mainly by subjective beliefs rather than immediate objective performance, productivity gains from AI tools depend on belief calibration. Over-reliance in low-accuracy tasks or under-reliance in high-accuracy tasks can misallocate human effort, affecting labor substitution decisions and firm-level productivity estimates.
- Liability, regulation, and consumer protection:
- Regulators should consider requiring standardized, per-task accuracy disclosures or audits for multipurpose AI systems, because global reputational signals mislead users about per-task reliability. Policies that mandate transparent, task-specific performance reporting (or default resets of reputational cues between tasks) could reduce market failures due to misinformed delegation.
- Contracting and SLAs:
- For B2B deployments, contracts and SLAs should specify per-task performance metrics and remedies. Given belief inertia, providers might be able to leverage good performance in one task to expand offerings, but buyers should insist on task-specific guarantees to avoid hidden externalities.
- Design of interventions and value of information:
- Investments in user interfaces that provide clear task-level feedback, calibrated confidence estimates, or hands-on demonstrations can have high economic value by reducing miscalibration. The social value of such interventions includes better allocation of human attention and reduced error externalities.
- Research and evaluation of multipurpose AI markets:
- Empirical economic models of AI adoption should incorporate belief spillovers and conservative updating. Ignoring cross-task belief dynamics will misestimate adoption speeds, welfare effects, and the competitive landscape.
Limitations and next steps (brief): - The experiment used pre-scripted simulated AI outputs and a binary delegation decision; richer mixed-initiative interactions with real LLMs may show different dynamics. - External validity: participant pool and lab setting may not capture workplace stakes that change updating and delegation. - Future work: test interventions (per-task signals, reset mechanisms), longer-horizon learning, heterogeneous user populations, and incorporate findings into formal economic models of adoption, competition, and platform design.
Assessment
Claims (7)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| We ran a preregistered experiment (N=240; 7,200 trials) in which participants interacted with a controlled AI simulation across grammar checking, travel planning, and visual question answering, each with fixed, domain-typical accuracy levels. Other | positive | experimental implementation and sampling (study design description) |
Reading fidelity
high
Study strength
high
|
n=240
|
| Delegation was operationalized as a binary reliance decision: accepting the AI's output versus acting independently. Other | positive | delegation operationalization (binary reliance decision) |
Reading fidelity
high
Study strength
high
|
n=240
|
| Participants do not reset beliefs between tasks: priors in a new task depend on posteriors from the previous task, with a 10-point increase predicting a 3–4 point higher subsequent prior. Task Allocation | positive | prior belief about AI accuracy in a new task (points on reported scale) |
Reading fidelity
high
Study strength
high
|
n=240
10-point increase predicts a 3-4 point higher subsequent prior
|
| Within tasks, belief updating follows the Bayesian direction but is substantially conservative, proceeding at roughly half the normative Bayesian rate. Decision Quality | mixed | magnitude and direction of belief updating relative to Bayesian normative benchmark |
Reading fidelity
high
Study strength
high
|
n=240
roughly half the normative Bayesian rate
|
| Delegation is driven primarily by subjective beliefs about AI accuracy rather than self-confidence. Task Allocation | positive | probability of delegating to the AI (binary reliance decision) |
Reading fidelity
high
Study strength
high
|
n=240
|
| Confidence independently reduces reliance when beliefs are held constant. Task Allocation | negative | probability of delegating to the AI (binary reliance decision) conditional on subjective beliefs |
Reading fidelity
high
Study strength
high
|
n=240
|
| Together, these findings show that users form global, path-dependent expectations about multipurpose AI systems, update them conservatively, and rely on AI primarily based on subjective beliefs rather than objective performance. Task Allocation | mixed | overall behavioral pattern: global/path-dependent expectations, conservative updating, reliance determinants |
Reading fidelity
high
Study strength
medium
|
n=240
|