The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

People carry expectations about a multipurpose AI across tasks and update them cautiously: belief adjustments run at roughly half the Bayesian rate and a 10-point higher posterior in one task raises priors in the next by 3–4 points. Delegation hinges more on users' subjective accuracy beliefs than on self-confidence, implying persistent, path-dependent reliance patterns.

Belief Updating and Delegation in Multi-Task Human-AI Interaction: Evidence from Controlled Simulations
Shreyan Biswas, Alexander Erlei, Ujwal Gadiraju · February 02, 2026
arxiv quasi_experimental medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Shreyan Biswas unresolved corpus identity
  2. Alexander Erlei unresolved corpus identity
  3. Ujwal Gadiraju unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Shreyan Biswas provider ID
  2. Alexander Erlei provider ID
  3. U. Gadiraju provider ID
In a preregistered experiment, users carry posteriors across tasks when interacting with a multipurpose AI, update beliefs conservatively at about half the Bayesian rate, and decide to delegate primarily based on subjective beliefs about AI accuracy rather than their own confidence.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Large language models (LLMs) increasingly support heterogeneous tasks within a single interface, requiring users to form, update, and act upon beliefs about one system across domains with different reliability profiles. Understanding how such beliefs transfer across tasks and shape delegation is therefore critical for the design of multipurpose AI systems. We report a preregistered experiment (N=240; 7,200 trials) in which participants interacted with a controlled AI simulation across grammar checking, travel planning, and visual question answering, each with fixed, domain-typical accuracy levels. Delegation was operationalized as a binary reliance decision: accepting the AI's output versus acting independently, and belief dynamics were evaluated against Bayesian benchmarks. We find three main results. First, participants do not reset beliefs between tasks: priors in a new task depend on posteriors from the previous task, with a 10-point increase predicting a 3-4 point higher subsequent prior. Second, within tasks, belief updating follows the Bayesian direction but is substantially conservative, proceeding at roughly half the normative Bayesian rate. Third, delegation is driven primarily by subjective beliefs about AI accuracy rather than self-confidence, though confidence independently reduces reliance when beliefs are held constant. Together, these findings show that users form global, path-dependent expectations about multipurpose AI systems, update them conservatively, and rely on AI primarily based on subjective beliefs rather than objective performance. We discuss implications for expectation calibration, reliance design, and the risks of belief spillovers in deployed LLM-based interfaces.

Summary

Main Finding

Users interacting with a single multipurpose AI do not treat tasks independently. Beliefs about the AI’s accuracy carry over across tasks (priors are anchored to previous posteriors), within-task belief updating is directionally Bayesian but substantially conservative (about half the normative Bayesian updating rate), and delegation (accepting AI output vs. acting independently) is driven mainly by subjective beliefs about AI accuracy rather than by users’ self-confidence. Dispositional trust and AI literacy raise initial priors.

Key Points

  • Experiment: preregistered, N = 240 participants, 7,200 trials (30 trials per participant). Tasks: grammar checking, travel planning, visual question answering. AI outputs were pre-scripted to enforce fixed, domain-typical accuracy levels.
  • Cross-task spillover: participants do not reset priors when starting a new task. A 10-point increase in posterior belief from the previous task predicts a ~3–4 point higher prior in the next task.
  • Conservative updating: within tasks, belief updates move in the Bayesian-predicted direction but at roughly ~50% of the Bayesian (Beta–Binomial) benchmark — i.e., systematic under-reaction (conservatism bias).
  • Delegation operationalized as a binary accept/reject decision. Delegation is primarily explained by subjective beliefs about the AI’s accuracy; self-confidence has an independent (negative) effect on reliance only when beliefs are held constant.
  • Individual differences: dispositional trust in automation predicts higher initial priors; AI literacy also independently increases initial priors. Need for cognition and related traits were measured as moderators.
  • The study intentionally isolates a one-shot-style delegation decision (accept AI output vs act independently) to provide a clean mechanism-level account; richer mixed-initiative interactions were not modeled here.

Data & Methods

  • Design: Sequential, multi-task behavioral experiment with controlled AI simulation (pre-scripted outputs) so ground-truth task accuracies were known and stable by task.
  • Participants: 240 (preregistered). Total trials = 7,200 (30 trials per participant across three tasks).
  • Tasks: grammar correction, travel planning, visual question answering — chosen for cross-domain heterogeneity and domain-typical accuracy differences.
  • Measures:
    • Trial-level subjective belief ratings about AI accuracy (priors and posteriors).
    • Binary delegation choice: accept AI output vs choose own answer.
    • Self-confidence (task-level), dispositional trust in automation (TiA), AI literacy (MAILS), need for cognition (NCS-6).
  • Benchmarks & analysis:
    • Normative benchmark: Bayesian updating modeled with a Beta–Binomial framework (pseudocounts).
    • Main tests: (a) regression of next-task priors on lagged posteriors to quantify carryover; (b) comparison of observed within-task update magnitudes to Bayesian predictions to estimate conservatism; (c) logistic regressions predicting delegation from lagged beliefs and confidence; (d) tests linking dispositional measures to initial priors.
  • Key quantitative results reported: 10-point posterior rise → ~3–4 point higher next-task prior; within-task updating roughly half the Bayesian rate.

Implications for AI Economics

  • Reputation and spillovers across product lines:
    • Multipurpose AI providers face strong cross-task reputation externalities. Good performance in one domain will raise user expectations in other domains (increasing initial adoption), while failures in any domain can depress demand broadly. Firms should anticipate and manage these cross-domain reputation effects when deploying bundled or multi-capability products.
  • Pricing, bundling, and product strategy:
    • Because users generalize beliefs across tasks, bundling multiple capabilities into one interface can create value beyond direct functionality (positive spillovers) but also increases downside risk (one domain’s failure harms others). This affects optimal bundling and pricing strategies: firms may charge a premium for bundled convenience but must invest more heavily in per-task guarantees or monitoring to avoid systemic reputation losses.
  • Platform design and signaling:
    • Per-task reliability signals, explicit per-task performance metrics, and segmented reputational systems (separate accuracy indicators or certifications by task) can reduce harmful belief spillovers and miscalibration. Platforms that fail to provide per-task signalization may induce over- or under-utilization of capabilities, reducing social welfare.
  • Adoption dynamics and market competition:
    • Conservative updating implies slower convergence of user beliefs to true per-task performance. Entrants with superior performance may face slow demand growth (users underweight new evidence), favoring incumbents with existing positive priors. Conversely, incumbents suffering observable errors could lose cross-domain trust rapidly. These dynamics influence diffusion models, investment in quality improvement, and marketing strategies.
  • Labor market and productivity externalities:
    • Since delegation is driven mainly by subjective beliefs rather than immediate objective performance, productivity gains from AI tools depend on belief calibration. Over-reliance in low-accuracy tasks or under-reliance in high-accuracy tasks can misallocate human effort, affecting labor substitution decisions and firm-level productivity estimates.
  • Liability, regulation, and consumer protection:
    • Regulators should consider requiring standardized, per-task accuracy disclosures or audits for multipurpose AI systems, because global reputational signals mislead users about per-task reliability. Policies that mandate transparent, task-specific performance reporting (or default resets of reputational cues between tasks) could reduce market failures due to misinformed delegation.
  • Contracting and SLAs:
    • For B2B deployments, contracts and SLAs should specify per-task performance metrics and remedies. Given belief inertia, providers might be able to leverage good performance in one task to expand offerings, but buyers should insist on task-specific guarantees to avoid hidden externalities.
  • Design of interventions and value of information:
    • Investments in user interfaces that provide clear task-level feedback, calibrated confidence estimates, or hands-on demonstrations can have high economic value by reducing miscalibration. The social value of such interventions includes better allocation of human attention and reduced error externalities.
  • Research and evaluation of multipurpose AI markets:
    • Empirical economic models of AI adoption should incorporate belief spillovers and conservative updating. Ignoring cross-task belief dynamics will misestimate adoption speeds, welfare effects, and the competitive landscape.

Limitations and next steps (brief): - The experiment used pre-scripted simulated AI outputs and a binary delegation decision; richer mixed-initiative interactions with real LLMs may show different dynamics. - External validity: participant pool and lab setting may not capture workplace stakes that change updating and delegation. - Future work: test interventions (per-task signals, reset mechanisms), longer-horizon learning, heterogeneous user populations, and incorporate findings into formal economic models of adoption, competition, and platform design.

Assessment

Paper Typequasi_experimental Evidence Strengthmedium — The study is preregistered, uses a substantial sample (N=240) and many repeated trials (7,200), and exercises tight experimental control over AI behavior, which supports credible internal validity for the measured behavioral effects; however, external validity is limited by the artificial simulation (not deployed, real-world LLMs), a small set of tasks, likely low-stakes laboratory/online participants, and binary reliance measures that may not map directly to real-world economic outcomes. Methods Rigorhigh — Methods are rigorous: preregistration, large number of trials, within-subject design, explicit manipulation of domain accuracies, and analysis against Bayesian benchmarks; potential limitations include reliance on simulated AI rather than live LLMs, limited task set, reliance on self-reported beliefs and binary choices, and possible order effects if not fully randomized. Sample240 human participants (online experimental sample) completing 7,200 trials total in a within-subject design across three tasks (grammar checking, travel planning, visual question answering); participants reported subjective beliefs about AI accuracy and made binary delegation decisions (accept AI output vs act independently); tasks had fixed, domain-typical accuracy rates imposed by the simulation. Themeshuman_ai_collab adoption IdentificationPreregistered within-subject behavioral experiment using a controlled AI simulation with fixed, domain-typical accuracies across three tasks; causal inference is supported by experimental control of AI output quality and repeated measurement of beliefs and binary delegation choices, and by comparing observed belief updates to normative Bayesian benchmarks. GeneralizabilityUses a simulated AI with fixed accuracies rather than deployed LLMs whose performance and error modes differ, Only three tasks studied (grammar, travel planning, visual Q&A) — findings may not hold for other domains or higher-stakes decisions, Likely short-term, low-stakes laboratory/online interactions that may not reflect long-run organizational adoption or real-world behavior under incentives, Participant sample (online) may not represent professional users, firms, or workers making consequential delegation decisions, Binary reliance measure (accept vs act independently) may oversimplify nuanced real-world delegation and mixed-initiative collaboration

Claims (7)

ClaimDirectionOutcomeConfidence & EvidenceDetails
We ran a preregistered experiment (N=240; 7,200 trials) in which participants interacted with a controlled AI simulation across grammar checking, travel planning, and visual question answering, each with fixed, domain-typical accuracy levels. Other positive experimental implementation and sampling (study design description)
Reading fidelity high
Study strength high
n=240
0.8
Delegation was operationalized as a binary reliance decision: accepting the AI's output versus acting independently. Other positive delegation operationalization (binary reliance decision)
Reading fidelity high
Study strength high
n=240
0.8
Participants do not reset beliefs between tasks: priors in a new task depend on posteriors from the previous task, with a 10-point increase predicting a 3–4 point higher subsequent prior. Task Allocation positive prior belief about AI accuracy in a new task (points on reported scale)
Reading fidelity high
Study strength high
n=240
10-point increase predicts a 3-4 point higher subsequent prior
0.8
Within tasks, belief updating follows the Bayesian direction but is substantially conservative, proceeding at roughly half the normative Bayesian rate. Decision Quality mixed magnitude and direction of belief updating relative to Bayesian normative benchmark
Reading fidelity high
Study strength high
n=240
roughly half the normative Bayesian rate
0.8
Delegation is driven primarily by subjective beliefs about AI accuracy rather than self-confidence. Task Allocation positive probability of delegating to the AI (binary reliance decision)
Reading fidelity high
Study strength high
n=240
0.8
Confidence independently reduces reliance when beliefs are held constant. Task Allocation negative probability of delegating to the AI (binary reliance decision) conditional on subjective beliefs
Reading fidelity high
Study strength high
n=240
0.8
Together, these findings show that users form global, path-dependent expectations about multipurpose AI systems, update them conservatively, and rely on AI primarily based on subjective beliefs rather than objective performance. Task Allocation mixed overall behavioral pattern: global/path-dependent expectations, conservative updating, reliance determinants
Reading fidelity high
Study strength medium
n=240
0.48

Notes