The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Visible unverified use can tip communities from cautious checking to unchecked dependence on AI, producing collective overreliance even when individuals could learn to be calibrated; redesigning feedback—showing verification or reducing social proof—can avert the cascade.

Modeling AI Overreliance as a Complex Adaptive System
Ahana Biswas · August 20, 2026
arxiv theoretical low evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Ahana Biswas unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Ahana Biswas provider ID
A formal agent-based model and analytic results show that while task difficulty and AI quality set baseline reliance, visible unverified peer use (social proof) can create a verification-suppression feedback that tips populations into collective overreliance, and interface or feedback-design changes (making verification visible or dampening social proof) can prevent or reverse such collapse.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Whether AI assistance helps or harms a population depends less on the model's accuracy than on whether people rely on it appropriately trusting it when it is right and checking it when it is not. Yet reliance is usually studied one user at a time. We model it as a population process: agents repeatedly solve a task alone, accept an AI answer, or verify it, updating a Bayesian belief about AI quality and, when networked, learning from peers. Four results form one story. The environment sets the baseline: task difficulty and AI quality fix both overreliance and calibration regret. Social learning creates consensus, not overreliance: a mean-preservation theorem, confirmed by a 2*2 topology*tagging design, shows connectivity moves the aggregate only when influence transmits beliefs. Social proof turns reliance into a feedback cascade: visible unverified use suppresses verification and tips the population into collective overreliance. Feedback design can prevent collapse: making verification visible or dampening social proof reverses it. Together, the results frame AI reliance as a computational social dynamics problem, where individual learning, peer observation, and feedback exposure jointly shape whether a population remains calibrated.

Summary

Main Finding

Collective overreliance on AI can arise not from model accuracy alone but from social-feedback dynamics: when peers’ visible unverified use raises the immediate utility of accepting AI (a “verification‑suppression” feedback), populations can tip from routine verification into pervasive, unverified AI use. Network connectivity by itself creates consensus (reduces trust dispersion) but does not change the aggregate mean trust unless influential nodes transmit correlated beliefs; feedback visibility and interface design determine whether the system cascades toward overreliance or remains calibrated.

Key Points

  • Conceptual framing

    • Appropriate reliance is a population-level, dynamic problem: users learn about AI quality over repeated interactions and from peers, so reliance is a coupled social-adaptive system.
    • Two central failures: overreliance (accepting wrong AI outputs) and underreliance (discarding useful AI after seeing errors). The paper stresses calibration (trust tracking true reliability) and uses regret as the primary normative metric.
  • Model contributions

    • Micro-founded agent-based model (ABM) where agents choose each period among: solve alone (S), accept AI unverified (A), or accept AI but verify (V).
    • Agents hold Dirichlet beliefs over outcome categories (Correct / Partial / Wrong); verifying gives higher-precision signals, unverified use yields optimism-upgraded (lower-precision) signals.
    • Choice is multinomial logit with utilities incorporating belief, verification cost (friction), and observed peer behavior (social proof).
    • Networked version: ER (random) and BA (hub-heavy) topologies, two peer-learning channels—experience-based (neighbors’ outcomes) and opinion dynamics (neighbors’ beliefs)—and a feedback channel that raises UA by s·π (π = fraction of neighbors who used AI unverified).
  • Theoretical results

  • Oracle proposition: a perfectly calibrated agent (correct belief) chooses the oracle action and attains zero expected regret; behavioral proxies (e.g., observed unverified use) can differ from regret because of stochastic outcomes.
  • Mean-preservation theorem: under DeGroot-style averaging with exchangeable peer signals and weights independent of signal values, population mean trust is invariant and cross-sectional variance falls—connectivity creates consensus, not mean change.
  • Cascade tipping (mean-field): the unverified-use fraction solves a fixed-point equation; saddle-node (fold) bifurcation yields hysteresis and bistability when coupling θs exceeds a threshold (θs > 4) and parameters lie in the fold interval. Finite population heterogeneity and stochasticity can smooth or shift this.

  • Computational findings (four empirical results)

  • Environment sets baseline: task difficulty and AI quality are the largest drivers of overreliance and regret; harder tasks and lower AI quality raise baseline reliance errors.
  • Social learning compresses trust (consensus) but does not change aggregate reliance unless influence transmits correlated beliefs (opinion dynamics) or neighbor signals are non-exchangeable; high-trust hubs matter only when they transmit beliefs directly.
  • Social proof creates feedback cascades: making neighbors’ unverified use visible (increasing s) can collapse verification rates (e.g., verification 0.29 → 0.002 as s increases) and raise overreliance markedly—this is endogenous and can occur even without changes in AI quality or individual rationality.
  • Feedback-aware interventions prevent collapse: making verification visible to peers (so verified use is also visible), dampening social proof (reducing effective s), or lowering verification friction can reverse or prevent cascades and produce beneficial counter-cascades.

  • Metrics and evaluation

    • Primary normative metric: regret (utility lost vs. oracle).
    • Behavioral proxies: severity-weighted unverified use (behavioral overreliance), underreliance (solving alone when AI would be better).
    • Additional: trust τ (expected AI score under belief), RAIR/RSR (relative AI-/self-reliance on disagreement cases).

Data & Methods

  • Model type: Agent-based model with Dirichlet belief updating and multinomial logit choice.
  • Population & runs: default N = 400 agents, T = 100 time steps, repeated seeds (typically 20), final-period averages reported.
  • Key parameters:
    • Choice decisiveness θ = 10 (explored up to 30 in sensitivity checks).
    • Verification friction λV default 0.20 (interventions reduced to 0.05).
    • Social proof coupling s swept in [0, 0.6]; damping ρ ∈ [0,1].
    • Difficulty sampled per task: “easy” Beta(2,6), “hard” Beta(6,2).
    • AI quality q interpolated between discrete Low/High parameter sets to create a continuous quality axis [0,1].
    • Agent heterogeneity: abilities, verification skills ~ Beta(2,2).
  • Network structures:
    • ER (Erdős–Rényi) random networks and BA (Barabási–Albert) hub-heavy networks.
    • Source tagging option prevents double-counting of hubs; experience-based vs opinion-dynamics peer-learning channels distinguished.
  • Theoretical analysis:
    • DeGroot averaging and mean-preservation proof under exchangeability.
    • Mean-field reduction (Brock–Durlauf style) to derive fold bifurcation conditions.
  • Validation & sensitivity:
    • Pattern-oriented validation: reproduces algorithm aversion/appreciation patterns.
    • Global sensitivity via Morris elementary effects to identify dominant parameters (difficulty and AI quality foremost).
  • Counterfactuals:
    • For evaluation only, the model samples counterfactual outcomes for unchosen actions (AI outcome and solo outcome) to compute behavioral proxies and appropriate reliance metrics; agents do not observe unchosen outcomes.

Implications for AI Economics

  • Externalities and network effects

    • AI reliance has social externalities: one user’s visible unverified acceptance raises others’ propensity to accept unverified outputs, generating positive feedback that can produce aggregate harm (higher regret).
    • Platform-level visibility design therefore internalizes—or amplifies—these externalities. Platforms that surface unverified use create a coordination game with potential negative equilibria (verification collapse).
  • Market adoption & welfare

    • Adoption decisions and aggregate welfare depend not only on per-user accuracy but on how social feedback shapes verification behavior; columns of seemingly small interface choices can shift populations from welfare-enhancing calibrated use to welfare-degrading collective overreliance.
    • Regret, not raw acceptance rates, should be used to measure welfare implications; high AI quality + high deference can raise regret if self-reliance is systematically better in some tasks.
  • Policy and regulatory design

    • Disclosure/visibility rules matter: regulators can mandate or encourage visibility of verification activity (e.g., badges, audit trails showing that outputs were checked) to counteract social-proof cascades.
    • Interventions that target the feedback channel (change what is visible or damp social-proof signals) are more powerful than those that only change individual costs (e.g., small subsidies for verification), because cascades are driven by endogenous utility shifts.
    • Antitrust/competition implications: hub influencers (platforms, high-reach users) can disproportionately shape aggregate reliance when opinion dynamics dominate; oversight should consider not just concentration of market power but concentration of epistemic influence.
  • Platform design & incentives

    • Product designers should treat verification as an observable action and consider surfacing verified behavior (e.g., “checked this answer”) to create a stabilizing counter-feedback.
    • Reducing verification friction (better search integration, one-click source checking) helps but may be insufficient if social-proof remains strong; combining friction reduction with visibility/damping is likely most effective.
    • Tagging or crediting sources (preventing hub double-counting) matters for truthful aggregation—platforms should avoid designs that over-weight high-activity users’ outcomes without accounting for dependence.
  • Empirical research agenda for AI economics

    • Measure social-proof effects empirically: field or lab experiments that vary visibility of peers’ unverified/verified use and measure verification rates and downstream welfare.
    • Study heterogeneity in influence: when do opinion dynamics (belief transmission) dominate experience-based learning in real settings? Characterize contexts (e.g., professional networks vs. casual consumer settings).
    • Quantify policy trade-offs: simulate welfare under alternative disclosure, friction, and damping designs to inform regulation and platform choices.

Overall, the paper reframes appropriate AI reliance as a social-computational dynamics problem: economics of AI adoption and welfare must account for visibility-driven feedbacks and networked learning, not only model accuracy or per-user incentives. Designing interfaces and policies that reshape social feedback channels is central to preventing collective overreliance and preserving calibration at scale.

Assessment

Paper Typetheoretical Evidence Strengthlow — The paper provides rigorous theoretical analysis and simulation evidence that certain mechanisms can produce population-level overreliance, but it contains no empirical validation (lab, field, or observational) connecting the model parameters or observed dynamics to real-world data; causal claims about deployed systems therefore remain speculative and mechanistic rather than empirically established. Methods Rigorhigh — The modeling is carefully specified: micro-founded agent decision rules, Dirichlet belief updates, explicit utility terms, analytic propositions (including proofs), mean-field approximations, ODD+D model documentation, global sensitivity (Morris) analysis, and systematic simulation experiments (network topologies, tagging, interventions, counterfactual outcome draws). The design thoroughly explores parameter sensitivity and boundary conditions, though it remains a simulation/theory exercise without empirical calibration. SampleSimulated population of N=400 heterogeneous agents over T=100 discrete steps (20 random seeds for averaging); agent traits (ability, verification skill, initial beliefs) drawn from Beta(2,2); environments with easy/hard difficulty regimes (Beta(2,6) and Beta(6,2)); AI quality parameterized as discrete (High/Low) and continuous interpolation; networks are Erdős–Rényi or Barabási–Albert with experiments on source tagging and degree-correlated initial trust; primary metrics averaged over final 10 steps across seeds. Themeshuman_ai_collab adoption org_design IdentificationCausal mechanisms are identified via formal theoretical propositions (oracle benchmark, mean-preservation theorem, mean-field cascade tipping condition) and by agent-based simulation (micro-founded ABM with Dirichlet beliefs, multinomial logit choice, ER/BA networks) that demonstrates sufficiency of specific mechanisms (verification-suppression feedback / social proof) to generate collective overreliance; robustness is checked with ODD+D documentation, Morris global sensitivity, topology/tagging experiments, and counterfactual draws within the simulation. GeneralizabilityModel is a minimal ABM meant to show mechanism sufficiency and is not calibrated to real-world usage data or specific domains (e.g., healthcare, coding, legal), limiting external validity., Simplified task environment and stylized AI error model may not capture domain-specific error structures or user incentives., Network topologies (ER/BA) and tagging rules are simple abstractions of real organizational networks and information displays; real-world social signals are messier and multi-channel., Agent decision rules (Dirichlet belief updates, logit choice) and parameter choices (e.g., θ, verification costs) are plausible but may not match actual human cognition or institutional constraints., Scale and temporal structure (N=400, T=100) may not capture long-run institutional adaptation, policy interventions, or cross-organization effects., Cultural, regulatory, and organizational heterogeneity that shape verification norms and visibility are not represented.

Claims (11)

ClaimDirectionOutcomeConfidence & EvidenceDetails
In the simulations, task difficulty increases behavioral overreliance. Automation Exposure negative Behavioral overreliance, defined as severity-weighted unverified AI use
Reading fidelity high
Study strength medium
n=400
≈0.02 → 0.38
0.12
For hard tasks, higher AI quality reduces behavioral overreliance in the simulations. Automation Exposure positive Behavioral overreliance
Reading fidelity high
Study strength medium
n=400
0.38 → 0.16
0.12
Calibration regret and behavioral overreliance can diverge: high-quality AI on hard tasks produced the highest regret despite not producing the highest overreliance. Decision Quality mixed Calibration regret and behavioral overreliance
Reading fidelity high
Study strength medium
n=400
Regret 0.441; overreliance 0.119
0.12
Under exchangeable peer signals and without feedback, social learning produces consensus in trust but does not change aggregate overreliance. Automation Exposure null_result Aggregate behavioral overreliance and cross-sectional trust dispersion
Reading fidelity high
Study strength medium
n=400
Overreliance 0.30–0.31 across cells
0.12
Topology can shift aggregate reliance when network influence transmits correlated beliefs rather than exchangeable experience signals. Automation Exposure mixed Consensus trust and behavioral overreliance
Reading fidelity high
Study strength medium
n=400
Overreliance 0.299 → 0.272 under careful hubs; consensus trust 0.530 → 0.549 or 0.455
0.12
Increasing the social-proof strength causes verification to collapse and behavioral overreliance to rise. Automation Exposure negative Verification rate and behavioral overreliance
Reading fidelity high
Study strength medium
n=400
Verification 0.29 → 0.002; overreliance 0.30 → 0.52
0.12
The model's social-proof cascade is generated by peer behavior entering the utility of unverified AI use, rather than by worsening AI quality or user ability. Automation Exposure negative Collective overreliance and verification behavior
Reading fidelity high
Study strength medium
n=400
0.12
The agent-based model exhibits a smooth crossover rather than hysteresis in the tested social-proof sweeps. Automation Exposure null_result Path dependence and hysteresis in collective verification/reliance
Reading fidelity high
Study strength medium
n=400
No hysteresis observed
0.12
In the mean-field cascade model, bistability and hysteresis require choice decisiveness times social-proof coupling to exceed 4, although this condition is not sufficient by itself. Automation Exposure negative Existence of multiple stable reliance equilibria and hysteresis
Reading fidelity high
Study strength high
θs > 4
0.2
The model predicts that making verification visible can trigger a counter-cascade that prevents verification collapse. Automation Exposure positive Verification rate and collective overreliance
Reading fidelity high
Study strength medium
n=400
0.12
The paper's mean-preservation theorem states that DeGroot social learning with a doubly stochastic network preserves the population mean of trust and does not increase cross-sectional variance. Automation Exposure null_result Population mean trust and cross-sectional trust variance
Reading fidelity high
Study strength high
not reported
0.2

Notes