The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

AI agents draw different conclusions from the same data depending on framing, favoring outcomes that align with their latent priors; in some cases they also search and select analyses to support those priors, creating delegation risks in high‑stakes domains.

Bayesian and Motivated Reasoning in AI Agents
Eddie Yang · July 31, 2026
arxiv quasi_experimental medium evidence 8/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Eddie Yang unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Eddie Yang provider ID
When given identical numerical evidence, AI agents’ conclusions shift with substantive framing in ways aligned with their elicited priors, and some agents exhibit asymmetric search and specification choices consistent with motivated reasoning.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

AI agents increasingly perform open-ended tasks in settings where their conclusions can guide consequential decisions. We provide evidence that AI agents draw different conclusions from identical numerical data when the substantive framing changes. We demonstrate this behavior in high-stakes domains in medicine, election forensics, and geopolitical forecasting by holding the evidence fixed while changing the scenario in which the evidence appears. Across twelve agent-domain comparisons, agents' conclusions are strongly influenced by their prior beliefs. They are more likely to reach an affirmative conclusion when it is framed around a proposition they already regard as likely, while the reverse holds when the framing conflicts with their prior. The framing also changes how some agents work: they search more extensively, choose different analytical specifications, and evaluate the same evidence differently. These results identify a particular risk of delegating decision-making to AI agents, as their decisions may depend on prior beliefs that are neither specified in the task nor visible in the decision record.

Summary

Main Finding

AI agents draw different conclusions from identical numerical data when the substantive framing changes. Across medicine, election forensics, and geopolitical forecasting, agents show consistent prior-aligned framing effects (they are more likely to reach an affirmative conclusion when the frame matches their prior and less likely when it conflicts). There is also suggestive evidence of motivated reasoning — agents sometimes change their search behavior, specification choices, and evaluation of the same evidence in ways that favor prior-supported conclusions, especially in the medical tasks.

Key Points

  • Three-level behavioral framework:
    • Framing effects: conclusions change when the substantive frame changes even with identical evidence.
    • Prior-aligned framing effects (Bayesian): framing effects systematically move in the direction of the agent’s prior beliefs.
    • Motivated reasoning: the analytic process is biased (selective search, use, or interpretation of evidence) to protect priors.
  • Empirical patterns:
    • In all three domains (medicine, elections, geopolitics) affirmative conclusions and point estimates were highest in the prior-aligned frame and lowest in the prior-opposed frame, with neutral in between.
    • The medical domain produced the clearest evidence of motivated reasoning: agents ran more analyses, selected different model specifications, and treated post-exposure covariates differently depending on framing.
    • Election and geopolitical tasks showed robust prior-aligned framing but weaker and more mixed evidence for motivated (directional) analytic behavior.
  • Discretion matters:
    • Shrinking agents’ degrees of freedom (by clarifying estimands or highlighting post-exposure covariates) reduced framing and motivated-reasoning patterns, though some residual frame dependence remained.
  • Conceptual nuance:
    • Prior-aligned outcomes can be rational (Bayesian updating) if priors are genuine and known. The problem arises when priors are unelicited, unexpected, or unrecorded and so change outputs without transparency.

Data & Methods

  • Design overview:
    • Two-stage design: (1) elicit agents’ priors by pairwise comparisons among three propositions per domain; (2) matched-data framing experiments where numerical evidence (dataset, codebook, files) is held fixed and only the substantive label/frame changes (prior-aligned, neutral, prior-opposed).
  • Prior elicitation:
    • 162 pairwise comparisons per agent–domain (9 prompt versions × 3 phrasings × all pairs), pooled into Bradley–Terry log-strength scores with 0.5 pseudo-wins for numerical stability.
  • Experimental scope:
    • Domains: medicine, election fraud detection, geopolitical forecasting.
    • Ground truths in synthetic data:
      • Medicine (harmful-effect design): true adjusted odds ratio ≈ 1.8–1.9 (also a null-effect variant where post-exposure adjustment inflates harm).
      • Elections: true manipulation ≈ 0.9 percentage points (threshold tested at 0.5 pp).
      • Geopolitics: true probability of decisive success = 0.58.
    • Agents: Claude Opus 4.8, Gemini 3.5 Flash, GLM-5.2, GPT-5.6 Sol.
    • Experimental scale: 3 domains × 10 dataset versions × 3 frames × 5 runs × 4 agents = 1,800 runs.
  • Agent capabilities and sandbox:
    • Agents could inspect files, write/execute code, fit models, run robustness checks; runs recorded final answers (preferred specification, point estimate, binary decision, confidence) and analytical traces (tool calls, executed commands, reasoning where exposed).
  • Tests for motivated reasoning:
    • Asymmetric search: whether agents increase search effort or robustness checks when evidence conflicts with priors.
    • Asymmetric updating/fishing: whether agents pick different “preferred” specifications across frames to obtain prior-consistent outcomes.
  • Outcome analysis:
    • Compare proportions of affirmative conclusions and average point estimates across frames; correlate framing-induced differences with elicited priors; examine behavioral traces for asymmetric search/fishing evidence.

Implications for AI Economics

  • Delegation risk and information aggregation:
    • When economists or policy actors delegate data analysis to AI agents, unelicited agent priors can materially affect conclusions even with identical data. This undermines the transparency of inference and can bias economic decision-making, forecasting, and policy evaluation.
  • Research reproducibility and researcher degrees of freedom:
    • AI agents introduce a new, opaque source of researcher degrees of freedom: implicit priors and automated analytic search strategies. This aggravates reproducibility issues unless priors and analytic traces are recorded and standardized.
  • Market and institutional design:
    • Products, platforms, and contracting around analytic agents should require disclosure or specification of priors, standard estimands, and constrained analytic protocols to avoid hidden bias. For marketplaces using AI for forecasting or risk assessment, certification or calibration of agent priors could become a competitive dimension.
  • Regulation, liability, and accountability:
    • Economic policies that rely on agent-produced analyses (e.g., algorithmic procurement, regulatory risk assessments) should consider rules for auditing analytical traces and requiring provenance for priors and model choices.
  • Guidance for practitioners and policy:
    • Elicit and document agent priors before delegating tasks where impartial evidence synthesis is required.
    • Constrain degrees of freedom: fix estimands, pre-specify covariates and robustness checks, or require formal model selection criteria to limit fishing.
    • Preserve and audit analytical traces: tool calls, code, model specifications, and intermediate outputs should be recorded for post hoc review.
    • Use ensemble and calibration methods: combine multiple agents with different elicited priors, or require simulation-based calibration to detect undue sensitivity to framing.
  • Research directions important for AI economics:
    • Quantify welfare consequences of prior-induced decision heterogeneity in market/policy settings.
    • Design incentive-compatible mechanisms to reveal or align agent priors with principal objectives.
    • Develop standards and tools for priors disclosure, traceable analysis pipelines, and robustness certification for agent-produced inferences.

Summary takeaway: AI agents can behave like Bayesian analysts whose priors matter — and sometimes like motivated reasoners who tilt their analytic process to protect priors. For economists and policymakers who rely on delegated AI analyses, the remedy is not banning priors but making them visible, constraining discretionary analysis, and auditing analytical traces so that conclusions can be interpreted correctly and decisions remain accountable.

Assessment

Paper Typequasi_experimental Evidence Strengthmedium — The design tightly controls evidence and measures priors, enabling credible identification of framing and prior-aligned effects; multiple agents, dataset versions, and repeated runs increase robustness. However, the use of synthetic datasets, a small set of contemporary agent models, and the lab sandbox setting limit external validity and the strength of causal claims about deployed systems. Methods Rigorhigh — Careful prior elicitation (many randomized pairwise prompts, Bradley–Terry aggregation), fixed-data matched-frame experiments, multiple dataset draws and repeated runs, and recording of tool calls and code executed to trace analytical process are strong design features that directly address alternative explanations; pre-specified variations to reduce discretion further probe mechanisms. Limitations include synthetic DGP choices that instantiate specific 'forking paths' and dependence on provider-exposed traces and model versions. SampleControlled experiments using synthetic datasets: medical domain — 10 dataset versions with 20,000 synthetic health records each (harmful-effect and null-effect DGP variants, plus two restricted-discretion variants); election domain — 10 dataset versions with ~12,000–16,000 reporting units across eight regions; geopolitical domain — 10 versions with 843 synthetic records from analogs, expert assessments, and simulations. Four agent models (Claude Opus 4.8, Gemini 3.5 Flash, GLM-5.2, GPT-5.6 Sol), five independent runs per version-frame cell, three frames per domain (prior-aligned, neutral, prior-opposed), yielding 1,800 main runs; prior elicitation uses 162 randomized pairwise comparisons per domain/agent. Themeshuman_ai_collab governance IdentificationWithin-agent framed experiment that holds numerical evidence fixed while varying only the substantive frame; independent elicitation of each agent's priors via many randomized pairwise comparisons (Bradley–Terry scores); multiple synthetic dataset versions and repeated agent runs; comparison of binary conclusions and point estimates across prior-aligned, neutral, and prior-opposed frames, plus analysis of recorded analytical traces (tool calls, specs) to distinguish Bayesian updating from motivated reasoning. GeneralizabilitySynthetic datasets may not capture complexities and noise of real-world case data and institutional workflows., Only four specific agent models and particular provider sandbox configurations were tested; results may not hold for other or future models or different tool integrations., Laboratory framing and prompt designs may differ from how human users describe tasks in practice, affecting external validity., Behavior recorded in a read-only sandbox with provider-exposed traces may differ from deployed agents with hidden state, tool access, or product-level prompt engineering., Time-sensitivity: model priors and behaviors can change rapidly with model updates and fine-tuning, limiting longitudinal applicability.

Claims (9)

ClaimDirectionOutcomeConfidence & EvidenceDetails
AI agents draw different conclusions from identical numerical data when the substantive framing changes. Ai Safety And Ethics negative Variation in agents' binary conclusions and point estimates across substantively different frames with identical data
Reading fidelity high
Study strength high
n=1800
0.8
Across the three domains, agents' affirmative conclusions were most frequent in the prior-aligned frame, least frequent in the prior-opposed frame, and intermediate in the neutral frame. Decision Quality positive Rate of affirmative binary conclusions
Reading fidelity high
Study strength medium
n=1800
0.48
The framing pattern also appears in the point estimates that agents derive from the data. Decision Quality positive Frame-specific point estimates, including odds ratios, election margins, and success probabilities
Reading fidelity high
Study strength medium
n=1800
0.48
Agents' framing effects were aligned with their prior beliefs: agents were more likely to reach an affirmative conclusion when the proposition was framed around a proposition they regarded as more likely. Ai Safety And Ethics positive Alignment between elicited prior-belief ordering and final affirmative conclusions
Reading fidelity high
Study strength medium
n=4
0.48
Some agents changed their analytical behavior across frames, including search effort, analytical specifications, and evaluation or selection of evidence. Ai Safety And Ethics mixed Analytical search behavior, model specification choices, and evidence use
Reading fidelity high
Study strength medium
n=1800
0.48
Evidence of motivated reasoning was strongest in the medical task and weaker in the election and geopolitical tasks. Ai Safety And Ethics mixed Strength of directional, prior-favoring analytical-process behavior across domains
Reading fidelity high
Study strength medium
n=1800
0.48
Reducing the space of agent discretion in two altered versions of the medical experiment reduced the observed framing effects, although some residual differences remained. Ai Safety And Ethics negative Magnitude or persistence of framing differences under reduced analytical discretion
Reading fidelity high
Study strength medium
n=4
0.48
In the medical harmful-effect design, adjustment for demographic variables and pre-exposure clinical confounders recovered the harmful ground truth, with an odds ratio of approximately 1.8 to 1.9. Decision Quality positive Adjusted odds ratio for the causal effect of exposure on incident pharyngeal cancer
Reading fidelity high
Study strength high
n=20000
odds ratio of approximately 1.8 to 1.9
0.8
In the medical harmful-effect design, adjusting for post-exposure follow-up visits and medication use attenuated the estimated effect, and a propensity-score analysis including those post-exposure measures could produce a null or protective estimate. Decision Quality negative Estimated causal effect of exposure on incident pharyngeal cancer
Reading fidelity high
Study strength high
n=20000
0.8

Notes