The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Large language models' ethical advice swings with surface framing: scarcity cues increase approval by 12 percentage points while high-stakes outcomes and fair-process prompts reduce it by about 10–11 points. When multiple contextual cues align, recommendations can shift by more than half the probability scale, indicating advisory outputs are highly sensitive to framing rather than fixed principles.

When Algorithms Meet Ethics: Systematic Evidence of Framing Effects in LLM Organizational Decision-Making
Jonathan H. Westover · February 15, 2026 · Preprints.org
openalex rct high evidence 7/10 relevance Summary only summary available; pdf_status=paywall DOI Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

OpenAlex

Latest observation:

  1. Jonathan H. Westover provider ID
In a preregistered factorial experiment on 14,306 LLM responses, contextual framing substantially shifts ethical recommendations—resource scarcity raises endorsement probability by 12pp while outcome severity and procedural justice reduce it by ~11pp and ~10pp respectively, with cumulative framing effects up to ~27pp and a maximum-to-minimum range of ~54pp.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Large language models (LLMs) are increasingly deployed as decision-support tools in organizational contexts, yet their susceptibility to contextual framing remains poorly understood. This preregistered experimental study systematically examines how six framing dimensions—procedural justice, outcome severity, stakeholder power, resource scarcity, temporal urgency, and transparency requirements—influence ethical recommendations from three frontier models: Claude 3.5 Sonnet, GPT-4o, and Gemini 1.5 Pro. We developed 5,000 unique organizational vignettes using a fractional factorial experimental design with balanced industry representation, generating 15,000 total model responses. After excluding responses without clear recommendations (n=694, 4.6%), we analyzed 14,306 responses using logistic regression with robust and clustered standard errors. We find that resource scarcity increases endorsement probability by 12.0 percentage points (pp) (OR = 1.67, 95% CI [1.45, 1.93], p < .001), while outcome severity reduces it by 11.3pp (OR = 0.62, 95% CI [0.54, 0.71], p < .001), and procedural justice reduces it by 10.1pp (OR = 0.66, 95% CI [0.57, 0.76], p < .001). These effect sizes are comparable to classical framing research (Tversky & Kahneman: 22pp; McNeil et al.: 18pp) and represent substantial shifts in organizational decision contexts. When multiple framing dimensions align in ethically unfavorable directions, cumulative effects reach approximately 27pp from baseline (range: 25-28pp depending on interaction assumptions), with maximum-to-minimum framing creating a 54-percentage-point total range approaching complete recommendation reversals. Effects appear consistently across all three models, with no significant Dimension × Model interactions, suggesting fundamental architectural properties rather than implementation-specific artifacts. Topic modeling of justification text from the 14,306 analyzed responses reveals systematic "adaptive rationalization"—models invoke utilitarian reasoning when contexts emphasize constraints (+6.7pp in high resource scarcity), deontological reasoning when contexts emphasize high stakes (+2.4pp in high outcome severity), and virtue/justice ethics when contexts emphasize fair processes (+4.4pp in high procedural justice). This suggests models select ethical frameworks to justify contextually appropriate conclusions rather than applying consistent principles across situations. Human validation confirms these patterns reflect genuine framing sensitivity rather than measurement artifacts. Crowdworker validation (n=7,500 responses, one rater each) achieved substantial agreement (Fleiss' κ = 0.71) and 81.3% concordance with expert codings. Subject matter expert evaluation (n=24 experts, 100 vignette pairs each including 20 control pairs, 2,400 total comparisons) detected framing-driven differences in 48.9% of pairs (net of 18.3% baseline false-positive rate), but correctly attributed differences to manipulated dimensions in only 41.3% of cases. Most detected differences (58.7%) were judged problematic for AI advisory systems. These findings raise fundamental questions about deploying LLMs for consequential organizational decisions where surface features may inappropriately influence outcomes. We discuss implications for AI governance, organizational ethics, and the design of more robust decision-support systems.

Summary

Main Finding

Frontier LLMs (Claude 3.5 Sonnet, GPT-4o, Gemini 1.5 Pro) show large, systematic sensitivity to contextual framings in organizational ethical vignettes. Specific framing dimensions (resource scarcity, outcome severity, procedural justice) change the probability models endorse ethically consequential recommendations by substantial margins (≈10–12 percentage points each), with cumulative and extreme framing creating up to a ~54 percentage-point swing. These effects are consistent across models and are accompanied by systematic shifts in the ethical rationales models produce (“adaptive rationalization”).

Key Points

  • Scope and design: Preregistered factorial experiment generating 5,000 unique organizational vignettes (balanced across industries) and collecting 15,000 model responses (3 models × 5,000 vignettes).
  • Analysis sample: 694 responses (4.6%) without clear recommendations were excluded, leaving 14,306 responses analyzed via logistic regression with robust, clustered standard errors.
  • Principal effects (relative to baseline endorsement probability):
    • Resource scarcity: +12.0 percentage points (pp); OR = 1.67, 95% CI [1.45, 1.93], p < .001.
    • Outcome severity: −11.3 pp; OR = 0.62, 95% CI [0.54, 0.71], p < .001.
    • Procedural justice (emphasizing fair process): −10.1 pp; OR = 0.66, 95% CI [0.57, 0.76], p < .001.
  • Effect sizes comparable to classic human framing results (Tversky & Kahneman ~22pp; McNeil et al. ~18pp) and thus substantively meaningful for organizational decision contexts.
  • Cumulative and extreme framing:
    • When multiple dimensions align in ethically unfavorable directions, net effects ≈ 27pp from baseline (range 25–28pp under different interaction assumptions).
    • Maximum-to-minimum framing scenarios produce ~54pp total range—approaching full reversal of recommendations.
  • Model-general phenomenon: No significant Dimension × Model interactions—effects were consistent across the three models, suggesting core architectural or training-related sensitivity rather than model-specific quirks.
  • Adaptive rationalization (from topic modeling of justification text):
    • Under high resource scarcity, models more often invoke utilitarian reasoning (+6.7 pp).
    • Under high outcome severity, models more often invoke deontological reasoning (+2.4 pp).
    • Under high procedural justice cues, models more often invoke virtue/justice ethics (+4.4 pp).
    • Interpretation: models adjust their ethical framing to match contextually salient constraints rather than applying a single consistent ethical principle.
  • Human validation:
    • Crowdworker coding: n = 7,500 responses (one rater each), Fleiss’ κ = 0.71; 81.3% concordance with expert codings.
    • Subject-matter experts: n = 24 experts, 100 vignette pairs each (including controls) → 2,400 comparisons. Experts detected framing-driven differences in 48.9% of pairs (after accounting for an 18.3% baseline false-positive rate) but correctly attributed differences to the manipulated dimensions in only 41.3% of cases. Of detected differences, 58.7% were judged problematic for AI advisory systems.

Data & Methods

  • Experimental design: Pre-registered fractional factorial design producing 5,000 unique vignettes varying six framing dimensions:
  • Procedural justice
  • Outcome severity
  • Stakeholder power
  • Resource scarcity
  • Temporal urgency
  • Transparency requirements
  • Balanced industry sampling to improve external validity across organizational contexts.
  • Models tested: Claude 3.5 Sonnet, GPT-4o, Gemini 1.5 Pro.
  • Responses: Each vignette was run through each model → 15,000 total responses; 694 ambiguous/no-recommendation responses excluded (4.6%).
  • Outcome variable: Binary endorsement recommendation (model recommends the ethically questionable action vs. not).
  • Statistical analysis: Logistic regression with robust and clustered standard errors; reported odds ratios, percentage-point changes (marginal effects), 95% CIs, and p-values. Interaction tests for Dimension × Model were performed.
  • Qualitative/interpretive analysis: Topic modeling of justification text to identify ethical frameworks; human coding and validation (crowdworkers and subject-matter experts) to verify framing sensitivity is substantive and not an artifact of the measurement pipeline.

Implications for AI Economics

  • Decision-support reliability: The magnitude of framing effects implies LLM-based advisors can be substantially swayed by superficial contextual cues. For organizational economics, this threatens consistent policy recommendations, risk assessments, and allocation decisions—introducing endogenous variability tied to prompt/context rather than underlying fundamentals.
  • Governance and auditing: Effective deployment requires mandatory stress-testing across realistic framing variations (a “framing robustness” dimension in audits), standardized evaluation protocols, and regulatory disclosure of model sensitivity to contextual features.
  • Organizational design & incentives: Firms relying on LLM recommendations should avoid single-model point estimates for consequential decisions. Use ensembles, counterfactual framing checks, and calibrated decision rules (e.g., decision thresholds invariant to framing) to reduce manipulable variation.
  • Market design and strategic behavior: If LLM advice is frame-sensitive, actors (internal or external) can potentially game outcomes by manipulating presentation or context cues. This raises new considerations for incentive-compatible mechanism design and governance structures to deter strategic framing.
  • Model alignment and training priorities: Consistent cross-model sensitivity suggests training and alignment procedures must prioritize invariance to irrelevant contextual framing and foster stable ethical reasoning rather than adaptive rationalization that mirrors the most salient cues.
  • Cost-benefit trade-offs and regulation: Economists and policymakers should incorporate framing-robustness metrics into cost–benefit analyses of automated decision-support adoption, and consider regulation that mandates robustness testing for high-stakes applications (finance, health, HR, safety-critical operations).
  • Research directions: Quantify economic impacts of framing-induced recommendation variability (e.g., error rates, welfare consequences), develop formal measures of "framing elasticity" for models, and test mitigation strategies (prompt engineering standards, constrained decoding, supplementary rule-based checks, human-in-the-loop safeguards).

Limitations to note: laboratory vignette setting may not capture all field complexities; only three models were tested (though effects were consistent across them); 4.6% of outputs were uninterpretable and excluded; attribution accuracy by experts was imperfect—practical detection of framing effects remains nontrivial. These caveats reinforce the need for operational robustness checks in real deployments.

Assessment

Paper Typerct Evidence Strengthhigh — Large preregistered experimental sample (15,000 model responses, 14,306 analyzed), randomized manipulation of key framing dimensions enabling causal interpretation of framing → model recommendation effects, precise effect estimates with confidence intervals, replication across three models, and convergent human-validation evidence. Methods Rigorhigh — Study uses a preregistered fractional-factorial design, balanced sampling, multiple state-of-the-art LLMs, appropriate exclusion rules, clustered robust inference, topic-modeling for mechanism exploration, and both crowdworker and expert validation; limitations (single-rater crowd validation, some excluded/no-decision responses, and potential prompt/implementation sensitivity) are acknowledged but do not undermine core internal validity. Sample5,000 unique organizational vignettes constructed via a fractional-factorial design with balanced industry representation; 3 LLMs (Claude 3.5 Sonnet, GPT-4o, Gemini 1.5 Pro) produced 15,000 responses, 694 responses without clear recommendations were excluded, leaving 14,306 responses analyzed; human validation comprised 7,500 crowdworker ratings (one rater each) and 24 subject-matter experts who evaluated 100 vignette pairs each (2,400 comparisons). Themeshuman_ai_collab governance IdentificationPreregistered fractional-factorial randomized experiment manipulating six framing dimensions across 5,000 unique organizational vignettes (balanced by industry) and measuring model recommendations from three LLMs; causal effects estimated via logistic regression with robust, clustered standard errors and supplemented by human validation and topic-modeling of justifications. GeneralizabilityFindings are limited to the three specific model versions and may not hold for other models or future updates/finetuned variants, Vignettes are synthetic experimental scenarios and may not capture full complexity of real organizational decision-making, Prompting style, model temperature/settings, and deployment guardrails could materially change results, Balanced industry sampling may not capture geographic, cultural, or regulatory context variation, Human validation used one crowdworker rater per item and a modest expert sample, which may limit external validity of human-label concordance estimates

Claims (15)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Resource scarcity increases endorsement probability by 12.0 percentage points (OR = 1.67, 95% CI [1.45, 1.93], p < .001). Decision Quality positive probability that the model endorses the ethically questionable action (endorsement probability)
Reading fidelity high
Study strength high
n=14306
12.0 percentage points (OR = 1.67, 95% CI [1.45, 1.93], p < .001)
1.0
Outcome severity reduces endorsement probability by 11.3 percentage points (OR = 0.62, 95% CI [0.54, 0.71], p < .001). Decision Quality negative probability that the model endorses the ethically questionable action (endorsement probability)
Reading fidelity high
Study strength high
n=14306
11.3 percentage points (OR = 0.62, 95% CI [0.54, 0.71], p < .001)
1.0
Procedural justice reduces endorsement probability by 10.1 percentage points (OR = 0.66, 95% CI [0.57, 0.76], p < .001). Decision Quality negative probability that the model endorses the ethically questionable action (endorsement probability)
Reading fidelity high
Study strength high
n=14306
10.1 percentage points (OR = 0.66, 95% CI [0.57, 0.76], p < .001)
1.0
When multiple framing dimensions align in ethically unfavorable directions, cumulative effects reach approximately 27 percentage points from baseline (range: 25–28pp depending on interaction assumptions). Decision Quality positive net change in endorsement probability from baseline under aligned framing
Reading fidelity high
Study strength high
n=14306
approximately 27 percentage points (range: 25-28pp depending on interaction assumptions)
1.0
Maximum-to-minimum framing creates a 54-percentage-point total range in recommendations (approaching complete recommendation reversals). Decision Quality mixed difference in endorsement probability between most and least favorable framing configurations
Reading fidelity high
Study strength high
n=14306
54-percentage-point total range
1.0
Effects appear consistently across all three models, with no significant Dimension × Model interactions. Decision Quality null_result presence/absence of significant interactions between framing dimensions and model identity on endorsement probability
Reading fidelity high
Study strength high
n=14306
no significant Dimension × Model interactions (as reported)
1.0
Topic modeling of justification text reveals adaptive rationalization: models invoke utilitarian reasoning when contexts emphasize constraints (+6.7pp in high resource scarcity). Decision Quality positive change in prevalence of utilitarian-style justifications in model-generated explanations
Reading fidelity high
Study strength medium
n=14306
+6.7 percentage points in utilitarian reasoning prevalence under high resource scarcity
0.6
Topic modeling shows models invoke deontological reasoning when contexts emphasize high stakes (+2.4pp in high outcome severity). Decision Quality positive change in prevalence of deontological-style justifications in model-generated explanations
Reading fidelity high
Study strength medium
n=14306
+2.4 percentage points in deontological reasoning prevalence under high outcome severity
0.6
Topic modeling shows models invoke virtue/justice reasoning when contexts emphasize fair processes (+4.4pp in high procedural justice). Decision Quality positive change in prevalence of virtue/justice-style justifications in model explanations
Reading fidelity high
Study strength medium
n=14306
+4.4 percentage points in virtue/justice reasoning prevalence under high procedural justice
0.6
Human (crowdworker) validation: n=7,500 responses, one rater each, achieved substantial agreement (Fleiss' κ = 0.71) and 81.3% concordance with expert codings. Output Quality positive inter-rater agreement (Fleiss' κ) and concordance rate with expert labels
Reading fidelity high
Study strength medium
n=7500
Fleiss' κ = 0.71; 81.3% concordance with expert codings
0.6
Subject matter expert (SME) evaluation: n=24 experts, 100 vignette pairs each (2,400 total comparisons); SMEs detected framing-driven differences in 48.9% of pairs (net of 18.3% baseline false-positive rate), but correctly attributed differences to manipulated dimensions in only 41.3% of cases. Decision Quality positive rate of detected framing-driven differences and correct attribution to manipulated dimensions by SMEs
Reading fidelity high
Study strength medium
n=2400
48.9% detection (net of 18.3% baseline false-positive); 41.3% correct attribution
0.6
Most detected differences (58.7%) were judged problematic for AI advisory systems. Decision Quality negative proportion of detected framing-driven differences judged problematic by experts
Reading fidelity high
Study strength medium
n=2400
58.7% of detected differences judged problematic
0.6
The experimental study generated 5,000 unique organizational vignettes using a fractional factorial design with balanced industry representation, producing 15,000 total model responses; 694 responses (4.6%) lacked clear recommendations and were excluded, leaving 14,306 responses analyzed with logistic regression using robust and clustered standard errors. Other null_result dataset composition and analytic approach (sample sizes, exclusions, regression method)
Reading fidelity high
Study strength high
n=14306
15,000 total responses generated; 694 excluded (4.6%); 14,306 responses analyzed
1.0
These effect sizes are comparable to classical framing research (Tversky & Kahneman: 22pp; McNeil et al.: 18pp). Decision Quality mixed magnitude of framing-effect percentage-point shifts compared to prior classical studies
Reading fidelity high
Study strength medium
comparative claim to Tversky & Kahneman (22pp) and McNeil et al. (18pp)
0.6
Models select ethical frameworks to justify contextually appropriate conclusions rather than applying consistent principles across situations (interpretation of adaptive rationalization patterns). Decision Quality mixed consistency vs. context-sensitivity of ethical frameworks in model justifications
Reading fidelity high
Study strength medium
n=14306
interpretive claim supported by topic-model prevalence shifts (e.g., +6.7pp utilitarian under scarcity, +2.4pp deontological under high stakes, +4.4pp virtue/justice under procedural justice)
0.6

Notes