The Commonplace
Home Three-study pilot Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Survey layout, not model creed: pinning options to fixed answer slots generated a robust ‘market‑rationalist’ cluster across LLMs, but a balanced randomized-order replication shows the apparent ideology collapses—answer‑slot bias, not underlying model beliefs, explains most of the effect.

Ideology by Alphabet: Option Order and the Measurement of Machine Political Preferences
Tamas Olah, Laszlo Erdey, Tibor Tokes, Levente Nadasi · September 17, 2026 · Research Square
openalex rct high evidence 7/10 relevance Full text usable extracted full text DOI Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

OpenAlex

Latest observation:

  1. Tamas Olah provider ID
  2. Laszlo Erdey provider ID
  3. Tibor Tokes provider ID
  4. Levente Nadasi provider ID
Fixed-order forced-choice instruments can manufacture coherent 'ideological' clusters in LLMs because persistent answer-slot (position) bias, not substantive content preference, drives much of the measured political orientation, a result that disappears under balanced randomized ordering.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Summary

Main Finding

Fixed-order forced-choice instruments can manufacture apparent political “ideologies” in large language models. In a 400-question governance benchmark administered to 15 open-weight models, a conventional fixed-option-order analysis produced a neat two-cluster ideological map (a small “market-rationalist” cluster vs. an “institutionalist” majority). That structure almost entirely disappeared when option order was randomized in a fully balanced way: the measured ideology correlated essentially zero with the de-confounded (content-only) profile and instead tracked a strong, model-specific position bias (especially first-slot preference).

Key Points

  • Experiment scale and outcome
    • 400 governance scenarios, 4-option forced-choice each, administered to 15 open-weight models.
    • Fixed-order run: 640,000 responses (multiple temperatures and repeats). Randomized-order replication: ≈48k parsed responses (balanced permutations).
  • Spurious ideology under fixed-order
    • Fixed-order analysis yielded two stable clusters (market-rationalist minority of 3 models vs. institutionalist majority of 12), with split‑half reliability ρ ≈ 0.91.
    • Cluster differences: utility +23.3, metrics +20.4, incentives +15.0 (market-rationalists higher).
  • Position bias drives the apparent structure
    • Balanced randomized-order design (each content occupies each display slot equally within every question) produces:
    • De-confounded map reliable (split-half ρ ≈ 0.87) but uncorrelated with the fixed-order map (ρ = 0.07).
    • The market-rationalist signature correlates r = 0.91 with a model’s first-slot bias.
    • Position bias is large and extremely stable across runs (slot bias split-half r ≈ 0.998 for pooled slot means).
  • Practical impact on recommendations
    • Changing only option order (holding model, question, and content fixed) alters the top recommendation on 71% of model–question pairs vs. a 33% run-to-run noise floor.
    • Self-reported confidence barely changes despite frequent reversals.
    • Reversals concentrate on low-signal questions, consistent with satisficing/primacy effects.
  • Multiverse sensitivity
    • Because randomized data sampled all content-slot pairings, authors re-ran the benchmark under all 13,824 alternative content-to-slot assignments.
    • The originally chosen arrangement (which listed market-favored options first) was the single most misleading assignment.
    • Nine of 15 models would be labeled “market‑rationalist” under at least one slot arrangement; the model most often so labeled (qwen3:8b) was not among the three from the fixed-order run.
    • About a quarter of slot assignments would produce no coherent cluster at all.
  • What survives de-confounding
    • A state-vs-market index (constructed from responsibility and instrument frames) is relatively robust (ρ ≈ 0.79), but even it can be materially shifted for many models by slot assignment.
  • Other diagnostics
    • Temperature had negligible effect on the fixed-order map.
    • Confidence and rationale text provide only weak, inconclusive signals about the position-confound.
  • Boundaries and caveats
    • All evidence is from reproducible open-weight models in the 1–20B parameter range; results are not tested on closed frontier commercial models.
    • The forced-spread (≥30-point spread) scoring instruction may inflate magnitude of bias but cannot explain directional, model-specific slot habits.

Data & Methods

  • Stimuli and framing
    • 400 scenarios across four frame families (100 each): responsibility locus (who), policy instrument (how), ethical principle (why), evidence stance (on what basis).
    • Each family had four consistent options (canonical order used in fixed-order run).
  • Elicitation protocol
    • Prompt: score each of four labeled options 0–100, name top option, report confidence 0–100, and provide a short justification. Prompt requested at least a 30-point spread to discourage ties.
  • Fixed-order run
    • Canonical mapping: each option content fixed to one display slot across all scenarios.
    • Crossed with two temperatures (T = 0.2, 0.8) and repeated runs to assess reliability.
  • Randomized replication (de-confounding)
    • For each question, 8 permutations of the 4 contents across slots were administered such that each content occupied each slot exactly twice (content × slot orthogonal within every question).
    • This yields a clean estimate of position bias (mean centered score by display slot) and a de-confounded content-driven fingerprint (mean centered score by canonical content).
  • Analysis
    • Scores row-centered within each response to remove scale differences.
    • Distances between model fingerprints are the main inference target.
    • Dimensionality reduction and grouping: Multiple Factor Analysis (MFA), hierarchical clustering, archetypal analysis.
    • Reliability: split-half correlations reported for fixed-order (≈0.91) and randomized (≈0.87) fingerprints; position bias shown extremely stable (≈0.998).
    • Multiverse: exhaustive re-computation over all 13,824 content-to-slot assignments enabled by balanced randomized data.
  • Robustness checks
    • Specification curves, alternative encodings, leave-one-out perturbations, temperature variation.
    • Permutation tests to assess model identity vs. temperature effects.

Implications for AI Economics

  • Measurement validity
    • Forced-choice, fixed-order instruments can produce coherent-seeming ideological structure that primarily reflects display position, not substantive model preference. Studies that infer machine political leanings from single, fixed-order multiple-choice batteries risk conflating position bias with content preference.
  • Audit and certification practice
    • Model audits, regulatory assessments, or certifications that use forced-choice instruments should:
    • Randomize option order and use balanced placement so each content appears in each slot sufficiently often (or report why not).
    • Report position-bias diagnostics (mean centered scores by display slot) alongside substantive results.
    • Include multiverse or sensitivity analyses over content-to-slot assignments where possible.
  • Research design recommendations
    • Prefer elicitation formats less vulnerable to simple positional primacy (e.g., balanced randomization, many repeated permutations, pairwise comparisons, or well-designed Likert scales that control acquiescence); when forced-choice is used, ensure content-slot orthogonality.
    • Report split-half reliabilities for de-confounded fingerprints and assess how much variation is explained by slot vs. content.
    • Where publication or policy claims rely on a single fixed-order administration, explicitly acknowledge the slot-assignment degree of freedom and quantify its possible effects.
  • Theoretical and policy consequences
    • Treating language models as “revealed-preference” agents in political-economy settings requires careful instrument design: apparent “governance philosophies” may be joint properties of models plus arbitrary measurement choices rather than intrinsic model commitments.
    • Because cheap, widely reused LLMs can act as low-cost advisors, measurement artifacts could mislead policy makers and the public about the distribution of machine advice—potentially affecting delegation, regulatory focus, or model-selection decisions.
  • Limits and future work
    • Need to test whether these position effects and magnitudes generalize to larger closed proprietary models and to other elicitation families.
    • Explore corrective analytics (e.g., statistical adjustment for position bias when balanced randomization is infeasible) and alternative elicitation protocols that minimize satisficing or primacy in models.

Overall takeaway: do not trust fixed-order forced-choice ideological maps of LLMs without balanced randomization and explicit diagnostics for position bias.

Assessment

Paper Typerct Evidence Strengthhigh — Large-scale, pre-registered-style empirical exercise with 640,000 fixed-order responses and a balanced randomized-order replication (≈48k parsed responses) that fully orthogonalizes content and slot within each question; high internal reliability (split-half ρ ≈ 0.87–0.91), permutation tests, sensitivity analyses across temperatures, and a complete multiverse over all content-to-slot permutations provide convergent evidence that position bias is a major driver of the fixed-order result. Limitations (model sample, prompt choices) are acknowledged but do not undermine the primary identification. Methods Rigorhigh — Design uses within-question balanced randomization so every content appears in every slot exactly twice (cleanly separating content and position effects), records both slot-indexed and content-indexed outcomes, runs many repeats, reports split-half reliabilities, conducts permutation/Freedman–Lane tests, enumerates all alternative slot assignments (multiverse), and performs robustness checks (temperatures, leave-one-out). Transparent about exclusions and scope. Remaining concerns are limited: sample restricted to open-weight 1–20B models, one prompt feature (forced 30-point spread) may amplify magnitudes, and scenarios/options authored by a single team. SampleBenchmark of 400 governance scenarios (100 per of 4 frame families) administered to a panel of 16 open-weight LLMs (one later excluded for parse failures, leaving 15 models in main analyses) in the 1–20 billion-parameter range; fixed-order run: 4 families × 100 questions × 16 models × 2 temperatures × 50 repeats = 640,000 raw responses (one model excluded post-hoc); randomized-order replication: balanced within-question permutations across 15 retained models at T=0.2 (≈47,966 parsed responses); each prompt elicited 0–100 scores for four labeled options, a lead option, confidence (0–100), and a short rationale. Themesgovernance adoption IdentificationRandomized within-question permutation of option-to-slot assignments (balanced so each content occupies each display slot equally within every question), combined with a fixed-order baseline; contrasts between fixed-order and balanced randomized-order runs isolate position (slot) bias from content preferences, and a full multiverse enumeration of all possible slot assignments quantifies how design choice affects measured ideology. GeneralizabilityResults are established for open-weight models in the 1–20B parameter range and may not hold quantitatively for larger commercial frontier models or heavily fine-tuned/production systems., Findings pertain to forced-choice, four-option, scored instruments with a required spread (the 30-point instruction) and may differ for Likert, free-text, or conversational elicitation formats., Scenarios and option phrasings were authored by a single research team; different content sets or culturally specific items might interact differently with position bias., The study estimates measurement artifacts (position bias) in a controlled benchmark, not normative correctness of model positions or real-world political influence of deployed systems., Presentation modality (visual GUI, conversational turn-taking, voice interfaces) could change order effects and is not tested here.

Claims (14)

ClaimDirectionOutcomeConfidence & EvidenceDetails
A conventional fixed-order analysis of the benchmark produces a two-cluster map of the 15 models, consisting of a 12-model “institutionalist” cluster and a three-model “market-rationalist” cluster. Ai Safety And Ethics positive Model-level political/governance preference clustering under fixed option order
Reading fidelity high
Study strength medium
n=15
12 versus 3 models
0.6
Under fixed option order, the apparent market-rationalist cluster scores higher than the institutionalist cluster on utility ethics, metrics-first evidence, and incentive-based instruments. Ai Safety And Ethics positive Relative model preference for governance frames
Reading fidelity high
Study strength medium
n=15
utility +23.3, metrics-first +20.4, incentives +15.0 centered points
0.6
Balanced randomization of option order dissolves the fixed-order ideological structure: the de-confounded map is nearly uncorrelated with the fixed-order map. Ai Safety And Ethics negative Correspondence between fixed-order and position-controlled model ideology maps
Reading fidelity high
Study strength high
n=15
ρ = 0.07
1.0
The apparent fixed-order ideological signature is strongly associated with first-slot bias rather than option content. Ai Safety And Ethics positive Association between apparent ideology and first-slot response bias
Reading fidelity high
Study strength high
n=15
r = 0.91
1.0
Eleven of the 15 models exhibit a position bias exceeding five centered score points, with the strongest first-slot bias reaching +20.1 centered points. Ai Safety And Ethics positive Display-slot preference
Reading fidelity high
Study strength high
n=15
11 of 15 models above 5 points; maximum +20.1 centered points
1.0
Changing only the display order changes the model’s top recommendation on approximately 71% of model-question pairs. Decision Quality positive Stability of the top recommendation under option-order changes
Reading fidelity high
Study strength high
n=15
70.6% of model-question pairs
1.0
The randomized-order de-confounded map remains internally reliable, with split-half reliability of 0.87. Ai Safety And Ethics positive Reliability of content-based model preference profiles
Reading fidelity high
Study strength medium
n=15
split-half ρ = 0.87
0.6
The randomized-order data permit reconstruction of the benchmark under all 13,824 possible assignments of option contents to display slots. Governance And Regulation mixed Sensitivity of measured model ideology to instrument design
Reading fidelity high
Study strength high
n=15
13,824 alternative arrangements
1.0
The particular fixed-order arrangement used by the authors was the most misleading of the 13,824 arrangements examined. Governance And Regulation negative Instrument-design-induced misleading ideological classification
Reading fidelity high
Study strength medium
n=15
1st most misleading of 13,824 arrangements
0.6
Nine of the 15 models are classified as market-rationalist under at least one alternative option arrangement, and approximately one quarter of arrangements produce no cluster at all. Ai Safety And Ethics mixed Robustness of model ideological classification to option arrangement
Reading fidelity high
Study strength medium
n=15
9 of 15 models; roughly one quarter of arrangements
0.6
The model most frequently classified as market-rationalist across alternative arrangements is qwen3:8b, which receives that label in 69% of arrangements, despite not being selected by the authors’ original fixed-order arrangement. Ai Safety And Ethics positive Frequency of market-rationalist model classification across instrument specifications
Reading fidelity high
Study strength medium
n=15
69% of arrangements
0.6
Option-order reversals are more common on weak-signal questions than on strong-signal questions. Decision Quality negative Recommendation reversal rate by question signal strength
Reading fidelity high
Study strength medium
n=15
83% on weak-signal questions versus 59% on strong-signal questions
0.6
Self-reported confidence changes little when option order changes, despite the large change in recommendations. Ai Safety And Ethics null_result Self-reported confidence under option-order changes
Reading fidelity high
Study strength medium
n=15
about one point on a 100-point scale; mean confidence 86
0.6
The study’s evidence is limited to open-weight models with 1–20 billion parameters, so whether the same effect magnitudes hold for commercial frontier models remains untested. Ai Safety And Ethics mixed Generalizability of option-order effects to commercial frontier models
Reading fidelity high
Study strength high
n=15
1.0

Notes