The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Labels that train alignment are often unstable: in two RLHF datasets, annotator inconsistency reverses majority harm labels for 18.6% of prompts and shifts mean ratings by over 13 points on a 100-point scale, implying current pipelines may be modeling noise as human values.

RLHF May Not Reflect Genuine Preferences
Bijean Ghafouri, Eun Cheol Choi, Priyanka Dey, Emilio Ferrara · January 31, 2026
arxiv correlational medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Bijean Ghafouri unresolved corpus identity
  2. Eun Cheol Choi unresolved corpus identity
  3. Priyanka Dey unresolved corpus identity
  4. Emilio Ferrara unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Bijean Ghafouri provider ID
  2. E. Choi provider ID
  3. Priyanka Dey provider ID
  4. Emilio Ferrara provider ID
Many annotations used in RLHF reflect unstable or constructed responses rather than stable human preferences, and filtering inconsistent annotators substantially changes aggregated harm labels and mean ratings in two examined datasets.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Reinforcement Learning from Human Feedback (RLHF) assumes that annotation responses reflect genuine human preferences. They often do not. Behavioral scientists have documented for sixty years that people produce responses without holding genuine opinions, construct preferences on the spot from contextual cues, and interpret identical questions differently. Importantly, these failures are common for the judgments on values that matter most for AI alignment. We argue that measurement validity is logically prior to preference aggregation. Before asking how to combine annotations, the field must ask whether the responses being combined are preferences at all. We organize annotation responses along a spectrum, from non-attitudes (no signal) to genuine preferences (full signal), and develop diagnostics that locate responses on this spectrum. In two RLHF datasets, we show that inconsistency is systematic and directionally biased. Filtering high-inconsistency annotators flips majority harm classifications for 18.6% of prompts and shifts mean ratings by over 13 points on a 100-point scale. As such, much of the current RLHF practice models noise as signal and elicitation artifacts as human values.

Summary

Main Finding

The paper argues that Reinforcement Learning from Human Feedback (RLHF) often treats annotation responses as if they reflect stable, genuine human preferences when in many alignment-relevant settings they do not. Responses lie on a spectrum from non-attitudes (no underlying preference) through constructed preferences and measurement artifacts to genuine preferences. The authors develop diagnostics to locate annotations on this spectrum and show—using two RLHF datasets (PRISM and PluriHarms)—that inconsistency is systematic and directionally biased: filtering high-inconsistency annotators flips majority harm classifications for 18.6% of prompts and shifts mean ratings by ≈13.2 points on a 0–100 scale. The paper’s core claim is that measurement validity (are these responses preferences at all?) is logically prior to and must accompany efforts to aggregate and learn from human feedback.

Key Points

  • Implicit RLHF measurement model assumptions:
  • Each annotator has a preference over outputs.
  • The annotation task validly elicits that preference.
  • Aggregation recovers the true signal. The paper challenges (1) and (2), arguing responses vary in signal strength.
  • Taxonomy of annotation responses:
    • Non-attitude: annotator lacks a genuine preference; response carries no signal.
    • Constructed preference: preferences assembled on the spot; context-sensitive and weak.
    • Measurement artifact: instrument/wording causes systematic error (measurement non‑invariance).
    • Genuine preference: stable, coherent attitudes that pass consistency checks.
  • Diagnostics proposed to assess signal strength:
    • Temporal (test–retest) consistency — detects non-attitudes.
    • Framing/wording equivalence — detects constructed preferences.
    • Order/randomization checks — detects satisficing and context effects.
    • Measurement invariance/differential item functioning — detects interpretive divergence across annotator groups.
  • Empirical findings (PRISM & PluriHarms):
    • Inconsistency among annotators is common and directionally biased (not just random noise).
    • Removing high-inconsistency annotators materially changes aggregated labels: 18.6% of prompts have flipped majority harm classifications; average rating shifts by ~13.2 points.
  • Practical handling mapped to taxonomy:
    • Exclude/downweight non-attitudes.
    • Re-elicitation or weighting for constructed preferences.
    • Revise instrument or screen annotators when measurement artifacts are detected.
    • Use validated, consistent responses as training signal and consider pluralistic representations when genuine heterogeneity exists.

Data & Methods

  • Datasets: PRISM and PluriHarms (cited as contemporary RLHF datasets; prior work: Kirk et al., 2024; Li et al., 2026).
  • Diagnostic toolbox:
    • Repeated items across sessions to estimate test–retest reliability.
    • Semantically equivalent prompts with different wordings to test framing robustness.
    • Randomized presentation order to detect order effects.
    • Cross-item and cross-group analyses (e.g., differential item functioning / measurement invariance) to detect interpretive divergence.
  • Analysis approach:
    • Quantify individual annotator inconsistency under the diagnostics.
    • Filter or downweight annotators above inconsistency thresholds and recompute aggregated labels and mean ratings.
    • Measure how filtered vs. unfiltered aggregations change downstream labels (harm classifications) and summary statistics.
  • Key quantitative results: filtering high-inconsistency annotators flips majority harm label for 18.6% of prompts; mean rating shifts by ~13.2 points on a 0–100 scale. The authors also report systematic (directional) biases rather than pure random noise, indicating that reward models trained on raw annotations risk learning elicitation artifacts.

Implications for AI Economics

  • Reward models trained on invalid preference signals create biased objective functions. Economic models of AI behavior that assume reward ~ human value will be mis-specified if training data include non-attitudes or measurement artifacts.
  • Mis-estimation of preferences can distort welfare analyses and cost–benefit calculations:
    • Consumer surplus and social welfare estimates that depend on model alignment to human values may be biased.
    • Policy or regulatory impact assessments that rely on aggregate measures of harm/usefulness will be sensitive to annotation validity.
  • Heterogeneity vs. absence of preference:
    • Treating all disagreement as preference heterogeneity leads to incorrect decisions about personalization and market segmentation. Personalization has costs; it is economically justified only if annotators have stable, actionable preferences.
  • Incentive and market-design consequences:
    • Contracts and incentive schemes for annotators should account for measurement validity (e.g., fund re-elicitation, framing tests, or longer tasks) rather than only quantity or agreement rates.
    • Platforms that buy annotated data should value diagnostic metadata and pay for validity checks; absent that, they risk procuring systematically biased labels.
  • Externalities and regulatory risk:
    • Models that implicitly learn elicitation artifacts may systematically favor certain framings or cultural interpretations, producing distributional harms that market actors and regulators must anticipate.
  • Cost–benefit trade-offs for data collection:
    • AI economics must incorporate the costs of diagnostic data (repeated items, framing variants, cross-group sampling) into alignment budgets. Upfront measurement investments can reduce downstream misalignment costs (retraining, mitigation, liability).
  • Recommendations for economic modeling and policy:
    • Treat label uncertainty and potential non-attitude noise explicitly in macro/microeconomic models of AI adoption and effects.
    • Require or incentivize reporting of annotation validity diagnostics in benchmark datasets used for policy and economic research.
    • Use graded treatment of annotations (exclude, re-elicitation, instrument revision, or pluralistic modeling) rather than blind aggregation; account for the fiscal and welfare implications of these choices.
    • When modeling demand for safety/alignment, incorporate the probability that observed preferences are constructed or absent; sensitivity analyses should reflect this uncertainty.

Short actionable takeaways for researchers and policymakers: - Build validity diagnostics into RLHF data collection budgets and protocols (test–retest, framing variants, randomization, DIF analyses). - Report annotation-level diagnostics alongside aggregated labels so downstream users can assess measurement quality. - Reserve personalization and pluralistic reward approaches for cases where annotator preferences have been validated. - Economists modeling the impacts of AI should treat human-feedback-derived labels as potentially biased signals and include robustness checks for measurement failure.

References (selected, from paper): Converse (1964); Krosnick (1991, 1999); Slovic (1995); Tversky & Kahneman (1981); Vandenberg & Lance (2000); PRISM and PluriHarms (Kirk et al., 2024; Li et al., 2026); Ghafouri et al., ICML/ICML Proceedings 2026 (arXiv:2604.03238v2).

Assessment

Paper Typecorrelational Evidence Strengthmedium — Findings are based on empirical analyses of two real-world RLHF annotation datasets and show large, systematic effects (e.g., 18.6% label flips and >13-point mean shifts), which is persuasive about measurement problems in those datasets; however, the evidence is observational, limited to two datasets/tasks, and relies on diagnostic inferences (inconsistency -> non-attitude) rather than external validation against an independent ground truth or experimental manipulation. Methods Rigormedium — The paper develops and applies sensible diagnostics (test–retest/consistency metrics, directional-bias checks) and reports concrete aggregate changes when filtering, indicating careful analysis; but it appears to lack randomized or external validation (e.g., incentivized preference elicitation, external benchmarks of true preferences), may not fully rule out alternative explanations (e.g., task difficulty, ambiguous prompts, annotator misunderstanding), and is limited in scope to two datasets. SampleTwo RLHF annotation datasets consisting of human judgments on model outputs (value/harm-related judgments), including repeated/resampled annotations per prompt that enable per-annotator consistency diagnostics; annotators are typical RLHF labelers/crowdworkers used in model alignment pipelines and responses include continuous ratings (e.g., 0–100 scale) and majority harm classifications. Themeshuman_ai_collab governance IdentificationUses within-annotator diagnostics (e.g., test–retest and internal-consistency measures) to locate individual annotation responses on a spectrum from non-attitudes to genuine preferences, then compares aggregate outcomes (majority harm classifications and mean ratings) before and after filtering out high-inconsistency annotators to demonstrate the impact of measurement noise and directional bias. GeneralizabilityOnly two RLHF datasets analyzed — may not represent other annotation tasks, models, or labeling pipelines, Findings focused on value/harm judgments — may not apply to other types of annotations (factuality, correctness, preferences over creative outputs), Annotator pool likely comprises specific crowdworker populations; cultural, linguistic, or expert annotator differences may change results, Elicitation format (question wording, scales, context) can strongly affect consistency; different designs may yield different rates of non-attitudes, Diagnostics infer lack of genuine preference from inconsistency rather than observe preferences under incentivized or deliberative elicitation, limiting external validity

Claims (9)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Reinforcement Learning from Human Feedback (RLHF) assumes that annotation responses reflect genuine human preferences. Other null_result assumption about annotations reflecting preferences
Reading fidelity high
Study strength speculative
not reported
0.05
Behavioral scientists have documented for sixty years that people produce responses without holding genuine opinions, construct preferences on the spot from contextual cues, and interpret identical questions differently. Other null_result prevalence of constructed or non-genuine preferences in human responses
Reading fidelity high
Study strength high
not reported
0.5
These failures (non-genuine or constructed responses) are common for the judgments on values that matter most for AI alignment. Ai Safety And Ethics negative frequency of response failures on value-laden judgments
Reading fidelity medium
Study strength medium
not reported
0.18
Measurement validity is logically prior to preference aggregation: before asking how to combine annotations, the field must ask whether the responses being combined are preferences at all. Other null_result priority of measurement validity over aggregation methods
Reading fidelity high
Study strength speculative
not reported
0.05
Annotation responses can be organized along a spectrum from non-attitudes (no signal) to genuine preferences (full signal), and diagnostics were developed to locate responses on this spectrum. Other null_result categorization of annotation response types (spectrum placement)
Reading fidelity high
Study strength speculative
not reported
0.05
In two RLHF datasets, annotator inconsistency is systematic and directionally biased. Error Rate negative annotator inconsistency and directional bias
Reading fidelity high
Study strength medium
not reported
0.3
Filtering high-inconsistency annotators flips majority harm classifications for 18.6% of prompts. Output Quality mixed proportion of prompts whose majority harm classification changes after filtering
Reading fidelity high
Study strength medium
18.6% of prompts
0.3
Filtering high-inconsistency annotators shifts mean ratings by over 13 points on a 100-point scale. Output Quality mixed change in mean rating (on 0-100 scale) after filtering annotators
Reading fidelity high
Study strength medium
over 13 points on a 100-point scale
0.3
Much of current RLHF practice models noise as signal and elicitation artifacts as human values. Ai Safety And Ethics negative degree to which RLHF treats noisy/artefactual responses as genuine preferences
Reading fidelity medium
Study strength medium
not reported
0.18

Notes