1 cumulative citations
View corpus contextLabels that train alignment are often unstable: in two RLHF datasets, annotator inconsistency reverses majority harm labels for 18.6% of prompts and shifts mean ratings by over 13 points on a 100-point scale, implying current pipelines may be modeling noise as human values.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Reinforcement Learning from Human Feedback (RLHF) assumes that annotation responses reflect genuine human preferences. They often do not. Behavioral scientists have documented for sixty years that people produce responses without holding genuine opinions, construct preferences on the spot from contextual cues, and interpret identical questions differently. Importantly, these failures are common for the judgments on values that matter most for AI alignment. We argue that measurement validity is logically prior to preference aggregation. Before asking how to combine annotations, the field must ask whether the responses being combined are preferences at all. We organize annotation responses along a spectrum, from non-attitudes (no signal) to genuine preferences (full signal), and develop diagnostics that locate responses on this spectrum. In two RLHF datasets, we show that inconsistency is systematic and directionally biased. Filtering high-inconsistency annotators flips majority harm classifications for 18.6% of prompts and shifts mean ratings by over 13 points on a 100-point scale. As such, much of the current RLHF practice models noise as signal and elicitation artifacts as human values.
Summary
Main Finding
The paper argues that Reinforcement Learning from Human Feedback (RLHF) often treats annotation responses as if they reflect stable, genuine human preferences when in many alignment-relevant settings they do not. Responses lie on a spectrum from non-attitudes (no underlying preference) through constructed preferences and measurement artifacts to genuine preferences. The authors develop diagnostics to locate annotations on this spectrum and show—using two RLHF datasets (PRISM and PluriHarms)—that inconsistency is systematic and directionally biased: filtering high-inconsistency annotators flips majority harm classifications for 18.6% of prompts and shifts mean ratings by ≈13.2 points on a 0–100 scale. The paper’s core claim is that measurement validity (are these responses preferences at all?) is logically prior to and must accompany efforts to aggregate and learn from human feedback.
Key Points
- Implicit RLHF measurement model assumptions:
- Each annotator has a preference over outputs.
- The annotation task validly elicits that preference.
- Aggregation recovers the true signal. The paper challenges (1) and (2), arguing responses vary in signal strength.
- Taxonomy of annotation responses:
- Non-attitude: annotator lacks a genuine preference; response carries no signal.
- Constructed preference: preferences assembled on the spot; context-sensitive and weak.
- Measurement artifact: instrument/wording causes systematic error (measurement non‑invariance).
- Genuine preference: stable, coherent attitudes that pass consistency checks.
- Diagnostics proposed to assess signal strength:
- Temporal (test–retest) consistency — detects non-attitudes.
- Framing/wording equivalence — detects constructed preferences.
- Order/randomization checks — detects satisficing and context effects.
- Measurement invariance/differential item functioning — detects interpretive divergence across annotator groups.
- Empirical findings (PRISM & PluriHarms):
- Inconsistency among annotators is common and directionally biased (not just random noise).
- Removing high-inconsistency annotators materially changes aggregated labels: 18.6% of prompts have flipped majority harm classifications; average rating shifts by ~13.2 points.
- Practical handling mapped to taxonomy:
- Exclude/downweight non-attitudes.
- Re-elicitation or weighting for constructed preferences.
- Revise instrument or screen annotators when measurement artifacts are detected.
- Use validated, consistent responses as training signal and consider pluralistic representations when genuine heterogeneity exists.
Data & Methods
- Datasets: PRISM and PluriHarms (cited as contemporary RLHF datasets; prior work: Kirk et al., 2024; Li et al., 2026).
- Diagnostic toolbox:
- Repeated items across sessions to estimate test–retest reliability.
- Semantically equivalent prompts with different wordings to test framing robustness.
- Randomized presentation order to detect order effects.
- Cross-item and cross-group analyses (e.g., differential item functioning / measurement invariance) to detect interpretive divergence.
- Analysis approach:
- Quantify individual annotator inconsistency under the diagnostics.
- Filter or downweight annotators above inconsistency thresholds and recompute aggregated labels and mean ratings.
- Measure how filtered vs. unfiltered aggregations change downstream labels (harm classifications) and summary statistics.
- Key quantitative results: filtering high-inconsistency annotators flips majority harm label for 18.6% of prompts; mean rating shifts by ~13.2 points on a 0–100 scale. The authors also report systematic (directional) biases rather than pure random noise, indicating that reward models trained on raw annotations risk learning elicitation artifacts.
Implications for AI Economics
- Reward models trained on invalid preference signals create biased objective functions. Economic models of AI behavior that assume reward ~ human value will be mis-specified if training data include non-attitudes or measurement artifacts.
- Mis-estimation of preferences can distort welfare analyses and cost–benefit calculations:
- Consumer surplus and social welfare estimates that depend on model alignment to human values may be biased.
- Policy or regulatory impact assessments that rely on aggregate measures of harm/usefulness will be sensitive to annotation validity.
- Heterogeneity vs. absence of preference:
- Treating all disagreement as preference heterogeneity leads to incorrect decisions about personalization and market segmentation. Personalization has costs; it is economically justified only if annotators have stable, actionable preferences.
- Incentive and market-design consequences:
- Contracts and incentive schemes for annotators should account for measurement validity (e.g., fund re-elicitation, framing tests, or longer tasks) rather than only quantity or agreement rates.
- Platforms that buy annotated data should value diagnostic metadata and pay for validity checks; absent that, they risk procuring systematically biased labels.
- Externalities and regulatory risk:
- Models that implicitly learn elicitation artifacts may systematically favor certain framings or cultural interpretations, producing distributional harms that market actors and regulators must anticipate.
- Cost–benefit trade-offs for data collection:
- AI economics must incorporate the costs of diagnostic data (repeated items, framing variants, cross-group sampling) into alignment budgets. Upfront measurement investments can reduce downstream misalignment costs (retraining, mitigation, liability).
- Recommendations for economic modeling and policy:
- Treat label uncertainty and potential non-attitude noise explicitly in macro/microeconomic models of AI adoption and effects.
- Require or incentivize reporting of annotation validity diagnostics in benchmark datasets used for policy and economic research.
- Use graded treatment of annotations (exclude, re-elicitation, instrument revision, or pluralistic modeling) rather than blind aggregation; account for the fiscal and welfare implications of these choices.
- When modeling demand for safety/alignment, incorporate the probability that observed preferences are constructed or absent; sensitivity analyses should reflect this uncertainty.
Short actionable takeaways for researchers and policymakers: - Build validity diagnostics into RLHF data collection budgets and protocols (test–retest, framing variants, randomization, DIF analyses). - Report annotation-level diagnostics alongside aggregated labels so downstream users can assess measurement quality. - Reserve personalization and pluralistic reward approaches for cases where annotator preferences have been validated. - Economists modeling the impacts of AI should treat human-feedback-derived labels as potentially biased signals and include robustness checks for measurement failure.
References (selected, from paper): Converse (1964); Krosnick (1991, 1999); Slovic (1995); Tversky & Kahneman (1981); Vandenberg & Lance (2000); PRISM and PluriHarms (Kirk et al., 2024; Li et al., 2026); Ghafouri et al., ICML/ICML Proceedings 2026 (arXiv:2604.03238v2).
Assessment
Claims (9)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Reinforcement Learning from Human Feedback (RLHF) assumes that annotation responses reflect genuine human preferences. Other | null_result | assumption about annotations reflecting preferences |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| Behavioral scientists have documented for sixty years that people produce responses without holding genuine opinions, construct preferences on the spot from contextual cues, and interpret identical questions differently. Other | null_result | prevalence of constructed or non-genuine preferences in human responses |
Reading fidelity
high
Study strength
high
|
not reported
|
| These failures (non-genuine or constructed responses) are common for the judgments on values that matter most for AI alignment. Ai Safety And Ethics | negative | frequency of response failures on value-laden judgments |
Reading fidelity
medium
Study strength
medium
|
not reported
|
| Measurement validity is logically prior to preference aggregation: before asking how to combine annotations, the field must ask whether the responses being combined are preferences at all. Other | null_result | priority of measurement validity over aggregation methods |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| Annotation responses can be organized along a spectrum from non-attitudes (no signal) to genuine preferences (full signal), and diagnostics were developed to locate responses on this spectrum. Other | null_result | categorization of annotation response types (spectrum placement) |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| In two RLHF datasets, annotator inconsistency is systematic and directionally biased. Error Rate | negative | annotator inconsistency and directional bias |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Filtering high-inconsistency annotators flips majority harm classifications for 18.6% of prompts. Output Quality | mixed | proportion of prompts whose majority harm classification changes after filtering |
Reading fidelity
high
Study strength
medium
|
18.6% of prompts
|
| Filtering high-inconsistency annotators shifts mean ratings by over 13 points on a 100-point scale. Output Quality | mixed | change in mean rating (on 0-100 scale) after filtering annotators |
Reading fidelity
high
Study strength
medium
|
over 13 points on a 100-point scale
|
| Much of current RLHF practice models noise as signal and elicitation artifacts as human values. Ai Safety And Ethics | negative | degree to which RLHF treats noisy/artefactual responses as genuine preferences |
Reading fidelity
medium
Study strength
medium
|
not reported
|