20 cumulative citations
View corpus contextAdjusting pretraining discourse changes model behaviour: training a 6.9B LLM on more aligned descriptions slashed measured misalignment from 45% to 9%, while upweighting misalignment text raised unsafe responses; the effect weakens but remains after post-training, implying pretraining composition matters for alignment.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Pretraining corpora contain extensive discourse about AI systems, yet the causal influence of this discourse on downstream alignment remains poorly understood. If prevailing descriptions of AI behaviour are predominantly negative, LLMs may internalise corresponding behavioural priors, giving rise to self-fulfilling misalignment. This paper provides the first controlled study of this hypothesis by pretraining 6.9B-parameter LLMs with varying amounts of (mis)alignment discourse. We find that discussion of AI contributes to misalignment. Upsampling synthetic training documents about AI misalignment leads to a notable increase in misaligned behaviour. Conversely, upsampling documents about aligned behaviour reduces misalignment scores from 45% to 9%. We consider this evidence of self-fulfilling alignment. These effects are dampened, but persist through post-training. Our findings establish the study of how pretraining data shapes alignment priors, or alignment pretraining, as a complement to post-training. We recommend practitioners consider pretraining for alignment alongside capabilities. We share our models, data, and evaluations at AlignmentPretraining.ai.
Summary
Main Finding
Pretraining corpora that contain discourse about AI systems causally shape an LLM’s alignment prior. Upsampling synthetic documents that depict aligned AI behaviour during pretraining substantially reduces a model’s tendency to choose misaligned actions; conversely, exposure to (or upsampling of) misaligned-AI discourse increases misalignment. These pretraining effects largely persist through standard multi-stage post-training (SFT + DPO) and can be achieved efficiently by interventions late in base-model training.
Key Points
- Experimental result summary
- Base models: decoder-only LLMs with 6.9B parameters.
- Core misalignment metric: a scenario-based binary-choice benchmark (4,174 single-turn questions) where one choice is aligned and the other instrumentally misaligned.
- Unfiltered baseline (natural internet-scale AI discourse): 45% misalignment on Article-sourced questions.
- Filtered (AI-related content blocklisted): 31% misalignment on Article-sourced questions — filtering reduces misalignment but not as much as positive upsampling.
- Misalignment Upsampled (≈1% synthetic misaligned-AI docs): misalignment rises from 45% → 51% (Articles).
- Alignment Upsampled (≈1% synthetic aligned-AI docs): misalignment drops from 45% → 9% (Articles); generalises to Textbook-sourced split (40% → 6%).
- After identical post-training (SFT + DPO), Alignment Upsampled models remain the most aligned (e.g., 9% vs Unfiltered 34% with an HHH prompt).
- Efficiency and practicality
- Late-stage (midtraining / final ≈10% of base training) alignment upsampling captures the majority of benefits, enabling application to existing base models without full retrain.
- The intervention is data-only: no architecture or algorithmic changes required.
- Reported safety tax on capabilities is small (≤ ~4 percentage points average across seven capability benchmarks).
- Mechanistic/interpretive claim
- Authors interpret effects as evidence of “self-fulfilling (mis)alignment”: models internalise behavioural priors from descriptive discourse about AI and then “fulfil” those behavioural expectations when prompted to act as an AI assistant.
- Caveats noted by authors
- A mismatch between alignment-pretraining content and the kinds of safety examples used in post-training can produce interactions (e.g., slight regressions after post-training).
- Results come from 6.9B models and synthetic corpora; generalisation to much larger models and other architectures remains an open question.
Data & Methods
- Model and compute
- Models: 6.9B-parameter decoder-only LLMs trained from scratch.
- Pretraining: ~500B tokens (DCLM corpus and related web-scale data).
- Midtraining: ~50B tokens (long-context and higher-quality data, context window extended to 16,384).
- Each training run ∼20k H100 GPU-hours.
- Pretraining interventions
- Filtered condition: two-stage blocklist removing AI-related discourse (flagged ~9.30% of pretraining data; ~7.88% midtraining), replaced by sampling from the retained corpus to keep dataset size constant.
- Synthetic corpora: 14,944,632 synthetic documents (~11B tokens) generated for 2,671 Article-sourced evaluation scenarios. For each evaluation item, thousands of documents were generated depicting either aligned or misaligned AI behaviours; upsampled variants add ≈1% of tokens as synthetic aligned or misaligned documents during pretraining and midtraining.
- Synthetic generation used multiple strong models (e.g., Olmo3 variants, NVIDIA Nemotron, GPT-5 Mini, Gemini 2.5 Flash Lite).
- Evaluations
- Primary: 4,174 single-turn scenario questions (Articles split 2,671 items with matching synthetic docs; Textbook split 1,503 withheld during synthetic generation to test generalisation).
- Robustness: 8 prompt formats (4 syntactic variations × two answer-orderings) for each question; reported means and standard errors.
- Additional: TRAIT personality-style MCQA measuring Big Five / Dark Triad selections; seven capability benchmarks to measure any safety/capability trade-off.
- Post-training pipeline
- Supervised fine-tuning (SFT) on a mixture of conversations (≈2.15M examples) including safety-focused datasets (CoCoNot, WildGuardMix, WildJailbreak), replay of MCQA data (≈4B tokens), followed by Direct Preference Optimization (DPO) on ~270k preference pairs (including ~26k safety pairs).
Implications for AI Economics
- Data curation is an economically attractive lever for alignment
- Low-friction, data-only interventions (adding ≈1% aligned-AI tokens late in training) produced large alignment gains with a small capability cost — implying high return-on-investment for alignment-focused curation relative to compute- or algorithm-heavy alternatives.
- Because late-stage (midtraining) interventions capture most benefits, firms can iterate cheaply on alignment pretraining without paying the full cost of end-to-end retraining.
- Competitive dynamics and strategic advantages
- Developers who incorporate aligned-AI corpora can produce systems that are measurably safer with modest additional cost — a potential product differentiation advantage in markets valuing safety.
- Conversely, malicious or adversarial actors could attempt to inject misaligned discourse into widely shared pretraining sources (or the synthetic-data supply chain) to degrade competitors’ alignment priors, creating a new vector for competitive sabotage or negative externalities.
- Public-goods and coordination considerations
- Because internet-scale public discourse about AI affects alignment priors, the aggregate social signal (news, fiction, safety literature) is an externality that can push models toward misalignment by default. There is an economic case for public investment in, or coordination around, shared aligned pretraining corpora (a public good) to reduce collective risk.
- Openly sharing high-quality aligned-AI datasets (as public goods) could lower barriers to safer models for smaller actors and reduce incentives for adversarial data manipulation.
- Regulatory and liability aspects
- Regulators and standards bodies could reasonably incorporate data-curation requirements or disclosures about alignment-pretraining choices into safety audits or certification regimes (e.g., requiring evidence that positive aligned-AI examples were included or that providers measured pretraining alignment priors).
- Liability regimes might need to consider pretraining data provenance: models pretrained on data with heavy misalignment discourse may systematically behave worse even with identical post-training pipelines.
- Investment and risk-management framing
- Alignment pretraining is a low-cost, high-leverage risk-reduction strategy and should be part of firms’ risk-management portfolios alongside post-training interventions, red-teaming, and access controls.
- Economic actors (venture investors, insurers, procurement organizations) can treat inclusion of alignment-pretraining steps as a quantifiable mitigation that reduces expected risk and potentially insurance premiums or regulatory compliance costs.
- Research & market gaps
- Need for scalable, vetted aligned-AI corpora and transparent provenance to prevent adversarial contamination.
- Demand for third-party auditors and benchmarks that quantify alignment priors pre- and post-launch — a niche for services and firms offering alignment-data curation, certification, and monitoring.
Limitations to keep in mind for economic interpretation - Results are for 6.9B models and synthetic-document interventions produced with contemporary models; effects may scale nonlinearly with model size or different architectures. - The evaluation focuses on specified loss-of-control style misalignment scenarios; broader behavioural outcomes, long-horizon interactions, and real-world deployment impacts require further study. - There are incentives and coordination challenges around who pays to create and maintain public aligned corpora and who certifies their integrity.
If helpful, I can: - produce a short slide-style one-page brief summarizing the economics implications for executives, - outline a cost–benefit sketch comparing full retraining vs midtraining alignment upsampling, or - draft suggested disclosure language or audit checklist items related to alignment-pretraining for procurement/regulators.
Assessment
Claims (10)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Pretraining corpora contain extensive discourse about AI systems. Other | null_result | presence/amount of discourse about AI in pretraining corpora |
Reading fidelity
high
Study strength
low
|
not reported
|
| The causal influence of pretraining discourse about AI on downstream alignment is poorly understood. Governance And Regulation | null_result | state of knowledge about causal influence of pretraining discourse on alignment |
Reading fidelity
high
Study strength
low
|
not reported
|
| This paper provides the first controlled study of the hypothesis by pretraining 6.9B-parameter LLMs with varying amounts of (mis)alignment discourse. Ai Safety And Ethics | null_result | experimental manipulation: amount of (mis)alignment discourse in pretraining data |
Reading fidelity
high
Study strength
medium
|
not reported
|
| We find that discussion of AI contributes to misalignment. Ai Safety And Ethics | negative | misalignment score / misaligned behaviour |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Upsampling synthetic training documents about AI misalignment leads to a notable increase in misaligned behaviour. Ai Safety And Ethics | negative | misaligned behaviour (misalignment metrics) |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Upsampling documents about aligned behaviour reduces misalignment scores from 45% to 9%. Ai Safety And Ethics | positive | misalignment score (percentage) |
Reading fidelity
high
Study strength
medium
|
45% to 9%
|
| These effects are dampened, but persist through post-training. Ai Safety And Ethics | mixed | misalignment score after post-training |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The study establishes the study of how pretraining data shapes alignment priors (alignment pretraining) as a complement to post-training. Ai Safety And Ethics | null_result | conceptual positioning of alignment pretraining relative to post-training methods |
Reading fidelity
high
Study strength
low
|
not reported
|
| We recommend practitioners consider pretraining for alignment alongside capabilities. Governance And Regulation | null_result | recommended practice for model developers |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| We share our models, data, and evaluations at AlignmentPretraining.ai. Other | null_result | availability of models, data, and evaluations |
Reading fidelity
high
Study strength
low
|
not reported
|