The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests About 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Aligning an AI to its operator’s wishes is not enough: without institutions that embody shared values, even 'perfectly aligned' systems can produce harmful societal outcomes. The authors propose 'full‑stack alignment'—thick, structured models of value embedded in institutions and systems—to enable normatively competent agents, stewardship, win‑win negotiation, meaning‑preserving economic mechanisms, and democratic regulation.

Full-Stack Alignment: Co-Aligning AI and Institutions with Thick Models of Value
Edelman, Joe, Zhi-Xuan, Tan, Lowe, Ryan, Klingefjord, Oliver, Wang-Mascianica, Vincent, Franklin, Matija, Kearns, Ryan Othniel, Hain, Ellie, Sarkar, Atrisha, Bakker, Michiel, Barez, Fazl, Duvenaud, David, Foerster, Jakob, Gabriel, Iason, Gubbels, Joseph, Goodman, Bryce, Haupt, Andreas, Heitzig, Jobst, Jara-Ettinger, Julian, Kasirzadeh, Atoosa, Kirkpatrick, James Ravi, Koh, Andrew, Knox, W. Bradley, Koralus, Philipp, Lehman, Joel, Levine, Sydney, Marro, Samuele, Revel, Manon, Shorin, Toby, Sutherland, Morgan, Tessler, Michael Henry, Vendrov, Ivan, Wilken-Smith, James · December 03, 2025 · arXiv (Cornell University)
openalex theoretical n/a evidence 7/10 relevance Full text usable extracted full text DOI Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

OpenAlex

Latest observation:

  1. Edelman, Joe provider ID
  2. Zhi-Xuan, Tan provider ID
  3. Lowe, Ryan provider ID
  4. Klingefjord, Oliver provider ID
  5. Wang-Mascianica, Vincent provider ID
  6. Franklin, Matija provider ID
  7. Kearns, Ryan Othniel unresolved corpus identity
  8. Hain, Ellie provider ID
  9. Sarkar, Atrisha unresolved corpus identity
  10. Bakker, Michiel provider ID
  11. Barez, Fazl provider ID
  12. Duvenaud, David provider ID
  13. Foerster, Jakob provider ID
  14. Gabriel, Iason provider ID
  15. Gubbels, Joseph provider ID
  16. Goodman, Bryce provider ID
  17. Haupt, Andreas provider ID
  18. Heitzig, Jobst provider ID
  19. Jara-Ettinger, Julian provider ID
  20. Kasirzadeh, Atoosa provider ID
  21. Kirkpatrick, James Ravi provider ID
  22. Koh, Andrew provider ID
  23. Knox, W. Bradley provider ID
  24. Koralus, Philipp provider ID
  25. Lehman, Joel provider ID
  26. Levine, Sydney provider ID
  27. Marro, Samuele provider ID
  28. Revel, Manon provider ID
  29. Shorin, Toby provider ID
  30. Sutherland, Morgan unresolved corpus identity
  31. Tessler, Michael Henry provider ID
  32. Vendrov, Ivan provider ID
  33. Wilken-Smith, James provider ID

Semantic Scholar

Latest observation:

  1. Joe Edelman provider ID
  2. Zhi-Xuan Tan provider ID
  3. Ryan Lowe provider ID
  4. Oliver Klingefjord provider ID
  5. Vincent Wang-Mascianica provider ID
  6. Matija Franklin provider ID
  7. R. Kearns provider ID
  8. Ellie Hain provider ID
  9. Atrisha Sarkar provider ID
  10. Michiel A. Bakker provider ID
  11. Fazl Barez provider ID
  12. D. Duvenaud provider ID
  13. Jakob N. Foerster provider ID
  14. Iason Gabriel provider ID
  15. J. Gubbels provider ID
  16. B. Goodman provider ID
  17. Andreas Haupt provider ID
  18. J. Heitzig provider ID
  19. J. Jara-Ettinger provider ID
  20. Atoosa Kasirzadeh provider ID
  21. J. Kirkpatrick provider ID
  22. Andrew Koh provider ID
  23. W. B. Knox provider ID
  24. Philipp Koralus provider ID
  25. Joel Lehman provider ID
  26. Sydney Levine provider ID
  27. Samuele G. Marro provider ID
  28. Manon Revel provider ID
  29. Toby Shorin provider ID
  30. M. Sutherland provider ID
  31. Michael Henry Tessler provider ID
  32. Ivan Vendrov provider ID
  33. James Wilken-Smith provider ID
The paper argues that aligning individual AI systems to operator intentions is insufficient and proposes 'full-stack alignment'—structured 'thick' models of value plus institutional design—to ensure AI advances societally beneficial outcomes.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Beneficial societal outcomes cannot be guaranteed by aligning individual AI systems with the intentions of their operators or users. Even an AI system that is perfectly aligned to the intentions of its operating organization can lead to bad outcomes if the goals of that organization are misaligned with those of other institutions and individuals. For this reason, we need full-stack alignment, the concurrent alignment of AI systems and the institutions that shape them with what people value. This can be done without imposing a particular vision of individual or collective flourishing. We argue that current approaches for representing values, such as utility functions, preference orderings, or unstructured text, struggle to address these and other issues effectively. They struggle to distinguish values from other signals, to support principled normative reasoning, and to model collective goods. We propose thick models of value will be needed. These structure the way values and norms are represented, enabling systems to distinguish enduring values from fleeting preferences, to model the social embedding of individual choices, and to reason normatively, applying values in new domains. We demonstrate this approach in five areas: AI value stewardship, normatively competent agents, win-win negotiation systems, meaning-preserving economic mechanisms, and democratic regulatory institutions.

Summary

Main Finding

The paper argues that aligning AI with individual operators or users is insufficient for beneficial societal outcomes; instead we need "full-stack alignment" (FSA): the concurrent alignment of AI systems and the institutions that shape them with what people value. Current representational toolkits—preferentist models of value (PMV) and values-as-text (VAT)—are inadequate because they conflate values with fleeting preferences, lack structure for normative reasoning, and fail to model collective goods. The authors propose Thick Models of Value (TMV): structured, philosophically grounded representations of values and norms that preserve justificatory structure, social embedding, and generalization, and that can be operationalized across AI systems and institutions.

Key Points

  • Problem diagnosis
    • AI systems do not operate in isolation; institutional incentives (platforms, markets, states) amplify misalignments.
    • PMV (utility functions, preference orderings) bundles together durable values, momentary tastes, manipulation, addiction, and social pressures without mechanisms to distinguish them.
    • VAT (unstructured natural language specifications) lacks commitment to what values are and often fails to support reliable normative reasoning or robust generalization.
  • Desiderata for representations that enable FSA
  • Robustness to distortions: distinguish enduring, endorsed values from fleeting or manipulated preferences and detect illegitimate preference change.
  • Better treatment of collective values and norms: represent shared norms, roles, and public goods so institutions can coordinate and provide them.
  • Better generalization: capture justificatory structure so values can be applied in novel contexts and renegotiation can be faster and traceable.
  • The TMV proposal
    • TMV are structured representations (neither thin preference lists nor free-form text) that encode things like justificatory relations between values, social roles, normativity, and endorsed value-change processes.
    • TMV are pluralistic (do not prescribe a single vision of the good) but constrain what counts as a legitimate value concept and how value-change is evaluated.
    • TMV draws on philosophy (thick evaluative concepts, moral epistemology), formal models (contractualist/role-based models, mechanism design, causal representations), and computational tools (ontologies, structured representations, RL with normative feedback).
  • Exemplary application areas (paper develops five domains)
  • AI value stewardship (organizational processes and accountability that preserve endorsed values across deployment).
  • Normatively competent agents (agents that can reason with justificatory structures, deliberate, and request clarification).
  • Win–win negotiation systems (systems that seek Pareto-improving, norm-sensitive agreements rather than one-shot preference aggregation).
  • Meaning‑preserving economic mechanisms (market and mechanism designs that explicitly price or preserve value-laden public goods and social meanings).
  • Democratic regulatory institutions (faster, deliberative feedback mechanisms and institutional designs that embed TMV into governance).
  • Early work and proofs of concept exist, but a coordinated, interdisciplinary research program is required to scale TMV across the full sociotechnical stack.

Data & Methods

  • Nature of the work: primarily conceptual and interdisciplinary synthesis combining
    • normative philosophy (thick evaluative concepts, moral epistemology, contractarianism),
    • formal economic and game-theoretic tools (mechanism design, market design, social choice),
    • computational alignment methods (RLHF, constitution-based LLM techniques, structured ontologies),
    • social-science perspectives (norms, roles, institutions).
  • Methods used in the paper:
    • Analytical critique of PMV and VAT, highlighting formal and empirical failure modes (e.g., revealed-preference pathologies, capture by engagement metrics).
    • Proposal of desiderata and architectural principles for TMV grounded in philosophical theory.
    • Mapping of TMV concepts to concrete technical and institutional levers (e.g., norm-aware mechanisms, contract-like specifications, deliberative procedures).
    • Illustrative examples and early proofs-of-concept (drawn from related work such as constitutional AI, meaning-preserving intermediaries, and domain-specific mechanism design like kidney exchanges).
  • Empirical content: limited; the paper is largely a theoretical and design roadmap rather than an experimental evaluation. It points to emerging implementations and suggests empirical research directions (benchmarks, field experiments, institutional pilots).

Implications for AI Economics

  • Rethinking welfare measurement and market design
    • Replace narrow revealed-preference or engagement-based objectives with TMV-aware metrics that preserve social meanings and public goods (e.g., trust, community cohesion).
    • Design market mechanisms and pricing schemes that internalize value-laden externalities (not just transaction efficiency), enabling markets that can allocate scarce resources while preserving meaningful social outcomes.
  • Mechanism design and incentives
    • Build mechanisms that explicitly represent normative constraints, role-based entitlements, and endorsed value-change processes—reducing exploitative arbitrage of regulatory or normative gaps (e.g., trading bots that game spirit-of-law).
    • Create markets for "meaning-preserving intermediaries" (platform services that mediate transactions while preserving or translating thick value signals upward in institutional stacks).
  • Platform and competition regulation
    • Regulatory frameworks should require or incentivize TMV-based auditing and stewardship (platforms must demonstrate how design choices preserve endorsed values across user communities).
    • Antitrust and competition policy may need to account for platform impacts on plural values (not only prices and outputs but civic and social goods).
  • Faster institutional feedback and governance
    • TMV facilitates quicker, traceable renegotiation of norms and rules because justification structures allow generalization to novel contexts—this reduces lag between technological change and appropriate institutional response.
    • Democratic institutions can use TMV-informed deliberative procedures and computational tools to aggregate reasons (not only preferences), improving legitimacy and robustness of collective choices.
  • Labor, firm strategy, and distributional effects
    • Firms that integrate TMV into product and incentive design may avoid negative externalities (e.g., addictive engagement) and capture market segments that value meaning-preserving interactions.
    • There are distributional consequences: pricing and allocation mechanisms that preserve public goods may alter surplus distribution—requiring careful policy design and empirical testing.
  • Research & policy agenda for AI economics
    • Develop operational TMV representations and measurement tools (benchmarks, proxies for thick concepts).
    • Field experiments and pilot mechanism deployments to quantify tradeoffs (efficiency vs. value-preservation), detect strategic responses, and refine institutional incentives.
    • Interdisciplinary collaboration (economists, philosophers, computer scientists, political scientists) to integrate normative structure into tractable economic models and deployable mechanisms.

Overall, the paper calls for a shift in AI economics from optimizing thin preference-based objectives toward building instruments and institutions that encode and protect richer, socially embedded values—so markets and AI together can produce outcomes aligned with plural, durable human goods.

Assessment

Paper Typetheoretical Evidence Strengthn/a — The paper is conceptual and argumentative rather than empirical; it does not present causal identification or empirical evidence that could be graded as high/medium/low. Methods Rigormedium — The authors provide a structured critique of existing value-representation approaches and propose a coherent conceptual framework ('thick models of value') with illustrative applications across five domains, but they do not formalize the models mathematically nor provide empirical or experimental validation of the proposals. SampleNot applicable — the paper is normative/theoretical and draws on conceptual analysis, literature critique, and illustrative thought experiments across domains (AI stewardship, agents, negotiation, economic mechanisms, regulation) rather than observational or experimental data. Themesgovernance org_design human_ai_collab innovation GeneralizabilityNo empirical testing — proposed models are unvalidated in real-world settings or cultures, Operationalization risk — unclear how 'thick' value models would be encoded, measured, or implemented across diverse institutions, Political and institutional constraints — applicability depends on legal, regulatory, and organizational contexts that vary across countries, Scale and complexity — challenges may arise when scaling from illustrative examples to large, heterogeneous socio-technical systems, Value pluralism — proposals may struggle to reconcile deeply conflicting normative frameworks across populations

Claims (15)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Beneficial societal outcomes cannot be guaranteed by aligning individual AI systems with the intentions of their operators or users. Governance And Regulation negative beneficial societal outcomes (ability to guarantee positive societal outcomes via system-level alignment)
Reading fidelity high
Study strength speculative
not reported
0.02
Even an AI system that is perfectly aligned to the intentions of its operating organization can lead to bad outcomes if the goals of that organization are misaligned with those of other institutions and individuals. Governance And Regulation negative likelihood of bad outcomes from organizationally-aligned AI systems
Reading fidelity high
Study strength speculative
not reported
0.02
We need full-stack alignment: the concurrent alignment of AI systems and the institutions that shape them with what people value. Governance And Regulation positive alignment of AI systems and institutions with people's values
Reading fidelity high
Study strength speculative
not reported
0.02
Full-stack alignment can be done without imposing a particular vision of individual or collective flourishing. Governance And Regulation positive feasibility of pluralistic full-stack alignment (ability to align without imposing a single vision of flourishing)
Reading fidelity medium
Study strength speculative
not reported
0.01
Current approaches for representing values, such as utility functions, preference orderings, or unstructured text, struggle to distinguish values from other signals, to support principled normative reasoning, and to model collective goods. Ai Safety And Ethics negative effectiveness/adequacy of current value-representation approaches (utility functions, preference orderings, unstructured text)
Reading fidelity high
Study strength medium
not reported
0.12
Thick models of value will be needed. Ai Safety And Ethics positive adequacy of value representation approaches (necessity of thick models)
Reading fidelity high
Study strength speculative
not reported
0.02
Thick models structure the way values and norms are represented, enabling systems to distinguish enduring values from fleeting preferences. Ai Safety And Ethics positive ability to distinguish enduring values from fleeting preferences
Reading fidelity high
Study strength speculative
not reported
0.02
Thick models enable systems to model the social embedding of individual choices. Ai Safety And Ethics positive ability to represent or model social embedding of individual choices
Reading fidelity high
Study strength speculative
not reported
0.02
Thick models allow systems to reason normatively and to apply values in new domains. Ai Safety And Ethics positive normative reasoning ability and cross-domain application of values
Reading fidelity high
Study strength speculative
not reported
0.02
We demonstrate this approach in five areas: AI value stewardship, normatively competent agents, win-win negotiation systems, meaning-preserving economic mechanisms, and democratic regulatory institutions. Governance And Regulation positive applicability of the thick-model approach across five named domains
Reading fidelity high
Study strength low
not reported
0.06
The paper's approach can support AI value stewardship (i.e., institutional practices for managing AI-aligned values). Governance And Regulation positive feasibility of AI value stewardship using the proposed approach
Reading fidelity high
Study strength low
not reported
0.06
The paper's approach can produce normatively competent agents (agents that can reason using normative principles). Ai Safety And Ethics positive normative competence of agents (ability to reason with norms)
Reading fidelity high
Study strength low
not reported
0.06
The approach can enable win-win negotiation systems. Decision Quality positive ability of negotiation systems to achieve mutually beneficial (win-win) outcomes
Reading fidelity high
Study strength low
not reported
0.06
The approach can produce meaning-preserving economic mechanisms. Market Structure positive ability of economic mechanisms to preserve meaning/values across transactions and allocations
Reading fidelity high
Study strength low
not reported
0.06
The approach can inform the design of democratic regulatory institutions. Governance And Regulation positive design and function of democratic regulatory institutions informed by the proposed thick-value approach
Reading fidelity high
Study strength low
not reported
0.06

Notes