0 cumulative citations
View corpus contextAligning an AI to its operator’s wishes is not enough: without institutions that embody shared values, even 'perfectly aligned' systems can produce harmful societal outcomes. The authors propose 'full‑stack alignment'—thick, structured models of value embedded in institutions and systems—to enable normatively competent agents, stewardship, win‑win negotiation, meaning‑preserving economic mechanisms, and democratic regulation.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
10 cumulative citations
View corpus contextBeneficial societal outcomes cannot be guaranteed by aligning individual AI systems with the intentions of their operators or users. Even an AI system that is perfectly aligned to the intentions of its operating organization can lead to bad outcomes if the goals of that organization are misaligned with those of other institutions and individuals. For this reason, we need full-stack alignment, the concurrent alignment of AI systems and the institutions that shape them with what people value. This can be done without imposing a particular vision of individual or collective flourishing. We argue that current approaches for representing values, such as utility functions, preference orderings, or unstructured text, struggle to address these and other issues effectively. They struggle to distinguish values from other signals, to support principled normative reasoning, and to model collective goods. We propose thick models of value will be needed. These structure the way values and norms are represented, enabling systems to distinguish enduring values from fleeting preferences, to model the social embedding of individual choices, and to reason normatively, applying values in new domains. We demonstrate this approach in five areas: AI value stewardship, normatively competent agents, win-win negotiation systems, meaning-preserving economic mechanisms, and democratic regulatory institutions.
Summary
Main Finding
The paper argues that aligning AI with individual operators or users is insufficient for beneficial societal outcomes; instead we need "full-stack alignment" (FSA): the concurrent alignment of AI systems and the institutions that shape them with what people value. Current representational toolkits—preferentist models of value (PMV) and values-as-text (VAT)—are inadequate because they conflate values with fleeting preferences, lack structure for normative reasoning, and fail to model collective goods. The authors propose Thick Models of Value (TMV): structured, philosophically grounded representations of values and norms that preserve justificatory structure, social embedding, and generalization, and that can be operationalized across AI systems and institutions.
Key Points
- Problem diagnosis
- AI systems do not operate in isolation; institutional incentives (platforms, markets, states) amplify misalignments.
- PMV (utility functions, preference orderings) bundles together durable values, momentary tastes, manipulation, addiction, and social pressures without mechanisms to distinguish them.
- VAT (unstructured natural language specifications) lacks commitment to what values are and often fails to support reliable normative reasoning or robust generalization.
- Desiderata for representations that enable FSA
- Robustness to distortions: distinguish enduring, endorsed values from fleeting or manipulated preferences and detect illegitimate preference change.
- Better treatment of collective values and norms: represent shared norms, roles, and public goods so institutions can coordinate and provide them.
- Better generalization: capture justificatory structure so values can be applied in novel contexts and renegotiation can be faster and traceable.
- The TMV proposal
- TMV are structured representations (neither thin preference lists nor free-form text) that encode things like justificatory relations between values, social roles, normativity, and endorsed value-change processes.
- TMV are pluralistic (do not prescribe a single vision of the good) but constrain what counts as a legitimate value concept and how value-change is evaluated.
- TMV draws on philosophy (thick evaluative concepts, moral epistemology), formal models (contractualist/role-based models, mechanism design, causal representations), and computational tools (ontologies, structured representations, RL with normative feedback).
- Exemplary application areas (paper develops five domains)
- AI value stewardship (organizational processes and accountability that preserve endorsed values across deployment).
- Normatively competent agents (agents that can reason with justificatory structures, deliberate, and request clarification).
- Win–win negotiation systems (systems that seek Pareto-improving, norm-sensitive agreements rather than one-shot preference aggregation).
- Meaning‑preserving economic mechanisms (market and mechanism designs that explicitly price or preserve value-laden public goods and social meanings).
- Democratic regulatory institutions (faster, deliberative feedback mechanisms and institutional designs that embed TMV into governance).
- Early work and proofs of concept exist, but a coordinated, interdisciplinary research program is required to scale TMV across the full sociotechnical stack.
Data & Methods
- Nature of the work: primarily conceptual and interdisciplinary synthesis combining
- normative philosophy (thick evaluative concepts, moral epistemology, contractarianism),
- formal economic and game-theoretic tools (mechanism design, market design, social choice),
- computational alignment methods (RLHF, constitution-based LLM techniques, structured ontologies),
- social-science perspectives (norms, roles, institutions).
- Methods used in the paper:
- Analytical critique of PMV and VAT, highlighting formal and empirical failure modes (e.g., revealed-preference pathologies, capture by engagement metrics).
- Proposal of desiderata and architectural principles for TMV grounded in philosophical theory.
- Mapping of TMV concepts to concrete technical and institutional levers (e.g., norm-aware mechanisms, contract-like specifications, deliberative procedures).
- Illustrative examples and early proofs-of-concept (drawn from related work such as constitutional AI, meaning-preserving intermediaries, and domain-specific mechanism design like kidney exchanges).
- Empirical content: limited; the paper is largely a theoretical and design roadmap rather than an experimental evaluation. It points to emerging implementations and suggests empirical research directions (benchmarks, field experiments, institutional pilots).
Implications for AI Economics
- Rethinking welfare measurement and market design
- Replace narrow revealed-preference or engagement-based objectives with TMV-aware metrics that preserve social meanings and public goods (e.g., trust, community cohesion).
- Design market mechanisms and pricing schemes that internalize value-laden externalities (not just transaction efficiency), enabling markets that can allocate scarce resources while preserving meaningful social outcomes.
- Mechanism design and incentives
- Build mechanisms that explicitly represent normative constraints, role-based entitlements, and endorsed value-change processes—reducing exploitative arbitrage of regulatory or normative gaps (e.g., trading bots that game spirit-of-law).
- Create markets for "meaning-preserving intermediaries" (platform services that mediate transactions while preserving or translating thick value signals upward in institutional stacks).
- Platform and competition regulation
- Regulatory frameworks should require or incentivize TMV-based auditing and stewardship (platforms must demonstrate how design choices preserve endorsed values across user communities).
- Antitrust and competition policy may need to account for platform impacts on plural values (not only prices and outputs but civic and social goods).
- Faster institutional feedback and governance
- TMV facilitates quicker, traceable renegotiation of norms and rules because justification structures allow generalization to novel contexts—this reduces lag between technological change and appropriate institutional response.
- Democratic institutions can use TMV-informed deliberative procedures and computational tools to aggregate reasons (not only preferences), improving legitimacy and robustness of collective choices.
- Labor, firm strategy, and distributional effects
- Firms that integrate TMV into product and incentive design may avoid negative externalities (e.g., addictive engagement) and capture market segments that value meaning-preserving interactions.
- There are distributional consequences: pricing and allocation mechanisms that preserve public goods may alter surplus distribution—requiring careful policy design and empirical testing.
- Research & policy agenda for AI economics
- Develop operational TMV representations and measurement tools (benchmarks, proxies for thick concepts).
- Field experiments and pilot mechanism deployments to quantify tradeoffs (efficiency vs. value-preservation), detect strategic responses, and refine institutional incentives.
- Interdisciplinary collaboration (economists, philosophers, computer scientists, political scientists) to integrate normative structure into tractable economic models and deployable mechanisms.
Overall, the paper calls for a shift in AI economics from optimizing thin preference-based objectives toward building instruments and institutions that encode and protect richer, socially embedded values—so markets and AI together can produce outcomes aligned with plural, durable human goods.
Assessment
Claims (15)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Beneficial societal outcomes cannot be guaranteed by aligning individual AI systems with the intentions of their operators or users. Governance And Regulation | negative | beneficial societal outcomes (ability to guarantee positive societal outcomes via system-level alignment) |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| Even an AI system that is perfectly aligned to the intentions of its operating organization can lead to bad outcomes if the goals of that organization are misaligned with those of other institutions and individuals. Governance And Regulation | negative | likelihood of bad outcomes from organizationally-aligned AI systems |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| We need full-stack alignment: the concurrent alignment of AI systems and the institutions that shape them with what people value. Governance And Regulation | positive | alignment of AI systems and institutions with people's values |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| Full-stack alignment can be done without imposing a particular vision of individual or collective flourishing. Governance And Regulation | positive | feasibility of pluralistic full-stack alignment (ability to align without imposing a single vision of flourishing) |
Reading fidelity
medium
Study strength
speculative
|
not reported
|
| Current approaches for representing values, such as utility functions, preference orderings, or unstructured text, struggle to distinguish values from other signals, to support principled normative reasoning, and to model collective goods. Ai Safety And Ethics | negative | effectiveness/adequacy of current value-representation approaches (utility functions, preference orderings, unstructured text) |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Thick models of value will be needed. Ai Safety And Ethics | positive | adequacy of value representation approaches (necessity of thick models) |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| Thick models structure the way values and norms are represented, enabling systems to distinguish enduring values from fleeting preferences. Ai Safety And Ethics | positive | ability to distinguish enduring values from fleeting preferences |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| Thick models enable systems to model the social embedding of individual choices. Ai Safety And Ethics | positive | ability to represent or model social embedding of individual choices |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| Thick models allow systems to reason normatively and to apply values in new domains. Ai Safety And Ethics | positive | normative reasoning ability and cross-domain application of values |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| We demonstrate this approach in five areas: AI value stewardship, normatively competent agents, win-win negotiation systems, meaning-preserving economic mechanisms, and democratic regulatory institutions. Governance And Regulation | positive | applicability of the thick-model approach across five named domains |
Reading fidelity
high
Study strength
low
|
not reported
|
| The paper's approach can support AI value stewardship (i.e., institutional practices for managing AI-aligned values). Governance And Regulation | positive | feasibility of AI value stewardship using the proposed approach |
Reading fidelity
high
Study strength
low
|
not reported
|
| The paper's approach can produce normatively competent agents (agents that can reason using normative principles). Ai Safety And Ethics | positive | normative competence of agents (ability to reason with norms) |
Reading fidelity
high
Study strength
low
|
not reported
|
| The approach can enable win-win negotiation systems. Decision Quality | positive | ability of negotiation systems to achieve mutually beneficial (win-win) outcomes |
Reading fidelity
high
Study strength
low
|
not reported
|
| The approach can produce meaning-preserving economic mechanisms. Market Structure | positive | ability of economic mechanisms to preserve meaning/values across transactions and allocations |
Reading fidelity
high
Study strength
low
|
not reported
|
| The approach can inform the design of democratic regulatory institutions. Governance And Regulation | positive | design and function of democratic regulatory institutions informed by the proposed thick-value approach |
Reading fidelity
high
Study strength
low
|
not reported
|