0 cumulative citations
View corpus contextAn axiomatic objective would make AI protect and fairly allocate human capability: by maximizing a risk- and inequality-averse aggregate of humans' goal-attainment probabilities, an AI is incentivized to communicate, follow correction, avoid irreversible changes and limit disempowerment.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
This paper explores the idea of promoting well-being and safety in human-AI interactions by forcing AI agents explicitly to empower humans and to manage the power balance between humans and AI agents in a desirable way. Using a principled, partially axiomatic approach based on desirable properties, we design a parametrizable and decomposable objective function for AI systems that represents an inequality- and risk-averse long-term aggregate of human power. It can take into account models of human bounded rationality and social norms, and crucially, considers a wide variety of possible human goals. We prove how certain desiderata enforce particular functional forms and restrict parameter ranges. We exemplify the consequences of softly maximizing this metric in several paradigmatic situations and describe what instrumental sub-goals it will likely imply.
Summary
Main Finding
The paper derives an axiomatically motivated, decomposable objective for AI agents that (softly) maximizes a long-term, inequality- and uncertainty-averse aggregate of human “power” (capability to attain goals). The metric is goal-agnostic about humans’ actual preferences (it uses a space of possible goals plus a goal-conditioned human behavior model), and yields behavioral incentives for AI—e.g., corrigibility, norm-following, avoiding irreversible changes, communicating well, and protecting less-powerful humans—rather than relying on inferring individual human utilities.
Key Points
- Hierarchy of metrics (from micro to macro):
- Cgh h (s): goal-attainment capability for human h and goal gh in state s (a truncated Bellman-like discounted probability of reaching gh).
- Ih(s): individual human power = log2 ∑_{gh∈Gh} Cgh h (s)^ζ (bits of effective attainable goals; ζ>1 enforces reliability preference).
- P(s): present aggregate human power = −log2 ∑_{h∈H} 2^{-ξ Ih(s)} (ξ>0 enforces strong inequality aversion; very powerful individuals get negligible marginal weight).
- T(s≥t): trajectory-specific aggregate (temporal aggregation with intertemporal inequality aversion parameter η>0 and robot discount γr∈(0,1)).
- L(st): long-term aggregate across uncertainty = −log2 E_{s≥t} [ (∑_{u≥t} γr^{u−t} 2^{-η P(su)})^ρ ] (ρ>0 enforces uncertainty aversion).
- Closed-form/representative equations (as derived):
- Cgh h (s) = 1_{s∈gh} + 1_{s∉gh∪S⊤} γh E_{s'∼…}[ Cgh h (s') ]
- Ih(s) = log2 ∑_{gh∈Gh} (Cgh h (s))^ζ
- P(s) = −log2 ∑_{h∈H} 2^{−ξ Ih(s)}
- L(st) = −log2 E_{s≥t} [ (∑_{u≥t} γr^{u−t} 2^{−η P(su)})^ρ ]
- A soft-max / Boltzmann-like robot policy: πr(s)(ar) ∝ ( E_{s'∼s,ar,πH} 2^{−L(s')} )^{−βr} (robot tends to choose actions that increase long-term L; βr controls softness)
- Axiomatic derivation: each functional form follows from a set of desiderata (neutrality, anonymity, independence, delay/discount invariances, Pigou–Dalton–style inequality aversion, disempowerment focus, composition additivity under special cases). Propositions 1–6 formalize these results.
- Relationship to prior concepts:
- Generalizes/extends the “empowerment” (information‑theoretic channel-capacity) notion by grounding power in capability to attain many possible goals and by incorporating bounded rationality and social-norm effects.
- Connects to assistance games, impact measures, attainable-utility preservation and corrigibility, but produces them as emergent incentives from a single unified objective rather than as separate constraints.
- Behavioral consequences highlighted: incentives to communicate, accept/comply with corrective inputs, avoid irreversible large-impact actions, protect humans (especially less powerful ones), follow relevant social norms (since πH can encode norm-adapted human behavior), and allocate resources fairly across people and time.
Data & Methods
- Model setting:
- World model: finite, acyclic stochastic game form Γ with states S, terminal states S⊤, transition kernel Pr(s'|s,a). Humans H and robot r have action sets; no extrinsic rewards are assumed.
- Human behavior model: given goal-conditioned policy πH(s,g) (no requirement that the robot knows humans’ actual goals).
- Goal space Gh: each gh ⊆ S denotes the set of terminal states considered goal-attaining; coverage axiom requires for every s there exists a gh containing s.
- Methodology:
- Principled axiomatic approach: for each aggregation level (goal → individual → population → temporal → uncertainty) the authors specify plausible desiderata (continuity, monotonicity, neutrality/anonymity, various invariances, inequality/disempowerment focus) and derive functional forms consistent with those axioms.
- Mathematical results: Propositions 1–6 show that, under the axioms, the metric at each level must take the forms summarized above (up to scaling and normative parameters). Some choices (e.g., fG = id and FG = γh id) are made pragmatically to obtain operational Bellman-like equations.
- Normative parameters: γh, γr ∈(0,1) (discounting of future goal attainment), ζ>1 (reliability preference across goals), ξ>0 (inequality aversion across humans), η>0 (intertemporal inequality aversion), ρ>0 (uncertainty aversion), βr (policy softness).
- Limitations / modeling assumptions:
- Requires a goal-conditioned human behavior model πH and structural world model Γ (the robot does not infer actual current human goals).
- Assumes finite, acyclic games for analytic convenience.
- Several formal choices are normative (ξ, η, ρ, ζ), not uniquely determined by axioms; different choices change incentives.
- Empirical/implementation status:
- This is a theory paper; a companion paper is promised to study scalable estimation and empirical validation.
Implications for AI Economics
- Distributional and welfare implications:
- The objective explicitly encodes inequality aversion (ξ) and prioritizes protecting low-power humans. If adopted by deployed AI (e.g., platforms, automation systems), this can shift surplus from high-ability / high-leverage agents toward broader empowerment—potentially reducing extreme concentration of influence and economic rents tied to automation.
- The long-term and uncertainty aversion parameters (η, ρ) embed precautionary motives into AI decisions; this may reduce throughput-maximizing but risky strategies, affecting productivity-growth trade-offs.
- Firm incentives, competition, and strategic adoption:
- Firms that adopt empowerment-preserving AI objectives may face short-term competitiveness costs relative to firms optimizing conventional profit proxies, creating competitive dynamics and potential first-mover disadvantages unless regulated or market norms change.
- Regulation or standards could internalize the externality (preserve human agency as a public good): policymakers might require or certify such objective functions for deployed systems where human empowerment is socially valuable.
- Labor markets and innovation:
- By protecting human capability to act and to pursue diverse goals, these objectives can mitigate some displacement harms (workers retaining more real options). This could alter skill investment incentives and the shape of labor demand (favoring human-in-the-loop roles).
- Conversely, stronger constraints on irreversible or high-impact automation may slow certain types of technological deployment, shifting the timing and distribution of productivity gains.
- Mechanism design and markets:
- The P(s) aggregator (−log-sum-exp across individuals) resembles social-welfare weights that heavily prioritize the tail; integrating such AI objectives into market mechanisms (auctions, matching, recommender systems) will change allocation rules and may require redesign of incentives to preserve efficiency while upholding empowerment.
- Measurement, implementation, and compliance costs:
- Operationalizing the metric requires estimating Gh, πH, and transition models—data- and modeling-intensive tasks. There are measurement risks (mis-specified πH or Gh could produce perverse incentives).
- Firms may face costs to build or verify required world/human models; these are economic considerations for adoption and regulation.
- Strategic responses and gaming:
- Agents (humans, firms) might seek to change observable behavior or environments to appear more empowered under the model, potentially producing strategic manipulation unless the human-modeling component is robust.
- Research and policy directions:
- Empirical work needed: evaluate how objective choice (ξ, η, ρ, ζ) affects economic outcomes—inequality, growth, robustness—and find operational estimators for Gh and πH.
- Cost–benefit and general equilibrium analyses to assess long-run macroeconomic impacts of broad adoption.
- Designing regulatory instruments and market incentives to align firm behavior with empowerment-preserving objectives without stifling innovation.
Short concluding note: the paper provides a coherent, axiomatic formalism for embedding human-empowerment preservation in AI objectives. For AI economics, its main value is supplying a transparent, parameterized bridge between normative social-welfare concerns (inequality, precaution, intertemporal fairness) and concrete algorithmic incentives—enabling explicit analysis of the trade-offs and market consequences of committing AI systems to preserve human agency.
Assessment
Claims (12)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Under the stated axioms (C1–C6), the goal-attainment capability metric must have a recursive form based on a continuous, strictly increasing transformation of the expected successor capability. Other | positive | Functional form of goal-attainment capability |
Reading fidelity
high
Study strength
medium
|
not reported
|
| With the paper's pragmatic functional-form choice, individual goal-attainment capability is a discounted probability of attaining the goal. Decision Quality | positive | Probability-weighted goal attainment |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The individual human power metric derived from the axioms is the logarithm of the sum of goal-attainment capabilities raised to the reliability-preference parameter ζ. Other | positive | Individual human power across possible goals |
Reading fidelity
high
Study strength
medium
|
not reported
|
| If a human can choose between attaining k different goals with certainty, the proposed individual power metric equals log2 k. Skill Acquisition | positive | Number of reliably attainable goals |
Reading fidelity
high
Study strength
medium
|
log2 k bits
|
| Under axioms P0–P8 and either inequality-aversion condition P6 or limited-trade-off condition P7, present aggregate human power has the form P(s) = −log2 Σh∈H 2^(−ξIh(s)), where ξ > 0 controls inequality aversion. Inequality | positive | Inequality-averse aggregate human power across people |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The proposed population metric gives an additional all-powerful human no effect in the limit, implementing a focus on preventing disempowerment rather than rewarding already powerful individuals. Inequality | positive | Marginal contribution of highly powerful individuals to aggregate power |
Reading fidelity
high
Study strength
medium
|
not reported
|
| For sufficiently strong inequality aversion, reducing one person's power from one bit to zero cannot be offset by arbitrarily increasing the power of a bounded number of other people. Inequality | negative | Compensability of individual power losses |
Reading fidelity
high
Study strength
medium
|
k ≤ 2^ξ − 1
|
| The temporal aggregation axioms yield a discounted aggregation of present aggregate human power or a logarithmic inequality-averse alternative; the authors select the form that penalizes intertemporal power inequality. Social Protection | mixed | Longitudinal distribution of aggregate human power |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The uncertainty-averse long-term aggregate human-power metric is L(st) = −log2 E[(Σu≥t γr^(u−t) 2^(−ηP(su)))^ρ], with ρ > 0. Organizational Efficiency | positive | Expected long-term aggregate human power under uncertainty |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The authors argue that softly maximizing the proposed human-power metric will likely incentivize an AI system to communicate well, follow orders, remain corrigible, avoid irreversible environmental changes, protect humans and itself from harm and disempowerment, follow relevant social norms, and allocate resources fairly and sustainably. Ai Safety And Ethics | positive | AI behavioral incentives related to human empowerment and safety |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| The proposed approach is goal-agnostic in the sense that the AI system does not model or infer humans' actual current goals; instead, it uses a model of goal-conditioned human behavior and the environment's dynamics. Ai Safety And Ethics | positive | Dependence of AI policy on human goal information |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The paper does not report empirical validation or a human/AI experimental sample; it presents a theoretical framework and states that empirical validation is reserved for a companion paper. Other | null_result | Empirical validation of the proposed metric |
Reading fidelity
high
Study strength
high
|
not reported
|