The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

An axiomatic objective would make AI protect and fairly allocate human capability: by maximizing a risk- and inequality-averse aggregate of humans' goal-attainment probabilities, an AI is incentivized to communicate, follow correction, avoid irreversible changes and limit disempowerment.

A Fair Objective for Human-Empowerment-Preserving AI: Desiderata, Design, and Likely Behavioral Consequences
Jobst Heitzig, Ram Potham · August 08, 2026
arxiv theoretical n/a evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Jobst Heitzig unresolved corpus identity
  2. Ram Potham unresolved corpus identity

Semantic Scholar

Latest observation:

  1. J. Heitzig provider ID
  2. Ram Potham provider ID
The paper develops an axiomatic, goal-agnostic objective for AI that aggregates humans' goal-attainment capabilities into a risk- and inequality-averse long-term metric and shows how softly maximizing that metric yields incentives (e.g., communication, corrigibility, avoidance of irreversible harm) that preserve human empowerment.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

This paper explores the idea of promoting well-being and safety in human-AI interactions by forcing AI agents explicitly to empower humans and to manage the power balance between humans and AI agents in a desirable way. Using a principled, partially axiomatic approach based on desirable properties, we design a parametrizable and decomposable objective function for AI systems that represents an inequality- and risk-averse long-term aggregate of human power. It can take into account models of human bounded rationality and social norms, and crucially, considers a wide variety of possible human goals. We prove how certain desiderata enforce particular functional forms and restrict parameter ranges. We exemplify the consequences of softly maximizing this metric in several paradigmatic situations and describe what instrumental sub-goals it will likely imply.

Summary

Main Finding

The paper derives an axiomatically motivated, decomposable objective for AI agents that (softly) maximizes a long-term, inequality- and uncertainty-averse aggregate of human “power” (capability to attain goals). The metric is goal-agnostic about humans’ actual preferences (it uses a space of possible goals plus a goal-conditioned human behavior model), and yields behavioral incentives for AI—e.g., corrigibility, norm-following, avoiding irreversible changes, communicating well, and protecting less-powerful humans—rather than relying on inferring individual human utilities.

Key Points

  • Hierarchy of metrics (from micro to macro):
    • Cgh h (s): goal-attainment capability for human h and goal gh in state s (a truncated Bellman-like discounted probability of reaching gh).
    • Ih(s): individual human power = log2 ∑_{gh∈Gh} Cgh h (s)^ζ (bits of effective attainable goals; ζ>1 enforces reliability preference).
    • P(s): present aggregate human power = −log2 ∑_{h∈H} 2^{-ξ Ih(s)} (ξ>0 enforces strong inequality aversion; very powerful individuals get negligible marginal weight).
    • T(s≥t): trajectory-specific aggregate (temporal aggregation with intertemporal inequality aversion parameter η>0 and robot discount γr∈(0,1)).
    • L(st): long-term aggregate across uncertainty = −log2 E_{s≥t} [ (∑_{u≥t} γr^{u−t} 2^{-η P(su)})^ρ ] (ρ>0 enforces uncertainty aversion).
  • Closed-form/representative equations (as derived):
    • Cgh h (s) = 1_{s∈gh} + 1_{s∉gh∪S⊤} γh E_{s'∼…}[ Cgh h (s') ]
    • Ih(s) = log2 ∑_{gh∈Gh} (Cgh h (s))^ζ
    • P(s) = −log2 ∑_{h∈H} 2^{−ξ Ih(s)}
    • L(st) = −log2 E_{s≥t} [ (∑_{u≥t} γr^{u−t} 2^{−η P(su)})^ρ ]
    • A soft-max / Boltzmann-like robot policy: πr(s)(ar) ∝ ( E_{s'∼s,ar,πH} 2^{−L(s')} )^{−βr} (robot tends to choose actions that increase long-term L; βr controls softness)
  • Axiomatic derivation: each functional form follows from a set of desiderata (neutrality, anonymity, independence, delay/discount invariances, Pigou–Dalton–style inequality aversion, disempowerment focus, composition additivity under special cases). Propositions 1–6 formalize these results.
  • Relationship to prior concepts:
    • Generalizes/extends the “empowerment” (information‑theoretic channel-capacity) notion by grounding power in capability to attain many possible goals and by incorporating bounded rationality and social-norm effects.
    • Connects to assistance games, impact measures, attainable-utility preservation and corrigibility, but produces them as emergent incentives from a single unified objective rather than as separate constraints.
  • Behavioral consequences highlighted: incentives to communicate, accept/comply with corrective inputs, avoid irreversible large-impact actions, protect humans (especially less powerful ones), follow relevant social norms (since πH can encode norm-adapted human behavior), and allocate resources fairly across people and time.

Data & Methods

  • Model setting:
    • World model: finite, acyclic stochastic game form Γ with states S, terminal states S⊤, transition kernel Pr(s'|s,a). Humans H and robot r have action sets; no extrinsic rewards are assumed.
    • Human behavior model: given goal-conditioned policy πH(s,g) (no requirement that the robot knows humans’ actual goals).
    • Goal space Gh: each gh ⊆ S denotes the set of terminal states considered goal-attaining; coverage axiom requires for every s there exists a gh containing s.
  • Methodology:
    • Principled axiomatic approach: for each aggregation level (goal → individual → population → temporal → uncertainty) the authors specify plausible desiderata (continuity, monotonicity, neutrality/anonymity, various invariances, inequality/disempowerment focus) and derive functional forms consistent with those axioms.
    • Mathematical results: Propositions 1–6 show that, under the axioms, the metric at each level must take the forms summarized above (up to scaling and normative parameters). Some choices (e.g., fG = id and FG = γh id) are made pragmatically to obtain operational Bellman-like equations.
    • Normative parameters: γh, γr ∈(0,1) (discounting of future goal attainment), ζ>1 (reliability preference across goals), ξ>0 (inequality aversion across humans), η>0 (intertemporal inequality aversion), ρ>0 (uncertainty aversion), βr (policy softness).
  • Limitations / modeling assumptions:
    • Requires a goal-conditioned human behavior model πH and structural world model Γ (the robot does not infer actual current human goals).
    • Assumes finite, acyclic games for analytic convenience.
    • Several formal choices are normative (ξ, η, ρ, ζ), not uniquely determined by axioms; different choices change incentives.
  • Empirical/implementation status:
    • This is a theory paper; a companion paper is promised to study scalable estimation and empirical validation.

Implications for AI Economics

  • Distributional and welfare implications:
    • The objective explicitly encodes inequality aversion (ξ) and prioritizes protecting low-power humans. If adopted by deployed AI (e.g., platforms, automation systems), this can shift surplus from high-ability / high-leverage agents toward broader empowerment—potentially reducing extreme concentration of influence and economic rents tied to automation.
    • The long-term and uncertainty aversion parameters (η, ρ) embed precautionary motives into AI decisions; this may reduce throughput-maximizing but risky strategies, affecting productivity-growth trade-offs.
  • Firm incentives, competition, and strategic adoption:
    • Firms that adopt empowerment-preserving AI objectives may face short-term competitiveness costs relative to firms optimizing conventional profit proxies, creating competitive dynamics and potential first-mover disadvantages unless regulated or market norms change.
    • Regulation or standards could internalize the externality (preserve human agency as a public good): policymakers might require or certify such objective functions for deployed systems where human empowerment is socially valuable.
  • Labor markets and innovation:
    • By protecting human capability to act and to pursue diverse goals, these objectives can mitigate some displacement harms (workers retaining more real options). This could alter skill investment incentives and the shape of labor demand (favoring human-in-the-loop roles).
    • Conversely, stronger constraints on irreversible or high-impact automation may slow certain types of technological deployment, shifting the timing and distribution of productivity gains.
  • Mechanism design and markets:
    • The P(s) aggregator (−log-sum-exp across individuals) resembles social-welfare weights that heavily prioritize the tail; integrating such AI objectives into market mechanisms (auctions, matching, recommender systems) will change allocation rules and may require redesign of incentives to preserve efficiency while upholding empowerment.
  • Measurement, implementation, and compliance costs:
    • Operationalizing the metric requires estimating Gh, πH, and transition models—data- and modeling-intensive tasks. There are measurement risks (mis-specified πH or Gh could produce perverse incentives).
    • Firms may face costs to build or verify required world/human models; these are economic considerations for adoption and regulation.
  • Strategic responses and gaming:
    • Agents (humans, firms) might seek to change observable behavior or environments to appear more empowered under the model, potentially producing strategic manipulation unless the human-modeling component is robust.
  • Research and policy directions:
    • Empirical work needed: evaluate how objective choice (ξ, η, ρ, ζ) affects economic outcomes—inequality, growth, robustness—and find operational estimators for Gh and πH.
    • Cost–benefit and general equilibrium analyses to assess long-run macroeconomic impacts of broad adoption.
    • Designing regulatory instruments and market incentives to align firm behavior with empowerment-preserving objectives without stifling innovation.

Short concluding note: the paper provides a coherent, axiomatic formalism for embedding human-empowerment preservation in AI objectives. For AI economics, its main value is supplying a transparent, parameterized bridge between normative social-welfare concerns (inequality, precaution, intertemporal fairness) and concrete algorithmic incentives—enabling explicit analysis of the trade-offs and market consequences of committing AI systems to preserve human agency.

Assessment

Paper Typetheoretical Evidence Strengthn/a — Paper is a theoretical/axiomatic derivation proposing an objective for AI; it presents formal propositions and proofs but contains no empirical or causal identification or measurement of real-world effects. Methods Rigormedium — The paper uses a clear axiomatic approach with formal propositions and proofs and derives closed-form objective functions; however it relies on multiple normative choices (e.g., choice of functional forms, specific parameterizations), strong modelling assumptions (finite acyclic stochastic games, availability of a goal-conditioned human policy), and pragmatic simplifications whose real-world applicability and robustness are not demonstrated. SampleNo empirical sample — theoretical model only: a finite, acyclic stochastic game with an AI agent (robot) and multiple humans; humans characterized by a set of possible goals Gh and a goal-conditioned behavioral policy πH; world transitions given by a known stochastic kernel; derivations produce per-goal capability C, per-person power I, present aggregate P, trajectory aggregate T, and long-term aggregate L, parameterized by normative discount/aversion parameters. Themeshuman_ai_collab governance inequality GeneralizabilityAssumes a finite, acyclic game form and accurate structural world model; real-world interactions are often cyclic/continuous and partially observed., Requires a goal-conditioned human policy πH and a specification of possible goal sets Gh — obtaining these in practice is difficult and may be misspecified., Normative parameters (γh, γr, ζ, ξ, η, ρ) are chosen by design and may affect incentives; paper does not prescribe how to set them empirically., Relies on structural (non-semantic) representations of goals; may miss context-dependent or latent human values and institutions., No empirical or experimental validation; scalability to large, open-ended, multi-agent environments is untested.

Claims (12)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Under the stated axioms (C1–C6), the goal-attainment capability metric must have a recursive form based on a continuous, strictly increasing transformation of the expected successor capability. Other positive Functional form of goal-attainment capability
Reading fidelity high
Study strength medium
not reported
0.12
With the paper's pragmatic functional-form choice, individual goal-attainment capability is a discounted probability of attaining the goal. Decision Quality positive Probability-weighted goal attainment
Reading fidelity high
Study strength medium
not reported
0.12
The individual human power metric derived from the axioms is the logarithm of the sum of goal-attainment capabilities raised to the reliability-preference parameter ζ. Other positive Individual human power across possible goals
Reading fidelity high
Study strength medium
not reported
0.12
If a human can choose between attaining k different goals with certainty, the proposed individual power metric equals log2 k. Skill Acquisition positive Number of reliably attainable goals
Reading fidelity high
Study strength medium
log2 k bits
0.12
Under axioms P0–P8 and either inequality-aversion condition P6 or limited-trade-off condition P7, present aggregate human power has the form P(s) = −log2 Σh∈H 2^(−ξIh(s)), where ξ > 0 controls inequality aversion. Inequality positive Inequality-averse aggregate human power across people
Reading fidelity high
Study strength medium
not reported
0.12
The proposed population metric gives an additional all-powerful human no effect in the limit, implementing a focus on preventing disempowerment rather than rewarding already powerful individuals. Inequality positive Marginal contribution of highly powerful individuals to aggregate power
Reading fidelity high
Study strength medium
not reported
0.12
For sufficiently strong inequality aversion, reducing one person's power from one bit to zero cannot be offset by arbitrarily increasing the power of a bounded number of other people. Inequality negative Compensability of individual power losses
Reading fidelity high
Study strength medium
k ≤ 2^ξ − 1
0.12
The temporal aggregation axioms yield a discounted aggregation of present aggregate human power or a logarithmic inequality-averse alternative; the authors select the form that penalizes intertemporal power inequality. Social Protection mixed Longitudinal distribution of aggregate human power
Reading fidelity high
Study strength medium
not reported
0.12
The uncertainty-averse long-term aggregate human-power metric is L(st) = −log2 E[(Σu≥t γr^(u−t) 2^(−ηP(su)))^ρ], with ρ > 0. Organizational Efficiency positive Expected long-term aggregate human power under uncertainty
Reading fidelity high
Study strength medium
not reported
0.12
The authors argue that softly maximizing the proposed human-power metric will likely incentivize an AI system to communicate well, follow orders, remain corrigible, avoid irreversible environmental changes, protect humans and itself from harm and disempowerment, follow relevant social norms, and allocate resources fairly and sustainably. Ai Safety And Ethics positive AI behavioral incentives related to human empowerment and safety
Reading fidelity high
Study strength speculative
not reported
0.02
The proposed approach is goal-agnostic in the sense that the AI system does not model or infer humans' actual current goals; instead, it uses a model of goal-conditioned human behavior and the environment's dynamics. Ai Safety And Ethics positive Dependence of AI policy on human goal information
Reading fidelity high
Study strength medium
not reported
0.12
The paper does not report empirical validation or a human/AI experimental sample; it presents a theoretical framework and states that empirical validation is reserved for a companion paper. Other null_result Empirical validation of the proposed metric
Reading fidelity high
Study strength high
not reported
0.2

Notes