0 cumulative citations
View corpus contextA theoretical model finds that simple acyclicity can rationalize AI recommendations, but does not secure alignment; only stronger double-monotonicity and an idempotence constraint make the AI's inferred preferences unique and keep recommendations within the original feasible set.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
This paper proposes a model of choice via agentic artificial intelligence (AI). A key feature is that the AI may misinterpret a menu before recommending what to choose. A single acyclicity condition guarantees that there is a monotonic interpretation and a strict preference relation that together rationalize the AI's recommendations. Since this preference is in general not unique, there is no safeguard against it misaligning with that of a decision maker. What enables the verification of such AI alignment is interpretations satisfying double monotonicity. Indeed, double monotonicity ensures full identifiability and internal consistency. But, an additional idempotence property is required to guarantee that recommendations are fully rational and remain grounded within the original feasible set.
Summary
Main Finding
The paper develops a choice-theoretic model of an agentic AI that may misinterpret menus before recommending an alternative. It shows (constructively and axiomatically) when observed recommendations can be rationalized as resulting from (i) a strict preference ranking over alternatives and (ii) an interpretation operator that maps actual menus to the AI’s internal menus. Key results:
- A single acyclicity condition (No Shifted Cycles, NSC) characterizes when recommendations are rationalizable by some monotonic interpretation operator and some strict preference (AIC model).
- Strengthening the interpretation operator to satisfy double monotonicity (order isomorphism between actual menus and interpreted menus) yields full identification of both the AI’s preferences and its interpretation (RAIC model); this is characterized behaviorally by three axioms: No Binary Cycles (NBC), C‑Contraction Independence (CCI), and Noticeable Difference (ND).
- Adding idempotence of the interpretation operator (no further reinterpretation loops) makes recommendations grounded in the original feasible sets and equivalent to classical rational choice satisfying WARP (GRAIC model).
Key Points
-
Objects and primitives:
- Finite universe X of alternatives; P(X) non-empty menus; observed choice function c: P(X) → X.
- Interpretation operator I: P(X) → P(X). Minimal assumption (IM): monotonicity (S ⊆ T ⇒ I(S) ⊆ I(T)).
- AI chooses the best element in I(S) according to a strict preference ≻.
-
AIC (AI agent’s choice):
- Definition: c(S) = the ≻-maximal element of I(S).
- Behavioral characterization: NSC (No Shifted Cycles). Intuition: rules out cycles that could not arise from a strict (acyclic) preference after interpretation distortions.
-
Revealed objects under AIC:
- Revealed preference ≻*: constructed as transitive closure of the binary relation x ≻ y whenever x is chosen from some T and y chosen from a strict subset S ⊂ T.
- Revealed interpretation I: x ∈ I(T) iff x is chosen from some S ⊂ T.
- These capture what is identifiable under mere monotonic interpretation: preferences and interpretations are in general only partially identified.
-
RAIC (Rational AI agent’s choice):
- Strengthened requirement on I: Double Monotonicity (TDM): I(S) ⊆ I(T) ⇔ S ⊆ T. This makes the mapping between posets (P(X), ⊆) and ({I(S)}, ⊆) an order isomorphism.
- Behavioral characterization: three axioms
- NBC (No Binary Cycles): rules out certain binary inconsistencies.
- CCI (C‑Contraction Independence): contraction-consistency relative to interpreted sets.
- ND (Noticeable Difference): singleton choices reflect distinct alternatives.
- Under these axioms both I and ≻ are fully identified:
- I(T) = {c({x}) | x ∈ T}
- x ≻ y iff there exist S ⊂ T with c(T)=x and c(S)=y
-
GRAIC (Grounded RAIC):
- Add Idempotence (IIP): I(I(S)) = I(S) for all S; prevents interpretation loops.
- Equivalent behavioral condition: WARP (weak axiom of revealed preference).
- Consequence: choices are classical rational choices (grounded in the true menus) and both interpretation and preference reduce to the usual revealed constructions (I(T)=T; x ≻ y iff x chosen from a menu containing y).
-
Illustrative examples:
- Example showing NSC-satisfying data where preference ranking is not unique (partial identifiability).
- Example showing RAIC without grounding — the AI may recommend alternatives not in the original menu (c(S) ∉ S) even though behavior is internally consistent on interpreted menus.
Data & Methods
- This is a theoretical/axiomatic paper; no empirical data is used.
- Methods:
- Formal model: finite alternative set, deterministic choice function, interpretation operator, strict preferences.
- Axiomatic analysis: introduce behavioral axioms (NSC, NBC, CCI, ND, WARP) and prove equivalence to underlying representational structures (existence/uniqueness of ≻ and I with given properties).
- Constructive identification: explicit formulae/algorithms to recover revealed preference ≻ and interpretation I (and full identification under stronger conditions).
- Proofs and technical lemmas are provided in appendices (existence, asymmetry, acyclicity, transitive closure arguments).
- Assumptions and modeling choices:
- Finite X; deterministic choices; AI’s interpretation operator acts on menus and preserves monotonicity; preferences are strict and acyclic.
- Interpretation distortions may cause chosen alternatives to lie outside the true menu unless additional structure (idempotence) is imposed.
Implications for AI Economics
-
Formalizing AI recommendations: The paper gives a clean formal language to model how an AI’s implicit interpretation of a decision context can confound observed recommendations. This is useful for economics models that incorporate AI advisors or platform recommendation systems.
-
Testable audit criteria:
- NSC provides a minimal, testable necessary-and-sufficient condition for whether an AI’s recommendations can be rationalized by some consistent preference plus a monotonic interpretation — so auditors can check whether observed recommendation data is consistent with any internally coherent AI decision rule.
- Stronger tests (NBC + CCI + ND) detect when preferences and interpretations are fully recoverable (RAIC). WARP identifies when recommendations are both rational and grounded (GRAIC).
-
Alignment diagnostics and limits:
- Under monotonic I alone, preferences are only partially identifiable: many underlying ≻ can produce the same recommendations. Therefore, observing aligned recommendations does not guarantee alignment of underlying AI preferences.
- Double monotonicity is the key property that enables full identifiability of both preference and interpretation; designing models/systems/finetuning to enforce or approximate double monotonicity would make alignment verification possible from recommendations alone.
- Groundedness (idempotence) is required to ensure an AI’s recommendations are actually chosen from the original feasible set (important for legal/regulatory/consumer-protection concerns).
-
Design and policy implications:
- Training objectives and fine-tuning should prioritize learning consistent set-relations (monotonicity) and, if possible, inject constraints that approximate double monotonicity or idempotence to improve interpretability and verifiability of recommendations.
- Platform audits can operationalize the paper’s axioms: collect choice/recommendation logs over menus and run the NSC/NBC/CCI/WARP checks to detect misinterpretation, non‑identifiability, or ungrounded behavior.
- Economic models of markets with AI advisors should account for possible recommendation distortions (AI considering items outside the feasible set) and for the identifiability failure under weak interpretation assumptions.
-
Research directions:
- Extend to stochastic choice, large/infinite alternative spaces, dynamic/learning AIs, or richer models of misinterpretation (context-dependent/randomized I).
- Empirical work: apply axiomatic tests to real AI recommendation logs (LLM agents, search/recommendation systems) to measure prevalence of non‑grounded or non‑identifiable behavior.
- Mechanism/policy design: build incentive/verification schemes that encourage double monotonicity/idempotence (e.g., architectures or constraints that preserve menu structure).
Summary takeaway: the paper gives a compact axiomatic toolkit to separate misinterpretation effects from true irrationality in AI recommendations, shows when and how alignment and interpretability are (or are not) identifiable from observed recommendations, and links the behavioral tests (NSC, NBC+CCI+ND, WARP) to progressively stronger guarantees (rationalizability, full identification, grounded rationality).
Assessment
Claims (7)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| This paper proposes a model of choice via agentic artificial intelligence (AI). Other | positive | existence of a formal model of choice mediated by agentic AI |
Reading fidelity
high
Study strength
high
|
not reported
|
| A key feature is that the AI may misinterpret a menu before recommending what to choose. Decision Quality | negative | AI misinterpretation of a menu prior to recommendation |
Reading fidelity
high
Study strength
high
|
not reported
|
| A single acyclicity condition guarantees that there is a monotonic interpretation and a strict preference relation that together rationalize the AI's recommendations. Decision Quality | positive | existence of a monotonic interpretation and strict preference relation that rationalize AI recommendations under an acyclicity condition |
Reading fidelity
high
Study strength
high
|
not reported
|
| Since this preference is in general not unique, there is no safeguard against it misaligning with that of a decision maker. Decision Quality | negative | risk of preference misalignment between AI-inferred preference and decision maker's true preference due to non-uniqueness |
Reading fidelity
high
Study strength
medium
|
not reported
|
| What enables the verification of such AI alignment is interpretations satisfying double monotonicity. Decision Quality | positive | verifiability of AI alignment when interpretations satisfy double monotonicity |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Double monotonicity ensures full identifiability and internal consistency. Decision Quality | positive | identifiability of interpretation/preference and internal consistency of recommendations under double monotonicity |
Reading fidelity
high
Study strength
high
|
not reported
|
| An additional idempotence property is required to guarantee that recommendations are fully rational and remain grounded within the original feasible set. Decision Quality | positive | full rationality and feasibility (grounding in original feasible set) of AI recommendations when idempotence holds |
Reading fidelity
high
Study strength
medium
|
not reported
|