0 cumulative citations
View corpus contextA hybrid of classical choice models and modern ML best advances mobility behavior modeling: preserve network feasibility and interpretable rewards, and use IRL/IL and graph/sequence models to scale, integrate data, and capture heterogeneity; but learned rewards or policies are not by themselves interchangeable with preferences or valid policy counterfactuals.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Route and activity choice are connected levels of a common sequential mobility decision problem: activity choice determines what people do, where, and when, while route choice governs how they move between activities. This review develops a unified framework connecting transportation choice modeling with inverse reinforcement learning (IRL) and imitation learning (IL). Under explicit assumptions, recursive logit, logit dynamic discrete choice, and maximum-entropy IRL share a soft Bellman representation, while trajectory occupancies and network flows satisfy related conservation laws. However, utility, reward, policy, occupancy, constraints, and observation errors remain different estimands with different behavioral and counterfactual interpretations. We review constrained and inverse-constrained learning, occupancy-ratio and DICE methods, incomplete and mixed-quality demonstrations, graph and sequence learning, transfer, data fusion, multi-agent choice, and large language models. Our central message is that machine learning adds the greatest value when embedded within a behaviorally disciplined framework: exact transitions enforce feasibility, structured rewards preserve interpretable trade-offs, observation models address heterogeneous data sources, and network or equilibrium solvers produce coherent system outcomes. Such hybrid models can improve scalability and prediction without sacrificing behavioral identification or policy relevance.
Summary
Main Finding
Route and activity choice are instances of a single sequential-mobility-choice problem: people make feasible, purposeful sequences of travel and activity decisions on spatial, temporal, and resource‑constrained networks. The most useful and policy‑relevant models combine the structural identification and counterfactual logic of discrete‑choice/structural models with the representation, scale, and data‑integration strengths of inverse reinforcement learning (IRL) and imitation learning (IL). Equivalences (a “soft Bellman” form and occupancy–flow conservation) make computational cross‑fertilization possible, but learned objects (utility, reward, policy, occupancy) are not interchangeable for economic interpretation or welfare/counterfactual analysis. The recommended approach is a behaviorally disciplined hybrid: enforce exact feasibility and constraints, keep interpretable utility/reward components, use learned modules for high‑dimensional context and heterogeneity, and use source‑aware observation models to fuse diverse data.
Key Points
-
Conceptual unification
- Defines sequential mobility choice as context‑conditioned sequences of feasible travel/activity actions through an augmented state (space, time, resources).
- Shows recursive logit (transportation), logit dynamic discrete choice (econometrics), and maximum‑entropy IRL (ML) share a soft Bellman recursion under explicit assumptions: same computational core but different interpretations.
-
What equivalences imply and do not imply
- Soft Bellman and occupancy–flow relations are computationally valuable: they enable trajectory distributions and network flows to be computed without full path enumeration.
- Equivalence across formalisms does not make the estimands identical: a learned reward that reproduces trajectories need not represent preferences, welfare, or credible policy counterfactuals unless identification conventions and constraints are specified.
-
Taxonomy of learning targets and methods
- Direct policy learning (IL), reward learning (IRL), occupancy matching / adversarial imitation, DICE and density‑ratio correction methods for off‑policy data, inverse constraint learning for constrained MDPs, offline RL, graph / sequence models (GNNs, Transformers), generative trajectory models (diffusion / flow matching), and LLMs for semantic layers.
- Special attention to incomplete and mixed‑quality demonstrations, multi‑source fusion, transfer learning, and multi‑agent strategic models.
-
Practical hybrid architecture
- Four distinct layers: environment (exact transitions & feasibility), behavioral objective (interpretable utility/reward & constraints), stochastic choice mechanism (policy/entropy structure), and observation process (source‑aware noise/selection).
- Use ML where it adds scale, representation, or integration, but keep behavioral identification explicit for counterfactuals and welfare.
-
Evaluation and research agenda
- Proposes conservative evidence dimensions: predictive performance, behavioral interpretation, reward/parameter recovery, transfer, and system/intervention validity.
- Calls for benchmarks, solver‑aware representation learning, source‑aware fusion, separating preferences/constraints/observations, linking individual learning to equilibrium, and verification workflows for LLMs.
-
Risks and cautions
- Learned rewards/policies can absorb observation error, selection bias, congestion, and unavailable alternatives, producing misleading counterfactuals.
- Reward functions are non‑unique (shaping, scaling, constraints matter); must choose estimand by scientific question.
- Privacy concerns in fused mobility data; ML priorities (prediction wins vs. causal validity) can misalign with policy needs.
Data & Methods
-
Review scope & search
- Critical integrative review (not exhaustive); focuses on individual/household sequential behavior across route, multi‑modal journeys, and daily activity scheduling.
- Literature screened via combined search terms (transportation + sequential model + learning) and updated through 15 Aug 2026; inclusion required learning/estimating utility, reward, policy, occupancy, or trajectory distributions or methodological identification/reproducibility results.
-
Formal framework used
- Four‑layer model: (1) Environment (network, schedule, resources, feasibility), (2) Behavioral objective (utility or reward and constraints), (3) Stochastic choice mechanism (extreme‑value shocks / entropy regularization → soft Bellman recursion), (4) Observation process (sampling, missingness, measurement error).
- Demonstrates mathematically: recursive logit, logit dynamic discrete choice, and max‑entropy IRL produce the same soft recursion when assumptions hold; trajectory occupancies satisfy conservation laws analogous to flow models.
-
Methods surveyed and evaluated
- Transportation foundations: path‑based random utility models (path enumeration, choice‑set sampling issues), recursive link‑choice (recursive logit), flow‑based assignment, and activity/schedule choice (augmented state time‑space networks).
- ML/IRL/IL methods: maximum‑entropy IRL, direct policy imitation (behavior cloning, adversarial IL), occupancy‑matching and DICE (density‑ratio corrected off‑policy matching), inverse constraint learning (recovering constraints/budgets), offline RL, graph neural networks and transformer sequence models for long/higher‑dimensional contexts, generative trajectory models (diffusion / flow matching), plus LLMs for semantics and preprocessing.
- Practical concerns treated: incomplete / suboptimal demonstrations, multi‑source fusion (surveys + passive traces), transfer learning, multi‑agent strategic interaction and equilibrium, risk/uncertainty quantification, evaluation metrics, and computational scaling (accelerator‑aware solvers).
-
Evaluation protocol advocated
- Report network size, solver tolerances, runtime, and data provenance; evaluate along five conservative dimensions: in‑sample fit (representation), held‑out trajectory prediction, domain transfer, parameter/reward recovery (requires synthetic/known truth or interventions), and system/intervention validity (policy counterfactuals and equilibrium effects).
Implications for AI Economics
-
On identification and estimands
- Economists should not conflate a policy or reward learned to match observed behavior with recovered structural preferences or welfare metrics. IRL/IL outputs need normalization, constraints, and explicit identification choices to be interpretable for policy (e.g., value of time, compensating variation).
- Choosing the estimand (utility parameters, reward shaping, policy, occupancy distribution, or constraints) must be driven by the economic question (prediction vs. causal policy vs. welfare analysis).
-
For policy and counterfactual analysis
- Structural feasibility (feasible choice sets, schedule/resource constraints, and equilibrium feedback) must be enforced to obtain credible counterfactuals. Purely policy‑optimizing or policy‑imitating ML models can recommend actions that violate feasibility or mispredict system equilibrium responses.
- Inverse constraint learning and constrained MDPs are promising for separating hard bounds (e.g., budgets, time windows) from preferences — crucial for realistic policy simulation.
-
For applied research design
- When using ML to model mobility behavior, explicitly model observation/selection processes (platform bias, missing trips) to avoid absorbing these effects into estimated preferences or rewards.
- Use solver‑aware representation learning: representations and amortized value functions should be designed so they can be converted back into paths, schedules, flows, welfare measures, and policy counterfactuals.
-
For equilibrium and multi‑agent settings
- Multi‑agent learning is necessary to capture congestion and strategic responses; individual‑level learned policies must be linked to system equilibrium solvers to assess aggregate outcomes and distributional effects.
- Offline or logged‑data methods (DICE, occupancy correction) enable learning without risky online experimentation but require careful density‑ratio correction and evaluation under covariate shift.
-
For data, transfer, and privacy
- ML enables fusing sparse, semantic survey data with dense passive traces—improving scale and context—only if source‑aware observation models preserve behavioral meaning.
- Privacy risks increase with high‑fidelity fusion; economists must balance richer inference against re‑identification and distributional harms.
-
For broader AI economics research
- The paper’s agenda (benchmarks, transfer tests, verification workflows for LLMs, separation of preferences/constraints/observation, linking individual learning to equilibrium) lays out concrete priorities for empirical validation and reproducibility in AI‑augmented economic modeling.
- Adoption of these standards would help ensure ML contributions are assessed by their ability to improve policy‑relevant inference (welfare, elasticity, distributional impacts), not just predictive accuracy.
Summary takeaway for AI economists: IRL and imitation methods expand reach and scalability for sequential mobility problems, but economic interpretation and policy validity demand a structured hybrid design that preserves feasibility, explicit behavioral primitives, and careful identification of what the model learns.
Assessment
Claims (12)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Route choice and activity choice can be represented as different instances of a common sequential mobility-choice problem involving feasible decisions on spatial, temporal, and resource-constrained networks. Decision Quality | positive | Unified representation of sequential mobility behavior |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Under explicit assumptions, recursive logit, logit dynamic discrete choice, and maximum-entropy inverse reinforcement learning share a soft Bellman representation. Organizational Efficiency | positive | Model-form equivalence and computational representation |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The equivalence between recursive choice, dynamic discrete choice, and maximum-entropy IRL does not by itself establish that the models identify the same behavioral quantities. Decision Quality | mixed | Behavioral identification and validity of policy counterfactuals |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Activity choice is a harder sequential-learning problem than route choice because activity goals may be endogenous, the horizon is longer, durations matter, and resource and coordination constraints can bind across episodes. Task Completion Time | negative | Complexity of learning feasible activity sequences |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Imitation learning can reproduce observed behavior without explaining the underlying preferences or behavioral mechanism. Decision Quality | mixed | Interpretability and structural explanation of observed mobility behavior |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Offline reinforcement learning based on logged behavior is prescriptive rather than automatically descriptive of the behavior-generating process. Decision Quality | mixed | Descriptive validity versus policy improvement from logged mobility data |
Reading fidelity
high
Study strength
medium
|
not reported
|
| A behaviorally disciplined hybrid architecture should use exact network, schedule, and resource transitions for feasibility; structured rewards for interpretable trade-offs; learned components for context and heterogeneity; and source-aware observation models for combining surveys with passive traces. Organizational Efficiency | positive | Behavioral coherence and scalability of mobility-choice models |
Reading fidelity
high
Study strength
low
|
not reported
|
| Machine-learning capabilities such as occupancy-ratio methods, graph and sequence models, constrained learning, transfer, data fusion, and multi-agent learning do not by themselves establish preference recovery or valid policy counterfactuals. Decision Quality | mixed | Preference recovery and policy-counterfactual validity |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Recursive route-choice models generate a distribution over feasible paths through sequences of local link choices, avoiding explicit enumeration of all paths. Organizational Efficiency | positive | Computational handling of large route choice sets |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Path-based random-utility models provide interpretable marginal utilities, elasticities, values of time, and welfare analysis, but their use can be limited by enormous path spaces and unknown or misspecified choice-set generation. Decision Quality | mixed | Interpretability and validity of route-choice estimation |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The review is a critical integrative review rather than a systematic review claiming exhaustive coverage, so it does not support PRISMA-style prevalence claims. Research Productivity | null_result | Exhaustiveness and generalizability of the literature review |
Reading fidelity
high
Study strength
high
|
not reported
|
| In-sample fit supports claims about representation, held-out trajectories support prediction, domain shifts support transfer, and interventions or known synthetic truths are needed to support structural recovery. Decision Quality | positive | Strength of evidence for prediction, transfer, and structural recovery |
Reading fidelity
high
Study strength
medium
|
not reported
|