0 cumulative citations
View corpus contextDesigning AI interactions around reachable targets and actionable counterfactuals can steer people toward genuine improvement while limiting gaming, and algorithmic constraints enable trade-offs between incentives and classifier accuracy; these claims are supported by formal guarantees, dataset experiments, and a small online experiment of parental responses to device mistreatment.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
When an AI system is deployed, the individuals who use and or are evaluated by it form beliefs about how the system operates and use those beliefs to strategically present their preferences, behaviors, or attributes. The system then responds with feedback or a decision outcome, thereby creating a human-AI interaction loop. This thesis studies how to design and shape such interactions to achieve three goals: (1) help individuals develop accurate beliefs about the AI systems so they can improve and or secure favorable outcomes at minimal cost, (2) encourage improvement and or discourage gaming behaviors, and (3) ensure that the AI system continues to achieve its intended objectives, such as maximizing accuracy. To address these goals, the thesis is organized into three complementary parts that examine and study human-AI interactions from the perspectives of both evaluated individuals and AI systems. Together, the work presented in this thesis advances human-centered machine learning by providing principles and methods for designing AI systems that align with human needs, values, and capabilities. Methodologically, this thesis integrates theoretical analysis, data-driven modeling, human-subject experiments, and empirical evaluations on real-world and semi-synthetic datasets.
Summary
Main Finding
This thesis develops theory, algorithms, and human-subject evidence for designing human–AI interaction protocols that (a) give people realistic, actionable pathways to improve outcomes at low cost, (b) discourage gaming while encouraging genuine improvement, and (c) allow decision systems to balance competing objectives (e.g., social welfare, fairness, and predictive accuracy). The work shows concrete mechanisms — reachable-target setting, scalable counterfactual/explanation generators, role-model disclosure policies, and classification algorithms that account for agents who can improve or game — together produce provable guarantees and practical gains in simulated and real-data evaluations, while highlighting trade-offs between incentive design, learnability, and system accuracy.
Key Points
- Human behavior matters: deployed models interact with users who form beliefs and then strategically act (improve or game). Interaction design must shape these behaviors.
- Three-part approach:
- Empirical HRI evidence (Part I): an online experiment studying parental responses to children aggressing on robots vs. smart speakers/tablets to inform design of AI-mediated developmental pathways.
- Enabling improvement (Part II): algorithmic methods to (i) set reachable improvement targets to maximize overall (or minimum) improvement, (ii) generate actionable, resource-efficient counterfactual explanations (CFEs) in large / discrete state spaces, and (iii) selectively reveal positive/negative role models to maximize social welfare under diffusion-like influence models.
- Balancing competing objectives (Part III): formal models and algorithms for PAC learning when agents can improve, a theory of classification when agents can both improve and game (including hardness, approximation, and sample-complexity results), plus empirical studies of classifier performance and agent utility under diverse scenarios.
- Reachable-targets problem: formalizes selecting a small set of target levels that individuals can reach with limited effort; establishes monotonicity/submodularity properties, greedy algorithms with approximation guarantees, and generalization bounds.
- Scalable CFEs: introduces hierarchical low-level (hl) continuous and discrete CFEs and data-driven CFE generators that are accurate and computationally efficient for large state spaces; empirically superior to bruteforce low-level CFEs.
- Role-model disclosure: models selective revelation as a submodular influence-style optimization; greedy policies get approximation guarantees; extensions cover targeted interventions and heterogeneous coverage radii; empirical results show practical gains.
- Learning with improvements: introduces PAC learning with improvements, separating it from standard and strategic PAC models. Gives constructive learning algorithms (zero-error in some geometric/graph settings), sample-complexity bounds, and mechanisms to enable improvement when helpful.
- Classification of improving-and-gaming agents: defines formal models (general discrete and linear), studies algorithmic objectives (maximize true positives subject to no false positives), proves hardness results, derives sample-complexity bounds in full and partial information settings, and analyzes optimal linear classifiers (including 2D characterizations).
- Empirical evaluation: simulations and experiments on real and semi-synthetic datasets show how classifier design, agent risk-aversion, geometric constraints, and the presence of gaming affect utility, accuracy, and social-welfare trade-offs.
Data & Methods
- Human-subject experiment (Part I)
- Online video-based protocol presenting parents with scenarios of children being aggressive to three device types (embodied robot, smart speaker, tablet).
- Measures: parental concern, likely interventions, perceptions of mistreatment and sympathy, and device attributions.
- Sample: recruited participants (details in chapter; ethical considerations discussed).
- Theoretical modeling & algorithm design (Parts II & III)
- Formal models for reachable targets, CFEs, role-model influence (submodular welfare), PAC learning with improvements, and improving/gaming agents (discrete and linear).
- Tools: submodular optimization (greedy algorithms), linear programming, PAC-analysis, sample-complexity proofs, hardness reductions.
- Approximation algorithms with provable guarantees; FPTAS for some variants; bounds on generalization from samples.
- Scalable counterfactual/explanation generation
- Hierarchical CFE constructs (hl-continuous, hl-discrete), plus data-driven predictors that learn where CFEs are feasible.
- Experimental evaluation uses real-world and semi-synthetic (agent–CFE) datasets; metrics include coverage, fidelity, cost-efficiency, and runtime.
- Empirical classification studies
- Simulated agents who may invest effort to genuinely improve or engage in gaming; utility defined by benefit minus cost and risk-aversion parameters.
- Experiments explore classifier choices (linear vs general), agent capabilities, and resulting trade-offs (accuracy, utility, false positives).
- Datasets: real and semi-/fully-synthetic datasets derived from real distributions (chapter provides dataset extraction/preprocessing and synthetic generation procedures).
Implications for AI Economics
- Incentive-aware deployment: Systems should be designed as mechanisms that explicitly trade off predictive accuracy and the externalities of information/feedback they provide (e.g., enabling improvement vs facilitating gaming). This reframes recourse/explanation design as an economic policy decision.
- Social-welfare optimization: Selecting whom to show targets, role models, or recourse should be treated as a resource allocation problem with submodular structure — enabling near-optimal greedy policies for maximizing aggregate improvement or fairness objectives.
- Cost of improvement vs gaming: Quantifying agents’ effort costs and risk aversion is crucial. Policies that lower true-improvement costs (easier, reachable targets, actionable CFEs) can increase productive investment and reduce socially costly gaming. Conversely, naive transparency can create negative incentives (cheap gaming).
- Regulator and firm trade-offs: Firms optimizing accuracy must account for user adaptation dynamics. The thesis gives tools to certify classifiers that minimize false positives while maximizing true positives in presence of strategic behavior — important for regulated domains (credit, hiring, etc.). Regulators may require bounded false-positive guarantees under strategic responses.
- Data collection and auditing economics: Sample-complexity results show how many examples (including post-intervention outcomes) are needed to learn reliably when agents can adapt. This affects the cost of auditing, monitoring, and updating deployed systems under strategic populations.
- Targeted support and equity: Designing reachable-target sets and role-model disclosures can be used to prioritize disadvantaged groups cheaply by maximizing minimum improvements or by constrained welfare objectives; this connects to economic allocation of remediation resources and fairness interventions.
- Operational guidance for product design: Practical, data-driven CFE generators make it economically feasible for firms to provide individualized, low-cost recourse at scale, altering the marginal cost-benefit calculation of offering explanations versus the risk of gaming.
- Measurement & policy cautions: The work highlights that zero classification error on observed (post-adaptation) data does not guarantee optimal social utility — economic evaluation must include costs borne by agents, long-term improvement dynamics, and distributional effects.
Limitations and ethical notes: the thesis discusses limitations (modeling assumptions about agent costs/capacities, risk aversion, and simulated/synthetic data) and ethical considerations (responsible design of recourse so as not to empower harmful gaming or manipulate vulnerable users). Future work needs richer behavioral data, longitudinal studies of improvement, and careful regulatory framing.
Assessment
Claims (8)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| The thesis develops methods to help individuals form accurate beliefs about AI systems so that they can improve their attributes or secure favorable outcomes at minimal cost. Decision Quality | positive | Individuals' ability to improve or obtain favorable AI-mediated decisions at low cost |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The thesis develops methods for setting reachable targets that maximize individual improvement in interactions with AI-driven decision systems. Skill Acquisition | positive | Amount of improvement achieved by individuals |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The thesis develops actionable guidance intended to help individuals reverse unfavorable outcomes from AI-driven decisions. Decision Quality | positive | Ability to reverse an unfavorable decision outcome |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The thesis proposes selectively revealing information about positive and negative role models to maximize the expected number of individuals who emulate positive role models. Training Effectiveness | positive | Expected number of individuals emulating positive role models |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The thesis analyzes how individuals' capacity for improvement affects the learnability and algorithmic design of accurate classification systems. Decision Quality | mixed | Classification learnability and predictive accuracy under individual improvement |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The thesis develops theoretical and empirical foundations for classifying individuals while maximizing true positives, minimizing false positives, penalizing gaming, and incentivizing genuine improvement. Decision Quality | mixed | True-positive rate, false-positive rate, gaming behavior, and genuine improvement under classification |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The data-driven counterfactual-explanation generators developed in the thesis are described as accurate and resource-efficient. Decision Quality | positive | Accuracy and resource efficiency of data-driven counterfactual-explanation generation |
Reading fidelity
high
Study strength
low
|
not reported
|
| The thesis uses an online human-subject experiment to investigate how parents perceive and respond to children's aggressive behavior toward embodied and disembodied AI-driven home devices. Ai Safety And Ethics | mixed | Parental perceptions and responses to children's aggressive behavior toward AI devices |
Reading fidelity
high
Study strength
medium
|
not reported
|