The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Designing AI interactions around reachable targets and actionable counterfactuals can steer people toward genuine improvement while limiting gaming, and algorithmic constraints enable trade-offs between incentives and classifier accuracy; these claims are supported by formal guarantees, dataset experiments, and a small online experiment of parental responses to device mistreatment.

Shaping Human-AI Interactions to Provide Improvement Pathways and Balance Competing Objectives
Keziah Naggita · August 06, 2026
arxiv other medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Keziah Naggita unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Keziah Naggita provider ID
This thesis develops theoretical models, algorithms, and empirical methods to design human-AI interactions that provide actionable improvement pathways, discourage gaming, and balance individual incentives with classifier accuracy, supported by a human-subject vignette study and dataset evaluations.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

When an AI system is deployed, the individuals who use and or are evaluated by it form beliefs about how the system operates and use those beliefs to strategically present their preferences, behaviors, or attributes. The system then responds with feedback or a decision outcome, thereby creating a human-AI interaction loop. This thesis studies how to design and shape such interactions to achieve three goals: (1) help individuals develop accurate beliefs about the AI systems so they can improve and or secure favorable outcomes at minimal cost, (2) encourage improvement and or discourage gaming behaviors, and (3) ensure that the AI system continues to achieve its intended objectives, such as maximizing accuracy. To address these goals, the thesis is organized into three complementary parts that examine and study human-AI interactions from the perspectives of both evaluated individuals and AI systems. Together, the work presented in this thesis advances human-centered machine learning by providing principles and methods for designing AI systems that align with human needs, values, and capabilities. Methodologically, this thesis integrates theoretical analysis, data-driven modeling, human-subject experiments, and empirical evaluations on real-world and semi-synthetic datasets.

Summary

Main Finding

This thesis develops theory, algorithms, and human-subject evidence for designing human–AI interaction protocols that (a) give people realistic, actionable pathways to improve outcomes at low cost, (b) discourage gaming while encouraging genuine improvement, and (c) allow decision systems to balance competing objectives (e.g., social welfare, fairness, and predictive accuracy). The work shows concrete mechanisms — reachable-target setting, scalable counterfactual/explanation generators, role-model disclosure policies, and classification algorithms that account for agents who can improve or game — together produce provable guarantees and practical gains in simulated and real-data evaluations, while highlighting trade-offs between incentive design, learnability, and system accuracy.

Key Points

  • Human behavior matters: deployed models interact with users who form beliefs and then strategically act (improve or game). Interaction design must shape these behaviors.
  • Three-part approach:
  • Empirical HRI evidence (Part I): an online experiment studying parental responses to children aggressing on robots vs. smart speakers/tablets to inform design of AI-mediated developmental pathways.
  • Enabling improvement (Part II): algorithmic methods to (i) set reachable improvement targets to maximize overall (or minimum) improvement, (ii) generate actionable, resource-efficient counterfactual explanations (CFEs) in large / discrete state spaces, and (iii) selectively reveal positive/negative role models to maximize social welfare under diffusion-like influence models.
  • Balancing competing objectives (Part III): formal models and algorithms for PAC learning when agents can improve, a theory of classification when agents can both improve and game (including hardness, approximation, and sample-complexity results), plus empirical studies of classifier performance and agent utility under diverse scenarios.
  • Reachable-targets problem: formalizes selecting a small set of target levels that individuals can reach with limited effort; establishes monotonicity/submodularity properties, greedy algorithms with approximation guarantees, and generalization bounds.
  • Scalable CFEs: introduces hierarchical low-level (hl) continuous and discrete CFEs and data-driven CFE generators that are accurate and computationally efficient for large state spaces; empirically superior to bruteforce low-level CFEs.
  • Role-model disclosure: models selective revelation as a submodular influence-style optimization; greedy policies get approximation guarantees; extensions cover targeted interventions and heterogeneous coverage radii; empirical results show practical gains.
  • Learning with improvements: introduces PAC learning with improvements, separating it from standard and strategic PAC models. Gives constructive learning algorithms (zero-error in some geometric/graph settings), sample-complexity bounds, and mechanisms to enable improvement when helpful.
  • Classification of improving-and-gaming agents: defines formal models (general discrete and linear), studies algorithmic objectives (maximize true positives subject to no false positives), proves hardness results, derives sample-complexity bounds in full and partial information settings, and analyzes optimal linear classifiers (including 2D characterizations).
  • Empirical evaluation: simulations and experiments on real and semi-synthetic datasets show how classifier design, agent risk-aversion, geometric constraints, and the presence of gaming affect utility, accuracy, and social-welfare trade-offs.

Data & Methods

  • Human-subject experiment (Part I)
    • Online video-based protocol presenting parents with scenarios of children being aggressive to three device types (embodied robot, smart speaker, tablet).
    • Measures: parental concern, likely interventions, perceptions of mistreatment and sympathy, and device attributions.
    • Sample: recruited participants (details in chapter; ethical considerations discussed).
  • Theoretical modeling & algorithm design (Parts II & III)
    • Formal models for reachable targets, CFEs, role-model influence (submodular welfare), PAC learning with improvements, and improving/gaming agents (discrete and linear).
    • Tools: submodular optimization (greedy algorithms), linear programming, PAC-analysis, sample-complexity proofs, hardness reductions.
    • Approximation algorithms with provable guarantees; FPTAS for some variants; bounds on generalization from samples.
  • Scalable counterfactual/explanation generation
    • Hierarchical CFE constructs (hl-continuous, hl-discrete), plus data-driven predictors that learn where CFEs are feasible.
    • Experimental evaluation uses real-world and semi-synthetic (agent–CFE) datasets; metrics include coverage, fidelity, cost-efficiency, and runtime.
  • Empirical classification studies
    • Simulated agents who may invest effort to genuinely improve or engage in gaming; utility defined by benefit minus cost and risk-aversion parameters.
    • Experiments explore classifier choices (linear vs general), agent capabilities, and resulting trade-offs (accuracy, utility, false positives).
    • Datasets: real and semi-/fully-synthetic datasets derived from real distributions (chapter provides dataset extraction/preprocessing and synthetic generation procedures).

Implications for AI Economics

  • Incentive-aware deployment: Systems should be designed as mechanisms that explicitly trade off predictive accuracy and the externalities of information/feedback they provide (e.g., enabling improvement vs facilitating gaming). This reframes recourse/explanation design as an economic policy decision.
  • Social-welfare optimization: Selecting whom to show targets, role models, or recourse should be treated as a resource allocation problem with submodular structure — enabling near-optimal greedy policies for maximizing aggregate improvement or fairness objectives.
  • Cost of improvement vs gaming: Quantifying agents’ effort costs and risk aversion is crucial. Policies that lower true-improvement costs (easier, reachable targets, actionable CFEs) can increase productive investment and reduce socially costly gaming. Conversely, naive transparency can create negative incentives (cheap gaming).
  • Regulator and firm trade-offs: Firms optimizing accuracy must account for user adaptation dynamics. The thesis gives tools to certify classifiers that minimize false positives while maximizing true positives in presence of strategic behavior — important for regulated domains (credit, hiring, etc.). Regulators may require bounded false-positive guarantees under strategic responses.
  • Data collection and auditing economics: Sample-complexity results show how many examples (including post-intervention outcomes) are needed to learn reliably when agents can adapt. This affects the cost of auditing, monitoring, and updating deployed systems under strategic populations.
  • Targeted support and equity: Designing reachable-target sets and role-model disclosures can be used to prioritize disadvantaged groups cheaply by maximizing minimum improvements or by constrained welfare objectives; this connects to economic allocation of remediation resources and fairness interventions.
  • Operational guidance for product design: Practical, data-driven CFE generators make it economically feasible for firms to provide individualized, low-cost recourse at scale, altering the marginal cost-benefit calculation of offering explanations versus the risk of gaming.
  • Measurement & policy cautions: The work highlights that zero classification error on observed (post-adaptation) data does not guarantee optimal social utility — economic evaluation must include costs borne by agents, long-term improvement dynamics, and distributional effects.

Limitations and ethical notes: the thesis discusses limitations (modeling assumptions about agent costs/capacities, risk aversion, and simulated/synthetic data) and ethical considerations (responsible design of recourse so as not to empower harmful gaming or manipulate vulnerable users). Future work needs richer behavioral data, longitudinal studies of improvement, and careful regulatory framing.

Assessment

Paper Typeother Evidence Strengthmedium — The thesis combines rigorous theoretical results (strong internal validity for formal claims) with controlled lab-style human-subject evidence and dataset experiments; however, the empirical components are limited to an online vignette experiment and dataset/simulation studies rather than large-scale field or longitudinal data that would support strong causal claims about economic outcomes or real-world productivity impacts. Methods Rigorhigh — Methods span formal theorem-proving, algorithm design and complexity/hardness analyses, PAC-style sample-complexity bounds, and carefully described experimental procedures (human-subject vignette experiment, semi-synthetic and synthetic data experiments). Theoretical sections appear thorough; empirical work uses multiple datasets and evaluation metrics, though human-subject and external-validity aspects are relatively narrow. SampleOnline human-subject experiment with parents viewing scripted video scenarios comparing child aggressive behavior toward embodied and disembodied devices (robots, smart speakers, tablets) under randomized experimental conditions; additionally, empirical evaluations use several real-world datasets plus semi-synthetic and fully synthetic agent–counterfactual-example (CFE) datasets to test CFE generators and classifiers; much of the remaining work is theoretical/algorithmic and validated via simulation. Themeshuman_ai_collab skills_training governance IdentificationMixed: Part I uses a randomized online human-subject experiment (vignette/video treatments varying device type and behavior) to identify how parents respond to child aggression toward devices; Parts II–III are primarily theoretical/algorithmic with formal proofs, worst-case guarantees, and simulation/empirical evaluations on real, semi-synthetic, and fully synthetic datasets rather than field causal identification. GeneralizabilityOnline vignette experiment (videos) may not reflect parents' real-world behavior in naturalistic settings (limited ecological validity)., Participant sample details (size, recruitment source, demographics) are not provided here; may be a convenience/online sample limiting representativeness., Empirical dataset experiments rely on semi-synthetic/fully synthetic agents and specific datasets, which may not generalize to all decision contexts or sectors (e.g., hiring, lending)., Theoretical models make structural assumptions (e.g., linear models, graph/coverage models, specific improvement/gaming cost structures) that may not hold in all real-world settings., Work focuses on micro-level interactions and classification incentives, so macroeconomic/general equilibrium implications for productivity and labor markets are not directly addressed.

Claims (8)

ClaimDirectionOutcomeConfidence & EvidenceDetails
The thesis develops methods to help individuals form accurate beliefs about AI systems so that they can improve their attributes or secure favorable outcomes at minimal cost. Decision Quality positive Individuals' ability to improve or obtain favorable AI-mediated decisions at low cost
Reading fidelity high
Study strength medium
not reported
0.12
The thesis develops methods for setting reachable targets that maximize individual improvement in interactions with AI-driven decision systems. Skill Acquisition positive Amount of improvement achieved by individuals
Reading fidelity high
Study strength medium
not reported
0.12
The thesis develops actionable guidance intended to help individuals reverse unfavorable outcomes from AI-driven decisions. Decision Quality positive Ability to reverse an unfavorable decision outcome
Reading fidelity high
Study strength medium
not reported
0.12
The thesis proposes selectively revealing information about positive and negative role models to maximize the expected number of individuals who emulate positive role models. Training Effectiveness positive Expected number of individuals emulating positive role models
Reading fidelity high
Study strength medium
not reported
0.12
The thesis analyzes how individuals' capacity for improvement affects the learnability and algorithmic design of accurate classification systems. Decision Quality mixed Classification learnability and predictive accuracy under individual improvement
Reading fidelity high
Study strength medium
not reported
0.12
The thesis develops theoretical and empirical foundations for classifying individuals while maximizing true positives, minimizing false positives, penalizing gaming, and incentivizing genuine improvement. Decision Quality mixed True-positive rate, false-positive rate, gaming behavior, and genuine improvement under classification
Reading fidelity high
Study strength medium
not reported
0.12
The data-driven counterfactual-explanation generators developed in the thesis are described as accurate and resource-efficient. Decision Quality positive Accuracy and resource efficiency of data-driven counterfactual-explanation generation
Reading fidelity high
Study strength low
not reported
0.06
The thesis uses an online human-subject experiment to investigate how parents perceive and respond to children's aggressive behavior toward embodied and disembodied AI-driven home devices. Ai Safety And Ethics mixed Parental perceptions and responses to children's aggressive behavior toward AI devices
Reading fidelity high
Study strength medium
not reported
0.12

Notes