Designing human–AI interfaces reshapes future capabilities: workflows that boost short‑term performance can erode independent skills or diversity, reversing their advantage later, so optimal architectures should be chosen with the state they create and the value of adaptation in mind.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
No provider observation is available for this paper.
Missing data, not a zero citation count.
Human-AI interaction can improve current performance while changing the capabilities and relationships on which future performance depends. We develop adaptive complementarity, a framework for choosing interaction architecture with these state consequences in view. Access, information exposure, task allocation, timing, and communication can alter which arrangement will be valuable later; their settings can often be reset faster than the capabilities, search patterns, or conventions they create. Three mechanisms organize the argument: information exposure and collective search, delegation and capability evolution, and strategic interdependence and information governance. Their integration yields cross-mechanism implications, including conditions under which a loss of expertise heterogeneity increases the information differentiation required to preserve independent search. We distinguish strong human-AI complementarity from advantage over another workflow and from advantage over an evolving reference policy. A knowledge-coverage illustration shows how different interaction histories can reverse current workflow rankings even at equal human competence. It also separates that result from the incremental value of state feedback, which can be small when a well-chosen stable workflow anticipates learning. The framework directs evaluation toward the states present interaction creates, their consequences for later architectural fit, and the conditions under which observing and responding to them is worthwhile.
Summary
Main Finding
The paper introduces "adaptive complementarity," a decision framework showing that choices about human–AI interaction architecture (who has access, what information is exposed, task allocation, timing, communication) shape not only immediate performance but also the state—human skills, problem structure, and relations—that determines future performance. Because architecture often can be changed faster than the latent states it creates, designers must weigh immediate gains against the continuation value of the successor state. Short-term improvements can undermine long-term fit; conversely, short-term losses can be optimal if they create valuable future states. The framework distinguishes three distinct evaluation comparisons (strong complementarity, architectural/workflow advantage, and trajectory/policy advantage) and identifies three mechanisms through which architecture produces state: (1) information exposure and collective search, (2) delegation and capability evolution, and (3) strategic interdependence and information governance.
Key Points
-
Architecture versus state
- Interaction architecture (It) is the designer-controllable organization of access, information exposure, task allocation, decision rights, timing, and communication.
- System state Xt = (At, Pt, Zt) is what interaction produces: agent capabilities (At), problem structure (Pt: epistemic and strategic features), and relational state (Zt: reliance, correlated search, conventions).
- Architecture can be changed more quickly than the latent state it creates; prior interactions leave persistent state effects.
-
Three evaluation comparisons (separate questions)
- Strong complementarity (Ct): Does the joint system outperform the better standalone constituent (human or AI) at the current state?
- Architectural/workflow advantage (Mt): Does this architecture improve on a specified alternative workflow at the current state?
- Trajectory/policy advantage (Atraj): Does the policy (sequence of architectural choices) produce more total discounted value over time from a common initial condition than alternative policies, each of which generates its own state trajectory?
- These comparisons can diverge: an architecture can improve over a status quo while not producing strong complementarity; a widening gap can accompany falling absolute system value if benchmarks deteriorate faster.
-
Dynamic decision rule and path dependence
- The paper frames architecture choice as a dynamic programming problem: choose It to maximize expected discounted value Jκ(π|x0,i−1), net of switching costs κ.
- The choice compares immediate payoff difference (∆) plus discounted continuation value difference (γL) to switching costs: choose I1 over I0 if ∆ + γL > switching cost.
- L captures how the chosen architecture influences future state and hence future value—this is the formal place where path dependence and state-creation matter.
-
Three mechanisms that produce adaptive complementarity
- Information exposure and collective search: shared AI suggestions can improve immediate outcomes but reduce independent exploration, which matters when collective search for alternatives is valuable (e.g., design, creativity).
- Delegation and capability evolution: repeated delegation can erode individual execution skills (practice declines) while increasing supervisory competence; this changes which architectures complement human capabilities later.
-
Strategic interdependence and information governance: who knows what and when affects coordination, competition, and externalities; information sharing can create beneficial conventions or harmful coordination that reduces welfare for others.
-
Illustrative results and nuances
- An exact knowledge-coverage illustration shows that different interaction histories can reverse current workflow rankings even when human competence is equal across histories.
- The paper separates the effect of changing workflow history from the incremental value of observing state and adapting: sometimes a well-chosen static workflow anticipates learning sufficiently that ongoing monitoring adds little value.
- Empirically, many experiments show short-term gains but longer-term tradeoffs (e.g., tutoring vs. generic assistance in education; reduced variety in AI-assisted creative tasks).
Data & Methods
- Conceptual and theoretical framework:
- Formal state representation Xt = (At, Pt, Zt); architecture It; dynamics Xt+1 = F(Xt, It, εt+1).
- Objective Jκ(π|x0,i−1) = expected discounted sum of stakeholder value w(Xt, It) minus switching costs, evaluated under a behavioral-response model and feasible governance constraints.
- Definitions for Ct (constituent complementarity), Mt (architectural advantage), and Atraj (trajectory advantage), and an exact dynamic comparison condition: choose I1 over I0 if immediate gain plus discounted continuation value exceeds switching cost.
- Mechanism analysis:
- Three mechanism-driven analytical sections linking problem features to architecture choices and state transitions.
- Formal appendices include governance formulation, derivations, and a complete computational specification.
- Illustrations and computational examples:
- An "exact knowledge-coverage" analytical illustration demonstrates how history-dependent effects can reverse rankings of workflows.
- Computational specification (in appendix) provides numeric examples and clarifies when the incremental value of state feedback is large versus small.
- Empirical guidance (proposed designs):
- Distinguishes empirical tests needed for the three claims: state creation, history-dependent fit, and operational value of feedback.
- Recommends withdrawal tests (measure what people can do after a specific interaction history) and trajectory comparisons (randomize policies that each generate their own learning histories).
- Note on data: the paper is primarily theoretical and conceptual; it synthesizes prior experimental findings (e.g., Bastani et al. 2025; Doshi & Hauser 2024; Bansal et al. 2021; others cited) and proposes empirical strategies rather than reporting a new large-scale empirical dataset.
Implications for AI Economics
-
Evaluation and measurement
- Researchers and practitioners should distinguish short-run assisted performance from long-run effects on human capabilities and problem structure; standard cross-sectional A/B evaluations may miss path-dependent losses or gains.
- Comparative metrics should state the reference clearly (human-only, AI-only, alternative workflow, or policy trajectory), because the same observed numbers can imply different welfare conclusions when benchmarks evolve.
- Longitudinal and trajectory-based experimental designs (randomized interaction policies, withdrawal tests) are essential to estimate the continuation value L and the operational value of monitoring/adaptation.
-
Design and organizational choice
- Firms should treat interaction architecture as a strategic choice with dynamic tradeoffs: immediate productivity gains from centralizing AI suggestions or delegating tasks may reduce exploratory diversity or execution skills, changing future returns to AI deployment.
- Switching costs and governance constraints make inaction or simpler fixed architectures rational in many settings; optimality of ongoing adaptation depends on the size of L relative to monitoring and switching costs.
- Anticipatory design matters: choosing an architecture that deliberately supports desired long-run capabilities can reduce the need for costly ongoing adaptation.
-
Labor-market and skill dynamics
- Repeated delegation can generate endogenous skill erosion (or skewed skill evolution toward supervisory/verification roles), with implications for wages, task allocation, and inequality across workers and firms.
- If architectures converge people toward similar representations (reduced heterogeneity), markets may need to create stronger information differentiation (more tailored signals) to preserve independent search and innovation.
-
Competition, externalities, and regulation
- Information governance choices (who sees what) have strategic externalities. Private incentives to share or hide AI insights can create coordination that harms consumers or third parties; regulators should consider these second-order effects, not just immediate accuracy gains.
- Platform-level architecture (e.g., defaults for exposure, sharing among users, or multi-agent communication) can shape industry-wide capability distributions, affecting innovation and competition.
-
Investment and policy recommendations
- Cost–benefit analysis of AI deployment should include long-run state impacts and switching/monitoring costs; policies that evaluate only assisted outputs risk endorsing architectures that degrade human skill or social welfare over time.
- Encourage randomized, longitudinal evaluations before scaling architectures that centralize decision-making or remove practice opportunities; support mechanisms that maintain heterogeneity of search (e.g., differentiated prompts, parallel independent exploration).
In short, Heydari’s adaptive complementarity reframes human–AI system design as a dynamic, state-dependent choice problem: architects must trade off immediate gains against the structural consequences they leave behind. For AI economics, this implies rethinking evaluation, organizational design, labor dynamics, and regulatory standards to account for path dependence and the continuation value of interaction-induced states.
Assessment
Claims (12)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Students using a conventional GPT-4 interface performed better while AI assistance was available but subsequently performed worse without it than students in a control group. Skill Obsolescence | mixed | Performance with AI assistance and subsequent unaided performance |
Reading fidelity
high
Study strength
medium
|
not reported
|
| A tutoring interface using the same underlying model, with teacher-informed guidance, largely mitigated the subsequent learning penalty associated with conventional GPT-4 assistance. Skill Acquisition | positive | Learning and subsequent unaided performance |
Reading fidelity
high
Study strength
medium
|
not reported
|
| AI suggestions improved the evaluations of individual stories but made stories produced by different writers more similar to one another. Creativity | mixed | Individual story quality and cross-writer story diversity |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Exposure to AI-generated design examples can increase fixation on those examples and limit exploration of alternatives. Creativity | negative | Exploration of alternative design ideas |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Adding explanations to AI recommendations increased acceptance of those recommendations without improving people’s ability to distinguish correct recommendations from incorrect ones. Decision Quality | mixed | Recommendation acceptance and discrimination between correct and incorrect recommendations |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Interfaces requiring more deliberate engagement reduced overreliance on AI recommendations but required additional effort from users. Decision Quality | mixed | AI overreliance and user effort |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Across 106 experiments, human–AI combinations substantially outperformed humans alone on average but underperformed the better standalone participant on average, with considerable variation across tasks. Decision Quality | mixed | Performance of human–AI combinations relative to humans alone and the better standalone participant |
Reading fidelity
high
Study strength
high
|
n=106
substantially outperformed humans alone; underperformed the better standalone participant on average
|
| Common AI suggestions can improve an individual’s next decision while reducing the number of independent search trajectories and imposing a collective search cost. Decision Quality | mixed | Individual decision performance and collective diversity of search |
Reading fidelity
high
Study strength
medium
|
not reported
|
| An interaction history can reverse the ranking of alternative workflows even when human competence is equal across histories. Task Allocation | mixed | Relative workflow performance after different interaction histories |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The value of responding to evolving system states depends on both immediate performance differences and the future value of the states created, net of switching costs. Organizational Efficiency | mixed | Expected discounted system performance from alternative interaction architectures |
Reading fidelity
high
Study strength
medium
|
not reported
|
| A workflow’s apparent improvement relative to a human benchmark can increase even while assisted system performance declines if the human benchmark deteriorates faster. Output Quality | mixed | Performance gap between an assisted workflow and evolving benchmarks |
Reading fidelity
high
Study strength
high
|
assisted performance falls from 90 to 85 while unaided human performance falls from 80 to 70; gap rises from 10 to 15
|
| The MASAI mammography-screening trial addressed whether an AI-supported workflow improved on established screening practice, rather than whether repeated AI use changed radiologists’ independent reading capability. Task Completion Time | positive | Radiologist reading workload and clinical screening outcomes |
Reading fidelity
high
Study strength
medium
|
substantial reduction in reading workload
|