0 cumulative citations
View corpus contextAn antagonistic AI agent that ‘plays devil’s advocate’ makes novice interaction designers rethink and materially revise their proposals more than simple self-reflection, and it spurs more conflictual perspectives and idea turnover than equivalent written guidance.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Generative AI agents are increasingly used in interaction design to facilitate ideation and offer critique, often following their own internal reasoning. These interactions tend to add design ideas and expand the design space. Our work explores an antagonistic role for design agents, prompting designers to engage with stakeholder tension. We built an AI agent inspired by adversarial design theory that enacts constructive conflict. We examine the agent's influence in a between-subjects experiment with 45 design students across three conditions: Self Reflection (unsupported review of the design proposal), Stepwise Guidance (written prompts that walk designers through a constructive-conflict framework), and Interactive Engagement (an AI agent that enacts the constructive-conflict framework interactively by synthesizing stakeholder pushback). The latter two conditions share the framework but differ in whether it is self-enacted or agent-enacted. Results show that, compared with Self Reflection, both the Stepwise Guidance and Interactive Engagement groups reported significantly higher self-reconsideration and made more improvements to their design proposals. Compared with Stepwise Guidance, the antagonistic agent introduced more conflictual perspectives, and participants in the Interactive Engagement condition generated and discarded more ideas. These findings suggest that agent-enacted constructive conflict can turn reconsideration into concrete design actions and deepen engagement with divergent stakeholder perspectives.
Summary
Main Finding
An antagonistic, LLM-powered “constructive conflict” agent that enacts stakeholder pushback (vs. no structure or written prompts) increases novice interaction designers’ self-reconsideration and produces more concrete changes to design proposals. Compared with just reading stepwise guidance, the agent surfaced more conflictual perspectives and led participants to generate and discard more ideas—turning reconsideration into actionable edits and deeper engagement with divergent stakeholder viewpoints.
Key Points
- Intervention types (between-subjects, N=45 design students):
- Self Reflection: unsupported review of one’s design.
- Stepwise Guidance: written stepwise prompts describing a constructive-conflict framework.
- Interactive Engagement: interactive adversarial agent that synthesizes stakeholder pushback (two rounds of 4 pushback points; participants tag/use frames to steer the agent).
- Main quantitative outcomes:
- Both Stepwise Guidance and Interactive Engagement led to significantly higher self-reconsideration and more improvements to proposals than Self Reflection.
- Interactive Engagement (agent) introduced more conflictual perspectives and, compared to Stepwise Guidance, increased idea generation and idea discarding—suggesting more active exploration and iteration.
- Stepwise Guidance alone encouraged reconsideration but tended to reduce new idea generation.
- Agent design highlights:
- Grounded in interviews with six public-sector designers; four design goals (DG1–DG4) centered on stakeholder grounding, designer agency, stress-testing commitments, and making trade-offs explicit.
- Implemented in Miro; agent passively collects multimodal context, generates first-round pushback, designers tag/steer (use Consensus or Command frames), then agent generates a second round focused on steered directions.
- Tone intentionally adversarial; retrieval-augmented with stored exemplars of constructive conflict.
- GPT-4.1 was used after pilot comparisons with Gemini 2.5 Pro.
- Measures and analysis:
- 45 proposals, 90 iterations, ~60 hours transcript; idea-units coded as added/edited/deleted (edits split into revisions vs enhancements).
- Think-aloud and end walkthroughs used to disambiguate new vs edited ideas.
- Statistical tests: Shapiro–Wilk for normality, Levene for variances, Kruskal–Wallis/ANOVA with Dunn’s or Tukey HSD post-hoc, effect sizes (ε2, ω2, Hedges’ g).
- Qualitative findings:
- Agent-enacted conflict made trade-offs explicit and prompted reframing of goals.
- Participants used tagging and frames to retain agency; the interactive back-and-forth often led to concrete changes in the second iteration.
- Limitations noted:
- Novice designers in lab task (redesign of a 311 civic-reporting site); short-duration sessions (~90 minutes) limit direct generalization to professional practice.
- Some Miro actions were researcher-mediated due to API limits (protocol to avoid bias).
- Unknown long-term impacts and effects with experienced practitioners.
Data & Methods
- Participants: 45 interaction/design students (mean age 25.08; mean ~4 years studying design/HCI), randomized into three conditions; study approved by IRB; ~90-minute remote sessions on Zoom using Miro.
- Task: two-iteration redesign of a local government non-emergency civic-reporting website (“311”).
- Agent implementation:
- Web stack: Next.js + Miro SDK/API, Firestore; Whisper for voice; GPT-4o for parsing visual content; GPT-4.1 for critique generation.
- Workflow: passive context collection during initial ideation → agent produces 4 pushback points → participant tags/frames (useful/not useful or custom tags; Consensus vs Command) → second round of 4 pushback points refined by participant steering.
- Pushback points pair a stakeholder perspective with a concrete challenge to the participant’s proposal; retrieval-augmented from a curated database of constructive-conflict examples.
- Measurement and coding:
- Idea-units identified and classified (added, deleted, edited [revision/enhancement]); think-aloud used to disambiguate.
- Post-survey scales for perceived challenge, satisfaction, self-efficacy, amount of design change, critical design thinking; additional agent-specific ratings for Interactive Engagement group.
- Qualitative coding: mixed deductive–inductive, affinity diagramming, team consensus to derive patterns.
- Analysis: normality tests, appropriate omnibus tests, corrected pairwise comparisons, and effect size reporting.
Implications for AI Economics
- Productivity and quality of design labor
- Agents that enact constructive conflict can function as complementary tools that increase the effective productivity of novice (and potentially mid-level) designers by turning deliberation into concrete edits. This suggests a pathway where AI augments design quality earlier in the lifecycle, possibly reducing costly downstream rework.
- The agent’s effect is not just additive suggestion; it changes iteration dynamics (more prototyping/idea turnover). Economic models of labor-AI complementarity should include agent-induced process improvements (not only time-savings).
- Human capital formation and training
- For novices, adversarial agents offer on-demand practice in negotiating stakeholder tensions—potentially accelerating skill accumulation. This has implications for the market value of junior designers and training investments: firms or educational institutions might adopt such agents to scale mentoring at lower marginal cost.
- Allocation of decision-making and bargaining power
- By surfacing diverse stakeholder perspectives and making trade-offs explicit, agent-enacted conflict could shift how public-sector or multi-stakeholder organizations negotiate priorities. Economically, this can change the distribution of influence among stakeholders and alter policy/design outcomes—raising questions about who controls the agent’s objectives and retrieval corpora.
- Product-market implications
- A design tool that reliably induces productive reconsideration could become a differentiator in design platforms or consultancy workflows, creating potential product-market value for AI systems that intentionally introduce friction rather than maximize agreement.
- Welfare and distributional considerations
- Benefits (better-explored solutions, fewer missed trade-offs) may accrue unevenly: organizations that can integrate and interpret adversarial outputs will gain more. There is potential for both welfare improvements (better public services) and harms (if agent adversarial framing amplifies polarizing stakeholder frames).
- Adoption risks and externalities
- Overuse or misuse of adversarial agents might create longer iteration cycles or unnecessary conflict in contexts where consensus is pragmatically required. There is also a risk of adversarial bias (agents systematically privileging certain stakeholder framings) — an externality for downstream users.
- If adversarial agents are perceived as intrusive or reduce trust, uptake could be limited despite productivity gains. Economic adoption models should account for perceived psychological costs and transaction frictions.
- Research and policy priorities for economists
- Measure downstream outcomes: beyond lab-change counts, evaluate long-run project quality, time-to-deploy, maintenance costs, and stakeholder satisfaction in field trials.
- Incorporate adversarial-AI into models of task automation: expand beyond “assist vs replace” to formalize how AI-mediated friction affects task complementarities, skill premiums, and wage dynamics.
- Distributional analysis: who benefits from improved design deliberation—citizens, agencies, vendors—and how agent design choices (training data, retrieval corpora, tone) shift welfare.
- Regulation and governance: because agents can materially shape public-sector design trade-offs, consider transparency requirements about agent objectives, provenance of example databases, and mechanisms to audit stakeholder representativeness.
- Practical takeaways for practitioners and economists
- Value is not only in agent correctness but in agent behavior design (adversarial framing, steerable interactions). Pricing and procurement of AI tools should account for behavioral affordances that affect process outcomes.
- Pilot deployments in real-world civic design projects are needed to quantify net benefits and potential negative externalities before scaling.
- When modeling AI’s economic impact on creative or deliberative tasks, include process-level effects (rate of iteration, diversity of options considered, and negotiation outcomes), not only static productivity gains.
Summary: This paper provides experimental evidence that interactive, adversarial AI agents—designed to enact constructive conflict—change how novice designers think and iterate, producing more concrete proposal changes and richer engagement with stakeholder tension. For AI economics, the key implication is that AI can shift not only productivity but also decision-process dynamics and human capital formation; economic analyses and policies should account for these process-level, distributional, and governance effects.
Assessment
Claims (9)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Compared with the Self Reflection condition, participants in both the Stepwise Guidance and Interactive Engagement conditions reported significantly higher self-reconsideration. Decision Quality | positive | Self-reconsideration of design decisions |
Reading fidelity
high
Study strength
medium
|
n=45
|
| Compared with Self Reflection, participants in both the Stepwise Guidance and Interactive Engagement conditions made more improvements to their design proposals. Output Quality | positive | Improvements and iterative changes to design proposals |
Reading fidelity
high
Study strength
medium
|
n=45
|
| Compared with Stepwise Guidance, the Interactive Engagement condition introduced more conflictual perspectives. Creativity | positive | Introduction of conflictual stakeholder perspectives |
Reading fidelity
high
Study strength
medium
|
n=45
|
| Compared with Stepwise Guidance, participants in the Interactive Engagement condition generated and discarded more ideas. Creativity | positive | Number of ideas generated and discarded between design iterations |
Reading fidelity
high
Study strength
medium
|
n=45
|
| Interactive Engagement led students to report more reconsideration of design decisions, more shifts in design thinking, and more iterative edits than Self Reflection. Decision Quality | positive | Reconsideration, shifts in design thinking, and iterative edits |
Reading fidelity
high
Study strength
medium
|
n=45
|
| Stepwise Guidance encouraged reconsideration but did not shift overall thinking compared with Self Reflection. Decision Quality | mixed | Reconsideration and overall shifts in design thinking |
Reading fidelity
high
Study strength
medium
|
n=45
|
| Stepwise Guidance reduced new idea generation relative to the other conditions. Creativity | negative | Generation of new design ideas |
Reading fidelity
high
Study strength
medium
|
n=45
|
| The agent design was informed by four recurring priorities identified through formative interviews with six experienced public-sector designers. Ai Safety And Ethics | positive | Agent design requirements for constructive-conflict interaction |
Reading fidelity
high
Study strength
low
|
n=6
|
| In preliminary evaluation, GPT-4.1 and Gemini 2.5 Pro produced comparable critique outputs that participants judged reasonable, with no model consistently outperforming the other. Ai Safety And Ethics | null_result | Perceived reasonableness and usefulness of AI-generated critique |
Reading fidelity
high
Study strength
low
|
n=3
|