0 cumulative citations
View corpus contextHybrid human–AI teams outperform either humans or AI alone on a controlled creative task, delivering higher accuracy without sacrificing diversity; both parties change behavior when partnered, and heterogeneous AI collaboration partially—but not fully—reproduces the benefit.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Generative AI is increasingly transforming creativity into a hybrid human-artificial process, but its impact on the quality and diversity of creative output remains unclear. We study collective creativity using a controlled word-guessing task that balances open-endedness with an objective measure of task performance. Participants attempt to infer a hidden target word, scored based on the semantic similarity of their guesses to the target, while also observing the best guess from previous players. We compare performance and outcome diversity across human-only, AI-only, and hybrid human-AI groups. Hybrid groups achieve the highest performance while preserving high diversity of guesses. Within hybrid groups, both humans and AI agents systematically adjust their strategies relative to single-agent conditions, suggesting higher-order interaction effects, whereby agents adapt to each other's presence. Although some performance benefits can be reproduced through collaboration between heterogeneous AI systems, human-AI collaboration remains superior, underscoring complementary roles in collective creativity.
Summary
Main Finding
Hybrid human–AI groups outperform homogeneous groups (human-only or AI-only) on a controlled collective creative search task. Benefits arise from complementary strategies and mutual adaptation: humans provide broad exploratory signals while AI exploits promising regions; AI behavior becomes more diverse and higher-quality when exposed to humans. Some synergy can be reproduced by mixing different LLMs, but human–AI hybrids remain superior.
Key Points
- Task and result summary:
- Participants iteratively guess a hidden target word and receive a semantic-similarity score; the best guess from prior rounds is shown as a hint.
- Four conditions: Human Social (many humans each doing one round), Human Asocial (single human does all rounds), AI-only (Gemini 2.5 agents), Human–AI Hybrid (rounds randomly assigned to human or Gemini).
- Hybrid condition achieved the highest average maximum similarity across rounds; AI-only performed worst.
- Mechanism:
- Humans tend to explore (high lexical diversity); AI tends to exploit (narrow clusters).
- In hybrid groups, AI increases lexical diversity and search quality; humans make more unique guesses when paired with AI.
- Mutual adaptation (second-order effects) — behavior change of both agent types driven by observing each other’s outputs — underlies gains, not mere additive effects.
- Heterogeneity vs. agent type:
- A hybrid composed of two different LLMs (Gemini 2.5 + GPT-5.1) produced some synergy but overall performance remained below human–AI hybrids, suggesting unique value from human cognition.
- Robustness:
- Effects hold across control studies varying AI systems, communication channels, and decoding temperatures.
- Strategy annotation:
- An LLM (Claude Sonnet 4) classified rounds into Exploit/Directed Explore/Undirected Explore/Mixed; humans adapt strategies to hint quality, AI-only does not.
- Key statistics & scale:
- Human participants N = 503 (US native English); total games = 200 (4 batches × 50 games); 100 guesses per game (10 rounds × 10 turns).
- AI queries: ~165k to Gemini 2.5; 10k to GPT‑5.1 for controls; significance levels reported p < .001 for main hybrid advantages; AI performance improved in hybrid (p < .001, d = 0.618).
Data & Methods
- Task design:
- Word-guessing game inspired by Semantle. Ten target words covering a difficulty range (e.g., harbor, satellite, metamorphosis).
- Each game: 10 rounds × 10 guesses = 100 guesses. Best guess per round passed onward as hint.
- Scoring and embeddings:
- Similarity score = 201.69 × cosine similarity of Word2Vec embeddings between guess and target.
- Vocabulary filtered to ~663k common English tokens; absent words scored zero.
- Experimental conditions:
- Human Social: different human per round (N = 210 participants used across rounds).
- Human Asocial: single human completes all rounds for a game (N = 179).
- AI-only: Gemini 2.5 Flash agents simulating rounds.
- Hybrid: rounds randomly assigned to humans or Gemini 2.5 (≈5 human, 5 AI rounds/game).
- Analysis:
- Semantic trajectories visualized via Word2Vec + UMAP.
- Performance metric: average of per-round maxima (best guess scores).
- Diversity: lexical diversity = proportion of unique guesses within a game.
- Strategy labeling: LLM-as-judge pipeline (Claude Sonnet 4) induced seven strategy types, collapsed into four categories; per-round soft labels then discretized.
- Statistical testing with corrections (Bonferroni) and bootstrap CIs for trends.
- Controls:
- Hybrid AI (Gemini+GPT‑5.1), different AI temperatures, alternate communication channels; overall robustness checks performed.
Implications for AI Economics
- Complementarity and returns to mixing inputs:
- Evidence of positive complementarities between human and AI contributors suggests higher marginal returns to mixed teams in creative search tasks than to homogeneous teams. Economic models of productivity should incorporate non‑linear gains from heterogeneity and interaction effects (second-order adaptation).
- Labor and task allocation:
- Optimal division of labor may assign exploratory, open-ended search to humans (or human-guided processes) and targeted exploitation to AI. Firms can improve R&D and ideation output by designing workflows that routinize interplay (humans seed exploration; AI refines).
- Platform and market design:
- Platforms that combine human crowdsourcing with AI agents (and dynamically route hints/results between them) can increase collective output. Marketplace pricing and contracting should value not only per-unit contribution but the system-level complementarities generated by mixed participation.
- Competition among AI models vs. human–AI mixes:
- Mixing different LLMs yields some gains, but human inputs deliver unique benefits. This implies that competition solely among AI providers may not substitute for human capital — investments in human-AI integration (tools, interfaces, incentives) remain valuable.
- Innovation policy and public goods:
- Human–AI hybrids reduce risks of premature homogenization relative to AI-only workflows, potentially preserving diversity of ideas important for long-run innovation. Policy and procurement should encourage hybrid architectures when public R&D or creative discovery is a goal.
- Second-order effects and externalities:
- Deploying AI at scale can change behavior of human collaborators (and vice versa). Regulatory assessment of AI’s economic impacts must account for these endogenous behavioral responses, not only one-off productivity estimates.
- Cost–benefit and scaling considerations:
- The study used substantial API calls (cost-bearing). Economic deployment requires weighing API/model costs vs. increased output; the hybrid benefit implies potential efficiency gains but also additional coordination costs (platform design, human recruitment, latency).
- Research & evaluation recommendations:
- Empirical economic work should measure system-level outcomes across varying mixes of human and AI agents, network topologies, and incentive schemes to estimate optimal team compositions and pricing. Structural models could quantify welfare gains from hybrid systems and guide investment and regulation.
Limitations to keep in mind - Task is artificial (word-guessing with Word2Vec scoring); external validity to complex real-world creative tasks (e.g., scientific discovery, product design) needs testing. - Models used (Gemini 2.5, GPT‑5.1) and a single scoring embedding (Word2Vec) may bias results; behaviors may differ with other architectures or tasks. - Human sample: US-based, native English speakers; cross-cultural and domain-expert settings could alter complementarities. - The experiment did not disclose AI participation to humans; transparency effects (trust, effort) are an important economic variable for future work.
Suggested follow-ups for AI economics researchers - Cost–benefit analysis of hybrid teams in domain-specific R&D with monetized outcomes. - Structural models estimating production functions with human–AI interaction terms (second-order effects). - Experiments varying network structure, incentive schemes (payoffs for uniqueness vs. accuracy), and disclosure of AI involvement to study equilibrium behavior and welfare.
Assessment
Claims (5)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Hybrid human-AI groups achieve the highest performance on the word-guessing task compared to human-only and AI-only groups. Output Quality | positive | semantic similarity of guesses to the hidden target (task performance) |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Hybrid groups preserve a high diversity of guesses (outcomes) while achieving superior performance. Creativity | positive | diversity of guesses (semantic/lexical diversity of outputs) |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Within hybrid groups, both human and AI agents systematically adjust their strategies relative to single-agent (human-only or AI-only) conditions, indicating higher-order interaction effects where agents adapt to each other's presence. Team Performance | mixed | agents' strategy/behavioral adjustments (qualitative/systematic changes in guessing strategy) |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Some of the performance benefits observed in hybrid groups can be reproduced through collaboration between heterogeneous AI systems. Output Quality | positive | task performance (semantic similarity of guesses) when heterogeneous AI systems collaborate |
Reading fidelity
medium
Study strength
medium
|
not reported
|
| Despite some gains from heterogeneous-AI collaboration, human-AI collaboration remains superior to AI-only collaboration, highlighting complementary roles of humans and AI in collective creativity. Creativity | positive | relative task performance and creative output quality across group types |
Reading fidelity
high
Study strength
medium
|
not reported
|