The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Hybrid human–AI teams outperform either humans or AI alone on a controlled creative task, delivering higher accuracy without sacrificing diversity; both parties change behavior when partnered, and heterogeneous AI collaboration partially—but not fully—reproduces the benefit.

Human-AI Synergy Supports Collective Creative Search
Chenyi Li, Raja Marjieh, Haoyu Hu, Mark Steyvers, Katherine M. Collins, Ilia Sucholutsky, Nori Jacoby · February 10, 2026
arxiv rct high evidence 8/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Chenyi Li unresolved corpus identity
  2. Raja Marjieh unresolved corpus identity
  3. Haoyu Hu unresolved corpus identity
  4. Mark Steyvers unresolved corpus identity
  5. Katherine M. Collins unresolved corpus identity
  6. Ilia Sucholutsky unresolved corpus identity
  7. Nori Jacoby unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Chenyi Li provider ID
  2. Raja Marjieh provider ID
  3. Haoyu Hu provider ID
  4. M. Steyvers provider ID
  5. K. M. Collins provider ID
  6. Ilia Sucholutsky provider ID
  7. Nori Jacoby provider ID
In a controlled word-guessing experiment, hybrid human-AI groups achieved the highest accuracy while maintaining high diversity of guesses, with both humans and AI systematically adapting strategies when paired.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Generative AI is increasingly transforming creativity into a hybrid human-artificial process, but its impact on the quality and diversity of creative output remains unclear. We study collective creativity using a controlled word-guessing task that balances open-endedness with an objective measure of task performance. Participants attempt to infer a hidden target word, scored based on the semantic similarity of their guesses to the target, while also observing the best guess from previous players. We compare performance and outcome diversity across human-only, AI-only, and hybrid human-AI groups. Hybrid groups achieve the highest performance while preserving high diversity of guesses. Within hybrid groups, both humans and AI agents systematically adjust their strategies relative to single-agent conditions, suggesting higher-order interaction effects, whereby agents adapt to each other's presence. Although some performance benefits can be reproduced through collaboration between heterogeneous AI systems, human-AI collaboration remains superior, underscoring complementary roles in collective creativity.

Summary

Main Finding

Hybrid human–AI groups outperform homogeneous groups (human-only or AI-only) on a controlled collective creative search task. Benefits arise from complementary strategies and mutual adaptation: humans provide broad exploratory signals while AI exploits promising regions; AI behavior becomes more diverse and higher-quality when exposed to humans. Some synergy can be reproduced by mixing different LLMs, but human–AI hybrids remain superior.

Key Points

  • Task and result summary:
    • Participants iteratively guess a hidden target word and receive a semantic-similarity score; the best guess from prior rounds is shown as a hint.
    • Four conditions: Human Social (many humans each doing one round), Human Asocial (single human does all rounds), AI-only (Gemini 2.5 agents), Human–AI Hybrid (rounds randomly assigned to human or Gemini).
    • Hybrid condition achieved the highest average maximum similarity across rounds; AI-only performed worst.
  • Mechanism:
    • Humans tend to explore (high lexical diversity); AI tends to exploit (narrow clusters).
    • In hybrid groups, AI increases lexical diversity and search quality; humans make more unique guesses when paired with AI.
    • Mutual adaptation (second-order effects) — behavior change of both agent types driven by observing each other’s outputs — underlies gains, not mere additive effects.
  • Heterogeneity vs. agent type:
    • A hybrid composed of two different LLMs (Gemini 2.5 + GPT-5.1) produced some synergy but overall performance remained below human–AI hybrids, suggesting unique value from human cognition.
  • Robustness:
    • Effects hold across control studies varying AI systems, communication channels, and decoding temperatures.
  • Strategy annotation:
    • An LLM (Claude Sonnet 4) classified rounds into Exploit/Directed Explore/Undirected Explore/Mixed; humans adapt strategies to hint quality, AI-only does not.
  • Key statistics & scale:
    • Human participants N = 503 (US native English); total games = 200 (4 batches × 50 games); 100 guesses per game (10 rounds × 10 turns).
    • AI queries: ~165k to Gemini 2.5; 10k to GPT‑5.1 for controls; significance levels reported p < .001 for main hybrid advantages; AI performance improved in hybrid (p < .001, d = 0.618).

Data & Methods

  • Task design:
    • Word-guessing game inspired by Semantle. Ten target words covering a difficulty range (e.g., harbor, satellite, metamorphosis).
    • Each game: 10 rounds × 10 guesses = 100 guesses. Best guess per round passed onward as hint.
  • Scoring and embeddings:
    • Similarity score = 201.69 × cosine similarity of Word2Vec embeddings between guess and target.
    • Vocabulary filtered to ~663k common English tokens; absent words scored zero.
  • Experimental conditions:
    • Human Social: different human per round (N = 210 participants used across rounds).
    • Human Asocial: single human completes all rounds for a game (N = 179).
    • AI-only: Gemini 2.5 Flash agents simulating rounds.
    • Hybrid: rounds randomly assigned to humans or Gemini 2.5 (≈5 human, 5 AI rounds/game).
  • Analysis:
    • Semantic trajectories visualized via Word2Vec + UMAP.
    • Performance metric: average of per-round maxima (best guess scores).
    • Diversity: lexical diversity = proportion of unique guesses within a game.
    • Strategy labeling: LLM-as-judge pipeline (Claude Sonnet 4) induced seven strategy types, collapsed into four categories; per-round soft labels then discretized.
    • Statistical testing with corrections (Bonferroni) and bootstrap CIs for trends.
  • Controls:
    • Hybrid AI (Gemini+GPT‑5.1), different AI temperatures, alternate communication channels; overall robustness checks performed.

Implications for AI Economics

  • Complementarity and returns to mixing inputs:
    • Evidence of positive complementarities between human and AI contributors suggests higher marginal returns to mixed teams in creative search tasks than to homogeneous teams. Economic models of productivity should incorporate non‑linear gains from heterogeneity and interaction effects (second-order adaptation).
  • Labor and task allocation:
    • Optimal division of labor may assign exploratory, open-ended search to humans (or human-guided processes) and targeted exploitation to AI. Firms can improve R&D and ideation output by designing workflows that routinize interplay (humans seed exploration; AI refines).
  • Platform and market design:
    • Platforms that combine human crowdsourcing with AI agents (and dynamically route hints/results between them) can increase collective output. Marketplace pricing and contracting should value not only per-unit contribution but the system-level complementarities generated by mixed participation.
  • Competition among AI models vs. human–AI mixes:
    • Mixing different LLMs yields some gains, but human inputs deliver unique benefits. This implies that competition solely among AI providers may not substitute for human capital — investments in human-AI integration (tools, interfaces, incentives) remain valuable.
  • Innovation policy and public goods:
    • Human–AI hybrids reduce risks of premature homogenization relative to AI-only workflows, potentially preserving diversity of ideas important for long-run innovation. Policy and procurement should encourage hybrid architectures when public R&D or creative discovery is a goal.
  • Second-order effects and externalities:
    • Deploying AI at scale can change behavior of human collaborators (and vice versa). Regulatory assessment of AI’s economic impacts must account for these endogenous behavioral responses, not only one-off productivity estimates.
  • Cost–benefit and scaling considerations:
    • The study used substantial API calls (cost-bearing). Economic deployment requires weighing API/model costs vs. increased output; the hybrid benefit implies potential efficiency gains but also additional coordination costs (platform design, human recruitment, latency).
  • Research & evaluation recommendations:
    • Empirical economic work should measure system-level outcomes across varying mixes of human and AI agents, network topologies, and incentive schemes to estimate optimal team compositions and pricing. Structural models could quantify welfare gains from hybrid systems and guide investment and regulation.

Limitations to keep in mind - Task is artificial (word-guessing with Word2Vec scoring); external validity to complex real-world creative tasks (e.g., scientific discovery, product design) needs testing. - Models used (Gemini 2.5, GPT‑5.1) and a single scoring embedding (Word2Vec) may bias results; behaviors may differ with other architectures or tasks. - Human sample: US-based, native English speakers; cross-cultural and domain-expert settings could alter complementarities. - The experiment did not disclose AI participation to humans; transparency effects (trust, effort) are an important economic variable for future work.

Suggested follow-ups for AI economics researchers - Cost–benefit analysis of hybrid teams in domain-specific R&D with monetized outcomes. - Structural models estimating production functions with human–AI interaction terms (second-order effects). - Experiments varying network structure, incentive schemes (payoffs for uniqueness vs. accuracy), and disclosure of AI involvement to study equilibrium behavior and welfare.

Assessment

Paper Typerct Evidence Strengthhigh — The study uses a controlled, randomized design with objective performance metrics, allowing clean internal causal inference about the effect of group composition (human, AI, hybrid) on task performance and diversity; however, claims about broader creative work are still limited by task abstraction and choice of models/participants. Methods Rigormedium — Strong internal design (randomization, objective scoring, multiple conditions) and analysis of adaptation effects, but key methodological details are not provided in the summary (sample size, pre-registration, randomization checks, model versions, incentives, and robustness checks), and the artificial word-guessing task may omit important real-world complexities. SampleHuman participants (recruited for the experiment) and AI agents (one or more generative models) engaged in a controlled word-guessing task where each player attempts to infer a hidden target word and sees prior best guesses; outcomes are measured by semantic-similarity between guesses and the target and by diversity metrics of guesses across rounds; exact sample sizes, recruitment platform, demographic composition, and model specifications are not specified in the summary. Themeshuman_ai_collab productivity innovation org_design IdentificationRandomized controlled experiment: participants (and AI agents) were assigned to human-only, AI-only, or hybrid group conditions and performed the same controlled word-guessing task; causal effects are identified by comparing objective outcome measures (semantic-similarity scores and diversity metrics) across randomly assigned conditions while holding the task and information structure constant. GeneralizabilityArtificial lab task (word-guessing) may not map to complex, open-ended creative work in firms or markets, Semantic-similarity score is an imperfect proxy for real-world creative value or economic productivity, Unknown representativeness of human participants (e.g., crowdworkers vs professionals) limits external validity, Results depend on specific AI model(s) and training/data; other models may behave differently, Short-term interactions in the experiment may not capture longer-term adaptation, learning, or organizational processes

Claims (5)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Hybrid human-AI groups achieve the highest performance on the word-guessing task compared to human-only and AI-only groups. Output Quality positive semantic similarity of guesses to the hidden target (task performance)
Reading fidelity high
Study strength medium
not reported
0.6
Hybrid groups preserve a high diversity of guesses (outcomes) while achieving superior performance. Creativity positive diversity of guesses (semantic/lexical diversity of outputs)
Reading fidelity high
Study strength medium
not reported
0.6
Within hybrid groups, both human and AI agents systematically adjust their strategies relative to single-agent (human-only or AI-only) conditions, indicating higher-order interaction effects where agents adapt to each other's presence. Team Performance mixed agents' strategy/behavioral adjustments (qualitative/systematic changes in guessing strategy)
Reading fidelity high
Study strength medium
not reported
0.6
Some of the performance benefits observed in hybrid groups can be reproduced through collaboration between heterogeneous AI systems. Output Quality positive task performance (semantic similarity of guesses) when heterogeneous AI systems collaborate
Reading fidelity medium
Study strength medium
not reported
0.36
Despite some gains from heterogeneous-AI collaboration, human-AI collaboration remains superior to AI-only collaboration, highlighting complementary roles of humans and AI in collective creativity. Creativity positive relative task performance and creative output quality across group types
Reading fidelity high
Study strength medium
not reported
0.6

Notes