The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References About 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

AI-supplied ideas shrink the variety of creative output and erase the non-native writers’ diversity edge, but using AI only to polish human ideas preserves collective creative breadth; simulating diverse humans with current LLMs fails to reproduce human-level variety unless fluency collapses.

Human diversity fuels collective creativity that large language models cannot simulate or sustain
Mengchen Dong, Hiromu Yakura · July 29, 2026
arxiv rct medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Mengchen Dong unresolved corpus identity
  2. Hiromu Yakura unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Mengchen Dong provider ID
  2. Hiromu Yakura provider ID
In a preregistered randomized experiment, AI-provided ideation homogenized metaphor outputs and eliminated non-native writers' collective diversity advantage, whereas using AI only to refine human-generated ideas preserved human collective diversity, and LLM-based persona simulations could not match human-level diversity without producing incoherent text.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Diverse human groups produce diverse ideas, the raw material of innovation. Generative AI challenges this engine twice over: everyday AI assistance may homogenize what diverse people create, and AI-simulated diversity may replace the people altogether. We tested both challenges in a preregistered creative metaphor experiment with native (L1) and non-native (L2) English writers, who wrote without AI, with AI-generated ideas (AI ideation), or with AI refining their own ideas (AI refinement). L2 writers contributed more collective diversity than L1 writers, with native-language ideation showing the most diverse pools. AI ideation compressed collective diversity for everyone and left the L2 advantage undetectable, whereas AI refinement preserved both. We then simulated the entire writer pool using personas built from participants' real backgrounds, three model families, native-language prompting, and elevated sampling temperatures. Every simulated pool fell below every human pool, and pushing models further induced diversity only through degenerate text. However, at the individual level, AI ideation raised writers' ratings, pitting private incentives against the collective good, except when L2 writers used their native language, which benefited both. Human diversity remains a valuable creative resource that current AI cannot simulate or sustain; the design of human-AI collaborative workflows determines whether it survives.

Summary

Main Finding

Human cultural-linguistic diversity produces greater collective creative variety than native-only groups, and current LLMs cannot simulate or sustainably replace that collective diversity. How AI is used matters: when models supply ideas (AI ideation) they substantially compress collective diversity and erase the advantage contributed by non-native (L2) writers; when models only refine human-generated ideas (AI refinement), collective diversity is largely preserved (though expression becomes more stylistically uniform). Attempts to simulate human pools via persona prompting, multilingual prompting, multiple model families, and high sampling temperature fail to reach human-level diversity except at the cost of fluency/coherence. At the same time, AI ideation tends to raise individual-level creativity ratings, creating a misalignment between private incentives and the collective good — with one important exception: L2 writers using their native language benefit both individually and collectively.

Key Points

  • Human diversity operationalized as native (L1) vs non-native (L2) English backgrounds:
    • Without AI, L2 writers produced a more diverse pool of metaphors than L1 writers (centroid-similarity difference Δ ≈ 0.024, p ≈ 0.003).
  • Workflow matters:
    • AI ideation (model supplies metaphor + explanation) compressed collective diversity for all writers (vs human-only: b = 0.036, p < 0.001) and erased the L2 advantage (L1 vs L2 no longer significant).
    • AI refinement (model polishes explanations of human-generated metaphors) preserved idea diversity; however, including AI-polished explanations in the embeddings showed stylistic homogenization.
  • Language of AI interaction matters:
    • L2 participants who interacted with AI in their native language produced the most diverse idea pools (directional, significant linear contrast across L1 → L2-English → L2-native: b = −0.025, p = 0.015).
  • LLM simulations of the full human pool fell short:
    • Simulated persona pools (GPT-4o mini, Claude Sonnet 4.6, Qwen3.5-27b), even when given native-language prompting and high temperature, were less diverse than every human pool (e.g., Claude recovered ~75.3% of human-only diversity).
    • Increasing temperature raised diversity but plateaued below human ranges; extreme temperature that approaches parity produced many incoherent/gibberish outputs (e.g., one high-temp config had ~14% gibberish).
    • Trade-off observed: model diversity increases come at the cost of fluency (higher perplexity).
  • Individual vs collective incentives:
    • AI ideation increased individual creativity ratings (novelty/surprise and aptness/explanation) — creating private incentives to use ideation workflows that nonetheless harm collective diversity.
    • Exception: L2 writers ideating in their native language realized both individual gains and higher collective diversity.

Data & Methods

  • Design:
    • Preregistered creative metaphor experiment with three conditions:
    • Human-only (no AI assistance)
    • AI ideation (LLM provides metaphors + short explanations that participants could adopt)
    • AI refinement (participants generate metaphors; LLM suggests refinements/explanations)
    • Participants: native (L1) and non-native (L2) English writers; L2 participants could choose AI interaction language (English or their native language). All final metaphors were submitted in English (translations provided when needed).
  • Outcomes and measurements:
    • Collective diversity: measured as mean cosine similarity of each metaphor to the centroid of metaphors in the same pool (higher mean similarity = less diverse).
    • Individual creative quality: independent human evaluators (N = 351) rated each metaphor on four items; combined into divergent (novelty, surprise) and convergent (aptness, explanation quality) subscales.
    • Statistical analysis: multilevel regression models (full specs in SI).
  • Simulation adversarial test:
    • For every human participant, a persona was built from background metadata (age, sex, ethnicity, education, first language, etc.).
    • Models used: GPT-4o mini; Claude Sonnet 4.6; Qwen3.5-27b.
    • Simulation levers: native-language prompting (L2 personas ideate in their native tongue then translate), sampling temperature sweep (0.5–2.0), and multilingual model selection.
    • Additional metrics: fluency measured by GPT-2 perplexity; format adherence and gibberish rates tracked.
  • Robustness:
    • Embeddings tested for metaphors alone and metaphors + explanations (showing expression homogenization when explanations included).

Implications for AI Economics

  • Human diversity is an irreplaceable economic input to collective creativity:
    • Firms and research organizations should continue investing in diverse human teams; LLM-driven “silicon sampling” cannot substitute for the breadth of ideas that diverse humans generate without degrading fluency.
  • Workflow design matters for organizational innovation outcomes:
    • Using LLMs as idea-generators (AI ideation) is likely to homogenize outputs across employees, diminishing the value of hiring diverse talent and potentially lowering the rate of radical/novel recombinations that drive breakthrough innovation.
    • Prefer workflows where humans ideate and LLMs serve as refiners/polishers (AI refinement), which preserve idea diversity while capturing productivity gains in expression.
  • Incentive misalignment and potential market failure:
    • Because AI ideation raises individual productivity/quality metrics, individuals (and managers optimizing short-term KPIs) may rationally prefer ideation workflows even though they reduce collective diversity — a classic private-versus-public-good tension that can produce suboptimal collective innovation.
    • Organizations should align incentives (promotion, compensation, performance metrics) to reward original human ideation and not just polished outputs.
  • Practical design and policy recommendations:
    • Product design: build interfaces that separate ideation and refinement stages, expose provenance (who ideated vs who refined), enable native-language ideation and high-quality translation pipelines, and surface diversity metrics to managers.
    • HR/organization: keep diverse hiring and multilingual teams; measure collective diversity (e.g., embedding-based dispersion metrics) as an innovation KPI; create rewards for original human contributions.
    • Platform/policy: support multilingual model development and translation tools that let L2 workers ideate in native languages; consider standards for documenting AI involvement to maintain credit/allocation of rewards.
  • Labor-market implications:
    • LLMs will shift task content (more polishing and scaling of expression) but are unlikely to eliminate roles that rely on culturally grounded, diverse conceptual repertoires (creative R&D, cross-cultural product design, high-novelty brainstorming).
    • Policymakers and firms should focus reskilling on ideation, cross-cultural competencies, and evaluation/curation skills that complement models rather than naively replacing diverse human inputs.
  • Metrics and monitoring:
    • Organizations should track both individual performance gains from AI and collective diversity outcomes; absence of such monitoring risks optimizing for short-term individual productivity at the expense of long-term innovation capacity.

Takeaway: AI can enhance individual productivity and expression, but it is not a costless substitute for diverse human contributors when the goal is maintaining a broad pool of ideas. The economic challenge is organizational and institutional: adopt AI workflows that protect where humans add unique value (ideation, culturally grounded framing) while using models to amplify expression, and realign incentives so private gains from AI do not erode the collective creativity that underpins long-run innovation.

Assessment

Paper Typerct Evidence Strengthmedium — High internal validity for the causal impact of where AI enters the writing workflow on collective diversity due to preregistration and randomized assignment; however, claims are constrained by an artificial laboratory creative task (metaphor generation), embedding-based outcome measures that depend on modelling choices, and limited external validity for broader economic outcomes. Methods Rigorhigh — Preregistered experimental design with random assignment, use of multilevel models, independent human evaluator panel for creativity ratings, and adversarial robustness checks (multiple model families, language modes, temperature sweeps); remaining concerns include reliance on embedding metrics for 'collective diversity' and potential sensitivity to choice of embedding model and hyperparameters. SampleHuman participants composed of native English (L1) and non-native English (L2) writers completing four creative metaphor tasks under one of three conditions; preregistered experiment with at least 228 L2 participants (84 of whom used native-language AI interaction in AI conditions); independent evaluator panel N=351 rated creativity; simulations used personas built from participants' background information and three model families (GPT-4o mini, Claude Sonnet 4.6, Qwen3.5-27b) with temperature swept from 0.5–2.0. Themeshuman_ai_collab innovation IdentificationPreregistered between-subjects randomized assignment of participants to three writing conditions (human-only, AI ideation, AI refinement), with multilevel regression models estimating condition and language-background effects; simulation comparisons use persona-based prompt engineering across multiple model families and sampling temperatures (adversarial, non-randomized). GeneralizabilityTask is creative metaphor generation (lab-style), which may not generalize to workplace innovation, complex real-world creative tasks, or productivity outcomes., Outcome measure (collective diversity) is operationalized via text embeddings and centroid similarity and is sensitive to embedding/model choice., Participant pool likely non-representative (online/experimental sample) and concentrated on English-language contexts., Simulations evaluate specific model families and prompt-engineering choices; results may not hold for future models or different prompting techniques., Short-term experimental setting does not capture long-run behavioral adaptation, team dynamics, or organizational incentives.

Claims (11)

ClaimDirectionOutcomeConfidence & EvidenceDetails
AI ideation produced the most homogeneous human metaphor pools, whereas AI refinement produced pools that were statistically indistinguishable from the human-only condition. Creativity negative Collective diversity of metaphor ideas
Reading fidelity high
Study strength high
AI ideation vs. human-only: b = 0.036; AI refinement vs. human-only: b = 0.007
1.0
Without AI assistance, non-native English writers produced a more diverse collective metaphor pool than native English writers. Creativity positive Collective diversity of metaphor ideas
Reading fidelity high
Study strength high
L1 vs. L2: M = 0.421 vs. 0.398, Δ = 0.024, p = 0.003
1.0
AI ideation compressed both L1 and L2 writers' metaphor pools and made the L2 collective-diversity advantage statistically undetectable. Creativity negative Collective diversity of metaphor ideas
Reading fidelity high
Study strength medium
L1 vs. L2 under AI ideation: M = 0.457 vs. 0.451, Δ = 0.006, p = 0.481
0.6
AI refinement preserved the L2 writers' collective-diversity advantage relative to L1 writers. Creativity positive Collective diversity of metaphor ideas
Reading fidelity high
Study strength medium
L1 vs. L2 under AI refinement: M = 0.426 vs. 0.409, Δ = 0.018, p = 0.032
0.6
Among L2 writers using AI, interaction in their native language was associated with greater collective metaphor diversity than interaction in English. Creativity positive Collective diversity of metaphor ideas
Reading fidelity high
Study strength medium
n=228
Linear contrast b = -0.025, p = 0.015; centroid similarity M = 0.442 vs. 0.427 vs. 0.421
0.6
AI refinement preserved diversity in the metaphors themselves but made the explanations more similar to one another. Creativity mixed Diversity of metaphor ideas and their written explanations
Reading fidelity high
Study strength medium
AI refinement with explanations vs. human-only: b = 0.046, p < 0.001
0.6
Every LLM-simulated metaphor pool was less diverse than every human pool, including the most homogenized human AI-ideation pool. Creativity negative Collective diversity of simulated metaphor pools
Reading fidelity high
Study strength medium
Claude Δ = 0.102, Qwen Δ = 0.111, GPT-4o mini Δ = 0.241 versus human AI ideation; all p < 0.001
0.6
The strongest reported simulation, Claude Sonnet 4.6, recovered only 75.3% of the diversity of the human-only pool. Creativity negative Relative collective diversity of simulated versus human metaphor pools
Reading fidelity high
Study strength medium
recovered only 75.3% of the diversity
0.6
Prompting simulated L2 personas in their native languages increased simulated diversity but did not close the gap with human pools. Creativity mixed Collective diversity of simulated metaphor pools
Reading fidelity high
Study strength medium
GPT-4o mini Δ = -0.092, p < 0.001; Claude native-language simulation vs. human AI ideation: M = 0.528 vs. 0.457, Δ = 0.071, p < 0.001
0.6
Increasing sampling temperature increased model output diversity only up to a plateau, and the configuration that exceeded the plateau did so with substantial gibberish and poor format adherence. Creativity mixed Simulated collective diversity, fluency, and format adherence
Reading fidelity high
Study strength medium
14.3% of output words were gibberish; format adherence below 75%
0.6
AI ideation increased individual creativity ratings, creating a conflict between private incentives and collective diversity. Creativity positive Individual perceived creativity of metaphors and explanations
Reading fidelity high
Study strength medium
n=351
0.6

Notes