The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Off‑the‑shelf language models tend to generate narrower idea pools than groups of humans because early outputs and a unified knowledge distribution induce convergence; simple prompt strategies—structured reasoning and varied ordinary personas—counteract those forces and can produce more diverse ideation than humans.

Examining and Addressing Barriers to Diversity in LLM-Generated Ideas
Yuting Deng, Melanie Brucks, Olivier Toubia · February 23, 2026
arxiv rct medium evidence 8/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Yuting Deng unresolved corpus identity
  2. Melanie Brucks unresolved corpus identity
  3. Olivier Toubia unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Yuting Deng provider ID
  2. Melanie S. Brucks provider ID
  3. Olivier Toubia provider ID
LLMs produce less diverse idea sets than independent human samples due to sequential fixation and unified knowledge aggregation, but targeted prompting—Chain-of-Thought to reduce fixation and varied ordinary-persona prompts to induce knowledge partitioning—restores and can even exceed human-level idea diversity when combined.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Ideas generated by independent samples of humans tend to be more diverse than ideas generated from independent LLM samples, raising concerns that widespread reliance on LLMs could homogenize ideation and undermine innovation at a societal level. Drawing on cognitive psychology, we identify (both theoretically and empirically) two mechanisms undermining LLM idea diversity. First, at the individual level, LLMs exhibit fixation just as humans do, where early outputs constrain subsequent ideation. Second, at the collective level, LLMs aggregate knowledge into a unified distribution rather than exhibiting the knowledge partitioning inherent to human populations, where each person occupies a distinct region of the knowledge space. Through four studies, we demonstrate that targeted prompting interventions can address each mechanism independently: Chain-of-Thought (CoT) prompting reduces fixation by encouraging structured reasoning (only in LLMs, not humans), while ordinary personas (versus "creative entrepreneurs" such as Steve Jobs) improve knowledge partitioning by serving as diverse sampling cues, anchoring generation in distinct regions of the semantic space. Combining both approaches produces the highest idea diversity, outperforming humans. These findings offer a theoretically grounded framework for understanding LLM idea diversity and practical strategies for human-AI collaborations that leverage AI's efficiency without compromising the diversity essential to a healthy innovation ecosystem.

Summary

Main Finding

LLMs tend to produce less collectively diverse idea sets than humans because of two separable mechanisms: (1) fixation at the individual-session level (early outputs constrain later ones), and (2) knowledge aggregation at the population level (LLMs sample from a more unified distribution rather than the partitioned, idiosyncratic knowledge of distinct humans). Targeted prompting interventions—Chain-of-Thought (CoT) to reduce fixation and ordinary persona prompts to induce knowledge partitioning—each address one mechanism. Combined, they produce the highest idea diversity and, under the authors’ experimental conditions, allow LLMs to surpass humans in collective idea diversity.

Key Points

  • Two mechanisms explain LLM homogeneity:
    • Fixation (individual-level): autoregressive generation and alignment (e.g., RLHF) make early tokens disproportionately shape later outputs, analogously to human fixation.
    • Knowledge aggregation (collective-level): LLMs collapse across sources into a unified sampling distribution, whereas human populations naturally partition knowledge across distinct mental models, producing diverse explorations.
  • Prompting interventions:
    • Chain-of-Thought (CoT) prompting (generate short titles, require making them distinct, then expand) reduces fixation in LLMs by limiting early elaboration and forcing deliberate diversification. CoT does not similarly reduce fixation in humans.
    • Persona prompts that specify ordinary, distinct personas (not just famous “creative entrepreneur” figures) act as sampling cues that anchor generations in different semantic regions and recover knowledge partitioning across LLM instances.
    • Combining CoT + ordinary personas yields the largest gains; this combination outperforms human groups on collective idea diversity in the authors’ experiments.
  • Other prompting strategies:
    • Increasing temperature mildly improves diversity but reduces quality and can produce nonsensical outputs.
    • Hybrid prompting (mixing pools) is less effective than targeted CoT and persona strategies.
  • Methodological advance: a novel hierarchical LLM-based content-categorization pipeline that classifies ideas by underlying meaning (not just lexical/embedding similarity), improving measurement of idea diversity and interpretability.

Data & Methods

  • Experimental design:
    • Four empirical studies comparing human and LLM idea generation under matched, “apples-to-apples” conditions.
    • Human controls and LLM sessions were given equivalent instructions and structures to isolate whether prompting specifically closes the human–LLM diversity gap.
  • Manipulations:
    • Chain-of-Thought (CoT) prompting vs. standard prompts.
    • Persona prompts (ordinary personas vs. creative-entrepreneur-type personas).
    • Combined CoT + persona condition.
    • Secondary comparisons with temperature adjustments and hybrid prompting.
  • Key metrics and measurement approaches:
    • Within-individual fixation measured by the slope at which new idea categories accumulate across successive ideas in a session.
    • Collective-level diversity measured by the breadth of semantic categories across independent sessions/agents.
    • Hierarchical content-categorization pipeline (LLM-assisted) to cluster ideas by conceptual meaning rather than lexical similarity—used to compute diversity more reliably across humans and LLMs.
  • Findings across studies:
    • Evidence of fixation-like patterns in LLMs similar to those in humans, but CoT reduces fixation only in LLMs.
    • Ordinary personas produced greater inter-instance dispersion (better partitioning) than creative-entrepreneur personas.
    • The CoT + ordinary persona combination yielded the highest measured collective diversity, exceeding human groups in the authors’ tasks.

Implications for AI Economics

  • Collective externalities and the “tragedy of the commons”:
    • Widespread use of off-the-shelf LLM ideation without diversity-preserving interventions can homogenize market- and research-level idea pools, reducing the exploratory breadth that underpins breakthrough innovation and risk management.
    • Empirical evidence (cited in the paper) already suggests AI tool adoption can increase individual productivity while narrowing topic diversity across scientific fields.
  • Policy and platform design:
    • Platform defaults and API primitives should support diversity-preserving workflows (e.g., easy CoT-enabled pipelines, persona sampling tools) so individual efficiency gains do not aggregate into collective losses.
    • Regulators and institutions might incentivize or require diversity-oriented features in AI-aided ideation used in public R&D, grant proposal triage, or corporate strategy processes.
  • Firm strategy and competitive dynamics:
    • Firms relying on LLMs for ideation should incorporate targeted prompting (CoT + varied persona sampling) to avoid convergent strategies across firms and to maintain a portfolio of diverse approaches—important for hedging and long-shot innovation.
    • Because ordinary personas (not just “celebrity creative” prompts) are effective at producing partitioning, firms can cheaply and scalably generate diverse idea pools internally.
  • Measurement and evaluation in AI economics:
    • The hierarchical meaning-based categorization method is useful for researchers measuring variety, specialization, and novelty in markets; it improves on embedding/lexical metrics when assessing innovation diversity.
  • Research and market implications:
    • Demonstrating conditions under which LLMs can exceed human diversity reframes the trade-off: LLMs need not be an inevitable source of homogenization—properly designed prompting and system-level choices can convert them into tools that both raise individual productivity and preserve or enhance collective exploration.
    • Further economic research should study adoption dynamics (how quickly diversity-preserving prompts diffuse), platform incentives (do default UX choices favor homogenization?), and system-level equilibria when many agents use similar LLM configurations.
  • Limitations and caution:
    • Results are conditional on models, alignment procedures, tasks, and prompting prescriptions; model architectures and RLHF practices evolve rapidly, so continual reassessment is needed.
    • Quality-diversity trade-offs persist when relying on high-temperature randomness; targeted prompting appears to avoid much of that trade-off, but practical deployments should validate quality alongside diversity.

If you want, I can (a) extract the specific experimental tasks and prompt templates used in the studies, (b) sketch a simple implementation plan (API calls + prompt scaffolding) for firms to operationalize CoT + persona sampling, or (c) outline an economic model of how LLM-induced homogenization propagates through an industry. Which would be most useful?

Assessment

Paper Typerct Evidence Strengthmedium — The paper uses controlled experiments that provide credible causal evidence for prompt effects on idea diversity in the lab; however, evidence is limited to specific tasks, prompts, and model(s) studied, and the link from measured idea diversity to real-world innovation and economic outcomes is indirect. Methods Rigormedium — Design includes multiple studies, clear interventions (CoT and persona prompting), and quantitative diversity metrics, but potential weaknesses include limited reporting on sample sizes/model variety, dependence on operationalization of 'diversity' (lexical/semantic measures), possibly narrow task domains, and uncertain robustness across LLM architectures and real-world collaborative settings. SampleFour experimental studies comparing independent human participants and outputs from one or more large language models on ideation tasks; interventions randomized at the participant/model prompt level to CoT prompting, ordinary-persona prompts, combined prompts, or baseline; outcomes measured via lexical/semantic diversity metrics, overlap statistics, and fixation tests (early-output anchoring). Exact participant counts, recruitment source, and specific LLM versions are not specified in the summary. Themesinnovation human_ai_collab IdentificationRandomized experimental manipulation of generation conditions: independent samples of humans and LLMs are assigned to baseline or intervention prompting conditions (Chain-of-Thought vs. control; persona prompts vs. control), with causal effects inferred from between-condition comparisons on pre-registered diversity metrics and within-session fixation tests (early outputs used as anchors to test sequential constraint). GeneralizabilityFindings may depend on the specific LLM architecture and version used; other models could show different fixation/partitioning behavior, Prompt wording and the particular persona set could materially affect results, limiting transferability to other prompting designs, Tasks used for ideation (e.g., product ideas) may not represent all innovation or creative domains (scientific research, policy, arts), Human samples in lab/online settings may not reflect professional ideation teams or organizational contexts, Measured semantic/lexical diversity may not map one-to-one to economically meaningful innovation outcomes (patents, firm productivity)

Claims (8)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Ideas generated by independent samples of humans tend to be more diverse than ideas generated from independent LLM samples. Creativity negative idea diversity
Reading fidelity high
Study strength medium
not reported
0.6
At the individual level, LLMs exhibit fixation: early outputs constrain subsequent ideation. Creativity negative fixation (constraint on subsequent ideation) / within-sequence idea diversity
Reading fidelity high
Study strength medium
not reported
0.6
At the collective level, LLMs aggregate knowledge into a unified distribution rather than exhibiting the knowledge partitioning inherent to human populations (where individuals occupy distinct regions of knowledge space). Creativity negative knowledge partitioning across generators / cross-sample idea diversity
Reading fidelity high
Study strength medium
not reported
0.6
Chain-of-Thought (CoT) prompting reduces fixation by encouraging structured reasoning in LLMs. Creativity positive fixation reduction / within-sequence idea diversity
Reading fidelity high
Study strength medium
not reported
0.6
Chain-of-Thought prompting does not reduce fixation in humans (i.e., CoT reduced fixation only in LLMs, not humans). Creativity null_result fixation reduction / within-sequence idea diversity
Reading fidelity high
Study strength medium
not reported
0.6
Using ordinary personas (versus 'creative entrepreneurs' such as Steve Jobs) improves knowledge partitioning by serving as diverse sampling cues and anchoring generation in distinct regions of semantic space. Creativity positive knowledge partitioning across generators / cross-sample idea diversity
Reading fidelity high
Study strength medium
not reported
0.6
Combining Chain-of-Thought prompting and ordinary personas produces the highest idea diversity and outperforms humans. Creativity positive idea diversity (combined-intervention effect)
Reading fidelity medium
Study strength medium
not reported
0.36
The two mechanisms (individual-level fixation and collective-level lack of knowledge partitioning) together explain why widespread reliance on LLMs could homogenize ideation and undermine innovation at a societal level. Innovation Output negative societal-level ideation homogenization / potential impact on innovation
Reading fidelity high
Study strength speculative
not reported
0.1

Notes