2 cumulative citations
View corpus contextOff‑the‑shelf language models tend to generate narrower idea pools than groups of humans because early outputs and a unified knowledge distribution induce convergence; simple prompt strategies—structured reasoning and varied ordinary personas—counteract those forces and can produce more diverse ideation than humans.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Ideas generated by independent samples of humans tend to be more diverse than ideas generated from independent LLM samples, raising concerns that widespread reliance on LLMs could homogenize ideation and undermine innovation at a societal level. Drawing on cognitive psychology, we identify (both theoretically and empirically) two mechanisms undermining LLM idea diversity. First, at the individual level, LLMs exhibit fixation just as humans do, where early outputs constrain subsequent ideation. Second, at the collective level, LLMs aggregate knowledge into a unified distribution rather than exhibiting the knowledge partitioning inherent to human populations, where each person occupies a distinct region of the knowledge space. Through four studies, we demonstrate that targeted prompting interventions can address each mechanism independently: Chain-of-Thought (CoT) prompting reduces fixation by encouraging structured reasoning (only in LLMs, not humans), while ordinary personas (versus "creative entrepreneurs" such as Steve Jobs) improve knowledge partitioning by serving as diverse sampling cues, anchoring generation in distinct regions of the semantic space. Combining both approaches produces the highest idea diversity, outperforming humans. These findings offer a theoretically grounded framework for understanding LLM idea diversity and practical strategies for human-AI collaborations that leverage AI's efficiency without compromising the diversity essential to a healthy innovation ecosystem.
Summary
Main Finding
LLMs tend to produce less collectively diverse idea sets than humans because of two separable mechanisms: (1) fixation at the individual-session level (early outputs constrain later ones), and (2) knowledge aggregation at the population level (LLMs sample from a more unified distribution rather than the partitioned, idiosyncratic knowledge of distinct humans). Targeted prompting interventions—Chain-of-Thought (CoT) to reduce fixation and ordinary persona prompts to induce knowledge partitioning—each address one mechanism. Combined, they produce the highest idea diversity and, under the authors’ experimental conditions, allow LLMs to surpass humans in collective idea diversity.
Key Points
- Two mechanisms explain LLM homogeneity:
- Fixation (individual-level): autoregressive generation and alignment (e.g., RLHF) make early tokens disproportionately shape later outputs, analogously to human fixation.
- Knowledge aggregation (collective-level): LLMs collapse across sources into a unified sampling distribution, whereas human populations naturally partition knowledge across distinct mental models, producing diverse explorations.
- Prompting interventions:
- Chain-of-Thought (CoT) prompting (generate short titles, require making them distinct, then expand) reduces fixation in LLMs by limiting early elaboration and forcing deliberate diversification. CoT does not similarly reduce fixation in humans.
- Persona prompts that specify ordinary, distinct personas (not just famous “creative entrepreneur” figures) act as sampling cues that anchor generations in different semantic regions and recover knowledge partitioning across LLM instances.
- Combining CoT + ordinary personas yields the largest gains; this combination outperforms human groups on collective idea diversity in the authors’ experiments.
- Other prompting strategies:
- Increasing temperature mildly improves diversity but reduces quality and can produce nonsensical outputs.
- Hybrid prompting (mixing pools) is less effective than targeted CoT and persona strategies.
- Methodological advance: a novel hierarchical LLM-based content-categorization pipeline that classifies ideas by underlying meaning (not just lexical/embedding similarity), improving measurement of idea diversity and interpretability.
Data & Methods
- Experimental design:
- Four empirical studies comparing human and LLM idea generation under matched, “apples-to-apples” conditions.
- Human controls and LLM sessions were given equivalent instructions and structures to isolate whether prompting specifically closes the human–LLM diversity gap.
- Manipulations:
- Chain-of-Thought (CoT) prompting vs. standard prompts.
- Persona prompts (ordinary personas vs. creative-entrepreneur-type personas).
- Combined CoT + persona condition.
- Secondary comparisons with temperature adjustments and hybrid prompting.
- Key metrics and measurement approaches:
- Within-individual fixation measured by the slope at which new idea categories accumulate across successive ideas in a session.
- Collective-level diversity measured by the breadth of semantic categories across independent sessions/agents.
- Hierarchical content-categorization pipeline (LLM-assisted) to cluster ideas by conceptual meaning rather than lexical similarity—used to compute diversity more reliably across humans and LLMs.
- Findings across studies:
- Evidence of fixation-like patterns in LLMs similar to those in humans, but CoT reduces fixation only in LLMs.
- Ordinary personas produced greater inter-instance dispersion (better partitioning) than creative-entrepreneur personas.
- The CoT + ordinary persona combination yielded the highest measured collective diversity, exceeding human groups in the authors’ tasks.
Implications for AI Economics
- Collective externalities and the “tragedy of the commons”:
- Widespread use of off-the-shelf LLM ideation without diversity-preserving interventions can homogenize market- and research-level idea pools, reducing the exploratory breadth that underpins breakthrough innovation and risk management.
- Empirical evidence (cited in the paper) already suggests AI tool adoption can increase individual productivity while narrowing topic diversity across scientific fields.
- Policy and platform design:
- Platform defaults and API primitives should support diversity-preserving workflows (e.g., easy CoT-enabled pipelines, persona sampling tools) so individual efficiency gains do not aggregate into collective losses.
- Regulators and institutions might incentivize or require diversity-oriented features in AI-aided ideation used in public R&D, grant proposal triage, or corporate strategy processes.
- Firm strategy and competitive dynamics:
- Firms relying on LLMs for ideation should incorporate targeted prompting (CoT + varied persona sampling) to avoid convergent strategies across firms and to maintain a portfolio of diverse approaches—important for hedging and long-shot innovation.
- Because ordinary personas (not just “celebrity creative” prompts) are effective at producing partitioning, firms can cheaply and scalably generate diverse idea pools internally.
- Measurement and evaluation in AI economics:
- The hierarchical meaning-based categorization method is useful for researchers measuring variety, specialization, and novelty in markets; it improves on embedding/lexical metrics when assessing innovation diversity.
- Research and market implications:
- Demonstrating conditions under which LLMs can exceed human diversity reframes the trade-off: LLMs need not be an inevitable source of homogenization—properly designed prompting and system-level choices can convert them into tools that both raise individual productivity and preserve or enhance collective exploration.
- Further economic research should study adoption dynamics (how quickly diversity-preserving prompts diffuse), platform incentives (do default UX choices favor homogenization?), and system-level equilibria when many agents use similar LLM configurations.
- Limitations and caution:
- Results are conditional on models, alignment procedures, tasks, and prompting prescriptions; model architectures and RLHF practices evolve rapidly, so continual reassessment is needed.
- Quality-diversity trade-offs persist when relying on high-temperature randomness; targeted prompting appears to avoid much of that trade-off, but practical deployments should validate quality alongside diversity.
If you want, I can (a) extract the specific experimental tasks and prompt templates used in the studies, (b) sketch a simple implementation plan (API calls + prompt scaffolding) for firms to operationalize CoT + persona sampling, or (c) outline an economic model of how LLM-induced homogenization propagates through an industry. Which would be most useful?
Assessment
Claims (8)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Ideas generated by independent samples of humans tend to be more diverse than ideas generated from independent LLM samples. Creativity | negative | idea diversity |
Reading fidelity
high
Study strength
medium
|
not reported
|
| At the individual level, LLMs exhibit fixation: early outputs constrain subsequent ideation. Creativity | negative | fixation (constraint on subsequent ideation) / within-sequence idea diversity |
Reading fidelity
high
Study strength
medium
|
not reported
|
| At the collective level, LLMs aggregate knowledge into a unified distribution rather than exhibiting the knowledge partitioning inherent to human populations (where individuals occupy distinct regions of knowledge space). Creativity | negative | knowledge partitioning across generators / cross-sample idea diversity |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Chain-of-Thought (CoT) prompting reduces fixation by encouraging structured reasoning in LLMs. Creativity | positive | fixation reduction / within-sequence idea diversity |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Chain-of-Thought prompting does not reduce fixation in humans (i.e., CoT reduced fixation only in LLMs, not humans). Creativity | null_result | fixation reduction / within-sequence idea diversity |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Using ordinary personas (versus 'creative entrepreneurs' such as Steve Jobs) improves knowledge partitioning by serving as diverse sampling cues and anchoring generation in distinct regions of semantic space. Creativity | positive | knowledge partitioning across generators / cross-sample idea diversity |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Combining Chain-of-Thought prompting and ordinary personas produces the highest idea diversity and outperforms humans. Creativity | positive | idea diversity (combined-intervention effect) |
Reading fidelity
medium
Study strength
medium
|
not reported
|
| The two mechanisms (individual-level fixation and collective-level lack of knowledge partitioning) together explain why widespread reliance on LLMs could homogenize ideation and undermine innovation at a societal level. Innovation Output | negative | societal-level ideation homogenization / potential impact on innovation |
Reading fidelity
high
Study strength
speculative
|
not reported
|