The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests About 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Large language models produce better individual product ideas and more top performers, but they converge on a narrower set of concepts than humans; careful model choice, pooling and prompt design largely restore human-level diversity.

AI and Its Impact on Creativity and Diversity: An Empirical Study of LLM-Generated Product Ideas
Christian Terwiesch, Lennart Meincke, Karan Girotra, Ethan Mollick, Gideon Nave, Karl T. Ulrich · July 30, 2026
arxiv quasi_experimental medium evidence 8/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Christian Terwiesch unresolved corpus identity
  2. Lennart Meincke unresolved corpus identity
  3. Karan Girotra unresolved corpus identity
  4. Ethan Mollick unresolved corpus identity
  5. Gideon Nave unresolved corpus identity
  6. Karl T. Ulrich unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Christian Terwiesch provider ID
  2. Lennart Meincke provider ID
  3. Karan Girotra provider ID
  4. Ethan Mollick provider ID
  5. G. Nave provider ID
  6. K. Ulrich provider ID
LLMs generate product ideas that score higher on average purchase-intent and are far more likely to produce top-tier ideas, but their idea sets are measurably less novel and less diverse than human-generated sets — a shortfall that can be largely mitigated by newer models, pooling across vendors, prompt engineering, creative agents, and scaling the number of AI-generated ideas.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

This research examines how well large language models, or LLMs, generate new product ideas for college students priced under $50. Across a series of studies, we identify key strengths and weaknesses of using LLMs for product innovation. Our first study shows that LLM-generated product ideas have higher average quality than human ideas, based on purchase intent, and are 7 times more likely to rank in the top 10%. Our second study shows that this AI-induced creativity boost is not explained by the LLM's more persuasive pitching skills. Our third and fourth studies identify a weakness of using LLMs for brainstorming: AI-generated ideas are less novel at the idea level and less diverse at the set level. In our fifth study, we analyze prior LLM-based creativity studies and find consistently lower idea diversity across all of them, demonstrating the generalizability of these findings. Our sixth and seventh studies investigate techniques to mitigate this diversity loss. We compare LLMs from different vendors and versions and find that more recent models generate more diverse ideas, though they still fall short of human-level diversity. We also demonstrate techniques that increase idea diversity almost to the level of human idea generation: pooling ideas across vendors; prompt engineering, including Chain-of-Thought prompting and injecting heterogeneous personas or constraints; and creative agents that broadly explore the solution landscape to restore diversity. Finally, in our eighth study, we show that exploiting the near-zero marginal cost of AI idea generation by scaling the number of ideas steadily improves coverage of the idea space, approaching human-level coverage. We conclude by presenting actionable recommendations for innovation managers who want to identify better new product ideas with the help of LLMs.

Summary

Main Finding

LLMs produce higher-average and higher-tail-quality product ideas than humans (AI ideas are 7× more likely to fall in the top 10% by purchase intent), but they generate substantially less novelty and set-level diversity. This diversity loss can be largely mitigated (though not fully eliminated) by model choice, pooling across vendors, prompt engineering (e.g., Chain-of-Thought, injected personas/constraints), creative-agent workflows, and—owing to AI’s near-zero marginal cost—by scaling up the number of ideas generated.

Key Points

  • Quality
    • Study 1a: AI-generated ideas (for physical products targeted at college students priced ≤ $50) have higher average quality measured by customer purchase intent; AI ideas are ~7× more likely to rank in the top 10%.
    • Study 1b: The quality advantage is not explained by superior AI “pitching” — rephrasing human ideas in AI writing style did not change purchase intent.
    • Study 1c: Human-rated novelty of AI ideas declined only slightly, indicating AI outputs are not simply verbatim regurgitations.
  • Diversity and Novelty
    • Study 2a: AI-generated idea sets are substantially less diverse (ideas are more semantically similar to each other) than human-generated sets.
    • Study 2b (meta-analysis): Across multiple prior LLM-creativity datasets, AI ideas show consistently lower diversity (effect sizes ranging small→very large) under multiple diversity metrics.
  • Mitigation strategies
    • Study 3a: Newer LLM versions and pooling across different vendors/models increase diversity; more recent models are more diverse but still below human-level diversity.
    • Study 3b: Procedural/prompt interventions—Chain-of-Thought prompting, injecting heterogeneous personas or constraints, and using specialized creative agents—meaningfully increase idea diversity.
    • Study 3c: Scaling the number of AI-generated ideas (e.g., from 100 → 1,000) steadily increases coverage of the idea space and can approach human-level coverage; some human-explored regions remain hard for AI to reach.
  • Practical conclusion: LLMs are powerful for finding high-quality outliers and for high-throughput ideation, but managers should combine technical and procedural interventions to avoid homogenized portfolios.

Data & Methods

  • Research design: Eight linked studies combining original experiments, controlled comparisons, meta-analysis, and mitigation experiments.
  • Primary setting: New-product ideation for college-student-targeted physical products priced ≤ $50. Human ideas came from an MBA innovation course (50 students; ideas collected 2021, pre-LLM widespread use).
  • Core outcome measures:
    • Idea quality: customer purchase-intent ratings (market-validated metric).
    • Novelty: human-judge novelty ratings.
    • Diversity: semantic-distance based metrics at the set level (multiple metrics applied; also applied in meta-analysis to external LLM-creativity datasets).
    • Coverage: extent of idea-space covered as number of AI ideas scaled.
  • Experiments and comparisons:
    • Study 1a: Direct comparison of human vs. LLM outputs (zero-shot and few-shot prompting variants).
    • Study 1b: Control for communication by rephrasing human ideas with LLMs and re-evaluating purchase intent.
    • Study 1c: Human novelty ratings to assess regurgitation concerns.
    • Study 2a: Quantified semantic diversity difference between human and AI idea sets.
    • Study 2b: Meta-analysis applying diversity metrics to three prior studies (Stevenson et al. 2022; Hubert et al. 2024; Boussioux et al. 2024).
    • Study 3a: Technical levers — model version/vendor comparisons, sampling parameters, pooling vendor outputs.
    • Study 3b: Procedural levers — Chain-of-Thought prompting, personas/constraints, creative-agent architectures combining human and AI exploration.
    • Study 3c: Scaling experiments — generating 100→1,000 AI ideas and measuring coverage gains.
  • Models evaluated: Multiple LLM vendors and versions (examples in paper include GPT-series and other vendors), with experiments showing newer models generally yield greater diversity.
  • Robustness checks: Controlled for persuasion/pitch effects, applied multiple diversity metrics, and cross-validated results via meta-analysis.

Implications for AI Economics

  • Productivity vs. Exploration trade-off
    • LLMs increase ideation productivity and raise the expected value of the best ideas (favorable for high-variance search problems), shifting managerial focus toward high-throughput search and selection.
    • Simultaneously, LLMs’ tendency to converge on high-probability associations narrows exploration, potentially reducing portfolio robustness and missing niche/novel opportunities.
  • Returns to scale and marginal cost
    • Near-zero marginal cost of LLM ideation turns quantity into a practical mitigation: scale can restore coverage of the idea space, implying new optimal investment rules for search effort (generate many inexpensive AI ideas and then filter/select).
  • Organizational strategy
    • Firms should combine AI ideation with methods that diversify the search (use multiple models/vendors, invest in prompt engineering, run creative-agent workflows, and retain human ideators for hard-to-reach solution regions).
    • “Model diversification” (pooling outputs across vendors/versions) is analogous to portfolio diversification and can be an inexpensive way to increase idea-space coverage.
  • Market structure and competition
    • If many firms rely on similar LLMs and prompts, product-market homogenization risks increase (less product variety, increased direct competition). Firms that invest in mitigation strategies (bespoke prompts, creative agents, proprietary data) can gain differentiation advantages.
  • Welfare and policy considerations
    • Consumer welfare: higher-quality mainstream products may increase welfare, but reduced diversity could harm consumers who prefer niche offerings.
    • Intellectual property and originality concerns: small but present reductions in assessed novelty and the possibility of training-data leakage warrant caution (due diligence, provenance checks).
    • Labor and task allocation: managers can reallocate human effort from quantity generation toward high-value curation, exploration of under-explored regions, and integration tasks.
  • Research implications for AI economics
    • Need for models of innovation search that incorporate AI’s productivity and homogenization tendencies.
    • Empirical work should quantify when scaling AI ideation is substitute vs. complement to human ideation across domains.
    • Policy and antitrust analyses should consider how commoditization of ideation via common LLMs affects product differentiation and industry dynamics.

Limitations noted by authors - Domain: experiments focused on low-cost college-student products; generalizability to other product types, industries, or more complex innovation tasks requires further study. - Training-data and model-update specifics: model behavior depends on vendor training and updates; results may change as models evolve. - Some human-explored idea regions remained difficult for LLMs to reach even after mitigation—human creativity still contributes unique value.

Actionable managerial takeaways (brief) - Use LLMs to rapidly generate large idea pools to surface high-quality outliers. - Counteract AI homogenization by: pooling vendors/models, using Chain-of-Thought and persona/constraint prompts, deploying creative-agent workflows, and allocating human effort to explore remaining blind spots.

Assessment

Paper Typequasi_experimental Evidence Strengthmedium — Multiple complementary studies, blinded customer purchase-intent as a market-validated outcome, mechanism checks (pitching, novelty) and a meta-analysis strengthen causal claims that LLMs produce higher-quality top ideas but lower within-set diversity; however the setting is narrow (student product ideas under $50), details about randomization and exact sample sizes for each study are incomplete in the provided text, and risks remain from training-data leakage and rapidly evolving models. Methods Rigormedium — Strengths: multi-study approach, use of an external market-relevant metric (purchase intent), blinded evaluations, targeted mechanism tests, meta-analysis and practical mitigation experiments (model choice, pooling, prompt engineering, scaling). Weaknesses: human idea sample is from a single course (possible selection bias), limited domain (cheap products for students), incomplete description of randomization/allocation and statistical controls in the excerpt, and potential confounds from model training data and changing model versions. SamplePrimary human idea sample: ~50 MBA students from a US university (ideas collected in a 2021 product-design/innovation course) who produced several hundred product ideas for physical products targeted at college students priced <= $50; LLM-generated ideas: multiple models/versions (vendors and newer models compared), generated via zero-shot and few-shot prompting and various prompt-engineering interventions; outcome raters: potential customers (college students/purchasers) who provided purchase-intent and novelty ratings; additional data: three external LLM creativity studies incorporated into a meta-analysis; some experiments scale AI-generated ideas from ~100 to ~1,000 to assess coverage. Themesinnovation human_ai_collab IdentificationComparative multi-study design in which ideas generated by LLMs are evaluated side-by-side with human-generated ideas using blinded market-relevant outcome measures (customer purchase-intent); mechanism tests include (a) rephrasing human ideas by the LLM to isolate pitching effects, (b) human novelty ratings to probe regurgitation, (c) cross-model comparisons and parameter variation (e.g., temperature, model/version) to test model-based heterogeneity, (d) pooling and prompt-engineering interventions to test mitigation strategies, and (e) a meta-analysis applying diversity metrics to external LLM creativity datasets for robustness. GeneralizabilityDomain-limited: physical products for college students priced under $50 — may not generalize to services, expensive or B2B products, or other demographic segments., Human sample is students from one course/university — limited representativeness of professional ideators., LLM landscape evolves rapidly — results tied to specific model versions and vendors available at time of study., Outcome proxy (purchase-intent) is a market-relevant but indirect measure of eventual commercial success., Possible cultural/geographic limits if raters and ideators are US-centric., Potential training-data leakage (LLMs may reproduce known products) could bias novelty assessments in other domains.

Claims (11)

ClaimDirectionOutcomeConfidence & EvidenceDetails
LLM-generated product ideas had higher average quality than human-generated ideas, as measured by customer purchase intent. Output Quality positive Customer purchase intent for generated product ideas
Reading fidelity high
Study strength medium
not reported
0.48
LLM-generated product ideas were seven times more likely than human-generated ideas to be evaluated as being among the best 10%. Output Quality positive Probability that a product idea ranks in the top 10% by purchase intent
Reading fidelity high
Study strength medium
7x more likely
0.48
The higher purchase intent of AI-generated ideas was not attributable to superior AI pitching or communication skills. Output Quality null_result Purchase intent for original versus AI-rephrased human ideas
Reading fidelity high
Study strength medium
no meaningful difference
0.48
AI-generated ideas had slightly lower novelty than human-generated ideas, but the decrease was small. Creativity negative Rated novelty of product ideas
Reading fidelity high
Study strength medium
small decrease
0.48
AI-generated idea sets were substantially less diverse than human-generated idea sets, meaning that the AI ideas were more similar to one another. Creativity negative Semantic diversity or distance among ideas within an idea set
Reading fidelity high
Study strength medium
substantially reduced
0.48
Across three additional AI-creativity studies and multiple diversity metrics, AI-generated ideas were significantly less diverse than human-generated ideas. Creativity negative Diversity of generated ideas across studies and diversity metrics
Reading fidelity high
Study strength high
effect sizes ranging from small to very large
0.8
More recent LLM models generated more diverse ideas than earlier models, although they still did not reach human-level diversity. Creativity positive Diversity of generated product ideas relative to human-generated ideas
Reading fidelity high
Study strength medium
more diverse, though still below human-level diversity
0.48
Pooling ideas generated by different model vendors increased idea diversity, bringing it close to the level of human idea generation. Creativity positive Diversity of pooled AI-generated ideas
Reading fidelity high
Study strength medium
almost to the level of human idea generation
0.48
Chain-of-Thought prompting and injecting heterogeneous personas or functional constraints increased the diversity of LLM-generated ideas. Creativity positive Diversity of generated product ideas
Reading fidelity high
Study strength medium
increase in idea diversity
0.48
Generating between 100 and 1,000 AI ideas steadily improved coverage of the idea space, approaching human-level coverage without reaching a plateau. Creativity positive Coverage of the product-idea solution space
Reading fidelity high
Study strength medium
from 100 to 1,000 ideas; almost reaching human-level coverage
0.48
The human comparison data came from 50 MBA students who generated physical-product ideas for college students priced at $50 or less. Other other Generation of product ideas under a price constraint
Reading fidelity high
Study strength medium
n=50
0.48

Notes