17 cumulative citations
View corpus contextLarge language models produce better individual product ideas and more top performers, but they converge on a narrower set of concepts than humans; careful model choice, pooling and prompt design largely restore human-level diversity.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
This research examines how well large language models, or LLMs, generate new product ideas for college students priced under $50. Across a series of studies, we identify key strengths and weaknesses of using LLMs for product innovation. Our first study shows that LLM-generated product ideas have higher average quality than human ideas, based on purchase intent, and are 7 times more likely to rank in the top 10%. Our second study shows that this AI-induced creativity boost is not explained by the LLM's more persuasive pitching skills. Our third and fourth studies identify a weakness of using LLMs for brainstorming: AI-generated ideas are less novel at the idea level and less diverse at the set level. In our fifth study, we analyze prior LLM-based creativity studies and find consistently lower idea diversity across all of them, demonstrating the generalizability of these findings. Our sixth and seventh studies investigate techniques to mitigate this diversity loss. We compare LLMs from different vendors and versions and find that more recent models generate more diverse ideas, though they still fall short of human-level diversity. We also demonstrate techniques that increase idea diversity almost to the level of human idea generation: pooling ideas across vendors; prompt engineering, including Chain-of-Thought prompting and injecting heterogeneous personas or constraints; and creative agents that broadly explore the solution landscape to restore diversity. Finally, in our eighth study, we show that exploiting the near-zero marginal cost of AI idea generation by scaling the number of ideas steadily improves coverage of the idea space, approaching human-level coverage. We conclude by presenting actionable recommendations for innovation managers who want to identify better new product ideas with the help of LLMs.
Summary
Main Finding
LLMs produce higher-average and higher-tail-quality product ideas than humans (AI ideas are 7× more likely to fall in the top 10% by purchase intent), but they generate substantially less novelty and set-level diversity. This diversity loss can be largely mitigated (though not fully eliminated) by model choice, pooling across vendors, prompt engineering (e.g., Chain-of-Thought, injected personas/constraints), creative-agent workflows, and—owing to AI’s near-zero marginal cost—by scaling up the number of ideas generated.
Key Points
- Quality
- Study 1a: AI-generated ideas (for physical products targeted at college students priced ≤ $50) have higher average quality measured by customer purchase intent; AI ideas are ~7× more likely to rank in the top 10%.
- Study 1b: The quality advantage is not explained by superior AI “pitching” — rephrasing human ideas in AI writing style did not change purchase intent.
- Study 1c: Human-rated novelty of AI ideas declined only slightly, indicating AI outputs are not simply verbatim regurgitations.
- Diversity and Novelty
- Study 2a: AI-generated idea sets are substantially less diverse (ideas are more semantically similar to each other) than human-generated sets.
- Study 2b (meta-analysis): Across multiple prior LLM-creativity datasets, AI ideas show consistently lower diversity (effect sizes ranging small→very large) under multiple diversity metrics.
- Mitigation strategies
- Study 3a: Newer LLM versions and pooling across different vendors/models increase diversity; more recent models are more diverse but still below human-level diversity.
- Study 3b: Procedural/prompt interventions—Chain-of-Thought prompting, injecting heterogeneous personas or constraints, and using specialized creative agents—meaningfully increase idea diversity.
- Study 3c: Scaling the number of AI-generated ideas (e.g., from 100 → 1,000) steadily increases coverage of the idea space and can approach human-level coverage; some human-explored regions remain hard for AI to reach.
- Practical conclusion: LLMs are powerful for finding high-quality outliers and for high-throughput ideation, but managers should combine technical and procedural interventions to avoid homogenized portfolios.
Data & Methods
- Research design: Eight linked studies combining original experiments, controlled comparisons, meta-analysis, and mitigation experiments.
- Primary setting: New-product ideation for college-student-targeted physical products priced ≤ $50. Human ideas came from an MBA innovation course (50 students; ideas collected 2021, pre-LLM widespread use).
- Core outcome measures:
- Idea quality: customer purchase-intent ratings (market-validated metric).
- Novelty: human-judge novelty ratings.
- Diversity: semantic-distance based metrics at the set level (multiple metrics applied; also applied in meta-analysis to external LLM-creativity datasets).
- Coverage: extent of idea-space covered as number of AI ideas scaled.
- Experiments and comparisons:
- Study 1a: Direct comparison of human vs. LLM outputs (zero-shot and few-shot prompting variants).
- Study 1b: Control for communication by rephrasing human ideas with LLMs and re-evaluating purchase intent.
- Study 1c: Human novelty ratings to assess regurgitation concerns.
- Study 2a: Quantified semantic diversity difference between human and AI idea sets.
- Study 2b: Meta-analysis applying diversity metrics to three prior studies (Stevenson et al. 2022; Hubert et al. 2024; Boussioux et al. 2024).
- Study 3a: Technical levers — model version/vendor comparisons, sampling parameters, pooling vendor outputs.
- Study 3b: Procedural levers — Chain-of-Thought prompting, personas/constraints, creative-agent architectures combining human and AI exploration.
- Study 3c: Scaling experiments — generating 100→1,000 AI ideas and measuring coverage gains.
- Models evaluated: Multiple LLM vendors and versions (examples in paper include GPT-series and other vendors), with experiments showing newer models generally yield greater diversity.
- Robustness checks: Controlled for persuasion/pitch effects, applied multiple diversity metrics, and cross-validated results via meta-analysis.
Implications for AI Economics
- Productivity vs. Exploration trade-off
- LLMs increase ideation productivity and raise the expected value of the best ideas (favorable for high-variance search problems), shifting managerial focus toward high-throughput search and selection.
- Simultaneously, LLMs’ tendency to converge on high-probability associations narrows exploration, potentially reducing portfolio robustness and missing niche/novel opportunities.
- Returns to scale and marginal cost
- Near-zero marginal cost of LLM ideation turns quantity into a practical mitigation: scale can restore coverage of the idea space, implying new optimal investment rules for search effort (generate many inexpensive AI ideas and then filter/select).
- Organizational strategy
- Firms should combine AI ideation with methods that diversify the search (use multiple models/vendors, invest in prompt engineering, run creative-agent workflows, and retain human ideators for hard-to-reach solution regions).
- “Model diversification” (pooling outputs across vendors/versions) is analogous to portfolio diversification and can be an inexpensive way to increase idea-space coverage.
- Market structure and competition
- If many firms rely on similar LLMs and prompts, product-market homogenization risks increase (less product variety, increased direct competition). Firms that invest in mitigation strategies (bespoke prompts, creative agents, proprietary data) can gain differentiation advantages.
- Welfare and policy considerations
- Consumer welfare: higher-quality mainstream products may increase welfare, but reduced diversity could harm consumers who prefer niche offerings.
- Intellectual property and originality concerns: small but present reductions in assessed novelty and the possibility of training-data leakage warrant caution (due diligence, provenance checks).
- Labor and task allocation: managers can reallocate human effort from quantity generation toward high-value curation, exploration of under-explored regions, and integration tasks.
- Research implications for AI economics
- Need for models of innovation search that incorporate AI’s productivity and homogenization tendencies.
- Empirical work should quantify when scaling AI ideation is substitute vs. complement to human ideation across domains.
- Policy and antitrust analyses should consider how commoditization of ideation via common LLMs affects product differentiation and industry dynamics.
Limitations noted by authors - Domain: experiments focused on low-cost college-student products; generalizability to other product types, industries, or more complex innovation tasks requires further study. - Training-data and model-update specifics: model behavior depends on vendor training and updates; results may change as models evolve. - Some human-explored idea regions remained difficult for LLMs to reach even after mitigation—human creativity still contributes unique value.
Actionable managerial takeaways (brief) - Use LLMs to rapidly generate large idea pools to surface high-quality outliers. - Counteract AI homogenization by: pooling vendors/models, using Chain-of-Thought and persona/constraint prompts, deploying creative-agent workflows, and allocating human effort to explore remaining blind spots.
Assessment
Claims (11)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| LLM-generated product ideas had higher average quality than human-generated ideas, as measured by customer purchase intent. Output Quality | positive | Customer purchase intent for generated product ideas |
Reading fidelity
high
Study strength
medium
|
not reported
|
| LLM-generated product ideas were seven times more likely than human-generated ideas to be evaluated as being among the best 10%. Output Quality | positive | Probability that a product idea ranks in the top 10% by purchase intent |
Reading fidelity
high
Study strength
medium
|
7x more likely
|
| The higher purchase intent of AI-generated ideas was not attributable to superior AI pitching or communication skills. Output Quality | null_result | Purchase intent for original versus AI-rephrased human ideas |
Reading fidelity
high
Study strength
medium
|
no meaningful difference
|
| AI-generated ideas had slightly lower novelty than human-generated ideas, but the decrease was small. Creativity | negative | Rated novelty of product ideas |
Reading fidelity
high
Study strength
medium
|
small decrease
|
| AI-generated idea sets were substantially less diverse than human-generated idea sets, meaning that the AI ideas were more similar to one another. Creativity | negative | Semantic diversity or distance among ideas within an idea set |
Reading fidelity
high
Study strength
medium
|
substantially reduced
|
| Across three additional AI-creativity studies and multiple diversity metrics, AI-generated ideas were significantly less diverse than human-generated ideas. Creativity | negative | Diversity of generated ideas across studies and diversity metrics |
Reading fidelity
high
Study strength
high
|
effect sizes ranging from small to very large
|
| More recent LLM models generated more diverse ideas than earlier models, although they still did not reach human-level diversity. Creativity | positive | Diversity of generated product ideas relative to human-generated ideas |
Reading fidelity
high
Study strength
medium
|
more diverse, though still below human-level diversity
|
| Pooling ideas generated by different model vendors increased idea diversity, bringing it close to the level of human idea generation. Creativity | positive | Diversity of pooled AI-generated ideas |
Reading fidelity
high
Study strength
medium
|
almost to the level of human idea generation
|
| Chain-of-Thought prompting and injecting heterogeneous personas or functional constraints increased the diversity of LLM-generated ideas. Creativity | positive | Diversity of generated product ideas |
Reading fidelity
high
Study strength
medium
|
increase in idea diversity
|
| Generating between 100 and 1,000 AI ideas steadily improved coverage of the idea space, approaching human-level coverage without reaching a plateau. Creativity | positive | Coverage of the product-idea solution space |
Reading fidelity
high
Study strength
medium
|
from 100 to 1,000 ideas; almost reaching human-level coverage
|
| The human comparison data came from 50 MBA students who generated physical-product ideas for college students priced at $50 or less. Other | other | Generation of product ideas under a price constraint |
Reading fidelity
high
Study strength
medium
|
n=50
|