0 cumulative citations
View corpus contextState-of-the-art LLMs out-forecast humans on startup crowdfunding: Gemini 2.5 Pro correctly orders about four in five venture pairs, beating 346 managers and MBA investors in a live Kickstarter tournament.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
3 cumulative citations
View corpus contextCan artificial intelligence outperform humans at strategic foresight -- the capacity to form accurate judgments about uncertain, high-stakes outcomes before they unfold? We address this question through a fully prospective prediction tournament using live Kickstarter crowdfunding projects. Thirty U.S.-based technology ventures, launched after the training cutoffs of all models studied, were evaluated while fundraising remained in progress and outcomes were unknown. A diverse suite of frontier and open-weight large language models (LLMs) completed 870 pairwise comparisons, producing complete rankings of predicted fundraising success. We benchmarked these forecasts against 346 experienced managers recruited via Prolific and three MBA-trained investors working under monitored conditions. The results are striking: human evaluators achieved rank correlations with actual outcomes between 0.04 and 0.45, while several frontier LLMs exceeded 0.60, with the best (Gemini 2.5 Pro) reaching 0.74 -- correctly ordering nearly four of every five venture pairs. These differences persist across multiple performance metrics and robustness checks. Neither wisdom-of-the-crowd ensembles nor human-AI hybrid teams outperformed the best standalone model.
Summary
Main Finding
Frontier large language models substantially outperformed experienced human evaluators in a fully prospective strategic-foresight benchmark. In a live prediction tournament over 30 U.S. technology Kickstarter projects (sampled after models' training cutoffs), several LLMs achieved rank correlations with realized fundraising outcomes > 0.60; the best model (Gemini 2.5 Pro) reached 0.74 (≈79% of pairwise orderings correct). Human managers (n=346) and three MBA investors achieved rank correlations between 0.04 and 0.45 (top human ≈60% pairwise accuracy). Ensembles and human–AI hybrids did not beat the best standalone model.
Key Points
- Prospective, real-world test: predictions were made while projects were fundraising and before outcomes were known; realized total funds raised provided objective ground truth.
- Task design: 30 Kickstarter technology ventures compared in a double round‑robin pairwise tournament (870 comparisons) to construct complete rankings of expected success.
- Models evaluated: a diverse suite of frontier and open-weight LLMs (GPT-5 variants, Claude 4.5, Gemini 2.5 variants, Grok 4, Gemma 3, Llama 3.1, DeepSeek 3.2, etc.).
- Human benchmarks: 346 employed U.S. managers recruited on Prolific plus three monitored MBA-trained investors performed the identical pairwise task.
- Performance gap: several LLMs >0.60 rank correlation; Gemini 2.5 Pro = 0.74. Human rank correlations ranged 0.04–0.45. Differences are statistically significant and robust across metrics and checks.
- Aggregation: neither “wisdom of silicon” ensembles nor aggregated human–AI teams outperformed the best single LLM.
- Method contribution: pairwise tournament scoring (double round‑robin) is an effective elicitation method for stable ranking from comparative judgments.
- Exploratory analysis: cross-model performance regressed on standard AI capability benchmarks to investigate drivers of LLM foresight performance.
Data & Methods
- Sample: 30 U.S.-based technology projects on Kickstarter; each project was launched after the training cutoff dates of all LLMs studied and was evaluated while fundraising remained in progress.
- Outcome measure: realized total funds raised at campaign close (objective market outcome).
- Prediction elicitation:
- LLMs: each model performed all pairwise comparisons between the 30 projects in both presentation orders (30 choose 2 = 435 pairs × 2 = 870 comparisons), producing a complete rank ordering via tournament aggregation.
- Humans: 346 experienced managers on Prolific completed the identical pairwise comparison task; three MBA-trained investors produced full rankings under monitored, no-technology conditions.
- Metrics: primary metric was rank correlation between predicted and realized rankings; pairwise accuracy (fraction of correctly ordered pairs) and other robustness checks reported.
- Controls/robustness: presentation-order controls (double round‑robin), multiple performance metrics, tests of aggregation strategies (crowd ensembles and hybrid teams), and regressions linking model accuracy to capability benchmarks.
- Key statistics: best LLM rank correlation = 0.74 (Gemini 2.5 Pro), several LLMs >0.60; human range 0.04–0.45. Best model ≈79% pairwise accuracy; top human ≈60%.
Implications for AI Economics
- Forecasting and market outcomes
- LLMs can materially improve ex ante forecasts of market-driven outcomes in strategic settings, potentially raising the precision of investment selection, product-launch evaluation, and real-time market signals.
- Widespread use may alter market efficiency: better foresight can change allocation decisions, price formation, and competition dynamics.
- Firm strategy and sources of advantage
- If foresight becomes more widely accessible via high‑performing LLMs, the value of pure predictive skill may commoditize.
- Competitive advantage is likely to shift toward complements: proprietary data, domain‑specific representations, problem framing, institutional processes that act on and learn from AI forecasts, and execution capabilities that convert forecasts into superior outcomes.
- Labor and organization
- Roles centered on unassisted strategic prediction may shrink or be reshaped; managerial value may increase in areas requiring judgment about using, contextualizing, or operationalizing model outputs.
- The limited benefit from simple human–AI aggregation suggests firms should invest in organizational design and governance to harness AI forecasts rather than rely on naïve human + model averaging.
- Policy and market design
- Regulators and platform designers should consider disclosure, accountability, and competitive implications of high‑accuracy AI forecasting tools—especially where forecasts affect allocations at scale (finance, venture funding, procurement).
- Benchmarks like this paper’s prospective design are valuable for evaluating model capability and for informing safe deployment standards.
- Research agendas and caveats
- Generalizability: results are strong for Kickstarter-style, market-driven venture outcomes but may not generalize automatically to strategic domains with high endogeneity, deliberate adversaries, or where actions substantially change the environment.
- Feedback and equilibrium effects: as firms adopt LLM foresight, behavior changes may erode historical regularities that models exploit (Goodhart-like effects), creating new dynamics requiring ongoing evaluation.
- Open questions: domain transferability, dynamic decision-making (sequential actions), human–AI interaction design that yields complementarity (not just aggregation), and the role of proprietary data in amplifying or mitigating model advantage.
- Practical takeaway for economists and managers
- High‑performing LLMs are a potent forecasting technology for many strategic foresight tasks; firms should prioritize (i) validating model accuracy in their domain, (ii) investing in the data and organizational routines that turn forecasts into profitable action, and (iii) monitoring for distributional and equilibrium effects when forecasts are deployed at scale.
Assessment
Claims (9)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| We conducted a fully prospective prediction tournament using live Kickstarter crowdfunding projects. Decision Quality | null_result | forecast accuracy for Kickstarter fundraising outcomes |
Reading fidelity
high
Study strength
high
|
n=30
|
| Thirty U.S.-based technology ventures, launched after the training cutoffs of all models studied, were evaluated while fundraising remained in progress and outcomes were unknown. Decision Quality | null_result | forecast accuracy for Kickstarter fundraising outcomes |
Reading fidelity
high
Study strength
high
|
n=30
|
| A diverse suite of frontier and open-weight large language models (LLMs) completed 870 pairwise comparisons, producing complete rankings of predicted fundraising success. Decision Quality | null_result | pairwise prediction/ranking of fundraising success |
Reading fidelity
high
Study strength
high
|
n=870
|
| We benchmarked these forecasts against 346 experienced managers recruited via Prolific and three MBA-trained investors working under monitored conditions. Decision Quality | null_result | forecast accuracy for Kickstarter fundraising outcomes by human evaluators |
Reading fidelity
high
Study strength
high
|
n=346
|
| Human evaluators achieved rank correlations with actual outcomes between 0.04 and 0.45. Decision Quality | positive | rank correlation between predicted and actual fundraising outcomes |
Reading fidelity
high
Study strength
medium
|
n=346
0.04 to 0.45 (rank correlation)
|
| Several frontier LLMs exceeded 0.60 rank correlation with actual outcomes. Decision Quality | positive | rank correlation between model-predicted and actual fundraising outcomes |
Reading fidelity
high
Study strength
medium
|
greater than 0.60 (rank correlation)
|
| The best model (Gemini 2.5 Pro) reached 0.74 rank correlation—correctly ordering nearly four of every five venture pairs. Decision Quality | positive | rank correlation (0.74) between Gemini 2.5 Pro predictions and actual fundraising outcomes; percent of correctly ordered venture pairs |
Reading fidelity
high
Study strength
medium
|
0.74 (rank correlation); correctly ordering nearly 4 of every 5 venture pairs
|
| These differences persist across multiple performance metrics and robustness checks. Decision Quality | positive | forecasting performance across multiple metrics and robustness checks |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Neither wisdom-of-the-crowd ensembles nor human-AI hybrid teams outperformed the best standalone model. Decision Quality | null_result | comparison of ensemble and hybrid team forecasting performance versus best standalone model |
Reading fidelity
high
Study strength
medium
|
not reported
|