The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

State-of-the-art LLMs out-forecast humans on startup crowdfunding: Gemini 2.5 Pro correctly orders about four in five venture pairs, beating 346 managers and MBA investors in a live Kickstarter tournament.

The Strategic Foresight of LLMs: Evidence from a Fully Prospective Venture Tournament
Csaszar, Felipe A., Peterson, Aticus, Wilde, Daniel · February 02, 2026 · arXiv (Cornell University)
openalex other medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

OpenAlex

Latest observation:

  1. Csaszar, Felipe A. provider ID
  2. Peterson, Aticus provider ID
  3. Wilde, Daniel provider ID

Semantic Scholar

Latest observation:

  1. Felipe A. Csaszar provider ID
  2. Aticus Peterson provider ID
  3. D. Wilde provider ID
In a prospective forecasting tournament of 30 live Kickstarter tech projects, several frontier LLMs—most notably Gemini 2.5 Pro—substantially outperformed human evaluators at ranking fundraising success, with the top model achieving a rank correlation of 0.74 versus humans' 0.04–0.45.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Can artificial intelligence outperform humans at strategic foresight -- the capacity to form accurate judgments about uncertain, high-stakes outcomes before they unfold? We address this question through a fully prospective prediction tournament using live Kickstarter crowdfunding projects. Thirty U.S.-based technology ventures, launched after the training cutoffs of all models studied, were evaluated while fundraising remained in progress and outcomes were unknown. A diverse suite of frontier and open-weight large language models (LLMs) completed 870 pairwise comparisons, producing complete rankings of predicted fundraising success. We benchmarked these forecasts against 346 experienced managers recruited via Prolific and three MBA-trained investors working under monitored conditions. The results are striking: human evaluators achieved rank correlations with actual outcomes between 0.04 and 0.45, while several frontier LLMs exceeded 0.60, with the best (Gemini 2.5 Pro) reaching 0.74 -- correctly ordering nearly four of every five venture pairs. These differences persist across multiple performance metrics and robustness checks. Neither wisdom-of-the-crowd ensembles nor human-AI hybrid teams outperformed the best standalone model.

Summary

Main Finding

Frontier large language models substantially outperformed experienced human evaluators in a fully prospective strategic-foresight benchmark. In a live prediction tournament over 30 U.S. technology Kickstarter projects (sampled after models' training cutoffs), several LLMs achieved rank correlations with realized fundraising outcomes > 0.60; the best model (Gemini 2.5 Pro) reached 0.74 (≈79% of pairwise orderings correct). Human managers (n=346) and three MBA investors achieved rank correlations between 0.04 and 0.45 (top human ≈60% pairwise accuracy). Ensembles and human–AI hybrids did not beat the best standalone model.

Key Points

  • Prospective, real-world test: predictions were made while projects were fundraising and before outcomes were known; realized total funds raised provided objective ground truth.
  • Task design: 30 Kickstarter technology ventures compared in a double round‑robin pairwise tournament (870 comparisons) to construct complete rankings of expected success.
  • Models evaluated: a diverse suite of frontier and open-weight LLMs (GPT-5 variants, Claude 4.5, Gemini 2.5 variants, Grok 4, Gemma 3, Llama 3.1, DeepSeek 3.2, etc.).
  • Human benchmarks: 346 employed U.S. managers recruited on Prolific plus three monitored MBA-trained investors performed the identical pairwise task.
  • Performance gap: several LLMs >0.60 rank correlation; Gemini 2.5 Pro = 0.74. Human rank correlations ranged 0.04–0.45. Differences are statistically significant and robust across metrics and checks.
  • Aggregation: neither “wisdom of silicon” ensembles nor aggregated human–AI teams outperformed the best single LLM.
  • Method contribution: pairwise tournament scoring (double round‑robin) is an effective elicitation method for stable ranking from comparative judgments.
  • Exploratory analysis: cross-model performance regressed on standard AI capability benchmarks to investigate drivers of LLM fore­sight performance.

Data & Methods

  • Sample: 30 U.S.-based technology projects on Kickstarter; each project was launched after the training cutoff dates of all LLMs studied and was evaluated while fundraising remained in progress.
  • Outcome measure: realized total funds raised at campaign close (objective market outcome).
  • Prediction elicitation:
    • LLMs: each model performed all pairwise comparisons between the 30 projects in both presentation orders (30 choose 2 = 435 pairs × 2 = 870 comparisons), producing a complete rank ordering via tournament aggregation.
    • Humans: 346 experienced managers on Prolific completed the identical pairwise comparison task; three MBA-trained investors produced full rankings under monitored, no-technology conditions.
  • Metrics: primary metric was rank correlation between predicted and realized rankings; pairwise accuracy (fraction of correctly ordered pairs) and other robustness checks reported.
  • Controls/robustness: presentation-order controls (double round‑robin), multiple performance metrics, tests of aggregation strategies (crowd ensembles and hybrid teams), and regressions linking model accuracy to capability benchmarks.
  • Key statistics: best LLM rank correlation = 0.74 (Gemini 2.5 Pro), several LLMs >0.60; human range 0.04–0.45. Best model ≈79% pairwise accuracy; top human ≈60%.

Implications for AI Economics

  • Forecasting and market outcomes
    • LLMs can materially improve ex ante forecasts of market-driven outcomes in strategic settings, potentially raising the precision of investment selection, product-launch evaluation, and real-time market signals.
    • Widespread use may alter market efficiency: better foresight can change allocation decisions, price formation, and competition dynamics.
  • Firm strategy and sources of advantage
    • If foresight becomes more widely accessible via high‑performing LLMs, the value of pure predictive skill may commoditize.
    • Competitive advantage is likely to shift toward complements: proprietary data, domain‑specific representations, problem framing, institutional processes that act on and learn from AI forecasts, and execution capabilities that convert forecasts into superior outcomes.
  • Labor and organization
    • Roles centered on unassisted strategic prediction may shrink or be reshaped; managerial value may increase in areas requiring judgment about using, contextualizing, or operationalizing model outputs.
    • The limited benefit from simple human–AI aggregation suggests firms should invest in organizational design and governance to harness AI forecasts rather than rely on naïve human + model averaging.
  • Policy and market design
    • Regulators and platform designers should consider disclosure, accountability, and competitive implications of high‑accuracy AI forecasting tools—especially where forecasts affect allocations at scale (finance, venture funding, procurement).
    • Benchmarks like this paper’s prospective design are valuable for evaluating model capability and for informing safe deployment standards.
  • Research agendas and caveats
    • Generalizability: results are strong for Kickstarter-style, market-driven venture outcomes but may not generalize automatically to strategic domains with high endogeneity, deliberate adversaries, or where actions substantially change the environment.
    • Feedback and equilibrium effects: as firms adopt LLM foresight, behavior changes may erode historical regularities that models exploit (Goodhart-like effects), creating new dynamics requiring ongoing evaluation.
    • Open questions: domain transferability, dynamic decision-making (sequential actions), human–AI interaction design that yields complementarity (not just aggregation), and the role of proprietary data in amplifying or mitigating model advantage.
  • Practical takeaway for economists and managers
    • High‑performing LLMs are a potent forecasting technology for many strategic foresight tasks; firms should prioritize (i) validating model accuracy in their domain, (ii) investing in the data and organizational routines that turn forecasts into profitable action, and (iii) monitoring for distributional and equilibrium effects when forecasts are deployed at scale.

Assessment

Paper Typeother Evidence Strengthmedium — The study uses a prospective, out-of-sample prediction tournament (projects launched after model training cutoffs) with many model forecasts and a large pool of human raters, providing strong within-domain evidence that some LLMs outperform humans on Kickstarter fundraising predictions. However, the number of projects is small (n=30), the domain is narrow (U.S. tech projects on Kickstarter), and the human expert sample and MBA comparator are limited, reducing confidence in broader claims. Methods Rigormedium — Design strengths include prospective forecasting on live projects, pre-specified out-of-sample targets, multiple frontier and open models, pairwise comparisons, and robustness checks including ensemble and hybrid tests. Missing or unclear elements (based on the abstract) are details on project selection procedures, pre-registration/blinding, representativeness and screening of human experts (beyond Prolific recruitment), and potential model access to auxiliary web data; these raise risks of selection and information biases. Sample30 U.S.-based technology ventures on Kickstarter launched after the models' training cutoffs; forecasts by a suite of frontier and open-weight LLMs (870 pairwise comparisons total); 346 experienced managers recruited via Prolific plus three MBA-trained investors under monitored conditions; outcome measured = realized fundraising success on Kickstarter; comparisons used rank correlations and other performance metrics. Themeshuman_ai_collab innovation GeneralizabilitySmall number of target projects (n=30) limits statistical generality., Single platform (Kickstarter) and sector (U.S. technology ventures) — results may not extend to other industries, geographic contexts, or types of strategic foresight., Outcome is fundraising success (a specific, observable metric), not broader high-stakes strategic outcomes (macroeconomic events, firm survival, long-run tech adoption)., Human comparison sample (Prolific managers and three MBAs) may not represent domain experts such as seasoned VCs or entrepreneurs., Models tested at a fixed point in time; future model versions or access to live web data could change performance., Potential differences in prompt engineering and interface across models could affect comparability.

Claims (9)

ClaimDirectionOutcomeConfidence & EvidenceDetails
We conducted a fully prospective prediction tournament using live Kickstarter crowdfunding projects. Decision Quality null_result forecast accuracy for Kickstarter fundraising outcomes
Reading fidelity high
Study strength high
n=30
0.2
Thirty U.S.-based technology ventures, launched after the training cutoffs of all models studied, were evaluated while fundraising remained in progress and outcomes were unknown. Decision Quality null_result forecast accuracy for Kickstarter fundraising outcomes
Reading fidelity high
Study strength high
n=30
0.2
A diverse suite of frontier and open-weight large language models (LLMs) completed 870 pairwise comparisons, producing complete rankings of predicted fundraising success. Decision Quality null_result pairwise prediction/ranking of fundraising success
Reading fidelity high
Study strength high
n=870
0.2
We benchmarked these forecasts against 346 experienced managers recruited via Prolific and three MBA-trained investors working under monitored conditions. Decision Quality null_result forecast accuracy for Kickstarter fundraising outcomes by human evaluators
Reading fidelity high
Study strength high
n=346
0.2
Human evaluators achieved rank correlations with actual outcomes between 0.04 and 0.45. Decision Quality positive rank correlation between predicted and actual fundraising outcomes
Reading fidelity high
Study strength medium
n=346
0.04 to 0.45 (rank correlation)
0.12
Several frontier LLMs exceeded 0.60 rank correlation with actual outcomes. Decision Quality positive rank correlation between model-predicted and actual fundraising outcomes
Reading fidelity high
Study strength medium
greater than 0.60 (rank correlation)
0.12
The best model (Gemini 2.5 Pro) reached 0.74 rank correlation—correctly ordering nearly four of every five venture pairs. Decision Quality positive rank correlation (0.74) between Gemini 2.5 Pro predictions and actual fundraising outcomes; percent of correctly ordered venture pairs
Reading fidelity high
Study strength medium
0.74 (rank correlation); correctly ordering nearly 4 of every 5 venture pairs
0.12
These differences persist across multiple performance metrics and robustness checks. Decision Quality positive forecasting performance across multiple metrics and robustness checks
Reading fidelity high
Study strength medium
not reported
0.12
Neither wisdom-of-the-crowd ensembles nor human-AI hybrid teams outperformed the best standalone model. Decision Quality null_result comparison of ensemble and hybrid team forecasting performance versus best standalone model
Reading fidelity high
Study strength medium
not reported
0.12

Notes