The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Human teams using a Hyperchat AI deliberation platform forecast NBA point-spreads more accurately than Vegas odds over 50 games, yielding an 18.4% hypothetical ROI; greater conversational engagement predicted stronger performance, though the single-arm, fan-based study limits causal claims and wider applicability.

Conversational Forecasting Across Large Human Groups Using A Swarm of Surrogate AI Agents
Louis Rosenberg, Hans Schumann, Ganesh Mani, Gregg Willcox · February 20, 2026
arxiv quasi_experimental low evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Louis Rosenberg unresolved corpus identity
  2. Hans Schumann unresolved corpus identity
  3. Ganesh Mani unresolved corpus identity
  4. Gregg Willcox unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Louis B. Rosenberg provider ID
  2. Hans Schumann provider ID
  3. Ganesh Mani provider ID
  4. G. Willcox provider ID
Teams of 25–30 basketball fans using a Hyperchat AI platform achieved 62% accuracy forecasting 50 NBA spreads over 12 weeks—outperforming Vegas odds and yielding a hypothetical 18.4% ROI—with higher conversational activity correlated with better predictions, but the study lacks randomized controls and a broader sample.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Hyperchat AI is a communication and collaboration architecture that employs intervening AI agents to enable real-time conversational deliberations among networked human teams of unlimited size. Prior work has shown that teams as large as 250 people can hold productive real-time conversations by text, voice, or video using Hyperchat AI to discuss complex problems, brainstorm solutions, surface risks, assess alternatives, prioritize options, and converge on optimized results. Building on this prior work, this new study tasked groups of 25 to 30 basketball fans with conversationally forecasting NBA games (against the spread) over a 12-week period. Results show that when discussing and debating NBA games (for five minutes each) using a Hyperchat AI enabled platform called Thinkscape, human teams were 62% accurate across a set of 50 forecasted NBA games. This is an impressive result versus the Vegas odds of 50% (p=0.059). Furthermore, had the participants wagered on the games, they would have produced an 18.4% ROI over the 12-week period. In addition, this study found that the group's conversation rate during each forecast was positively correlated with their prediction accuracy. In fact, when excluding the 12 forecasts in the bottom 25th percentile by average conversation rate, the remaining 38 forecasts recorded a 68% accuracy, significantly better than the 50% Vegas odds (p=0.017). This result also outperformed the well-known prediction market Polymarket (p=0.062) across the same set of NBA games. These outcomes suggest that real-time conversational deliberations, when facilitated by Surrogate AI agents, can significantly amplify groupwise collective intelligence during human forecasting tasks.

Summary

Main Finding

Hyperchat AI — a conversational, agent-mediated platform (Thinkscape) that links many small human subgroups via a swarm of AI “conversational surrogate” agents — materially amplified group forecasting accuracy on NBA point-spread bets. Across 56 collaborative forecasts (50 scored), groups of 25–30 sports fans deliberating for ~5 minutes per game achieved 62% accuracy (31/50), outperforming a 50% baseline (p = 0.059). Excluding the 12 forecasts in the bottom conversation-rate quartile raised accuracy to 68% (38 forecasts; p = 0.017). If participants had bet at standard –110 lines, the group outcome implied an 18.4% ROI over the 12-week period.

Key Points

  • Sample and performance
    • 56 forecasts attempted over 12 weeks; 6 classified as toss-ups → 50 scored forecasts.
    • Group (Hyperchat AI) accuracy: 62% (31/50); p = 0.059 vs 50% Vegas baseline.
    • Removing lowest-25%-by-conversation-rate forecasts: 68% accuracy (38 forecasts); p = 0.017 vs 50%.
    • Compared to individuals: mean individual accuracy ≈ 52%; Hyperchat groups outperformed 36/43 individuals (84%); bootstrapped comparison p ≈ 0.071.
    • Compared to Polymarket on the same games: group outperformed (reported p = 0.062).
  • Economic outcome
    • Using standard sportsbook payouts (–110), the group’s picks imply an 18.4% ROI for the 12-week sample.
  • Relation to conversation dynamics
    • Higher conversational engagement (conversation rate) correlated positively with forecasting accuracy; low-engagement sessions drove much of the error.
  • Platform behavior and design features
    • Participants were split into 5–6 small “Thinktanks” (4–5 people each).
    • Each Thinktank included a surrogate AI agent that (a) monitors local conversation, (b) extracts and expresses local insights, (c) shares novel/challenging insights from other Thinktanks.
    • A Deliberative Matching Engine (DME) prioritized sharing novel and maximally challenging insights to foster productive push–pull deliberation and mitigate early social influence.
    • Real-time LLM processing inferred for each participant: preferred option (team + wager level), conviction strength, and rationale; these produced per-user Support Values (0–100% across four options).
    • A Weighted Collective Forecast (WCF) aggregated supports into a -2 to +2 scale to produce the group pick and confidence.

Data & Methods

  • Participants
    • Recruited via Prolific; self-identified “basketball fans.”
    • Payment ≈ $6.50 per forecasting session.
    • 43 participants contributed at least 12 forecasts across the study.
  • Experimental setup
    • Thinkscape platform implemented text-chat deliberations (anonymity preserved), 4 games per ~20-minute session, ~5 minutes deliberation per game.
    • Each forecast question included the two teams, the current DraftKings spread, and four betting-choice options (Team A/B × risk 10 or 20 points) to elicit both direction and confidence.
    • LLM processed live chat to tag each participant’s momentary support among the four options; these support values were aggregated to form the WCF.
    • Toss-up threshold: WCF in [-0.08, +0.08] (6 toss-ups).
  • Analysis
    • Primary comparison: group accuracy vs. 50% baseline implied by spread (binomial testing reported).
    • Bootstrapping of individual forecasts (10,000 resamples) used to compare aggregated individual performance to group performance.
    • Subgroup analyses: favorite vs underdog accuracy reported (favorites 59%, underdogs 73%).
    • Reported p-values: group vs baseline (p=0.059); group excluding lowest conversation quartile vs baseline (p=0.017); group vs Polymarket (p=0.062); bootstrapped individuals vs group (p≈0.071).
  • System mechanics of note
    • DME ensures novel and challenging cross-group transfers to speed deliberation and surface counterarguments.
    • Support Values per user sum to 100% across four options; WCF mapping converts supports to signed score indicating collective leaning and strength.

Implications for AI Economics

  • Value creation and arbitrage potential
    • Agent-mediated conversational deliberation produced a nontrivial ROI in this sports forecasting sample, suggesting practical economic value from AI-facilitated collective forecasting that could be monetized (betting, trading, corporate forecasting).
    • If robust and reproducible, such systems could enable coordinated groups to systematically capture mispricings in prediction markets or betting markets, with implications for market efficiency and liquidity.
  • Organizational forecasting and productivity
    • Hyperchat-style architectures may increase accuracy of internal forecasts (sales, demand, risk) while preserving scale; the approach could substitute for or augment traditional elicitation (surveys, polls) by enabling rapid deliberative aggregation of diverse expertise.
    • The staged information-sharing design (local → regional → global) is an economically important mechanism: it mitigates premature conformity and potentially improves the information content of aggregate signals.
  • Labor and coordination economics
    • Surrogate agents reduce the coordination costs of large-group deliberation, effectively amplifying marginal productivity of dispersed human contributors. This changes the cost structure of collective forecasting: fixed costs of platform/agents vs variable labor inputs.
    • There are substitution and complementarity effects: agents complement human reasoning (curation, routing, synthesis) rather than fully replacing domain expertise.
  • Market design and policy considerations
    • Widespread use could alter how prices reflect information (faster/stronger aggregation), but also raise concerns of coordinated strategies that could be used to manipulate thin markets or exploit insider-like advantages.
    • Transparency, auditability, and reproducibility of agent decision rules and inference models (LLMs) become important for regulatory oversight if such systems interact with financial/political markets.
  • Research and deployment priorities for economists
    • Need cost–benefit analyses: platform build/operation costs, participant recruitment incentives, and expected economic returns across domains beyond sports.
    • Replication and generalizability: test on larger samples, diverse domains (macroeconomic forecasting, sales forecasting), and against more/broader baselines (prediction markets, professional forecasters).
    • Study externalities: how agent-mediated groups interact with prediction-market prices, possibility of endogenous equilibrium effects, and strategic behavior when agents or groups can participate in markets.
  • Limitations affecting economic interpretation
    • Sample size and domain specificity (sports): results may not generalize to other forecasting domains with different signal structures or incentives.
    • Marginal p-values and limited number of scored forecasts caution against overclaiming robustness.
    • Participant pool (paid Prolific users, self-identified fans) differs from professional or market-participating forecasters; hence external validity to market-arbitrage contexts is uncertain.
    • Proprietary/platform effects: LLM processing, DME heuristics, and surrogate-agent behavior are not fully open, which complicates reproducibility and economic valuation.

Suggested next steps for economic research and deployment - Larger-scale replication across multiple domains and longer horizons. - Controlled comparisons vs established baselines (professional forecasters, larger prediction markets) and against other coordination mechanisms (Delphi, tournaments). - Detailed cost accounting to compute net expected economic returns per unit of investment (platform costs, incentives). - Analysis of strategic dynamics if such groups enter real markets (market impact, detectability, regulation).

Assessment

Paper Typequasi_experimental Evidence Strengthlow — The study reports statistically suggestive differences versus market benchmarks but relies on a single-arm design with a small number of events (50 games), marginal p-values, potential selection and novelty effects, and no randomized or placebo control to establish causality or rule out alternative explanations. Methods Rigorlow — Analyses appear limited to comparisons versus external benchmarks and correlational tests; key design elements are missing or unclear (randomization, pre-registration, power calculations, handling of multiple comparisons, participant selection procedures), raising risks of bias, confounding, and overclaiming of causal effects. SampleHuman teams composed of basketball fans (groups of roughly 25–30 participants) used the Thinkscape Hyperchat AI platform to jointly forecast 50 NBA games across a 12-week period; outcomes compared to Vegas point-spread implied probabilities and contemporaneous Polymarket prices; hypothetical wagering ROI was computed ex post. Participants appear non-random, self-selected (fan base), and the number of distinct teams/sessions beyond the reported group size is not fully specified. Themeshuman_ai_collab productivity adoption IdentificationSingle-arm intervention: teams used the Hyperchat AI (Thinkscape) platform and their forecast accuracy was compared to external benchmarks (Vegas point-spread odds and Polymarket prices); correlation analysis between conversation rate and accuracy was used to support mechanism claims. No randomized control group, no pre-registered counterfactual assignment, and no within-study randomization of treatment intensity. Generalizabilitydomain_limited: forecasts restricted to NBA point-spread betting and may not generalize to other forecasting or economic tasks, sample_limited: participants were basketball fans (self-selected) rather than a representative population, small_event_sample: only 50 games over 12 weeks, limiting statistical power and robustness, no_randomization: lack of control group reduces causal generalizability, short_time_horizon: 12 weeks may include novelty or learning effects that differ over longer horizons, geographic_demographics_unclear: participant demographics and location not reported, limiting cross-population generalization

Claims (9)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Prior work has shown that teams as large as 250 people can hold productive real-time conversations by text, voice, or video using Hyperchat AI to discuss complex problems, brainstorm solutions, surface risks, assess alternatives, prioritize options, and converge on optimized results. Team Performance positive ability to hold productive real-time conversations (productivity of team deliberation)
Reading fidelity high
Study strength low
n=250
0.24
In a new study, groups of 25 to 30 basketball fans conversationally forecasted NBA games (against the spread) over a 12-week period. Other null_result study design / participant task (forecasting NBA games)
Reading fidelity high
Study strength medium
not reported
0.48
When discussing and debating NBA games (for five minutes each) using a Hyperchat AI enabled platform called Thinkscape, human teams were 62% accurate across a set of 50 forecasted NBA games. Decision Quality positive forecast accuracy (proportion correct vs. spread)
Reading fidelity high
Study strength medium
n=50
62% accurate
0.48
The teams' 62% accuracy is an impressive result versus the Vegas odds of 50% (p=0.059). Decision Quality positive difference in accuracy relative to Vegas baseline
Reading fidelity high
Study strength medium
n=50
62% vs 50% (p=0.059)
0.48
Had the participants wagered on the games, they would have produced an 18.4% ROI over the 12-week period. Other positive return on investment (hypothetical wagering)
Reading fidelity high
Study strength medium
n=50
18.4% ROI
0.48
The group's conversation rate during each forecast was positively correlated with their prediction accuracy. Decision Quality positive relationship between conversation rate and forecast accuracy
Reading fidelity high
Study strength medium
n=50
0.48
Excluding the 12 forecasts in the bottom 25th percentile by average conversation rate, the remaining 38 forecasts recorded a 68% accuracy, significantly better than the 50% Vegas odds (p=0.017). Decision Quality positive forecast accuracy in subset of higher-conversation-rate forecasts
Reading fidelity high
Study strength medium
n=38
68% accuracy (p=0.017)
0.48
The 68% accuracy result (excluding low conversation-rate forecasts) also outperformed the well-known prediction market Polymarket (p=0.062) across the same set of NBA games. Decision Quality positive performance comparison vs. Polymarket on same games
Reading fidelity high
Study strength low
n=38
p=0.062
0.24
These outcomes suggest that real-time conversational deliberations, when facilitated by Surrogate AI agents, can significantly amplify groupwise collective intelligence during human forecasting tasks. Decision Quality positive amplification of group collective intelligence in forecasting
Reading fidelity high
Study strength speculative
n=50
0.08

Notes