The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

A specialized live benchmark shows state-of-the-art agentic LLMs handle open-domain search but lack the accuracy and domain grounding required for high-value vertical prediction tasks such as market forecasting, supply-chain demand, epidemic tracking and disaster monitoring; the gap raises caution for immediate industrial deployment.

FutureX-Pro: Extending Future Prediction to High-Value Vertical Domains
Jiashuo Liu, Siyuan Chen, Zaiyuan Wang, Zhiyuan Zeng, Jiacheng Guo, Liang Hu, Lingyue Yin, Suozhi Huang, Wenxin Hao, Yang Yang, Zerui Cheng, Zixin Yao, Lingyue Yin, Haoxin Liu, Jiayi Cheng, Yuzhen Li, Zezhong Ma, Bingjie Wang, Bingsen Qiu, Xiao Liu, Zeyang Zhang, Zijian Liu, Jinpeng Wang, Mingren Yin, Tianci He, Yali Liao, Yixiao Tian, Zhenwei Zhu, Anqi Dai, Ge Zhang, Jingkai Liu, Kaiyuan Zhang, Wenlong Wu, Xiang Gao, Xinjie Chen, Zhixin Yao, Zhoufutu Wen, B. Aditya Prakash, Jose Blanchet, Mengdi Wang, Nian Si, Wenhao Huang · January 18, 2026
arxiv descriptive n/a evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Jiashuo Liu unresolved corpus identity
  2. Siyuan Chen unresolved corpus identity
  3. Zaiyuan Wang unresolved corpus identity
  4. Zhiyuan Zeng unresolved corpus identity
  5. Jiacheng Guo unresolved corpus identity
  6. Liang Hu unresolved corpus identity
  7. Lingyue Yin unresolved corpus identity
  8. Suozhi Huang unresolved corpus identity
  9. Wenxin Hao unresolved corpus identity
  10. Yang Yang unresolved corpus identity
  11. Zerui Cheng unresolved corpus identity
  12. Zixin Yao unresolved corpus identity
  13. Lingyue Yin unresolved corpus identity
  14. Haoxin Liu unresolved corpus identity
  15. Jiayi Cheng unresolved corpus identity
  16. Yuzhen Li unresolved corpus identity
  17. Zezhong Ma unresolved corpus identity
  18. Bingjie Wang unresolved corpus identity
  19. Bingsen Qiu unresolved corpus identity
  20. Xiao Liu unresolved corpus identity
  21. Zeyang Zhang unresolved corpus identity
  22. Zijian Liu unresolved corpus identity
  23. Jinpeng Wang unresolved corpus identity
  24. Mingren Yin unresolved corpus identity
  25. Tianci He unresolved corpus identity
  26. Yali Liao unresolved corpus identity
  27. Yixiao Tian unresolved corpus identity
  28. Zhenwei Zhu unresolved corpus identity
  29. Anqi Dai unresolved corpus identity
  30. Ge Zhang unresolved corpus identity
  31. Jingkai Liu unresolved corpus identity
  32. Kaiyuan Zhang unresolved corpus identity
  33. Wenlong Wu unresolved corpus identity
  34. Xiang Gao unresolved corpus identity
  35. Xinjie Chen unresolved corpus identity
  36. Zhixin Yao unresolved corpus identity
  37. Zhoufutu Wen unresolved corpus identity
  38. B. Aditya Prakash unresolved corpus identity
  39. Jose Blanchet unresolved corpus identity
  40. Mengdi Wang unresolved corpus identity
  41. Nian Si unresolved corpus identity
  42. Wenhao Huang unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Jiashuo Liu provider ID
  2. Siyuan Chen provider ID
  3. Zaiyuan Wang provider ID
  4. Zhiyuan Zeng provider ID
  5. Jiacheng Guo provider ID
  6. Liang Hu provider ID
  7. Lingyue Yin provider ID
  8. Suozhi Huang provider ID
  9. Wenxing Hao provider ID
  10. Yang Yang provider ID
  11. Zerui Cheng provider ID
  12. Zixin Yao provider ID
  13. Haoxin Liu provider ID
  14. Jiayi Cheng provider ID
  15. Yuzhen Li provider ID
  16. Zezhong Ma provider ID
  17. Bingjie Wang provider ID
  18. Bingsen Qiu provider ID
  19. Xiao Liu provider ID
  20. Zeyang Zhang provider ID
  21. Zijian Liu provider ID
  22. Jinpeng Wang provider ID
  23. Mingren Yin provider ID
  24. Tianci He provider ID
  25. Yali Liao provider ID
  26. Yi Tian provider ID
  27. Zhenwei Zhu provider ID
  28. Anqi Dai provider ID
  29. Ge Zhang provider ID
  30. Jingkai Liu provider ID
  31. Kai Zhang provider ID
  32. Wenlong Wu provider ID
  33. Xiang Gao provider ID
  34. Xinjie Chen provider ID
  35. Zhixin Yao provider ID
  36. Zhou Wen provider ID
  37. B. Prakash provider ID
  38. Jose H. Blanchet provider ID
  39. Mengdi Wang provider ID
  40. Nian Si provider ID
  41. Wenhao Huang provider ID
FutureX-Pro extends a live, contamination-free benchmark into five high-value verticals and finds that current SOTA agentic LLMs fall short of the precision and reliability needed for finance, retail, public health, and natural-disaster forecasting tasks.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Building upon FutureX, which established a live benchmark for general-purpose future prediction, this report introduces FutureX-Pro, including FutureX-Finance, FutureX-Retail, FutureX-PublicHealth, FutureX-NaturalDisaster, and FutureX-Search. These together form a specialized framework extending agentic future prediction to high-value vertical domains. While generalist agents demonstrate proficiency in open-domain search, their reliability in capital-intensive and safety-critical sectors remains under-explored. FutureX-Pro targets four economically and socially pivotal verticals: Finance, Retail, Public Health, and Natural Disaster. We benchmark agentic Large Language Models (LLMs) on entry-level yet foundational prediction tasks -- ranging from forecasting market indicators and supply chain demands to tracking epidemic trends and natural disasters. By adapting the contamination-free, live-evaluation pipeline of FutureX, we assess whether current State-of-the-Art (SOTA) agentic LLMs possess the domain grounding necessary for industrial deployment. Our findings reveal the performance gap between generalist reasoning and the precision required for high-value vertical applications.

Summary

Main Finding

FutureX-Pro extends the original FutureX live, contamination-free benchmark into high-value verticals (Finance, Retail, Public Health, Natural Disaster) plus a retrieval-oriented FutureX-Search. Across domain-specific, entry-level prediction tasks, state-of-the-art agentic LLMs (e.g., GPT-5 variants, Grok-4) show competence in broad reasoning and calibrated uncertainty but fall short of the numerical precision, reliability, and domain grounding required for capital-intensive or safety-critical industrial deployment.

Key Points

  • Scope and motivation
    • Focuses on four verticals with high economic or societal impact: Finance (market indicators), Retail (demand forecasting), Public Health (epidemiological metrics), Natural Disaster (hazard impacts). Adds FutureX-Search to turn resolved events into pure retrieval tasks.
    • Uses a contamination-impossible live-evaluation pipeline to ensure verifiable, out-of-sample future events.
  • Finance (FutureX-Finance)
    • Dataset: 150 equities (100 US — NASDAQ/S&P constituents; 50 China A-shares).
    • Tasks: Spot prediction (single-day exact value), Window extremum (max absolute value across N trading days), Directional momentum (largest positive change in window).
    • Metric: High-sensitivity linear penalty S = max(0, 1 − 20 × |ŷ − y| / |y|). Errors >5% score zero (strict tolerance).
    • Result highlights: Performance tiers emerged; GPT-5-High and Grok-4 lead (Type 1 scores ≈46.37 and 41.25 respectively). No model averaged above 50 — indicating substantial difficulty converting information into <5% numerical forecasts.
  • Retail (FutureX-Retail)
    • Dataset: 240 cross-border e-commerce products (4 major categories, 24 subcategories) sourced from Temu.
    • Task grid: 2 (input conditions: snapshot-only vs. sparse time-series) × 3 (output granularity: point, top‑K probabilistic, full distribution).
    • Scoring: Hybrid — linear-decay relative-error metric for exact ground truth (ε = 0.05) and binary range checks for coarse badge-like counts (>999). Probabilistic tasks scored as expected value under predicted distribution.
    • Results highlights:
      • "Probabilistic advantage": models often score better when asked for distributions (Top‑3 or full distribution) than single-point estimates.
      • Historical context substantially improves accuracy.
      • In-context examples strongly influence output granularity and reported confidence (models mimic example counts/probabilities).
      • Model personalities: Grok-4 tends to output long-tail distributions; GPT-series tends to be more conservative/calibrated.
  • Public Health (FutureX-PublicHealth)
    • Data strictly from authoritative national bulletins (US CDC, China CDC).
    • Scope: 70 weekly event templates → 382 variables with high temporal frequency (weekly).
    • Goal: Evaluate agents on structured epidemiological report interpretation and near-term metric tracking (high temporal resolution).
  • FutureX-Search
    • Transforms predictive tasks with known outcomes into retrieval/search challenges to benchmark pure information retrieval and grounding separately from forecasting ability.
  • Models evaluated
    • Broad evaluation across ~18 leading agentic LLMs (GPT-5.1/GPT-5 variants, Grok-4, Claude-Opus, Kimi-K2 variants, Qwen3-Max, DeepSeek, GLM, etc.). SOTA models show relative stability but absolute precision remains limited for high-stakes thresholds.

Data & Methods

  • Contamination-free live evaluation
    • All tasks are constructed to be verifiably about future events so model training data contamination is effectively prevented.
    • Live updates and a weekly competition pipeline (Hugging Face dataset integration).
  • Domain datasets and design
    • Finance: 150 tracked companies, sector-balanced across US/China; tasks mimic junior quantitative analyst workflows.
    • Retail: 240 products, HTML snapshots (T−7) and sparse time-series input (T−14) used to simulate proprietary/ephemeral sales data scenarios.
    • Public Health: Official weekly bulletins only; fine-grained demographic and geographic stratifications.
    • Natural Disaster and FutureX-Search are defined at the benchmark level (Natural Disaster aims at meteorological/hazard forecasts; FutureX-Search uses historical resolved events), though full construction details are in the complete report and prior FutureX methodology.
  • Task formats and evaluation
    • Finance: strict linear penalty with 5% tolerance cutoff; distinct task types to probe point vs. temporal/extrema reasoning.
    • Retail: deterministic, top‑K probabilistic, and full-distribution forecasts; scoring accounts for exact vs. coarse labels and rewards proper probability mass assignment.
    • Public Health: high-frequency weekly variables; evaluation focuses on correct recent-period identification and numerical tracking.
  • Experimental protocol
    • Agents were given search tools in many evaluations to simulate agentic behavior (retrieval + reasoning).
    • Analyses included ablations on historical context, in-context learning sensitivity (few-shot example design), and cross-model behavior comparisons.

Implications for AI Economics

  • Economic value vs. required precision
    • High-value verticals impose narrow precision requirements (e.g., <5% for market predictions) where current SOTA models are not yet reliable; errors translate into real economic losses if used directly for trading, inventory optimization, or health responses.
    • The observed "probabilistic advantage" suggests commercial value in supplying calibrated distributions (risk-aware decision support) rather than brittle point predictions—this is especially useful for inventory managers, risk desks, and public-health triage.
  • Automation opportunity and limits
    • Agents show promise for junior-analyst augmentation (information synthesis, preliminary screening, uncertainty quantification) but not for fully autonomous decision-making in capital- or safety-critical workflows without additional safeguards, domain-specific models, or human oversight.
  • Importance of domain grounding and authoritative data
    • Performance gains from historical/contextual inputs and from access to authoritative structured sources imply that firms should prioritize integrating vetted domain data pipelines, specialized tools (time-series modules, epidemiological models), and domain-specific fine-tuning to reach production-grade reliability.
  • Market and policy considerations
    • Firms deploying agentic LLMs in finance, health, or disaster response will need robust model evaluation, explainability, and regulatory-compliant validation processes; live contamination-free benchmarks like FutureX-Pro can serve as an independent validation tool.
    • Policymakers should treat generalist-agent outputs as decision-support (not authoritative) and consider standards for model accuracy, auditability, and human-in-the-loop requirements in high-stakes domains.
  • Research directions that matter economically
    • Improving numerical precision, temporal-extrema reasoning, and robust calibration under sparse/proprietary data regimes.
    • Hybrid architectures combining LLM reasoning with domain-specific model components (statistical time-series, epidemiological simulators, meteorological ensembles).
    • Better in-context and tool-use protocols to reduce sensitivity to example design and to stabilize outputs for operational adoption.
  • Value of benchmarks
    • FutureX-Pro demonstrates that live, uncontaminated, domain-focused benchmarks are essential for measuring economically relevant progress and for de-risking adoption decisions by firms and regulators.

If you want, I can: - Produce a one-page slide-ready summary of this benchmark for stakeholders (investors, product managers, regulators). - Extract a short list of recommended operational guards and engineering practices for deploying agentic LLMs in each vertical.

Assessment

Paper Typedescriptive Evidence Strengthn/a — This is a benchmark/evaluation study rather than a causal inference study; it documents model performance on prediction tasks but does not attempt to identify causal effects of AI on economic outcomes. Methods Rigormedium — The report adapts a contamination-free, live-evaluation pipeline which increases realism and reduces data leakage, and covers multiple verticals and tasks; however it appears limited to entry-level prediction tasks, lacks detail on model selection, prompt/agent configurations, metrics and statistical uncertainty, and does not validate against downstream economic decisions, reducing overall methodological conclusiveness. SampleA multi-vertical live benchmark (FutureX-Pro) composed of specialized subsets: FutureX-Finance, FutureX-Retail, FutureX-PublicHealth, FutureX-NaturalDisaster, and FutureX-Search; evaluates several state-of-the-art agentic LLMs on entry-level forecasting tasks (market indicators, supply-chain demand, epidemic trends, natural disaster tracking, and search-related predictions) using a contamination-free live-evaluation pipeline and real-time ground truth streams (exact model list, sample sizes, time windows, geographic scope, and metric definitions not specified in the summary). Themesadoption productivity human_ai_collab innovation GeneralizabilityTasks are entry-level and may not capture full complexity of production/industrial decision processes, Results reflect specific SOTA agentic LLMs and configurations at evaluation time; model updates could change outcomes, Verticals covered (finance, retail, public health, natural disaster, search) exclude many industry-specific constraints and regulatory/local variations, Live-evaluation realism may still omit organizational integration, downstream human oversight, and economic outcome measures, Benchmark outcomes depend on ground-truth sources and time windows; performance may vary across geographies and market regimes

Claims (6)

ClaimDirectionOutcomeConfidence & EvidenceDetails
This report introduces FutureX-Pro, a specialized framework extending agentic future prediction to high-value vertical domains. Adoption Rate positive existence and scope of the FutureX-Pro framework
Reading fidelity high
Study strength high
not reported
0.3
FutureX-Pro includes FutureX-Finance, FutureX-Retail, FutureX-PublicHealth, FutureX-NaturalDisaster, and FutureX-Search as domain-specific components. Adoption Rate positive coverage of verticals in the benchmark suite
Reading fidelity high
Study strength high
not reported
0.3
We benchmark agentic Large Language Models (LLMs) on entry-level yet foundational prediction tasks — ranging from forecasting market indicators and supply chain demands to tracking epidemic trends and natural disasters. Output Quality neutral forecasting/prediction performance on domain-specific tasks
Reading fidelity high
Study strength medium
not reported
0.18
Generalist agents demonstrate proficiency in open-domain search, but their reliability in capital-intensive and safety-critical sectors remains under-explored. Adoption Rate mixed reliability / evaluation coverage in capital-intensive and safety-critical sectors
Reading fidelity high
Study strength medium
not reported
0.18
By adapting the contamination-free, live-evaluation pipeline of FutureX, we assess whether current SOTA agentic LLMs possess the domain grounding necessary for industrial deployment. Adoption Rate neutral domain grounding / readiness for industrial deployment as assessed by live evaluation
Reading fidelity high
Study strength medium
not reported
0.18
Our findings reveal the performance gap between generalist reasoning and the precision required for high-value vertical applications. Output Quality negative difference between generalist reasoning performance and domain-required precision (forecasting accuracy / reliability)
Reading fidelity high
Study strength medium
not reported
0.18

Notes