The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

A new evaluation framework, TruthTensor, shows that LLMs with similar forecast accuracy can diverge sharply on calibration, drift, and risk preferences when tested in 500+ live prediction markets, arguing for multi-dimensional, reproducible testing tied to real-world decision contexts.

TruthTensor: Evaluating LLMs through Human Imitation on Prediction Market under Drift and Holistic Reasoning
Shirin Shahabi, Spencer Graham, Haruna Isah · January 20, 2026
arxiv descriptive medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Shirin Shahabi unresolved corpus identity
  2. Spencer Graham unresolved corpus identity
  3. Haruna Isah unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Shirin Shahabi provider ID
  2. S. Graham provider ID
  3. Haruna Isah provider ID
TruthTensor is a reproducible evaluation paradigm that uses live prediction markets and multi-axial diagnostics (accuracy, calibration, drift, risk-sensitivity, cost) to show that models with similar forecast accuracy can behave very differently in socially-grounded, high-entropy decision settings.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Evaluating language models and AI agents remains fundamentally challenging because static benchmarks fail to capture real-world uncertainty, distribution shift, and the gap between isolated task accuracy and human-aligned decision-making under evolving conditions. This paper introduces TruthTensor, a novel, reproducible evaluation paradigm that measures reasoning models not only as prediction engines but as human-imitation systems operating in socially-grounded, high-entropy environments. Building on forward-looking, contamination-free tasks, our framework anchors evaluation to live prediction markets and combines probabilistic scoring to provide a holistic view of model behavior. TruthTensor complements traditional correctness metrics with drift-centric diagnostics and explicit robustness checks for reproducibility. It specify human vs. automated evaluation roles, annotation protocols, and statistical testing procedures to ensure interpretability and replicability of results. In experiments across 500+ real markets (political, economic, cultural, technological), TruthTensor demonstrates that models with similar forecast accuracy can diverge markedly in calibration, drift, and risk-sensitivity, underscoring the need to evaluate models along multiple axes (accuracy, calibration, narrative stability, cost, and resource efficiency). TruthTensor therefore operationalizes modern evaluation best practices, clear hypothesis framing, careful metric selection, transparent compute/cost reporting, human-in-the-loop validation, and open, versioned evaluation contracts, to produce defensible assessments of LLMs in real-world decision contexts. We publicly released TruthTensor at https://truthtensor.com.

Summary

Main Finding

TruthTensor reframes LLM evaluation from static prediction accuracy to a human-imitation, market-grounded testbed. By anchoring models to live prediction markets and measuring forward-looking probabilistic forecasts, calibration, narrative drift, and risk-sensitivity, TruthTensor shows that models with similar headline accuracy can behave very differently on economically relevant dimensions (calibration, temporal stability, and risk-aware decision quality). The framework is contamination-free, reproducible (instruction locking, versioned contracts), and released publicly at https://truthtensor.com.

Key Points

  • Evaluation objective: measure how well LLMs imitate human probabilistic reasoning and narrative updates in socially grounded, high-entropy environments (prediction markets), not just whether they hit the correct outcome.
  • Drift-centric: primary emphasis on measuring narrative drift, temporal inconsistency, and confidence decay over time—dimensions often ignored by static benchmarks.
  • Contamination-free design: only forward-looking events are used so outcomes were unknown at training time; prompts are versioned/locked to prevent prompt-engineering leakage.
  • Holistic metrics: combines accuracy with calibration (e.g., Brier/log scoring), narrative stability, risk-sensitivity, resource/cost efficiency, and behavioral/temporal diagnostics.
  • Architecture & protocol: specifies prompt templates, baseline construction, agent deployment, market-linked execution, human-in-the-loop validation, and statistical testing to ensure interpretability and reproducibility.
  • Empirical result: across 500+ real markets (political, economic, cultural, technological) models that tie on accuracy diverge markedly in calibration, drift magnitude, and risk preferences—supporting multi-axis evaluation for deployment decisions.

Data & Methods

  • Data
    • 500+ live, real-world prediction markets spanning political, economic, cultural, and technological events.
    • Events are forward-looking (no outcome contamination), categorized by domain, risk profile, temporal horizon, and market liquidity.
  • Evaluation design
    • Agents produce probabilistic forecasts over time tied to market timestamps; market-implied probabilities act as the human-imitation ground truth for aggregated human expectations.
    • Instruction locking and versioned prompt templates prevent experimental drift from prompt changes.
    • Baselines are constructed to be independent of rolling-window calibration so comparisons remain fair across models with different training histories.
  • Metrics & diagnostics
    • Proper scoring rules (Brier score, log-likelihood) measure forecast quality and calibration.
    • Drift diagnostics quantify temporal inconsistency and narrative shifts in model outputs relative to markets.
    • Risk-sensitivity and cost/resource-efficiency metrics capture economic decision quality and operational burden.
    • Token-constraint and prompt-robustness evaluations to test resource-limited deployments.
    • Statistical testing procedures and human-in-the-loop annotation protocols for interpretability and reproducibility.
  • System & algorithms
    • Layered architecture: instruction/prompt lock, baseline construction, agent deployment, market-linked execution with drift tracking, and an integrated evaluation loop.
    • Versioned evaluation contracts, reproducible reporting, and open release of agents and templates (A.1/A.2 and algorithms included).
  • Experimental setup
    • Multiple LLM agents (commercial and open families) were run against the same market streams and baselines.
    • Aggregation procedures and behavioral/temporal diagnostics used to compare models across the multi-dimensional metric suite.

Implications for AI Economics

  • Markets as social ground truth: Prediction markets provide a dynamic, aggregated human signal that can be used to assess whether LLMs behave like market participants — valuable for any economic forecasting application.
  • Calibration matters for economic decisions: Models with similar accuracy but different calibration/risk profiles will produce different economic actions (e.g., portfolio sizing, hedging); procurement and deployment should prioritize calibration and temporal stability, not only point accuracy.
  • Drift undermines economic reliability: Narrative drift and temporal inconsistency degrade the usefulness of forecasts for repeated-decision settings (trading, policy signaling). Continuous drift diagnostics are needed for operational monitoring and model governance.
  • Model selection & risk management: Multi-axis evaluation (accuracy, calibration, drift, cost) better aligns selection to economic objectives—e.g., risk-averse traders prefer well-calibrated, stable forecasts even at slight accuracy cost.
  • Cost and resource trade-offs: Token/compute constraints and cost-efficiency metrics are essential when evaluating models for real-time market-facing tasks where latency and execution cost matter.
  • Regulatory & auditability benefits: Instruction locking, versioned evaluation contracts, and reproducible scoring provide audit trails useful for regulators, compliance, and institutional deployment in finance and policy contexts.
  • Design of benchmarking for economic forecasting: TruthTensor demonstrates an operational blueprint for forward-looking, contamination-free benchmarks that can be adopted by investors, central banks, and research labs to evaluate agentic forecasting tools.
  • Cautions for economic use: Prediction markets themselves have limits (liquidity, representativeness, manipulation risk, domain coverage). Market-linked evaluation should be complemented by domain-specific validation and human oversight when stakes are high.

If you’d like, I can extract the specific drift-measurement formulas and the listed evaluation metrics (Brier, log-likelihood, drift statistics) into a one-page quick reference for decision-makers.

Assessment

Paper Typedescriptive Evidence Strengthmedium — The paper reports experiments across 500+ live prediction markets and uses probabilistic scoring and diagnostics to show substantive differences across models beyond point accuracy, which provides empirical support for its claims; however, it does not make or test causal claims about economic outcomes, relies on markets as proxies for real-world decision contexts, and may be subject to selection and measurement biases that limit how strongly results generalize. Methods Rigormedium — The framework emphasizes reproducibility (versioned evaluation contracts, contamination controls), prescribes annotation and testing protocols, and anchors evaluation to live, forward-looking tasks with probabilistic scoring—features indicative of careful methodological design; nevertheless, the paper appears focused on an evaluation paradigm rather than on rigorous identification of confounders, and details on model selection, market sampling strategy, statistical power for specific diagnostics, and robustness to market microstructure are not fully spelled out or experimentally varied. SampleEvaluation applied to 500+ forward-looking, 'contamination-free' real prediction markets spanning political, economic, cultural, and technological topics; models (various reasoning/LLM systems) produce probabilistic forecasts which are scored against market outcomes, with human-in-the-loop annotation and drift/robustness diagnostics; publicly released evaluation code and contracts via TruthTensor. Themeshuman_ai_collab adoption GeneralizabilityPrediction markets may not represent the full set of real-world decision contexts (e.g., enterprise workflows, regulatory decisions, non-forecasting tasks)., Results depend on the specific markets sampled (topics, geographies, liquidity) and may not hold in low-liquidity or proprietary markets., Models tested are not exhaustively described—findings may vary across model architectures, fine-tuning regimes, or access to external data., Forward-looking market outcomes take time to resolve, limiting rapid evaluation and possibly biasing toward events with clearer resolution criteria., Behavior in forecasting tasks (probabilistic prediction) may not translate directly to other economic outcomes such as productivity, wages, or firm performance.

Claims (9)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Static benchmarks fail to capture real-world uncertainty, distribution shift, and the gap between isolated task accuracy and human-aligned decision-making under evolving conditions. Decision Quality negative ability of benchmarks to capture real-world uncertainty and distribution shift
Reading fidelity high
Study strength medium
not reported
0.18
This paper introduces TruthTensor, a novel, reproducible evaluation paradigm that measures reasoning models not only as prediction engines but as human-imitation systems operating in socially-grounded, high-entropy environments. Decision Quality positive evaluation of reasoning models as human-imitation systems in socially-grounded environments
Reading fidelity high
Study strength low
not reported
0.09
The framework builds on forward-looking, contamination-free tasks and anchors evaluation to live prediction markets, combining probabilistic scoring to provide a holistic view of model behavior. Decision Quality positive holistic measurement of model behavior via probabilistic scores and live-market alignment
Reading fidelity high
Study strength low
not reported
0.09
TruthTensor complements traditional correctness metrics with drift-centric diagnostics and explicit robustness checks for reproducibility. Decision Quality positive robustness and drift diagnostics relative to traditional correctness metrics
Reading fidelity high
Study strength low
not reported
0.09
The framework specifies human vs. automated evaluation roles, annotation protocols, and statistical testing procedures to ensure interpretability and replicability of results. Decision Quality positive interpretability and replicability of evaluation results
Reading fidelity high
Study strength low
not reported
0.09
In experiments across 500+ real markets (political, economic, cultural, technological), TruthTensor demonstrates that models with similar forecast accuracy can diverge markedly in calibration, drift, and risk-sensitivity. Decision Quality mixed calibration, drift, risk-sensitivity (despite similar forecast accuracy)
Reading fidelity high
Study strength medium
n=500
0.18
TruthTensor underscores the need to evaluate models along multiple axes (accuracy, calibration, narrative stability, cost, and resource efficiency). Decision Quality positive necessity of multi-axis evaluation (accuracy, calibration, narrative stability, cost, resource efficiency)
Reading fidelity high
Study strength low
not reported
0.09
TruthTensor operationalizes modern evaluation best practices: clear hypothesis framing, careful metric selection, transparent compute/cost reporting, human-in-the-loop validation, and open, versioned evaluation contracts to produce defensible assessments of LLMs in real-world decision contexts. Decision Quality positive quality and defensibility of LLM assessments in real-world decision contexts
Reading fidelity high
Study strength low
not reported
0.09
TruthTensor was publicly released at https://truthtensor.com. Other positive public availability of the TruthTensor resource
Reading fidelity high
Study strength high
not reported
0.3

Notes