1 cumulative citations
View corpus contextA new evaluation framework, TruthTensor, shows that LLMs with similar forecast accuracy can diverge sharply on calibration, drift, and risk preferences when tested in 500+ live prediction markets, arguing for multi-dimensional, reproducible testing tied to real-world decision contexts.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Evaluating language models and AI agents remains fundamentally challenging because static benchmarks fail to capture real-world uncertainty, distribution shift, and the gap between isolated task accuracy and human-aligned decision-making under evolving conditions. This paper introduces TruthTensor, a novel, reproducible evaluation paradigm that measures reasoning models not only as prediction engines but as human-imitation systems operating in socially-grounded, high-entropy environments. Building on forward-looking, contamination-free tasks, our framework anchors evaluation to live prediction markets and combines probabilistic scoring to provide a holistic view of model behavior. TruthTensor complements traditional correctness metrics with drift-centric diagnostics and explicit robustness checks for reproducibility. It specify human vs. automated evaluation roles, annotation protocols, and statistical testing procedures to ensure interpretability and replicability of results. In experiments across 500+ real markets (political, economic, cultural, technological), TruthTensor demonstrates that models with similar forecast accuracy can diverge markedly in calibration, drift, and risk-sensitivity, underscoring the need to evaluate models along multiple axes (accuracy, calibration, narrative stability, cost, and resource efficiency). TruthTensor therefore operationalizes modern evaluation best practices, clear hypothesis framing, careful metric selection, transparent compute/cost reporting, human-in-the-loop validation, and open, versioned evaluation contracts, to produce defensible assessments of LLMs in real-world decision contexts. We publicly released TruthTensor at https://truthtensor.com.
Summary
Main Finding
TruthTensor reframes LLM evaluation from static prediction accuracy to a human-imitation, market-grounded testbed. By anchoring models to live prediction markets and measuring forward-looking probabilistic forecasts, calibration, narrative drift, and risk-sensitivity, TruthTensor shows that models with similar headline accuracy can behave very differently on economically relevant dimensions (calibration, temporal stability, and risk-aware decision quality). The framework is contamination-free, reproducible (instruction locking, versioned contracts), and released publicly at https://truthtensor.com.
Key Points
- Evaluation objective: measure how well LLMs imitate human probabilistic reasoning and narrative updates in socially grounded, high-entropy environments (prediction markets), not just whether they hit the correct outcome.
- Drift-centric: primary emphasis on measuring narrative drift, temporal inconsistency, and confidence decay over time—dimensions often ignored by static benchmarks.
- Contamination-free design: only forward-looking events are used so outcomes were unknown at training time; prompts are versioned/locked to prevent prompt-engineering leakage.
- Holistic metrics: combines accuracy with calibration (e.g., Brier/log scoring), narrative stability, risk-sensitivity, resource/cost efficiency, and behavioral/temporal diagnostics.
- Architecture & protocol: specifies prompt templates, baseline construction, agent deployment, market-linked execution, human-in-the-loop validation, and statistical testing to ensure interpretability and reproducibility.
- Empirical result: across 500+ real markets (political, economic, cultural, technological) models that tie on accuracy diverge markedly in calibration, drift magnitude, and risk preferences—supporting multi-axis evaluation for deployment decisions.
Data & Methods
- Data
- 500+ live, real-world prediction markets spanning political, economic, cultural, and technological events.
- Events are forward-looking (no outcome contamination), categorized by domain, risk profile, temporal horizon, and market liquidity.
- Evaluation design
- Agents produce probabilistic forecasts over time tied to market timestamps; market-implied probabilities act as the human-imitation ground truth for aggregated human expectations.
- Instruction locking and versioned prompt templates prevent experimental drift from prompt changes.
- Baselines are constructed to be independent of rolling-window calibration so comparisons remain fair across models with different training histories.
- Metrics & diagnostics
- Proper scoring rules (Brier score, log-likelihood) measure forecast quality and calibration.
- Drift diagnostics quantify temporal inconsistency and narrative shifts in model outputs relative to markets.
- Risk-sensitivity and cost/resource-efficiency metrics capture economic decision quality and operational burden.
- Token-constraint and prompt-robustness evaluations to test resource-limited deployments.
- Statistical testing procedures and human-in-the-loop annotation protocols for interpretability and reproducibility.
- System & algorithms
- Layered architecture: instruction/prompt lock, baseline construction, agent deployment, market-linked execution with drift tracking, and an integrated evaluation loop.
- Versioned evaluation contracts, reproducible reporting, and open release of agents and templates (A.1/A.2 and algorithms included).
- Experimental setup
- Multiple LLM agents (commercial and open families) were run against the same market streams and baselines.
- Aggregation procedures and behavioral/temporal diagnostics used to compare models across the multi-dimensional metric suite.
Implications for AI Economics
- Markets as social ground truth: Prediction markets provide a dynamic, aggregated human signal that can be used to assess whether LLMs behave like market participants — valuable for any economic forecasting application.
- Calibration matters for economic decisions: Models with similar accuracy but different calibration/risk profiles will produce different economic actions (e.g., portfolio sizing, hedging); procurement and deployment should prioritize calibration and temporal stability, not only point accuracy.
- Drift undermines economic reliability: Narrative drift and temporal inconsistency degrade the usefulness of forecasts for repeated-decision settings (trading, policy signaling). Continuous drift diagnostics are needed for operational monitoring and model governance.
- Model selection & risk management: Multi-axis evaluation (accuracy, calibration, drift, cost) better aligns selection to economic objectives—e.g., risk-averse traders prefer well-calibrated, stable forecasts even at slight accuracy cost.
- Cost and resource trade-offs: Token/compute constraints and cost-efficiency metrics are essential when evaluating models for real-time market-facing tasks where latency and execution cost matter.
- Regulatory & auditability benefits: Instruction locking, versioned evaluation contracts, and reproducible scoring provide audit trails useful for regulators, compliance, and institutional deployment in finance and policy contexts.
- Design of benchmarking for economic forecasting: TruthTensor demonstrates an operational blueprint for forward-looking, contamination-free benchmarks that can be adopted by investors, central banks, and research labs to evaluate agentic forecasting tools.
- Cautions for economic use: Prediction markets themselves have limits (liquidity, representativeness, manipulation risk, domain coverage). Market-linked evaluation should be complemented by domain-specific validation and human oversight when stakes are high.
If you’d like, I can extract the specific drift-measurement formulas and the listed evaluation metrics (Brier, log-likelihood, drift statistics) into a one-page quick reference for decision-makers.
Assessment
Claims (9)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Static benchmarks fail to capture real-world uncertainty, distribution shift, and the gap between isolated task accuracy and human-aligned decision-making under evolving conditions. Decision Quality | negative | ability of benchmarks to capture real-world uncertainty and distribution shift |
Reading fidelity
high
Study strength
medium
|
not reported
|
| This paper introduces TruthTensor, a novel, reproducible evaluation paradigm that measures reasoning models not only as prediction engines but as human-imitation systems operating in socially-grounded, high-entropy environments. Decision Quality | positive | evaluation of reasoning models as human-imitation systems in socially-grounded environments |
Reading fidelity
high
Study strength
low
|
not reported
|
| The framework builds on forward-looking, contamination-free tasks and anchors evaluation to live prediction markets, combining probabilistic scoring to provide a holistic view of model behavior. Decision Quality | positive | holistic measurement of model behavior via probabilistic scores and live-market alignment |
Reading fidelity
high
Study strength
low
|
not reported
|
| TruthTensor complements traditional correctness metrics with drift-centric diagnostics and explicit robustness checks for reproducibility. Decision Quality | positive | robustness and drift diagnostics relative to traditional correctness metrics |
Reading fidelity
high
Study strength
low
|
not reported
|
| The framework specifies human vs. automated evaluation roles, annotation protocols, and statistical testing procedures to ensure interpretability and replicability of results. Decision Quality | positive | interpretability and replicability of evaluation results |
Reading fidelity
high
Study strength
low
|
not reported
|
| In experiments across 500+ real markets (political, economic, cultural, technological), TruthTensor demonstrates that models with similar forecast accuracy can diverge markedly in calibration, drift, and risk-sensitivity. Decision Quality | mixed | calibration, drift, risk-sensitivity (despite similar forecast accuracy) |
Reading fidelity
high
Study strength
medium
|
n=500
|
| TruthTensor underscores the need to evaluate models along multiple axes (accuracy, calibration, narrative stability, cost, and resource efficiency). Decision Quality | positive | necessity of multi-axis evaluation (accuracy, calibration, narrative stability, cost, resource efficiency) |
Reading fidelity
high
Study strength
low
|
not reported
|
| TruthTensor operationalizes modern evaluation best practices: clear hypothesis framing, careful metric selection, transparent compute/cost reporting, human-in-the-loop validation, and open, versioned evaluation contracts to produce defensible assessments of LLMs in real-world decision contexts. Decision Quality | positive | quality and defensibility of LLM assessments in real-world decision contexts |
Reading fidelity
high
Study strength
low
|
not reported
|
| TruthTensor was publicly released at https://truthtensor.com. Other | positive | public availability of the TruthTensor resource |
Reading fidelity
high
Study strength
high
|
not reported
|