The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

AI agents can build spreadsheet models nearly correctly but falter at valuation judgment: on GAUGE the top agent outperforms finance students yet falls short of senior analysts, passing most mechanical checks but missing a substantial share of judgment facets.

GAUGE: Grading Agent-Built Financial Models Without a Golden Answer
Jiacheng Lu, Sinuo Wang, Wentao Zhao, Rui Sun, Cheng Hua, Tao Song, Hui Cai, Beidi Luan, Zhengze Wu, Lingjing Teng, Yijia He, Jing Li, Daxin Jiang, Zuo Bai, Haibing Guan · July 27, 2026
arxiv descriptive medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Jiacheng Lu unresolved corpus identity
  2. Sinuo Wang unresolved corpus identity
  3. Wentao Zhao unresolved corpus identity
  4. Rui Sun unresolved corpus identity
  5. Cheng Hua unresolved corpus identity
  6. Tao Song unresolved corpus identity
  7. Hui Cai unresolved corpus identity
  8. Beidi Luan unresolved corpus identity
  9. Zhengze Wu unresolved corpus identity
  10. Lingjing Teng unresolved corpus identity
  11. Yijia He unresolved corpus identity
  12. Jing Li unresolved corpus identity
  13. Daxin Jiang unresolved corpus identity
  14. Zuo Bai unresolved corpus identity
  15. Haibing Guan unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Jiacheng Lu provider ID
  2. Sinuo Wang provider ID
  3. Wentao Zhao provider ID
  4. Rui Sun provider ID
  5. Cheng Hua provider ID
  6. Tao Song provider ID
  7. Hui Cai provider ID
  8. Beidi Luan provider ID
  9. Zhengze Wu provider ID
  10. Lingjing Teng provider ID
  11. Yijiao He provider ID
  12. Jing Li provider ID
  13. Daxin Jiang provider ID
  14. Zuo Bai provider ID
  15. Haibing Guan provider ID
GAUGE shows that modern LLM agents reliably reconstruct the mechanical structure of analyst Excel models but perform substantially worse on valuation judgment: the best agent scores above finance students but below senior analysts, with a pronounced mechanical–judgment performance gap.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Financial models combine public disclosures with analyst assumptions to produce forecasts and valuations. While some components can be checked mechanically, forecasts, discount rates, and target prices often admit multiple reasonable answers. Existing benchmarks nevertheless tend to grade such outputs against a single expert reference. Using independently built analyst models for the same companies, we find that across 108 directed pairs covering 65 companies, the median single-reference score is 0.33, 92.6% score below 0.70, and no same-vintage pair agrees on implied price within 10%. Point-tolerance grading can therefore penalize disagreement already present among professionals. We introduce GAUGE, a benchmark for evaluating agent-built valuation models against observed analyst practice rather than a single point answer. GAUGE uses 1,001 vendor-classified analyst workbooks and a 196-task evaluation set, with a three-layer observed-practice envelope, 56 auditable facets, eight validity gates, and deterministic structural checks. We validate the benchmark with a 55-participant known-groups study, company-grouped cross-fitting, and judge-stability audits. On the failure-aware score $φ_0$, senior analysts average 88.3, juniors 66.0, and finance students 43.2. Across 24 agents and 1,011 scored generations, the best agent scores 53.4, above the student mean but below every senior and most juniors. It passes 93% of mechanical facets and 78% of judgment facets, with a fleet-median gap of 26 points. Current agents are substantially stronger at model construction than valuation judgment. We release the methodology, a gated de-identified data tier, a controlled training split, a versioned 48-task evaluation core, and a withheld refresh pool.

Summary

Main Finding

GAUGE introduces a defensibility-aware benchmark for end-to-end agent-built financial valuation models that grades agent outputs against observed analyst practice rather than a single expert “golden” answer. Using 1,001 analyst workbooks and a 196-task evaluation bank, GAUGE shows (1) point-tolerance scoring against one reference penalizes disagreement that already exists among analysts, and (2) current LLM agents are substantially stronger at mechanically constructing models than at making valuation judgments.

Key Points

  • Peer-audit of analyst workbooks:
    • Corpus: 1,001 vendor-classified analyst workbooks (922 tickers, 25 industries).
    • Multi-covered subset: 137 workbooks covering 65 companies used to test the single-reference assumption.
    • Under conventional point tolerances (revenue ±5%, WACC ±50bp, price ±10%): median single-reference score across 108 directed same-company pairs = 0.33; 92.6% of pairs score < 0.70. No same-vintage pair agreed on implied share price within ±10%.
    • Observed dispersion: median absolute ΔWACC = 147 bp (p90 = 374 bp); median absolute Δ implied price = 25% (p90 = 107%).
  • GAUGE scoring design:
    • Three-layer defensibility envelope (from most to least specific): E-method (same workbook multi-methods/sensitivity), E-industry (industry p10–p90 distributions), E-company (same-company analyst dispersion used to widen bands).
    • 56 auditable facets across mechanical vs. judgment dimensions: 29 deterministic, 23 human-judged, 4 direct-rule facets.
    • Eight hard validity gates (structural checks) that cap scores if the model is unusable (e.g., unbalanced balance sheet, circular projections, missing valuation).
    • Facet scoring: extracted value ∈ envelope → 2; in widened band only → 1; outside → 0. Mapped to benchmark scale φ: 0→0, 1→60 (Pass), 2→100 (Excellent). Final score is mean over active facets, then gated by the lowest triggered ceiling.
    • Failure-aware score φ0: non-completions (capability failures) are counted as zero.
  • Evaluation results:
    • Human baseline (55 participants): seniors avg φ0 = 88.3; juniors = 66.0; finance students = 43.2.
    • Agents: 24 agents, 1,011 scored generations on a frozen 48-task core (drawn from the 196 tasks).
    • Best agent (Claude Fable 5 in paper) φ0 = 53.4 — above the student mean but below all senior analysts and most juniors.
    • Mechanical vs. judgment gap: top agent passes 93% of mechanical facets vs. 78% of judgment facets; fleet-median gap = 26 points. Mechanical checks are more reliably handled by agents than judgmental assumptions/valuation.
  • Validity and measurement controls:
    • Company-grouped cross-fitting prevents reuse of a company in the envelope calibration for its own evaluation fold.
    • Known-groups study, judge-vote sampling (k=5) with high stability (Kendall τ ≈ 0.944), and an external expert-consensus comparison (reported agreement ≈ 86.7% vs. experts).
    • Deterministic detectors/code grade ~54% of facet outcomes and all gate decisions (reduces run variance).
  • Resources released: methodology, gated de-identified data tier, controlled training split, versioned 48-task evaluation core, withheld refresh pool for longitudinal evaluation.

Data & Methods

  • Corpus and task construction:
    • 1,001 analyst-built workbooks (vendor-classified, sizes from ~2K to ~35K cells).
    • 196 verification-validated evaluation tasks: Excel-in / Model-out design; each task supplies 3 FY historicals, IS/BS/CF skeletons, as-of-date guidance, and requires a formula-driven model, valuation, sensitivity analysis, assumptions file, and memo.
    • Extractor: provenance-tracked extractor reads models, checks accounting identities, and extracts assumption values and outputs for grading.
  • Envelope construction:
    • E-method: uses multi-method values and sensitivity grids inside the same workbook where available.
    • E-industry: industry-level empirical distributions (GICS p10–p90) for extractable assumptions.
    • E-company: same-company analyst dispersion used to define a widened/near band (e.g., widen by p90 cross-analyst disagreement).
    • Scoring rule: value in E → score 2; in widened-only region → score 1; outside → 0. Unmeasurable facets recorded N/A (not penalized).
  • Facet taxonomy and aggregation:
    • 56 facets across 5 pillars and 21 sub-capabilities; facets activated conditionally by industry and instance-level availability.
    • φ mapping: raw 0/1/2 → {0, 60, 100}; pre-gate score = mean over active facets; final score = min(pre-gate, gate ceilings).
  • Validity checks:
    • Peer-workbook audit: tests single-reference assumption by scoring analyst A vs analyst B for same company under point tolerances.
    • Known-groups human study: senior/junior/student ordering as expected.
    • Cross-fit calibration: removes evaluated company from calibration pool.
    • Judge stability: five draws from frozen judge, majority reduction; low flip rate.
  • Evaluation harness:
    • Fixed tool-calling harness, same prompts, tools and turn budgets for all agents.
    • Single-run-per-task to avoid best-case bias; capability failures counted as zero in φ0.
    • 48-task frozen core for leaderboard and longitudinal comparability; ∼600 workbooks withheld for refresh waves.

Implications for AI Economics

  • Benchmarks for valuation must account for non-identifiability of judgmental quantities:
    • Point-reference grading (single golden answer) systematically penalizes defensible variation already present among professional analysts. Benchmark design in AI-for-finance should use empirically derived practice envelopes and validity gates to avoid conflating reasonable divergence with agent error.
  • Agents are approaching human competence on mechanical modeling but lag on valuation judgment:
    • High mechanical pass rates show LLMs can reliably assemble spreadsheet models, link statements, and implement calculations — tasks where correctness is deterministic and verifiable.
    • Lower performance on judgment facets (assumptions, discount rates, terminal values, price implications) implies agents are not yet trustworthy as autonomous valuation decision-makers. Human oversight remains necessary for judgmental, strategic, or regulatory-significant valuation choices.
  • Deployment and market impact:
    • Near-term productive uses: agents as model-building assistants (data extraction, formula wiring, sensitivity scaffolding) to improve analyst productivity and reduce mechanical errors.
    • Limits on automation: clients, compliance teams, and regulators should treat agent-generated target prices and investment recommendations with caution until judgment capabilities and calibration to professional norms improve.
    • Governance: use of observed-practice envelopes and hard structural gates can be operationalized as part of model-risk controls to detect unusable outputs and limit overreliance on single-point valuations.
  • Research directions valuable to AI economics:
    • Better uncertainty representation and multi-reference scoring: improving how agents express defensible ranges and justify assumptions that map into observed-practice envelopes.
    • Learning from heterogeneous expert distributions: methods that learn and condition on analyst-style clusters or objectives (sell-side vs buy-side, conservative vs aggressive) could yield customizable agent behavior aligned with institutional preferences.
    • Improving economic judgment: incorporate causal reasoning, scenario analysis, and macro-financial linkages to reduce the valuation gap.
    • Continual evaluation and contamination control: GAUGE’s withheld refresh pool and versioning illustrate the need for living benchmarks and contamination-aware training splits in financial domains.
  • Limitations and cautions:
    • GAUGE’s observed-practice envelope is defensibility-aware, not a claim of “ground truth.” Being in-band means consistent with practice, not necessarily correct.
    • Corpus vendor labels do not verify author credentials or analytical quality; envelope calibration is an empirical regularization, not a normative standard.
    • External validity: calibration and validation are internal to the corpus; generalization beyond the sampled coverage remains an empirical question.

Summary takeaway: GAUGE provides a principled, reproducible alternative to single-golden grading for financial-model valuation tasks, revealing that LLM agents can reliably build models but still fall substantially short of professional analysts on valuation judgment. For AI economics research and deployment, the paper argues for defensibility-aware evaluation, continued human oversight on judgmental outputs, and research that closes the judgment gap while preserving structural checks and contamination controls.

Assessment

Paper Typedescriptive Evidence Strengthmedium — The paper presents a large, well-documented benchmark (1,001 analyst workbooks, 196-task evaluation bank, a 48-task frozen core) and multi-modal validation (peer-workbook audit, a 55-participant known-groups human baseline, cross-fold calibration and judge-stability checks). These provide substantial empirical support for claims about current agent capabilities on the benchmark tasks. However, claims about broader agent competence are limited by calibration on a 65-company multi-covered subset, potential provenance gaps in human expert audits, a fixed 48-task core (rather than full-bank sampling), and possible training-data contamination concerns, so evidence is not yet conclusive at the highest level. Methods Rigormedium — The measurement design is careful and comprehensive: deterministic structural checks, a three-layer observed-practice envelope derived from multi-analyst workbooks, 56 granular facets, 8 hard validity gates, known-groups human validation, cross-fold calibration (company-grouped), and judge sampling stability analyses. Methodological caveats include reliance on vendor-classified workbooks without verified author credentials, a calibration pool limited to 65 multi-covered tickers for envelope construction, incomplete provenance for the human expert labeling used to validate judged facets, and a fixed small frozen core (48 tasks) for leaderboard comparisons which may limit representativeness. SampleCorpus of 1,001 vendor-classified analyst-built Excel valuation workbooks covering 922 tickers and 25 GICS groups (vendor tiers: ~404/347/200/50). A multi-covered subset contains 137 workbooks covering 65 companies (2–3 workbooks per ticker) used for peer-audit and envelope calibration. The evaluation bank contains 196 extraction-verified tasks; a frozen 48-task core (stratified by tier and GICS) is used for flagship leaderboard reporting. Experiments: 24 agents evaluated (1,011 scored runs across core tasks), with a 55-participant human baseline (senior analysts, junior analysts, finance students) and additional withheld refresh set (~600 workbooks). Facet-level scoring: 56 facets (29 deterministic, 23 judged, 4 direct-rule) across mechanical/assumptions/valuation pillars; 8 deterministic validity gates. Themeshuman_ai_collab productivity labor_markets adoption GeneralizabilityCorpus is vendor-classified and authorship/analyst credentials are not independently verified, which may limit external validity relative to true sell-side or buy-side research populations., Envelope calibration relies on 65 multi-covered companies; industry/company coverage may not generalize to all sectors, geographies, or uncommon corporate structures., Evaluation focuses on equity valuation Excel workbooks (three-statement DCF/triangulation archetypes) and may not generalize to other financial tasks (e.g., fixed income, structured products, portfolio construction) or non-Excel workflows., Agents were run under a single tool-calling harness and turn budget; results may depend on harness/tooling choices and would differ under alternative toolchains or more/less permissive budgets., The 48-task frozen core used for leaderboards is small relative to the full bank and may not capture the full heterogeneity of real-world modeling challenges., Potential for benchmark contamination/training overlap with some agents cannot be fully ruled out, despite withhold/refreshed pools.

Claims (12)

ClaimDirectionOutcomeConfidence & EvidenceDetails
When one analyst's workbook is used as the single golden reference, the median score across 108 directed same-company analyst pairs is 0.33. Decision Quality negative Agreement between independently built analyst valuation models under single-reference grading
Reading fidelity high
Study strength medium
n=108
median score of 0.33
0.18
Under standard single-reference tolerances, 92.6% of directed same-company analyst pairs score below 0.70. Decision Quality negative Proportion of analyst-model comparisons meeting a single-reference score threshold
Reading fidelity high
Study strength medium
n=108
92.6% below 0.70
0.18
No same-vintage analyst pair agrees on implied share price within 10% under the benchmark's base price tolerance. Decision Quality null_result Agreement in implied share price across same-vintage analyst models
Reading fidelity high
Study strength low
n=14
0% of same-vintage pairs passed within 10%
0.09
Among observed same-company analyst models, the median absolute difference in WACC is 147 basis points and the median absolute difference in implied price is 25%. Decision Quality negative Cross-analyst dispersion in WACC and implied share price
Reading fidelity high
Study strength medium
n=78
147 bp median WACC difference; 25% median implied-price difference
0.18
GAUGE is constructed from 1,001 vendor-classified analyst-built valuation workbooks covering 922 tickers and 25 GICS industry groups. Training Effectiveness positive Benchmark corpus coverage
Reading fidelity high
Study strength medium
n=1001
0.18
GAUGE's evaluation set contains 196 extraction-verified modeling tasks, with 98% of candidate workbooks certified under the benchmark's schemas. Training Effectiveness positive Benchmark task verification and coverage
Reading fidelity high
Study strength medium
n=200
196/200 certified (98%)
0.18
On GAUGE's failure-aware score, senior analysts average 88.3, junior analysts average 66.0, and finance students average 43.2. Decision Quality positive End-to-end financial-modeling benchmark performance
Reading fidelity high
Study strength medium
n=55
senior 88.3; junior 66.0; students 43.2
0.18
Across 24 evaluated agents and 1,011 scored generations, the best agent scores 53.4 on GAUGE's failure-aware score, exceeding the student mean but remaining below every senior analyst and most junior analysts. Decision Quality mixed Agent performance on end-to-end financial-model construction and valuation
Reading fidelity high
Study strength medium
n=1011
best-agent score 53.4
0.18
Agents perform better on mechanical model-construction facets than on judgment facets: the best agent passes 93% of mechanical facets versus 78% of judgment facets. Decision Quality mixed Pass rates on mechanical construction versus financial judgment facets
Reading fidelity high
Study strength medium
n=1011
93% mechanical versus 78% judgment; 15 percentage-point gap
0.18
The median mechanical-versus-judgment performance gap across agents is 26 points. Decision Quality negative Difference between agent performance on mechanical and judgment facets
Reading fidelity high
Study strength medium
n=24
26-point median gap
0.18
Company-grouped cross-fitting found that the method-based envelope strictly covered 53.8% of eligible peer prices and covered 91.2% under the p90 near rule; strict held-out industry-based WACC coverage was 82.4%. Decision Quality positive Coverage of analyst-observed valuation ranges by GAUGE's grading envelope
Reading fidelity high
Study strength low
53.8% strict price coverage; 91.2% p90-near price coverage; 82.4% strict WACC coverage
0.09
Using five draws from one frozen judge with majority reduction, GAUGE achieves Kendall's tau of 0.944 and a facet flip rate of 2.2%. Ai Safety And Ethics positive Stability and repeatability of qualitative facet judgments
Reading fidelity high
Study strength medium
n=5
Kendall's τ=0.944; facet flip rate=2.2%
0.18

Notes