0 cumulative citations
View corpus contextExpert regrading reveals that flawed questions and grader errors have substantially understated frontier LLMs' physics abilities: after auditing, GPT-5.6-Sol's scores rise by tens of percentage points on several leading benchmarks, suggesting near-saturation on these closed-ended tasks.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Low reported scores on leading physics benchmarks, including those featured in the Artificial Analysis Intelligence Index (2026), suggest that frontier language models still struggle with advanced physics, a demanding test of their scientific reasoning and quantitative problem-solving abilities. Yet this impression does not always align with domain experts' experiences using these models in their work. We revisit these reported findings by evaluating frontier models on six widely used physics benchmarks and auditing them with experts, focusing on text-only problems with verifiable final answers. For each subfield of physics, faculty and graduate researchers with relevant expertise carefully review problem statements, reference solutions, and model responses to distinguish genuine model errors from grader errors, incorrect reference solutions, and ambiguous or underspecified questions. Most audited cases initially evaluated as incorrect reflect these benchmarking issues rather than errors in the models' physics reasoning. We then ask experts to address these benchmarking issues by correcting erroneous reference solutions and repairing or excluding flawed questions. We find that GPT-5.6-Sol's measured mean@4 rises from 47.3% to 78.7% on HLE-Physics and from 61.0% to 87.2% on CMT-Benchmark, while its corrected pass@4 reaches 94.4% on the 54 retained CritPt challenges. Corrected scores are computed on the retained evaluation subsets following expert review. Scores on the audited subsets of UGPhysics, PRISM-Physics, and PHYBench also rise substantially after correction. These findings suggest that current benchmarks substantially understate frontier models' ability to solve well-posed physics problems. Near-saturation on these closed-ended tasks highlights the need for more demanding, expert-validated evaluations.
Summary
Main Finding
Expert re-grading of six major physics benchmarks shows that most apparent failures of frontier language models are due to flawed benchmarks or graders, not model incapability. After correcting reference solutions, repairing or excluding ill-posed questions, and re-evaluating with expert oversight, measured performance of top models (notably GPT-5.6‑Sol) on closed-ended, text-only physics problems rises dramatically—often from tens of percent to the mid/upper 80s–90s. This implies many widely cited benchmark-based claims that frontier models “struggle” with physics substantially understate their ability on well-posed problems.
Key Points
- Benchmarks audited: UGPhysics, PHYBench, PRISM‑Physics (drawn from public sources) and HLE‑Physics, CMT‑Benchmark, CritPt (expert‑authored). All are closed‑ended, text-only tasks with verifiable final answers.
- Models evaluated: GPT‑5.6‑Sol, Claude Fable 5, Gemini 3.1 Pro. Primary metric: mean@4 (except existing pre-audit CritPt mean@5).
- Error taxonomy used by auditors: model error (true model failure), grader error (correct model answer marked wrong by evaluator), benchmark error (ill-posed question, incorrect reference solution, missing assumptions).
- Audit procedure: faculty and graduate researchers matched by subfield inspected problem statements, reference solutions, and model outputs. They repaired or excluded broken questions, derived missing solutions where needed, and re-scored responses with a standardized judge pipeline adapted from HLE.
- Scale of benchmarking failures:
- Across public-source audit subsets, 148 of 152 audited failures (≈97.4%) were grader or benchmark errors; only ~2.6% were genuine model errors.
- Representative evaluator failure: rule-based Expression Edit Distance (EED) marked algebraically equivalent answers as incorrect.
- Large score changes after correction (examples, mean@4 unless noted):
- HLE‑Physics (GPT‑5.6‑Sol): pre-audit 47.3% → corrected 78.7%; pass@4 up to ~91.4%.
- CMT‑Benchmark (GPT‑5.6‑Sol): 61.0% → 87.2%; pass@4 to ~98.0%.
- CritPt: pre-audit mean@5 32.3% (70 items) → corrected mean@4 87.5% on 54 retained challenges; pass@4 94.4%.
- Public-source sets (PHYBench, PRISM, UGPhysics) similarly rose from low pre-audit numbers to ≈85–95% on retained/ repaired subsets.
- Conclusion: On well-posed, closed-ended physics problems, frontier LLMs are often near saturation. The reported low scores largely reflect evaluation defects rather than fundamental physics reasoning failures.
Data & Methods
- Benchmarks: 6 widely used physics evaluation suites covering undergraduate through advanced/expert problems. Some benchmarks (UGPhysics, PHYBench, PRISM) reuse public exercises (risk of training-data contamination); others (HLE, CMT, CritPt) are expert-authored.
- Sampling: For efficiency, audits focused on cases where GPT‑5.6‑Sol had all attempts marked incorrect for certain benchmarks; full auditing performed for the expert-authored sets as described.
- Auditors: Teams of faculty and graduate researchers in relevant subfields reviewed per-question materials and model outputs; each audited case received a single label (model/grader/benchmark error).
- Repairs: For benchmark errors, auditors either corrected reference solutions or clarified/added missing assumptions when defensible; irreparable items were excluded (retained sets vary by benchmark).
- Evaluation pipeline: Pre-audit used each benchmark’s provided evaluator where available. Corrected evaluations used a unified pipeline (adapted from HLE) with an LLM judge prompt to reduce grader errors and standardize equivalence checks across benchmarks.
- Metrics reported: mean@4 (average correctness across four attempts), pass@4 (fraction solved in at least one of four attempts). CritPt pre-audit used the originally reported mean@5 when applicable.
- Key quantitative finding: Extremely high prevalence of non-model errors in audited failure cases (e.g., 97.4% on public-source audit subsets).
Implications for AI Economics
- Capability assessments and forecasts: Economics analyses and forecasts that rely on published benchmark scores to infer model ability (for productivity, task automation, or displacement risk) may systematically understate the capabilities of frontier LLMs in quantitative, domain‑specific tasks. Updating forecasts to reflect expert-validated performance can materially change estimates of near-term impacts.
- Labor productivity and substitution: If frontier models can reliably solve well-posed, closed-ended scientific and technical problems at high accuracy, downstream productivity gains in research, engineering, finance, and quantitative analytics could materialize faster than suggested by uncorrected benchmark scores. This increases the plausibility of earlier and larger labor reallocation effects in quantitative occupations.
- Valuation, investment, and R&D strategy: Investors and firms using benchmark-based capability signals to allocate R&D or product investment may misprice opportunities. Expert-validated evaluations suggest stronger model utility for technical tasks, which should influence investment in AI-enabled tools, complementary human capital, and retraining programs.
- Policy and regulation: Policymakers using benchmark performance as an input for safety, procurement, or regulatory timing should require evaluations that are expert‑audited and robust to grader/bias errors. Underestimating capability can delay needed governance, while overreliance on flawed benchmarks can produce misaligned regulation.
- Benchmark design and resource allocation: The paper highlights the economic return to investing in higher-quality, expert‑validated benchmarks (and dynamic, open-ended evaluations). As models near saturation on closed-ended tasks, marginal value of more examples declines; resources may be better spent developing complex, process-level, or open-ended assessments (and tooling for human-in-the-loop verification) that better predict real-world impact.
- Cautions for economists and modelers:
- Training-data contamination: Public-source benchmarks risk leakage into model training; corrected high performance might partly reflect memorization. Economic impact analyses must distinguish genuine generalization from contamination-driven performance.
- Evaluation saturation: Near-saturation on closed-ended tasks makes raw accuracy an unreliable discriminant of future progress—economic models should use richer performance measures (e.g., robustness, generalization to misspecified problems, time-to-solution, need for human edits).
- Need for stress tests: For policy and market decisions, emphasize stress tests, adversarial or underspecified scenarios, and process-level evaluations that reveal weaknesses not captured by closed-ended benchmarks.
Recommended actions for AI economics stakeholders: - Re-examine modeling assumptions that use un-audited benchmark scores as inputs for capability or impact estimates. - Support and use expert‑validated, open-ended, and process-oriented evaluations when assessing models for high-stakes economic applications. - Incorporate uncertainty about benchmark quality into scenario analyses (sensitivity to corrected performance). - Monitor both corrected benchmark results and indicators of training-data contamination to better infer genuine generalization vs. memorization.
If helpful, I can extract a concise table of the pre-audit vs corrected scores for the three models and all six benchmarks to use directly in economic models or presentations.
Assessment
Claims (8)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Expert correction increased GPT-5.6-Sol High's HLE-Physics mean@4 score from 47.28% to 78.66%. Output Quality | positive | Accuracy on closed-ended physics questions |
Reading fidelity
high
Study strength
medium
|
n=116
31.38 percentage-point increase
|
| Expert correction increased GPT-5.6-Sol High's HLE-Physics pass@4 from 55.94% to 91.38%. Task Completion Time | positive | Fraction of physics questions solved in at least one of four attempts |
Reading fidelity
high
Study strength
medium
|
n=116
35.44 percentage-point increase
|
| Expert correction increased GPT-5.6-Sol High's CMT-Benchmark mean@4 from 61.00% to 87.24%. Output Quality | positive | Accuracy on advanced condensed-matter physics questions |
Reading fidelity
high
Study strength
medium
|
n=49
26.24 percentage-point increase
|
| Expert correction increased GPT-5.6-Sol High's CMT-Benchmark pass@4 from 72.00% to 97.96%. Output Quality | positive | Fraction of condensed-matter physics questions solved in at least one of four attempts |
Reading fidelity
high
Study strength
medium
|
n=49
25.96 percentage-point increase
|
| On the retained CritPt challenges, GPT-5.6-Sol Max achieved a corrected mean@4 of 87.50% and a corrected pass@4 of 94.44%. Output Quality | positive | Accuracy and at-least-once solution rate on expert-curated physics challenges |
Reading fidelity
high
Study strength
medium
|
n=54
87.50% mean@4; 94.44% pass@4
|
| Among 152 audited cases from the three public-source benchmarks, 148 cases (97.37%) were attributed to benchmark or grader errors and 4 cases (2.63%) to model errors. Error Rate | negative | Rate and attribution of apparent evaluation errors |
Reading fidelity
high
Study strength
medium
|
n=152
97.37% benchmark or grader errors; 2.63% model errors
|
| For GPT-5.6-Sol High, corrected scores on the public-source benchmarks were substantially higher than pre-audit scores: 90.23% versus 26.50% on PHYBench, 94.59% versus 13.00% on PRISM-Physics, and 92.07% versus 83.00% on UGPhysics. Output Quality | positive | Mean@4 accuracy on public-source physics benchmarks |
Reading fidelity
high
Study strength
medium
|
n=243
63.73, 81.59, and 9.07 percentage-point increases, respectively
|
| The shared HLE-adapted evaluator used for corrected evaluations had a grader error rate of 4.08%, the lowest among the evaluators audited by the authors. Error Rate | positive | Evaluator/grader error rate |
Reading fidelity
high
Study strength
low
|
4.08%
|