The Commonplace
Home Three-study pilot Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Expert regrading reveals that flawed questions and grader errors have substantially understated frontier LLMs' physics abilities: after auditing, GPT-5.6-Sol's scores rise by tens of percentage points on several leading benchmarks, suggesting near-saturation on these closed-ended tasks.

How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks
Ali Ansari, Haoran Sun, Andy Zeyi Liu, Mark Jabbour, Yongshan Ding, Steven Girvin, Yu He, Sohrab Ismail-Beigi, Aleksander Kubica, Owen D. Miller, Corey O'Hern, Vidvuds Ozolins, David Poland, A. Douglas Stone, Frank C. van den Bosch, Logan Wright, Navid Akbari, Santanu Antu, Kangle Cai, Andrew Calabrese-Day, Mateo Cárdenes Wuttig, Meng Cheng, Barry T. Chiang, Ali Ghorashi, Shouzhen Gu, Haoyang Huang, Zhibo Kang, Lukas Kienesberger, Hantian Liu, Charles Lomba, Zhongling Lu, Wenchao Ma, Rohin E. McIntosh, Evan McKinney, Ivan Rojkov, Xulei Sun, Yarone Meir Tokayer, Naveen Balaji Umasankar, Mira Varma, Leda Wang, Qimin Wang, Tyler Wang, Haoyu Wei, Jinming Yang, Jinchen Zhao, Sherlock Tingrui Zhao, Qinyuan Zheng, Jay S. Zou, Lucas Baker, Arman Cohan, John Sous · September 11, 2026
arxiv descriptive medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Ali Ansari unresolved corpus identity
  2. Haoran Sun unresolved corpus identity
  3. Andy Zeyi Liu unresolved corpus identity
  4. Mark Jabbour unresolved corpus identity
  5. Yongshan Ding unresolved corpus identity
  6. Steven Girvin unresolved corpus identity
  7. Yu He unresolved corpus identity
  8. Sohrab Ismail-Beigi unresolved corpus identity
  9. Aleksander Kubica unresolved corpus identity
  10. Owen D. Miller unresolved corpus identity
  11. Corey O'Hern unresolved corpus identity
  12. Vidvuds Ozolins unresolved corpus identity
  13. David Poland unresolved corpus identity
  14. A. Douglas Stone unresolved corpus identity
  15. Frank C. van den Bosch unresolved corpus identity
  16. Logan Wright unresolved corpus identity
  17. Navid Akbari unresolved corpus identity
  18. Santanu Antu unresolved corpus identity
  19. Kangle Cai unresolved corpus identity
  20. Andrew Calabrese-Day unresolved corpus identity
  21. Mateo Cárdenes Wuttig unresolved corpus identity
  22. Meng Cheng unresolved corpus identity
  23. Barry T. Chiang unresolved corpus identity
  24. Ali Ghorashi unresolved corpus identity
  25. Shouzhen Gu unresolved corpus identity
  26. Haoyang Huang unresolved corpus identity
  27. Zhibo Kang unresolved corpus identity
  28. Lukas Kienesberger unresolved corpus identity
  29. Hantian Liu unresolved corpus identity
  30. Charles Lomba unresolved corpus identity
  31. Zhongling Lu unresolved corpus identity
  32. Wenchao Ma unresolved corpus identity
  33. Rohin E. McIntosh unresolved corpus identity
  34. Evan McKinney unresolved corpus identity
  35. Ivan Rojkov unresolved corpus identity
  36. Xulei Sun unresolved corpus identity
  37. Yarone Meir Tokayer unresolved corpus identity
  38. Naveen Balaji Umasankar unresolved corpus identity
  39. Mira Varma unresolved corpus identity
  40. Leda Wang unresolved corpus identity
  41. Qimin Wang unresolved corpus identity
  42. Tyler Wang unresolved corpus identity
  43. Haoyu Wei unresolved corpus identity
  44. Jinming Yang unresolved corpus identity
  45. Jinchen Zhao unresolved corpus identity
  46. Sherlock Tingrui Zhao unresolved corpus identity
  47. Qinyuan Zheng unresolved corpus identity
  48. Jay S. Zou unresolved corpus identity
  49. Lucas Baker unresolved corpus identity
  50. Arman Cohan unresolved corpus identity
  51. John Sous unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Ali Ansari provider ID
  2. Hao-Ran Sun provider ID
  3. Andy Zeyi Liu provider ID
  4. Mark Jabbour provider ID
  5. Yong-Shan Ding provider ID
  6. S. Girvin provider ID
  7. Yu He provider ID
  8. S. Ismail-Beigi provider ID
  9. Aleksander Kubica provider ID
  10. Owen D. Miller provider ID
  11. C. O'Hern provider ID
  12. Vidvuds Ozolins provider ID
  13. David Poland provider ID
  14. A. Stone provider ID
  15. F. C. van den Bosch provider ID
  16. Logan Wright provider ID
  17. Navid Akbari provider ID
  18. Santanu Antu provider ID
  19. Kang-Le Cai provider ID
  20. Andrew Calabrese-Day provider ID
  21. M. Wuttig provider ID
  22. Meng-Dan Cheng provider ID
  23. Barry T. Chiang provider ID
  24. Ali Ghorashi provider ID
  25. Shou-Zhen Gu provider ID
  26. Hao-Yan Huang provider ID
  27. Zhi-Bo Kang provider ID
  28. Lukas Kienesberger provider ID
  29. Han-Tian Liu provider ID
  30. Charles J. Lomba provider ID
  31. Zhong-Lin Lu provider ID
  32. Wen-Chao Ma provider ID
  33. Rohin E. McIntosh provider ID
  34. Evan McKinney provider ID
  35. I. Rojkov provider ID
  36. Xu-Lei Sun provider ID
  37. Y. M. Tokayer provider ID
  38. Naveen Balaji Umasankar provider ID
  39. Mira Varma provider ID
  40. Le-Da Wang provider ID
  41. Qi-Ming Wang provider ID
  42. Tyler Wang provider ID
  43. Hao T. Wei provider ID
  44. Jin-Ming Yang provider ID
  45. Jin-Chen Zhao provider ID
  46. Sherlock Tingrui Zhao provider ID
  47. Qinan Zheng provider ID
  48. Jay S. Zou provider ID
  49. L. Baker provider ID
  50. Arman Cohan provider ID
  51. J. Sous provider ID
An expert audit shows many errors in physics benchmark materials and graders; after correcting and repairing questions and reference solutions, frontier LLMs (e.g., GPT-5.6-Sol) achieve substantially higher — often near-saturated — scores on closed-ended physics benchmarks.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Low reported scores on leading physics benchmarks, including those featured in the Artificial Analysis Intelligence Index (2026), suggest that frontier language models still struggle with advanced physics, a demanding test of their scientific reasoning and quantitative problem-solving abilities. Yet this impression does not always align with domain experts' experiences using these models in their work. We revisit these reported findings by evaluating frontier models on six widely used physics benchmarks and auditing them with experts, focusing on text-only problems with verifiable final answers. For each subfield of physics, faculty and graduate researchers with relevant expertise carefully review problem statements, reference solutions, and model responses to distinguish genuine model errors from grader errors, incorrect reference solutions, and ambiguous or underspecified questions. Most audited cases initially evaluated as incorrect reflect these benchmarking issues rather than errors in the models' physics reasoning. We then ask experts to address these benchmarking issues by correcting erroneous reference solutions and repairing or excluding flawed questions. We find that GPT-5.6-Sol's measured mean@4 rises from 47.3% to 78.7% on HLE-Physics and from 61.0% to 87.2% on CMT-Benchmark, while its corrected pass@4 reaches 94.4% on the 54 retained CritPt challenges. Corrected scores are computed on the retained evaluation subsets following expert review. Scores on the audited subsets of UGPhysics, PRISM-Physics, and PHYBench also rise substantially after correction. These findings suggest that current benchmarks substantially understate frontier models' ability to solve well-posed physics problems. Near-saturation on these closed-ended tasks highlights the need for more demanding, expert-validated evaluations.

Summary

Main Finding

Expert re-grading of six major physics benchmarks shows that most apparent failures of frontier language models are due to flawed benchmarks or graders, not model incapability. After correcting reference solutions, repairing or excluding ill-posed questions, and re-evaluating with expert oversight, measured performance of top models (notably GPT-5.6‑Sol) on closed-ended, text-only physics problems rises dramatically—often from tens of percent to the mid/upper 80s–90s. This implies many widely cited benchmark-based claims that frontier models “struggle” with physics substantially understate their ability on well-posed problems.

Key Points

  • Benchmarks audited: UGPhysics, PHYBench, PRISM‑Physics (drawn from public sources) and HLE‑Physics, CMT‑Benchmark, CritPt (expert‑authored). All are closed‑ended, text-only tasks with verifiable final answers.
  • Models evaluated: GPT‑5.6‑Sol, Claude Fable 5, Gemini 3.1 Pro. Primary metric: mean@4 (except existing pre-audit CritPt mean@5).
  • Error taxonomy used by auditors: model error (true model failure), grader error (correct model answer marked wrong by evaluator), benchmark error (ill-posed question, incorrect reference solution, missing assumptions).
  • Audit procedure: faculty and graduate researchers matched by subfield inspected problem statements, reference solutions, and model outputs. They repaired or excluded broken questions, derived missing solutions where needed, and re-scored responses with a standardized judge pipeline adapted from HLE.
  • Scale of benchmarking failures:
    • Across public-source audit subsets, 148 of 152 audited failures (≈97.4%) were grader or benchmark errors; only ~2.6% were genuine model errors.
    • Representative evaluator failure: rule-based Expression Edit Distance (EED) marked algebraically equivalent answers as incorrect.
  • Large score changes after correction (examples, mean@4 unless noted):
    • HLE‑Physics (GPT‑5.6‑Sol): pre-audit 47.3% → corrected 78.7%; pass@4 up to ~91.4%.
    • CMT‑Benchmark (GPT‑5.6‑Sol): 61.0% → 87.2%; pass@4 to ~98.0%.
    • CritPt: pre-audit mean@5 32.3% (70 items) → corrected mean@4 87.5% on 54 retained challenges; pass@4 94.4%.
    • Public-source sets (PHYBench, PRISM, UGPhysics) similarly rose from low pre-audit numbers to ≈85–95% on retained/ repaired subsets.
  • Conclusion: On well-posed, closed-ended physics problems, frontier LLMs are often near saturation. The reported low scores largely reflect evaluation defects rather than fundamental physics reasoning failures.

Data & Methods

  • Benchmarks: 6 widely used physics evaluation suites covering undergraduate through advanced/expert problems. Some benchmarks (UGPhysics, PHYBench, PRISM) reuse public exercises (risk of training-data contamination); others (HLE, CMT, CritPt) are expert-authored.
  • Sampling: For efficiency, audits focused on cases where GPT‑5.6‑Sol had all attempts marked incorrect for certain benchmarks; full auditing performed for the expert-authored sets as described.
  • Auditors: Teams of faculty and graduate researchers in relevant subfields reviewed per-question materials and model outputs; each audited case received a single label (model/grader/benchmark error).
  • Repairs: For benchmark errors, auditors either corrected reference solutions or clarified/added missing assumptions when defensible; irreparable items were excluded (retained sets vary by benchmark).
  • Evaluation pipeline: Pre-audit used each benchmark’s provided evaluator where available. Corrected evaluations used a unified pipeline (adapted from HLE) with an LLM judge prompt to reduce grader errors and standardize equivalence checks across benchmarks.
  • Metrics reported: mean@4 (average correctness across four attempts), pass@4 (fraction solved in at least one of four attempts). CritPt pre-audit used the originally reported mean@5 when applicable.
  • Key quantitative finding: Extremely high prevalence of non-model errors in audited failure cases (e.g., 97.4% on public-source audit subsets).

Implications for AI Economics

  • Capability assessments and forecasts: Economics analyses and forecasts that rely on published benchmark scores to infer model ability (for productivity, task automation, or displacement risk) may systematically understate the capabilities of frontier LLMs in quantitative, domain‑specific tasks. Updating forecasts to reflect expert-validated performance can materially change estimates of near-term impacts.
  • Labor productivity and substitution: If frontier models can reliably solve well-posed, closed-ended scientific and technical problems at high accuracy, downstream productivity gains in research, engineering, finance, and quantitative analytics could materialize faster than suggested by uncorrected benchmark scores. This increases the plausibility of earlier and larger labor reallocation effects in quantitative occupations.
  • Valuation, investment, and R&D strategy: Investors and firms using benchmark-based capability signals to allocate R&D or product investment may misprice opportunities. Expert-validated evaluations suggest stronger model utility for technical tasks, which should influence investment in AI-enabled tools, complementary human capital, and retraining programs.
  • Policy and regulation: Policymakers using benchmark performance as an input for safety, procurement, or regulatory timing should require evaluations that are expert‑audited and robust to grader/bias errors. Underestimating capability can delay needed governance, while overreliance on flawed benchmarks can produce misaligned regulation.
  • Benchmark design and resource allocation: The paper highlights the economic return to investing in higher-quality, expert‑validated benchmarks (and dynamic, open-ended evaluations). As models near saturation on closed-ended tasks, marginal value of more examples declines; resources may be better spent developing complex, process-level, or open-ended assessments (and tooling for human-in-the-loop verification) that better predict real-world impact.
  • Cautions for economists and modelers:
    • Training-data contamination: Public-source benchmarks risk leakage into model training; corrected high performance might partly reflect memorization. Economic impact analyses must distinguish genuine generalization from contamination-driven performance.
    • Evaluation saturation: Near-saturation on closed-ended tasks makes raw accuracy an unreliable discriminant of future progress—economic models should use richer performance measures (e.g., robustness, generalization to misspecified problems, time-to-solution, need for human edits).
    • Need for stress tests: For policy and market decisions, emphasize stress tests, adversarial or underspecified scenarios, and process-level evaluations that reveal weaknesses not captured by closed-ended benchmarks.

Recommended actions for AI economics stakeholders: - Re-examine modeling assumptions that use un-audited benchmark scores as inputs for capability or impact estimates. - Support and use expert‑validated, open-ended, and process-oriented evaluations when assessing models for high-stakes economic applications. - Incorporate uncertainty about benchmark quality into scenario analyses (sensitivity to corrected performance). - Monitor both corrected benchmark results and indicators of training-data contamination to better infer genuine generalization vs. memorization.

If helpful, I can extract a concise table of the pre-audit vs corrected scores for the three models and all six benchmarks to use directly in economic models or presentations.

Assessment

Paper Typedescriptive Evidence Strengthmedium — The paper provides systematic empirical auditing across six widely used physics benchmarks, using domain experts to relabel, repair, or exclude flawed questions and re-evaluate multiple frontier models, which yields large and internally consistent shifts in measured performance. However, limitations remain: selection rules (audits often focused on items the model originally failed), potential training-data contamination for public-source benchmarks, non-blinded expert reviews and possible institutional concentration of auditors, and corrected scores are computed on retained/repaired subsets rather than exact held-out originals — reducing the strength of generalization beyond the evaluated closed-ended tasks. Methods Rigormedium — The authors describe a reproducible protocol: pre-audit evaluation with original graders, a labeled taxonomy of error types (model, grader, benchmark), per-question expert assignment by subfield, repair/exclusion rules, and a unified post-audit evaluator pipeline. This is strong practice for an audit study. Weaknesses include audit sampling choices (e.g., auditing only responses marked incorrect for some benchmarks), limited transparency about inter-rater agreement and auditor independence, potential bias from adapting an evaluator pipeline that the authors prefer, and reliance on expert-derived reference solutions for some datasets. SampleSix closed-ended, text-only physics benchmarks were evaluated: three drawn from public sources (UGPhysics: 100 questions, 82 retained after validation; PHYBench: 100 questions, 87 retained; PRISM-Physics: 100 questions, 74 retained) and three expert-authored (HLE-Physics: 202 original, 116 retained; CMT-Benchmark: 50 original, 49 retained; CritPt: 70 original, 54 retained after audit). Three frontier models were tested (GPT-5.6-Sol, Claude Fable 5, Gemini 3.1 Pro) under multiple reasoning settings; pre-audit scores used original evaluators (or adapted HLE evaluator when none available) and corrected scores were computed after expert audits that repaired or excluded flawed questions and corrected reference solutions. Auditors were primarily Yale faculty and graduate researchers matched to subfields and they labeled errors as model, grader, or benchmark errors; repaired problems were retained when a defensible repair existed. Themesproductivity human_ai_collab GeneralizabilityResults apply to closed-ended, text-only physics problems with verifiable final answers and do not automatically generalize to open-ended, process-level, experimental, or multi-step research tasks., Public-source benchmarks may suffer from training-data contamination, so high corrected scores on those sets may partly reflect memorization or exposure., Audit focused on items originally marked incorrect (in some benchmarks), which can bias the error-attribution sample and change comparability to pre-audit aggregate scores., Expert repairs and derived reference solutions introduce subjectivity and possible institutional or reviewer bias; independent, blinded regrading is not shown., Performance depends on evaluation settings and tool access (some runs used tools), so scores may not reflect models running in more constrained environments.

Claims (8)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Expert correction increased GPT-5.6-Sol High's HLE-Physics mean@4 score from 47.28% to 78.66%. Output Quality positive Accuracy on closed-ended physics questions
Reading fidelity high
Study strength medium
n=116
31.38 percentage-point increase
0.18
Expert correction increased GPT-5.6-Sol High's HLE-Physics pass@4 from 55.94% to 91.38%. Task Completion Time positive Fraction of physics questions solved in at least one of four attempts
Reading fidelity high
Study strength medium
n=116
35.44 percentage-point increase
0.18
Expert correction increased GPT-5.6-Sol High's CMT-Benchmark mean@4 from 61.00% to 87.24%. Output Quality positive Accuracy on advanced condensed-matter physics questions
Reading fidelity high
Study strength medium
n=49
26.24 percentage-point increase
0.18
Expert correction increased GPT-5.6-Sol High's CMT-Benchmark pass@4 from 72.00% to 97.96%. Output Quality positive Fraction of condensed-matter physics questions solved in at least one of four attempts
Reading fidelity high
Study strength medium
n=49
25.96 percentage-point increase
0.18
On the retained CritPt challenges, GPT-5.6-Sol Max achieved a corrected mean@4 of 87.50% and a corrected pass@4 of 94.44%. Output Quality positive Accuracy and at-least-once solution rate on expert-curated physics challenges
Reading fidelity high
Study strength medium
n=54
87.50% mean@4; 94.44% pass@4
0.18
Among 152 audited cases from the three public-source benchmarks, 148 cases (97.37%) were attributed to benchmark or grader errors and 4 cases (2.63%) to model errors. Error Rate negative Rate and attribution of apparent evaluation errors
Reading fidelity high
Study strength medium
n=152
97.37% benchmark or grader errors; 2.63% model errors
0.18
For GPT-5.6-Sol High, corrected scores on the public-source benchmarks were substantially higher than pre-audit scores: 90.23% versus 26.50% on PHYBench, 94.59% versus 13.00% on PRISM-Physics, and 92.07% versus 83.00% on UGPhysics. Output Quality positive Mean@4 accuracy on public-source physics benchmarks
Reading fidelity high
Study strength medium
n=243
63.73, 81.59, and 9.07 percentage-point increases, respectively
0.18
The shared HLE-adapted evaluator used for corrected evaluations had a grader error rate of 4.08%, the lowest among the evaluators audited by the authors. Error Rate positive Evaluator/grader error rate
Reading fidelity high
Study strength low
4.08%
0.09

Notes