The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

A new benchmark for autonomous scientific research finds current AI agents can implement prescribed procedures but struggle to design and execute research independently: average scores fall from about 51% with full guidance to 27% when methods must be chosen by the AI, highlighting a substantial gap to truly autonomous scientific discovery.

ASI-Bench: At the Dawn of Artificial Superintelligence
Junwei Zhou, Zhen Sun, Binyu Li, Jiangyu Zhou, Yuexi Pan, Hengyu Wang, Honghe Ren, Xiaohan Jia, Xueyang Zhou, Xiaoyu Cao, Yongchao Chen, Yuanning Feng, Junhao Wu, Cheng Zhang, Sijia Chen, Haoyu Xue, Chengsong You, Huan Wang, Koutian Wu, Peigan Gao, Jiakun Wu, Wenzhe Li, Ergan Shang, Qingyuan Zheng, Jingjing Zhou, Ruixuan Jia, Yan Xu, Hongrui Zhang, Xiao-Han Ma, Zhengxiang Cheng, Yuexing Hao, Liting Mai, Xianglin Ji, Wenjun Zhang, Zhuofan Chen, Yixiao Huang, Chi Wang, Wenyue Hua, Yilun Hao, Yuantao Zhai, Ziyan Zhao, Jingyan Xie · August 18, 2026
arxiv descriptive medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Junwei Zhou unresolved corpus identity
  2. Zhen Sun unresolved corpus identity
  3. Binyu Li unresolved corpus identity
  4. Jiangyu Zhou unresolved corpus identity
  5. Yuexi Pan unresolved corpus identity
  6. Hengyu Wang unresolved corpus identity
  7. Honghe Ren unresolved corpus identity
  8. Xiaohan Jia unresolved corpus identity
  9. Xueyang Zhou unresolved corpus identity
  10. Xiaoyu Cao unresolved corpus identity
  11. Yongchao Chen unresolved corpus identity
  12. Yuanning Feng unresolved corpus identity
  13. Junhao Wu unresolved corpus identity
  14. Cheng Zhang unresolved corpus identity
  15. Sijia Chen unresolved corpus identity
  16. Haoyu Xue unresolved corpus identity
  17. Chengsong You unresolved corpus identity
  18. Huan Wang unresolved corpus identity
  19. Koutian Wu unresolved corpus identity
  20. Peigan Gao unresolved corpus identity
  21. Jiakun Wu unresolved corpus identity
  22. Wenzhe Li unresolved corpus identity
  23. Ergan Shang unresolved corpus identity
  24. Qingyuan Zheng unresolved corpus identity
  25. Jingjing Zhou unresolved corpus identity
  26. Ruixuan Jia unresolved corpus identity
  27. Yan Xu unresolved corpus identity
  28. Hongrui Zhang unresolved corpus identity
  29. Xiao-Han Ma unresolved corpus identity
  30. Zhengxiang Cheng unresolved corpus identity
  31. Yuexing Hao unresolved corpus identity
  32. Liting Mai unresolved corpus identity
  33. Xianglin Ji unresolved corpus identity
  34. Wenjun Zhang unresolved corpus identity
  35. Zhuofan Chen unresolved corpus identity
  36. Yixiao Huang unresolved corpus identity
  37. Chi Wang unresolved corpus identity
  38. Wenyue Hua unresolved corpus identity
  39. Yilun Hao unresolved corpus identity
  40. Yuantao Zhai unresolved corpus identity
  41. Ziyan Zhao unresolved corpus identity
  42. Jingyan Xie unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Junwei Zhou provider ID
  2. Zhen Sun provider ID
  3. Binyu Li provider ID
  4. Jiangyue Zhou provider ID
  5. Yue Pan provider ID
  6. He Wang provider ID
  7. Hong Ren provider ID
  8. Xiaohan Jia provider ID
  9. Xueyang Zhou provider ID
  10. Xiaoyu Cao provider ID
  11. Yongchao Chen provider ID
  12. Yuanning Feng provider ID
  13. Jun Wu provider ID
  14. Cheng Zhang provider ID
  15. Si Chen provider ID
  16. Haoyu Xue provider ID
  17. Chengsong You provider ID
  18. Huanhuan Wang provider ID
  19. Ko-Hui Wu provider ID
  20. Peigan Gao provider ID
  21. Jiakun Wu provider ID
  22. Wenzhe Li provider ID
  23. Ergan Shang provider ID
  24. Qing Zheng provider ID
  25. Jingjing Zhou provider ID
  26. Ruixuan Jia provider ID
  27. Yan Xu provider ID
  28. Hongrui Zhang provider ID
  29. Xiaohui Ma provider ID
  30. Zhengxiang Cheng provider ID
  31. Yuexing Hao provider ID
  32. Liting Mai provider ID
  33. Xianglin Ji provider ID
  34. Wenjun Zhang provider ID
  35. Zhuo Chen provider ID
  36. Yixiao Huang provider ID
  37. Chi Wang provider ID
  38. Wenyue Hua provider ID
  39. Yilun Hao provider ID
  40. Yu Zhai provider ID
  41. Ziyan Zhao provider ID
  42. Jing-Dun Xie provider ID
ASI-Bench is a rigorously constructed 60-task, cross-domain benchmark that measures AI agents' ability to autonomously conduct project-level scientific research and finds a large drop in performance as human methodological guidance is withdrawn.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Artificial superintelligence (ASI) requires AI to move beyond mastering existing knowledge toward exploring the unknown, creating new knowledge, and turning new ideas into verifiable results. However, the capabilities of today's AI systems are still largely built on learning, compressing, and applying existing human knowledge. Accordingly, existing benchmarks primarily test whether AI can produce correct answers based on learned knowledge, or whether it can complete tasks under extensive human guidance. We therefore introduce ASI-Bench, the first benchmark to jointly evaluate AI systems' capabilities of innovative exploration and autonomous scientific execution across general research domains, and the first to progressively withdraw human methodological guidance within the same research project to test how far AI can proceed on its own. Built by over 40 experts with the cost of 31,000+ human hours, ASI-Bench contains 60 project-level research tasks across 11 scientific domains and progressively reduces methodological guidance to test whether AI can independently select methods, conduct research, and produce verifiable results. All tasks undergo expert review, AI-assisted auditing, sandbox execution, and scorer validation. Across 18 state-of-the-art agent--model configurations, the average score drops from 50.91 with full methodological guidance to 29.10 with only the method specified and 26.62 when agents must determine the method themselves. This sharp decline shows that current systems remain heavily dependent on human guidance and are still far from autonomously conducting end-to-end, project-level scientific research. ASI-Bench is open to the world. We invite researchers and builders everywhere to contribute new tasks, challenge the limits of today's AI, and help accelerate humanity's collective path toward artificial superintelligence at https://asibench.apexin.ai/submit.

Summary

Main Finding

ASI-Bench is a new, large-scale benchmark for measuring AI systems’ ability to perform open-ended, project‑level scientific research as human methodological guidance is progressively withdrawn. Evaluated across 60 cross‑domain research tasks and 18 Agent×Model configurations, performance drops sharply when detailed procedural guidance is removed (mean score: B1 = 50.91 → B2 = 29.10 → B3 = 26.62), revealing that current systems remain heavily dependent on human-prescribed procedures and are far from reliably autonomous scientific discovery.

Key Points

  • Benchmark structure (guidance gradient):
    • B1: Full methodological guidance (detailed step-by-step procedure provided).
    • B2: Only the method class/name provided (no procedural steps).
    • B3: No methodological guidance—agent must choose method and workflow.
    • B4: B3 plus plausible but irrelevant distractors.
  • Main quantitative results (macro-averaged over tasks; 3 runs unless noted):
    • Mean scores: B1 = 50.91, B2 = 29.10, B3 = 26.62, B4 = 26.99.
    • Largest drop is B1→B2: −21.82 points (method operationalization loss).
    • Smaller drop B2→B3: −2.48 points (method selection less limiting than operationalization).
    • Only one configuration (GPT-5.6 Sol (ultra) / Codex harness) exceeded 50 in B3 (51.60).
  • Harness (agent system) matters: same backbone can perform substantially differently under different harnesses; capability emerges from model × harness interaction, not model scale alone.
  • Computational cost increases when guidance is reduced:
    • Avg. tokens per task: B1 = 4.4M, B2 = 6.9M (+59%), B3 = 5.4M (+25%), B4 = 5.7M (+30%).
    • Avg. time per task: B1 = 37.8 min, B2 = 49.7 min (+32%), B3 = 45.9 min (+22%).
  • Current systems are robust to irrelevant distractors (B3 vs B4 nearly unchanged), but fragile in turning a method into a working, end-to-end research procedure.

Data & Methods

  • Benchmark composition:
    • 60 project-level research tasks spanning 11 domains: mathematics, physics, chemistry, biology, astronomy, materials science, earth science, medicine & biostatistics, computer science, robotics, electrical engineering.
    • Each task: scientific objective, task-specific data, executable sandbox environment, verifiable artifacts, and fixed scoring criteria.
    • Tasks are long-horizon and multi-stage (over 2,600 interaction turns, ~2,400 execution steps, >35 hours total agent execution across tasks).
  • Construction & validation pipeline:
    • Started from >1,300 candidate ideas; iterative process produced final 60 tasks.
    • Extensive human effort: >31,000 human-hours, >1,100 review assignments, >2,000 task revisions, 5 review rounds.
    • 1,500 sandbox runs to verify reproducibility, scoring, and to eliminate shortcuts or leakage.

  • Evaluation:
    • 18 representative Agent×Model configurations tested (various backbone LMs and harnesses).
    • Scores macro-averaged across tasks; most results are means over three independent runs (some single-run results noted).
    • Diagnostics computed per run then aggregated (e.g., B2−B1, B3−B1, B4−B3).
  • Transparency & extensibility:
    • ASI-Bench is public and invites external task contributions and leaderboard submissions.

Implications for AI Economics

  • Where to allocate R&D investment:
    • High returns likely from investing in agent harness design, tool integration, and method operationalization infrastructure (turning high‑level methods into robust end‑to‑end procedures) rather than solely scaling backbone models.
    • Engineering effort that reduces the human-to-agent procedural gap could produce larger practical gains for autonomous research than equivalent expenditures on model scale alone.
  • Labor and task substitution:
    • Current systems’ heavy dependence on procedural guidance implies AI will more immediately augment human researchers (improving productivity, automating laborious implementation/testing) rather than fully replace them in open-ended research.
    • Demand will remain for human domain experts who craft procedures, validate results, and close evaluation gaps—supporting complementary human capital and high-skill employment in research ecosystems.
  • Market structure and concentration:
    • High compute, data, and engineering costs (large token/time per task, sandbox runs) favor well‑resourced organizations. Institutions that can invest jointly in models, harnesses, and validated sandboxes may obtain durable competitive advantages, increasing concentration in AI-driven scientific discovery.
    • Leaderboards and public benchmarks (like ASI-Bench) can mitigate information asymmetries but only if they remain well-maintained and resistant to gaming.
  • Pricing and ROI for AI-for-Science products:
    • Buyers should account for substantial integration and engineering costs to realize autonomous behavior; product pricing and adoption models should reflect combined model + systems development rather than treating model access in isolation.
    • For investors, cost–performance tradeoffs shown in ASI-Bench imply careful evaluation of marginal gains from compute spend: higher cost systems can improve B3 performance, but the operationalization bottleneck may limit returns from pure compute increases.
  • Policy and regulation economics:
    • Reproducibility, audit trails, and sandboxed verification (as implemented by ASI-Bench) are crucial public goods; funding and standards for independent validation infrastructure could have large social value.
    • Because the benchmark reveals limited true autonomy today, regulatory focus can prioritize standards for claimed autonomous scientific systems (verification, transparency, human oversight) to avoid premature deployment claims and misaligned incentives.
  • Emerging markets:
    • Demand for high-quality, domain-specific sandboxes, verified datasets, automated validation tooling, and benchmark‑engineering expertise will grow—creating opportunities for specialized firms and service providers that bridge models and domain workflows.

Summary: ASI-Bench quantifies a key gap in present AI capability: converting a methodological idea into a robust, end‑to‑end research workflow. For AI economics, this implies most near‑term value comes from systems engineering, validation infrastructure, and complementary human expertise rather than model scaling alone—shaping R&D priorities, labor markets, investment strategies, and regulatory needs.

Assessment

Paper Typedescriptive Evidence Strengthmedium — The paper provides systematic empirical evaluation across 60 expert-curated, cross-domain tasks with multiple runs and validation, producing clear performance metrics; however, it does not make causal claims about economic outcomes, is limited to curated sandbox tasks (no external tool access), and its findings are about benchmark performance rather than real-world productivity or economic impact. Methods Rigorhigh — Benchmark construction and validation are thorough (multi-stage reviews, AI-assisted auditing, 1,500+ sandbox runs, multiple independent runs for models, attention to information leakage and scoring); however, task selection is curated and finite (60 tasks), some model results use single runs, and evaluation excludes external-tool-enabled research workflows, which constrain external validity. Sample60 project-level research tasks spanning 11 scientific domains (mathematics, physics, chemistry, biology, astronomy, materials science, earth science, medicine/biostatistics, computer science, robotics, electrical engineering); evaluated with 18 Agent×Model configurations (various backbone models and harnesses) with typically three independent runs per configuration; tasks include >2,600 interaction turns and >2,400 execution steps; benchmark construction involved >1,300 candidate ideas, ~1,500 sandbox runs, and 31,000+ human hours; reported experiments generally run without external tool access. Themesinnovation productivity human_ai_collab adoption GeneralizabilityTasks are curated and finite (60), so results may not generalize to the full space of scientific projects., Evaluations were performed in sandboxed environments without external tool or internet access, limiting applicability to real-world research workflows that use external data and tools., Performance depends strongly on agent harness and backbone model; results may not transfer to different system architectures or future models., Scoring and task formulation involve expert judgment and may favor certain methods or domains., Benchmark measures capability to perform scientific research, not downstream economic impacts (productivity, wages, firm performance).

Claims (9)

ClaimDirectionOutcomeConfidence & EvidenceDetails
ASI-Bench contains 60 project-level research tasks spanning 11 scientific domains. Research Productivity positive Breadth of benchmark coverage for autonomous scientific research
Reading fidelity high
Study strength medium
n=60
60 tasks across 11 domains
0.18
Across 18 evaluated Agent×Model configurations, average performance declined from 50.91 under full methodological guidance to 29.10 when only the method was specified and to 26.62 when agents had to determine the method themselves. Research Productivity negative Scientific research score under progressively reduced methodological guidance
Reading fidelity high
Study strength high
n=18
−24.29 points from B1 to B3 (50.91 to 26.62)
0.3
Current AI systems remain far from reliable autonomous scientific discovery under the benchmark's low-guidance condition. Research Productivity negative Autonomous scientific research performance
Reading fidelity high
Study strength medium
n=18
Average B3 score of 26.62
0.18
The strongest evaluated configuration, Codex with GPT-5.6 Sol (ultra), achieved a B3 score of 51.60 and was the only evaluated system to exceed 50 when agents had to independently select methods and construct research workflows. Research Productivity positive Best-performing system's autonomous scientific research score
Reading fidelity high
Study strength medium
n=18
51.60 B3 score
0.18
Increasing GPT-5.6 Sol inference-time reasoning from xhigh to ultra increased the B3 score from 40.86 to 51.60, a gain of 10.74 points. Research Productivity positive B3 autonomous scientific research score
Reading fidelity high
Study strength medium
n=60
10.74 points
0.18
Removing detailed procedural guidance while retaining the method caused a larger performance decline than removing the method itself: average performance fell by 21.82 points from B1 to B2, compared with a further 2.48-point decline from B2 to B3. Research Productivity negative Performance loss from removing procedural and methodological guidance
Reading fidelity high
Study strength high
n=18
−21.82 points from B1 to B2; −2.48 points from B2 to B3
0.3
Adding task-irrelevant contextual information under B4 had little effect on performance relative to B3: the average score increased from 26.62 to 26.99, a difference of 0.36 points. Research Productivity null_result Scientific research score under distraction
Reading fidelity high
Study strength medium
n=18
+0.36 points
0.18
The agent harness can substantially change the capability expressed by the same underlying model. Research Productivity mixed Scientific research score as a function of agent harness
Reading fidelity high
Study strength medium
n=4
MiMo V2.5 Pro: +7.08 points; Kimi K2.7: +7.62 points
0.18
Reducing methodological guidance increased average per-task token consumption and execution time: relative to B1, token use rose by 59% in B2, 25% in B3, and 30% in B4, while execution time rose by 32%, 22%, and 18%, respectively. Task Completion Time negative Computational cost and execution time per research task
Reading fidelity high
Study strength medium
n=60
B2: +59% tokens and +32% time; B3: +25% tokens and +22% time; B4: +30% tokens and +18% time
0.18

Notes