1 cumulative citations
View corpus contextA new benchmark for autonomous scientific research finds current AI agents can implement prescribed procedures but struggle to design and execute research independently: average scores fall from about 51% with full guidance to 27% when methods must be chosen by the AI, highlighting a substantial gap to truly autonomous scientific discovery.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Artificial superintelligence (ASI) requires AI to move beyond mastering existing knowledge toward exploring the unknown, creating new knowledge, and turning new ideas into verifiable results. However, the capabilities of today's AI systems are still largely built on learning, compressing, and applying existing human knowledge. Accordingly, existing benchmarks primarily test whether AI can produce correct answers based on learned knowledge, or whether it can complete tasks under extensive human guidance. We therefore introduce ASI-Bench, the first benchmark to jointly evaluate AI systems' capabilities of innovative exploration and autonomous scientific execution across general research domains, and the first to progressively withdraw human methodological guidance within the same research project to test how far AI can proceed on its own. Built by over 40 experts with the cost of 31,000+ human hours, ASI-Bench contains 60 project-level research tasks across 11 scientific domains and progressively reduces methodological guidance to test whether AI can independently select methods, conduct research, and produce verifiable results. All tasks undergo expert review, AI-assisted auditing, sandbox execution, and scorer validation. Across 18 state-of-the-art agent--model configurations, the average score drops from 50.91 with full methodological guidance to 29.10 with only the method specified and 26.62 when agents must determine the method themselves. This sharp decline shows that current systems remain heavily dependent on human guidance and are still far from autonomously conducting end-to-end, project-level scientific research. ASI-Bench is open to the world. We invite researchers and builders everywhere to contribute new tasks, challenge the limits of today's AI, and help accelerate humanity's collective path toward artificial superintelligence at https://asibench.apexin.ai/submit.
Summary
Main Finding
ASI-Bench is a new, large-scale benchmark for measuring AI systems’ ability to perform open-ended, project‑level scientific research as human methodological guidance is progressively withdrawn. Evaluated across 60 cross‑domain research tasks and 18 Agent×Model configurations, performance drops sharply when detailed procedural guidance is removed (mean score: B1 = 50.91 → B2 = 29.10 → B3 = 26.62), revealing that current systems remain heavily dependent on human-prescribed procedures and are far from reliably autonomous scientific discovery.
Key Points
- Benchmark structure (guidance gradient):
- B1: Full methodological guidance (detailed step-by-step procedure provided).
- B2: Only the method class/name provided (no procedural steps).
- B3: No methodological guidance—agent must choose method and workflow.
- B4: B3 plus plausible but irrelevant distractors.
- Main quantitative results (macro-averaged over tasks; 3 runs unless noted):
- Mean scores: B1 = 50.91, B2 = 29.10, B3 = 26.62, B4 = 26.99.
- Largest drop is B1→B2: −21.82 points (method operationalization loss).
- Smaller drop B2→B3: −2.48 points (method selection less limiting than operationalization).
- Only one configuration (GPT-5.6 Sol (ultra) / Codex harness) exceeded 50 in B3 (51.60).
- Harness (agent system) matters: same backbone can perform substantially differently under different harnesses; capability emerges from model × harness interaction, not model scale alone.
- Computational cost increases when guidance is reduced:
- Avg. tokens per task: B1 = 4.4M, B2 = 6.9M (+59%), B3 = 5.4M (+25%), B4 = 5.7M (+30%).
- Avg. time per task: B1 = 37.8 min, B2 = 49.7 min (+32%), B3 = 45.9 min (+22%).
- Current systems are robust to irrelevant distractors (B3 vs B4 nearly unchanged), but fragile in turning a method into a working, end-to-end research procedure.
Data & Methods
- Benchmark composition:
- 60 project-level research tasks spanning 11 domains: mathematics, physics, chemistry, biology, astronomy, materials science, earth science, medicine & biostatistics, computer science, robotics, electrical engineering.
- Each task: scientific objective, task-specific data, executable sandbox environment, verifiable artifacts, and fixed scoring criteria.
- Tasks are long-horizon and multi-stage (over 2,600 interaction turns, ~2,400 execution steps, >35 hours total agent execution across tasks).
- Construction & validation pipeline:
- Started from >1,300 candidate ideas; iterative process produced final 60 tasks.
- Extensive human effort: >31,000 human-hours, >1,100 review assignments, >2,000 task revisions, 5 review rounds.
-
1,500 sandbox runs to verify reproducibility, scoring, and to eliminate shortcuts or leakage.
- Evaluation:
- 18 representative Agent×Model configurations tested (various backbone LMs and harnesses).
- Scores macro-averaged across tasks; most results are means over three independent runs (some single-run results noted).
- Diagnostics computed per run then aggregated (e.g., B2−B1, B3−B1, B4−B3).
- Transparency & extensibility:
- ASI-Bench is public and invites external task contributions and leaderboard submissions.
Implications for AI Economics
- Where to allocate R&D investment:
- High returns likely from investing in agent harness design, tool integration, and method operationalization infrastructure (turning high‑level methods into robust end‑to‑end procedures) rather than solely scaling backbone models.
- Engineering effort that reduces the human-to-agent procedural gap could produce larger practical gains for autonomous research than equivalent expenditures on model scale alone.
- Labor and task substitution:
- Current systems’ heavy dependence on procedural guidance implies AI will more immediately augment human researchers (improving productivity, automating laborious implementation/testing) rather than fully replace them in open-ended research.
- Demand will remain for human domain experts who craft procedures, validate results, and close evaluation gaps—supporting complementary human capital and high-skill employment in research ecosystems.
- Market structure and concentration:
- High compute, data, and engineering costs (large token/time per task, sandbox runs) favor well‑resourced organizations. Institutions that can invest jointly in models, harnesses, and validated sandboxes may obtain durable competitive advantages, increasing concentration in AI-driven scientific discovery.
- Leaderboards and public benchmarks (like ASI-Bench) can mitigate information asymmetries but only if they remain well-maintained and resistant to gaming.
- Pricing and ROI for AI-for-Science products:
- Buyers should account for substantial integration and engineering costs to realize autonomous behavior; product pricing and adoption models should reflect combined model + systems development rather than treating model access in isolation.
- For investors, cost–performance tradeoffs shown in ASI-Bench imply careful evaluation of marginal gains from compute spend: higher cost systems can improve B3 performance, but the operationalization bottleneck may limit returns from pure compute increases.
- Policy and regulation economics:
- Reproducibility, audit trails, and sandboxed verification (as implemented by ASI-Bench) are crucial public goods; funding and standards for independent validation infrastructure could have large social value.
- Because the benchmark reveals limited true autonomy today, regulatory focus can prioritize standards for claimed autonomous scientific systems (verification, transparency, human oversight) to avoid premature deployment claims and misaligned incentives.
- Emerging markets:
- Demand for high-quality, domain-specific sandboxes, verified datasets, automated validation tooling, and benchmark‑engineering expertise will grow—creating opportunities for specialized firms and service providers that bridge models and domain workflows.
Summary: ASI-Bench quantifies a key gap in present AI capability: converting a methodological idea into a robust, end‑to‑end research workflow. For AI economics, this implies most near‑term value comes from systems engineering, validation infrastructure, and complementary human expertise rather than model scaling alone—shaping R&D priorities, labor markets, investment strategies, and regulatory needs.
Assessment
Claims (9)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| ASI-Bench contains 60 project-level research tasks spanning 11 scientific domains. Research Productivity | positive | Breadth of benchmark coverage for autonomous scientific research |
Reading fidelity
high
Study strength
medium
|
n=60
60 tasks across 11 domains
|
| Across 18 evaluated Agent×Model configurations, average performance declined from 50.91 under full methodological guidance to 29.10 when only the method was specified and to 26.62 when agents had to determine the method themselves. Research Productivity | negative | Scientific research score under progressively reduced methodological guidance |
Reading fidelity
high
Study strength
high
|
n=18
−24.29 points from B1 to B3 (50.91 to 26.62)
|
| Current AI systems remain far from reliable autonomous scientific discovery under the benchmark's low-guidance condition. Research Productivity | negative | Autonomous scientific research performance |
Reading fidelity
high
Study strength
medium
|
n=18
Average B3 score of 26.62
|
| The strongest evaluated configuration, Codex with GPT-5.6 Sol (ultra), achieved a B3 score of 51.60 and was the only evaluated system to exceed 50 when agents had to independently select methods and construct research workflows. Research Productivity | positive | Best-performing system's autonomous scientific research score |
Reading fidelity
high
Study strength
medium
|
n=18
51.60 B3 score
|
| Increasing GPT-5.6 Sol inference-time reasoning from xhigh to ultra increased the B3 score from 40.86 to 51.60, a gain of 10.74 points. Research Productivity | positive | B3 autonomous scientific research score |
Reading fidelity
high
Study strength
medium
|
n=60
10.74 points
|
| Removing detailed procedural guidance while retaining the method caused a larger performance decline than removing the method itself: average performance fell by 21.82 points from B1 to B2, compared with a further 2.48-point decline from B2 to B3. Research Productivity | negative | Performance loss from removing procedural and methodological guidance |
Reading fidelity
high
Study strength
high
|
n=18
−21.82 points from B1 to B2; −2.48 points from B2 to B3
|
| Adding task-irrelevant contextual information under B4 had little effect on performance relative to B3: the average score increased from 26.62 to 26.99, a difference of 0.36 points. Research Productivity | null_result | Scientific research score under distraction |
Reading fidelity
high
Study strength
medium
|
n=18
+0.36 points
|
| The agent harness can substantially change the capability expressed by the same underlying model. Research Productivity | mixed | Scientific research score as a function of agent harness |
Reading fidelity
high
Study strength
medium
|
n=4
MiMo V2.5 Pro: +7.08 points; Kimi K2.7: +7.62 points
|
| Reducing methodological guidance increased average per-task token consumption and execution time: relative to B1, token use rose by 59% in B2, 25% in B3, and 30% in B4, while execution time rose by 32%, 22%, and 18%, respectively. Task Completion Time | negative | Computational cost and execution time per research task |
Reading fidelity
high
Study strength
medium
|
n=60
B2: +59% tokens and +32% time; B3: +25% tokens and +22% time; B4: +30% tokens and +18% time
|