Computer research lab with multiple monitors

A New AI Benchmark Finds a Big Gap in Autonomous Scientific Research

Computer research lab with multiple monitors
Photo by 铮 夏 on Unsplash

A new benchmark for AI agents is testing a harder question than whether a model can answer a difficult science question: can an AI system decide how to investigate a research problem and carry the project through with much less human guidance?

The paper, ASI-Bench: At the Dawn of Artificial Superintelligence, evaluates project-level scientific research rather than isolated questions. The distinction is important because real research requires choosing methods, interpreting evidence, revising plans and producing results that can be checked.

What ASI-Bench measures

The benchmark contains 60 project-level research tasks across 11 scientific domains. According to the paper, more than 40 experts spent over 31,000 human hours building and validating the benchmark.

The researchers tested 18 agent-model configurations while progressively removing methodological guidance. In the most guided setting, agents received substantial direction about how to approach a project. In harder settings, the systems had to make more of the methodological decisions themselves.

The autonomy gap

The reported average score fell from 50.91 with full methodological guidance to 29.10 when only the method was specified, and to 26.62 when agents had to determine the method themselves.

That is a large drop. It suggests that current systems can benefit substantially from humans who already understand the research problem and can provide a useful methodological path.

In other words, strong performance with an expert-designed workflow should not automatically be interpreted as autonomous scientific discovery.

Why this matters for AGI

AGI discussions often focus on whether AI can solve increasingly difficult problems. But an important part of general intelligence is deciding which problems to investigate, choosing sensible strategies and adapting when the first strategy fails.

Scientific research is a particularly useful stress test because it combines knowledge, reasoning, experimentation, uncertainty and long-horizon planning. A system that can produce a good answer to a known question is different from a system that can independently turn an open-ended question into a defensible research result.

What the benchmark does—and does not—prove

ASI-Bench is evidence about a specific class of research tasks. It is not a universal test of intelligence, and a benchmark score cannot establish whether AGI has or has not been achieved.

There are also practical limitations. Research agents may perform differently when given better tools, larger context windows, stronger models, domain-specific software or more opportunities to interact with humans. Benchmark design also influences what capabilities are rewarded.

The useful takeaway is narrower: removing human methodological guidance exposes a capability gap that can be hidden by heavily scaffolded workflows.

What to watch next

  • Whether future agents can choose research methods more reliably.
  • Whether they can detect and recover from flawed experimental plans.
  • Whether results remain reproducible when humans provide less supervision.
  • Whether independent evaluations agree with benchmark-specific gains.
  • Whether agents can transfer research strategies between scientific domains.

Our take

The most interesting part of this work is not the headline number. It is the attempt to measure the difference between executing a research workflow and designing one.

That distinction will matter increasingly as AI moves from chat interfaces toward agents that can browse, write code, run experiments and operate software. The more autonomy a system receives, the more important it becomes to measure not just final answers, but the quality of the decisions made along the way.

Primary source: ASI-Bench: At the Dawn of Artificial Superintelligence.

For more context, read our guide to AGI and follow our AI Research coverage.

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *