A New AI Benchmark Finds a Big Gap in Autonomous Scientific Research
A new benchmark tests whether AI agents can conduct scientific research with progressively less human methodological guidance. The results reveal a major autonomy gap.
A new benchmark tests whether AI agents can conduct scientific research with progressively less human methodological guidance. The results reveal a major autonomy gap.