ASI-Bench

ASI-Bench was submitted to arXiv on August 18, 2026 (paper 2608.17271) by a group of more than forty authors led by Junwei Zhou. Its stated aim is to test whether AI systems can, in the paper’s words, “move beyond mastering existing knowledge toward exploring the unknown, creating new knowledge, and turning new ideas into verifiable results.” Where most agent benchmarks score discrete tasks, ASI-Bench scores 60 project-level research tasks spanning 11 scientific domains, each built to require a full investigation rather than a single answer.

The design choice that gives the benchmark its teeth is a graduated reduction in methodological guidance. When agents are given complete guidance on how to approach a task, the reported score is 50.91; when they must determine the method themselves, it falls to 26.62. That gap is the measurement. It separates execution ability - carrying out a research plan someone else supplied - from research ability, which is deciding what to do in the first place. The authors report evaluating 18 agent-model configurations, and say the construction of the benchmark consumed roughly 31,000 human hours of expert effort from more than 40 experts.

The authors’ conclusion is that current systems remain heavily dependent on human direction and cannot yet autonomously carry out comprehensive, project-level scientific investigation. That finding lands in the same week as a broader argument in the literature and press that AI is not ready to do its own research, and it is consistent with earlier scientific-agent evaluations in this library that found agents strong at execution and weak at judgment.

For a leader deciding where to put an AI-for-science budget, the useful takeaway is the shape of the gap, not the absolute numbers. A halving of performance when method selection is handed to the agent suggests the practical deployment pattern for the near term is a human-specified protocol executed at machine scale, not an autonomous research program. It also suggests a specific procurement question: any vendor claiming autonomous discovery should be asked what its system scores when the method is not supplied. As with any single benchmark, the scores are self-reported by the benchmark’s authors and the 60-task set is small enough that domain coverage, not agent capability, could drive part of the result.

Sources

Last verified August 24, 2026