GeneBench-Pro: Testing AI Research Judgment in Biology

On June 30, 2026, OpenAI introduced GeneBench-Pro, a research-level benchmark that measures whether AI agents can make the higher-order judgment calls that real computational biology requires, rather than merely recalling facts or running a fixed workflow. The benchmark contains 129 problems spanning genomics, quantitative biology, and translational medicine. Each problem hands the model a realistic, messy dataset, brief experimental context, and a target quantity to estimate, then asks it to choose an analysis path, revise assumptions when diagnostics warrant, and decide when a result is decision-ready.

To avoid the ambiguity that plagues benchmarks built on historical datasets, every GeneBench-Pro problem is generated synthetically from a fully known causal structure. That lets OpenAI simulate the data-generating process, tune difficulty, and grade correctness deterministically against known targets. The team sent 82 of the 129 problems to external domain experts, including graduate students, postdocs, industry scientists, and professors, who assessed realism and whether the target answer was identifiable.

The headline result: OpenAI’s strongest model, GPT-5.6 Sol, reached a 28.7 percent pass rate at its highest reasoning level, rising to 31.5 percent with Pro mode enabled. That is a sharp jump from the earlier GeneBench era, when the best frontier model scored below 5 percent, but it still means frontier models solve fewer than a third of the problems. Reviewers estimated a typical problem would take a human expert roughly 20 to 40 hours, placing the human labor cost in the thousands of dollars per problem.

OpenAI reported that models most often fail by not being cautious enough about data issues such as ancestry swaps or quality-control artifacts, making partial progress but struggling to close the inferential loop the way an experienced researcher would. The company is open-sourcing 10 representative questions and plans to provide a 50-question subset for independent third-party benchmarking.

Sources

Last verified July 6, 2026