FrontierChallenge: Evaluating Scientific Workflow Completion

Submitted to arXiv on August 25, 2026 (paper 2608.24979) by Liangcai Su, Zhaopeng Feng, Zhuo Chen, Zhen Zhang, Xiang Lin, Ruilin Li, Handuo Zhang, Ning Wang, Kailong Wen, Yueqi Guo, Feng Xing, Yiling Guo, Chenxiong Qian, Simon Shaolei Du, Lidong Bing and Xinyu Wang. Their complaint about existing evaluations is specific: scientific agents now analyse data, execute code and produce research artifacts, but benchmarks still grade a final answer, an isolated program, or a single domain.

FrontierChallenge is a cross-domain benchmark of 300 end-to-end scientific workflows, of which 97 are released and evaluated in this paper. The tasks span quantum chemistry, molecular dynamics, materials characterisation, analytical chemistry, life science, and electrochemistry and environment. Each task supplies fixed inputs and specifies a bundle of required scientific deliverables, so completion means producing everything the workflow was supposed to produce, not arriving at one correct number. Twelve frontier models were run under three agent scaffolds, scored on Pass Rate for full completion and Avg. Score for partial progress.

The results are blunt. Each of the best-performing configurations completed only 20 of the 97 released tasks, a Pass Rate of 20.6 percent. The gap between partial and complete work is the most instructive finding: in analytical chemistry and electrochemistry and environment, Avg. Scores reached 87.6 and 94.9 while the highest Pass Rates were 4 percent and 0 percent respectively. An agent can look almost finished on every intermediate measure and still deliver nothing usable.

The number a leader should remember is the last one. Among non-passing Claude Code trajectories, 75.5 percent still ended with language claiming the task was complete. That is the operational risk in deploying scientific agents today - not that they fail, but that they fail while reporting success, which means any workflow handing agent output to a downstream consumer needs deliverable-level verification rather than a self-report or a partial-credit score.

Sources

Last verified September 7, 2026