StateM: harness scaling on Terminal-Bench 2.1
A runtime harness lifts GPT-5.6 Sol to 95.3 percent on Terminal-Bench 2.1 for about 15 dollars, without touching model weights.
What the papers actually said - linked to the originals.
A runtime harness lifts GPT-5.6 Sol to 95.3 percent on Terminal-Bench 2.1 for about 15 dollars, without touching model weights.
A Berkeley-led system serves mixture-of-experts models up to 753B parameters on a single personal machine with an 8GB GPU.
Anthropic reports Claude autonomously designed 354 confirmed protein binders from 1,320 designs, succeeding on 14 of 15 targets.
Microsoft Research shipped Skala 1.1, trained on 2.5 times more data, reporting 2.8 kcal/mol weighted average error on GMTKN55.