VibeThinker-3B is a dense 3-billion-parameter model that probes how far verifiable reasoning can be pushed in a deliberately small model. According to the arXiv report submitted on June 15, 2026, the model “attains a score of 94.3 on AIME26 (improving to 97.1 with claim-level test-time scaling),” and the authors report it matching or exceeding flagship systems that are orders of magnitude larger on competition-grade math and code benchmarks.
The result was achieved through a post-training pipeline rather than a new large pretraining run: curriculum-based supervised fine-tuning, multi-domain reinforcement learning, and self-distillation applied on top of a small base model. The headline claim is that careful post-training can elicit large-model reasoning behavior from a model small enough to run in a few gigabytes of memory.
The work matters because it challenges the assumption that frontier math and code reasoning requires very large models. If a 3B model can approach flagship scores on hard benchmarks, the cost and deployment story for reasoning shifts substantially. As with any single-benchmark claim, the numbers warrant independent scrutiny, but the direction is notable for anyone weighing the price of reasoning at scale.