On September 22, 2026, Epoch AI published “The plunging price of thought” by Luke Emberson and David Roodman. The report asks how fast the cost of reaching a fixed level of AI performance is falling. Rather than re-run every model at many token budgets, which would be slow and expensive, the authors followed a procedure developed by the federal Center for AI Standards and Innovation (CAISI) that uses a transcript from a high-budget benchmark run to predict performance under tighter budgets. They then fitted models to the resulting cost-performance frontiers across five benchmarks covering math, science and games of skill, over roughly three years of releases.
The headline: the cost of a given level of performance has fallen about 47 percent per quarter since 2023, or about 13x per year. The rate varies by domain, from roughly 39 to 43 percent per quarter (7 to 10x per year) on game-based puzzles to 50 to 52 percent per quarter on math problems. Declines are steepest right after a level first becomes state of the art: averaged across the five benchmarks, cost falls 66 percent per quarter (75x per year) for newly state-of-the-art performance, slowing to 32 percent per quarter (4.7x per year) two years later. As an illustration, the report says GPT-5.6 Luna matched an earlier GPQA Diamond score at about 0.0004 dollars per question, a 725-fold drop in under 18 months.
Epoch compares the rate with other technologies: four times faster than the fall in DNA sequencing costs, six times faster than compute, 18 times faster than lithium batteries and 54 times faster than electricity in the century to 1973. Data and code are published on GitHub.
The practical point is that a capability which is expensive at launch should be assumed cheap within a year or two, which changes what is worth building now versus waiting for. The caveats are the authors’ own and they are substantial: labs may be optimizing for these specific benchmarks, benchmarks are an imperfect proxy for useful work, users do not always switch to the cheapest adequate model, three years is a short series, and reasonable averaging choices move the quarterly rate within a band of several percentage points. It is a measure of benchmark price-performance, not of what any customer actually pays for real work.