OpenAI released GPT-6 Sol and GPT-6 Luna on 2026-09-22, nineteen days after GPT-6 Astra. The company says the two were “trained with similar methods as GPT-6 Astra” and positions them as the cost-efficient tiers of the GPT-6 family, with Astra remaining its best model overall. API prices are half the GPT-5.6 promotional rates: Sol at 2 dollars per million input tokens and 10 dollars per million output (from 4 and 20), Luna at 0.10 dollars input and 0.50 output (from 0.20 and 1.20). They are available in the API as gpt-6-sol and gpt-6-luna and in ChatGPT Work and Codex for paid plans, with Free and Go users getting Luna in the desktop app.
OpenAI’s published numbers lean on cost per task. On AutomationBench 1.0.6, Sol at xhigh effort scores 33.2 percent at 0.27 dollars per task, which OpenAI says beats Claude Opus 5 at max effort (26.9 percent) at 9 percent of its cost. On DeepSWE v1.1, Sol at max effort scores 68.8 percent, within 1.1 points of Claude Fable 5’s 69.9 at roughly 80 percent lower cost per task, and Luna reaches 66.6 percent. On OSWorld 2.0 offline, Sol at xhigh scores 60.5 percent against Opus 5 medium at 60.3. On an internal factuality evaluation built from conversations where users flagged errors, OpenAI says Sol makes about half as many mistakes as its predecessor.
The release came with prompt-caching changes for GPT-6: higher default cache hit rates, 90 percent discounts on cached input reads, explicit cache breakpoints, and the ability to change reasoning effort or toggle tools without breaking the cache. OpenAI quotes GitHub as saying these improvements cut the share of prompt tokens needing fresh processing by more than 50 percent across billions of Copilot requests. The post also discloses that daily token usage at OpenAI, valued at API prices, exceeds 600 dollars for the median researcher and 7,000 dollars at the 90th percentile.
Why it matters: Sol and Luna landed the same day as Anthropic’s Claude Opus 5.5, and both launches argued on price per unit of capability rather than on raw frontier scores, a sign that competition in the middle tier is now about cost curves. What it does not show: competitor scores are taken from other labs’ published reports, some comparisons are against Fable 5 rather than Fable 5.1 where 5.1 numbers were unavailable, and the factuality evaluation is internal and deliberately drawn from error-inducing conversations rather than typical use.