SpaceXAI released Grok 4.6 on 2026-08-12, three and a half weeks after Grok 4.5. The company frames it as an incremental build on the same line with “a particular focus on long-running agents and more ambitious interactive and visual work.” Pricing starts at 2 dollars per million input tokens and 6 dollars per million output tokens, with a fast variant at twice that. It shipped the same day into Cursor, Grok Build, the SpaceXAI API, and partners including OpenRouter, Vercel and Cloudflare, with double included usage in Cursor and Grok Build for the first week.
The headline claim is parity rather than a lead. SpaceXAI says Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index, matching GPT-5.6 Sol at 61 and trailing Claude Fable 5 at 62, up from 56 for Grok 4.5. Its published table has Grok 4.6 ahead on a handful of evals - CursorBench v3.2 at 69.9 percent, GDPVal-AA v2 at 1753, AA-Briefcase at 1577, Harvey LAB at 15.8 percent - and clearly behind on others, notably Terminal-Bench v3.0 at 26 percent against 34.6 for GPT-5.6 Sol and 34.1 for Fable 5, and DeepSWE v1.1 at 65.9 percent against 73 for Sol. The company notes competitor figures are drawn from those developers’ own system cards or leaderboards, which is a weaker comparison basis than independent re-evaluation.
On method, SpaceXAI says Grok 4.6 had a longer supplemental training run than 4.5 using curated model-generated data, then used Grok 4.5 itself to regenerate the supervised fine-tuning trajectories across reasoning efforts and agent harnesses, filtering bad traces with model-based checks. Reinforcement learning covered agentic tasks including kernel optimization, web development and computer-aided design. The qualitative claim the company leans on is that on longer trajectories the model started “self-testing and verification,” checking its own work before moving on, and that it produces stronger first passes on visual and interactive projects.
For a buyer, the interesting fact is not the benchmark table but the distribution. Grok 4.6 launched inside Cursor on day one, which is now a sibling product following SpaceX’s acquisition of Anysphere, and reached Amazon Bedrock, Google’s enterprise agent platform and GitHub Copilot within a week. A model that is roughly at parity on a composite index and materially behind on terminal-agent benchmarks competes on price and placement rather than on capability, and the vertical integration with a leading coding tool is the sharper part of this release. Treat the self-verification claim as unproven until third parties measure it.