SpaceXAI released Grok 4.6 on August 12, 2026, a 1.5-trillion-parameter model built for long-running agentic tasks and visual/interactive work. The headline benchmark is that Grok 4.6 matches GPT-5.6 Sol Max at 61 on the Artificial Analysis Intelligence Index (a composite of 9 benchmarks). The model scores 1753 ELO on GDPVal-AA v2 (knowledge work) and leads on CursorBench 3.2 (69.9%) and FrontierCode. Pricing is $2 per million input tokens, $6 per million output tokens—roughly 60% cheaper than Claude Opus 5 ($5/$25) or GPT-5.6 Sol ($5/$30).
Grok 4.6 is trained for multi-step agentic workflows: research, codebase analysis, CAD, web development, kernel optimization. xAI applied longer supplemental training than Grok 4.5, curated model-generated data for reasoning, and used Grok 4.5 itself to regenerate SFT trajectories across domains (STEM, software engineering, knowledge work), filtering problematic traces with model-based checks. RL was applied across agentic tasks including domain-specific environments. xAI reports improved self-testing and verification behavior on longer trajectories, and stronger visual/interactive first passes.
The launch strategy targets developers in Cursor (the AI-native code editor) with 2x usage credit for the first week. Grok 4.6 is live in Cursor, Grok Build, and the API; available via OpenRouter, Vercel, and Cloudflare. Elon Musk confirmed Grok 4.7—a larger 2.1T-parameter model—is already in training, expected within weeks; xAI is signaling a 2-3 week release cadence.
For agent builders and platform architects: Grok 4.6 competes on cost and agentic step-efficiency (reportedly ~53 steps vs ~103 for Claude Opus 5 on multi-step tasks), not raw throughput. The 1.5T foundation reused from 4.5 keeps latency and token overhead low. Independent verification (LMSYS Chatbot Arena) is pending; treat the AA Intelligence Index parity as provisional. Watch whether the Cursor distribution strategy translates to meaningful adoption shifts, and whether Grok 4.7's aggressive timeline holds.