Overview
<cite index="2-12,2-13">Grok 4.6 is xAI's newest flagship Large Language Model (LLM), released on August 12, 2026, as a post-training upgrade to Grok 4.5 that retains a 500,000-token context window while improving reported reasoning and coding results.</cite> <cite index="5-1,5-2">The model is described by xAI as a frontier system built for coding, agentic tasks, and knowledge work, with a particular focus on long-running agents and more ambitious interactive work.</cite> <cite index="6-3">xAI did not publish a parameter count for Grok 4.6.</cite>
Benchmark Performance
<cite index="6-4">On xAI's launch table, Grok 4.6 (High reasoning setting) scores 61 on the Artificial Analysis Intelligence Index, up from 56 for Grok 4.5 and tied with GPT-5.6 Sol Max.</cite> <cite index="10-6,10-7,10-8,10-9,10-10">The composite index score of 61 is tied with GPT-5.6 Sol Max and sits one point behind Fable 5 Max (62), while on the GDPVal-AA v2 professional knowledge benchmark Grok 4.6 leads all listed models at an Elo of 1,753. On CursorBench v3.2, the model scores 69.9%, ahead of GPT-5.6 Sol Max (67.2%) but behind Fable 5 Max (70.5%). On FrontierCode v1.1 Extended, it scores 61.3%, edging GPT-5.6 Sol Max (60.6%). On the DeepSWE v1.1 autonomous software engineering benchmark, Grok 4.6 reaches 65.9% — a gain of 11.9 percentage points over Grok 4.5, but trailing GPT-5.6 Sol Max at 73%.</cite>
<cite index="10-13">On the APEX-Agents benchmark — which evaluates multi-step agentic task completion — Grok 4.6 scores 57.5%, placing it between GPT-5.6 Sol Max and Fable 5 Max.</cite> <cite index="11-11">This represents a gain of 10.4 percentage points over Grok 4.5's score of 47.1% on the same benchmark.</cite> <cite index="14-2">Terminal-Bench v3.0 reaches 26%, nearly double Grok 4.5's 15.7%, though still last among the four models listed in xAI's launch table.</cite>
<cite index="14-5,14-6">Independent analysis notes that bolded wins for Grok 4.6 on GDPval-AA v2 and AA-Briefcase sit inside Artificial Analysis' published confidence intervals, making them statistical ties rather than confirmed leads. The comparison set also excludes Anthropic's Claude Opus 5, which currently tops the Artificial Analysis Intelligence Index.</cite>
Pricing Structure
<cite index="4-14,4-15">Below 200,000 prompt tokens, Grok 4.6 is priced at $2 per million input tokens, $0.50 per million cached input tokens, and $6 per million output tokens. In xAI's long-context band — activated when prompt input reaches 200,000 tokens — those rates double to $4, $1, and $12 per million tokens respectively, and the higher rates apply to all tokens in the request, not just those exceeding the threshold.</cite>
<cite index="17-8,17-9">Once a prompt enters xAI's long-context band, the full request is repriced at the elevated tier — a caveat that analysts note is the difference between a useful price comparison and a misleading one.</cite> <cite index="16-14">Between Grok 4.5 and 4.6, input and output rates are identical at the headline level; the one structural difference is cached input, where Grok 4.5 is cheaper at $0.30 per million against Grok 4.6's $0.50, meaning cache-heavy workloads cost slightly more on the newer model.</cite>
Reasoning and Distribution
<cite index="6-2">The model's `reasoning_effort` parameter now supports four levels — low, medium, high (default), and a new `xhigh` setting.</cite> <cite index="8-2">Independent evaluation finds that `xhigh` reasoning has diminishing returns, scoring 70.8% on CursorBench at $2.81 per task compared to 69.9% at $2.34 for the High setting.</cite>
<cite index="8-3">Grok 4.6 is the default model in Grok Build, available across all Cursor plans, listed by OpenRouter, and supported in Vercel AI SDK 4.0.37 with `xhigh` reasoning.</cite> <cite index="3-11">Amazon Web Services added the model to its Bedrock managed inference service on August 19, 2026, with four configurable reasoning tiers and the 500,000-token context window, with Low set as the default reasoning level on that platform.</cite>
Roadmap Context
<cite index="13-3">A follow-up model, Grok 4.7, is expected within three to four weeks of the 4.6 launch, according to statements attributed to Elon Musk, with Grok 5 targeted before year-end.</cite> <cite index="10-21,10-22">Analysts characterize Grok 4.6 as the post-training release in xAI's current cycle, with Grok 4.7 expected to represent a more substantial architectural scale jump — suggesting xAI is simultaneously exploring how much performance can be extracted from an existing base model and how much a larger base adds on top.</cite>