xAI Launches Grok 4.6, Ranking in Top Tier on Agentic Benchmarks at Roughly Half Competitors' Pricing
Nashnova编辑部
Elon Musk's xAI released Grok 4.6, posting top-tier scores on multiple agent benchmarks while pricing its API at roughly half of competing frontier models — performance parity meets a price war.
How did Grok 4.6 actually score?
On GDPVal-AA v2, Grok 4.6 hit 1,753 Elo — above GPT-5.6 Sol Max at 1,728 and Claude Fable 5 Max at 1,741, and a major jump from Grok 4.5 High's 1,526.
On the Artificial Analysis Intelligence Index it scored 61, tying GPT-5.6 Sol Max.
It does not lead everywhere: on DeepSWE 1.1 and Terminal-Bench v3.0, GPT-5.6 Sol Max still scores higher. This means → Grok 4.6 has entered the top tier, but has not swept every benchmark.
Half the price — is that real?
API starting price: $2 per million input tokens, $6 per million output tokens, with a faster, double-priced tier also available. xAI says this is roughly half of other frontier models.
Atreides Management CIO Gavin Baker noted Grok 4.6 performs roughly on par with Fable 5 Max but costs 80% less on input and 88% less on output. Analyst Kim Monnis added that it undercuts even Sonnet 5 while matching top-tier performance.
In plain terms = "half the price" compares list-price API rates. Actual cost per task varies — different models consume tokens differently and take different reasoning paths, so a 50% saving on every job is not guaranteed.
What changed in agent capability?
Grok 4.6 went through longer supplementary training than its predecessor, incorporating curated model-generated reasoning data, high-quality engineering data, and an improved optimizer.
The key shift: on extended tasks the model now self-tests and self-verifies — it checks earlier work before moving to the next step. In plain terms = it went from "finish first, hand in later" to "check as you go."
Benchmark gains: CursorBench v3.2 reached 69.9% (up from 66.7%); APEX-Agents hit 57.5% (up from 47.1%); AA-Briefcase scored 1,577 (up from 1,313).
Where is Grok 4.6 available — and what comes next?
Grok 4.6 is live on Cursor and Grok Build, and available through OpenRouter, Vercel, and Cloudflare. First-week usage on Cursor and Grok Build includes a 2× included-usage boost.
Gavin Baker expects the next iteration, Grok 4.7, to be a significantly larger model that may fold in Cursor and SpaceX data for pre-training — but that is an investor's projection, not a confirmed xAI product roadmap.
This reflects a broader dynamic: as frontier models converge on capability, the price-to-performance balance is becoming the decisive variable for developers. Whether Grok 4.6's low-price strategy converts into real market share is the key test to watch.
Content is for reference only, not financial advice.