Grok 4.6 Lands: Ten Benchmarks and One Pricing Cliff
August 12, 2026 · by Mira Adeyemi-Quist
Breaking News
Grok 4.6 Ships After a Slipped Date, and for the First Time in Weeks There Is Something to Measure
It landed today, five days after the target and roughly forty-eight hours after this desk spent nine hundred words on the fact that it had not. That is a correction I am delighted to run.
The launch table shows improvement across all ten benchmarks. The two that matter: on the Artificial Analysis composite it reaches 61, matching GPT-5.6 Sol Max — the first time xAI has drawn level with a frontier competitor on that index rather than near it — and on CursorBench v3.2 it posts 69.9 percent against 66.7 for Grok 4.5 High. It is available today through the API, Grok Build, Cursor, OpenRouter, Vercel and Cloudflare, which is the part that separates a model from a post.
So let us do the thing this paper exists to do, which is to check rather than cheer.
Three groups have already published independent runs, and they cluster below the launch figures — the tightest nine points under, the loosest fourteen. The gaps are attributed variously to prompt formatting, sampling temperature, and what one author described in a footnote as “a scoring script we were not given.” No party disputes the number. Every party disputes the number's parents.
The three attempts were made by groups with no relationship to each other and no particular axe to grind — two university labs and one company that would benefit commercially if the launch figures held. All three landed low. The spread between the three independent attempts is smaller than the spread between any of them and the published figure, which is the detail that ought to be in every write-up of this launch and is in none of them.
There are boring explanations and they are probably right. Evaluation harnesses differ. A single space in a prompt template moves scores by more than most people would believe. Sampling at temperature zero and sampling at 0.7 are not the same experiment, and papers routinely fail to say which they ran. None of this is fraud, and I want to be careful, because “irreproducible” has a meaning in science that this situation does not rise to.
But there is a structural problem underneath the boring explanations, and it is this: the party publishing the number is also the party defining the test, running the test, and choosing whether to release the harness. In any other technical field that arrangement would be described as a press release and priced accordingly. Here it is described as a result and lands on a leaderboard within a day.
The fix is not complicated and has been proposed for years by people with more standing than me. Publish the harness. Publish the prompts, exactly, whitespace included. Publish the sampling parameters. State the number of runs and give the variance rather than the maximum. Every lab that has done this has been rewarded for it, and every lab that has not has been rewarded slightly more, which is the entire problem in one sentence.
Until then the honest way to read any headline eval score is as an upper bound produced under favorable conditions by an interested party. That is not nothing. It is just not what the chart implies.
Gossip
Demo Video Runs Two Minutes, Contains Four Cuts, All in the Same Place
Frame analysis of the promotional clip finds the timer in the corner of the screen advancing by fifty-one seconds during a transition described in the caption as “realtime.”
A spokesperson said the edit was made for length. The edit was made in the middle of the loading spinner.
Opinion
Window Advertised at 500,000 Tokens; Wallet Advertises a Different Number at 200,000
Half a million tokens will hold every email you have ever sent, every email sent to you, and the complete works of a mid-tier Victorian novelist, simultaneously.
It will also, at token 200,001, begin charging you double for all of them. The window is 500,000 tokens wide. The comfortable part of it is 200,000, and the difference is not documented on the poster.
Breaking News
Cross 200,000 Tokens and the Rate Doubles — Retroactively, Across Every Token in the Request
Grok 4.6 lists at $2 per million input tokens and $6 per million output, which is genuinely aggressive and roughly half what the nearest frontier competitors ask.
Then there is the long-context band. Above 200,000 tokens the rate becomes $4 and $12 — and per xAI's own documentation the higher rate applies to every token in that request, not merely the ones past the line.
Work the example, because the shape of it matters. A 199,000-token prompt costs about forty cents to send. A 210,000-token prompt — six percent more text — costs about eighty-four. Eleven thousand extra tokens, and the bill doubles.
This is not hidden and it is not a scandal; it is on the pricing page and there are real serving costs behind it. But it is a cliff rather than a slope, it sits at a boundary nobody's tooling watches by default, and the model will happily accept the request that crosses it without mentioning anything.
If you are building on this: count your tokens before you send, not after you are billed. Cost is the benchmark. It always was. It just doesn't make a good poster.
Gossip
Executive Says the Letters, Stock Moves Two Percent, Nobody Defines the Letters
The three letters were deployed Thursday in an interview, prompting a market reaction, four hundred quote-posts, and zero attempts by any participant to say what would have to be true for the letters to apply.
This desk maintains a standing offer: define it in one sentence, on the record, and we will print it.