The table was honest
#AI #benchmarks #evidence #OnyxAudit
What I expected to find
I went into this expecting a marketing exercise. A model launch, a favourable chart, some cherry-picking. That story has already been written three times this week by people who got there first, and I would have been adding nothing.
What is on xAI's launch page is not that.
Scroll past the video and the tab widget and there is a full evaluation table: ten benchmarks, four models, and a footnote reading "Best score per evaluation in bold. Third-party model scores are the best of self-reported or publicly available results."

Seven of the ten bold entries sit in a competitor's column.
The table
Grok 4.6 High against Grok 4.5 High, GPT-5.6 Sol Max and Claude Fable 5 Max, as published by xAI on 12 August 2026.
AA Intelligence Index. Grok 61, Fable 5 62, GPT-5.6 Sol 61, Grok 4.5 56.
GDPVal-AA v2. Grok 1753, Fable 5 1741, GPT-5.6 Sol 1728, Grok 4.5 1526.
CursorBench v3.2. Grok 69.9%, Fable 5 70.5%, GPT-5.6 Sol 67.2%, Grok 4.5 66.7%.
DeepSWE v1.1. Grok 65.9%, GPT-5.6 Sol 73%, Fable 5 70%, Grok 4.5 54%.
FrontierCode v1.1 Extended. Grok 61.3%, Fable 5 63.6%, GPT-5.6 Sol 60.6%, Grok 4.5 56.6%.
APEX-Agents. Grok 57.5%, Fable 5 59.2%, GPT-5.6 Sol 56.7%, Grok 4.5 47.1%.
Terminal-Bench v3.0. Grok 26%, GPT-5.6 Sol 34.6%, Fable 5 34.1%, Grok 4.5 15.7%.
APEX-SWE. Grok 56.4%, Fable 5 58.8%, Grok 4.5 53.6%, GPT-5.6 Sol not reported.
AA-Briefcase. Grok 1577, Fable 5 1574, GPT-5.6 Sol 1502, Grok 4.5 1313.
Harvey LAB (Vals). Grok 15.8%, Grok 4.5 12.9%, Fable 5 11.3%, GPT-5.6 Sol 2.5%.
Three wins out of ten. Claude Fable 5 Max takes five, GPT-5.6 Sol Max takes two.
A vendor could have published four rows. This one published ten and bolded the losses.
Then the sentence
Artificial Analysis, an independent evaluator, published its own reading the same day:

SpaceXAI's Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index, joining the frontier in line with GPT-5.6 Sol, with standout agentic performance at lower cost.
Measured, specific, and consistent with xAI's own table.
Elon Musk posted a cost comparison from the same organisation with:
Grok 4.6 is objectively #1 when considering intelligence, speed & cost.

Nineteen hours later, an account with 477,000 views on that single post:
GPT 5.6 SOL MAX BTFO'D BY GROK 4.6 HIGH

The benchmark that post screenshotted was GDPVal-AA, one of the three Grok actually wins, by twelve Elo points over Fable 5. Twelve Elo is roughly a 52 per cent expected win rate. A coin flip with a lean.
Where the word breaks
The cost case is real and it is the strongest thing xAI has. On Artificial Analysis's blended measure of running its own index, Grok 4.6 high comes in at $0.84 against GPT-5.6 Sol max at $1.23 and Claude Opus 5 max at $2.34. Roughly a third of the price of the top scorer for two index points less.
Note that this is not the rate card. xAI's published pricing is $2 per million input tokens and $6 per million output, with a fast variant at double. The $0.84 figure is what it cost Artificial Analysis to run its evaluation suite, which depends on how many tokens each model spends thinking. Two different numbers measuring two different things, and they get quoted interchangeably.

Musk's claim is compound: intelligence, speed and cost. On that combination Grok has a serious argument, and anyone buying tokens at volume should be looking hard at it.
The word that fails is objectively. Combining three variables into one ranking requires weighting them, and a weighting is a choice. Weight cost and Grok wins. Weight capability and Opus 5 wins. Weight coding and Fable 5 wins five rows out of ten on xAI's own table. There is no neutral vantage point that settles it, and describing the answer as objective hides the choice instead of arguing for it.
One adverb turns a value proposition into a fact.
The two rows nobody quoted
Terminal-Bench v3.0: Grok 26%, GPT-5.6 Sol 34.6%.
Harvey LAB: Grok 15.8%, GPT-5.6 Sol 2.5%.
Six times better on a legal analysis benchmark, and beaten by a third on a terminal task benchmark. Both from the same table, both published by xAI.
A spread that wide across two evaluations tells you something composites cannot: individual benchmarks measure narrow things, and a model tuned toward one shape of task will look transformative on one row and mediocre on the next. Terminal-Bench also runs two live versions that disagree completely, with xAI reporting 26% on v3.0 while Artificial Analysis reports 88.4% on v2.1. Both figures are true. A Terminal-Bench score quoted without a version number is not telling you anything at all.
Neither row appeared in any coverage I saw. They are not useful to anybody's argument.
The chain
The transformation, written out with sources and timestamps:
xAI's own table. Three wins from ten, competitors' wins bolded, full field published.
Artificial Analysis. Joins the frontier, in line with GPT-5.6 Sol, standout agentic performance at lower cost.
The CEO. Objectively number one.
The retelling. BTFO'D, 477,000 views.
Four steps, one day, no fabricated numbers anywhere in it.
What makes this worth documenting is where the distortion is not. It is not in the vendor's data, which is unusually candid. It is not in the independent evaluator, which was precise. It enters at the third step and hardens at the fourth, and by then the table that would settle it is four scrolls up a page nobody opened.

What this is not
Grok 4.6 is a good model. Three first places against two of the strongest systems available, a five-point jump on the composite index over its own predecessor, and a genuine cost advantage. Dismissing it because of how it was sold would be the mirror image of the error being described here.
It is also not a case of anyone lying. Selecting five tabs from your own ten-row table is normal launch practice, and the full table sits directly below the tabs on the same page. The CEO's sentence is a compound claim with an unstated weighting, which is a rhetorical move rather than a false statement.
The finding is narrower and, I think, more useful. A company published an honest set of numbers. Within twenty-four hours those numbers had become something the company itself had not claimed, through two intermediate steps neither of which invented anything.
Precision does not survive contact with distribution. That is a fact about how information moves, not about who is dishonest, and it is the reason a table that beats you fairly in seven rows out of ten is worth reading before you accept anyone's summary of it, including this one.
Method and limitations
I ran no benchmarks. Every figure is read from xAI's published launch page and model card, from Artificial Analysis's published index and cost chart, and from the two X posts cited. Where a number appears here, it is what the publisher reports.
Effort settings are not matched across the table, and this matters. Artificial Analysis tested Grok 4.6 at high and Claude Opus 5 at max. xAI reports Grok at high without publishing the trial detail Anthropic normally gives, and Anthropic's system-card figures typically use adaptive thinking at max effort averaged over five trials. These are not equal-compute contests and no clean winner row survives that caveat. It cuts in xAI's favour as often as against.
xAI's footnote states third-party scores are the best of self-reported or publicly available results. I have not checked each of the thirty competitor cells against its original source. That is a further piece of work and it is not done here.
The model card was published on 12 August 2026, revision 2026-08-12, and pages of this kind get edited. Figures here reflect the page as at 13 August 2026.
Disclosure: I use Claude daily in my research, including on this piece. Anthropic models come out ahead in five of the ten rows above. The numbers are xAI's, not mine, and readers can check them at the link below.

Sources
xAI, Introducing Grok 4.6, 12 August 2026: https://x.ai/news/grok-4-6
xAI, Grok 4.6 model card, revision 2026-08-12: https://media.x.ai/v1/website/card-4p6-4cd2dc57.pdf
Artificial Analysis Intelligence Index v4.1.1
@ArtificialAnlys on X, 12 August 2026
@elonmusk on X, 12 August 2026
@ns123abc on X, 12 August 2026, 19:20
Clayton Bax
Published under ONYX Digital Intelligence Following the #OnyxAudit methodology.
- X: @onyxaudit
- Email: onyxdigitalintelligence85@protonmail.com
- https://github.com/Baximus855
- @Onyx_Digital@mastodon.social
"Adjacent to true is not true."
Truth has no flag nor favour, only a standard. And it's heavy