Xai Grok 4 7 · AI Security

Grok 4.7 Matches Claude Fable 5.1 Max — on SpaceXAI's Own Scorecard

Head-to-head data poster from SpaceXAI's Grok 4.7 launch table: Grok 4.7 xHigh against Claude Fable 5.1 Max on DeepSWE v1.1 71.0 to 70.0, EEBench 64.0 to 56.4, Harvey Legal 19.6 to 6.7, and Terminal-Bench 38.0 to 57.9, stamped VENDOR SCORED.
AK

Threat intelligence editor · Updated Sep 22, 2026, 5:27 AM EDT

SpaceXAI's launch table shows Grok 4.7 edging Claude Fable 5.1 Max on DeepSWE at a fraction of the price. On an independent harness, the gap is 29 points.

SpaceXAI released Grok 4.7 on Monday, 21 September 2026, and the launch scorecard makes a claim the company has not been able to make before: on its own published benchmark table, Grok 4.7 goes toe to toe with Claude Fable 5.1 Max. It edges Fable on DeepSWE v1.1, beats it comfortably on two more, and sits within 21 Elo points on a fourth — while costing, by most estimates, something like a fifth to an eighth as much per token.

That is a real result, and it deserves to be read carefully. Every number in it comes from the scorecard SpaceXAI wrote. Where an independent harness exists for the same named benchmark, the picture changes sharply.

Parity on the vendor's own table

SpaceXAI's launch comparison runs Grok 4.7 at its xHigh configuration against Claude Fable 5.1 and GPT-5.6 Sol at their Max tiers:

BenchmarkGrok 4.7 (xHigh)Claude Fable 5.1 MaxGPT-5.6 Sol Max
DeepSWE v1.171.0%70.0%72.7%
EEBench64.0%56.4%39.4%
Harvey Legal Agent19.6%6.7%2.5%
AA-Briefcase v1.11,6571,6781,487
CursorBench 4.046.3%51.8%41.7%
HealthBench Professional56.7%62.1%60.5%
Terminal-Bench 4.038.0%57.9%37.3%

Grok 4.7 takes three of the seven outright and loses AA-Briefcase by 21 Elo — close enough to call a tie on an agentic knowledge-work benchmark that SpaceXAI has historically lost badly. Against GPT-5.6 Sol Max it wins five of seven. Priced at $2 per million input tokens and $6 per million output, that is the cheapest model to have credibly reached this table.

Two caveats travel with it. The first is that CursorBench 4.0 — one of the two benchmarks SpaceXAI leads its table with — is built around Cursor, a company SpaceXAI acquired in August 2026, four months after SpaceX completed its all-stock acquisition of xAI and roughly two months after the combined entity took the SpaceXAI name. A first-party benchmark, run in a first-party harness, inside a first-party product.

The second is Terminal-Bench.

Three harnesses, three Terminal-Bench scores

Terminal-Bench 4.0 is the one benchmark on this table where an independent, standardized measurement exists. Grok 4.7 has three published scores on it, depending entirely on who ran it and how:

  • 26% — Artificial Analysis, standardized harness, the one applied identically across every model in its Intelligence Index
  • 33% — Artificial Analysis, run through Grok Build, SpaceXAI's own first-party coding agent, up from 18% for Grok 4.6
  • 38.0% — SpaceXAI's own launch table, at xHigh effort, against 20.3% for Grok 4.6

Three Terminal-Bench 4.0 scores for Grok 4.7: 26% on a standardized harness, 33% through Grok Build, and 38.0% self-reported by SpaceXAI, with GPT-6 Astra at 60% and Claude Fable 5.1 at 55% for reference.

One named benchmark, three harnesses: twelve points separate SpaceXAI's own figure from the standardized measurement.

Twelve points separate the vendor's figure from the standardized one on the same named benchmark. Artificial Analysis is explicit that its Grok Build results are held apart from the Intelligence Index precisely because the Index standardizes the harness across models.

The gap that matters is the head-to-head. On SpaceXAI's table, Grok 4.7 trails Fable 5.1 Max on Terminal-Bench by 20 points. On the standardized harness, Grok 4.7's 26% sits against 55% for Claude Fable 5.1 and 60% for GPT-6 Astra — a 29-point deficit, and barely ahead of DeepSeek V4.1 Flash at 27%, a model priced as a budget option.

The broader independent picture is consistent with that. Grok 4.7 scores 46 on the Artificial Analysis Intelligence Index v4.3.2, up two points from Grok 4.6's 44, against 53 for both GPT-6 and Claude Fable 5.1. On the Coding Agent Index, Grok 4.7 paired with Grok Build reaches 56, up nine points from 47 — a genuine gain, and still fourth behind Claude Fable 5.1, GPT-6 Astra and Claude Opus 5. Under that harness: DeepSWE v1.1 at 73% from 65%, SWE-Atlas-QnA at 63% from 58%, Terminal-Bench at 33% from 18%. Hallucination rate improved to 29% from 34%, while measured accuracy stayed flat at 47% against 48%.

Parity on the vendor's table; mid-pack on everyone else's.

What actually shipped

Grok 4.7 is a larger base model than Grok 4.6, trained with a longer reinforcement-learning run weighted toward multi-hour problems and tuned to check its own work more carefully. The context window stays at 500,000 tokens and the knowledge cutoff is May 2026. SpaceXAI has not disclosed a parameter count; a figure of 2.1 trillion circulated in secondary coverage on launch day, but it does not appear in the company's own launch material and should be treated as unconfirmed.

Pricing is unchanged from Grok 4.6: $2 and $6 per million input and output tokens, rising to $4 and $12 beyond 200,000 tokens of context, with cached input at $0.50 and $1.00. A fast variant doubles output speed at double the token price.

Distribution is the part that matters operationally. There is no waitlist. Grok 4.7 went live immediately in the Grok app, on X, in Tesla dashboard voice assistants, in Cursor, in Grok Build, through the xAI API, across third-party model routers and cloud platforms, and in GitHub Copilot for Pro, Pro+, Max, Business and Enterprise subscribers.

The safety claims carry the same problem

Grok 4.7 ships with what SpaceXAI calls an entirely new safeguard stack, and the company describes it as the strongest model it has tested on refusals and jailbreak resistance. The headline figure for security teams: on HackerBench v0.3, an internal benchmark of risky and malicious cyber tasks, Grok 4.7 allowed through just 3.3% of dangerous dual-use prompts while, SpaceXAI says, rarely blocking legitimate security work. It also reports topping LatchBio's biosafety benchmark at 62.4%, and leading on both benign-task utility and dangerous-task refusal across cybersecurity and biological domains — two axes that usually trade against each other. SpaceXAI has additionally begun giving a small number of cybersecurity partners invite-only access to Grok 4.7's red-team capabilities for defensive research.

That is a remarkable turn for a lineage that spent two years as the cautionary example. It is also, right now, entirely unverifiable — and for the same structural reason the benchmark table is.

HackerBench v0.3 is a version-0 benchmark with no public methodology, no published prompt set, and no independent reproduction. A refusal rate is not like a coding score: it is a property of a policy layer the vendor controls, can tune per deployment, and can change without a version bump. "3.3% of risky prompts allowed" describes a system under test on a given day, not a durable property of the weights an enterprise calls through Copilot.

The red-team partner programme compounds this rather than resolving it. Controlled access for defensive research is the responsible way to distribute offensive capability, but a closed programme that produces no public findings also produces no external check on the safety numbers. Defenders are asked to accept the capability claim and the containment claim on the same authority.

The price did not change; the bill did

The unchanged $2/$6 headline is doing a lot of work in the parity story. Grok 4.7 at xHigh consumes roughly 81,000 output tokens per Intelligence Index task against about 38,000 for Grok 4.6, at an average of 7.1 minutes per task.

On the coding-agent side the compounding is sharper. Artificial Analysis measured Grok Build with Grok 4.7 at $8.82 in pay-per-token API cost and 39.2 minutes per task, against $3.57 and 19.5 minutes for Grok 4.6 — about 2.47 times the cost and twice the wall time for nine index points. On CursorBench, xHigh runs $6.01 per task for 46.3% against $4.69 for 43.9% at high effort.

Flat token rates do not mean flat invoices when the model is trained to think longer. The fifth-to-an-eighth price advantage over Fable 5.1 is a per-token figure, and per-token is not the unit anyone is billed in.

What defenders should do with this

Take the parity claim at exactly its stated scope: Grok 4.7 matches Claude Fable 5.1 Max on a table SpaceXAI published, under configurations SpaceXAI chose, on benchmarks that in one case SpaceXAI owns. That is not nothing — three outright wins against a frontier leader is a real result at this price. It is also not the same claim as matching Fable 5.1 Max in your codebase.

Treat the safety numbers the same way you would any vendor's self-assessment of its own control plane: as a statement of intent, not a control. If your organisation has Copilot Business or Enterprise, Cursor, or a model router, Grok 4.7 is already reachable, and the 3.3% figure is not something a risk register can cite. Test refusal behaviour against your own dual-use prompt set, in your own harness, on the deployment you actually call — the Terminal-Bench spread is a direct demonstration that the harness moves the number, by as much as 29 points against the same competitor.

Musk has sketched the next three releases: Grok 4.8 as a meaningful step up, Grok 4.9 in what he called the Astra/Fable class, and Grok 5 as a possible frontier leader. On the evidence of 4.7, the interesting question is not whether the scores climb, but whether anyone outside SpaceXAI gets to check them.