Field intelligence for AI-first professionalsVol. II · Nº 61 · Thursday, September 24, 2026
← Phantom Notes
Nº 58AI ModelsSeptember 24, 202612 min read

Grok 4.7 Took Five Delays to Land in 21st Place. It Is Not in the Grok App Yet.

xAI, now SpaceXAI, shipped Grok 4.7 on Monday, September 21, 2026, after a summer of slipping dates on Elon Musk's feed. Artificial Analysis scores it 46 on the current v4.3.2 Intelligence Index, 21st of 210, and measured about 81,000 output tokens per task at a slow 40 tokens per second on its default workload. The price is unchanged at $2 in / $6 out. The consumer Grok app is not, per the model card, one of the places you can use it yet.

By T.W. Ghost


On Monday, September 21, 2026, xAI, now doing business as SpaceXAI, shipped Grok 4.7. The launch post from the @SpaceXAI handle that replaced @xai opens: "Grok 4.7 is here. It's a notable improvement over Grok 4.6 at the same price and speed."

It arrived 31 days after the first date Elon Musk implied for it, five walk-backs later, and one day before Anthropic shipped Claude Opus 5.5 and took the top of the independent leaderboard, a story we cover separately. When we last wrote about xAI's engine, Grok 4.6 in August, it had just reached the intelligence frontier. Since then the frontier moved, Artificial Analysis rebuilt its index, and SpaceX bought Cursor, which now ships this model to every plan tier. We read the 30-page model card, the docs, and Artificial Analysis's benchmarking run. Here is what survived.


Five Dates, One Ship

Every entry below is a dated Musk post on X, verbatim where quoted.

DateWhat Musk wroteImplied ship
Jul 24"Grok 4.6 in 2 weeks and Grok 4.7 in 4 weeks"~Aug 21
Jul 28"Grok 4.7 will be the 2.1T model released a few weeks later. This will be better than 4.6 in every way, except slightly slower to serve, albeit with even better token efficiency"Late Aug
Aug 12"Grok 4.7 is significantly better than 4.6 and should be ready in 3 to 4 weeks. Initial training is complete and now we're adding a massive amount of SpaceX company data in supplemental training. This will be something special."Sep 2 to 9
Sep 1 (US)"Grok 4.7 comes out in 10 days"Sep 11 to 12
Sep 11"Grok 4.7 needs a few more days to cook." He added that RL might have penalized response length too much, so the model still gave up on hard tasks too early and was not yet rigorous enough in checking its work.Unspecified
Sep 21Shipped

The September 11 post is the useful one, because it names the fix: the RL run had been pushing response length down. As the token bill below shows, the correction swung hard the other way. The September 1 thread also predicted that "Grok 4.7 will exceed all current models. That said, Anthropic is a great company and will probably release improved models soon..." The second sentence aged better than the first.


What xAI States, and What Only Musk States

Parameter count. The 2.1 trillion figure in nearly every headline comes from Musk's July 28 post. It does not appear in xAI's announcement, the API docs, or the model card, which says only that Grok 4.7 uses "a new, larger base model" than Grok 4.6. The "up 40% from 1.5T" arithmetic is Musk's number divided by Musk's other number. We print it as Musk-stated, unconfirmed by xAI.

SpaceX data. Musk's August 12 post about "a massive amount of SpaceX company data in supplemental training" is verified as a Musk statement. The Starlink telemetry, manufacturing records, and failure-log detail appears in Decrypt and AndroidHeadlines without a citation; neither xAI nor Musk has said it, and the words SpaceX and Starlink do not appear in the model card. What the card does disclose, in a footnote, is a different corpus: "Grok 4.7 received supplemental training on anonymized Cursor workflow data to improve coding and agentic performance." The rocket story is Musk's.

Grok 4.7
AnnouncedSeptember 21, 2026 (Monday)
Model IDgrok-4.7
API price (under 200K prompt tokens)$2.00 in / $0.50 cached / $6.00 out per 1M
API price (200K prompt tokens and up)$4.00 in / $1.00 cached / $12.00 out per 1M, on the whole request
Context500,000 tokens; "No text output limit"
ModalitiesText and image in (20 MiB image cap); text out
Knowledge cutoffMay 2026 (API docs) / June 2026 pretraining cutoff (model card)
Effort levelslow, medium, high (default), xhigh
AA Intelligence Index v4.3.246 at xhigh and at high; #21 of 210
AA output speed40.4 tok/s (xhigh), 52.4 tok/s (high)
Consumer Grok app"at a later date" (model card)

The Benchmarks, Sorted by Who Reported Them

House rules: Artificial Analysis runs its own harness, so its numbers are stated as fact. The second half of this section is xAI's self-reported model card and blog, labeled as such.

One correction first, and it applies to our own August coverage. Artificial Analysis re-versioned its Intelligence Index three times this month: v4.2 on September 4, v4.3 on September 7 (Terminal-Bench 2.1 replaced by 4.0, tau3-Banking replaced by AutomationBench-AA), and v4.3.2 on September 19. Grok 4.6's August score of 61, which we printed at the time, was on the retired scale. On the current v4.3.2 index it reads 44. Opus 5 went from 63 to 51 in the same rebuild.

Independent: Artificial Analysis

On the current v4.3.2 index, Grok 4.7 scores 46, identical at xhigh and high effort, 21st of 210 models. AA's launch-day post: "Grok 4.7 scores 46 on the Artificial Analysis Intelligence Index to bring SpaceXAI into the top 4 AI labs".

Model (effort)AA Intelligence Index v4.3.2
Claude Opus 5.5 (max)58
Claude Fable 5.1 (max)53
GPT-6 Astra (max)53
Claude Opus 5 (max)51
GPT-6 Sol (max)48
GPT-5.6 Sol (max)47
Grok 4.7 (xhigh)46
Grok 4.6 (high)44
Gemini 3.8 Flash (high)41

Two points over its predecessor, twelve behind Opus 5.5. The component deltas versus Grok 4.6 (high) explain the small net: Terminal-Bench 4.0 up 4.5 points and GDP.pdf up 3.0, but AA-LCR long-context down 3.7 and AutomationBench-AA down 1.1. Hallucination improved to 29% from 34% on AA-Omniscience, on flat accuracy (47% versus 48%).

The agentic knowledge-work evals are where Grok 4.7 looks best, and they are still AA's own runs:

Model (effort)GDPval-AA v2.1 EloAA-Briefcase v1.1 Elo
Claude Opus 5.5 (max)18461822
Claude Fable 5.1 (max)17351678
Claude Opus 5 (max)17081673
Grok 4.7 (xhigh)16951657
Grok 4.6 (xhigh)16321555
GPT-5.6 Sol (max)15881487
GPT-6 Astra (max)15421569

Sixth on GDPval-AA v2.1, ahead of every OpenAI model, and just behind Opus 5 and Fable 5.1 on AA-Briefcase. On AA's Coding Agent Index, Grok 4.7 (xhigh) inside Grok Build scores 56, up from 47 for Grok 4.6 (xhigh).

Vendor: xAI's model card and blog

These are xAI's self-reported numbers. Grok 4.7 runs in the Grok Build harness; peer models "use their respective provider harnesses". Effort settings are as xAI printed them.

BenchmarkGrok 4.7 (xhigh)Grok 4.6 (high)GPT-5.6 Sol (max)Fable 5.1 (max)
DeepSWE v1.171.0%*65.2%72.7%70.0%
Terminal-Bench 4.038.0%20.3%37.3%57.9%
CursorBench 4.046.3%40.4%41.7%51.8%
Harvey Legal Agent Benchmark19.6%15.8%2.5%6.7%
EEBench (Atopile)64.0% blog / 66.0% card53.0%39.4%56.4%

*xAI footnotes DeepSWE as run at high effort, not xhigh.

From the model card only: FrontierSWE V2 29.0% at xhigh against Fable 5.1 (max) at 56.3%. SWE-Marathon v1.1 46.0% at high, against Opus 5 (max) at 50.0%. CADGenBench 44.4% at high, leading Grok 4.6 (40.9%), GPT-5.6 Sol (37.1%), and Opus 5 (36.6%). CVE-Bench 36.6% at xhigh, below Grok 4.6's 39.8%.

Four honest observations the launch post does not lead with:

  • EEBench is first or second depending on which xAI document you open. The blog table prints 64.0% with no GPT-6 Astra column, so Grok 4.7 is first, and several outlets wrote that it tops every model. The model card prints 66.0% at xhigh and adds GPT-6 Astra (max) at 69.3%, which puts Grok 4.7 second. We print the card.
  • Terminal-Bench 4.0 has two vendor readings too. The blog served us 37.6% twice; the model card says 38.0%. We print the card. Either way it is 19.9 points behind Fable 5.1 on xAI's own table.
  • The "fewer output tokens" claim did not survive measurement. The card describes Grok 4.7 "reaching results with fewer steps and fewer output tokens than other frontier models". AA measured about 81,000 output tokens per Intelligence Index task, against 36,000 for Grok 4.6 (high) and 27,000 for GPT-6 Astra (max).
  • Several dual-use science scores went down, by design. VCT 63.0% (Grok 4.6: 67.4%), Biosecurity VCT 41.5% (47.8%), WMDP-Bio 88.1% (90.0%), LAB-Bench 76.8% (80.7%), ProtocolQA 70.4% (79.6%), BixBench 88.4% (93.8%), attributed to "safer RL environments and better selectivity of training data". If your work touches biology or security tooling, read those rows before you migrate.

And one absence: no image or multimodal benchmark appears anywhere in the 30 pages.


Cheap Per Token, Expensive Per Task

The price is unchanged from Grok 4.6: $2.00 in / $6.00 out per million tokens, cache hits at $0.50. AA's blended rate (7:2:1 cache:input:output) is $1.35 per million for both models. Against Opus 5.5 at $4/$20 and Fable 5.1 at $10/$50, it is the cheapest frontier-adjacent list price on the board.

The per-token price is not what you pay for. AA measured Grok 4.7 at xhigh using about 81,000 output tokens per Intelligence Index task, at 40.4 tokens per second, which AA labels "notably slow" and "very verbose", about 7.1 minutes per task. AA's benchmarking article separately measured about 188 tokens per second on long prompts, so the 40.4 figure is the default-workload reading, not a ceiling. The result is a cost per index task of $3.74 at xhigh and $2.73 at high. Grok 4.6 (high) cost $1.86 on the same run. GPT-6 Astra (max) costs $3.26, and Opus 5.5 $1.82 at high or $3.46 at xhigh. AA's sentence on this is the one to remember: "A model charging less per token can still be more expensive on a finished workload if it needs substantially more reasoning tokens."

Theo Browne reported that every Claude Opus 5 run he tried cost less than a quarter of Grok 4.7's, and described the result as last-generation performance at a current-generation price (via BigGo, September 22). VentureBeat's launch-day headline: "high token consumption threatens real-world ROI".

Two more lines of fine print:

The long-context tier is a cliff, not a surcharge. At or above 200,000 prompt tokens, every token in the request bills at $4 / $1 cached / $12, not just the overflow.

"Same speed" is xAI's claim, not AA's. AA measured Grok 4.6 (high) at 60.9 tokens per second and Grok 4.7 (high) at 52.4, closer to Musk's July 28 "slightly slower to serve" than to the launch post. If you need speed, Grok 4.7 Fast exists only inside Cursor and Grok Build, at "twice the output speed at twice the price": $4 / $1 / $12 under 200K. We could not confirm OpenRouter's rate (one outlet prints $1.60 / $4.80, another says pass-through), so we do not print one.


Where You Can Run It (and Where You Cannot)

Live on launch day, per the model card: the SpaceXAI API at console.x.ai; Grok Build, where 4.7 is the default with a free entry point; Cursor, for "all users, on every plan tier"; the Grok add-ins for Word, PowerPoint, and Excel, also as default; and the gateways OpenRouter, Vercel, Cloudflare, Snowflake, and Databricks Mosaic. GitHub Copilot added it the same day for paid plans, with a gradual rollout.

Not live: the Grok app. Decrypt and AndroidHeadlines both said Grok 4.7 was available in the consumer app. The model card says otherwise, verbatim: "SpaceXAI plans to add Grok 4.7 to its consumer surfaces (web, mobile apps, and Grok-in-X on the X platform) at a later date." As of September 24 we found no rollout announcement and no SuperGrok or X Premium tier statement from xAI. If you pay for SuperGrok, you are on 4.6 until told otherwise.

On Grok Bot, the always-on agent product from August: xAI says it "also trained Grok 4.7 to natively understand the Grok Bot harness", but does not say whether Grok Bot now runs 4.7 by default.


The Roadmap, in Musk's Words

On September 14, a week before launch, Musk set expectations lower than the launch post did: "Grok 4.7 should be roughly on par with Opus 5.0, not 5.1. Better in some ways, worse in others. We need to fix multimodal performance. Grok 4.8 will be a noticeable improvement. Grok 4.9 is probably Astra/Fable class. Grok 5 maybe better than anything. We shall see."

One factual note: there is no Claude Opus 5.1; his "5.1" most plausibly means Fable 5.1. On the AA index, "on par with Opus 5.0" is generous by five points (46 versus 51), though the GDPval-AA gap is narrower (1695 versus 1708). The same weekend he wrote that "Grok 4.8, which is a 2.5T model trained with our new C++ software stack, will finish training this week and start RL". No dates were given. Given the table at the top of this post, we would not pencil any in.


Verified vs Unconfirmed: The Scorecard

ClaimVerdict
Released Sep 21, 2026 at $2 / $0.50 cached / $6 per 1M, 500K context, doubling at 200K prompt tokensVerified (xAI docs and model card)
"Same price and speed" as Grok 4.6Price verified; speed contradicted (AA: 52.4 vs 60.9 tok/s at high)
2.1 trillion parameters, up from 1.5TUnverified (Musk's posts only; not in any xAI document)
Supplemental training on SpaceX company dataVerified (Musk, Aug 12); the Starlink / manufacturing / failure-log detail is Unverified (press only)
Live in the Grok consumer appContradicted (model card: consumer surfaces "at a later date")
AA Intelligence Index v4.3.2 score 46, #21; GDPval-AA v2.1 1695; AA-Briefcase v1.1 1657Verified (Artificial Analysis)
"Fewer steps and fewer output tokens than other frontier models"Contradicted (AA: ~81K output tokens per task vs 27K for GPT-6 Astra)
First on EEBenchContradicted (blog 64.0% omits Astra; model card 66.0% is second to GPT-6 Astra 69.3%)
Knowledge cutoffTwo vendor answers (May 2026 in API docs, June 2026 pretraining in model card)

Who Should Use It

Use Grok 4.7 if you live in Cursor or Grok Build and your work is agentic knowledge work: documents, spreadsheets, legal and engineering research, the GDPval-AA and AA-Briefcase territory where it sits ahead of every OpenAI model at a list price a fraction of Claude's. Run it at high, not xhigh: same index score, 27% cheaper per task.

Skip it for anything billed by the task rather than the token, because at $3.74 per index task it costs more than Opus 5.5 at high or GPT-6 Astra at max while scoring below both. Skip it for image-heavy work (no multimodal benchmarks, and Musk says the fix is pending), for the hardest software engineering (29.0% on FrontierSWE V2 against Fable 5.1's 56.3%, on xAI's own card), and for anything you planned to do in the Grok app, because it is not there.

The frontier is elsewhere: Claude Opus 5.5 at 58 on the current index, with Claude Fable 5.1 and GPT-6 Astra tied at 53 behind it, and Gemini 3.7 Flash and its successor holding the budget tier. Our model comparison has the current v4.3.2 standings side by side.


Sources


*Not sure whether a cheap-per-token, expensive-per-task model fits how you actually work? Take the free 2-minute quiz and get matched. Then read what shipped the next day and why it reset the leaderboard: GPT-6 Sol and Luna vs Claude Opus 5.5.*

Which model should you be using?

Three minutes, twelve questions, one defensible answer.

Take the quiz →