Grok 4.7 is out. SpaceXAI published Introducing Grok 4.7 on September 21, 2026, calling it "a notable improvement over Grok 4.6 at the same price and speed." That is a narrower claim than the pre-launch marketing around the 2.1-trillion-parameter model suggested — and the shipped eval table backs the narrower version, not the louder one.
Grok 4.7 works longer on difficult tasks, checks its own work more carefully, and comes with what SpaceXAI calls its best-calibrated safeguards to date. It is priced identically to Grok 4.6: $2 per million input tokens, $6 per million output tokens, with a fast variant at twice the speed and twice the price.
TL;DR — what people are actually asking
| Question | Direct answer |
|---|---|
| Is Grok 4.7 out? | Yes. SpaceXAI announced it September 21, 2026 |
| Same price as 4.6? | Yes — $2/M in, $6/M out. Fast variant is 2x ($4/$12) |
| Where do I run it? | Cursor, Grok Build, SpaceXAI API, third-party harnesses and routers |
| Does it beat 4.6? | Yes on every published row vs Grok 4.6 High |
| Does it beat Fable 5.1 / Sol? | No sweep. Fable 5.1 Max leads CursorBench, AA-Briefcase, Terminal-Bench |
| What does Grok 4.7 win? | EEBench, Harvey Legal Agent Benchmark |
| What is new under the hood? | Larger base model, longer RL run on hard/long-horizon tasks, better self-verification |
| Is this Grok Bot? | No. Bot is the persistent-agent product; 4.7 is the model it now natively understands |
| Any safety numbers? | HackerBench v0.3: 3.3% risky prompts allowed through. LatchBio biosafety: 62.4% |
What SpaceXAI actually shipped
Grok 4.7 is SpaceXAI's most capable model yet for coding and knowledge work, per the announcement, built on three changes: a larger base model than 4.6, a longer reinforcement learning run on a harder task mix weighted toward problems that take many hours to complete, and better self-verification — checking its own work and managing longer context on multi-step trajectories.
SpaceXAI also says 4.7 was trained to natively understand the Grok Bot harness, making it better at conversational tasks and general knowledge work — not just agentic coding. That is a continuation of the pattern set by Grok 4.6's harness-aware training, extended now to the persistent-agent product rather than just Cursor and Grok Build.
The launch post's own framing is worth quoting directly: "Grok 4.7 works longer on difficult tasks, checks its own work more carefully, and comes with our best-calibrated safeguards to date." That is a harness-and-reliability pitch, not a raw-intelligence pitch — and the eval table underneath it matches that framing more than it matches "beats everything," which nothing on the sheet actually claims this time.
Official evals: Grok 4.7 xHigh vs 4.6 vs Sol vs Fable 5.1

Table from SpaceXAI's September 21, 2026 Grok 4.7 announcement. An asterisk on Grok 4.7's DeepSWE score marks a high-effort run rather than xHigh.
| Eval | Grok 4.7 xHigh | Grok 4.6 High | GPT-5.6 Sol Max | Fable 5.1 Max |
|---|---|---|---|---|
| Input price ($/M tokens) | $2 | $2 | $4 | $10 |
| Output price ($/M tokens) | $6 | $6 | $20 | $50 |
| CursorBench 4.0 | 46.3% | 40.4% | 41.7% | 51.8% |
| DeepSWE v1.1 | 71.0%* | 65.2% | 72.7% | 70.0% |
| EEBench | 64.0% | 53.0% | 39.4% | 56.4% |
| AA-Briefcase v1.1 | 1,657 | 1,546 | 1,487 | 1,678 |
| Terminal-Bench 4.0 | 38.0% | 20.3% | 37.3% | 57.9% |
| Harvey Legal Agent Benchmark | 19.6% | 15.8% | 2.5% | 6.7% |
| HealthBench Professional | 56.7% | 48.5% | 60.5% | 62.1% |
Who actually leads which row
Grok 4.7 xHigh leads EEBench (64.0% — a wide 7-to-25-point gap over the field) and Harvey Legal Agent Benchmark (19.6% vs Fable's 6.7% and Sol's 2.5%). Electrical engineering and legal-agent work are the two categories where this card is unambiguous.
Fable 5.1 Max leads CursorBench 4.0 (51.8% vs Grok's 46.3%), AA-Briefcase v1.1 (1,678 vs Grok's 1,657), Terminal-Bench 4.0 (57.9% — a full 19.9-point gap over Grok's 38.0%), and HealthBench Professional (62.1%). Terminal-Bench 4.0 is Grok 4.7's clearest weak row on this sheet, similar to the gap Grok 4.6 showed on Terminal-Bench v3.0 a month earlier — this looks like a recurring category, not a one-off.
GPT-5.6 Sol Max leads DeepSWE v1.1 at high effort (72.7%, just ahead of Grok's 71.0% xHigh run) and HealthBench Professional (60.5%, second to Fable). Sol is also the only model priced above Grok on this table at $4/$20 versus Grok's $2/$6, while still trailing Grok on five of the seven rows — the closest thing to a clean price-performance argument for Grok 4.7 on this specific sheet.
Versus Grok 4.6 High, the upgrade is consistent across every row: +5.9 CursorBench, +5.8 DeepSWE, +11.0 EEBench, +111 AA-Briefcase, +17.7 Terminal-Bench, +3.8 Harvey, +8.2 HealthBench. Same price, same speed tier. That is the clean, uncontested comparison on this card — the multi-model comparison is where the picture gets mixed.
One caveat that applies to every cell here: SpaceXAI's own chart credits itself with best-of-self-reported figures for its own model and cites competitor scores from public results, not a single lab re-running all four in one harness. For how to weigh that kind of vendor table generally, see the AI benchmarks guide.
GDPval, AA Briefcase, and EEBench: knowledge-work Elo
SpaceXAI also published an Elo-style comparison on GDPval, AA-Briefcase, and EEBench — the "professional knowledge work" cluster where AI is asked to do tasks done by lawyers, nurses, and financial analysts. Grok 4.7 improves on Grok 4.6 on both GDPval and AA-Briefcase, and lands comparable to other frontier models overall:
| Model | GDPval Elo |
|---|---|
| Fable 5.1 (max) | 1735 |
| Grok 4.7 (xhigh) | 1695 |
| Grok 4.6 (high) | 1605 |
| GPT-6 Astra (max) | 1542 |
Grok 4.7 sits between Fable 5.1 and its own predecessor on this composite — a real gain over 4.6, not yet a lead over Fable 5.1, and ahead of GPT-6 Astra on this specific chart despite Astra's later ship date and larger claimed scope.
The new safeguard stack
SpaceXAI says Grok 4.7 shipped with an entirely new safeguard stack, calling it the strongest model it has tested on refusals and jailbreak resistance. Two numbers back that claim:
- HackerBench v0.3 (risky and malicious cyber tasks): Grok 4.7 allowed only 3.3% of risky dual-use prompts through, which SpaceXAI frames as the highest safety score on this bench, while — per the announcement — rarely blocking legitimate security work.
- LatchBio's biosafety benchmark: Grok 4.7 topped it at 62.4%.
SpaceXAI also says it has begun giving select cybersecurity partners invite-only access to Grok 4.7's red-team capabilities for defense research. This is the second consecutive Grok release with an independent biosecurity number attached — Grok 4.6's own LatchBio evaluation used a different pair of benchmarks (BioSecBench-Refusal and BioSecBench-Surveillance) roughly three weeks earlier, so the 62.4% figure here is not directly comparable to that post's 59.2%/64.8% split — different suite, same general claim of leading the field on dual-use safety.
As with the 4.6 launch, none of these figures come with an independently reproduced result yet. Treat "strongest model we've tested" as SpaceXAI's own characterization until a third party runs HackerBench v0.3 or the LatchBio suite against Grok 4.7 directly.
Pricing and availability
| Detail | Number |
|---|---|
| Input | $2 / million tokens |
| Output | $6 / million tokens |
| Fast variant | 2x speed, 2x price ($4/$6 base → 2x output speed at 2x price) |
| vs Grok 4.6 list | Unchanged |
| Surfaces | Cursor, Grok Build, SpaceXAI API, third-party harnesses, model routers, cloud platforms |
The economic pitch is the same one SpaceXAI has run since Grok 4.5: hold price and speed flat, ship the eval gains on top. No week-one usage promo was announced this time, unlike 4.6's 2x included-usage week in Cursor and Grok Build.
Install path from the announcement, unchanged from prior releases:
curl -fsSL https://x.ai/cli/install.sh | bash
For the harness layer this model runs inside, see what an agent harness is and Grok Build's open-sourced tree.
What this does and doesn't settle about the pre-launch claims
The pre-launch reporting on Grok 4.7 centered on a 2.1-trillion-parameter figure and SpaceX-engineering-data training angle, plus a Musk claim that it would "surpass all existing AI models." None of the parameter count, the SpaceX-data mechanism, or the "surpass everything" framing appears in SpaceXAI's own launch post — the official announcement talks about a larger base model and longer RL, without a parameter count, and the eval table shows Grok 4.7 losing four of seven rows to Fable 5.1. That is the same pattern Grok 4.6's launch showed against its own pre-announcement hype: real, measurable gains over the prior version, at the same price, without a clean sweep of the field.
Musk's own mid-September comments — made while Grok 4.8 was already in training — had already tempered expectations for 4.7 relative to Opus 5.0, which lines up with what shipped here more than the earlier "surpass every model" framing did.
Should you switch?
If you're on Grok 4.6: switch. Every published row improves, at identical price and speed, with a materially better safety story on cyber and biosafety benchmarks.
If you're on Fable 5.1 for coding agents: don't switch on this table alone. Fable 5.1 Max still leads CursorBench 4.0, Terminal-Bench 4.0 by a wide margin, and AA-Briefcase. Grok 4.7's price advantage ($2/$6 vs Fable's $10/$50) is real and large, so the switch decision is a cost-per-completed-task question, not a raw-score one — run your own loop against both before deciding.
If you care about electrical engineering or legal-agent work: EEBench and Harvey Legal Agent Benchmark are the two rows that justify a serious look at Grok 4.7 specifically, with double-digit-point leads over every other model on the sheet.
Limitations worth stating plainly:
- Self-reported vs. public competitor scores, same caveat as every prior Grok launch table.
- No independent reproduction yet — this is launch-day coverage.
- Terminal-Bench 4.0 is a real, recurring gap, echoing Grok 4.6's Terminal-Bench v3.0 shortfall.
- The 2.1T-parameter and SpaceX-data claims from pre-launch reporting were not repeated in the official post — treat them as unconfirmed until SpaceXAI publishes a model card with a parameter count.
- HackerBench v0.3 and LatchBio numbers are SpaceXAI's own, not third-party reproduced.
Related on explainx.ai
Update — September 22, 2026: Grok 4.7 has now officially launched with a full eval table — see this post for the shipped numbers, safeguard stack, and pricing, superseding the pre-launch parameter-count reporting below.
- Grok 4.7: 2.1T-Parameter Pre-Launch Claims (Sept 3) — what Musk claimed before launch, and how it compares to what actually shipped
- Grok 4.8: 2.5T Parameters and a New C++ Training Stack — the next model already in training as of mid-September
- Grok 4.6 Launch: Official Evals, Same $2/$6, Cursor Access — the prior release this one directly upgrades
- Grok 4.6's LatchBio Biosecurity Evaluation — the earlier BioSecBench numbers to compare against this post's 62.4% LatchBio figure
- Grok 4.6/4.7 Release Timeline: Musk's July Announcement
- GPT-6 Astra: Launch, Benchmarks, and Pricing
- Claude Fable 5.1 / Mythos 5.1: Launch, Benchmarks, and Pricing
- Grok Build open source — Apache 2.0 harness
- What is an agent harness?
- Loop engineering for coding agents
- AI benchmarks complete guide
Primary source: SpaceXAI, Introducing Grok 4.7 (September 21, 2026)
Evals, pricing, and availability are accurate as of September 22, 2026 and reflect SpaceXAI's September 21 announcement. Competitor scores on that table are the best of self-reported or publicly available results, not a single-harness re-run by explainx.ai. Follow @explainx_ai for updates.
