Grok 4.6 is out. SpaceXAI published Introducing Grok 4.6 on August 12, 2026 — a post-training upgrade of Grok 4.5 aimed at long-running agents and more ambitious interactive and visual work, at the same $2 per million input / $6 per million output list price.
Elon Musk's July 25 timeline had put 4.6 "in 2 weeks," roughly August 8. The official model card arrived four days later. Musk's launch-day posts called it "now out" and "smart, fast & amazing bang for buck." The evals underneath that slogan are more mixed than the slogan: Grok 4.6 High ties GPT-5.6 Sol Max at 61 on the Artificial Analysis Intelligence Index, trails Fable 5 Max (62), and does not sweep the board.
This post is the model launch and the official numbers. Grok Bot — the persistent-agent product that signs into your tools — shipped the day before and is a separate story.
TL;DR — what people are actually asking
| Question | Direct answer |
|---|---|
| Is Grok 4.6 out? | Yes. SpaceXAI announced it August 12, 2026 |
| Same price as 4.5? | Yes — $2/M in, $6/M out. Fast variant is 2x ($4/$12) |
| Where do I run it? | Cursor, Grok Build, SpaceXAI API, OpenRouter, Vercel, Cloudflare |
| Week-one promo? | 2x included usage in Grok Build and Cursor for the first week |
| Does it beat 4.5? | On SpaceXAI's table, yes — every published row vs Grok 4.5 High |
| Does it beat Fable 5 / Sol? | No sweep. Ties Sol at 61 AA; Fable leads several coding benches |
| What does Grok win? | GDPVal-AA v2, AA-Briefcase, Harvey LAB (Vals) |
| What does it lose? | CursorBench, FrontierCode, APEX-Agents, APEX-SWE (Fable); DeepSWE and Terminal-Bench v3.0 (Sol) |
| Is this Grok Bot? | No. Bot is the persistent-agent product; 4.6 is the model |
| What about 4.7? | Still ahead — Musk's next cadence claim; see the 4.7 training note |
What SpaceXAI actually shipped
The opening of the announcement is the product thesis, so it is worth quoting:
Today we are releasing Grok 4.6. Grok 4.6 builds on Grok 4.5 with a particular focus on long-running agents and more ambitious interactive and visual work. It stays with complex tasks across many steps, whether researching a topic, analyzing information, working across a codebase, or turning an idea into a polished application or work artifact.
That is a harness-era pitch, not a "smarter chatbot" pitch. The unit of work is a trajectory: research, code, visual first pass, then iterate. SpaceXAI says 4.6 is especially strong at turning a broad product idea into a working first version, and that on longer trajectories it started showing more self-testing and verification — checking its own work before moving on.
Grok 4.6 is available the same day in Cursor and Grok Build. It is also on the SpaceXAI API and the partner routes named in the post: OpenRouter, Vercel, and Cloudflare. SpaceXAI is offering 2x included usage inside Grok Build and Cursor for the first week.
Musk separately claimed Grok 4.6 is state of the art on Databricks' OfficeQA Pro V2 when run with the Genie harness — a document-understanding and data-reasoning eval. That number is not in SpaceXAI's published table. Treat it as a founder claim until Databricks or a third party reproduces it. The numbers that are on the card are below.
Official evals: Grok 4.6 vs 4.5 vs Sol vs Fable

Table from SpaceXAI's August 12, 2026 Grok 4.6 announcement. Best score per row in bold on the original; competitor figures are best of self-reported or publicly available results.
| Eval | Grok 4.6 High | Grok 4.5 High | GPT-5.6 Sol Max | Fable 5 Max |
|---|---|---|---|---|
| AA Intelligence Index | 61 | 56 | 61 | 62 |
| GDPVal-AA v2 | 1753 | 1526 | 1728 | 1741 |
| CursorBench v3.2 | 69.9% | 66.7% | 67.2% | 70.5% |
| DeepSWE v1.1 | 65.9% | 54% | 73% | 70% |
| FrontierCode v1.1 (Extended) | 61.3% | 56.6% | 60.6% | 63.6% |
| APEX-Agents | 57.5% | 47.1% | 56.7% | 59.2% |
| Terminal-Bench v3.0 | 26% | 15.7% | 34.6% | 34.1% |
| APEX-SWE | 56.4% | 53.6% | — | 58.8% |
| AA-Briefcase | 1577 | 1313 | 1502 | 1574 |
| Harvey LAB (Vals) | 15.8% | 12.9% | 2.5% | 11.3% |
SpaceXAI's headline sentence matches the composite, not the row-by-row: "Grok 4.6 achieves frontier intelligence across several agentic coding and knowledge work benchmarks. It matches GPT-5.6 Sol on the Artificial Analysis Intelligence Index, which is a composite score of nine benchmarks."
Who actually leads which row
Grok 4.6 High leads GDPVal-AA v2 (1753 vs Fable 1741 vs Sol 1728), AA-Briefcase (1577 vs Fable 1574), and Harvey LAB (15.8% vs Fable 11.3% vs Sol 2.5%). Those are the knowledge-work and professional-task columns. The Harvey gap versus Sol is the loudest single cell on the sheet.
Fable 5 Max leads AA Intelligence Index (62 vs 61), CursorBench v3.2 (70.5% vs 69.9%), FrontierCode v1.1 Extended (63.6% vs 61.3%), APEX-Agents (59.2% vs 57.5%), and APEX-SWE (58.8% vs 56.4%). If your default question is "which model wins the coding leaderboard," Fable still answers it on this card.
GPT-5.6 Sol Max leads DeepSWE v1.1 (73% vs Grok 65.9% vs Fable 70%) and Terminal-Bench v3.0 (34.6% vs Fable 34.1% vs Grok 26%). Terminal-Bench v3.0 is Grok 4.6's weakest published row — a 26% against a ~34% cluster. Do not confuse that with Grok 4.5's 83.3% on Terminal-Bench 2.1. Different suite, different scoring band; everyone is in the 15–35% range on v3.0.
Versus Grok 4.5 High, the upgrade is unambiguous on this sheet: +5 AA Intelligence, +227 GDPVal, +3.2 CursorBench, +11.9 DeepSWE, +4.7 FrontierCode, +10.4 APEX-Agents, +10.3 Terminal-Bench v3.0, +2.8 APEX-SWE, +264 AA-Briefcase, +2.9 Harvey. Same price. That is the clean comparison.
Two caveats belong next to every one of those numbers. First, SpaceXAI says competitor figures are the best of self-reported or publicly available results — this is not a single lab running all four models in one harness. Second, CursorBench was excluded from the Grok 4.5 launch card because a Cursor repo snapshot had leaked into training. Publishing CursorBench v3.2 now does not erase that history; treat in-IDE scores as directional until someone independent re-runs them. For how to read vendor tables in general, see the AI benchmarks guide.
Pricing, promo, and where it runs
| Detail | Number |
|---|---|
| Input | $2 / million tokens |
| Output | $6 / million tokens |
| Fast variant | 2x the base rate ($4 / $12) |
| vs Grok 4.5 list | Unchanged |
| Week-one promo | 2x included usage in Grok Build and Cursor |
| Surfaces | Cursor, Grok Build, SpaceXAI API, OpenRouter, Vercel, Cloudflare |
The economic pitch has not changed since 4.5: frontier-adjacent agent work at a mid-tier token price. What changed is the eval sheet and the distribution. Day-one Cursor access is the same play SpaceXAI ran for 4.5; the Cursor acquisition is the reason that play is available.
The 2x included-usage week is the only time-boxed reason to try it this week rather than after your current eval queue. It does not change API list price. If you already burned a week of double usage on 4.5's launch promo, this is a second, model-specific pulse — not a permanent pool increase.
Grok Build remains the in-house harness. If you want to inspect the agent loop rather than just call the model, the open-sourced Grok Build tree (Apache 2.0, read-only upstream) is still the artifact. Prompt-to-app flows live in Grok Build Mode. Install path from the announcement:
curl -fsSL https://x.ai/cli/install.sh | bash
For why the scaffolding around the model often moves the score more than a point-release does, see what an agent harness is and loop engineering. SpaceXAI's own OfficeQA-via-Genie claim is an accidental illustration of that thesis: the harness, not just the weights, is part of the number.
How they trained 4.6
SpaceXAI describes a longer supplemental training run than Grok 4.5, then a two-stage post-training stack:
- Supplemental pre/post mix — curated model-generated data for reasoning and advanced technical concepts, high-quality engineering data, and an improved optimizer and training recipe. The claim is a stronger foundation for the SFT and RL stages that followed, not a new pretrained base.
- SFT regenerated by Grok 4.5 — 4.5 was used to regenerate supervised-fine-tuning trajectories across reasoning efforts, agent harnesses, and domains (STEM, software engineering, knowledge work). Problematic traces were filtered with model-based checks.
- Agentic RL — knowledge work, general coding, and domain-specific environments including kernel optimization, web development, and computer-aided design.
That is the same "use the last model to write the next model's homework" pattern the industry has been running all year. The interesting part is the RL environment list. CAD and kernel work are not generic SWE-bench clones; they are a bet that 4.6 should survive long, tool-heavy trajectories in domains where a wrong step is expensive. Whether that shows up in your repo is an empirical question the GDPVal and APEX rows only approximate.
SpaceXAI also says safeguards were "improved and calibrated in line with the model's capabilities," with its widest-ever pre-deployment testing suite plus post-deployment and third-party testing. No public safety card with numbers landed alongside the eval table. Note the claim; do not treat it as a measured result.
Visual first passes and self-testing
The qualitative section of the announcement is doing as much work as the table, because the table does not measure "turns a product idea into a working first version."
SpaceXAI says 4.6 produces stronger first passes on visual and interactive projects than 4.5 typically did: given a concrete product idea, it can establish structure and visual language in one pass, then iterate. On longer trajectories, the lab started to see more self-testing — the model checking its own work before continuing.
That last bit is the one to verify yourself. Self-testing in a vendor write-up can mean anything from "it wrote a unit test" to "it ran the app and fixed the crash." The useful experiment this week, while 2x usage is on, is a single multi-step build: a small interactive UI, a CAD-ish layout, or a kernel-adjacent micro-optimization, with an instruction to verify before declaring done. Encode the verification as a skill or a loop, not a one-shot prompt, or you will not know whether 4.6 is actually checking work or just narrating that it did.
Grok Bot is the other half of this week
Grok 4.6 is the model. Grok Bot is the persistent-agent product SpaceXAI put into early beta on August 11 — teammates with their own VMs that sign into your tools and come back with finished work.
Musk's own gating sentence, posted around Bot launch: the company will widen the Grok Bot beta after it fixes basic early-beta issues and releases Grok 4.6. The second half of that sentence is now done. The first half — "basic issues" — is still the company's line, and no wider-access date has been published. SuperGrok Heavy, Cursor Ultra, and Cursor Teams Premium remain the named tiers.
One early-access anecdote that circulated with the Bot news: product newsletter author Lenny Rachitsky described using a bot as a talent matchmaker — pairing job seekers with hiring companies — alongside inbox auto-replies, subscription audits, and podcast-guest research briefs. That is the workload 4.6 is supposed to make cheaper and more reliable: multi-step knowledge work across real tools, not a single chat completion. It is also why the Bot post's credential warning still applies. A smarter model inside a VM that holds your logins is still a VM that holds your logins.
If you have Bot access, the practical move this week is to point one low-stakes bot at 4.6 (or whatever the product actually routes) and compare a repeated admin workflow against last week's 4.5 runs. If you do not have Bot access, the model is still usable in Cursor and Grok Build without handing it your email session.
What about Grok 4.7?
Musk's July 25 thread stacked two releases: 4.6 in two weeks, 4.7 in four. Four weeks from July 25 is roughly August 22. SpaceXAI's 4.6 post does not mention 4.7. The working read, pending the 4.7 training update, is that 4.6 is the post-training extract from the 4.5-era base and 4.7 is the next scale or training pulse.
Do not wait for 4.7 to decide whether 4.6 is usable. The 4.6 card is a shipped, priced, in-IDE model. 4.7 is still a calendar claim.
Should you switch, and what this table does not prove
If you are on Grok 4.5: switch the default for a week and keep a 4.5 baseline on one hard task. Same price, better numbers on every official row, 2x included usage. That is the easiest call on this card.
If you are on Fable 5 or GPT-5.6 Sol for coding agents: do not switch on a vendor table where Fable still leads CursorBench, FrontierCode, and both APEX suites, and Sol still leads DeepSWE and Terminal-Bench v3.0. Run your own loop against 4.6 during the promo week. Cost-per-completed-task will decide this, not AA Intelligence Index.
If you care about knowledge work and professional deliverables: GDPVal-AA v2, AA-Briefcase, and Harvey LAB are the rows that justify a serious look. They are also the rows most sensitive to harness and rubric. Reproduce on your documents before you rewrite a procurement default.
Limitations worth stating plainly:
- Self-reported competitor scores. SpaceXAI did not claim it ran Sol and Fable itself.
- No independent reproduction yet. This post is launch-day coverage, not a re-bench.
- Terminal-Bench v3.0 is a real gap at 26% versus ~34%. If your agent lives in a shell, this card does not make 4.6 the leader.
- CursorBench history. Contamination forced 4.5 to omit the bench; v3.2 is back; remain slightly skeptical of in-IDE numbers.
- OfficeQA Pro V2 / Genie SOTA is a Musk claim, not a cell on the official table.
- Grok Bot widening is still gated on unspecified beta fixes, even though 4.6 has shipped.
- Safety is described, not quantified, in the launch post.
Related on explainx.ai
- Why Grok 4.6 "Freaked Out" Over Tobi Lütke's GitHub ID — an informal agent-behavior moment from the same launch week, not part of the official evals above
- Grok Bot early beta — persistent agents that log into your tools
- Grok 4.6 and 4.7 timeline — Musk's July 25 "2 weeks" claim
- Grok 4.7: Musk's SpaceX training window
- Grok 4.5 public launch — Cursor, pricing, first evals
- Grok 4.5 vs Claude Opus 4.7 and 4.8
- Grok Build open source — Apache 2.0 harness
- Grok Build Mode — prompt to app
- GPT-5.6 vs Claude Fable 5
- What is an agent harness?
- Loop engineering for coding agents
- DeepSeek V4-Pro 0813 on Terminal-Bench 2.1 (not v3.0)
- AI benchmarks complete guide
- What are agent skills?
Primary sources: SpaceXAI, Introducing Grok 4.6 (August 12, 2026) · Elon Musk posts on X (August 12, 2026) · Grok Bot early-beta launch (August 11, 2026)
Evals, pricing, and availability are accurate as of August 13, 2026 and reflect SpaceXAI's August 12 announcement. Competitor scores on that table are the best of self-reported or publicly available results, not a single-harness re-run by explainx.ai. First-week 2x usage in Cursor and Grok Build is time-limited. Follow @explainx_ai for updates.
