explainx.ainewsletter3.5k
TrendingNewsPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

learn

pathways — start freeworkshopsbootcampscoursescertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsagentsllmsdesignsdictionaryagi trackerranks

company

aboutvisionmissionteaminstructorscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

On this page

  • TL;DR — What People Are Asking
  • What Actually Shipped to Consumers
  • VulcanBench v3 — The 91.3% Number, Unpacked
  • Graffiti Conjecture 284 — What Was Claimed
  • What Developers Are Actually Using It For
  • Honest Limitations
  • Practical Takeaway for Teams
  • Related on explainx.ai
← Back to blog

explainx / blog

Grok 4.5: VulcanBench Lead and Graffiti Conjecture 284

Grok 4.5 hits 91.3% on VulcanBench v3 at $0.32/solved task and a Capy agent on Grok 4.5 Medium reportedly refutes Graffiti Conjecture 284. What holds.

Jul 24, 2026·8 min read·Yash Thakker
Grok AIxAIBenchmarksGraph TheoryCoding Agents
go deep
Grok 4.5: VulcanBench Lead and Graffiti Conjecture 284

Two July stories collided into one Grok headline: consumer rollout of Grok 4.5 across Grok.com, X, and mobile apps, and a louder claim that the same model family is winning coding cost benchmarks while an agent on Grok 4.5 Medium helped poke a hole in a decades-old graph theory conjecture. explainx.ai separates launch marketing, VulcanBench numbers, and the Graffiti 284 claim so you can decide what is shipped, what is scored, and what still needs a mathematician’s pen.

We already covered the Cursor/SpaceXAI public launch and the Opus comparison. This piece is the late-July update: wide consumer access, VulcanBench Report No. 4, and the math thread.

TL;DR — What People Are Asking

table · 2 cols
QuestionShort answer
What’s new this week?Grok 4.5 called out as live on Grok.com, X, iOS, Android
Coding claim?91.3% on VulcanBench v3 (21/23) at medium effort
Cost claim?~$0.32 per solved task in that report (suite ~$6.67)
Speed claim?Vendor: ~80 TPS; efficiency pitch is steps + $/task
Math claim?Agent on Grok 4.5 Medium → counterexample to Graffiti 284
Peer reviewed?No public math paper yet — traces + checks reported
Which model in the app?UI often opaque — use API IDs for serious work
Replace Claude/GPT?Suite-specific win; not a universal crown
Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

What Actually Shipped to Consumers

xAI’s Grok account framed Grok 4.5 as “our most capable model yet,” available across the consumer surfaces people already use: the web app, X, and the mobile apps. That matters because earlier July coverage was skewed toward Cursor, Grok Build, and API/console access — developer channels — not every free or Premium X user.

The practical effect:

  1. More people hit 4.5 by default without reading a model picker.
  2. Anecdotes get noisier — game prototyping, agentic coding, research assistants — without shared eval harnesses.
  3. Support confusion rises when the UI does not print grok-4.5 next to the reply (a recurring X complaint: “it said 4.5 last week; today it says 4.0”).

For builders, treat consumer Grok like any other chat surface: great for exploration, weak for reproducible agent loops. If you need deterministic model selection, stay on the console/API path documented in our launch guide and pair it with a proper agent harness.

VulcanBench v3 — The 91.3% Number, Unpacked

VulcanBench is not a vibes leaderboard. The July Technical Report No. 4 run compared Grok 4.5, Claude Fable 5, and GPT-5.6 Sol across 23 software tasks at low/medium/high effort — 207 graded runs with hidden tests. Tasks are drawn from real merged PRs across multiple languages.

Headline row that went viral:

table · 4 cols
ConfigScorepass@1Approx $/solved
Grok 4.5 (medium)21/2391.3%~$0.32
Grok 4.5 (high)21/2391.3%~$0.42
Fable 5 (low)†20/2387.0%~$0.61
GPT-5.6 Sol (high)20/2387.0%~$0.80
Grok 4.5 (low)19/2382.6%~$0.18

† Fable 5’s card notes safety refusals on some runs with Opus 4.8 fallback — read the report’s methodology before declaring a three-way knockout.

Why “medium” is the story

The report’s own takeaway: Grok plateaus at medium. High effort burned ~33% more tokens and ~31% more money for the same 21/23, failing the same two tasks. That is the efficiency narrative in one sentence — not “Grok is smarter at every temperature,” but “the Pareto front for this suite sits at medium.”

Two frontier-hard tasks went 0-for-27 across models and efforts. So 91.3% is a ceiling of the suite, not proof Grok clears every repo on Earth.

How to read this vs other benches

Cross-check against:

  • Our Grok vs Opus table (SWE-Bench Pro, Terminal-Bench, GDPval+)
  • Artificial Analysis / harness notes from the launch week
  • Your own loop engineering evals on your stack

VulcanBench rewards cheap, correct PR-style fixes. It does not measure long-horizon research agents, design taste, or cyber refusal behavior. If your workload is closer to Grok Build office agents or multi-hour math search, borrow the cost mindset — do not copy the percentage onto a slide as “best model.”

Graffiti Conjecture 284 — What Was Claimed

Separate from coding benches, July chatter claimed a ~30-year open graph-theory statement — Graffiti Conjecture 284, from Siemion Fajtlowicz’s automated conjecture generator — fell to a counterexample found with help from a Capy agent running Grok 4.5 Medium.

Reported shape of the claim:

  • Object: Hoffman–Singleton graph (classic strongly regular graph: 50 vertices, degree 7, diameter 2)
  • Method: Agent exploration during a Slack-style research thread; counterexample spotted in minutes per developer write-ups
  • Verification: Justin Sun’s group and others reportedly checked traces and cross-checked other models
  • Status: Not a peer-reviewed journal article as of this post — vendor/developer narrative first

Why Hoffman–Singleton shows up

Famous named graphs are exactly where automated conjecture systems and AI search collide: they are small enough to compute invariants, weird enough that hand proofs get sticky, and well documented enough that an agent can retrieve the right object. Finding that δ* (or whichever dual-degree quantity the conjecture bounds) violates Graffiti 284 on this graph is the kind of search + check win AI already demonstrated elsewhere in 2026.

How to place it next to other AI math news

table · 3 cols
ClaimWhoVerification bar
Planar unit distance / Erdős lineOpenAI internal modelExternal mathematicians + papers — see our coverage
Jacobian conjecture counterexampleHuman + Fable 5 assistExplainer: what the conjecture is
Expert ChatGPT sessionTerence TaoShared conversation analysis
Graffiti 284Capy + Grok 4.5 MediumDeveloper traces; await formal write-up

The honest product lesson matches “will AI replace mathematicians?”: models are getting better at finding counterexamples; humans still own definitions, publication, and social verification.

What Developers Are Actually Using It For

X threads around the wide release cluster into three jobs:

  1. Game prototyping — fast iteration loops where 80 TPS and cheap tokens beat max SWE-Bench.
  2. Agentic coding — multi-file fixes where VulcanBench-style PR tasks dominate.
  3. Research assistants — conjecture hunting, literature skim, graph invariant checks — high upside, high hallucination risk (Artificial Analysis previously flagged rising hallucination rates even as knowledge scores rose).

If you are wiring Grok into team workflows, steal identity and audit ideas from Block Buzz / Nostr agent rooms rather than dumping an unlabeled consumer chat into production.

Honest Limitations

  • Harness dependence: VulcanBench scores Grok under its harness and pricing file; Claude Code / Codex harnesses can reorder the ranking.
  • Refusal fallback: Fable’s reported score includes Opus fallback on some refused steps — apples-to-apples requires reading the card.
  • Math claim prematurity: No arXiv citation in the viral summary — do not teach Graffiti 284 as “settled by Grok” in a classroom yet.
  • Model opacity: Consumer apps that hide version labels poison A/B intuition and support tickets.
  • Not free of secrets risk: Treat any “upload the repo” agent path with the same care as Grok Build secret-scan issues.

Practical Takeaway for Teams

  1. Budget: If your evals look like VulcanBench, trial Grok 4.5 medium before burning Opus/Sol high effort.
  2. Routing: Keep a stronger verifier model for math and security; use Grok for volume coding.
  3. Labeling: Force model IDs in your harness logs — never trust the consumer chat chrome.
  4. Math: Celebrate candidate counterexamples; publish only after independent check (Lean or human referee).
  5. Corpus: Bookmark the launch post for pricing/context and this post for VulcanBench + Graffiti.

For agent product managers, the combined lesson is effort ladders beat brand loyalty. Medium Grok winning a PR suite does not cancel Opus on multilingual SWE or Fable on long research loops — it means your router should price the rung, not the logo. Pair that with the math-thread humility: a Slack agent finding Hoffman–Singleton is exciting; a refereed note is shipping.

Related on explainx.ai

  • Grok 4.5 public launch — Cursor, SpaceXAI, pricing
  • Grok 4.5 vs Claude Opus 4.7 / 4.8
  • Grok Build open-source harness
  • What is an agent harness?
  • Jacobian conjecture explained
  • OpenAI planar unit distance / Erdős
  • Will AI replace mathematicians?
  • India Bitchat GitHub block — same week’s policy story

Primary sources to verify: VulcanBench Technical Report No. 4 model card (July 2026 3-way run) · xAI / Grok product posts on X · developer write-ups on Graffiti 284 / Hoffman–Singleton · Artificial Analysis Grok 4.5 notes.


Benchmarks, pricing, and math claims reflect public materials and press summaries as of July 24, 2026. VulcanBench figures are suite-specific; Graffiti Conjecture 284 status awaits formal mathematical publication. Consumer availability and model labeling can change without notice — verify in your own console before committing workloads.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

Related posts

Jul 9, 2026

Grok 4.5 vs Claude Opus 4.7 and 4.8: Benchmarks, Price, and When to Switch

SpaceXAI's Grok 4.5 lands July 8–9 at $2/$6 — Musk says Opus 4.7-class, faster and cheaper. Opus 4.8 still leads SWE-Bench Pro and independent DeepSWE 1.1. Snorkel GDPval+ favors Grok on professional work. Full decision tree.

Aug 13, 2026

Grok 4.7 in 3–4 Weeks: SpaceX Training Data and the Timeline Slip

After Grok 4.6 shipped, Elon Musk said Grok 4.7 is significantly better and should be ready in 3 to 4 weeks. Initial training is complete; SpaceX company data is going into supplemental training. The July ~August 22 date has slipped to early-to-mid September 2026. Here's what that claim is, what it isn't, and whether you should wait.

Aug 13, 2026

Grok 4.6 Launch: Official Evals, Same $2/$6, Cursor Access

SpaceXAI released Grok 4.6 on August 12, 2026 — a long-running-agent upgrade at the same $2/$6 as Grok 4.5, with 2x included usage in Cursor and Grok Build for week one. Official evals tie GPT-5.6 Sol at 61 on the AA Intelligence Index. Fable 5 Max still leads several coding benches; Grok 4.6 High leads GDPVal-AA v2, AA-Briefcase, and Harvey LAB.