A hardware leak is making the rounds that matters more to people who run models at home than to gamers: Nvidia is reportedly done sending GB202 chips to GeForce boards. If true, the RTX 5090, the 32GB card behind a large share of single-GPU local AI builds, would be effectively discontinued. The report appeared on October 10, 2026, via VideoCardz, and was echoed by eTeknix and others. Nvidia has not confirmed anything.
Because the story is spreading before any official word, this is a "claimed versus verified" post. We lay out who said what, how credible each piece looks, and what local-AI builders can sensibly do while the facts are unsettled.
TL;DR: the rumor in one table
| Question | Answer |
|---|---|
| What is claimed? | All future GB202 supply goes to RTX PRO cards, which would end the RTX 5090 and its China variants. |
| Who says so? | Leaker MEGAsizeGPU (posted October 10), with detail from BenchLife and a second leaker, hongxing2020. |
| Has Nvidia confirmed it? | No. No announcement or statement found as of this writing. |
| Is it plausible? | Yes on economics, since the same die sells in workstation cards for far more. Timing and scope are unproven. |
| Would owners be affected? | Not directly. Existing cards keep working. Supply and price are the open questions. |
| What else is rumored? | A 24GB RTX 5080 in Q1 2027, replacing the 16GB model, from BenchLife and a leaked factory notice. |
What exactly is being claimed?
According to eTeknix's summary, MEGAsizeGPU wrote that "All future GB202 supplies are RTX Pro exclusive." Taiwanese outlet BenchLife reportedly says board partners have already stopped receiving RTX 5090 chips for production, and suggests Nvidia may not frame this as a formal end-of-life decision but as a reallocation of capacity to more profitable lines.
A second leaker, hongxing2020, shared what is described as a factory notice saying Nvidia will not sell any RTX 5090 models, including the China-specific D variants, and that an RTX 5080 24GB will become the top GeForce card while the 16GB version is discontinued. eTeknix notes that MEGAsizeGPU has not backed up this second part, so it rests on shakier ground. BenchLife adds that the 24GB 5080 might arrive in Q1 2027, with partners told only verbally.
What is verified, and what is not?
Verifiable facts:
- GB202 is the biggest desktop Blackwell die and is shared between the GeForce RTX 5090 and professional cards. ITHome, as cited by eTeknix, lists three GB202-based RTX PRO models on sale: the RTX PRO 6000, 5500 and 5000 Blackwell.
- Per eTeknix, the RTX PRO 6000 Blackwell has 24,064 CUDA cores and 96GB GDDR7; the RTX 5090 has 21,760 CUDA cores and 32GB on a 512-bit bus. The RTX 5080 uses a different chip, GB203, with 16GB.
Unverified:
- That chip allocation to GeForce has actually stopped.
- That the decision is permanent rather than a temporary reallocation.
- The 24GB RTX 5080 timing and the factory notice.
- Any Nvidia statement. We searched for an official post or partner announcement and found none.
A faded dotted card above a solid card with a check mark, illustrating a rumor versus a confirmed report
Stopping new chip allocation is also not identical to discontinuing a product: finished boards and chips already in the channel can keep selling for weeks or months. Nvidia has a history of pushing back on this kind of story. In September 2025, after RTX 5090 and 5080 Founders Edition listings disappeared, Nvidia told press that "GeForce RTX 50 series Founders Editions continue to be in production," as Club386 reported. That was a different situation, but it shows why waiting for an official statement is wise.
Why is the economics plausible?
When one die can go into a consumer card or a professional card with 96GB of memory at a far higher price, allocation tends to follow margin. eTeknix makes the same argument and points to the surrounding market: RTX 50 prices in Germany reportedly rising sharply and a global memory squeeze. Memory cost is also the stated reason the earlier RTX 50 SUPER launch was reportedly put on hold, since 3GB GDDR7 modules were too expensive, according to the same report on a July VideoCardz story.
For local AI, the logic loops back: demand from people buying cards for models is part of what has pushed prices up. One Hacker News commenter, dabinat, argued that 5090 prices are now high enough that more buyers want it for AI than for gaming. That is one reader's opinion, not data, but it matches the direction of the market.
What does this mean for running models locally?
A stack of blocks with one green block fitting a slot, representing GPU memory capacity for local models
The 32GB of VRAM on the RTX 5090 is the main reason it shows up in so many guides. Our own coverage includes benchmark-style numbers from it: the FreeToken paper's single-GPU MoE results used an RTX 5090 and reported tens of tokens per second on large mixture-of-experts models with host memory help. The Qwen 3.6 27B local setup guide and the Qwen 3.8 27B open-weight comparison sit squarely in the 24 to 32GB class.
If supply ends, the practical fallbacks are:
| Option | Memory | Trade-off |
|---|---|---|
| Used or remaining RTX 5090 | 32GB | Prices likely rise if new supply stops |
| Rumored RTX 5080 24GB | 24GB | Roughly half the CUDA cores of a 5090, per eTeknix, and not yet confirmed |
| RTX PRO 6000 Blackwell | 96GB | Far more memory and money; workstation pricing |
| Apple silicon unified memory | Varies | See DFlash 2 on MLX and M5 Max |
| Smaller or compressed models | Lower | Fits a 24GB card; see near-lossless 27B compression |
For a longer-term view, our gaming and AI hardware cost forecast and the local hardware projection thread discuss how memory pricing shapes what ordinary developers can afford.
A small house-shaped box with a glowing core, representing a local LLM home machine
How did we get here? A short timeline
The rumor lands on top of a year of supply stress, so it helps to line up what is on the record.
- September 2025: Nvidia pulled RTX 5090 and RTX 5080 Founders Edition listings from its store. Nvidia's statement to press said the cards stayed in production and were limited editions that go out of stock, as quoted by Club386.
- July 2026: VideoCardz reported that RTX 50 SUPER cards had reached at least one board partner, but that Nvidia put the launch on hold because 3GB GDDR7 modules were too expensive, per the eTeknix recap. Those modules are what a 24GB card on the 5080's 256-bit bus would need.
- Early October 2026: eTeknix reported that RTX 50 prices had climbed 66 percent in Germany, with the RTX 5090 up 132 percent. We have not independently verified that figure, so treat it as one outlet's number.
- October 10, 2026: the GB202 allocation leak, plus the factory-notice claim about the 5080 24GB.
Read together, the pattern is tightening supply and pricey memory rather than a single decisive event. That is why a reallocation story sounds credible even before anyone confirms it, and also why it is easy to over-read. Rising prices fit a discontinuation, but they also fit ordinary scarcity.
What are people asking about this rumor?
Does "GB202 to RTX PRO only" kill the China variants too? The leak says yes, because the RTX 5090 D and 5090 D v2 use the same silicon, per eTeknix. That detail is part of the unverified claim.
Could the 5090 return later? Nothing in the reporting says. eTeknix notes there are separate reports that Nvidia and AMD may delay next-generation graphics cards until 2028, which would leave the top of the GeForce range thin for a long time, but that too is a report, not an announcement.
Is the 24GB RTX 5080 a real fix for local AI? It would add memory over the 16GB card, which helps with models in the 20 to 24GB range after quantization, but it would not match a 5090's compute. It also depends on memory pricing, since the earlier SUPER delay was blamed on exactly that.
Why should AI builders care about a gaming card? Because consumer GPUs are the cheapest way to get large VRAM into a developer's own machine. Anything that changes the availability of the top consumer card changes how many people can run open-weight models without renting cloud GPUs, which feeds back into which models the community builds tooling for.
How to read hardware leaks like this one
Leaks are not all equal. A useful checklist: does the claim name a specific mechanism (here, chip allocation to partners), is there a second independent source (BenchLife is a separate Taiwanese outlet, but it may share supply-chain sources), and does the claim make a falsifiable prediction (retail listings drying up, partner announcements)? This rumor scores moderately: a specific mechanism and a second outlet, but no document anyone can check, and a rumored second part (the 5080 24GB) that even the main leaker has not backed.
What should you do right now?
- Do not panic-buy on a rumor. Prices may already include speculation, and the report could be wrong or partial.
- Decide by workload. If your models fit in 24GB after quantization, the rumored 5080 24GB or a used card may be enough. If you need 32GB or more, weigh a 5090 against unified-memory machines.
- Watch for primary signals: an Nvidia statement, board-partner announcements, or retailer listings going empty without restock.
- Treat leakers as probabilistic. MEGAsizeGPU's claim has now been repeated by several outlets, but they are repeating the same source, not independent confirmation.
The bottom line
The rumor is plausible and consistent with how Nvidia has historically weighed professional against consumer supply, but it is not confirmed. For local-AI builders, the cautious reading is that 5090 availability could tighten, not that existing cards or software stop working. We will update this post if Nvidia or its partners say anything on the record.
Related reading
- FreeToken: 753B MoE on a single GPU
- Qwen 3.6 27B local development guide
- Qwen 3.8 27B open-weight model comparison
- Gaming and AI hardware cost forecast 2027
- Local hardware projection for 2028
- DFlash 2 on MLX and M5 Max
Details reflect reporting as of October 10, 2026 and may change if Nvidia responds.
