A 15-second video showing DeepSeek 4.1 Flash and GPT-6 Astra each generating a simulated planet picked up attention on X this week, framed as a head-to-head matchup: "DeepSeek 4.1 Flash VS GPT-6 Astra at simulating planets." It came from the official account of Flowith, the AI platform that rendered both outputs — and that detail matters more than the video itself.
TL;DR: what's actually in this video
| Question | Answer |
|---|---|
| Who posted it? | Flowith.io's own official account, promoting its platform |
| What's shown? | Two AI-generated planet simulations, one per model, no side-by-side scoring |
| Is a prompt disclosed? | No |
| Is a judging method disclosed? | No |
| Has anyone reproduced it independently? | Not as of this post |
| Is this a real benchmark? | No — treat it as a product demo, not a controlled comparison |
| How much engagement did it get? | Under 2,000 views — modest, not viral by the usual standard |
What the clip actually shows
Flowith's post pairs a short caption — "DeepSeek 4.1 Flash VS GPT-6 Astra at simulating planets, both rendered by https://flowith.io" — with a 15-second video. There's no prompt text shown, no explanation of what "simulating planets" means as a task (physics accuracy? visual fidelity? code correctness? an interactive demo versus a static render?), and no scoring criteria distinguishing a "win" for either model. One reply on the thread simply asks how to reproduce it — "how can I make it, is it through the web flow or what" — and gets no answer in the visible thread, which tells you even Flowith's own engaged audience doesn't have enough information to try it themselves.
This is worth naming plainly: a company that sells a multi-model comparison canvas has a direct incentive to post clips that make model comparisons look impressive and effortless, since that's the product's core pitch. That doesn't mean the outputs shown are fake — there's no reason to doubt DeepSeek 4.1 Flash and GPT-6 Astra both produced some kind of planet-simulation output on Flowith's platform — but it does mean the video functions as an advertisement for Flowith first, and as evidence about relative model capability a distant second.
The baseline problem: DeepSeek 4.1 Flash isn't a settled release
There's a second layer of weakness specific to this comparison. explainx.ai has covered DeepSeek V4.1 Flash's rollout in detail, and as of the most recent coverage, it has circulated primarily through beta API endpoints and architecture previews — including one temporary endpoint that expired September 10, 2026 — rather than shipping as a finalized model with a published model card and benchmark suite. A separate report covered a claimed ~75% KV-cache HBM reduction tied to its new architecture, again ahead of a full, stable public release.
That matters for this specific comparison because "vs" framings implicitly assume both sides are stable, comparable things. Comparing a fully shipped model like GPT-6 Astra against a model still moving through beta endpoints and architecture previews means you're not just missing the prompt and scoring method — you're not even fully sure which version of DeepSeek's model actually produced the output in the clip, or whether that exact configuration will still exist by the time someone else tries to reproduce it.
What both models actually are, separate from this clip
GPT-6 Astra is OpenAI's current flagship multimodal model, and explainx.ai has covered several of its capability claims this year with the same skepticism this post applies here — including a viral, similarly unverified claim that Astra beat the video game Portal unaided and grew a large procedural three.js forest. That earlier post's core lesson applies directly here too: single-source demo claims, however plausible, aren't the same thing as reproducible evidence, and grouping several impressive-sounding claims together in one thread tends to generate more social proof than any individual claim actually earns.
DeepSeek's models, meanwhile, have built a genuine reputation for aggressive price-performance positioning rather than headline-grabbing creative demos — the DeepSeek V4 line has been covered on cost-per-task and pricing terms more often than on "wow" factor generation demos, which makes this planet-simulation framing something of a departure from how DeepSeek typically gets discussed.
Why "simulating planets" is actually an interesting task, even if this demo doesn't prove much
It's worth separating skepticism about the demo's evidentiary value from skepticism about the underlying task. Generating a convincing planet simulation, whether as a static render or an interactive scene, genuinely tests several distinct model capabilities at once: writing correct code (whatever graphics library or shader approach the model chooses), getting the physics or visual approximation of things like axial tilt, atmospheric scattering, and terrain or cloud texturing plausibly right, and composing all of that into a scene that actually looks like a planet rather than a textured sphere. That combination — code correctness plus visual/scientific plausibility plus aesthetic composition — is exactly the kind of multi-dimensional creative-coding task that's genuinely hard to benchmark with a single number, which is part of why vendors gravitate toward these demo-style comparisons in the first place. A rigorous version of this test would fix the prompt, specify what counts as success (does the planet need correct relative sizing? realistic axial tilt? a specific programming language or library?), and have multiple independent people or an automated rubric score the outputs — none of which happened here.
This is also why creative-coding demos are simultaneously some of the most viscerally convincing AI capability evidence and some of the least rigorous. A viewer's gut reaction to "that looks like a real planet" is a genuine, hard-to-fake signal about output quality on that one attempt — but it says very little about consistency across different prompts, about how much prompt engineering or retry attempts happened off-camera, or about whether the same model would produce something equally impressive on a different but structurally similar task.
What a real comparison would need
If you wanted to actually settle whether DeepSeek 4.1 Flash or GPT-6 Astra is better suited to this kind of visual/procedural generation task, the missing pieces are specific and checkable: the exact prompt text given to each model, including any system instructions or platform-level scaffolding Flowith's canvas might inject automatically; whether both models were given the same number of attempts, or whether one output was cherry-picked from several tries; what programming language, library, or rendering approach each model was free to choose (a model that picks three.js versus one that hand-rolls WebGL isn't competing on equal footing for a "which model is smarter" claim, even if the visual result looks comparable); and some disclosed criteria for what "winning" this specific comparison would even mean. None of that accompanies Flowith's clip, which is a normal and unremarkable thing for a 15-second marketing video to omit — vendors aren't obligated to publish a methodology section under a promotional post — but it's exactly why the clip shouldn't be treated as evidence beyond "these two models can both produce planet-shaped outputs on this platform."
How to actually evaluate a clip like this
If you want to judge whether either model is genuinely strong at generating simulations, procedural 3D content, or scientific-visualization-style code, the useful move isn't watching a vendor's 15-second highlight reel — it's running the same prompt yourself, or checking whether someone with no stake in Flowith's product has reproduced comparable results independently, ideally with the actual prompt and generated code published alongside the output. A repeatable test with a disclosed methodology, even an informal one, is worth more than a polished demo clip from the company that built the comparison tool in the first place.
That's a generally useful filter for the growing genre of "Model A vs Model B" clips circulating on social platforms: ask who benefits from the comparison looking a certain way, and ask what you'd need to see to try it yourself. If those two things aren't available, treat the clip as marketing content that happens to feature two real AI models, not as a data point about which one is actually better.
This genre isn't going away — if anything, it's becoming a standard marketing format for any platform whose core pitch is letting users run multiple models side by side. Flowith's own account posted a similar format around the same time, pitting the video models Seedance 2.5 against Wan 3.0, following the same pattern: no disclosed prompt, no scoring rubric, just a compelling clip and an implicit invitation to sign up and try the comparison feature yourself. Recognizing the format matters more than debunking any single instance of it, since new "vs" clips from the same accounts will keep appearing on a regular cadence, each one carrying the same evidentiary weaknesses as this one.
Related reading
- DeepSeek V4.1 Flash: a two-day beta with a new multimodal architecture
- DeepSeek V4.1 Flash cuts KV cache HBM by ~75% — what changed
- GPT-6 Astra reportedly beat Portal and wrote a Bach chorale — unverified
- DeepSeek V4 Flash 0731: ARC-AGI cost-per-task breakdown
- Artificial Analysis Intelligence Index v4.2, September 2026
- Source: Flowith's original post on X, Flowith.io
This post evaluates a promotional clip and its framing as of September 15, 2026. No independent reproduction of the demonstrated task was available at the time of writing.
