On September 30, 2026, Bloomberg reporters Julia Love and Davey Alba described a familiar lab tension at Google: Gemini 4 looks strong on the coding benches the industry screens with, and less convincing when employees actually try to ship work. Google disputed that characterization the same day. It also published Gemini 4 Argon with a scorecard that already loses two of the three most relevant agentic-coding rows.
That combination — anonymous internal skepticism, a public denial, and a mixed official table — is the story builders should use. It is not a reason to treat one DeepSWE cell as a production swap.
explainx.ai's launch numbers post covers price, the 1 million token output cap, and Fairwind access. This companion is the bench-versus-real-work read: what Bloomberg attributed to people with access, what Google said on the record, and how to eval coding without outsourcing the decision to a vendor's best cell.
TL;DR
| Question | Direct answer |
|---|---|
| What did Bloomberg report? | Internal skepticism: Gemini 4 does well on industry benchmarks, less well when employees put it to work, including certain coding tasks. Sources requested anonymity. |
| What did Google say? | It would be inaccurate to say Gemini 4 is underperforming in coding. Kavukcuoglu was encouraged. One employee cited a "large consensus" that it is at the frontier. |
| Is this a published Google eval? | No. The skepticism is anonymous-source journalism. The denial is Google's. Neither is your repo. |
| Does Google's own table already split coding? | Yes. Argon leads DeepSWE v1.1 at 77.9%, loses FrontierSWE v2 55.0 vs Astra 65.5, loses Terminal-bench 4.0 57.4 vs Opus 5.5 66.4. |
| Front-end design / "too big"? | English-language recaps of Bloomberg add a front-end design weakness and a very large model cost note. Treat those as relayed claims, not Google specs. |
| Can I call Argon today? | No public model ID. Fairwind testers first. |
| Should I swap production? | No. Pin repo evals that include a FrontierSWE-shaped multi-file fix and a DeepSWE-shaped long ticket. |
| Is Vending-Bench 2 the coding proof? | No. #3 on Vending-Bench 2 is long-horizon business simulation, not coding quality. |
What Bloomberg actually attributed
Bloomberg's Gemini 4 piece is paywalled. The quotes circulating in English recaps, including the block that first spread on September 30, 2026, are consistent across multiple write-ups. explainx.ai is treating the following as primary attributed language, not as something we independently interviewed:
As Alphabet Inc.’s Google prepares for the coming launch of Gemini 4, it’s grappling with internal skepticism over how well the flagship artificial intelligence model performs in key areas, such as coding. While Gemini 4 has performed well on benchmarks the industry uses to gauge model efficacy, it does less well when employees actually put it to work, according to people with direct access to the effort. The model struggles to handle certain coding tasks, said the people, who requested anonymity to discuss an internal matter.
That is a gap claim, not a score. It says: public-style benches look good; some internal work does not. It does not name the tasks, the harness, the rival checkpoint, or a pass rate. Anyone quoting it as "Gemini 4 cannot code" is inflating the reporting.
The same recaps add texture that was not in the short block above. We include it only as relayed Bloomberg sourcing:
- Some people with access said the model is not particularly adept at front-end design — the look-and-feel layer of apps and sites.
- Some sources framed a benchmaxxing worry: optimizing for public evals more than for internal jobs.
- Bloomberg described Gemini 4 as a very large model, which recaps tied to serving-cost pressure. Google has not published a parameter count. The published economics are intro $2 / $10 per million input/output tokens, later $4 / $20 — documented in the Argon launch post, not in this rumor layer.
If you cannot open the paywalled original, do not treat those three bullets as stronger than the quoted paragraph. They are consistent relays, not a second on-record dataset.
What Google said — and what it already published
Google's rebuttal, as English recaps of the Bloomberg exchange put it, is specific: it would be inaccurate to say Gemini 4 is underperforming in areas such as coding. The company pointed to comments last week from Koray Kavukcuoglu, head of Google DeepMind, that he was encouraged by the model's performance — the same fast-track posture we covered when he said Gemini 4 was in post-training and Google wanted an early release as soon as possible.
A Google employee familiar with model development told Bloomberg there is "large consensus" internally that Gemini 4 is at the frontier, and denied that the models struggle with real-world coding tasks. That is still one named-role, unnamed person. It contradicts the other unnamed people. Journalism can hold both; a routing table cannot.
Google's own launch post is the on-record product story. On September 30 it announced Gemini 4 Argon, rolling out first to trusted cyber defenders in the Fairwind Program. Kavukcuoglu called it the company's next era of frontier intelligence. The post says Argon is already powering internal workflows, with thousands of Googlers highlighting strengths in specialized coding tasks, deeper research, and writing quality. It says Google engineers use Argon for everyday debugging, large-scale codebase migrations, and algorithm design, and it cites a state of the art on DeepSWE v1.1 at 77.9%.
Those sentences are Google's. They are not independent replication. They also do not erase the other rows on the same chart.
The coding split Google already printed
| Agentic coding row (Google's table) | Gemini 4 Argon | Who leads | Gap |
|---|---|---|---|
| DeepSWE v1.1 | 77.9% | Argon | Lead |
| Vibe Code Bench | 91.9% | Argon, photo finish | All four models above ~89% |
| FrontierSWE v2 | 55.0% | GPT-6 Astra 65.5% | −10.5 |
| Terminal-bench 4.0 | 57.4% | Claude Opus 5.5 66.4% | −9.0 |
This is the practitioner hook. You do not need an anonymous source to believe that one coding suite is not all coding. Google's chart already shows Argon as a DeepSWE high with FrontierSWE and Terminal-bench losses. If your production harness looks like a messy multi-file SWE ticket or a long terminal session, the official table is telling you to keep Astra or Opus 5.5 as the control — the same advice we gave on launch day.
The leak-week tables that called a rumored "Gemini 4 Pro" an 88.7 DeepSWE / 95.3 Terminal-Bench monster are not this SKU. That mismatch is already unpacked in the floatplane and arena leak post. Do not mix rumor cells with Argon's published rows when you brief a VP.
Vending-Bench 2 placing Argon third is independent and useful for long-horizon honesty in a simulated shop. It is not a coding-quality verdict. Do not cite it as proof Bloomberg was right or Google was right.
Timeline the reporting also filled in
Bloomberg's piece, via the same recaps, is not only about one week of vibes. It also restated product history that matters for how you read "flagship":
- Google had planned Gemini 3.5 Pro for June 2026 and abandoned that effort. Kavukcuoglu's September 24 comments already treated the skip as strategy; Bloomberg's sources treat it as a capability miss that became a calendar miss. Our fast-track post tracks the on-stage version.
- Some employees think Anthropic's Fable line and OpenAI's Astra are still improving faster. Other employees think Gemini 4 has caught up. Those are two camps inside the same company, reported as such. They cancel as a single "Google thinks X" headline.
- Versions of Gemini already sit under Search, Maps, Gmail, and Chrome — products with more than a billion users each, per the recaps. That is why internal coding quality is not a hobbyist argument. It is why you still should not confuse consumer Gemini surfaces with Argon-the-API-SKU. You cannot bake off a Fairwind-only checkpoint against production Search.
Google's pushback, as summarized in follow-on coverage of the launch, also highlighted multimodal, video, safety, and cybersecurity strengths and shared a table that matches or beats GPT-6 Astra and Claude Opus 5.5 on many non-coding rows (Vals Index, AutomationBench, long GraphWalks). That is real on the vendor sheet. It is also a reminder: knowledge-work wins do not close a FrontierSWE hole.
What people are asking
Is this just "benchmaxxing" gossip?
Maybe, and maybe not. The word showed up in recaps of Bloomberg as a source frame: engineers concentrating on a good public score more than on a product that does a job well. How to read AI benchmarks is the checklist for that worry — dataset, harness, retries, judge, exclusions, whether the task resembles your job.
Google's table is an unusually clean case study for that checklist because the lab published the losses. A benchmaxxed launch typically hides the weak rows. Argon did not. That does not prove internal day-to-day coding is excellent. It does prove that "Gemini 4 topped every coding bench" was never the official claim.
Does a DeepSWE lead contradict Bloomberg?
Only if you collapse all software work into one suite. DeepSWE v1.1 is Google's chosen long-horizon SWE headline. FrontierSWE v2 and Terminal-bench 4.0 are different jobs. Bloomberg's anonymous sources could be describing exactly the class of work those weaker rows measure — or they could be describing UI polish, review latency, or a task family that is not on any public chart.
Until someone publishes a matched internal-vs-public eval, the honest statement is: the official scorecard already predicts uneven coding. That is enough to refuse a one-cell swap. It is not enough to declare the anonymous sources proven.
What about Gemini 3.8 Flash — did Google already "fix coding" in September?
Gemini 3.8 Flash shipped DeepSWE and Flash Cyber numbers on September 2, 2026, at Flash pricing, with Fairwind as the cyber gate. That is a different SKU and price tier. Using 3.8 Flash's coding story as a substitute for Argon's flagship story is a category error. Keep Flash on cost-sensitive routes; treat Argon as the unreleased frontier bake-off.
If Fairwind testers like it, does that settle coding quality?
No. Fairwind is a vetted defender program. Testers are not a random sample of product engineers, and they get the model without public cyber guardrails. Positive Fairwind notes would tell you something about defensive security workflows. They would not automatically tell you Argon writes production front-end or survives a 40-file refactor in your monorepo.
Should I wait for an independent SWE leaderboard?
Yes, and you should not wait with an empty eval folder. Independent labs have not, as of October 1, 2026, reproduced Google's four-model coding table under a public harness we can rerun. Build the private suite now. When a versioned API ID appears, you are measuring delta against your baseline, not discovering your requirements in a panic.
What to do this week (copy-paste eval shape)
Stay on billed models you can call: Gemini 3.8 Flash, GPT-6 Astra, or Claude Opus 5.5. Do not rewrite routers because a flagship blog post exists.
Pin 30–50 tasks that look like your actual tickets. Minimum mix, matching Google's own split:
eval_folder/
01_frontierswe_shaped/ # multi-file bugfix, real tests, no golden patch in prompt
02_deepswe_shaped/ # long-horizon ticket, many tool steps
03_terminal_shaped/ # install, repro, patch, prove in a shell
04_frontend_shaped/ # layout, a11y, design-system constraints (Bloomberg relay)
05_review_shaped/ # "would you merge this?" with your CI
Score pass / fail / human-repair minutes / $ per passed task. That last column is how you absorb the "very large model" cost rumor without pretending you know the parameter count. If Argon later wins DeepSWE-shaped tasks and loses frontend-shaped ones, you have a route, not a religion.
Do not use Vending-Bench 2 as a proxy for any of those folders. If you run profit-max agents, keep it as a honesty check. Coding quality is a different axis.
Honest limitations
- Anonymous sources. Bloomberg's coding-struggle claims come from people who requested anonymity. We did not interview them. We did not see internal traces.
- Google denial. The company called the underperforming-in-coding framing inaccurate. One employee claimed large internal consensus the other way. Both can be true in different teams.
- Paywall and relays. The exact Bloomberg PDF is not our primary file. We used the widely quoted paragraph plus consistent English recaps. Front-end design, benchmaxxing, and "very large model" are relayed, not independently confirmed wording from a page we control.
- Fairwind-only access. explainx.ai has not run Argon. Neither has a typical API customer. Internal Google use is not your SLO.
- Vendor vs independent evals. DeepSWE / FrontierSWE / Terminal-bench numbers here are Google's launch table. Vending-Bench 2 is independent and not a coding test.
- No stock read. Product versions of Gemini sit under billion-user Google surfaces. That is distribution context, not a trading thesis.
Bottom line
Bloomberg reported a benchmarks-versus-desk gap on Gemini 4 coding, attributed to people with access. Google said that underperformance framing is inaccurate, pointed to Kavukcuoglu's encouragement, and published an internal-use story of migrations and specialized coding. Those two accounts do not resolve in a tweet.
They also do not need to. Google's own Argon table already refuses the "coding is solved" shortcut: DeepSWE lead, FrontierSWE and Terminal-bench losses. The builder move is the same as launch day: do not swap production on a vendor DeepSWE cell. Pin your repo evals. Follow @explainx_ai when a public model ID makes those folders runnable.
Related reading
- Gemini 4 Argon vs Opus 5.5 vs Grok 4.7 vs GPT-6 Astra
- Gemini 4 Argon launch: 1M output, $2/$10 intro, Fairwind
- Google fast-tracks Gemini 4 post-training (Sep 24, 2026)
- Gemini 3.8 Flash launch benchmarks
- How to read AI benchmarks
- GPT-6 Astra launch: benchmarks and pricing
- Claude Opus 5.5 launch benchmarks
- Leaked Gemini 4 Pro arena tests vs Opus 5.5
Sources
- Google — Gemini 4 Argon: our next era of frontier intelligence (September 30, 2026)
- Google DeepMind — Fairwind Program
- Bloomberg, Julia Love and Davey Alba, reporting dated about September 30, 2026 (paywalled; quoted paragraph as circulated in English recaps). Google's dispute, the Kavukcuoglu pointer, the abandoned Gemini 3.5 Pro June plan, the split employee views on Fable and Astra, the "large consensus" employee, and the Search / Maps / Gmail / Chrome distribution note follow those recaps. We are not linking the paywalled article.
Anonymous-source claims, Google's denial, Fairwind access rules, and vendor benchmark cells reflect reporting and Google's September 30, 2026 announcement as of October 1, 2026. Independent labs can still move the coding rows. explainx.ai has not run Argon.
