I did the same thing with Google’s AI team this summer. Written off. The last top model it shipped was Gemini 3.1 Pro in February, which I went through in “Google Just Doubled AI Reasoning in 90 Days. The Number Speaks for Itself.”, and by September that was over seven months ago. Then, on September 30, Google announced Gemini 4 Argon.
My take up front: the independent scores are real, and I still wouldn’t call this a comeback yet. Argon isn’t open to developers, enterprises, or Google AI Ultra subscribers. It’s rolling out first to trusted cyber defenders through Google’s Fairwind Program, and paid API customers and Ultra subscribers come next, with no date given.
Why Google scrapped Gemini 3.5 Pro
Google announced Gemini 3.5 Pro at its I/O conference in May and promised it for June. June came and went. According to Bloomberg, Google has since abandoned the model, not delayed it. Bloomberg Intelligence analyst Mandeep Singh put the cost of a training run for a model like that at as much as $400 million. That’s an analyst’s ceiling for the category, not a number Google has confirmed.
Walking away from that much money is the part I respect.
Economists point to the Concorde case as an example of the sunk cost fallacy, where continued funding is justified by the large previous investment.
Shipping a 3.5 Pro that couldn’t beat the competition would have been exactly that, with a press release attached. It still cost Google months it didn’t have.
The summer DeepMind lost its stars.
The talent story is the second half of the fiasco. In June, John Jumper, who shared the 2024 chemistry Nobel with Demis Hassabis for AlphaFold, left for Anthropic, and Noam Shazeer, Gemini’s co-lead, left for OpenAI the same week. On August 5, Jeff Dean co-founded a company called Discovery Loop with Oriol Vinyals, Quoc Le, and Sanjay Ghemawat, with Google as a founding investor. The same day, Hassabis stepped down as CEO of Google DeepMind to become chairman and Alphabet’s chief scientist, and Koray Kavukcuoglu took over the daily running of the lab. Axios called it Google’s biggest AI leadership shake-up since OpenAI’s 2023 upheaval.
The competition didn’t wait. OpenAI released GPT-6 Astra on September 3, and Anthropic released Claude Opus 5.5 on September 22, which I covered in “Anthropic’s CEO Said Slow Down. His New Model Beat Everything Anyway.” Three frontier launches in 27 days, and Google had nothing to put next to either of them until the 30th.
What Gemini 4 Argon scores on independent tests
Artificial Analysis, an independent evaluator, gave Argon a 53 on its Intelligence Index, matching GPT-6 Astra, one point ahead of GPT-6.1 Sol and 23 points above Gemini 3.1 Pro Preview’s 30. Claude Opus 5.5 still sits higher at 58. So, Argon is once again among the top group, but it isn’t leading it.

Elsewhere, the picture is better. Argon took first place on Arena’s text leaderboard with 1,525 points, though Arena marked the result as preliminary at 4,942 votes. Across the 18 benchmarks Google disclosed, Argon leads outright on 12, GPT-6 Astra on three, and Claude Opus 5.5 on two. On Artificial Analysis’s AutomationBench, it scores 78%, seven points ahead of Claude Sonnet 5.5. Coding is where it’s mixed. Argon manages 57% on Terminal-Bench 4, behind Sonnet 5.5 at 64%, Opus 5.5 at 60%, and Astra at 59%.
Now, the number I find most interesting. On AA-Omniscience, Argon’s hallucination rate is 15%, against 51% for GPT-6 Astra and 54% for GPT-6.1 Sol. But its accuracy is 50%, 13 points below Astra’s 63%. Argon doesn’t know more. It reduces guessing, and if you’ve ever suffered from a confidently wrong response, that’s a trade-off worth considering. Argon is also the noble gas that reacts with almost nothing, and I can’t tell whether the name was a joke, but the behavior fits.
The price is lower per token and higher per result.
Google’s introductory price is $2 per million input tokens and $10 per million output tokens, rising to $4 and $20 when the promotion ends, and Google hasn’t said when that is. Argon spends 62,000 output tokens per task on Artificial Analysis’s index, against 27,000 for GPT-6 Astra, at a discount that still works out to $1.99 per task against Astra’s $3.26. At standard prices, it’s $3.98, about 1.2 times Astra. A rate card isn’t a bill, and anyone budgeting around the launch price will rebuild the spreadsheet in a month or two.
What Google says Argon already does inside Google
These are Google’s own claims, so read them that way. In its announcement, Google says a team of Argon agents analyzed fleet-wide profiling data and applied memory optimizations that will free up over 300 TiB of memory across its data centers once rolled out, with an estimated 500 TiB to 1 PiB in total savings. Argon beat the published baseline by 40% on the spacetime cost of a quantum computing subroutine, in minutes. And it’s working on migrating C and C++ code to Rust, up to 800,000 lines for the Fuchsia Zircon kernel, with audits planned before anything reaches production.
The result that landed hardest for me is the security one. Google says Wiz, which uses Argon through its Scan for Good program, found:
a critical vulnerability exposing sensitive personal information across healthcare software used by hospitals worldwide, identifying a severe risk that previous frontier models had missed.
The announcement doesn’t name the software. That’s the sentence I’d want filled in before leaning on it.
There’s a counterweight. The same day, Bloomberg reported employee skepticism: Gemini 4 does well on the benchmarks the industry uses but less well when staff put it to work, and it struggles with certain coding tasks, according to people with direct access. Google told Bloomberg it would be inaccurate to say Gemini 4 underperforms in coding.
Pichai signed a voluntary pledge the day before
On September 29, a day before the launch, Pichai was at the White House with Trump and other AI executives signing a voluntary accord that includes internal and external reviews. AP’s headline called it an accord to self-police AI development. Google says it’s working with the US government on pre-release evaluations before wider access, which VentureBeat describes as part of a rollout that starts with cyber defenders.
I like the idea of a gate. I don’t trust a pledge nobody can enforce to be the only thing holding it, and a launch limited to cyber defenders is Google choosing its own gate, which is a distinct thing.
Is Google back?
Partly. It’s back in the top group on independent tests, with the lowest hallucination rate among the leading models and a lead in agentic automation. It isn’t back in the sense that matters most to me, which is a model you and I can pay for. The test comes when paid API customers and Ultra subscribers get access, and whether the coding complaints Bloomberg heard survive real workloads. Until then, I’d file Argon under promising and gated. Three frontier launches in 27 days says the race isn’t slowing down, whatever any accord says, and Google just proved it can still enter. Staying in is the harder part.
If the part about a voluntary pledge and a stop button stuck with you, the piece on an OpenAI model walking out of its sandbox goes deeper, including what happens when the stop button is the slowest thing in the building.
If you’re trying to work out whether Google’s comeback is real or just a leaderboard screenshot, follow Nov Tech, and don’t forget to subscribe because I’ll keep checking the numbers once Argon opens up.


