Cover: AI-generated editorial composition by TMRW, based on Google’s Gemini 4 Argon announcement.
On September 30, Google introduced Gemini 4 Argon and told almost everyone to wait. Koray Kavukcuoglu, who leads Google DeepMind, said the model is going first to a set of trusted cyber defenders through the company’s Fairwind program. Developers, companies, and consumers come “as soon as possible,” starting with paid API customers and Google AI Ultra subscribers. There is no date.
That order is the news. The chart is the advertisement.
The numbers Google is willing to print
Google says Argon sets a mark of 77.9 percent on DeepSWE v1.1, a test of long software work in real codebases, and calls that a new state of the art. It says Argon leads the Vals Index, which weights finance, coding, legal, and tax tasks by their share of the U.S. economy, and ranks first on Zapier’s AutomationBench at 51.3 percent. On LVBench, a long-video test, it claims 91.7 percent. The maximum output rises from 64,000 tokens to one million, so a single run can think for a very long time.
When paying customers are allowed in, the introductory price is $2 per million input tokens and $10 per million output tokens, with cached input 95 percent off. After that window, Google says the price becomes $4 and $20. Those introductory rates match the list price OpenAI posted the same week for GPT-6.1 Sol. The comparison cannot be run. Sol is on the API now. Argon is not.
What Google says it is already doing inside Google
The larger claims are about Google’s own systems, and they are still Google’s account. Kavukcuoglu says Argon agents found memory savings above 300 tebibytes in the company’s data centers, with an estimate of 500 tebibytes to one pebibyte if the work is fully applied. He says agents are helping rewrite C and C++ into Rust, including more than 800,000 lines of the Fuchsia Zircon kernel, and that those rewrites stay in automated and manual review before production. On libgav1, Google’s open-source video decoder, he says agents took an existing Rust port and made it 2.7 times faster, with the same video out. A quantum-computing example beat a published baseline by 40 percent in minutes.
None of that is a paper a reader can rerun this week. A rewrite that has not shipped is not a migration you can copy. It is evidence that Google trusts the model on its own code, under its own review. That is interesting. It is not a customer result.
Defenders first, on purpose
Google says it trained Argon to find, check, and patch software flaws, and that Fairwind partners and Google’s own teams will get the model without the cyber limits placed on everyone else. Through a Wiz program called Scan for Good, Google says an early run found a critical flaw in healthcare software used by hospitals, exposing personal information that earlier frontier models had missed. Google does not name the software or the bug. On CWE-bench v1, a test of fixing known classes of vulnerability, it says Argon ties for first at 68 percent.
The company also says it is in the U.S. government’s voluntary process for looking at a model before wide release, and that it is still hardening refusals, defenses against hidden instructions, and monitors that stop a run when the model steps past what the user asked. Holding a system with that profile is a coherent choice. It also means the model Google wants judged as its return to the front of the pack cannot yet be judged by the people who would use it for ordinary work.
What to do with a chart you cannot call
Until there is a public endpoint, treat the leaderboard as Google’s description of Google’s tests. A team picking a model this month can call GPT-6.1 Sol, Claude, and the Gemini models already on the API. Argon is a name on a waiting list. A score you cannot buy is not a reason to pause a project, and it is not a reason to ignore Google once the gate opens. Ask then for the same ten tasks, the same dollar cap, and a date. “As soon as possible” is not one.



