Google announces Gemini 4 Argon, with initial access for selected cyber defenders
Google says its new model can handle extended coding, knowledge-work and cyber defense tasks. Wider access is planned after further testing, but the company has given no public release date.
Google announced Gemini 4 Argon on September 30, 2026, saying the model is rolling out first to selected cyber defenders through its Fairwind Program. The company presents it as a tool for complex software engineering, professional knowledge work and cyber defense. Developers, businesses and consumers must wait for a broader release while Google gathers feedback and tests safeguards.
The announcement establishes an initial, restricted rollout rather than general availability. Google says paid API customers and Google AI Ultra subscribers will be among the first groups offered wider access. It has not given a date for that step or for a public release, leaving prospective users without a timetable for trying the model themselves.
Who can use Gemini 4 Argon now?
Google says trusted cyber defenders are receiving initial access through Fairwind. It says those defenders and its own internal teams will be able to use Argon without the cyber guardrails intended for broader access, so they can use its full defensive capabilities. That is a description of the company’s controlled rollout, not a statement that the model is already available to all security teams.
Google says Wiz is using Argon through its Scan for Good initiative, which seeks to identify and remediate exposures in public infrastructure. The company also says the model found a critical vulnerability affecting healthcare software used by hospitals worldwide. Its announcement does not identify the software or provide enough detail to independently assess that example, so the reported discovery remains a Google account of an early use case.
What Google’s coding and business benchmarks show
Google reports that Argon scored 77.9% on DeepSWE v1.1, a test it describes as measuring extended, real-world software engineering tasks. It also reports a 51.3% score on AutomationBench, which evaluates the completion of business workflows, and 91.7% on LVBench, a long-video understanding test. These are company-published results on particular evaluations; they do not show how reliably a customer’s own project will work.
For vulnerability repair, Google says Argon scored 68% on CWE-bench v1 and tied for first place. The result is relevant to its proposed cyber defense role, but it measures performance under that benchmark’s conditions. The announcement does not establish that Argon can find and fix every kind of vulnerability in a live system, or that its performance on the test transfers unchanged to a defender’s environment.
The company says it has expanded Argon’s output limit to one million tokens from 64,000. That specification gives the model room to produce much longer responses or work through extended tasks. It is a capacity claim, separate from evidence that a long answer is accurate, useful or safe throughout. Google connects the larger limit to its aim of sustaining multi-step work in coding and other professional settings.
What Google says it has used Argon for internally
Google gives examples from its own operations. It says Argon agents analysed data-centre telemetry and identified memory optimisations that would free more than 300 TiB once rolled out, with estimated total savings of 500 TiB to 1 PiB. Those figures describe a company-reported project and forecast savings; the announcement does not establish that all of the projected savings have already been realised.
In another example, Google says Argon agents replaced 32,000 lines of SIMD code in a Rust port of its libgav1 video decoder. It reports identical video output and a decoder 2.7 times faster than that Rust port. Google says larger migrations of C and C++ code to Rust are still undergoing automated and manual auditing, emulation tests and review before production rollout. Those checks matter because a benchmark result or a working port alone does not settle whether a large code change is ready for deployment.
Why benchmark scores need context
Independent research offers a useful limit on what software-task results mean. METR’s time-horizon research measures AI agents on more than a hundred software tasks and explains that its tasks are largely self-contained, well specified and automatically scored. METR cautions that professionals’ everyday work often depends on prior context and has less tidy success criteria. METR has not presented those findings as an evaluation of Argon; they explain why a score on a defined task set should not be read as a general measure of workplace performance.
Google says it is gathering feedback from early testers and strengthening safeguards before wider access. Its announcement describes work on resistance to misuse and prompt injection, monitoring for actions outside a user’s intentions, and more secure testing environments. Readers can assess the published claims and the scope of the restricted rollout now. Broader availability and evidence from independent, real-world use remain open questions.
Sources and context
- Gemini 4 Argon: our next era of frontier intelligenceGoogle
- Task-Completion Time Horizons of Frontier AI ModelsMETR
AI-assisted article checked against the listed sources. NewsJaws did not conduct interviews or attend the reported events.
About NewsJaws Desk
AI-assisted reporting and explainers reviewed against the linked source documents. No claim of on-scene reporting or original interviews.