Google began rolling out its frontier model Gemini 4 Argon to vetted cybersecurity defenders, holding back general developer access. Outside Google, only members of the Fairwind Program can use it, a group that now includes more than 650 organizations. Google says Argon leads rival models on 12 of 18 internal benchmarks. The move matters because the strongest coding and automation model is entering business processes through security teams first.
Phased rollout and benchmark results
Fairwind opened on Sept. 3 with the smaller Gemini 3.8 Flash Cyber model and has since added CrowdStrike Holdings and Palo Alto Networks among its members. Google chief AI architect Koray Kavukcuoglu said capabilities at this level require a phased approach. The company is participating in the U. S. government voluntary process for pre-release model access. It also says safeguards against misuse for cyberattacks or weapons development are still being hardened.
Google describes Argon as built for long-horizon work, with an output limit of 1 million tokens against 64,000 for earlier Gemini models. On DeepSWE v1.1 for long-horizon software engineering, Argon scored 77.9%, ahead of Anthropic Claude Opus 5.5 at 74.2% and OpenAI GPT-6 Astra just behind it. On AutomationBench for end-to-end business work, Argon reached 51.3% against 42.5% for Opus 5.5. Terminal-Bench 4.0 remains with Opus 5.5, while FrontierSWE v2 stayed among the few benchmarks kept by Astra.
For defense against manipulation, Google says Argon resists indirect prompt injection better than any model it has shipped. Separate monitors observe the chain of thought and actions and can stop the model if it moves beyond user intent. Fairwind members and internal Google teams receive a version with cyber guardrails removed, able to find and fix vulnerabilities autonomously. On CWE-bench v1 for remediation, that version tied GPT-6 Astra at 68%. Paying developers and Google AI Ultra subscribers are next in line, with no date given.
What this means for business use of AI
Internal deployment shows where large engineering organizations may apply such models first: migration and optimization of existing code. Google teams use Argon agents to move C and C++ code to Rust, from core libraries of tens of thousands of lines to the 800,000-line Zircon kernel in Fuchsia. On the libgav1 video decoder, agents rewrote 32,000 lines of speed-critical code into safe Rust, making it run 2.7 times faster than an earlier Rust port with unchanged output. A fleetwide memory review freed more than 300 tebibytes across data centers.
Pricing creates a temporary window for early planning of agent workloads. At launch Argon will cost $2 per million input tokens and $10 per million output tokens, with cached input priced 95% lower. Anthropic charges $4 and $20 for Opus 5.5, and Google says Argon will move to those rates when launch pricing ends. Small companies can test long outputs and cached context while rates are low, while large firms can compare vendor costs for sustained automation. The difference will show most clearly in projects with repeated runs over large codebases.
The security-first release also signals limits that buyers should verify before building around the model. The Wiz Scan for Good program found a critical flaw exposing personal information in hospital software that earlier frontier models had missed, but that result does not prove similar detection rates elsewhere. Questions to ask include which guardrails are removed for defenders, what monitoring remains, and when the general version will differ. It is also worth checking how benchmark leads in DeepSWE and AutomationBench translate into specific workflows, rather than assuming broad superiority.
