Google DeepMind introduced Gemini 4 Argon for coding, enterprise knowledge work and cyber defense, positioning it against GPT-6 Astra and Claude Opus 5.5 class models. The company claims first place on 13 of 19 published benchmarks and an industry-leading 1M-token output limit, up from 64K. Access starts only with government users and trusted cyber defenders in the Fairwind Program, with wider developer and enterprise access promised as soon as possible. For business, this matters because Google returns to the frontier tier after months without a larger-than-Flash release.

Gemini 4 Argon Brings 1M-Token Output in Limited Cyber Preview

What Argon offers at launch

Standard pricing is set at $4 per 1M input tokens and $20 per 1M output tokens, with a 50% introductory discount to $2 and $10 and no end date announced. Cached input carries a 95% discount. The 1M-token output is enabled through Long Decode Continuation, a new API feature that pauses long responses and resumes them across calls. Measurement differs by evaluator: Vals lists 262K maximum output, while Artificial Analysis reached 1M tokens through that continuation mechanism. Google says it will refine guardrails before opening access to developers, enterprises and consumers.

On claimed results, Argon scores 77.9% on DeepSWE, versus 74.2% for Opus 5.5 and 74.1% for Astra. Artificial Analysis gives it 53 on the Intelligence Index, matching GPT-6 Astra at 53 and edging GPT-6.1 Sol at 52. It ranks first on AutomationBench-AA at 77.5% and scores 57% on Terminal Bench 4, behind Sonnet 5.5, Opus 5.5 and Astra. Vals places it first on the Vals Index at 68.9%, with Terminal-Bench 4.0 rising from 19.0% to 57.6%, plus 70% on CyberBench proof-of-concept tasks and 100% on IOI 2024-2026. In arenas it is first in Text Arena at 1525 and eighth in Code Arena WebDev at 1679.

Argon arrives after Google DeepMind last shipped a larger-than-Flash model in February with 3.1 Pro, followed by incremental 3.x Flash versions and a management shakeup last month. Peers in the meantime launched Fable and Astra class models, leaving timing of a Google response as the open question. Google points to internal deployments as evidence of scale: Argon agents freed more than 300 TiB of data-center memory and are migrating more than 800K lines of C and C++ kernel code to Rust. Agents also replaced 32K lines of SIMD code with safe Rust, making the existing Rust port 2.7x faster with identical output. The team says internal agent loops built on Argon helped complete the CK conjecture.

What this means for enterprise AI use

For companies running coding and knowledge-work agents, the near-term effect is lower price per completed task if discounted pricing holds. Artificial Analysis measures $1.99 per task at discounted pricing versus $3.26 for Astra, while standard pricing would raise Argon to $3.98. The saving comes from price rather than efficiency: Argon averages 62K output tokens per task against Astra's 27K. On the Vals Index the average is $15.68 per task. Small teams gain cheaper access to long outputs for migration and refactoring work, while large organizations can test bulk code conversion where output volume dominates budgets.

The limits shape selection decisions. Availability remains restricted to Fairwind participants, so most enterprises cannot yet test guardrails, latency, or support terms in production. Hallucination on AA-Omniscience is 15% against 51% for Astra, but accuracy is lower at 50% versus Astra's 63%, a tradeoff to verify on domain data. Coding breadth looks strong with 30 perfect Vibe Code Bench apps against 25 for Opus 5 and 24 for Astra, yet some observers questioned published numbers, including DeepSWE and possible preference-data effects. On Harvey's legal benchmark, the reported 19.6% trails Muse Spark 1.2 at 25.42%, so regulated use needs separate validation.

The marker to follow is the move from Fairwind preview to general developer and enterprise access, alongside any end date for the $2 and $10 introductory pricing. If Artificial Analysis and Vals scores hold outside Google's test set and cache behavior proves stable, long-decode workflows for code migration and research synthesis become practical to budget. If access stays narrow or standard pricing returns, the cost advantage narrows quickly.