Google DeepMind launches Gemini 4 Argon with 1 million output tokens
Google DeepMind has unveiled Gemini 4 Argon, the first model in the fourth generation of Gemini, betting heavily on unprecedented output length: 1 million tokens in a single response, 16 times the previous 64 thousand-token ceiling. The model targets three primary use cases — long-horizon software engineering, enterprise knowledge work in legal and finance, and cyber defense — and is being released through a phased rollout that includes a voluntary U.S. government early-access process before broad availability.
1-million-token output window and aggressive pricing
The jump to 1 million output tokens puts Argon well ahead of immediate rivals: Claude Opus 5.5, Claude Fable 5.1 and GPT-6 Astra all stop at 128 thousand tokens. Google claims the ability to "think deeply" and emit hundreds of thousands of tokens in one run enables large-scale refactoring or long reports without splitting work across multiple turns. Launch pricing stands at $2 per million input tokens and $10 per million output tokens; cached input tokens receive a 95% discount, dropping to 10 cents per million (about 36 agorot). After the launch period, prices double to $4 and $20 respectively. Logan Kilpatrick confirmed the figures. Google has not yet disclosed the model's input-window size.
Benchmarks: leads in long-form engineering and legal, lags in computer use
In internal comparisons against the three competitors above, Argon leads on 12 of 18 tests and ties for first on one more. Standout wins: DeepSWE v1.1 (long-horizon software engineering) at 77.9%, a new SOTA result versus 74.2% for Opus 5.5 and 74.1% for GPT-6 Astra; Vals Index (economic impact across finance, code, legal and tax) at 68.9%; AutomationBench (end-to-end business execution in Zapier) at 51.3% versus Opus 5.5's 42.5%; Harvey Legal Agent Benchmark at 19.6% versus GPT-6 Astra's 5.4%; and LVBench (long-video understanding) at 91.7%, also a new SOTA. Conversely, Argon trails on FrontierSWE v2 (55% versus GPT-6 Astra's 65.5%), Terminal-Bench 4.0 (57.4% versus Opus 5.5's 66.4%) and OSWorld-2.0 (computer use, 69.2% versus 72.6%). An Artificial Analysis review found Argon matches GPT-6 Astra on their intelligence metric at 60% of the cost per task, using discounted prices.
Autonomous cyber defense and petabyte-scale internal use
Google trained Argon to autonomously discover, verify and remediate critical vulnerabilities; authorized defenders and internal teams receive a version without cyber guardrails. On CWE-bench v1, which tests vulnerability remediation, the model ties for first at 68%, while competitors run inside their own agent harnesses. Wiz is already using the model under its Scan for Good initiative and reported it uncovered a critical vulnerability in healthcare software deployed in hospitals worldwide that previous frontier models missed. Ahead of broad release, Google is hardening defenses in four areas: cyber and CBRN misuse protections including deployment monitoring under the Frontier Safety Framework; indirect prompt-injection resistance (leading on Gray Swan's IPI benchmark); chain-of-thought misalignment monitoring with the ability to halt runs; and signed, isolated sandboxes for high-risk training and evaluation. Inside Google, thousands of employees are already running Argon: agents have implemented memory optimizations in data centers and freed more than 300 tebibytes (TiB), with a projection of 500 TiB to 1 pebibyte (PiB); other agents replaced 32 thousand lines of code, a figure that was originally redacted.