Google announced Gemini 4 Argon, its next frontier model, and it is not a general release. The company is rolling it out to a set of trusted cyber defenders through its Fairwind Program, and describes the staged approach as deliberate: Google says it is working through the US government’s voluntary process for pre-release model access, gathering feedback from early testers and iterating on guardrails before the model reaches developers, enterprises and consumers.
The company’s own description of what the model is for comes from Koray Kavukcuoglu, the Google DeepMind senior vice president who also holds the chief AI architect title: frontier performance in complex workflows across software engineering, enterprise knowledge work like legal and finance, and cybersecurity defence. It is aimed at long-horizon work — the kind that runs for many steps rather than answering a single question — and one concrete change supports that. The output limit rises to a million tokens, up from 64K, which lets the model hold a much longer working trajectory in a single pass.
The performance numbers are Google’s own, published alongside the announcement. Gemini 4 Argon scores 77.9% on DeepSWE v1.1, ahead of the 74.2% and 74.1% the company reports for Claude Opus 5.5 and GPT-6 Astra. It takes 51.3% on Zapier’s AutomationBench, 91.7% on LVBench for long-video understanding, and ties for first at 68% on CWE-bench v1, which measures how well a model remediates security vulnerabilities. Pricing is set at an introductory $2 per million input tokens and $10 per million output tokens, with cached input billed at 95% off.
The most concrete part of the announcement is what Google says the model already does inside the company. Thousands of Googlers are using it, and Google cites three worked examples. Quantum researchers have used it to optimise the qubit-and-gate resources of subroutines that bottleneck important workloads, where it beat a published baseline by 40% in minutes. A fleet of Argon agents read profiling telemetry across Google’s data centres, applied memory optimisations, and freed more than 300 TiB of memory, with the company estimating 500 TiB to 1 PiB in total savings once rolled out. And agents are migrating C and C++ code to Rust — from tens of thousands of lines in core libraries up to more than 800,000 lines for the Fuchsia OS Zircon kernel, under automated and manual auditing before anything reaches production.
On safety, Google is giving trusted defenders the model without cyber guardrails so they can use its full defensive capability, while hardening the rest in parallel. The company says it is improving ways to monitor the model’s internal activations to catch misuse, claims its most resilient model yet against indirect prompt injection, is sealing and isolating sandboxed environments before high-risk training or evaluations, and is deploying misalignment mitigations that watch the model’s reasoning and actions and stop execution when needed.
The timing is pointed. The launch landed a day after OpenAI’s DevDay, where OpenAI shipped an agent product and a new model while shelving another planned release over safety concerns. Google’s answer is that a phased, defender-first rollout is how a model this capable gets shipped at all.