Google launches Gemini 4 Argon with 1 million-token outputs and a focus on cybersecurity

Google has unveiled Gemini 4 Argon, its new frontier AI model built for long-running tasks across software engineering, enterprise work and cybersecurity.

For now, access is limited. Google is rolling Argon out first to a group of trusted cybersecurity specialists through its Fairwind Program, with developers, businesses and consumers set to follow after additional testing.

One of Argon's most notable upgrades is how long it can keep generating. The model supports outputs of up to 1 million tokens in a single run, compared with 64,000 tokens for previous Gemini models. The idea is to give agents enough room to reason through lengthy problems and execute large sequences of actions without repeatedly breaking the work into separate requests.

Coding and professional work

Google is already using Argon internally for debugging, algorithm design and large-scale changes to its own codebases.

One major use case is migrating Google projects from C and C++ to memory-safe Rust. Argon-powered agents have worked on projects ranging from tens of thousands of lines of code to more than 800,000 lines in Fuchsia's Zircon kernel.

In another example, Argon rewrote roughly 32,000 lines of SIMD code in the libgav1 video decoder. Google says the resulting implementation ran 2.7 times faster than the previous Rust port while producing identical decoding results. The changes were still subjected to automated and human review before deployment.

Google has also put teams of Argon agents to work on its infrastructure. In one experiment, the agents analyzed data-center telemetry and identified memory optimizations that freed more than 300 TB. The company estimates the same approach could ultimately save between 500 TB and 1 PB.

Researchers have used Argon for scientific work as well. In one example involving quantum algorithms, Google says the model reduced the computational resources required by 40% compared with a previously published baseline, completing the optimization in minutes.

Google's benchmarks show Argon performing particularly well on long-running software engineering tasks:

  • DeepSWE v1.1: Argon scored 77.9%, compared with 74.1% for GPT-6 Astra and 74.2% for Claude Opus 5.5.
  • FrontierSWE v2: Argon scored 55%, behind GPT-6 Astra at 65.5%.
  • Terminal-bench 4.0: Argon reached 57.4%, compared with 66.4% for Claude Opus 5.5.

For professional knowledge work, Argon scored 68.9% on Vals Index, which covers areas including finance, law and software engineering. It reached 51.3% on Zapier's AutomationBench for multi-step business workflows and 91.7% on LVBench, a long-video understanding benchmark.

Cybersecurity is where Argon gets more autonomy

Google trained Argon to do more than flag potentially vulnerable code. The model can investigate suspected vulnerabilities, attempt to verify whether they are exploitable and propose fixes, giving it considerably more autonomy than a conventional code-scanning tool. On CWE-bench v1, Argon scored 68%, tying GPT-6 Astra, according to Google.

The company says it tested the model on complex software projects spanning 20 programming languages. In a separate evaluation with Wiz, Argon examined live web systems without access to their source code, searched for possible attack paths and produced evidence to validate the vulnerabilities it found.

Wiz is now using Argon through its Scan for Good initiative, which offers free security testing for critical infrastructure. During one assessment, Google says the model discovered a critical vulnerability in software used by healthcare organizations that previous frontier models had missed.

Those capabilities also create a difficult safety trade-off. Security researchers need a model capable of following an attack far enough to determine whether a vulnerability is real, but the same capabilities could be useful to an attacker. Google is therefore giving trusted specialists in its Fairwind Program access to Argon without the cyber restrictions imposed on general users, while broader releases will retain safeguards against potentially harmful cyber and CBRN requests.

Google says it has also strengthened Argon against prompt-injection attacks, including malicious instructions hidden in documents and other material the model encounters while working. The company is separately testing monitoring systems that can follow an agent’s actions and intervene if it begins operating outside the task assigned by the user.

Argon is starting with a limited rollout

Google is releasing Argon gradually, beginning with trusted cybersecurity specialists through its Fairwind Program. The company is also participating in the U.S. government’s voluntary program that gives officials early access to advanced AI models before broader release.

Developers, businesses and consumers are next, although Google has not given a date for general availability. The wider rollout will begin with paid Gemini API customers and Google AI Ultra subscribers after the company gathers more feedback and completes additional safety testing.

API access will initially cost $2 per million input tokens and $10 per million output tokens, with cached input discounted by 95%. Those rates are introductory: Google says pricing will later rise to $4 per million input tokens and $20 per million output tokens.