Google rolls out Gemini 4 Argon with 1M-token output limit and cybersecurity capabilities

Google has announced Gemini 4 Argon, its new frontier model built to sustain deep reasoning across complex, long-horizon workflows. The model covers coding, enterprise knowledge work, multimodality, cybersecurity defense, and creative writing, with Google highlighting its use in internal engineering workflows and security research.

Gemini 4 Argon for coding and research

Thousands of Googlers are already using Gemini 4 Argon for specialized coding tasks, deeper research, and writing quality. Google engineers are also using the model daily for debugging, large-scale codebase migrations, and algorithm design.

Google highlighted several internal applications:

  • Quantum algorithmic optimization: Argon is helping quantum computing researchers optimize the spacetime resources, measured as qubits × gates, of subroutines that bottleneck important applications. In one example, it beat a published baseline by 40% within minutes.
  • Memory efficiency: Argon agents analyzed fleet-wide profiling telemetry to autonomously identify and apply memory optimizations across Google data centers. Once rolled out, the changes freed more than 300 TiB of memory, with estimated total savings of 500 TiB to 1 PiB.
  • Large-scale codebase migrations: Argon agents are helping migrate C/C++ codebases to Rust across Google, ranging from tens of thousands of lines in core libraries such as re2 and libgav1 to more than 800,000 lines in the Fuchsia Zircon kernel.

Because many of these systems are critical, Google said the rewrites undergo rigorous automated and manual auditing, emulation testing, and review before production deployment.

For libgav1, Google’s open-source software for video decoding, Argon agents took an existing Rust port and replaced 32,000 lines of SIMD code through many rounds of profile-guided experiments. They studied compiler output and produced safe Rust so the compiler could automatically vectorize the code.

Google said the resulting memory-safe video decoder is 2.7 times faster than the Rust port while producing identical video output, bringing it closer to the optimized C++ version.

1M-token output limit and benchmark results

Gemini 4 Argon increases the output token limit to 1 million tokens, up from the previous 64,000-token limit. Google said the model can think deeply and generate hundreds of thousands of tokens in a single trajectory, giving it more room to work through complex problems.

Google reported the following performance results:

  • DeepSWE v1.1: 77.9%, measuring performance on real-world, long-horizon software engineering tasks.
  • AutomationBench: 51.3%, measuring end-to-end execution across core business functions. The benchmark is from Zapier.
  • LVBench: 91.7%, measuring long-video understanding.

Google also reported results across several enterprise and domain-specific evaluations:

  • Vals Index: Measures economic impact across finance, coding, legal, and tax work, with each sector weighted according to its contribution to U.S. GDP.
  • Vals Finance Agent v2: Evaluates multi-step financial research.
  • Harvey’s Legal Agent Benchmark: Evaluates legal research and drafting.

Gemini 4 Argon also supports multimodality for knowledge work. Google said it can perform professional chart analysis, identify details from long videos, and take action based on a series of documents.

Gemini 4 Argon for cybersecurity defense

Google has trained Gemini 4 Argon for cybersecurity defense, including the ability to autonomously find, validate, and patch critical software vulnerabilities. Trusted cyber defenders and Google’s internal teams will receive Argon without cyber guardrails so they can use its cybersecurity capabilities for defensive work.

Wiz is using Argon through Scan for Good, a program dedicated to protecting critical public infrastructure for free by finding and remediating high-risk exposures. In an early demonstration, Google said Argon uncovered a critical vulnerability exposing sensitive personal information across healthcare software used by hospitals worldwide. The company said previous frontier models had missed the vulnerability.

On CWE-bench v1, which evaluates remediation of security vulnerabilities, Argon tied for first with a 68% score. Google said this builds on 3.8 Flash Cyber’s performance on CWE-bench v0.

Google also reported the following vulnerability-discovery results:

  • Google internal vulnerability benchmark: Argon uncovered a wide range of exposures across complex codebases spanning 20 programming languages.
  • Wiz black-box penetration testing benchmark: Argon outperformed 3.8 Flash Cyber in discovering the attack surface, identifying vulnerabilities, and producing proof-of-concept evidence to validate them.

The Wiz benchmark tests a model’s ability to analyze live web systems without source code.

Safeguards before broader availability

Google is strengthening Gemini 4 Argon’s safeguards across four areas: misuse prevention, prompt injection protection, misalignment monitoring, and system security.

Misuse prevention

Argon is designed to refuse harmful requests involving cyber or chemical, biological, radiological, and nuclear (CBRN) attacks while preserving legitimate dual-use scientific research, in accordance with Google’s Frontier Safety Framework.

Google is also strengthening monitoring of model internal activations to identify potential misuse. Internal and external red teams have tested the safeguards using manual and automated attack methods.

Prompt injection protection

Google said Argon is its most resilient model yet against indirect prompt injections, where malicious instructions or context attempt to hijack a model’s behavior.

The company said these attacks require multiple layers of defense and continued monitoring. Google used automated red teaming and adversarial training to strengthen the model and reported that Argon led on Gray Swan’s Indirect Prompt Injection (IPI) benchmark.

Misalignment monitoring

Google is deploying mitigations that monitor Argon’s chain-of-thought and actions and can stop execution when necessary to prevent the model from going beyond a user’s intentions.

A similar system monitored training runs and sent alerts to a dedicated incident response team. Google said it took precautions against feeding those findings back into training so the model’s reasoning would not be shaped to evade monitoring.

The company also encouraged the industry to preserve reasoning transparency so model thoughts can remain useful for identifying and diagnosing misalignment.

System security

Google is hardening sandboxed environments by isolating and sealing them before high-risk training and evaluations, in line with its agent control roadmap. The company also plans to share agent security best practices with partners to improve security across the ecosystem.

Gemini 4 Argon pricing and availability

Gemini 4 Argon is initially rolling out to a set of trusted cyber defenders through Google’s Fairwind Program. Google is also participating in the U.S. government’s voluntary process for pre-release model access while gradually expanding access.

The company will gather feedback from early testers and continue iterating on its guardrails before broader availability to developers, enterprises, and consumers. Broader access will start with paid API customers and Google AI Ultra subscribers.

The introductory pricing is:

  • Input: $2 per 1 million tokens
  • Output: $10 per 1 million tokens
  • Cached input: 95% off the input token price

After the introductory period, pricing will be $4 per 1 million input tokens and $20 per 1 million output tokens.


Related Post