- Google has unveiled Gemini 4 Argon ahead of a broader rollout, positioning the flagship model for complex programming, enterprise and cybersecurity tasks.
- The model expands its context limit to 1 million tokens from 64,000 in the previous version.
- Google plans to initially offer access to paid API customers and Google AI Ultra subscribers.
- Gemini 4 Argon will cost $2 per 1 million input tokens and $10 per 1 million output tokens at launch.
Google has unveiled Gemini 4 Argon, a new flagship model designed for complex, multi-step tasks in programming, enterprise workflows and cybersecurity. Ahead of a broader rollout, Google is testing the model with a limited group of cyber defenders through its Fairwind Program while continuing to validate its safety mechanisms.
One of Argon’s main features is a 1 million-token context limit, up from 64,000 in the previous version. At launch, the model will cost $2 per 1 million input tokens and $10 per 1 million output tokens, with initial access planned for paid API customers and Google AI Ultra subscribers.
Google details internal development work
Google said it is already using Gemini 4 Argon internally for programming, research and work involving large codebases. According to the company, the model helped optimize a quantum algorithm and beat the published baseline result by 40%.
Argon also helped free more than 300 TiB of memory after analyzing data-center telemetry, Google said. The company estimated total potential savings of 500 TiB to 1 PiB.
Google said the model helped scale the migration of C and C++ codebases to Rust, including work on the Fuchsia Zircon kernel, which contains more than 800,000 lines of code.
In a separate example involving the open-source libgav1 project, Argon agents replaced 32,000 lines of SIMD code and produced a Rust version that runs 2.7 times faster than the previous Rust port while delivering identical video-decoding output, according to Google.
On DeepSWE v1.1, a benchmark evaluating performance on real, long-running programming tasks, Argon scored 77.9%. It scored 51.3% on AutomationBench, which tests end-to-end execution of business tasks.
Google said Argon is not yet broadly available. The company plans to expand access gradually as it gathers feedback from testers and strengthens the model’s safety mechanisms.
Cybersecurity testing and safeguards
Google said Gemini 4 Argon can autonomously identify, validate and patch critical software vulnerabilities. Wiz is already using the model through the Scan for Good initiative to protect critical infrastructure.
During one test, Argon discovered a critical vulnerability in medical software that could have exposed patients’ personal data at hospitals worldwide, according to Google.
Argon scored 68% on CWE-bench v1, which evaluates models’ ability to remediate vulnerabilities, tying for first place with another model.
Ahead of a large-scale rollout, Google is strengthening safeguards against abuse, prompt-injection attacks and failures to align the model’s actions with user intent. The company said it uses automated and manual red-team testing, monitoring of the model’s internal processes and isolated environments for high-risk tests.
A week earlier, Anthropic unveiled the cheaper and faster Claude Opus 5.5 model, while OpenAI introduced GPT-6 Sol and Luna with lower usage costs.
Source: Incrypted
