Google has released Gemini 4 Argon, its new frontier AI model. The company announced it on Wednesday, September 30, and says it is its most capable model for coding, office work and cyber defense.
The launch came one week after Anthropic released Claude Opus 5.5 and one day after OpenAI released GPT 6.1 Sol.
Gemini 4 Argon Benchmark Scores
On DeepSWE v1.1, a test of long, real-world software engineering tasks, Argon scored 77.9%. Claude Opus 5.5 scored 74.2%, GPT-6 Astra scored 74.1% and Claude Fable 5.1 scored 67.4%.
Google calculated its own DeepSWE score. Rival scores came from a public leaderboard and company reports.
In Google's own comparison table, Argon leads on 12 of 18 benchmarks. It ties on one and trails on five.
Argon ranked first on Zapier's AutomationBench with 51.3% and scored 91.7% on LVBench, which tests long video understanding. Google says it also leads the Vals Index, which measures work in finance, coding, legal and tax.
The model can now write up to 1 million tokens in one reply, up from 64,000. That equals about 750,000 words, compared with about 48,000 before.
Cybersecurity and Prompt Injection Results
Google trained Argon to find, check and fix software security flaws. On Gray Swan's Indirect Prompt Injection benchmark, which tests whether hidden instructions can hijack an AI, Argon had a 0.7% attack success rate. Lower is better.
Claude Opus 5.5 and Claude Fable 5.1 both scored 1.0%. GPT-6 Astra scored 8.5%, while Grok 4.6 and Kimi K3 were tricked just over half the time.
On the Wiz Penetration Test Benchmark, Argon solved 70.9% of tasks on the first try. Gemini 3.8 Flash Cyber, the restricted model Google launched in September, scored 58.2%.
Google says Argon helped security firm Wiz find a critical flaw in healthcare software used by hospitals worldwide. Earlier frontier models had missed it.
Argon is first going to vetted security teams through the Fairwind Program. The program launched September 2 with more than 650 partners, including governments and critical infrastructure operators.
These users get the model without cyber guardrails, the built-in refusals that normally block help with hacking. Google says a phased rollout is needed because the same skills could be misused.
The company is taking part in the U.S. government's voluntary process for pre-release model access. Argon launched the same day President Trump unveiled a voluntary AI accord that Google's leadership signed.
Google says it is also monitoring the model's reasoning and actions to stop it from going beyond what users intend.
The launch follows a rough summer for Google. In July, it shipped smaller Flash models but skipped the promised Gemini 3.5 Pro, and Alphabet shares fell about 4.4%.
Wider access will start with paid API customers and Google AI Ultra subscribers. Google has not given a date.
Introductory pricing is $2 per million input tokens and $10 per million output tokens. Standard rates will be $4 and $20, though Google has not said when the introductory period ends.