Gemini 4 Argon: pricing, benchmarks and rollout limits

In this article
Google announced Gemini 4 Argon, its newest frontier model, initially for trusted cyber defenders via the Fairwind Program. What its 1M output token limit, $2/$10 pricing and benchmark claims mean for teams comparing frontier models for coding and agentic workloads.
Google announced Gemini 4 Argon, a frontier model for long-horizon coding, enterprise work and cyber defense. Access starts with trusted defenders via the Fairwind Program. Pricing is $2 per million input and $10 per million output tokens, cached input 95% off; the output limit is 1M tokens. Google claims leading scores on DeepSWE and AutomationBench, but the public Vals Index still tops with Sonnet 5.5 and Opus 5.5. Treat claims as pending until independent results land.
Cover: original artwork. Gemini icon: Source: Google (Google Image Library). Google and Gemini are trademarks of Google LLC. This article is independent and is not affiliated with or endorsed by Google.
Teams that benchmark frontier models for coding and agentic workloads have been working through a crowded comparison cycle, with the Vals Index currently showing Claude Sonnet 5.5 and Claude Opus 5.5 effectively tied at the top. On September 30, 2026, Google added a new variable to that exercise: Gemini 4 Argon, announced on the company blog as its newest frontier model, with claimed state-of-the-art results in long-horizon software engineering, enterprise knowledge work and defensive cybersecurity. The complication for evaluation teams is that Argon is not generally available yet, so most of what can be planned today rests on announced specifications, prices and vendor-reported scores.

The launch drew immediate developer attention: the announcement collected 1,475 points and 981 comments on Hacker News. The interest is unsurprising, because the claimed strengths overlap with the workloads many teams are already benchmarking: multi-step coding tasks, autonomous agents and security operations.
What was announced, and who gets access first
According to the announcement post, Gemini 4 Argon is built to sustain deep reasoning across complex, long-horizon workflows, and it is rolling out first to a set of trusted cyber defenders through the Fairwind Program. Google says it is taking part in the U.S. government's voluntary process for pre-release model access, will gather feedback from early testers and will iterate on guardrails before making Argon available to developers, enterprises and consumers. No general-availability date was given.
Two specifications stand out for engineering teams. The output token limit has been expanded to 1 million tokens, up from 64K, which Google frames as headroom for thinking deeply and solving tough problems in a single trajectory. Argon also launches at an introductory price of $2 per million input tokens and $10 per million output tokens, with cached input tokens at 95% off the input price, which works out to about $0.10 per million cached tokens. For agent loops that reuse a large stable context, that cached-input rate may matter as much as the headline prices.
Benchmark claims versus the leaderboards teams use
Google reports a new state of the art on DeepSWE v1.1, which measures real-world long-horizon software engineering tasks, with a score of 77.9%. It also places Argon first on Zapier's AutomationBench at 51.3%, a benchmark of end-to-end execution across core business functions, and reports leading results on Vals Finance Agent v2 and Harvey's Legal Agent Benchmark. On LVBench, which measures long video understanding, Google reports a state-of-the-art 91.7%.
These claims land in a market where independent standings are closely watched. The public Vals Index leaderboard currently lists Claude Sonnet 5.5 at 67.04% and Claude Opus 5.5 at 66.97%, a 0.07-point gap well inside the benchmark's standard error of about ±0.9. Sonnet 5.5 reaches that score at $21.34 per test against $32.14 for Opus 5.5, and the two split the coding components: Sonnet 5.5 leads Vibe Code Bench and Code Migration, while Opus 5.5 leads Terminal-Bench 4.0. Argon does not appear among the models in the published standings, so Google's claim that it now leads the index should be treated as pending until the leaderboard itself is updated.
Methodology matters when comparing these numbers. The Vals Index aggregates six private and two public benchmarks across finance, coding, legal and tax, weighting each sector by its share of U.S. GDP, about 8.0%, 5.6%, 1.2% and 0.5% respectively, based on Bureau of Economic Analysis data. Vals also discloses that some component scores for Sonnet 5.5, Opus 5.5 and Claude Fable 5.1 include tasks served by a fallback model after provider refusals, and that counting those as failures would lower their overall results. Vals positions the index as a signal into the trade-offs between capability, latency and cost, so once Argon appears on the leaderboard, its cost-per-test figure there is the number to compare against, rather than raw token prices.
What Google's internal results suggest about agentic coding
Inside Google, Argon is already used by thousands of employees for coding, research and writing tasks. The announcement describes C and C++ to Rust migrations running from tens of thousands of lines in libraries like re2 and libgav1 up to more than 800,000 lines in the Fuchsia Zircon kernel, with automated and manual auditing, emulation testing and review before production rollout. On libgav1, Google says Argon agents replaced 32,000 lines of SIMD code through profile-guided experiments, producing safe Rust that the compiler vectorizes automatically: a memory-safe video decoder running 2.7x faster than the existing Rust port with identical output.
Two other internal examples point at infrastructure and research workloads. A team of Argon agents analyzed fleet-wide profiling telemetry and applied memory optimizations across Google's data centers, freeing over 300 TiB once rolled out, with estimated total savings of 500 TiB to 1 PiB. In quantum computing research, the model beat a published baseline by 40% on spacetime resource optimization for bottlenecking subroutines, in a matter of minutes. These are vendor-reported internal results rather than independently reproducible benchmarks, but they target exactly the categories, code migration, performance optimization and long autonomous runs, that many evaluation programs are now testing.

Cybersecurity is the first real deployment channel
Google trained Argon for defensive cybersecurity and says it can autonomously find, validate and patch critical software vulnerabilities. On CWE-bench v1, which measures vulnerability remediation, Google reports a tie for first place at 68%, building on the earlier frontier performance of 3.8 Flash Cyber on CWE-bench v0. The company also says Argon uncovered a wide range of exposures across codebases spanning 20 programming languages on an internal benchmark, and outperforms 3.8 Flash Cyber on Wiz's internal black-box penetration testing benchmark, which tests analysis of live web systems without source code.
The first external proof point runs through Wiz. Google says Wiz is using Argon for cyber defense in its Scan for Good initiative, a program that maps the internet-facing surface of critical infrastructure, public services, healthcare organizations and nonprofits, then validates high-impact findings before private disclosure. The program page counts 17,761 organization-linked domains, 326,891 monitored endpoints and 475 critical exposures found. Google cites an early result: Argon uncovered a critical vulnerability exposing sensitive personal information across healthcare software used by hospitals worldwide, a risk previous frontier models had missed.
One boundary deserves attention from security teams: Google will release Argon without cyber guardrails to trusted defenders and its own internal teams so they can use its full cybersecurity capabilities. That unrestricted variant is not part of the general rollout; the stated recipients are trusted defenders, who get access through the Fairwind Program, and Google's own internal teams.
Safeguards and the phased release
Before broad availability, Google says it is strengthening frontier safeguards across four main areas. The announcement details defenses against misuse, with the model designed to refuse harmful cyber and CBRN requests while preserving legitimate dual-use research under its Frontier Safety Framework, plus improved monitoring of the model's internal activations to spot misuse, tested by internal and external red teams with manual and automated methods. It also describes Argon as Google's most resilient model yet against indirect prompt injections, leading on Gray Swan's Indirect Prompt Injection benchmark after automated red teaming and adversarial training, and mentions monitoring for misalignment.
For teams building agentic systems, the prompt-injection claim is the most directly relevant, since agents that read documents, browse the web or operate tools are the main targets of such attacks. As with the capability benchmarks, robustness scores reported by the model's own vendor are a starting point, not a substitute for testing against your attack surface.
What benchmarking teams can do now
Until API access opens, evaluation teams can prepare rather than wait:
Treat announced scores as claims to reproduce: keep your current matrix, including Sonnet 5.5 and Opus 5.5 baselines, and add Argon as soon as developer access allows like-for-like runs.
Re-baseline long-horizon tasks: a 1 million token output limit makes single-trajectory runs feasible for work that previously required multi-call orchestration, so harnesses, context strategies and success criteria may need updating.
Model costs for agent patterns: with cached input at 95% off, loops that reuse large stable contexts change the economics, so price scenarios against the $2 and $10 per million token list prices.
Watch the Vals Index for Argon's entry: its GDP-weighted, cost-aware methodology is built for deployment decisions and differs from pure coding leaderboards.
If you evaluate security use cases, track Fairwind and Scan for Good outcomes, and note that the guardrail-free variant remains restricted to trusted defenders and Google's internal teams.
The practical picture: Gemini 4 Argon arrives with real deployment evidence behind it, but for most teams its relevance is prospective. The specifications that matter for planning, the 1M output window, the token prices and the defensive security focus, are public; the independent scores that would justify a switch are not. Teams in the middle of the current frontier comparison cycle should log the claims, adjust their evaluation plans, and wait for access before rewriting conclusions.
Key takeaways
- Gemini 4 Argon is announced but not generally available: access starts with trusted cyber defenders via the Fairwind Program, with no public GA date.
- The output token limit grows to 1 million tokens, up from 64K, enabling single-trajectory runs for long-horizon coding and agent tasks.
- Introductory pricing is $2 per million input tokens and $10 per million output tokens, with cached input at 95% off, about $0.10 per million.
- Google claims state-of-the-art results on DeepSWE v1.1 (77.9%), AutomationBench (51.3%) and LVBench (91.7%), plus a tie for first on CWE-bench v1 (68%).
- The public Vals Index still lists Claude Sonnet 5.5 (67.04%) and Opus 5.5 (66.97%) at the top; Argon's claimed lead there awaits independent confirmation.
- Internal Google results include C/C++ to Rust migrations up to 800K+ lines and a libgav1 decoder 2.7x faster than the Rust port; these are vendor-reported, not independent benchmarks.
- A guardrail-free variant with full cybersecurity capability is restricted to trusted defenders and Google internal teams.
- For eval teams, the immediate actions are planning around the 1M output window, cached-input pricing and adding Argon to matrices once API access opens.
Frequently asked questions
What is Gemini 4 Argon?
A frontier model announced by Google on September 30, 2026, built for deep reasoning across long-horizon workflows in software engineering, enterprise knowledge work such as legal and finance, and defensive cybersecurity.
Can developers use Gemini 4 Argon today?
Not broadly. It is rolling out to a set of trusted cyber defenders through the Fairwind Program, and Google says it will expand access to developers, enterprises and consumers after gathering feedback and iterating on guardrails. No general-availability date was given.
How much does Gemini 4 Argon cost?
Google announced an introductory price of $2 per million input tokens and $10 per million output tokens, with cached input tokens at 95% off the input price, which works out to about $0.10 per million cached tokens.
How does Argon score on coding benchmarks?
Google reports a new state of the art on DeepSWE v1.1 at 77.9% and first place on Zapier's AutomationBench at 51.3%. The public Vals Index has not yet published an Argon result.
Does Argon lead the Vals Index?
Google says it does, but the published Vals Index page still lists Claude Sonnet 5.5 at 67.04% and Claude Opus 5.5 at 66.97% at the top. Teams should wait for the leaderboard to add Argon before treating the claim as confirmed.
What is the Fairwind Program?
Google's phased access channel for this launch: Argon is initially rolling out to trusted cyber defenders through it, including a variant without cyber guardrails for full defensive capability.
What cybersecurity results does Google cite?
A tie for first place on CWE-bench v1 at 68%, improved vulnerability discovery over 3.8 Flash Cyber on internal and Wiz black-box penetration testing benchmarks, and a critical healthcare software vulnerability found through Wiz's Scan for Good.
Why does the 1 million token output limit matter?
It lets the model generate very long single trajectories, so tasks previously split across multiple calls can run in one pass. That changes harness design and cost modeling for agentic workloads.
Sources
- Google blog: Gemini 4 Argon announcement
- Google DeepMind: Fairwind Program
- Vals AI: Vals Index leaderboard and methodology
- Wiz: Scan for Good
- Hacker News discussion of the announcement
- Googleplex HQ (cropped)
- CC BY-SA 4.0
- Virginia Tech - data center
- CC BY-SA 2.0
- Google Image Library — Gemini icon and publication credit












