The headline needs a footnote

Google announced Gemini 4 Argon on September 30, 2026. The standout number is a one-million-token output limit, increased from the previous 64K limit. That is an output limit, not a claim about the size of its input context window. Google’s announcement

Why does the distinction matter? Reading a large document and generating a very long response are different requirements. A model can be strong at one without having the same headroom for the other. A maximum also describes capacity; it does not mean every request uses that many tokens.

What the announced pricing means

These are Google’s announced introductory rates, in US dollars per million tokens:

Token typeIntroductory rate
Input$2.00
Output$10.00
Cached input$0.10

The cached-input figure is calculated from the announced 95% discount on the $2 input rate. Google also states later input and output rates of $4 and $20. The announcement does not give an end date for the introductory period. Pricing announcement and footnote

The useful calculation is the cost of a completed task. Long outputs, repeated attempts and tool calls can change that cost substantially. A low input rate is only one part of the bill.

The coding comparison is mixed

Google’s published comparison reports these percentages:

BenchmarkGemini 4 ArgonGPT-6 AstraClaude Opus 5.5
DeepSWE v1.177.974.174.2
Terminal-Bench 4.057.458.266.4

Argon leads this DeepSWE comparison; Opus leads this Terminal-Bench comparison. There is no universal winner in these two rows. Google’s comparison table

The methodology also matters. The published results combine Argon’s evaluation settings with competitor-reported or leaderboard values, and the agent harnesses differ. These are provider-published comparisons, not independent tests by The Latent Notes. Evaluation methodology

Treat each score as a result for a particular benchmark and setup. It is evidence to investigate, rather than a guarantee for your own workflow.

Availability is part of the specification

At announcement, access was being phased through Google’s Fairwind Program for trusted cyber defenders, with broader availability planned. That should not be read as unrestricted public API access. Release announcement

The practical takeaway: separate capacity, cost, measured performance and access. Each answers a different question, and together they make a much more useful model comparison.

This note reflects the sources checked on October 11, 2026. Announced terms and availability can change.