Gemini 4 Argon: AI That Works Longer Than You Can Watch
RedHub AI Editorial8 min read

Jump to a section9
Gemini 4 Argon is Google's newest frontier model, announced on September 30, 2026, and the change that matters most for a business is how much it can produce in one go. Google raised its output limit from 64K tokens to 1 million tokens. Tokens are the small chunks of text a model reads and writes, and the output limit caps how many it can write in a single response. Google says the extra headroom lets Argon "generate hundreds of thousands of tokens in a single trajectory" and solve hard problems in one pass. Most businesses cannot use it yet: Google is releasing it first to vetted cyber defenders, then to paid API customers and Google AI Ultra subscribers, with no dates for either stage.
TL;DR: Google built Gemini 4 Argon "to sustain deep reasoning across complex, long-horizon workflows" in software engineering, legal and finance work, and cyber defense. Its introductory price is $2 per million input tokens and $10 per million output tokens, rising later to $4 and $20. Access starts with Google's Fairwind cyber-defense program. For a business, the useful question is which workflows could safely hand an AI a much bigger piece of work, and what has to be in place first. The free Decision Fit Check tells you what level of AI one task needs. Also see what the 1M-token output changes, Argon vs Claude Sonnet 5.5, why Google restricted access, and design rules for long-horizon agents.
What Google announced
The announcement came from Google DeepMind's Koray Kavukcuoglu on the Google blog (opens in a new tab). Google aims Argon at three kinds of work: real-world software engineering, enterprise knowledge work such as legal and finance, and cybersecurity defense. Here is what the company has said on the record, as of October 2, 2026.
| Item | What Google says |
|---|---|
| Output limit | 1 million tokens, up from 64K |
| Introductory price | $2 per million input tokens, $10 per million output tokens |
| Cached input | 95% off the input price |
| Price after the introductory period | $4 input, $20 output, per million tokens |
| How long the introductory price lasts | Not stated |
| Who gets it first | Trusted cyber defenders in the Fairwind Program |
| Who gets it next | Developers, enterprises and consumers, starting with paid API customers and Google AI Ultra subscribers, no dates given |
Two of those rows shape every decision below. The price will double, on a date nobody has named. And most readers of this post are, for now, on the waiting list.
Why a bigger output limit matters
A chat assistant answers a question and stops. An agent works in a loop: plan, act, look at the result, fix what went wrong, take the next step. Some of the hardest jobs need one long stretch of output inside that loop: a full set of code changes across a large repository, a complete review of a stack of contracts, or a long chain of reasoning before an answer. At 64K tokens, a job like that has to be broken into pieces and stitched back together, and the stitching is where detail gets lost.
Think of a library upgrade in a large codebase. A short assistant can explain how to do it. A model with room to write hundreds of thousands of tokens in one response can, in principle, map what depends on the old library, write the changes, and explain each one in a single pass, ready for a person to review. That is the bet Google is making. Whether Argon delivers it on your codebase is something nobody outside early access can say yet.
A bigger limit is also not an invitation to longer reports. Most business output should stay short, and every extra token is something a person has to read.
Who can use Gemini 4 Argon today
Almost nobody reading this. Google's first stage is its Fairwind Program, a limited cyber-defense program for governments, Google Cloud customers and security partners, and SecurityWeek reports that Google's own internal teams also have access. Google has not said when paid API customers or AI Ultra subscribers get in.
So the useful work this month is choosing which workflows deserve a test when access arrives, and what controls each one needs before an AI touches it.
The price is introductory
At $2 in and $10 out per million tokens, Argon's launch price matches Anthropic's list price for Claude Sonnet 5.5. That tie lasts only while the introductory period does. At $4 and $20, Argon costs twice as much per token.
Per-token price is also the wrong number to budget on. A business pays for finished work it can use, so the figure to track is cost per accepted result: everything a task costs, including retries, tool calls and the minutes a person spends checking it, divided by the results you keep.
Our guide to AI cost per task walks through that calculation step by step.
Where it could matter for a business
"Use AI for operations" is not a project. "Read every inbound supplier document, pull the required fields, flag what is missing and draft the follow-up" is. That second version has the shape a long-output model is built for: high volume, a pattern that repeats, a lot of reading, and a mistake a person can catch before it reaches a supplier.
Within the three areas Google names, the candidates that look most like that are large code migrations teams keep putting off, due-diligence and contract reviews across many documents, RFP drafts and policy comparisons, and security work such as alert triage. In each one, the right first job for the model is to prepare work for a person to approve. A drafted change set is fine. A merged change in production is not, at least not at first.
Longer work means bigger mistakes
Everything that makes this useful also scales the damage when it goes wrong. A model that can write a whole change set in one pass can also write a whole wrong one. An agent built on it can follow a bad assumption for many steps, repeat a failing tool call, run up a bill, or reach systems it should never touch. A short assistant's mistake is one bad paragraph. A long run's mistake can be a day of bad changes.
So the order matters: controls first, capability second. That means limited access, an approved tool list, versioned prompts, a way to escalate to a person, a log of what the agent did, and a tested way to stop it. Our post on long-horizon AI agents works through seven design rules for exactly this, and they apply whichever model you choose.
How much autonomy is the right amount has no clean answer. It depends on what the agent can touch and how fast you would notice it going wrong.
What we will not tell you
We are not going to tell you Argon is better or worse than the model you use today. Access is limited to Fairwind members and Google's own teams, so there are no independent results on business work yet, and a vendor benchmark says nothing true about your documents or your code. The honest test is your own: twenty to fifty real tasks, the same success criteria and the same review effort on each model you are considering, scored by the results you would keep.
Decide what level of AI a task needs before you pick a model
The free Decision Fit Check asks eight questions about one task and tells you whether it needs plain code, a decision model, a fast general model, a frontier reasoning model, or a person.
Run the free Decision Fit CheckPairs well with
The Agent Use-Case Fit & Proof-of-Value Gate ($99) scores a proposed agent on six fit dimensions and weighs its projected return against what its controls will cost, so a project can score well and still come back as not worth building. The Agent Side-Effect & Blast-Radius Checkpoint ($89) grades each agent action on how reversible it is and how many records one call can touch. The Agent Reliability Harness ($149) evaluates the full sequence of steps an agent took, not just its final answer.
More in this guide
What is Gemini 4 Argon?
It is Google's newest frontier AI model, announced on September 30, 2026. Google says it is built to sustain deep reasoning across complex, long-horizon workflows, aimed at software engineering, legal and finance knowledge work, and cyber defense. Its output limit is 1 million tokens, up from 64K.
Can my business use Gemini 4 Argon now?
Probably not yet. Google is releasing it first to trusted cyber defenders in its Fairwind Program, then to developers, enterprises and consumers, starting with paid API customers and Google AI Ultra subscribers. As of October 2, 2026, Google has given no dates for those later stages.
How much does Gemini 4 Argon cost?
Google lists an introductory price of $2 per million input tokens and $10 per million output tokens, with cached input 95% off. After the introductory period the price rises to $4 and $20. Google has not said how long the introductory period lasts.
What is a token?
A token is a small chunk of text, often part of a word, that an AI model reads or writes. Models are priced and limited in tokens. The output limit is the most tokens a model can write in one response.
Is Gemini 4 Argon better than other models?
No independent results on business work exist yet, because access is limited to Google's early-access program and its own teams. Vendor benchmarks do not tell you how a model performs on your own documents or code. The reliable test is your own set of real tasks.
What should a business do before using a long-running AI agent?
Put the controls in place first: limited access, an approved tool list, spending and step limits, a log of every action, a way to escalate to a person, and a tested way to stop the agent. Start with work the agent prepares for a person to approve.


The gate this post refers to, drawn from the tool’s own logic. See the tool.