Gemini 4 Argon's 1M-Token Output: What Agents Can Do Now

RedHub AI Editorial6 min read

A continuous paper printout runs from a printer across the carpet and out the doorway, lit red where it crosses.
Jump to a section7

Google just made a single AI response roughly 15 times longer. The Gemini 4 Argon 1 million-token output limit, up from 64K on the previous model, caps how much the model can write in one response, and Google says the headroom lets it "generate hundreds of thousands of tokens in a single trajectory" to solve hard problems in one go. For a business, that changes the size of a job you can hand over in one piece. It also changes the bill and the review burden, and both grow faster than most people expect.

TL;DR: The Gemini 4 Argon 1 million-token output limit lets one response carry a whole code change set, a full document review or a long chain of reasoning. A single response that uses the full limit costs $10 in output at the introductory price and $20 later, and on long jobs the input side is often the bigger bill. A bigger response is also more for a person to check. The Agent Side-Effect & Blast-Radius Checkpoint ($89) grades which actions are safe to run without a person watching. Start with the pillar: Gemini 4 Argon: AI That Works Longer Than You Can Watch.

Three limits that get mixed up

Coverage of this launch blurs three different things, and each one changes a different part of your plan.

TermWhat it capsWhat Argon changed
Context windowHow much the model can read and consider at once: instructions, documents, tool resultsGoogle's announcement does not state a figure
Output limitHow much the model can write in one responseRaised from 64K to 1 million tokens
Agent runThe whole job: many model calls, tool calls and checks, start to finishNo single cap; a run can span many responses

The output limit is the only one Google changed on the record. It does not make an agent run forever. It makes each step of the run able to carry far more.

What a long response is good for

Most AI tasks never come near 64K tokens of output, let alone a million. The limit matters for a narrower set of jobs where the work only makes sense in one unbroken piece.

Take a due-diligence review of a stack of supplier contracts. The output a reviewer wants is one memo: every clause that matters, compared across every contract, with the gaps flagged. Split into ten separate requests, each piece loses sight of the others, and someone has to stitch the memo together and check that nothing contradicts. In one response, the model can hold the whole comparison together. Google's three target areas are full of jobs with that shape:

  • Code changes that touch many files and have to stay consistent with each other
  • Reviews that compare many documents and report on all of them together
  • Security investigations that have to trace one thread through a lot of evidence
  • Long reasoning before a short answer, where Google says the extra room adds "depth in reasoning"

What a full response costs

At Argon's introductory price of $10 per million output tokens, one response that uses the whole limit costs $10 in output. After the introductory period, at $20 per million, it costs $20. Few teams can run that today: access is limited to Google's Fairwind cyber-defense program and its own teams, with paid API customers and Google AI Ultra subscribers next and no dates given.

That is not the cost of the job. On a long agent run, the model typically re-reads its instructions, documents and earlier results on every call, so input is often the bigger bill. Caching changes that picture: Google prices cached input at 95% off, so a job that reuses the same large documents across many calls costs far less than one that sends fresh input every time. A budget that only counts output tokens will be wrong in the direction that hurts.

A cost cap per run is the control that matters here, and it has to stop the run, not just send an email. Our post on why API spending limits don't stop runaway bills explains the difference.

The review problem nobody prices in

A longer response has to be checked by someone. A 300-word draft takes a minute to review. A response that runs to hundreds of thousands of tokens is a book, and nobody reads a book-length code change line by line before merging it.

That cuts against the main benefit. One long response keeps the work consistent, but it also means one long thing to trust or reject. If it is wrong in one place, you may have to throw out the whole response or pay someone to find the error. Splitting a job into reviewable pieces costs some consistency and buys back the ability to check each piece. Neither choice is free, and the right split depends on how expensive a missed error is in that workflow.

The practical answer is to decide where a person checks before the model writes anything. Our guide to long-horizon AI agents sets out seven rules for placing those checkpoints, along with budgets and recovery.

Bigger output, bigger mistakes

The same headroom that lets Argon finish a big job lets it make a big mistake. An agent built on it can follow a wrong assumption through an entire change set, repeat a failing tool call, or reach systems it was never meant to touch. Give the agent enough room to finish the job, and never enough unchecked authority to turn one mistake into a large incident. Those are separate settings, and the output limit only moved the first one.

Know which actions can run without a person watching

The Agent Side-Effect & Blast-Radius Checkpoint grades each action your agent takes on two separate axes, how reversible it is and how many records one call can touch, and returns RUN UNATTENDED, RUN WITH APPROVAL or DO NOT AUTOMATE.

Get the Blast-Radius Checkpoint — $89

Pairs well with

The AI Spend Runaway & Billing-Safeguard Gate ($49) checks whether a runaway agent bill would be stopped, or only reported in an email. The Token Economics Workbook ($59) adds a forecasting calculator and caching patterns, the two levers that decide what a long job costs. The Agent Reliability Harness ($149) evaluates the full sequence of steps an agent took, including tool choice and step efficiency.

More in this guide

What is Gemini 4 Argon's output limit?

Google lists it at 1 million tokens, up from 64K on the previous model. The output limit is the most the model can write in a single response.

Is the output limit the same as the context window?

No. The context window is how much the model can read at once. The output limit is how much it can write in one response. Google's announcement is about the output limit and does not state a context window figure.

Does a 1 million-token limit mean an agent can run forever?

No. It caps one response, not a whole job. An agent run is usually many responses and tool calls, and it still needs its own limits on time, cost, steps and retries.

How much does a full 1 million-token response cost?

At Google's introductory price of $10 per million output tokens, a response that uses the whole limit costs $10 in output. At the later price of $20 per million, it costs $20. Input tokens, retries and tool calls are extra, and on long jobs input is often the larger cost.

Does every task need a long output limit?

No. Most business tasks produce short output. The limit matters for jobs that only make sense in one piece, such as a consistent code change across many files or one memo comparing many documents.

How do you review a long AI response?

Decide before the model starts where a person will check the work, and break the job at those points. A response too long to review carefully is a risk, even when it is mostly right.

How it decides
Diagram of the Agent Side-Effect & Blast-Radius Checkpoint: six proposed actions graded on reversibility and blast radius, an undo-window gate, and the batch reading UNSAFE TO AUTOMATE.

The gate this post refers to, drawn from the tool’s own logic. See the tool.