AI Intelligence Costs: Your Automation Math Is Out of Date

RedHub AI Editorialupdated October 2, 20268 min read

A new copier still shrink-wrapped on its pallet and lit red, while two staff sort paper by hand at the counter behind it.
Jump to a section9

AI intelligence costs, meaning what it costs to have one of the most capable AI models do a piece of work, now list at $2 per million input tokens and $10 per million output tokens from both Google and Anthropic. Google announced Gemini 4 Argon at that rate on September 30, 2026, though most businesses cannot run it yet, and Anthropic lists Claude Sonnet 5.5, released two days earlier, at the same price. Argon's rate is introductory: Google says it rises to $4 and $20, and has not said when. When the cost of a run drops, three things move at once. Tasks that did not pay to automate start to pay, agents can afford longer runs, and some software gets built instead of bought. Automation math done before this fall is worth doing again.

TL;DR: As of October 2026, Gemini 4 Argon and Claude Sonnet 5.5 both list at $2 per million input tokens and $10 per million output, but Argon's price is introductory and rises to $4 and $20. Cheaper runs move the automation line three ways: thresholds fall, agents run longer, and some software gets built instead of bought. A line moved by an introductory price can move back, so price every workflow at the later rate too. The AI Unit-Economics & Token-Shock Exposure Kit ($59) grades that exposure workflow by workflow. In this guide: the true cost of an AI agent, when one workflow is worth automating, what cheap AI does to your attack surface, and the signs of AI vendor lock-in.

What frontier AI costs as of October 2026

"Frontier" is the industry's word for the most capable models a lab sells. They are billed in tokens, the small chunks of text a model reads (input) and writes (output). Google publishes its prices in its Argon announcement (opens in a new tab), and Anthropic on its Sonnet 5.5 page (opens in a new tab).

Gemini 4 ArgonClaude Sonnet 5.5
Input, per million tokens$2 (introductory)$2
Output, per million tokens$10 (introductory)$10
Cached input95% off the input price$0.20 per million for cache reads
Later price$4 input, $20 output, date not givenNo change announced
Who can use itFairwind cyber defenders first, then paid API customers and Google AI Ultra subscribers, no dates givenReleased September 28, 2026

Two rows shape everything below. Argon's price is set to double on a date nobody has named, and most businesses cannot run Argon yet. Our side-by-side of the two models covers the comparison itself.

A task can get cheaper while the token price stands still

Anthropic prices Sonnet 5.5 the same as Sonnet 5, at $2 and $10, yet says that in its own testing the new model "costs up to 30% less per task than its predecessor" and generates output more than 30% faster. Anthropic also reports that its early testers saw Sonnet 5.5 batch tool calls together more than Sonnet 5, "leading to fewer steps and lower costs." Those are claims Anthropic publishes about its own model, not measurements on your work.

They point at the number that matters. A task gets cheaper in two ways: the price per token falls, or the model uses fewer tokens and steps to finish the job. The second kind of saving never shows up on a rate card. So this post works in cost per run, and our post on the true cost of an AI agent shows how to get that figure right.

Effect 1: the automation line moves

Every repeated task has a line. Below it, running the task through AI costs more than the work is worth. Above it, each run pays for itself. When the cost per run falls, tasks cross the line without changing at all.

Illustrative example: a parts distributor gets 40,000 customer emails a month and wants each one tagged by topic and urgency so it lands in the right queue. Say a correct tag saves $0.05 of someone's sorting time, and assume, generously, that every tag comes out right. At $0.10 per run, tagging costs $4,000 a month to save $2,000. At $0.02 per run, it costs $800 to save the same $2,000, and the task clears by $1,200 a month. These numbers are made up to show the arithmetic.

The tasks that cross first are small, frequent and dull: tagging, routing, pulling fields off a form, checking an order for missing items. None ever justified a project on its own. At volume, a few cents decides it. Cost is only the first test, though. Our post on the AI automation threshold scores one workflow on six factors, from data readiness to who owns the result.

Effect 2: agents can afford longer runs

An agent is an AI system that works in a loop: it plans, uses a tool, checks the result and takes the next step, and every step costs tokens. If you cap what one job may cost, halving the cost of a step buys roughly twice as many steps under the same cap.

Longer is not the same as better. If each step succeeds 85% of the time, independently of the others, an 8-step job comes out right end to end only 27.2% of the time (0.85 multiplied by itself eight times). That example comes from our AI Agent Quiet-Failure & Drift Monitor Kit, and real steps are rarely fully independent, so read it as a direction. Cheaper steps are best spent on checks that catch errors, not on extra steps that compound them.

Cheap steps also make it practical to run many agents at once, and our post on cheap AI security risk covers what that does to your exposure.

Effect 3: some software gets built instead of bought

A lot of business software puts a fixed screen in front of a fixed process: a form, a set of rules, a report at the end. When a model can read instructions, look at the data and call the right tools for a few cents a run, a narrow workflow can sometimes be assembled from a model, a prompt and a couple of connections instead of another subscription.

Take a shop that pays for a tool that turns supplier price sheets into its own catalog format. A model that reads each sheet and writes the rows into a spreadsheet, with a person checking the changes, may do the same job. The catch is upkeep: a subscription vendor carries it, and a workflow you build carries it yourself, including model retirements. Anthropic's deprecation page says it gives "at least 60 days' notice" before retiring a publicly released model, and requests to a retired model fail. When one provider ends up holding the whole workflow, you have the problem in our post on AI vendor lock-in.

A line moved by an introductory price can move back

All three effects assume today's price holds. For Argon, Google has already said it will not hold, without saying when it ends.

Illustrative example, continued: suppose the email tagger's $0.02 per run is all Argon tokens at the introductory rate. At the later rate the same run costs $0.04, so the month costs $1,600 to save $2,000. It still clears, but by $400 instead of $1,200. Doubling the price adds an amount equal to the original token bill, so any task whose margin was smaller than its token bill falls back below the line.

So price each workflow twice, at today's rate and at the rate already announced. A workflow that pays only at the introductory price is a bet on a deadline nobody has published. Where prices go after that is not something anyone outside these companies can tell you, and we are not going to guess. What you can know is which of your workflows survive a doubling.

What leaders should do now

Falling prices touch five decisions, and each has its own number to watch.

DecisionWhat to measure
Select a modelAccepted outcomes, not headline tokens
Automate a taskVolume, repeatability, error cost, and review burden
Expand agent autonomyMarginal value versus added operational and security risk
Negotiate vendorsPricing changes, caching, rate limits, portability, and retirement terms
Report ROITime saved, quality improvement, throughput, error reduction, and realized financial value

The first row is the easiest to get wrong: a cheaper token can still buy a more expensive result if more of the output needs fixing. The fourth matters more while prices move. Caching discounts, rate limits (caps on requests or tokens per minute) and retirement notice all shape next year's bill, and you have leverage on them only if you can leave.

Start with the tasks you already rejected. List the repeated work you ruled out on cost in the past year and re-price each item at today's cost per run, then at any announced later rate. Run the survivors through the free Decision Fit Check, which tells you what level of AI a task needs, from plain code to a person. And set a hard spending cap before you add volume. A billing alert is not one, for reasons our API spending limits post lays out.

Find out which workflows survive a price change

The AI Unit-Economics & Token-Shock Exposure Kit computes the fully loaded cost per successful outcome for each workflow from your own numbers, including retries and growth, and bands it PREDICTABLE, DRIFTING or TOKEN SHOCK. With no usage cap and an uncapped true-up in the contract, a workflow reads TOKEN SHOCK even when today's cost is a penny. Its Model Swap tab checks whether a cheaper model is still cheaper per finished task once its misses are paid for.

Get the AI Unit-Economics & Token-Shock Exposure Kit — $59

Pairs well with

The Token Economics Workbook ($59) adds a forecasting calculator, a model-routing matrix, caching patterns and 15 production teardowns. The AI Burn-Rate & Budget Blowout Forecaster ($49) projects your AI spend across the budget year, names the month you blow the budget, and returns ON BUDGET, TRIM NOW or BLOWOUT AHEAD. The AI Agent Quiet-Failure & Drift Monitor Kit ($49) computes each multi-step agent's end-to-end success rate from its per-step reliability.

More in this guide

What are AI intelligence costs?

They are what it costs to have an AI model finish a piece of work: the token price, plus how many tokens and steps the task takes, retries, tool fees and the time a person spends checking the result.

What do Gemini 4 Argon and Claude Sonnet 5.5 cost as of October 2026?

Google lists Gemini 4 Argon at an introductory $2 per million input tokens and $10 per million output tokens, rising later to $4 and $20, with no date given. Anthropic lists Claude Sonnet 5.5 at $2 and $10, with cache reads at $0.20 per million.

Why does a lower AI price change what is worth automating?

Every repeated task has a cost per run below which automating it pays. When that cost falls, small, frequent tasks such as tagging and routing can cross the line without the task changing.

Will Gemini 4 Argon's price go up?

Yes. Google says the $2 and $10 rates are introductory and rise to $4 and $20 per million tokens. As of October 2026, it has not said when.

Does a cheaper model always cost less per task?

No. A model with a lower token price can need more steps, retries or human correction, and cost more per usable result.

What should I do with automation ideas I rejected on cost?

Price each one again at today's cost per run and at any announced later rate. Then score the ones that pass, one at a time, before building anything.

How it decides
Diagram of the AI Unit Economics & Token-Shock Exposure Kit: a cost-per-outcome score reading predictable and a no-usage-cap gate forcing TOKEN SHOCK.

The gate this post refers to, drawn from the tool’s own logic. See the tool.