The True Cost of an AI Agent (and How to Calculate It)

RedHub AI Editorialupdated October 2, 20266 min read

A worker with her arms folded looks up at a tied stack of paper on the counter, taller than she is and lit red.
Jump to a section9

The true cost of an AI agent is everything it takes to get usable results out of it each month, divided by the results your team accepts. That total covers the model calls, tool and data fees, hosting, monitoring, human review, upkeep, a reserve for failures, and a monthly share of what the agent cost to build. Token price is one line in that list. In the illustrative example below, an agent whose tokens cost $0.03 per request costs more than nine times that per accepted result once everything else is counted.

TL;DR: Total monthly agent cost = build amortization + model inference + tool and data fees + hosting + observability + human review + maintenance + a failure and incident reserve. Divide it by the results you accept, not the requests you send. Measure real runs before you estimate, and treat acceptance rate as a lever, because it can move the number more than the token price does. The AI Cost-Per-Task Calculator ($29) puts each task's AI cost next to its human cost. Start with the pillar: AI Intelligence Costs: Your Automation Math Is Out of Date.

Token cost is one line item

The common way to estimate an agent is to take the model's price, multiply by a rough number of requests, and call it the budget. That measures token spend, not the agent. An agent makes several model calls per request, calls tools that charge their own fees, runs on infrastructure someone pays for, and produces work a person checks. Some of that work gets thrown out. Our post on why AI agent costs climb fast covers how loops and retries inflate the token line itself. This post builds the whole monthly bill around it.

The full monthly cost model

Use this formula:

Total monthly agent cost = build amortization + model inference + tool and data fees + hosting + observability + human review + maintenance + failure and incident reserve.

Build amortization means spreading what the agent cost to build over the months you expect to use it. Observability means the logging and monitoring that let you see what the agent did. Then:

Cost per accepted result = total monthly agent cost ÷ results accepted that month.

The last term in the first formula is the one estimates leave out. A workflow that can send messages, change records or handle sensitive information needs an allowance for testing, investigation, exceptions and the controls that keep it safe.

Build cost versus run cost

Seven categories cover it. The first is paid once. The other six come back every month.

Cost categoryExamples
One-time buildWorkflow design, integrations, prompt and policy design, evaluation data, security review, testing
Model inferenceInput, output, caching, retries, fallback calls
Tools and dataSearch, retrieval, CRM, browser, API, database, or third-party fees
InfrastructureHosting, queues, vector storage, observability, identity, and secrets management
Human operationsReview, exception handling, quality assurance, escalation, support
MaintenanceModel upgrades, prompt changes, regression testing, connector updates
Risk reserveIncident response, rollback, compliance review, and remediation

Two rows surprise people. In the worked example below, human review is the biggest recurring line, at half the cost of each request. And maintenance rarely reaches zero, because models change underneath you. Our post on the hidden costs of AI tools goes through the ones teams forget to budget.

Measure the workload before you estimate

Generic token assumptions are fine for a first look and wrong for a production decision. Run a sample of real requests through the agent, including the messy ones, and record for each run:

  • Input and output tokens across every model call, not just the first one
  • Tool calls and their fees
  • Retries, fallback calls and failures
  • How long the run took
  • The minutes a person spent reviewing or fixing the result
  • Whether the result was accepted

Write down what "accepted" means before the sample runs. A result a reviewer had to rewrite from scratch is not accepted, even if the agent delivered it on time. With that sample you can compare the agent against a person, a software tool or a simpler automation on the same terms. Our guide to calculating AI cost per task covers where to find each number.

A worked example

Illustrative example: an agent handles 10,000 requests a month. Tokens cost $0.03 per run, tools add $0.02, infrastructure and monitoring add $0.01, and human review averages $0.06. That is $0.12 per request, or $1,200 a month, before maintenance and the risk reserve. If 80% of results are accepted, that is 8,000 accepted results, and $1,200 ÷ 8,000 = $0.15 per accepted result before any rework.

Now add the monthly fixed costs, still illustrative. The build cost $6,000, spread over 12 months, which is $500 a month. Maintenance runs $300 a month and the risk reserve $200. The month now costs $1,200 + $500 + $300 + $200 = $2,200, and $2,200 ÷ 8,000 is $0.275 per accepted result. The token line was $0.03.

Acceptance rate can outweigh token price

Run the same example two ways. Cut the token cost by a third, from $0.03 to $0.02 per run, and the month drops by $100 to $2,100, or about $0.26 per accepted result. Leave the tokens alone and raise acceptance from 80% to 90%, and the same $2,200 is spread over 9,000 accepted results, about $0.24 each.

In this example, ten more points of acceptance save more than twice as much per accepted result as a one-third cut in token price. That is why a cheaper model that needs more fixing can cost more.

The answer depends on how long the agent lasts

One input in this model is a guess, and it moves the answer. The $6,000 build was spread over 12 months. Spread it over 36 months instead and it adds about $167 a month, so the month costs about $1,867 and each accepted result about $0.23. Same agent, same work, a different number, based on a guess about its lifespan.

Model retirements shorten that life in ways you do not control. Anthropic, for one, promises "at least 60 days' notice" before a publicly released model retires, according to its model deprecations page. After that date, calls to the old model fail. A forced migration lands in the maintenance line, possibly halfway through the period you amortized over. There is no correct period. Pick one, write it next to the number, and redo the math when a model you depend on gets a retirement date.

Put each task's AI cost next to its human cost

The AI Cost-Per-Task Calculator compares human hours against AI for every task your team runs, with real AI cost and setup cost, and gives each task an Automate, Review or Skip verdict. Bring the cost per accepted result from your sample, not the token price.

Get the AI Cost-Per-Task Calculator — $29

Pairs well with

The AI Unit-Economics & Token-Shock Exposure Kit ($59) computes the fully loaded cost per successful outcome for each workflow, with retry and growth multipliers, and bands it PREDICTABLE, DRIFTING or TOKEN SHOCK. The Workslop Cost & Verification-Tax Calculator ($49) turns the time colleagues spend verifying or redoing polished but hollow AI output into a monthly dollar figure per team. When a model upgrade lands in your maintenance line, the Prompt Regression Lab ($89) snapshots a baseline and fails the build on any regression.

More in this guide

What is the true cost of an AI agent?

It is the full monthly cost of running the agent, including build amortization, model calls, tools, hosting, monitoring, human review, maintenance and a failure reserve, divided by the number of results your team accepts that month.

Why is token price not the cost of an agent?

An agent also uses tools with their own fees, runs on paid infrastructure and produces work people check. Token price covers only the model calls, and it counts every result, including the ones your team throws out.

How do I calculate cost per accepted result?

Add up everything the agent cost in a month and divide by the results accepted that month. In an illustrative example, $1,200 of direct cost across 10,000 requests at 80% acceptance gives $1,200 ÷ 8,000 = $0.15 per accepted result.

What costs do agent estimates usually leave out?

Human review, maintenance after model changes, and a reserve for failures and incidents. A monthly share of the build cost is often missing too, because it was paid up front.

How should I amortize the build cost?

Spread it over the months you expect to use the agent, write that period next to the number, and redo the math when a model you depend on gets a retirement date.

What matters more, token price or acceptance rate?

It depends on your numbers. In this post's example, raising acceptance from 80% to 90% lowers cost per accepted result more than cutting token cost by a third.

How it decides
Formula: 1,040 tasks a year times 12 minutes saved over 60 times $50 an hour, working one repetitive task to $10,400 a year of labor.

The gate this post refers to, drawn from the tool’s own logic. See the tool.