If you ask three vendors what it costs to build an AI agent, you will probably get three figures that are far apart, and none of them will be wrong. They are pricing different things: a different scope, a different number of systems to connect, a different level of testing, and a different assumption about who checks the agent's work.
This guide does not quote prices, because a number without your scope behind it would be a guess. It explains what you pay for when you build an AI agent, what you keep paying for once it is live, which design decisions move each cost, and how to question a vendor's estimate.
Quick answer
- There are two budgets. A one-time build budget and a recurring running budget. An estimate that shows only the first is incomplete.
- Build cost is mostly ordinary software work: scoping, integrations, data preparation, evaluation, guardrails and the screens people use to approve or correct the agent.
- Running cost has five parts: model usage billed by tokens, hosting, monitoring, maintenance and staff time spent reviewing the agent's work.
- The biggest levers are yours: how narrow the task is, how many systems the agent can write to, how much autonomy it has, and which model handles which step.
- A narrow pilot is the cheapest way to get real numbers. Measured usage from real cases is worth more than any forecast.
What drives the build cost
Scoping and discovery
Before anyone writes code, someone has to establish what the agent will do, what it must never do, and how you will know it is working. That means mapping the current process, listing the exceptions staff handle from memory, and agreeing what a correct outcome looks like. If the process has never been written down, this phase is longer, and skipping it moves the cost into rework later.
Scoping should also test whether you need an agent at all. A fixed workflow with an AI step is cheaper to build, test and run. Our guide to AI agents vs chatbots vs workflow automation covers that decision.
Integrations
An agent acts through tools, and each tool is a piece of software that connects to one of your systems: a CRM, an order database, a help desk, a calendar. This is usually the largest single block of build effort. Three things determine its size:
- How many systems. Each one needs its own connection, credentials, error handling and tests.
- Whether a usable API exists. A documented, modern API is quick to connect. A legacy system with no API may need one built around it first.
- Read versus write. Letting an agent look up an order is simple. Letting it change one requires validation, permission checks and safe handling of retries, so that a repeated request does not issue a refund twice.
Tool quality matters more than it sounds. In its guide Building Effective Agents, Anthropic reports that when building one of its own agents its engineers "spent more time optimizing our tools than the overall prompt".
Data preparation
Most agents need your content: policies, product information, past tickets, contract terms. If that material is scattered, out of date or contradictory, the agent will repeat the contradictions. Preparation means deciding which sources are authoritative, archiving stale content, setting up retrieval so the agent finds the right passage, and carrying over access rules so the agent cannot show a document to someone who is not allowed to see it.
Evaluation
Check any low estimate for this item first. An agent does not always take the same path twice, so one successful demo proves little. Evaluation means assembling realistic cases with known correct outcomes, running the agent against them, scoring the results, and repeating that every time a prompt, a tool or the model changes.
The effort grows with the number of case types and the cost of a wrong action. It also needs time from your subject experts, because only they can say whether an answer is right. Anthropic's guide recommends "extensive testing in sandboxed environments, along with the appropriate guardrails", which means a test copy of your systems has to exist or be created.
Guardrails
Guardrails are the controls that limit what the agent can do: permission checks enforced by the systems it calls, limits on steps and spend per task, input and output checks, and a log of every action. The amount of work depends on risk. An agent that drafts internal summaries needs far less than one that sends messages to customers or moves money.
Approval interfaces
If a person must approve some actions, that person needs a screen to do it on: a queue of pending actions, the evidence the agent used, controls to approve, edit or reject, and a record of who decided what. This is a small application in its own right, and its design determines how long each review takes, which feeds directly into running cost.
What drives the running cost
Model usage: how token billing works
Model providers charge by the token. A token is a fragment of text; Anthropic's pricing documentation gives a rough estimate of about 4 characters or 0.75 English words per token. Rates are published per million tokens, and the structure is similar across the major providers. The points below come from the official pricing pages of Anthropic, OpenAI and Google's Gemini API as read on October 4, 2026. Check them for current rates, since they change.
- Input and output are priced separately. Input is everything sent to the model. Output is what it writes back. On Anthropic's price table, output tokens cost more than input tokens for every model listed.
- Larger models cost more per token. Each provider offers several models at different price points. Anthropic's own guidance is to choose a smaller model for simple tasks and keep its more capable models for complex reasoning.
- Tools count as tokens. Anthropic states that tool definitions sent with a request and the results returned from tools are billed as input tokens. An agent with many tools pays for their descriptions on each request.
- Reasoning can count as output. Google's Gemini pricing lists its output price as "including thinking tokens", so a model that reasons at length before answering produces more billable output.
- Repeated context can be cached. All three providers price cached input separately, and Anthropic and OpenAI list it below their standard input rate. The details differ: Anthropic charges more than the base input rate to write to the cache and much less to read from it, and Google lists a storage charge for cached content.
- Non-urgent work can be batched. All three list discounted batch processing for requests that do not need an immediate answer.
- Some built-in tools are billed per use. Web search, for example, is charged per search on top of token costs by Anthropic and OpenAI, and Google charges for search grounding beyond a free allowance.
The reason agents cost more to run than chatbots follows from this. An agent works in a loop: it plans, calls a tool, reads the result and decides again. Each turn of the loop is a new model request that typically includes the conversation and tool results so far, so the input grows with every step. Anthropic puts it plainly: "The autonomous nature of agents means higher costs, and the potential for compounding errors."
A note for UK readers: Anthropic states that all its prices are in USD, and the other two providers also list rates in USD. If your revenue is in GBP, model usage carries exchange rate exposure, and you should ask how any tax on the provider's invoice is handled.
Hosting and monitoring
The agent's own software has to run somewhere: the orchestration code, the tool services, a database for state and logs, and often a search index for retrieval. This can usually share infrastructure with an existing application, and it grows with availability and data residency requirements. Some managed agent platforms add their own charge; Anthropic's Managed Agents, for instance, is billed on tokens plus session runtime.
You also need to see what the agent did on each task, what it cost and where it failed. That means storing traces, building or buying dashboards, and setting alerts on spend and error rates. Storage cost depends on how long your audit needs require you to keep logs.
Maintenance
Three things change underneath a live agent. Your business rules and content change. The systems it connects to change their APIs. And the model itself is retired on the provider's schedule. Anthropic's model deprecations page says that requests to retired models will fail and that it gives at least 60 days' notice before retiring publicly released models. OpenAI's deprecations page likewise says a model is no longer accessible after its shutdown date.
Moving to a newer model is not a one-line change. The agent has to be re-tested, and the bill can shift too: Anthropic's pricing page notes that its newer models use a tokenizer that produces more tokens for the same text. This is where a good evaluation set pays for itself: it turns a migration into a test run instead of a fresh project.
Human review time
Staff time spent approving actions and handling escalated cases is a running cost that never appears on a vendor invoice. It depends on what share of tasks need review and how long each review takes. Plan to review heavily at first and reduce review for low-risk actions as evidence builds.
How design choices move the budget
| Design choice | Effect on build cost | Effect on running cost |
|---|---|---|
| Fixed workflow with AI steps instead of a free-running agent | Lower: easier to test | Lower: predictable number of model calls |
| Read-only access first, write access later | Lower: fewer guardrails needed at the start | Higher human time, since staff still perform the action |
| Human approval on every action | Higher: approval interface required | Higher review time, lower risk of costly errors |
| One large model for every step | Lower: simpler to build | Higher token spend |
| Routing simple steps to a smaller model | Higher: more to build and evaluate | Lower token spend at volume |
| Several agents working together | Higher: more to coordinate and test | Higher: more model calls per task |
The pattern is that optimization costs effort up front and pays back with volume. At low volume, a simple design with a capable model is often the sensible choice. At high volume, engineering time spent on caching, routing and trimming context becomes worth it.
How to scope a first pilot
A pilot exists to replace assumptions with measurements. Keep it small enough that the answer arrives quickly.
- Pick one task. Choose something frequent, well understood and tolerant of mistakes, with a named owner in the business.
- Limit the systems. Connect one or two, and prefer read access plus drafts over direct writes.
- Define success before building. Agree the measures: share of cases handled correctly, time saved per case, share escalated to a person.
- Collect test cases from history. Past tickets, emails or orders with known outcomes become your evaluation set.
- Keep a human in the loop. Have the agent propose and a person approve, and record every correction.
- Instrument cost from day one. Log tokens, steps and review minutes so you can calculate a cost per completed task.
- Set a decision point. Agree a date and the criteria for expanding, changing or stopping.
An illustrative example, not a client case study: a distributor wants an agent to handle delivery change requests. The pilot reads from the order system only, drafts replies for the support team to send, and covers one product line. At the decision point the business knows the measured cost per request, how often staff edited the drafts, and whether write access is worth building.
Questions to ask a vendor about their estimate
- What is explicitly excluded from scope?
- Which integrations are included, and what did you assume about our APIs and test environments?
- How much of the estimate is evaluation, and who builds the test set?
- Does it include an approval interface, audit logging and monitoring, or are those extra?
- What usage did you assume when estimating model costs: tasks per month, steps per task, which model?
- Who holds the model provider account and pays the usage bill? If you re-bill it, is there a markup?
- What limits will stop a runaway task from running up cost?
- What does ongoing maintenance include, and does it cover migration when the model we launch on is retired?
- Who owns the code, prompts, test sets and logs if we change supplier?
- How many hours do you need from our staff?
If you are evaluating an offshore partner, our guide on how to outsource software development from the US or UK covers contracts and engagement models in more depth.
For technical readers: estimating token spend
A first approximation of monthly model cost is tasks per month, multiplied by average steps per task, multiplied by the cost of one step (input tokens at the input rate plus output tokens at the output rate). The trap is that input per step is not constant. Because each step usually resends the system prompt, tool definitions and accumulated history, total input across a task grows faster than the step count.
- Measure from traces. Per-task token counts from a pilot beat any spreadsheet estimate.
- Put stable content, such as the system prompt and tool definitions, at the start of the prompt so it can be cached.
- Return only the fields the agent needs from each tool call. Large raw API responses are billed as input on every later step.
- Set hard limits on iterations, output length and spend per task. Anthropic notes that stopping conditions "such as a maximum number of iterations" are common practice.
- Model the tail, not only the average. A small share of long-running tasks can account for a large share of spend.
Frequently asked questions
Why do AI agent development quotes vary so much?
Because the vendors are pricing different scopes. The main differences are usually the number of integrations, whether the agent can write to your systems, how much evaluation is included, and whether approval screens, monitoring and maintenance are in the estimate or left out.
Is the model usage bill the largest cost of running an agent?
Not necessarily. It depends on volume and design. For a low-volume agent with heavy human review, staff time and maintenance can outweigh token charges. For a high-volume agent with many steps per task, model usage can dominate. A pilot that logs both will tell you which applies to you.
Can we get a fixed price for an AI agent project?
The build can often be fixed-price once the scope is tightly defined, which is one reason to pay for a scoping phase first. Model usage is consumption-based by nature, so expect it to be estimated from stated assumptions and controlled with spending limits, not fixed.
Does using a smaller model always save money?
No. A smaller model is cheaper per token, but if it takes more steps, makes more mistakes or needs more human correction, the cost per completed task can be higher. Compare models on cost per correct outcome using your own test cases.
Next step
The cost of an AI agent is mostly decided before development starts, by how narrowly the task is defined, how many systems it touches and how much it is allowed to do unsupervised. Write down one candidate task, the systems involved, the monthly volume and what a mistake would cost. With that page, a vendor can give you an estimate with its assumptions stated, and you can compare estimates fairly.
Entrant Technologies builds web applications, mobile apps and custom software, including the integrations and admin screens that agent projects depend on. If you would like a second opinion on a scope or an estimate, you can get in touch.