"AI agent" is now attached to so many products that the term has stopped telling you much. If you are deciding whether your business should use one, you need a definition precise enough to test a vendor's claim against, and a realistic picture of what the technology can and cannot do today.
This article defines an AI agent in plain English, follows one task from start to finish so you can see how an agent actually works, and sets out what should be in place in your business before you rely on one.
Quick answer
An AI agent is software that is given a goal and a set of tools, and uses a large language model (LLM) to decide which steps to take, carry them out, check the results and continue until the goal is met or it needs a person.
- The defining feature is that the model chooses the steps. Nobody wrote them down in advance.
- An agent acts only through the tools it has been given, such as access to your CRM, order system or email.
- Agents are strongest on variable, multi-step tasks where the result can be checked. They are weakest on vague goals, long chains of unchecked steps and decisions that are expensive to get wrong.
- Before using one you need a documented process, systems that software can connect to, clear permissions, an approval step for high-impact actions, and a named owner.
A precise definition of an AI agent
The clearest short definitions come from the companies that build the underlying models. Anthropic's engineering guide Building Effective Agents describes agents as "systems where LLMs dynamically direct their own processes and tool usage, maintaining control over how they accomplish tasks". A later Anthropic article on context engineering compresses this to "LLMs autonomously using tools in a loop". OpenAI's developer documentation on agents describes them in similar terms: systems that plan and complete tasks using tools and keep context across steps.
Three words in those definitions do the work: the model directs the process, it uses tools, and it works in a loop. If a product has all three, it is an agent. If a developer fixed the sequence of steps, or the software only produces text for a person to act on, it is something else. We compare those alternatives in detail in AI agents vs chatbots vs workflow automation, so this article stays with the agent itself.
The seven parts of an agent
Every working agent is made of the same parts. Knowing them gives you a checklist for any proposal or product demo.
| Part | What it means in plain English | Question to ask a vendor or developer |
|---|---|---|
| Goal | The outcome the agent is asked to reach, with the rules it must follow | How is the goal stated, and how does the agent know it is done? |
| Planning | The model breaks the goal into steps and revises them as it learns more | What happens when the first plan turns out to be wrong? |
| Tools | The specific actions it can take in other systems, such as look up an order or draft an email | Exactly which tools does it have, and which of them change data? |
| Context | The information in front of the model for the current task: instructions, the request, and results so far | What does the agent see, and what is kept out of its view? |
| Memory | Information saved outside the current task so it can be used later | What is stored, where, for how long, and can it be deleted? |
| Evaluation | Checking each result against the goal before moving on, and testing the whole system before release | How do you measure whether it got the task right? |
| Human approval | Points where the agent must stop and wait for a person | Which actions need sign-off, and who gives it? |
Context and memory are often confused. Context is the agent's working desk: only what is on it right now can be used. Memory is the filing cabinet: notes the agent writes down and retrieves later. Anthropic's context engineering article describes this second technique as structured note-taking, where the agent records notes outside its working context and pulls them back in when needed.
How an AI agent works: one task, step by step
The scenario below is illustrative, not a client case study. A wholesale distributor receives this email: "We were charged twice for order 4812. Please sort this out." The agent's goal is to resolve billing queries in line with company policy. It has five tools: look up a customer, look up orders, look up payments, request a refund, and draft a reply.
- It receives the goal and the request. As Anthropic's guide puts it, agents begin their work with a command from, or an interactive discussion with, a human user. Here the trigger is the customer's email, which is added to the agent's context alongside its standing instructions and the refund policy.
- It plans. The model reasons that it must identify the customer, find the order, and compare the payments taken against what was ordered. Nothing has been done yet.
- It uses a tool. It calls the customer lookup with the sender's email address and gets back an account record.
- It reads the result and decides what to do next. With the account identified, it calls the order lookup, then the payment lookup. It finds two card payments of the same amount on the same day.
- It checks before it concludes. Two identical payments could be a duplicate charge, or two separate orders. The agent looks for a second order. This is the point where a fixed script would struggle, because the right next step depends on what was just found. In this case there is no second order, so the charge is a genuine duplicate.
- It reaches an action it is not allowed to take alone. Company policy says refunds above a set limit need approval. The agent prepares the refund request with the evidence attached and stops.
- A person approves. A member of the finance team sees the two payments and the single order in an approval queue and clicks approve. If the evidence looked wrong, they would reject it and the agent would take a different path.
- It acts and verifies. The agent submits the refund through the refund tool, then calls the payment lookup again to confirm that the refund is recorded.
- It closes the task. It drafts a reply to the customer, adds a note to the account so that the next query has the history, and finishes. Every step is logged.
Steps 2 to 5 are the loop: plan, act, observe, decide again. Anthropic describes the same pattern, noting that agents need "ground truth from the environment at each step", can pause for human feedback at checkpoints, and should have stopping conditions such as a maximum number of iterations. Had the agent found a second order, it would have skipped the refund entirely and written a different reply. Nobody programmed that branch. The model chose it from what it found.
What AI agents are good at today
Agents perform best when four conditions hold together.
- The path varies from case to case. Anthropic recommends agents for open-ended problems where the number of steps is hard to predict and a fixed path cannot be hardcoded.
- The work happens in systems the agent can reach. Looking things up, comparing records, updating fields and drafting messages all suit an agent, provided each system has a tool for it.
- Success can be checked. The same guide singles out customer support and software development as strong fits. In support, a resolution can be measured. In coding, automated tests show whether the work is correct, and the agent can use failed tests as feedback.
- A mistake is recoverable. A wrongly categorized ticket or a draft that needs editing costs little. That leaves room for the agent to be useful while it is still being tuned.
Typical business uses that meet these conditions include resolving routine support requests, researching and preparing a response to a sales inquiry, reconciling records between two systems, triaging IT requests, and assembling a report from several data sources.
Where AI agents still struggle
The limits matter as much as the strengths, because they decide how much supervision an agent needs.
- Errors compound. A wrong conclusion at step three shapes everything after it. Anthropic states that the autonomous nature of agents means higher costs and "the potential for compounding errors", and recommends extensive testing in sandboxed environments with appropriate guardrails.
- They are not perfectly repeatable. The same request can lead to a slightly different sequence of steps on different days. This is acceptable for many tasks and unacceptable where every case must be handled identically.
- They can be confidently wrong. A language model can produce a plausible statement that is not supported by the data. This is why checking results against a real system, as in step 8 above, is part of the design and not an extra.
- Long tasks degrade. The context engineering article describes "context rot": as the volume of information in front of the model grows, its ability to recall the right details falls. Tasks that run for many steps need deliberate techniques to stay on track.
- They can be manipulated by what they read. OWASP lists prompt injection as a top risk for LLM applications, including indirect injection, where content from an external source such as a website or file alters the model's behavior. An agent that reads customer emails is reading untrusted text.
- They cost more and take longer than a single AI answer. Each loop is another model call. Anthropic notes that agentic systems "often trade latency and cost for better task performance".
- They do not supply judgment you have not defined. If your own staff disagree about how a case should be handled, an agent will not settle it. It will pick one reading and proceed.
Good fit or poor fit?
| Signs of a good fit | Signs of a poor fit |
|---|---|
| Each case needs different lookups and actions | The steps are identical every time |
| The result can be verified in a system of record | Quality is a matter of opinion with no agreed standard |
| Mistakes can be reversed or caught at an approval step | One wrong action is costly or cannot be undone |
| The policy for handling cases is written down | The rules live in a few people's heads |
| The systems involved can be accessed by software | The work depends on tools with no way to connect |
If your process sits mostly in the right-hand column, a simpler approach is likely to serve you better. Anthropic's own advice to developers is to find the simplest solution possible and to increase complexity only when needed.
What your business needs in place before using an AI agent
Most agent projects that disappoint do so because of what surrounded the agent, not the model. Six things should be in place first.
- A written process and policy. The agent needs to be told what "resolved" means, what the limits are and when to hand over. If you cannot explain this to a new employee in writing, you cannot explain it to an agent.
- Systems the agent can connect to. Agents act through tools, and tools are usually built on APIs, the interfaces that let one piece of software talk to another. Older systems without one need integration work first, and this is often the largest part of the build.
- Data you trust. An agent that reads duplicate customer records or outdated policy documents will act on them.
- Narrow permissions. OWASP's entry on excessive agency traces damaging agent actions to excessive functionality, excessive permissions and excessive autonomy. Give the agent only the tools the task requires, and have the connected systems enforce what it may do.
- An approval step and an audit trail. Decide which actions wait for a person, build the screen where that person reviews them, and log every step so that any outcome can be reconstructed.
- A way to measure it, and someone who owns it. Collect a set of real past cases with known correct outcomes and test the agent against them before launch and after every change. Name the person who reviews failures and decides when to widen or narrow what the agent may do.
The cost of an agent is driven by these items more than by the model: the number of systems to integrate, the state of your data, the amount of testing, the approval and monitoring screens, and ongoing upkeep as models, APIs and policies change. If an outside team will build it, our guide on how to outsource software development from the US or UK covers how to structure that relationship.
If the agent makes decisions about people
Agents that affect individuals in significant ways, for example in hiring, lending or eligibility for a service, fall under data protection rules that differ by country. This is a high-level pointer, not legal advice.
In the United Kingdom, the Information Commissioner's Office explains that the Data (Use and Access) Act 2025 removed earlier restrictions on solely automated decisions while requiring safeguards: telling the person about the decision, and enabling them to make representations, obtain human intervention and contest it. In the United States there is no single federal rule and requirements vary by state. In California, for example, the privacy regulator has stated that businesses using automated decisionmaking technology to make significant decisions must comply with its requirements beginning January 1, 2027. Take legal advice before automating decisions of this kind in either country.
For technical readers
- The loop is ordinary code. An orchestration layer sends context to the model, executes the tool calls it requests, appends the results and repeats, with caps on iterations, time and spend.
- Tool design decides reliability. Anthropic suggests investing as much effort in the agent-computer interface as you would in a human-computer interface: clear names, unambiguous parameters and useful error messages.
- Connectors are standardizing. The Model Context Protocol is an open-source standard for connecting AI applications to external systems, which can reduce one-off integration work.
- Context is a budget. Techniques described in the context engineering article include compaction (summarizing a long history and continuing from the summary), structured notes held outside the context window, and loading data just in time through tools instead of up front.
- Authorization belongs downstream. OWASP recommends enforcing authorization in the backend systems, running tools in the context of the specific user, and requiring human approval before significant actions. A prompt is not an access control.
- Evaluation is a release gate. Keep a versioned set of test cases, run it on every change to prompts, tools or model version, and trace every production run.
Frequently asked questions
What is an AI agent in simple terms?
An AI agent is software that is given a goal and a set of tools and works out for itself how to reach the goal. It plans a step, takes an action in a connected system, looks at the result and decides what to do next, repeating until the task is complete or it needs a person.
What is an example of an AI agent in business?
A billing support agent is a typical example. It reads a customer's complaint, looks up the account, order and payment records, works out what went wrong, prepares a refund for a person to approve, and drafts the reply. The steps differ depending on what the records show.
Can an AI agent work without human supervision?
Technically yes, for some tasks, but it is rarely the right starting point. A sensible approach is to have the agent propose actions for a person to approve, then remove approval for low-risk actions once test results and production logs show it is reliable. High-impact actions such as payments usually keep an approval step.
Does an AI agent learn from its mistakes?
Not automatically. The underlying model does not change as your agent runs. An agent can appear to learn if it is designed to save notes to memory and read them on later tasks, but lasting improvement comes from people reviewing failures and updating the instructions, tools and test cases.
What is the difference between an AI agent and agentic AI?
The terms overlap. "AI agent" usually refers to a specific piece of software that pursues a goal with tools. "Agentic AI" is a looser label for the general approach of letting a model direct multi-step work. When you evaluate a product, ignore the label and ask who chooses the steps, which tools it has and where a person approves.
How long does it take to build an AI agent?
It depends less on the AI and more on what it connects to. The main drivers are the number of systems to integrate and whether they have usable APIs, the condition of your data, how many types of case the agent must handle, and how much testing and approval tooling the risk level demands. A narrow agent on well-connected systems is a much smaller project than a broad one on legacy software.
Conclusion and next step
An AI agent is not a smarter chatbot or a digital employee. It is a model working in a loop, choosing its own steps toward a goal you define, acting only through the tools you grant and stopping where you tell it to. That makes it useful for variable, checkable, multi-step work, and a poor choice for vague goals or decisions you cannot afford to get wrong.
A practical next step is to pick one process and run it through the seven-part table above. Write down the goal, list the tools it would need, mark which actions require approval, and gather twenty past cases you could test against. If you can complete that exercise, you have the outline of a project. If you cannot, you have found what to fix first.
Entrant Technologies builds web applications, mobile apps and custom software, including the APIs and admin screens an agent depends on. You can see our development services, or request a quote if you would like to discuss a specific process.