An AI agent is a software system that works toward a goal you give it by deciding its own next steps: it picks an action or a tool, carries it out, reads the result, and uses that to choose what to do next, repeating until it reaches a stopping point or hands back to a person. A chatbot, by contrast, is a conversational interface, a way to talk to a system in natural language. The two words describe different things: "chatbot" names an interface, while "agent" names how a system operates. A simple assistant waits for you and answers turn by turn; an agent chooses its own steps between your requests. The two can be combined, so one product can be a chat interface and an agent underneath at the same time.

Why this matters now

Companies are giving these systems real permissions, so the distinction has moved from a labelling question to a safety one. On 11 September 2026, Australia's Australian Signals Directorate (ASD) published guidance on agentic AI harnesses, the software layer above the model that actually runs an agent's tool calls and enforces its permissions, and warned that some risks cannot be addressed within the model alone. For a company weighing one, that means judging the whole system around the model, not the model alone.

The loop, in plain terms

A modern business agent runs a simple loop. You give it a goal. It chooses a step or a tool. It observes the result. It repeats. It stops when the goal is met, a budget runs out, or it reaches a point where a person should decide. Anthropic's engineering guide Building effective agents (19 December 2024) defines agents as systems where the model "dynamically direct[s] their own processes and tool usage," in contrast to workflows, where "LLMs and tools are orchestrated through predefined code paths." The model is the decision-maker inside the loop; the harness, the surrounding software, is what executes the tool calls, holds the credentials, and enforces what the agent is and is not allowed to do. Terminology varies across the industry; this article uses Anthropic's architectural distinction.

One request, a few ways to handle it

Picture a customer who asks to change the delivery address on an order. Whether that change is allowed at all is set by company policy and authorization; it is not something the system should assume. The example is hypothetical, meant only to show the shape of each approach.

A baseline first: ordinary automation is a fixed script that applies a rule the company already approved, and it may use no AI at all. An LLM workflow runs along code paths fixed in advance that can branch and can call a model at set points. An agent chooses its next step itself. Applied to the address change, the three look like this:

  • A conversational assistant answers. It explains the address-change policy and points to the right page, and may look up the order through a connected tool. The conversation is the product; a person still makes any change.
  • A workflow runs. Its code paths are fixed in advance but can branch and call a model, for example to read and classify the request, so a shipped order and an unshipped one follow different designed branches, each ending where policy says it should.
  • An agent chooses the next step. Within the permissions and policy it was given, it checks the order, sees it has partly shipped, proposes splitting the change across shipped and unshipped items, and routes the consequential part to a person for approval. The path was not fixed in advance; the agent chose it from what it found.

Only the last is an agent in the narrow sense used here, and a conversational assistant, ordinary automation, or a workflow is often the right answer.

 Conversational assistantWorkflowAgent
Who picks the stepsA person, in chatCode paths set aheadThe model, at run time
Path through a taskVaries by talkBranches defined aheadVaries with findings
Best whenAnswers, guidanceWell-specified rulesNeeds model-driven choices
Main trade-offLimited actionExceptions need designed handlingCan add cost and latency

These are three ways to build the same task, not mutually exclusive product classes; a real product can combine them. Ordinary automation, a fixed script with no AI, is a simpler fourth option. A conversational assistant can also use tools, and a well-built agent can pause for a person.

What counts as an agent, and what does not

A few boundaries keep the term useful. An agent is not the model on its own: a language model generates text and can produce structured output, but it cannot reach your systems until a harness gives it tools and permissions. A chat interface alone does not make an agent, and neither does connecting a single tool. A decision-making component, such as Jev, is one piece of an agent, not a complete agent on its own. The broader term "AI agent" also covers older systems that were never language models; this article is about the modern, LLM-powered kind a company is most likely to deploy now. And autonomy is bounded: an agent does not have to run in the background, keep a persistent memory, or hold permission to take consequential actions. A narrow, supervised pilot is a practical way to start.

Where an agent helps, and where it does not

The tasks where an agent earns its place share a shape: the goal is clear, but reaching it needs choices the model has to make case by case, rather than a path you can specify in advance. Research and document gathering, internal support that looks across several systems, and coding by exploring a codebase and proposing a change are common examples, and none of them come with a product guarantee. Where the task is well-defined, simpler options usually win: a conversational assistant for questions with a standard answer, and ordinary automation, which is often cheaper and more predictable, for steps you can specify up front. Compared with a simpler solution, an agent can add cost and latency, and because it acts in a loop, a wrong step early can compound. There is no guaranteed return and no guaranteed reliability; measure against a simpler baseline before you widen its reach.

Controls to assess before deployment

Because an agent acts, not just answers, these controls matter as much as the capability. They reduce risk rather than remove it; weigh them before you deploy.

  • Permissions and data boundaries. Give the agent the narrowest access its task needs, and decide in your own code, not by trusting the model, which actions it may take.
  • Validate tools independently. Check that an action is allowed in your access rules before it runs, rather than letting the model's say-so authorise it.
  • Human review of consequential actions. Require a person to approve anything destructive, financial, or hard to reverse.
  • Logs, with access control and redaction. Keep records of prompts, tool calls, and decisions, but limit who can read them and redact secrets, since full logs can hold sensitive data.
  • Stop conditions, budgets, and testing. Cap iterations and spend, and run adversarial tests against your own setup before you widen its reach.

One risk deserves its own line: prompt injection. Because an agent reads outside content, web pages, emails, documents, a planted instruction can steer it, and even a read-only agent can be made to disclose data it can see. Whether you are building an agent or designing the AI governance around it, decide in advance what the system may touch, and assume the content it reads may try to steer it.

Frequently asked questions

What is the difference between an AI agent and a chatbot?
A chatbot is a way to talk to a system in natural language, so the word describes an interface. An AI agent is a system that chooses its own next steps toward a goal, so the word describes how it operates. That is why they are not opposites: a simple assistant waits and answers turn by turn, an agent chooses steps on its own, and one product can be a chat interface and an agent underneath at the same time. The useful question is not chatbot or agent, but whether the system directs its own steps or waits for yours.

Is an AI agent the same as a large language model?
No. A large language model generates text and can produce structured output, but on its own it cannot reach your systems. An AI agent is the model plus a harness, the surrounding software that gives it tools, holds credentials, executes its tool calls, and enforces its permissions. The model is the decision-maker inside the loop; the harness is what lets it act, and where most of the controls live.

When does a company actually need an AI agent instead of a chatbot or automation?
When reaching the goal needs choices the model has to make case by case, rather than a path you can specify in advance. If a question has a standard answer, a conversational assistant is enough; if the task is well-defined, ordinary automation is often cheaper and more predictable. An agent can add cost and latency compared with a simpler solution, and a wrong step early can compound, so there is no guaranteed return or reliability. Pilot narrowly, measure against a simpler baseline, and keep human review on consequential actions.

Sources