AI Systems Research guide

When not to use an AI agent: five signs, and what to use instead

An agent earns its cost only when the next step depends on what it finds. Most tasks never get there.

A hand saw and mitre box cutting one oak board on a workbench while a large CNC router waits under a dust sheet behind
Some cuts do not need the big machine.

Direct answer

Do not use an AI agent when you can write down every step in advance, when a single prompt does the job, when nobody can say what a correct result looks like, when you cannot limit what it can touch, or when a mistake cannot be undone and no person approves first. Use a script, one model call, a draft-only assistant, or a person instead. An agent, a system that chooses its own next step and tools, earns its cost only when the next step depends on what it finds.

Do not use an AI agent when a simpler tool can do the job or when nobody can safely stop its mistakes. In practice that means five signs: you can write down every step in advance, one prompt does the job, nobody can say what a correct result looks like, you cannot limit what it can touch, or a mistake cannot be undone and no named person approves it first. Use a script, a single model call, a draft-only assistant, or a person instead.

An AI agent is a system that pursues a goal by choosing its own next step and its own tools, based on what it finds along the way. That freedom is the whole point, and the whole cost. What is an AI agent? has the full definition.

Even the companies that sell agents say this

Anthropic’s engineering guide Building effective agents (December 2024) is blunt: “When building applications with LLMs, we recommend finding the simplest solution possible, and only increasing complexity when needed. This might mean not building agentic systems at all.” (An LLM, or large language model, is the kind of AI behind ChatGPT and Claude.) The same guide warns that “the autonomous nature of agents means higher costs, and the potential for compounding errors.”

OpenAI’s A practical guide to building agents (2025) recommends agents for three kinds of work: decisions that need nuanced judgment, rule sets too tangled to maintain, and work that depends on reading messy text such as emails and documents. Then it adds: “Before committing to building an agent, validate that your use case can meet these criteria clearly. Otherwise, a deterministic solution may suffice.” Deterministic means the same input always leads to the same steps, as in a script.

Five signs, and what to use instead

SignUse this insteadExample
You can write down every step before the runA script or fixed workflowEvery Friday, copy approved invoices from the accounting tool into the cash-flow sheet
The job is one transformation of input you already haveOne model callSummarise this contract; translate this email
Nobody can say what a correct result looks likeA person, until someone writes the check”Handle customer operations”
You cannot limit its tools and access to what the job needsNo agent yet: a draft-only assistant without send or delete rightsThe only mail connection available can send as the founder
A mistake cannot be undone, and no named person approves firstA person takes the step; the agent may prepare itIssuing a refund, paying a supplier, deleting records

1. You can write down every step

A fixed workflow can still include conditions, checks, and error handling. What it does not do is invent the next step. Anthropic’s guide draws the same line: workflows are “systems where LLMs and tools are orchestrated through predefined code paths”, while agents “dynamically direct their own processes and tool usage”. If you can name every branch in advance, you want the first. An agentic workflow is not an automation goes deeper on that split.

2. One prompt does the job

Summarising, translating, classifying, or drafting from a document you already have is one call to a model, not a loop. Anthropic again: “For many applications, however, optimizing single LLM calls with retrieval and in-context examples is usually enough.” Retrieval means looking up the right document first; in-context examples means showing the model two or three good answers inside the prompt.

3. Nobody can say what a correct result looks like

“Handle customer operations” gives a reviewer nothing to pass or fail. “Draft a reply for each ticket tagged refund, or say why it cannot be drafted” does. If you cannot tell whether an output is right, letting a model take more steps does not add that judgment. It adds more output to check.

4. You cannot limit what it can touch

Separate the limits you wrote from the limits your software enforces. The Agent Dream Team, one of Len’s books (more on it below), names three levels: the platform itself blocks tools outside the allowed set; the agent can only propose, and a separate approval step executes; or only the instructions say “do not”, while the tool stays connected. For the third level, the book’s advice is to keep dangerous tools disconnected from that agent, or to route the work through a separate approval step. A “draft-only” assistant that holds a send credential still has a route past its instructions.

5. The mistake cannot be undone

OpenAI’s guide says actions that are “sensitive, irreversible, or have high stakes should trigger human oversight until confidence in the agent’s reliability grows. Examples include canceling user orders, authorizing large refunds, or making payments.”

The timing matters. A person who approves a refund before it goes out owns a decision. A person who reads the log afterwards can only investigate one. What an agent may never do lists four categories this site keeps with a human: spending above a ceiling, speaking as the company, changing production data, and closing a judgment the founder still owns.

Watch for one quiet version of this sign: retries. If a “send” times out, you do not know whether the message went. An agent that simply tries again may send it twice.

A worked example: the returns desk

Imagine a hypothetical four-person online shop that sells cookware. Returns arrive by email, and the founder wants “an agent for returns”. Walk the requests through the five signs.

Standard returns. The item is inside the return window, unused, and the order number is in the email. The policy already lists every step: check the order date, create a return label, send the template. That is sign 1. A fixed workflow handles it, and when it breaks, you can see which step broke.

“Can you explain the return policy in German?” One model call with the policy text attached. Sign 2.

“It arrived dented, photos attached.” Now the next step depends on what the system finds: the photos, the order history, the courier’s delivery record. Refund, replacement, or a partial refund? This is genuinely agent territory. But the outcome is a refund, which is money and cannot be undone (sign 5), and the only payment connection the shop has can refund any amount (sign 4).

So the damaged-goods assistant reads the case and drafts a recommendation that cites the policy clause. A person approves it and issues the refund.

The result: one script, one prompt, one draft-only assistant, and no agent with refund power. That is not a failure to adopt AI. It is the design.

Not sure? Ask the six ROSTER questions

ROSTER is a framework from The Agent Dream Team, one of Len’s books (available now). It is written for solo founders and small teams who run several specialist AI agents the way a manager runs a team with job descriptions. Each letter is a decision you make for every agent. Turned into questions, any blank answer means: not an agent yet.

  • R, Roles. Can you give it one job, with a one-page charter (a job description) that says what it owns and what it must hand back to you? The book’s working rule: if the job description runs past two sentences, the scope is too broad.
  • O, Objectives. Can you name a result you could grade on Friday without an argument?
  • S, Skills and tools. Can you give it only the tools and permissions this job needs, and does your platform enforce that?
  • T, Triggers. Can you name the event that starts it, so it does not run whenever someone happens to remember it?
  • E, Evaluation. Is there a scorecard, and a person who applies it before anything reaches a customer?
  • R, Rotation. Do you know when you would rework or retire it, for example when it keeps failing its scorecard or the model it runs on is withdrawn?

You do not need the book to use these six questions. The ROSTER framework page has the short version.

Applied to the returns desk: the damaged-goods assistant can answer Roles, Objectives, Triggers, Evaluation, and Rotation. It cannot answer Skills and tools while the only payment connection can refund any amount. That one blank is why it stays draft-only.

Try this today (15 minutes)

Pick one task you were about to hand to an agent, or already have.

  1. Run it past the five signs. Write down which alternative from the table fits, if any.
  2. If it still needs an agent, write down the first action it must stop before, and the name of the person who approves that action.
  3. Check whether your platform actually blocks that action, or whether only the instructions say so. If only the instructions say so, disconnect that tool today.

If the task passes all five signs, Using AI agents effectively covers how to run it: one job, a ceiling on steps and spend, and a log you can read.

Cite this:When not to use an AI agent: five signs, and what to use instead.Len P. van der Hof. https://lenvanderhof.com/en/blog/when-not-to-use-an-agent/ ·

Terminology

Sources

  1. Building effective agents · Anthropic
  2. A practical guide to building agents · OpenAI
  3. What is an AI agent?
  4. What an agent may never do
  5. ROSTER framework
  6. The Agent Dream Team

Further reading

Markdown for LLMs