A retry is not the same call again for free. In most agent setups it re-sends everything the agent has read so far, and it may be one of dozens stacked on top of each other.
An agent runtime budget is the most one task may consume before the agent has to stop and hand back: a ceiling on tokens, money, wall-clock time and tool calls. The choice that decides whether it works is scope. The ceiling has to cover the whole task, including every retry, fallback, sub-agent and rerun. If each attempt starts with a fresh allowance, you have a budget per attempt, not per task, and a task that fails five times costs five budgets.
Two definitions first. An AI agent is a program built on a language model that works toward a goal over several steps, calling tools such as a search, a database or an email draft along the way. A token is the unit language models read and write, roughly a piece of a word; Anthropic puts one Claude token at about 3.5 English characters (Anthropic). Providers charge per token, with separate prices for input (what you send) and output (what the model writes) (Anthropic pricing).
What goes into a runtime budget?
Four meters per task, set before the first run:
| Meter | What it caps | Why it matters |
|---|---|---|
| Tokens | Input plus output across all calls | The raw material of the bill |
| Money | Those tokens at your prices, plus paid tools | The number finance sees |
| Time | Wall-clock seconds from start to hand-back | A stuck run is not a slow run |
| Tool calls | Calls per tool, above all those that send, pay or write | Some actions must never repeat |
Using AI agents effectively covers limits like these as part of setting up an agent. This page is about the thing that quietly breaks them: retries.
Why do retries multiply cost?
Three mechanisms, each documented.
1. Every call re-sends the conversation. Anthropic’s documentation says its Messages API “is stateless, which means that you always send the full conversational history to the API” (Anthropic). A retry at step 12 therefore pays again for everything from steps 1 to 11. If the error message is added to the conversation before the next attempt, each retry is also bigger than the one before.
2. Retry layers multiply. Google’s Site Reliability Engineering book warns against retrying at several levels of one system: “a single request at the highest layer may produce a number of attempts as large as the product of the number of attempts at each layer to the lowest layer.” Its example: three layers that each retry three times (four attempts) can turn one user action into 64 attempts (Google SRE book). An agent stack has the same layers: the model SDK, a validator that re-asks when the output does not parse, the agent loop, and the job queue that reruns failed tasks. Some of these layers you never wrote. Anthropic’s official SDKs, for example, retry transient failures such as rate limits and server errors “twice by default” (Anthropic).
3. Caps often reset. Agent frameworks do ship limits, so check what they cover. The OpenAI Agents SDK raises a MaxTurnsExceeded error when a run passes its max_turns limit (OpenAI). Anthropic’s Claude Agent SDK has a dollar cap, maxBudgetUsd, which “counts only the call’s own spend: totals restored from a resumed session don’t count against it, and a /clear starts the budget over” (Anthropic). That is a clear, documented scope for one call. It also means that if your code retries a task by starting a new call, the cap starts again at zero unless you carry the running total yourself.
The same page warns that the SDK’s cost figures “are client-side estimates, not authoritative billing data.” Reconcile against your provider’s own usage reports.
What does this look like in euros? (hypothetical numbers)
Imagine a hypothetical agent that reads supplier invoices and books them. The prices are invented round numbers, not any provider’s list: €2 per million input tokens and €10 per million output tokens. Each task takes 10 steps. Step 1 sends 2,000 tokens, every step adds 2,000 tokens to the conversation, and each step writes 1,000 tokens.
A clean run costs €0.32. Step 10 alone costs €0.05, because it re-sends 20,000 tokens.
Now the output of step 10 fails validation: a field comes back in the wrong format. To keep the arithmetic simple, the retries below do not add the error message to the conversation. In practice they often do, which makes every attempt dearer.
| What happens | Model calls | Cost of the task |
|---|---|---|
| Clean run | 10 | €0.32 |
| Step 10 fails; a validator re-asks up to 3 times inside each of 4 agent-loop attempts | 25 | €1.07 |
| Same, and the job queue reruns the whole task twice, each run with a fresh €1.50 cap | 75 | €3.21 |
| One ceiling of €0.64 for the whole task, retries included | 16 | €0.62, then hands back |
| Same ceiling, retries in the agent loop only, at most 2 | 12 | €0.42, then escalates |
The third row costs ten times the clean run, and no cap fired, because every cap was per run.
Scale it up. At 2,000 invoices a month, a clean month costs €640. If a prompt change pushes 30% of tasks into that failure, the month costs €2,374. Under the last row’s rule it costs €700, and 600 invoices come back to a person with a trace instead of quietly burning money.
Is 30% far-fetched? The Model Portfolio, the book behind the ROUTE framework below, describes a composite case in which one prompt edit pushed roughly a third of an agent’s outputs out of their expected format, and the retry loop would have roughly tripled that route’s monthly cost had a per-route meter not caught it the same morning.
How do you keep retries inside one ceiling?
- Create one budget per task. Tokens, money, seconds and tool calls. Every attempt, retry, fallback model and sub-agent draws from it. Pass the remainder down; never give a sub-step a fresh budget.
- Cap output on every call. Set a maximum output length per call. Then you know the worst case of the next call before you make it: the input you are about to send plus the output cap.
- Check before you call. If what is left is less than that worst case, stop.
- Retry in one layer. Pick the agent loop or the client, not both, and set the others to zero. Retry only error types you have listed as safe, such as a timeout or a rate limit, a fixed number of times. Wait longer after each failure, with some randomness, so many clients do not retry at the same moment. The SRE book puts it as “Always use randomized exponential backoff when scheduling retries.”
- Stop on the unfamiliar. An error you have not seen before ends the run.
- Hand back, do not hide. When the ceiling is hit, return the partial result, the trace and the retry count to a named person.
- Count retries. Log a retry counter per task, and alert when calls per task rise while the number of incoming tasks stays flat.
A reasonable first ceiling is about twice the cost of a measured clean run. Review every task that hits it for the first few weeks, then adjust from what you see.
Where does this rule live in ROUTE?
ROUTE is a framework for running several AI models the way an investor runs a portfolio: every model has a job, a cost class and a fallback, and nothing runs without an owner. It has five steps:
- Register. Catalog every model with its job, cost class, privacy class and retirement status.
- Objective typing. Name what a task needs (depth of reasoning, format, speed, quality, privacy) before you choose a model.
- Utilize policy. Write the routing rules, cascades and fallbacks that decide which model handles which request.
- Track. Measure cost, speed and quality per route, so a waste pattern shows up before the invoice does.
- Evolve. Promote, demote and retire models on a schedule, so the end of a model is planned rather than a scramble.
ROUTE comes from The Model Portfolio, a book by the author of this site, which is available now. You do not need the book to use it. The framework page has the short version, and What is model routing? explains what a route is.
For the invoice agent, the runtime budget belongs in two steps:
- Utilize policy. The retry rules are part of the route, written next to the fallback: which errors retry, how often, in which layer, and the one task ceiling they all share. The book describes the failure this prevents in three short sentences: “The per-call timeout bounded each call. The retry policy bounded each attempt. Nothing bounded the route.” A model cascade that escalates to a bigger model needs the same ceiling around it.
- Track. Cost per task and a retry counter per route. Total monthly spend can hide a loop for days; a per-route meter shows it within hours.
When you set the money meter, take prices from the provider’s own pricing page. The author of this site also built Undominated.ai, an independent index of AI inference prices; it is his project, so weigh it with that in mind.
Try this today (15 minutes)
Pick one agent you run. List every place a retry can happen: the SDK’s settings, any output validator, the agent loop, the job queue, the scheduler. Next to each, write its number of attempts, then multiply them.
Then answer one question: when the task is retried from the top, does its budget start again at zero? If the product is above 10, or the answer is yes, choose one layer to own retries this week and give the task a single ceiling.
Cite this:Agent runtime budget: keep retries inside the same ceiling.Len P. van der Hof. https://lenvanderhof.com/en/blog/agent-runtime-budget/ ·