---
title: "Agent runtime budget: one ceiling, retries included | Len P. van der Hof"
description: An agent runtime budget caps tokens, money, time and tool calls per task. Why retries multiply cost, and how to keep them inside one ceiling.
image: "https://lenvanderhof.com/media/generated/blog-hero-agent-runtime-budget-v1.003b4fbaee34.wide.webp"
---

[AI Systems](https://lenvanderhof.com/en/blog/category/ai-systems/) Research guide

# Agent runtime budget: keep retries inside the same ceiling

A cap that resets on every retry is a cap per attempt, not per task.

Len P. van der HofPublished 3 October 20268 min read

Every bounce counts against the same beam.

Direct answer

An agent runtime budget is the most one task may consume before the agent must stop and hand back: a ceiling on tokens, money, wall-clock time and tool calls. The choice that decides whether it works is scope. The ceiling must cover the whole task, including every retry, fallback, sub-agent and rerun. Retries multiply cost because each call re-sends the conversation so far and retry layers stack. If each attempt starts with a fresh allowance, a task that fails five times costs five budgets.

## Key takeaways

- Budget per task, not per call: every retry, fallback and rerun draws from the same ceiling.
- A retry re-sends the whole conversation, and stacked retry layers multiply: three layers of four attempts can make 64 attempts.
- Cap output length on every call so you can check the worst case before you make it.
- Retry in one layer only, for error types you listed, a fixed number of times.
- When the ceiling is hit, stop, hand back the trace and let a named person decide.

A retry is not the same call again for free. In most agent setups it re-sends everything the agent has read so far, and it may be one of dozens stacked on top of each other.

**An agent runtime budget is the most one task may consume before the agent has to stop and hand back: a ceiling on tokens, money, wall-clock time and tool calls.** The choice that decides whether it works is scope. The ceiling has to cover the whole task, including every retry, fallback, sub-agent and rerun. If each attempt starts with a fresh allowance, you have a budget per attempt, not per task, and a task that fails five times costs five budgets.

Two definitions first. An **AI agent** is a program built on a language model that works toward a goal over several steps, calling tools such as a search, a database or an email draft along the way. A **token** is the unit language models read and write, roughly a piece of a word; Anthropic puts one Claude token at about 3.5 English characters ([Anthropic](https://platform.claude.com/docs/en/about-claude/glossary)). Providers charge per token, with separate prices for input (what you send) and output (what the model writes) ([Anthropic pricing](https://platform.claude.com/docs/en/about-claude/pricing)).

## What goes into a runtime budget?

Four meters per task, set before the first run:

MeterWhat it capsWhy it mattersTokensInput plus output across all callsThe raw material of the billMoneyThose tokens at your prices, plus paid toolsThe number finance seesTimeWall-clock seconds from start to hand-backA stuck run is not a slow runTool callsCalls per tool, above all those that send, pay or writeSome actions must never repeat

[Using AI agents effectively](https://lenvanderhof.com/en/blog/using-ai-agents-effectively/) covers limits like these as part of setting up an agent. This page is about the thing that quietly breaks them: retries.

## Why do retries multiply cost?

Three mechanisms, each documented.

**1. Every call re-sends the conversation.** Anthropic’s documentation says its Messages API “is stateless, which means that you always send the full conversational history to the API” ([Anthropic](https://platform.claude.com/docs/en/build-with-claude/working-with-messages)). A retry at step 12 therefore pays again for everything from steps 1 to 11. If the error message is added to the conversation before the next attempt, each retry is also bigger than the one before.

**2. Retry layers multiply.** Google’s Site Reliability Engineering book warns against retrying at several levels of one system: “a single request at the highest layer may produce a number of attempts as large as the product of the number of attempts at each layer to the lowest layer.” Its example: three layers that each retry three times (four attempts) can turn one user action into 64 attempts ([Google SRE book](https://sre.google/sre-book/addressing-cascading-failures/)). An agent stack has the same layers: the model SDK, a validator that re-asks when the output does not parse, the agent loop, and the job queue that reruns failed tasks. Some of these layers you never wrote. Anthropic’s official SDKs, for example, retry transient failures such as rate limits and server errors “twice by default” ([Anthropic](https://platform.claude.com/docs/en/api/errors)).

**3. Caps often reset.** Agent frameworks do ship limits, so check what they cover. The OpenAI Agents SDK raises a `MaxTurnsExceeded` error when a run passes its `max_turns` limit ([OpenAI](https://openai.github.io/openai-agents-python/running_agents/)). Anthropic’s Claude Agent SDK has a dollar cap, `maxBudgetUsd`, which “counts only the call’s own spend: totals restored from a resumed session don’t count against it, and a /clear starts the budget over” ([Anthropic](https://code.claude.com/docs/en/agent-sdk/cost-tracking)). That is a clear, documented scope for one call. It also means that if your code retries a task by starting a new call, the cap starts again at zero unless you carry the running total yourself.

The same page warns that the SDK’s cost figures “are client-side estimates, not authoritative billing data.” Reconcile against your provider’s own usage reports.

## What does this look like in euros? (hypothetical numbers)

Imagine a hypothetical agent that reads supplier invoices and books them. The prices are invented round numbers, not any provider’s list: €2 per million input tokens and €10 per million output tokens. Each task takes 10 steps. Step 1 sends 2,000 tokens, every step adds 2,000 tokens to the conversation, and each step writes 1,000 tokens.

A clean run costs €0.32. Step 10 alone costs €0.05, because it re-sends 20,000 tokens.

Now the output of step 10 fails validation: a field comes back in the wrong format. To keep the arithmetic simple, the retries below do not add the error message to the conversation. In practice they often do, which makes every attempt dearer.

What happensModel callsCost of the taskClean run10€0.32Step 10 fails; a validator re-asks up to 3 times inside each of 4 agent-loop attempts25€1.07Same, and the job queue reruns the whole task twice, each run with a fresh €1.50 cap75€3.21One ceiling of €0.64 for the whole task, retries included16€0.62, then hands backSame ceiling, retries in the agent loop only, at most 212€0.42, then escalates

The third row costs ten times the clean run, and no cap fired, because every cap was per run.

Scale it up. At 2,000 invoices a month, a clean month costs €640. If a prompt change pushes 30% of tasks into that failure, the month costs €2,374. Under the last row’s rule it costs €700, and 600 invoices come back to a person with a trace instead of quietly burning money.

Is 30% far-fetched? *The Model Portfolio*, the book behind the ROUTE framework below, describes a composite case in which one prompt edit pushed roughly a third of an agent’s outputs out of their expected format, and the retry loop would have roughly tripled that route’s monthly cost had a per-route meter not caught it the same morning.

## How do you keep retries inside one ceiling?

1. **Create one budget per task.** Tokens, money, seconds and tool calls. Every attempt, retry, fallback model and sub-agent draws from it. Pass the remainder down; never give a sub-step a fresh budget.
2. **Cap output on every call.** Set a maximum output length per call. Then you know the worst case of the next call before you make it: the input you are about to send plus the output cap.
3. **Check before you call.** If what is left is less than that worst case, stop.
4. **Retry in one layer.** Pick the agent loop or the client, not both, and set the others to zero. Retry only error types you have listed as safe, such as a timeout or a rate limit, a fixed number of times. Wait longer after each failure, with some randomness, so many clients do not retry at the same moment. The SRE book puts it as “Always use randomized exponential backoff when scheduling retries.”
5. **Stop on the unfamiliar.** An error you have not seen before ends the run.
6. **Hand back, do not hide.** When the ceiling is hit, return the partial result, the trace and the retry count to a named person.
7. **Count retries.** Log a retry counter per task, and alert when calls per task rise while the number of incoming tasks stays flat.

A reasonable first ceiling is about twice the cost of a measured clean run. Review every task that hits it for the first few weeks, then adjust from what you see.

## Where does this rule live in ROUTE?

**ROUTE** is a framework for running several AI models the way an investor runs a portfolio: every model has a job, a cost class and a fallback, and nothing runs without an owner. It has five steps:

1. **Register.** Catalog every model with its job, cost class, privacy class and retirement status.
2. **Objective typing.** Name what a task needs (depth of reasoning, format, speed, quality, privacy) before you choose a model.
3. **Utilize policy.** Write the routing rules, cascades and fallbacks that decide which model handles which request.
4. **Track.** Measure cost, speed and quality per route, so a waste pattern shows up before the invoice does.
5. **Evolve.** Promote, demote and retire models on a schedule, so the end of a model is planned rather than a scramble.

ROUTE comes from *[The Model Portfolio](https://lenvanderhof.com/books/the-model-portfolio/)*, a book by the author of this site, which is available now. You do not need the book to use it. The [framework page](https://lenvanderhof.com/frameworks/route/) has the short version, and [What is model routing?](https://lenvanderhof.com/en/blog/what-is-model-routing/) explains what a route is.

For the invoice agent, the runtime budget belongs in two steps:

- **Utilize policy.** The retry rules are part of the route, written next to the fallback: which errors retry, how often, in which layer, and the one task ceiling they all share. The book describes the failure this prevents in three short sentences: “The per-call timeout bounded each call. The retry policy bounded each attempt. Nothing bounded the route.” A [model cascade](https://lenvanderhof.com/en/blog/what-is-a-model-cascade/) that escalates to a bigger model needs the same ceiling around it.
- **Track.** Cost per task and a retry counter per route. Total monthly spend can hide a loop for days; a per-route meter shows it within hours.

When you set the money meter, take prices from the provider’s own pricing page. The author of this site also built [Undominated.ai](https://lenvanderhof.com/undominated/), an independent index of AI inference prices; it is his project, so weigh it with that in mind.

## Try this today (15 minutes)

Pick one agent you run. List every place a retry can happen: the SDK’s settings, any output validator, the agent loop, the job queue, the scheduler. Next to each, write its number of attempts, then multiply them.

Then answer one question: when the task is retried from the top, does its budget start again at zero? If the product is above 10, or the answer is yes, choose one layer to own retries this week and give the task a single ceiling.

Cite this:Agent runtime budget: keep retries inside the same ceiling.Len P. van der Hof. [https://lenvanderhof.com/en/blog/agent-runtime-budget/](https://lenvanderhof.com/en/blog/agent-runtime-budget/) · Published 3 October 2026.

## Terminology

- [ROUTE](https://lenvanderhof.com/glossary/route/)

## Sources

1. [Glossary: Tokens](https://platform.claude.com/docs/en/about-claude/glossary) · Anthropic (Claude Platform documentation)
2. [Pricing](https://platform.claude.com/docs/en/about-claude/pricing) · Anthropic (Claude Platform documentation)
3. [Using the Messages API](https://platform.claude.com/docs/en/build-with-claude/working-with-messages) · Anthropic (Claude Platform documentation)
4. [Errors](https://platform.claude.com/docs/en/api/errors) · Anthropic (Claude Platform documentation)
5. [Addressing Cascading Failures (Site Reliability Engineering, chapter 22)](https://sre.google/sre-book/addressing-cascading-failures/) · Google (Mike Ulrich, in Site Reliability Engineering)
6. [Running agents](https://openai.github.io/openai-agents-python/running_agents/) · OpenAI Agents SDK documentation
7. [Track cost and usage](https://code.claude.com/docs/en/agent-sdk/cost-tracking) · Anthropic (Claude Agent SDK documentation)
8. [ROUTE (framework)](https://lenvanderhof.com/frameworks/route/)
9. [The Model Portfolio](https://lenvanderhof.com/books/the-model-portfolio/)
10. [Using AI agents effectively: one job, a ceiling, a log](https://lenvanderhof.com/en/blog/using-ai-agents-effectively/)
11. [What is model routing? A policy, not a default](https://lenvanderhof.com/en/blog/what-is-model-routing/)

## Further reading

- [Using AI agents effectively: one job, a ceiling, a log](https://lenvanderhof.com/en/blog/using-ai-agents-effectively/)
- [What is a model cascade? Escalate with a stop rule](https://lenvanderhof.com/en/blog/what-is-a-model-cascade/)
- [ROUTE](https://lenvanderhof.com/frameworks/route/)

About the author

## [Len P. van der Hof](https://lenvanderhof.com/en/authors/len-p-van-der-hof/)

Entrepreneur, AI Innovator and Venture Builder

Len P. van der Hof builds practical AI systems, digital ventures and evidence-informed tools for founders.

```json
{
	"@context": "https://schema.org",
	"@graph": [
		{
			"@type": "Person",
			"@id": "https://lenvanderhof.com/#person",
			"name": "Len P. van der Hof",
			"alternateName": [
				"Len van der Hof",
				"L.P. van der Hof",
				"Leendert Pieter van der Hof"
			],
			"honorificSuffix": "MSc",
			"url": "https://lenvanderhof.com/",
			"image": [
				"https://lenvanderhof.com/photos/len-portrait-1.jpg",
				"https://lenvanderhof.com/photos/len-portrait-2.jpg",
				"https://lenvanderhof.com/photos/len-portrait-3.jpg",
				"https://lenvanderhof.com/photos/len-portrait-4.jpg",
				"https://lenvanderhof.com/photos/len-speaking.jpg",
				"https://lenvanderhof.com/photos/len-hero.jpg"
			],
			"jobTitle": "Entrepreneur, AI Innovator and Venture Builder",
			"description": "Len P. van der Hof, MSc, is a Dutch entrepreneur and AI innovator in Zwijndrecht. He builds ReasonKit, MindSesh, Undominated.ai, books under his name, the fiction imprint LPH98.lifestyle, and technology ventures through LPH98.ventures. Eleven titles in Systems for the Strategic Self are available now, in English and Dutch.",
			"address": {
				"@type": "PostalAddress",
				"addressLocality": "Zwijndrecht",
				"addressCountry": "NL"
			},
			"alumniOf": {
				"@type": "CollegeOrUniversity",
				"name": "Rotterdam School of Management, Erasmus University"
			},
			"knowsAbout": [
				"Artificial intelligence",
				"AI agents",
				"Agentic AI systems",
				"LLM routing",
				"SEO",
				"Generative engine optimization",
				"Answer engine optimization",
				"Venture building",
				"Founder performance",
				"Founder psychology",
				"Evidence-based decision-making"
			],
			"sameAs": [
				"https://www.linkedin.com/in/lenvanderhof/",
				"https://x.com/LenvanderHof",
				"https://www.youtube.com/channel/UCTG20buKqYYbitqqf7l3zJA",
				"https://www.instagram.com/Lenvanderhof/",
				"https://www.threads.com/@lenvanderhof",
				"https://github.com/Lenvanderhof",
				"https://huggingface.co/LPH98",
				"https://www.npmjs.com/~lenvanderhof",
				"https://www.goodreads.com/author/show/70983905.Len_P_van_der_Hof",
				"https://www.amazon.com/author/lenvanderhof",
				"https://www.bol.com/nl/nl/b/len-p-van-der-hof-msc/609879394/",
				"https://bsky.app/profile/lenvanderhof.com",
				"https://mastodon.social/@Lenvanderhof",
				"https://crates.io/users/Lenvanderhof",
				"https://cursor.com/@Lenvanderhof",
				"https://medium.com/@Lenvanderhof",
				"https://gitlab.com/Lenvanderhof",
				"https://hub.docker.com/u/lenvanderhof/",
				"https://dev.to/lenvanderhof",
				"https://www.facebook.com/Lenvanderhof",
				"https://soundcloud.com/Lenvanderhof"
			],
			"affiliation": [
				{
					"@id": "https://lenvanderhof.com/#publisher"
				},
				{
					"@id": "https://lenvanderhof.com/#mindsesh"
				},
				{
					"@id": "https://lenvanderhof.com/#lifestyle"
				}
			]
		},
		{
			"@type": "WebSite",
			"@id": "https://lenvanderhof.com/#website",
			"url": "https://lenvanderhof.com/",
			"name": "Len P. van der Hof",
			"description": "Len P. van der Hof, MSc, is a Dutch entrepreneur and AI innovator in Zwijndrecht. He builds ReasonKit, MindSesh, Undominated.ai, books under his name, the fiction imprint LPH98.lifestyle, and technology ventures through LPH98.ventures. Eleven titles in Systems for the Strategic Self are available now, in English and Dutch.",
			"inLanguage": [
				"en",
				"nl"
			],
			"publisher": {
				"@id": "https://lenvanderhof.com/#person"
			}
		},
		{
			"@type": "Organization",
			"@id": "https://lenvanderhof.com/#publisher",
			"name": "LPH98.ventures",
			"url": "https://lph98.ventures",
			"founder": {
				"@id": "https://lenvanderhof.com/#person"
			}
		},
		{
			"@type": "Organization",
			"@id": "https://lenvanderhof.com/#mindsesh",
			"name": "MindSesh",
			"url": "https://mindsesh.net",
			"founder": {
				"@id": "https://lenvanderhof.com/#person"
			}
		},
		{
			"@type": "SoftwareApplication",
			"@id": "https://lenvanderhof.com/#reasonkit",
			"name": "ReasonKit",
			"url": "https://reasonkit.sh",
			"creator": {
				"@id": "https://lenvanderhof.com/#person"
			}
		},
		{
			"@type": "SoftwareApplication",
			"@id": "https://lenvanderhof.com/#undominated",
			"name": "Undominated.ai",
			"url": "https://undominated.ai",
			"creator": {
				"@id": "https://lenvanderhof.com/#person"
			}
		},
		{
			"@type": "Organization",
			"@id": "https://lenvanderhof.com/#lifestyle",
			"name": "LPH98.lifestyle",
			"url": "https://lph98.lifestyle",
			"founder": {
				"@id": "https://lenvanderhof.com/#person"
			}
		},
		{
			"@type": "ImageObject",
			"@id": "https://lenvanderhof.com/en/blog/agent-runtime-budget/#primaryimage",
			"url": "https://lenvanderhof.com/media/generated/blog-hero-agent-runtime-budget-v1.003b4fbaee34.wide.webp",
			"contentUrl": "https://lenvanderhof.com/media/generated/blog-hero-agent-runtime-budget-v1.003b4fbaee34.wide.webp",
			"representativeOfPage": true
		},
		{
			"@type": "BreadcrumbList",
			"@id": "https://lenvanderhof.com/en/blog/agent-runtime-budget/#breadcrumb",
			"itemListElement": [
				{
					"@type": "ListItem",
					"position": 1,
					"name": "Home",
					"item": "https://lenvanderhof.com/"
				},
				{
					"@type": "ListItem",
					"position": 2,
					"name": "Blog",
					"item": "https://lenvanderhof.com/en/blog/"
				},
				{
					"@type": "ListItem",
					"position": 3,
					"name": "AI Systems",
					"item": "https://lenvanderhof.com/en/blog/category/ai-systems/"
				},
				{
					"@type": "ListItem",
					"position": 4,
					"name": "Agent runtime budget: keep retries inside the same ceiling",
					"item": "https://lenvanderhof.com/en/blog/agent-runtime-budget/"
				}
			]
		},
		{
			"@type": "WebPage",
			"@id": "https://lenvanderhof.com/en/blog/agent-runtime-budget/#webpage",
			"url": "https://lenvanderhof.com/en/blog/agent-runtime-budget/",
			"name": "Agent runtime budget: keep retries inside the same ceiling",
			"description": "An agent runtime budget caps tokens, money, time and tool calls per task. Why retries multiply cost, and how to keep them inside one ceiling.",
			"isPartOf": {
				"@id": "https://lenvanderhof.com/#website"
			},
			"primaryImageOfPage": {
				"@id": "https://lenvanderhof.com/en/blog/agent-runtime-budget/#primaryimage"
			},
			"breadcrumb": {
				"@id": "https://lenvanderhof.com/en/blog/agent-runtime-budget/#breadcrumb"
			},
			"inLanguage": "en-GB"
		},
		{
			"@type": "BlogPosting",
			"@id": "https://lenvanderhof.com/en/blog/agent-runtime-budget/#article",
			"mainEntityOfPage": {
				"@id": "https://lenvanderhof.com/en/blog/agent-runtime-budget/#webpage"
			},
			"headline": "Agent runtime budget: keep retries inside the same ceiling",
			"description": "An agent runtime budget caps tokens, money, time and tool calls per task. Why retries multiply cost, and how to keep them inside one ceiling.",
			"datePublished": "2026-10-03T14:00:00.000Z",
			"author": {
				"@id": "https://lenvanderhof.com/#person"
			},
			"publisher": {
				"@id": "https://lenvanderhof.com/#person"
			},
			"image": [
				"https://lenvanderhof.com/media/generated/blog-hero-agent-runtime-budget-v1.003b4fbaee34.square.webp",
				"https://lenvanderhof.com/media/generated/blog-hero-agent-runtime-budget-v1.003b4fbaee34.landscape.webp",
				"https://lenvanderhof.com/media/generated/blog-hero-agent-runtime-budget-v1.003b4fbaee34.wide.webp"
			],
			"articleSection": "AI Systems",
			"inLanguage": "en-GB"
		}
	]
}
```
