---
title: "Agent shadow mode: test proposals before action rights | Len P. van der Hof"
description: "Agent shadow mode: the agent proposes, a person still acts, and you compare. What to measure, how many cases you need, and a promotion rule you can copy."
image: "https://lenvanderhof.com/media/generated/blog-hero-agent-shadow-mode-v1.8ea88ca10401.wide.webp"
---

[AI Systems](https://lenvanderhof.com/en/blog/category/ai-systems/) Research guide

# Agent shadow mode: evaluate proposals before granting action rights

Let the agent be wrong on paper first, where a mistake costs a comparison instead of a customer.

Len P. van der HofPublished 4 October 20268 min read

The learner drives. The second pedal stays until the record says otherwise.

Direct answer

Agent shadow mode means an AI agent receives real work and records what it would do, but has no permission to do it. A person or the existing process still acts, and afterwards you compare the two. The idea comes from software shadow testing, where a new version gets a copy of live requests and its answers are discarded or logged. For agents, the write tools must be switched off, not merely ignored. Grant action rights one narrow class of cases at a time, under a written rule.

## Key takeaways

- Shadow mode: real inputs, no write rights, a comparison afterwards.
- Switch the write tools off. Discarding the output is not enough when the output is an action.
- Let the person decide before seeing the proposal, or you are measuring persuasion.
- Zero serious errors in 300 cases still allows a true rate of up to about 1 in 100.
- Promote one narrow class at a time, and write down what sends it back.

**Agent shadow mode means an AI agent receives real work and writes down what it would do, but has no permission to do it. A person or the existing process still acts, and afterwards you compare the two.** You grant the agent action rights only when the comparison clears a rule you wrote in advance.

An agent that made no serious mistake in 50 cases has shown you very little. This page covers why, what to measure instead, and a promotion rule you can copy.

An [AI agent](https://lenvanderhof.com/glossary/ai-agent/) here means software that takes a goal and uses tools, such as email, a database, or a planning system, over several steps without a person approving each one. Action rights are its permissions to change things in those systems.

## Where does the term come from?

From software deployment. Amazon’s machine-learning service, SageMaker, offers shadow tests: it “automatically deploys the new variant in shadow mode and routes a copy of the inference requests to it in real time within the same endpoint. Only the responses of the production variant are returned to the calling application.” In plain words: the new version receives a copy of the real questions sent to the live model, and users only ever get answers from the version already in production.

Istio, a service mesh (software that routes traffic between the parts of a cloud application), describes the same move: “Traffic mirroring, also called shadowing, is a powerful concept that allows feature teams to bring changes to production with as little risk as possible.” The copied requests are fire and forget. Their responses are thrown away.

That works because a prediction is only an answer. You can discard an answer. An agent’s output is an action: an email sent, a booking moved, a refund paid. A mirrored agent with a live send tool does not shadow. It sends.

So agent shadow mode is stricter than its software cousin. The agent gets the same inputs and the same read access. Every tool that writes is replaced by a proposal tool that records the intended action and does nothing else.

## What goes into a proposal?

Each proposal is a short record with the same case ID the person’s real decision will carry:

1. **The action.** The exact tool and arguments it would have used: “move delivery 5530 to Thursday 14:00 to 16:00.”
2. **The reason.** The message, record, or rule it relied on.
3. **Escalate or not.** Whether it would have handed the case to a person instead.
4. **The time.** When it proposed, so you can compare speed as well.

The person’s decision is recorded separately, by the system they already use, not typed into the agent’s log.

## Why must the comparison be blind?

If the person sees the proposal before deciding, agreement measures persuasion, not correctness. A tidy proposal on the screen is easy to accept.

The EU AI Act names this risk for systems it classes as high-risk. People overseeing them must be enabled “to remain aware of the possible tendency of automatically relying or over-relying on the output produced by a high-risk AI system (automation bias)”.

The fix is procedural. The person decides first, the proposal stays hidden, and the two are compared afterwards by someone with written criteria.

## What should you measure?

Agreement alone hides the cases that matter. Sort every disagreement into one of four bins before you count anything.

BinMeaningWhat it tells youAgent wrong, seriousExecuting the proposal would have caused harm as defined in the charterThe number that decides promotionAgent wrong, minorSuboptimal but harmless, for example a worse time slotTraining material, not a blockerPerson wrongThe agent was right and the person erredA useful side result; it is not a reason to promoteBoth acceptableTwo valid answersNot an error at all

Add two more numbers. Coverage: the share of cases where the agent proposed an action at all. And escalation quality: did it hand over the cases that people also found hard? An agent that escalates everything never errs and never helps.

Also read the reasons, not only the actions. A proposal can match the person for the wrong reason, which is the argument of [evaluate the reasoning, not the fluency](https://lenvanderhof.com/en/blog/evaluate-reasoning-not-fluency/).

## How many cases before you trust it?

More than you think. Statisticians call the shortcut the rule of three. Hanley and Lippman-Hand (JAMA, 1983) put it this way: “if none of n patients shows the event about which we are concerned, we can be 95% confident that the chance of this event is at most three in n (ie, 3/n).”

Clean cases in a rowSerious-error rate you cannot yet rule out50up to 6 in 100300up to 1 in 1003,000up to 1 in 1,000

Work backwards. Decide the serious-error rate you can live with, then divide 3 by it. A tolerable rate of 1 in 200 needs about 600 clean cases in that class. Two quiet weeks with 40 cases is a demo, not evidence.

## Where does CHORUS fit?

**CHORUS** is a coordination protocol for small teams of people and AI agents, from *[The Multi-Agent Organization](https://lenvanderhof.com/books/multi-agent-organization/)*, which is available now. Six skills:

1. **Charter.** A written contract for one worker, human or agent, with seven fields, including an authority ceiling: what it may do without asking.
2. **Handoff.** A context packet that carries a task across a boundary, so the receiver never guesses.
3. **Orchestrate.** Who does what in which order, where parallel work merges, and who a problem escalates to.
4. **Review.** A gate where output passes only against written acceptance criteria.
5. **Update.** What reviews find flows back into prompts, tools, and charters.
6. **Sync.** One shared view of who owns what and what is blocked, read weekly.

You do not need the book to use this. Shadow mode is Review and Update working before any write right exists.

**Review.** The comparison is a review gate. Written criteria decide which disagreements count as serious, and a named reviewer sorts them into the four bins, not the agent and not a gut feeling.

**Update.** Findings go back into the agent’s prompts, tools, and charter. Granting action rights is a change to the charter’s authority ceiling. The book is explicit that moving that boundary “is a charter event, not a prompt tweak”: log the change, name who approved it, tell the team. It also notes that a role with a clean log has earned a lighter touch, and a role you trusted may drift after a model change.

One collision to avoid. The book uses “shadow agent” for an agent working off the map, with no charter. Shadow mode is the opposite: chartered, visible, and without write rights.

## The promotion rule

Grant rights per class of cases, never for the agent as a whole.

StageThe agent mayMove up whenWho signsMove back whenShadowRead and propose; no writesThe clean-case count for the class is met (3 divided by the tolerable rate), in a blind comparisonThe process ownerNot applicableSupervisedAct in one narrow class after a named person approves each actionThe count is met again in real use, with no serious error caught by the approversThe process ownerAny serious error in the classSampledAct in that class; a named person reviews a written sample afterwardsNot applicableThe process ownerAny serious error, or a change of model, instructions, or tools

A demotion sends the class back to shadow, not the whole agent back to the drawing board. Some actions never leave the supervised stage: sending as the company, spending above a ceiling, deleting. The [HITL article](https://lenvanderhof.com/en/blog/hitl-in-ai-workflows/) calls those permanent gates, and the [agent charter](https://lenvanderhof.com/en/blog/what-is-an-ai-agent-charter/) is where you write the ceiling down.

## A worked example: rescheduling deliveries

Imagine a hypothetical regional delivery firm. Customers email to move a delivery, and planners make the change in the planning system, about 40 a day. An agent reads the same emails and proposes each change, with no write access. The planners never see the proposals.

After four weeks, about 800 requests, a reviewer sorts the disagreements. All figures are hypothetical.

Class of requestCasesSerious agent errorsDecisionLater, same week, unchilled goods5200Promote to supervisedEarlier than planned1502Stay in shadowChilled goods904Stay in shadow; add a cold-chain rule to the charterUnclear requests400 (31 escalated)Stay in shadow

The comparison also caught 9 planner typos, a pleasant side effect but not a reason to promote.

Why supervised and not sampled? Zero errors in 520 still allows a rate of about 1 in 170. At roughly 26 requests of that class a day, that could still mean about three serious errors every four weeks. So a planner approves each action while the count keeps running.

## Try this today (20 minutes)

1. Pick one action your agent wants to take, and narrow it to one class of cases.
2. Replace the write tool for that action with a proposal tool that only records.
3. Write the four-bin criteria: what counts as a serious error here?
4. Choose the tolerable serious-error rate and compute the clean-case count.
5. Name the reviewer who sorts disagreements, and write what sends the class back.

Shadow mode answers whether the agent should act. It does not tell you whether you could halt it once it does, or reconstruct later what it did. Settle both before the first promotion, not after.

Cite this:Agent shadow mode: evaluate proposals before granting action rights.Len P. van der Hof. [https://lenvanderhof.com/en/blog/agent-shadow-mode/](https://lenvanderhof.com/en/blog/agent-shadow-mode/) · Published 4 October 2026.

## Terminology

- [CHORUS](https://lenvanderhof.com/glossary/chorus/)

## Sources

1. [Shadow tests (Amazon SageMaker AI Developer Guide)](https://docs.aws.amazon.com/sagemaker/latest/dg/shadow-tests.html) · Amazon Web Services
2. [Mirroring (Istio documentation)](https://istio.io/latest/docs/tasks/traffic-management/mirroring/) · Istio
3. [If Nothing Goes Wrong, Is Everything All Right? Interpreting Zero Numerators (JAMA 1983;249(13):1743-1745)](https://jhanley.biostat.mcgill.ca/c607/ch08/zero_numerator.pdf) · JAMA (author copy hosted by James A. Hanley, McGill University)
4. [Regulation (EU) 2024/1689 (Artificial Intelligence Act), consolidated text](https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:02024R1689-20260727) · EUR-Lex, Publications Office of the European Union
5. [CHORUS (framework)](https://lenvanderhof.com/frameworks/chorus/)
6. [The Multi-Agent Organization](https://lenvanderhof.com/books/multi-agent-organization/)
7. [HITL in AI workflows: put the named person on the irreversible step](https://lenvanderhof.com/en/blog/hitl-in-ai-workflows/)

## Further reading

- [HITL in AI workflows: put the named person on the irreversible step](https://lenvanderhof.com/en/blog/hitl-in-ai-workflows/)
- [What is an AI agent charter? Seven fields or it is a demo](https://lenvanderhof.com/en/blog/what-is-an-ai-agent-charter/)
- [CHORUS](https://lenvanderhof.com/frameworks/chorus/)
- [The Multi-Agent Organization](https://lenvanderhof.com/books/multi-agent-organization/)

About the author

## [Len P. van der Hof](https://lenvanderhof.com/en/authors/len-p-van-der-hof/)

Entrepreneur, AI Innovator and Venture Builder

Len P. van der Hof builds practical AI systems, digital ventures and evidence-informed tools for founders.

```json
{
	"@context": "https://schema.org",
	"@graph": [
		{
			"@type": "Person",
			"@id": "https://lenvanderhof.com/#person",
			"name": "Len P. van der Hof",
			"alternateName": [
				"Len van der Hof",
				"L.P. van der Hof",
				"Leendert Pieter van der Hof"
			],
			"honorificSuffix": "MSc",
			"url": "https://lenvanderhof.com/",
			"image": [
				"https://lenvanderhof.com/photos/len-portrait-1.jpg",
				"https://lenvanderhof.com/photos/len-portrait-2.jpg",
				"https://lenvanderhof.com/photos/len-portrait-3.jpg",
				"https://lenvanderhof.com/photos/len-portrait-4.jpg",
				"https://lenvanderhof.com/photos/len-speaking.jpg",
				"https://lenvanderhof.com/photos/len-hero.jpg"
			],
			"jobTitle": "Entrepreneur, AI Innovator and Venture Builder",
			"description": "Len P. van der Hof, MSc, is a Dutch entrepreneur and AI innovator in Zwijndrecht. He builds ReasonKit, MindSesh, Undominated.ai, books under his name, the fiction imprint LPH98.lifestyle, and technology ventures through LPH98.ventures. Eleven titles in Systems for the Strategic Self are available now, in English and Dutch.",
			"address": {
				"@type": "PostalAddress",
				"addressLocality": "Zwijndrecht",
				"addressCountry": "NL"
			},
			"alumniOf": {
				"@type": "CollegeOrUniversity",
				"name": "Rotterdam School of Management, Erasmus University"
			},
			"knowsAbout": [
				"Artificial intelligence",
				"AI agents",
				"Agentic AI systems",
				"LLM routing",
				"SEO",
				"Generative engine optimization",
				"Answer engine optimization",
				"Venture building",
				"Founder performance",
				"Founder psychology",
				"Evidence-based decision-making"
			],
			"sameAs": [
				"https://www.linkedin.com/in/lenvanderhof/",
				"https://x.com/LenvanderHof",
				"https://www.youtube.com/channel/UCTG20buKqYYbitqqf7l3zJA",
				"https://www.instagram.com/Lenvanderhof/",
				"https://www.threads.com/@lenvanderhof",
				"https://github.com/Lenvanderhof",
				"https://huggingface.co/LPH98",
				"https://www.npmjs.com/~lenvanderhof",
				"https://www.goodreads.com/author/show/70983905.Len_P_van_der_Hof",
				"https://www.amazon.com/author/lenvanderhof",
				"https://www.bol.com/nl/nl/b/len-p-van-der-hof-msc/609879394/",
				"https://bsky.app/profile/lenvanderhof.com",
				"https://mastodon.social/@Lenvanderhof",
				"https://crates.io/users/Lenvanderhof",
				"https://cursor.com/@Lenvanderhof",
				"https://medium.com/@Lenvanderhof",
				"https://gitlab.com/Lenvanderhof",
				"https://hub.docker.com/u/lenvanderhof/",
				"https://dev.to/lenvanderhof",
				"https://www.facebook.com/Lenvanderhof",
				"https://soundcloud.com/Lenvanderhof"
			],
			"affiliation": [
				{
					"@id": "https://lenvanderhof.com/#publisher"
				},
				{
					"@id": "https://lenvanderhof.com/#mindsesh"
				},
				{
					"@id": "https://lenvanderhof.com/#lifestyle"
				}
			]
		},
		{
			"@type": "WebSite",
			"@id": "https://lenvanderhof.com/#website",
			"url": "https://lenvanderhof.com/",
			"name": "Len P. van der Hof",
			"description": "Len P. van der Hof, MSc, is a Dutch entrepreneur and AI innovator in Zwijndrecht. He builds ReasonKit, MindSesh, Undominated.ai, books under his name, the fiction imprint LPH98.lifestyle, and technology ventures through LPH98.ventures. Eleven titles in Systems for the Strategic Self are available now, in English and Dutch.",
			"inLanguage": [
				"en",
				"nl"
			],
			"publisher": {
				"@id": "https://lenvanderhof.com/#person"
			}
		},
		{
			"@type": "Organization",
			"@id": "https://lenvanderhof.com/#publisher",
			"name": "LPH98.ventures",
			"url": "https://lph98.ventures",
			"founder": {
				"@id": "https://lenvanderhof.com/#person"
			}
		},
		{
			"@type": "Organization",
			"@id": "https://lenvanderhof.com/#mindsesh",
			"name": "MindSesh",
			"url": "https://mindsesh.net",
			"founder": {
				"@id": "https://lenvanderhof.com/#person"
			}
		},
		{
			"@type": "SoftwareApplication",
			"@id": "https://lenvanderhof.com/#reasonkit",
			"name": "ReasonKit",
			"url": "https://reasonkit.sh",
			"creator": {
				"@id": "https://lenvanderhof.com/#person"
			}
		},
		{
			"@type": "SoftwareApplication",
			"@id": "https://lenvanderhof.com/#undominated",
			"name": "Undominated.ai",
			"url": "https://undominated.ai",
			"creator": {
				"@id": "https://lenvanderhof.com/#person"
			}
		},
		{
			"@type": "Organization",
			"@id": "https://lenvanderhof.com/#lifestyle",
			"name": "LPH98.lifestyle",
			"url": "https://lph98.lifestyle",
			"founder": {
				"@id": "https://lenvanderhof.com/#person"
			}
		},
		{
			"@type": "ImageObject",
			"@id": "https://lenvanderhof.com/en/blog/agent-shadow-mode/#primaryimage",
			"url": "https://lenvanderhof.com/media/generated/blog-hero-agent-shadow-mode-v1.8ea88ca10401.wide.webp",
			"contentUrl": "https://lenvanderhof.com/media/generated/blog-hero-agent-shadow-mode-v1.8ea88ca10401.wide.webp",
			"representativeOfPage": true
		},
		{
			"@type": "BreadcrumbList",
			"@id": "https://lenvanderhof.com/en/blog/agent-shadow-mode/#breadcrumb",
			"itemListElement": [
				{
					"@type": "ListItem",
					"position": 1,
					"name": "Home",
					"item": "https://lenvanderhof.com/"
				},
				{
					"@type": "ListItem",
					"position": 2,
					"name": "Blog",
					"item": "https://lenvanderhof.com/en/blog/"
				},
				{
					"@type": "ListItem",
					"position": 3,
					"name": "AI Systems",
					"item": "https://lenvanderhof.com/en/blog/category/ai-systems/"
				},
				{
					"@type": "ListItem",
					"position": 4,
					"name": "Agent shadow mode: evaluate proposals before granting action rights",
					"item": "https://lenvanderhof.com/en/blog/agent-shadow-mode/"
				}
			]
		},
		{
			"@type": "WebPage",
			"@id": "https://lenvanderhof.com/en/blog/agent-shadow-mode/#webpage",
			"url": "https://lenvanderhof.com/en/blog/agent-shadow-mode/",
			"name": "Agent shadow mode: evaluate proposals before granting action rights",
			"description": "Agent shadow mode: the agent proposes, a person still acts, and you compare. What to measure, how many cases you need, and a promotion rule you can copy.",
			"isPartOf": {
				"@id": "https://lenvanderhof.com/#website"
			},
			"primaryImageOfPage": {
				"@id": "https://lenvanderhof.com/en/blog/agent-shadow-mode/#primaryimage"
			},
			"breadcrumb": {
				"@id": "https://lenvanderhof.com/en/blog/agent-shadow-mode/#breadcrumb"
			},
			"inLanguage": "en-GB"
		},
		{
			"@type": "BlogPosting",
			"@id": "https://lenvanderhof.com/en/blog/agent-shadow-mode/#article",
			"mainEntityOfPage": {
				"@id": "https://lenvanderhof.com/en/blog/agent-shadow-mode/#webpage"
			},
			"headline": "Agent shadow mode: evaluate proposals before granting action rights",
			"description": "Agent shadow mode: the agent proposes, a person still acts, and you compare. What to measure, how many cases you need, and a promotion rule you can copy.",
			"datePublished": "2026-10-04T07:00:00.000Z",
			"author": {
				"@id": "https://lenvanderhof.com/#person"
			},
			"publisher": {
				"@id": "https://lenvanderhof.com/#person"
			},
			"image": [
				"https://lenvanderhof.com/media/generated/blog-hero-agent-shadow-mode-v1.8ea88ca10401.square.webp",
				"https://lenvanderhof.com/media/generated/blog-hero-agent-shadow-mode-v1.8ea88ca10401.landscape.webp",
				"https://lenvanderhof.com/media/generated/blog-hero-agent-shadow-mode-v1.8ea88ca10401.wide.webp"
			],
			"articleSection": "AI Systems",
			"inLanguage": "en-GB"
		}
	]
}
```
