Agent shadow mode means an AI agent receives real work and writes down what it would do, but has no permission to do it. A person or the existing process still acts, and afterwards you compare the two. You grant the agent action rights only when the comparison clears a rule you wrote in advance.
An agent that made no serious mistake in 50 cases has shown you very little. This page covers why, what to measure instead, and a promotion rule you can copy.
An AI agent here means software that takes a goal and uses tools, such as email, a database, or a planning system, over several steps without a person approving each one. Action rights are its permissions to change things in those systems.
Where does the term come from?
From software deployment. Amazon’s machine-learning service, SageMaker, offers shadow tests: it “automatically deploys the new variant in shadow mode and routes a copy of the inference requests to it in real time within the same endpoint. Only the responses of the production variant are returned to the calling application.” In plain words: the new version receives a copy of the real questions sent to the live model, and users only ever get answers from the version already in production.
Istio, a service mesh (software that routes traffic between the parts of a cloud application), describes the same move: “Traffic mirroring, also called shadowing, is a powerful concept that allows feature teams to bring changes to production with as little risk as possible.” The copied requests are fire and forget. Their responses are thrown away.
That works because a prediction is only an answer. You can discard an answer. An agent’s output is an action: an email sent, a booking moved, a refund paid. A mirrored agent with a live send tool does not shadow. It sends.
So agent shadow mode is stricter than its software cousin. The agent gets the same inputs and the same read access. Every tool that writes is replaced by a proposal tool that records the intended action and does nothing else.
What goes into a proposal?
Each proposal is a short record with the same case ID the person’s real decision will carry:
- The action. The exact tool and arguments it would have used: “move delivery 5530 to Thursday 14:00 to 16:00.”
- The reason. The message, record, or rule it relied on.
- Escalate or not. Whether it would have handed the case to a person instead.
- The time. When it proposed, so you can compare speed as well.
The person’s decision is recorded separately, by the system they already use, not typed into the agent’s log.
Why must the comparison be blind?
If the person sees the proposal before deciding, agreement measures persuasion, not correctness. A tidy proposal on the screen is easy to accept.
The EU AI Act names this risk for systems it classes as high-risk. People overseeing them must be enabled “to remain aware of the possible tendency of automatically relying or over-relying on the output produced by a high-risk AI system (automation bias)”.
The fix is procedural. The person decides first, the proposal stays hidden, and the two are compared afterwards by someone with written criteria.
What should you measure?
Agreement alone hides the cases that matter. Sort every disagreement into one of four bins before you count anything.
| Bin | Meaning | What it tells you |
|---|---|---|
| Agent wrong, serious | Executing the proposal would have caused harm as defined in the charter | The number that decides promotion |
| Agent wrong, minor | Suboptimal but harmless, for example a worse time slot | Training material, not a blocker |
| Person wrong | The agent was right and the person erred | A useful side result; it is not a reason to promote |
| Both acceptable | Two valid answers | Not an error at all |
Add two more numbers. Coverage: the share of cases where the agent proposed an action at all. And escalation quality: did it hand over the cases that people also found hard? An agent that escalates everything never errs and never helps.
Also read the reasons, not only the actions. A proposal can match the person for the wrong reason, which is the argument of evaluate the reasoning, not the fluency.
How many cases before you trust it?
More than you think. Statisticians call the shortcut the rule of three. Hanley and Lippman-Hand (JAMA, 1983) put it this way: “if none of n patients shows the event about which we are concerned, we can be 95% confident that the chance of this event is at most three in n (ie, 3/n).”
| Clean cases in a row | Serious-error rate you cannot yet rule out |
|---|---|
| 50 | up to 6 in 100 |
| 300 | up to 1 in 100 |
| 3,000 | up to 1 in 1,000 |
Work backwards. Decide the serious-error rate you can live with, then divide 3 by it. A tolerable rate of 1 in 200 needs about 600 clean cases in that class. Two quiet weeks with 40 cases is a demo, not evidence.
Where does CHORUS fit?
CHORUS is a coordination protocol for small teams of people and AI agents, from The Multi-Agent Organization, which is available now. Six skills:
- Charter. A written contract for one worker, human or agent, with seven fields, including an authority ceiling: what it may do without asking.
- Handoff. A context packet that carries a task across a boundary, so the receiver never guesses.
- Orchestrate. Who does what in which order, where parallel work merges, and who a problem escalates to.
- Review. A gate where output passes only against written acceptance criteria.
- Update. What reviews find flows back into prompts, tools, and charters.
- Sync. One shared view of who owns what and what is blocked, read weekly.
You do not need the book to use this. Shadow mode is Review and Update working before any write right exists.
Review. The comparison is a review gate. Written criteria decide which disagreements count as serious, and a named reviewer sorts them into the four bins, not the agent and not a gut feeling.
Update. Findings go back into the agent’s prompts, tools, and charter. Granting action rights is a change to the charter’s authority ceiling. The book is explicit that moving that boundary “is a charter event, not a prompt tweak”: log the change, name who approved it, tell the team. It also notes that a role with a clean log has earned a lighter touch, and a role you trusted may drift after a model change.
One collision to avoid. The book uses “shadow agent” for an agent working off the map, with no charter. Shadow mode is the opposite: chartered, visible, and without write rights.
The promotion rule
Grant rights per class of cases, never for the agent as a whole.
| Stage | The agent may | Move up when | Who signs | Move back when |
|---|---|---|---|---|
| Shadow | Read and propose; no writes | The clean-case count for the class is met (3 divided by the tolerable rate), in a blind comparison | The process owner | Not applicable |
| Supervised | Act in one narrow class after a named person approves each action | The count is met again in real use, with no serious error caught by the approvers | The process owner | Any serious error in the class |
| Sampled | Act in that class; a named person reviews a written sample afterwards | Not applicable | The process owner | Any serious error, or a change of model, instructions, or tools |
A demotion sends the class back to shadow, not the whole agent back to the drawing board. Some actions never leave the supervised stage: sending as the company, spending above a ceiling, deleting. The HITL article calls those permanent gates, and the agent charter is where you write the ceiling down.
A worked example: rescheduling deliveries
Imagine a hypothetical regional delivery firm. Customers email to move a delivery, and planners make the change in the planning system, about 40 a day. An agent reads the same emails and proposes each change, with no write access. The planners never see the proposals.
After four weeks, about 800 requests, a reviewer sorts the disagreements. All figures are hypothetical.
| Class of request | Cases | Serious agent errors | Decision |
|---|---|---|---|
| Later, same week, unchilled goods | 520 | 0 | Promote to supervised |
| Earlier than planned | 150 | 2 | Stay in shadow |
| Chilled goods | 90 | 4 | Stay in shadow; add a cold-chain rule to the charter |
| Unclear requests | 40 | 0 (31 escalated) | Stay in shadow |
The comparison also caught 9 planner typos, a pleasant side effect but not a reason to promote.
Why supervised and not sampled? Zero errors in 520 still allows a rate of about 1 in 170. At roughly 26 requests of that class a day, that could still mean about three serious errors every four weeks. So a planner approves each action while the count keeps running.
Try this today (20 minutes)
- Pick one action your agent wants to take, and narrow it to one class of cases.
- Replace the write tool for that action with a proposal tool that only records.
- Write the four-bin criteria: what counts as a serious error here?
- Choose the tolerable serious-error rate and compute the clean-case count.
- Name the reviewer who sorts disagreements, and write what sends the class back.
Shadow mode answers whether the agent should act. It does not tell you whether you could halt it once it does, or reconstruct later what it did. Settle both before the first promotion, not after.
Cite this:Agent shadow mode: evaluate proposals before granting action rights.Len P. van der Hof. https://lenvanderhof.com/en/blog/agent-shadow-mode/ ·