To retire a production agent, close every way it can start work, sign in or cost money, then prove each one is closed by trying it. A production agent is an AI agent (software that takes a goal, chooses its own steps and uses tools) that does real work on live systems: sending email, changing records, calling paid services. Pressing stop on its current run is not retirement. The next scheduled run, a saved key or a monthly plan can keep it working and billing.
The order: freeze new work, settle accepted work, redirect callers, revoke access, close spending, keep the record, then test every old path. The one-screen checklist is further down.
Why is “we switched it off” not enough?
Imagine a hypothetical invoice-reminder agent at a ten-person company. The accounting software now sends reminders itself, so the agent is no longer needed. Someone stops the running job and posts “reminder agent is off” in the team chat. Three things keep going.
- The schedule. On Monday at 09:00 the scheduler (the service that starts jobs on a timetable) launches a fresh run, and a client gets a duplicate reminder.
- A copied key. The agent’s API key (a password for software) was also pasted into a finance spreadsheet, which keeps calling the model under the agent’s name.
- The plan. The provider keeps billing a monthly plan bought for the agent, because stopping usage does not cancel a subscription.
None of this is exotic. Each is a path the stop button never touched.
Separate the three jobs
| Job | What it establishes | What it leaves open |
|---|---|---|
| Stop a run | This execution has ended | Another trigger can start another run |
| Name an owner | A person is accountable for the agent | Keys, schedules and charges can stay live |
| Retire the agent | Work, access and spend are closed and tested | Records and the final bill still need an owner |
Give each job its own status in one retirement ticket. An owner signing off on a stopped run does not revoke its key. This page covers the planned, permanent shutdown; halting a misbehaving agent in the middle of a run is a different drill.
Start from the agent’s charter, the written contract that says what it does and may not do. Its inputs and authority ceiling show what the agent was meant to reach. The inventory below shows what it actually reaches.
Step 1: find everything that can still start, sign in or bill
Build the list from the systems themselves. A registry can miss a copied key or a caller owned by another team. Check four groups.
- Triggers (whatever starts a run): platform schedules, cron entries (timed jobs on a server), queue consumers (processes that pick up waiting tasks), webhooks (a web address another system calls to start work) and inbox rules.
- Credentials (whatever lets it sign in or act): API keys, OAuth grants (permission a person gave an app to act on their account) and their refresh tokens, service accounts (accounts that belong to software, not people), cloud roles, and keys stored in the settings files that connect AI apps to tools.
- Callers (whatever sends it work): other agents, notebooks, spreadsheets, customer workflows and retry jobs.
- Costs: model-provider accounts, subscriptions, reserved capacity (computing paid for in advance), search indexes and other services bought for the agent.
For each item, write down the owner, the identifier, the intended action and the evidence that will close it. Trace each key through the secret store, the deployment settings, notebooks and local files. Record identifiers, never the secret values.
Mark shared keys. Deleting a key from one file leaves every other copy working. Revoking it stops every copy at once, including the ones live systems still depend on; the TrueFoundry playbook puts it plainly: the copies are the same credential. So move those systems to replacement access, confirm they work, then revoke. Keep the replacement out of the retired agent’s settings.
Step 2: withdraw the agent in order
TrueFoundry’s playbook, by Boyu Wang (8 August 2026), names six steps: inventory, redirect, revoke, retain, tombstone and verify. The order below follows it, splits “redirect” into three plain steps, and adds spending as a step of its own.
- Freeze intake. Disable the triggers that accept new work and record the cut-off time. Keep only the access needed to finish work already accepted.
- Resolve accepted work. List running jobs, queued jobs and pending retries. Complete, reassign or cancel each one, and record who owns any handoff. Read the queue itself, not only the dashboard.
- Redirect callers. Point each caller at a tested replacement, or have it receive a clear “retired” response. Remove retries aimed at the old agent. Ask each caller’s owner to confirm the new behaviour.
- Revoke access. Revoke the agent’s own keys and grants, including refresh tokens that can issue new access. Then check what happens to access tokens already issued. The OAuth revocation standard, RFC 7009, says a server that supports it SHOULD also invalidate access tokens from a revoked refresh token’s grant. “Should” and “supports” are the catch: assume an issued token may work until it expires unless your provider says otherwise. Remove role assignments and record the identifier, action and time for each credential.
- Close spend. Disable further usage and cancel the agent’s own subscriptions, seats and reservations. Check what each budget setting actually does, because many only warn. Google Cloud’s own documentation says an alerts-only budget “doesn’t automatically cap” usage or spending. For shared services, remove the agent’s share without closing something live systems still need.
- Retain the record. Keep the ticket and the logs, traces and prompts your policy requires or permits. Write down where the archive lives, who may open it, the retention rule and the deletion date if the policy sets one. Then mark the registry entry retired, with owner, date and reason. That entry, the tombstone, helps prevent accidental reuse. It does not disable a single key.
If a vendor runs part of the agent, add the vendor’s retirement action and confirmation to the ticket. Closing your side can leave their hosted account, schedule or subscription running.
Step 3: verify by trying
Call the old webhook and the old start path without re-enabling either. Present the old access token. Try the old refresh token. Call the old endpoint from a client that used to succeed. Use harmless requests you can trace, in case a control fails.
Check the effect as well as the response. Nothing should start, no old credential should authorise work, and the refresh token should not yield a usable token. If a caller now reaches a replacement, confirm which system handled the request.
Then read the logs from the cut-off onwards, as TrueFoundry’s verify step does. Successful traffic attributed to the retired agent should be zero, and attempts with revoked credentials should appear only as authentication failures. Any success means revocation is incomplete or another credential still works.
Check usage and billing separately. A final invoice can include work done before the cut-off. Match each remaining charge to its service period and account for any reservation or subscription still running. For every attempt, record the time, the expected result, the observed result and the supporting log line.
The retirement checklist
| Path | Close it by | Proof it is closed |
|---|---|---|
| Schedules and timed jobs | Disable, then delete | The next scheduled time passes with no run in the log |
| Webhooks and queues | Remove the endpoint or the consumer | A harmless test call starts nothing |
| Callers | Point them at the replacement or a clear “retired” response | Each caller’s owner confirms the new behaviour |
| Keys and grants | Move shared users first, then revoke | The old key and refresh token are refused |
| Accounts and roles | Disable the service account, remove its roles | Sign-in fails and no role assignment remains |
| Spend | Cancel subscriptions, seats and reservations | Every final invoice line matches a service period |
| Record | Archive per policy, mark the registry entry retired | Archive location and deletion date written down |
Close the ticket on evidence
Use one ticket: the agent’s registry ID, who authorised retirement, the intake cut-off, and a link to the evidence for each row above. Give every open item an owner and a next check; a cancellation that runs to the end of a billing period needs a follow-up date, and a submitted request closes nothing on its own. Close the ticket only when every old path has been tried and has failed, and every remaining charge is accounted for.
Where does retirement fit in ROSTER?
ROSTER is the framework from The Agent Dream Team, available now, for running AI agents as named specialists. It makes six decisions for every agent:
- Roles: one job per agent, with a charter that says what it owns and what it does not.
- Objectives: a result you can grade.
- Skills and tools: only the access the job needs.
- Triggers: the event that starts the agent.
- Evaluation: a written scorecard, graded weekly.
- Rotation: upgrade, merge or retire the role when the evidence says so.
You do not need the book to use it. Rotation decides whether a role is retired; this page covers how. The book’s rule is that one week below the scorecard floor starts a diagnosis, while two weeks in a row, or four of the last six, start a decision to upgrade, merge or retire. It adds a check that saves real work: if the provider changed or retired the model underneath, treat it as a forced migration (move the role to a new model and re-test it), not a retirement. When a role does retire, the book files its charter in an archive with a one-paragraph note on why, and removes it from the trigger map.
Apply that to the invoice-reminder agent. The Rotation decision reads “retire: the accounting software now does this job”. The Triggers entry names the Monday schedule to delete. The Skills and tools entry lists the key to revoke. The charter’s output line tells you who received its work. A good charter is your first inventory, which is why the ROSTER framework page puts it first.
Set the next retirement date when you create the agent
In his Forbes Technology Council essay, Akhilesh Sharma argues that every application, agent, automation and integration should get a time-to-live (an expiry date) when it is created, a named owner with the authority to renew, restrict or retire it, and a staged retirement.
Make those decisions when you write the next agent’s charter. Name the renewal owner, set the review date, and list the credentials, callers and billing accounts that retirement will touch. The exit rule in agent strategy is the same idea at portfolio level. And if an agent runs on access created by an employee, add it to that person’s offboarding checklist: closing their login can leave a service account active.
Try this today: a 15-minute retirement dry run
Pick one agent that is still running, even a good one.
- Write down its triggers, credentials (identifiers only), callers and costs, one line each.
- Next to each line, write who owns it and how you would close it.
- Any line without an owner is your first fix. Any line you cannot fill in is a path you would miss on retirement day.
Keep the page. On the day you retire the agent, it becomes the ticket.
Cite this:How to retire a production agent (stopping it is not enough).Len P. van der Hof. https://lenvanderhof.com/en/blog/retire-a-production-agent/ ·