AI Systems Research guide

Agent audit trail: reconstruct the action, approval, and outcome

If the only witness is the agent, you do not have an audit trail. You have a testimony.

Three stamped paper slips clipped together on a steel archive table under a single work lamp, a cardboard parcel beside them with a matching torn label
The action, the approval, the outcome. Clipped together, or it did not happen.

Direct answer

An agent audit trail is the set of records that lets someone who was not there reconstruct one action by an AI agent: what it did, with which inputs and instructions, who or what approved it, and what actually happened in the system it touched. The records are written by the software around the agent, not by the agent itself, share one ID, and sit where the agent cannot edit them. A chat transcript is not enough: it shows what the model said, not what the tools did.

An agent audit trail is the set of records that lets someone who was not there reconstruct one action by an AI agent: what it did, who or what approved it, and what actually happened in the system it touched. The records are written by the software around the agent, not by the agent. They share one ID. And they sit where the agent cannot edit them.

The test is simple. Six weeks from now, a customer disputes something the agent did. Can a colleague answer from records alone, without asking the agent and without asking you? If not, you have a memory, not a trail.

An AI agent here means software that takes a goal and uses tools, such as email, a database, or a payment system, over several steps. An audit trail is a set of records, in time order, that shows what happened, who did it, and on whose authority.

Why is a chat transcript not an audit trail?

The transcript is the conversation you see in the chat window. It feels like a record. It records the wrong thing.

It shows what the model said, not what the tools did. An agent can write “I’ve refunded the order” whether or not the refund went through. The words and the action are separate events.

Asking the agent afterwards does not fix that. A 2023 OpenAI white paper on governing agents puts it plainly: “It is unfortunately not possible to simply ‘ask’ the agent to retroactively justify its behavior, as this is likely to produce confabulated reasoning.” Confabulated means a plausible story made up after the fact. The same paper recommends giving users “a ledger of actions taken by the agent.”

It leaves out the approval. Who clicked approve, when, and what exactly was on their screen? The chat rarely knows.

It leaves out the outcome. The payment provider’s own record of the refund, with its own ID and status, lives in the payment system, not in the chat.

It is easy to lose. Chats can be edited, cut short, or deleted when a vendor’s retention period ends. And they carry no shared ID you can use to join them to the other systems.

What is the minimum record?

Three parts, joined by one correlation ID (one identifier stamped on every record that belongs to the same action). Here is the minimum, with a hypothetical refund as the example.

PartFieldWhat it answersExample (hypothetical)
ActionCorrelation IDWhich records belong togetheract-0917-0142
ActionTime (UTC)When the tool was called17 Sep, 14:03:22
ActionAgent and versionsWhich agent, model, instruction file, and tool versionrefund-agent, instructions at commit 4f2a
ActionTrigger and inputsWhat started it and what it readTicket 4471; order lookup returned 88213
ActionTool call as sentThe exact tool and arguments, or a safe summary if they hold secretsrefund(order 88213, €64.90)
ApprovalRule appliedWhich rule allowed it, or required a personRefunds over €50 need approval
ApprovalApprover and timeThe named person and the momentSupport lead, 14:05:10
ApprovalWhat the approver sawThe exact payload version on their screenOrder 88213, €64.90, customer name
OutcomeTarget system resultThe other system’s own ID and statusProvider refund re_71x, succeeded
OutcomeEffect checkWhat changed, checked in the system itselfOrder 88213 marked refunded
OutcomeFollow-upsRetries, reversals, complaints, linked by the same IDCustomer email, 29 Oct

Eleven fields. You do not need a platform. For a small team, a table with these columns, filled in automatically, is a working trail.

The HITL article on this site names the log as one of four objects a human gate needs. This table is what that log has to contain to survive a dispute.

Who writes the record: STACK

STACK is a five-layer map for code repositories where people and AI agents share the work. It comes from The Agentic Codebase, which is available now. The five layers:

  1. Structure. The layout, entry points, and boundaries an agent can find its way through in its first minute in the repository.
  2. Toolchain. The shells, commands, and command-line contracts an agent is allowed to run, written down instead of remembered.
  3. Agent configuration. The standing instructions (AGENTS.md, CLAUDE.md, rules, skills, role charters), kept under version control like code.
  4. Connection. The tool servers, tool contracts, hooks, and guardrails through which the agent touches the outside world, with the least access the job needs.
  5. Knowledge and quality. Memory, context limits, tests of the agent’s work (evals), and automated checks (CI), so a model upgrade does not quietly lower the bar.

You do not need the book to use this. The audit trail lives in the last two layers.

Connection: write the record where the call passes. Every tool call passes through a point you control: a tool server, a gateway, or a hook (a small program that runs automatically right before or right after a tool call). The book’s point about hooks is that they hold no matter what the model decides. That makes them the right place to write the action record and the approval, with the same ID. The book also lists “no secret in a log” among the rules a hook should enforce, so strip keys and passwords before anything is stored. The payment provider’s response should be recorded by the tool, not summarized by the model. If your tools connect through MCP (Model Context Protocol, a standard way for AI apps to reach tools), the MCP server is a natural place for this.

Knowledge and quality: record the versions, and read the trace. A record that says “the refund agent did it” is useless if the agent’s instructions changed twice that week. Store the model name, the commit of the instruction files, and the tool contract version on every action.

The book also recommends tracing runs with OpenTelemetry, an open standard for recording what software does as a tree of timed steps called spans. Its GenAI conventions define an execute_tool span carrying the tool’s name and a call ID. Arguments and results are opt-in and flagged as possibly sensitive, and the whole set is still marked Development, so pin your version. The attribute registry, as read for this page, has no field for a human approval. Add that yourself.

A worked example: “You refunded the wrong order”

Imagine a hypothetical online shop. Its refund agent reads support tickets, looks up the order, and calls the payment tool. Refunds over €50 need a support lead’s approval.

Six weeks later a customer writes: she returned order 88231, but the refund went to 88213, an order she kept.

With a transcript only, you find the agent’s message “I’ve refunded your order” and nothing else. The support lead remembers approving “something like that.”

With the trail, a colleague who was not involved searches by order number and needs about ten minutes:

  1. Action. At 14:03 the agent called the refund tool for order 88213. Its input shows the ticket text said 88231, and the lookup tool returned 88213: a fuzzy match, meaning a near miss accepted as a hit, with two digits swapped.
  2. Approval. The support lead approved at 14:05. The record of her screen shows the order number from the call, the amount, and the customer’s name. It did not show the number the customer typed.
  3. Outcome. The payment provider’s record confirms refund re_71x on 88213. Order 88231 was never refunded.

Two fixes follow, and neither is “be more careful.” The approval screen must show the customer’s own order number next to the one in the call. And the lookup tool must refuse fuzzy matches on order numbers. Both are changes to the Connection layer, and the trail pointed straight at them.

What does the law ask for?

This is not legal advice. For systems the EU AI Act classes as high-risk, Article 12(1) says: “High-risk AI systems shall technically allow for the automatic recording of events (logs) over the lifetime of the system.” Article 26(6) asks deployers (the organizations that use such a system) to keep the logs under their control for at least six months, unless other law says otherwise.

Many internal agents will not be high-risk under the Act. The disputes, refunds, and complaints arrive anyway.

Try this today: the reconstruction drill (20 minutes)

Pick one action by one agent from last month at random, not one you remember. Give a colleague who was not involved read access to the records, but no chance to ask the people involved. Ask them to answer seven questions within 20 minutes:

  1. What exactly did the agent do: which tool, which arguments?
  2. What inputs and which instruction versions was it working from?
  3. Which rule allowed the action, or who approved it?
  4. What did the approver see?
  5. What did the target system report back?
  6. What changed as a result, and is it still that way?
  7. Did anything follow: a retry, a reversal, a complaint?

Every “I would have to ask someone” is a missing field. Every answer that comes from the agent’s own words is a claim, not a record. Fix the first gap before the next action runs.

An audit trail tells you what happened. It does not halt anything, and it does not decide whether the agent should have been allowed to act in the first place. Those are separate tests, and each needs its own proof. For claims instead of actions, the same idea is an evidence ledger.

Cite this:Agent audit trail: reconstruct the action, approval, and outcome.Len P. van der Hof. https://lenvanderhof.com/en/blog/agent-audit-trail/ ·

Terminology

Sources

  1. Practices for Governing Agentic AI Systems (Shavit et al.) · OpenAI
  2. Semantic conventions for generative client AI spans · OpenTelemetry
  3. Gen AI attribute registry · OpenTelemetry
  4. Regulation (EU) 2024/1689 (Artificial Intelligence Act), consolidated text · EUR-Lex, Publications Office of the European Union
  5. STACK (framework)
  6. The Agentic Codebase
  7. HITL in AI workflows: put the named person on the irreversible step

Further reading

Markdown for LLMs