AI Agent Forensics: Prove What Your Agent Actually Did
RedHub AI Editorialupdated October 2, 20266 min read

Jump to a section9
After something goes wrong, AI agent forensics works out exactly what the agent did: which agent acted, on whose authority, with what instructions and context, through which tools, and what changed as a result. It moves a team from suspicion to evidence. It cannot be bolted on afterward: evidence not captured while the agent runs is gone by the time anyone asks.
TL;DR: No security program prevents every failure, so agents need reconstruction as well as prevention. Give every run one trace ID, attach every model call, tool call, policy decision and approval to it, and keep before-and-after values. Then an investigator can walk backward from a bad outcome to its cause. The AI Output Audit-Trail & Record-Keeping Kit ($79) audits whether each entry in your AI usage record would hold up. Start with the pillar: Zero Trust for AI Agents: Why One Check at the Door Fails.
Prevention is not enough
An agent can be hijacked by content it reads, misuse a tool it is allowed to use, act on stale data, or hold more authority than it should. Good controls make those rarer, never impossible. So a mature setup needs two capabilities: stopping bad actions, and reconstructing the ones that got through.
An investigation has to answer which agent acted and who delegated its authority, what goal it had, which model, prompt, memory and context it used, which tools and credentials were active, what it contacted, what changed, and whether a person approved anything. Without those answers, an incident becomes a debate about what the agent probably did.
A pricing agent, walked backward
An online store runs an agent that updates product prices overnight from supplier cost feeds. On a Wednesday morning, 312 products are listed at about a tenth of their usual price, and orders are coming in (illustrative numbers). The owner asks what it did, and why. If every event in the run shares one trace ID, the answer takes minutes, not days. Start with one wrong price and follow the chain back.
| Walking back from | What the trace shows |
|---|---|
| The outcome | One product's price changed from $41.00 to $4.10 at 2:14 a.m. |
| The tool call | The price-update tool, called with that product and $4.10 |
| The model's input | A supplier cost of $2.46, read from that night's feed |
| The retrieved source | The supplier's file, where the cost column had shifted one decimal place |
| The policy decision | The change rule allowed price moves of up to 95%, so it passed |
| The approval | None required for the nightly batch |
The diagnosis falls out of the trace. The agent's usual markup turns a $24.60 cost into a $41.00 price. The shifted file read $2.46, and the same markup gave $4.10, a 90% cut that a 95% limit let through. The fix is a tighter limit on price moves, with a hold for review above it. Without the trace, the team would still be arguing about whether the model had gone rogue.
What a defensible trail contains
That walk works only if the trail holds the right things. A trail you can defend contains:
- The agent's identity, and the person, system or schedule that started the work
- The business goal, the scope and the risk classification
- The model and its version, the prompt version and the policy version
- The documents, search results and memory the task used
- Every tool call, with its parameters, its response and the change it caused
- The permissions used, the credentials delegated and the outside destinations
- Human approvals, overrides, blocked actions, retries and errors
- Before-and-after values for important records, plus a rollback reference
Most of that is provenance: where each input came from and which version was in play. Our posts on what an AI audit trail needs and on why logs are not an audit trail cover the trail as a governance record. This post uses it for one job: rebuilding an incident.
One trace ID per run
Reconstruction needs an event stream, not scattered application logs. Each run gets a unique trace ID, and every model call, retrieval, tool call, policy decision, approval and result is written with that ID attached.
Scattered logs fail in a predictable way. The tool's log shows the price change, the model provider's log shows a request, and the feed system shows a file, but nothing joins them. Someone matches timestamps by hand. With one ID, the join is already made when the event is written, and an investigator can start from any outcome, an email sent, a record changed, a file created, and walk back through the whole decision path.
An investigation, step by step
- Contain the agent and revoke its credentials.
- Preserve task state, logs, queued jobs, tool records and affected files.
- Rebuild the timeline from the trace ID.
- Find the starting instruction, context, policy decision and tool sequence.
- Decide whether the failure involved identity, permissions, prompt injection, tool misuse, data quality or model behavior.
- Assess the affected data, systems, customers and downstream automations.
- Close the control gap, document the incident and test the fix.
Order matters at the start: preserve before you repair. Rolling the prices back first would have overwritten the values the investigation needed. The AI agent kill switch post makes evidence preservation one of the parts of a working stop.
The audit-trail test
Each "no" below names what you could not prove after an incident.
| Question | If the answer is no |
|---|---|
| Can you identify which agent acted? | You do not have accountable identity. |
| Can you see the exact tool parameters? | You cannot verify what action was actually attempted. |
| Can you link actions to an approval? | You cannot prove human oversight. |
| Can you reconstruct retrieved context? | You cannot understand why the agent reached its conclusion. |
| Can you view before-and-after state? | You cannot reliably investigate or roll back damage. |
| Can you preserve evidence after an incident? | You cannot support forensics or compliance review. |
The evidence is sensitive too
A complete trail creates its own risk. To reconstruct the pricing run, the trace has to hold the supplier's cost data. A support agent's trace holds customer messages. Keep everything forever, and the trail becomes a second, ungoverned copy of your most sensitive data, with weaker controls than the original.
So design the trail with retention rules, access controls, redaction, encryption and role-based access. That leaves a real tension: every field you redact is one an investigator may later need, and no setting satisfies both. Record enough to explain critical behavior, and keep sensitive payloads only where the decision depended on them. Retention can depend on rules in your industry, so ask counsel what applies to you.
Find out whether your AI record would hold up
The kit audits your record of AI usage, not the output. Each entry is scored on eight weighted provenance fields, such as who generated it, for what decision and who reviewed it, and returns LOGGED, PARTIAL or NO TRAIL. The whole record rolls up to TRACEABLE, GAPS or UNRECORDED. A missing required field, or a high-impact decision with no reviewer and no source, forces NO TRAIL and makes the whole record UNRECORDED. The shipped sample is 98% complete and still reads UNRECORDED. It checks the record. It does not capture traces.
Get the AI Output Audit-Trail & Record-Keeping Kit — $79Pairs well with
The AI Incident Postmortem & Readiness Gate ($79) grades a finished postmortem and gates the close, returning CLEARED TO CLOSE AS DESCRIBED, FINISH ACTIONS or NOT CLOSEABLE. The Audit Evidence Pack Assay ($119) grades whether an automated decision can be reconstructed by someone who was not there, returning ANSWERABLE, NEEDS A WITNESS or NO RECORD. The Deterministic Replay Warden ($109) grades whether a decision already made could be produced again, returning REPLAYABLE, RECONSTRUCTABLE or UNREPLAYABLE.
More in this guide
What is AI agent forensics?
It is the ability to work out exactly what an agent did after something went wrong: who acted, on whose authority, with what context, through which tools, and what changed.
How is it different from logging?
Logs record events, often in separate systems. Forensics needs them joined into one chain per run, usually by a shared trace ID.
What is a trace ID?
A unique identifier for one agent run. Every event in that run is recorded with it, which links them into a single timeline.
What should I do first after an agent incident?
Contain the agent, revoke its credentials and preserve its state, logs, queued jobs and affected records before repairing anything. A rollback done first can overwrite the evidence you need.
Can an audit trail be a privacy risk?
Yes. A detailed trail can copy sensitive data. Use retention rules, access controls, redaction and encryption, and ask counsel whether record-keeping rules apply in your industry.
Is agent forensics only useful after a disaster?
No. The same evidence shows which tools cause retries, which sources carry injection attempts, and how model versions compare.


The gate this post refers to, drawn from the tool’s own logic. See the tool.