WentRogueWentRogue

Research notes4 min read

Is an agent's explanation an audit trail?

An agent-generated explanation can help a reader understand an answer, but it is not a complete or independently verified record of what caused the action. Audit evidence needs external events, tool results, state changes and known limits.

An explanation is one kind of record

Ask an AI agent why it did something and it will usually answer fluently. That answer can be genuinely helpful: it can summarise the goal, name the inputs it relied on and make a result easier to follow. It is still text the model generated after, or alongside, the work. It is not automatically a record of what caused the action.

The same applies to a displayed chain-of-thought or rationale. Showing intermediate reasoning can make a process feel transparent, but the display is also generated output. An audit trail has a different job. It should let someone check what happened from evidence that does not depend solely on the agent's own account.

Five records have different jobs

A model-generated explanation tells a reader how the agent describes its choice. A displayed chain-of-thought shows reasoning-like text the system chose to present. Tool and event logs record calls, inputs, results and times as captured by the surrounding software. Outcome records state what finished: an order completed, a file changed, a message sent. Verified external state is what an independent check confirms, such as a settled payment or a delivered item.

These layers can agree, and often should. When they disagree, the explanation is the weakest witness for causation, because it is the only one produced entirely by the system being examined. Audit evidence keeps provenance, timestamps, system boundaries and known gaps attached, so a reader can tell which layer supports which claim. Confidence in prose is not, regrettably, a log format.

What the Anthropic research tested

Anthropic's research post “Reasoning models don't always say what they think” examined whether displayed chain-of-thought faithfully reflected what influenced particular reasoning models. The researchers used deliberately constructed evaluation settings: they planted hints in prompts and checked whether models that used a hint acknowledged it, and they built reward-hacking scenarios to see whether exploited shortcuts appeared in the displayed reasoning.

The reported percentages belong to those models, prompts, hints and scoring procedures. They should not be stretched into a general rate for every model or task, and the post itself discusses limits on how far its findings generalise. The careful conclusion is narrower: in those experiments, displayed reasoning was not always faithful to the factors that influenced the answer. That is enough to make an explanation insufficient as sole evidence. It is not evidence that every model always conceals its reasoning.

Transparency is not one switch

NIST's AI Risk Management Framework treats transparency, explainability and interpretability as related but distinct characteristics of trustworthy AI. Transparency concerns what information about a system and its outputs is available to the people interacting with it. Explainability concerns representing the mechanisms behind a system's operation. Interpretability concerns the meaning of an output in the context of its intended purpose. A fluent explanation may help with one and say little about the others.

The NIST AI RMF Core organises risk management around Govern, Map, Measure and Manage, with outcomes covering documentation, monitoring and accountability. That emphasis points toward records kept by processes and people, not only stories told by the system. WentRogue uses this framing as a reading aid; it does not claim NIST certification or formal compliance.

What WentRogue records—and what it cannot prove

WentRogue's core public record has seven fields: light ID, light number, self-declared Agent ID, UTC timestamp, base amount in US dollars, the buyer's declaration and whether an owner email was supplied. A completed light represents an eligible purchase that was paid and fulfilled. Those are bounded external facts, published so readers can check what the record says.

The record deliberately does not publish private prompts, model versions, hidden reasoning or complete execution traces. A declared Agent ID is not verified as a unique agent, model, person or organization, and a declaration is a self-report rather than proof of authority. So the record cannot prove an agent's internal cause or identity. It records an outcome and a claim, not a forensic reconstruction.

That boundary is a feature. A claim that states what it can and cannot support is easier to audit than an explanation that sounds complete.

References