The agent sent a message, changed a record, or refused a tool. When someone asks what happened, the team finds only the conversation and a trace full of spans. The link between the person who started the run, the active authority, the approved decision, and the system that confirmed the effect is missing.

To build an audit log for an AI agent, record structured events that connect identity, intent, policy, action, and outcome. Include blocked attempts, not only successful calls. The model may propose an action, but the runtime or the system that executes it should record what was allowed, what started, and what was actually confirmed.

This is different from storing every prompt. A trace of AI agent tool calls with OpenTelemetry explains execution order and duration. The audit log must answer whether the action had authority, which approval covered it, and what evidence remained afterward.

Diagram showing an AI agent passing through a policy decision before a tool call and a verified effect receipt.

Short answer

  • Record runId, traceId, the human or service identity, and the agent and model versions.
  • Record the proposed action, tool, target, policy, and decision, including denied.
  • Bind an approval to the exact payload that was approved, not only to the tool name.
  • Separate completed from unknown: a timeout does not prove that the effect did not happen.
  • Keep references, hashes, and redacted results when the full payload contains sensitive data.

A trace and an audit log answer different questions

The trace answers: “In what order did the operations happen, and where did the run fail?” The audit log answers: “Who authorized this action, which policy applied, and what result can be verified?” A system can produce both from one execution, but it should not treat them as the same record.

The OpenAI Agents SDK tracing guide describes traces that include model generations, tool calls, handoffs, guardrails, and custom events. The Agents API tracing guide also organizes tool calls by agent, session, turn, and span, with arguments, results, and status when those data are recorded. That is useful for diagnosis.

An audit adds a responsibility relationship. The event should say on whose behalf the action happened, which agent and version were active, which authority scope was used, and whether the action was proposed, allowed, denied, executed, or confirmed. The trace can point to the event with traceId; it does not replace the event.

Do not confuse the agent's final summary with proof. The model may say that it updated an order when the tool failed, was blocked, or timed out before anyone checked the result. The writer of the outcome event belongs at the boundary that received the real tool response or read the external system back.

Which questions should the record answer?

Start with the incident you would want to investigate six weeks later. A useful record answers the questions below without reconstructing the story from five incompatible logs.

Question Minimum fields Preferred source
Who started the work? actorId, identity type, runId gateway or runtime
Which agent made the decision? agentId, version, modelId, config version runtime
What was the goal? task reference or redacted intent application, not only the model
What action was proposed? tool, operation, target, actionHash runtime before execution
What authorized it? scope, rule, decision, and approvalId policy layer or gateway
Was the action blocked? decision: denied, rule, categorized reason policy boundary
What was executed? attempt, external request ID, start and end tool executor
What was the effect? status, receipt, resource reference, or read-back affected external system
What is still uncertain? outcome: unknown, reason, next step runtime after timeout or crash

These fields are not a universal standard. They are a minimum contract that links proposal, authorization, execution, and consequence. The OWASP Securing Agentic Applications Guide includes logging plans, validation, tool calls, approval interactions, errors, and state changes among its secure-operation practices. The list above turns that concern into a boundary a team can implement and test.

Do not record “the agent decided” as if that were an identity. The identity comes from the user, service, or credential that started the operation. The model and version explain which component generated the proposal. The policy explains why the executor accepted or refused it. These are different roles.

Model the action lifecycle, including what did not happen

A single tool_call: success event is too small for an action that passes through approval, a queue, a retry, and external confirmation. Use states that represent observable transitions. One possible set is:

proposed -> approved -> started -> completed
                 \-> denied
started -> failed
started -> unknown -> reconciled

proposed means the agent suggested an operation. approved means a policy or person authorized that payload. denied is useful and should be recorded even when no external system was called. started marks the point when the executor crossed the authorization boundary. completed should appear only when the tool or external system returned enough confirmation.

unknown matters for actions with effects. If a worker loses its connection after sending a request, the local process knows that it lost the response, not that the operation failed. The runtime can query the external system, use an external request ID, or apply an idempotent operation before changing unknown to reconciled or failed.

This complements the guide to pausing and resuming an AI agent. The checkpoint stores the state needed to continue. The audit log stores the transitions and evidence the runtime observed before the interruption. One should not replace the other.

Bind approval to the exact action

An approval that says “you can update the CRM” is too weak for an audit. The same tool name can receive different targets, filters, and values. The approval should point to a stable representation of the operation that will run.

An implementation can store a normalized action, a hash of the redacted payload, and the scope that was evaluated:

import { createHash, randomUUID } from "node:crypto";

type Decision = "proposed" | "approved" | "denied";
type Outcome = "started" | "completed" | "failed" | "unknown";

type AuditEvent = {
  id: string;
  createdAt: string;
  runId: string;
  traceId: string;
  actorId: string;
  agentId: string;
  agentVersion: string;
  eventType: "decision" | "execution";
  toolName: string;
  target: string;
  actionHash: string;
  decision?: Decision;
  outcome?: Outcome;
  policyRule?: string;
  approvalId?: string;
  externalRequestId?: string;
  evidenceRef?: string;
};

function hashAction(action: unknown): string {
  return createHash("sha256")
    .update(JSON.stringify(action))
    .digest("hex");
}

function makeAuditEvent(
  base: Omit<AuditEvent, "id" | "createdAt" | "actionHash">,
  action: unknown,
): AuditEvent {
  return {
    ...base,
    id: randomUUID(),
    createdAt: new Date().toISOString(),
    actionHash: hashAction(action),
  };
}

The code is illustrative. It shows the contract shape, but it does not implement append-only storage, authorization, redaction, concurrency control, or external confirmation. In production, the approval must be consumed by the executor together with the actionHash; if the payload changes, the approval no longer matches the action.

Do not use the hash as proof that the action ran. At most, it proves that a representation was associated with the event. The result needs the status returned by the tool, an external identifier, or a later read of the resource. The difference between “I sent it” and “the system confirmed it” is the part the incident will try to clarify.

The runtime should write the outcome, not the model

The model can suggest a reason, tool, and arguments. It should not be the author of the line that says a payment was sent, a file was deleted, or a permission was changed. The executor knows the request ID, return code, and exception. A policy gateway knows the decision and scope. The external system may provide the final receipt.

That does not mean throwing away model output. Keep a reference or redacted summary when it is needed to explain the proposal. Just distinguish model_output from execution_result. The first is an input to review. The second is an observation from the boundary that executed or checked the action.

The Microsoft Agent Framework observability guide shows why this separation matters: prompts, responses, arguments, and sensitive results are disabled by default, and enabling them can expose confidential information. Telemetry configuration does not replace a retention policy. Even when a trace contains an argument, the audit log may keep only a hash, a protected reference, and the minimum set needed for review.

A practical split is:

  • Model: proposes intent, tool, and arguments.
  • Policy: decides whether to allow, require approval, or deny.
  • Executor: records start, end, error, retry, and request ID.
  • External system: confirms the change or supports a reconciliation read.
  • Audit layer: joins the references and exposes a reviewable view.

This division reduces the chance that a plausible narrative becomes the only version of what happened.

What not to log by default

A complete audit log is not a prompt warehouse. Storing everything can widen the leak you meant to investigate, increase retention costs, and give access to people who only needed to see the decision.

Start by blocking these by default:

  • tokens, cookies, keys, and authorization headers;
  • complete system prompts when a versioned reference or hash answers the question;
  • personal and financial data that is not needed to identify the target;
  • the full body of large external responses;
  • private model reasoning treated as if it were a verifiable explanation.

Use identifiers, data classification, size, status, hash, a redacted artifact reference, and an access policy instead. Keep full content in separate storage only when its purpose, retention, and access controls are defined. Redacting it only in the dashboard is too late: the data may already have crossed the transport, exporter, and storage layers.

Do not promise an “immutable log” because the application writes JSONL. Append-only is a property of storage and authorization, not of the file format. If a review requires protection against edits, define who can write, who can read, how integrity is checked, and how long the record exists. This article does not turn those choices into a legal compliance conclusion.

How do you test whether the log tells the right story?

Make the audit fail in a controlled environment. Choose a tool with a fictional effect and verify every transition, not only the agent's final answer.

  • Run an allowed action and check proposed, approved, started, and completed.
  • Try an out-of-scope action and confirm that denied exists without an external call.
  • Change an argument after approval and verify that the actionHash no longer matches.
  • Make the tool return an error and confirm failed, with secrets removed from the message.
  • Cut the connection after the send and confirm unknown, without turning the timeout into “it did not run.”
  • Reconcile by external request ID and record the receipt or the reason it remains uncertain.
  • Cross a handoff and check that runId, traceId, authority, and agent version remain connected.
  • Send a sensitive argument and confirm that the exporter receives the expected redacted form.

The most important test is done by someone who did not write the runtime. Give that person a runId and a concrete question, such as “which agent tried to change this record, which rule allowed it, and what did the system confirm?” If they have to read the code to understand the sequence, the record is still internal telemetry, not a useful audit trail.

For operations that can be repeated, combine this test with verifying an AI agent tool call before retrying and testing API idempotency without duplicate side effects. The log should record the attempt and reconciliation, but it cannot fix an executor that repeats effects.

Trace, log, metric, or checkpoint?

Use each signal for the question it can answer:

Signal Main question Does not replace
Trace What was the order, and where did time or error appear? proof of authorization or external state
Audit log Who proposed, allowed, denied, and executed the action? run-state storage
Metric How many failures, denials, or actions occurred? explanation of one case
Checkpoint Where can the runtime continue? proof that an external effect was confirmed
External receipt Which system confirmed which change? the decision context

OpenTelemetry maintains conventions for traces, events, and GenAI signals, but a telemetry convention does not automatically define your application's audit policy. Use trace and span IDs for correlation when it is safe. Keep the audit contract stable even if you change observability backends.

Frequently asked questions

Do I need to store the entire prompt to audit an agent?

No. Keep a versioned reference, hash, or redacted summary when that is enough to identify the context. Preserve the full content only when the purpose and access policy justify it. The record should explain the decision without turning everything the agent saw into a permanent copy.

Is an OpenTelemetry trace already an audit log?

Not necessarily. A trace organizes operations, duration, relationships, and errors. An audit log adds identity, authority, policy decision, approval, exact action, and verifiable result. You can produce both together and connect them with traceId, but you should not assume that a trace dashboard preserves the evidence or retention needed for a review.

Should I record denied actions?

Yes, when the attempt helps answer what the agent or user tried to do and which rule blocked it. A denial may be the most important event in a prompt-injection incident or a bad configuration. Record a categorized reason and policy reference without copying unnecessary sensitive data.

Can the log prove the model's intent?

No. It can record the output that proposed the action, the model version, and the selected context. That helps investigation, but it does not turn a later-generated explanation into causal proof. Authorization, execution, and effect should be recorded by components that control those boundaries.

Conclusion

An agent is not auditable because it produces more text. It is auditable when each meaningful action leaves a verifiable link between identity, authority, policy, execution, and effect.

Start with a small event. Record what was proposed, allowed or denied, started, and confirmed. Preserve unknown when the response disappears. Redact content that does not need to be in the record, and test the story with someone who does not know the code.

The trace explains the path. The checkpoint lets the runtime continue. The audit log supports the question of responsibility. Mixing these jobs creates large records and still leaves the main question unanswered.

How this analysis was done

Samuel Fajreldines is the accountable author of this article. The research compared current documentation from the OpenAI Agents SDK, OpenAI Agents API, Microsoft Agent Framework, OpenTelemetry, and OWASP with recent public discussions and the site's existing cluster. The event contract and verification checklist are editorial synthesis. The TypeScript code is illustrative and was not run against an agent or audit backend. AI assistance supported discovery, drafting, image generation, localization, and consistency review; it did not provide production experience or replace source verification. The author also uses RemoteCode as a work tool.

Sources consulted