Agent in Production
ObservabilityLong read

Structured Logging Schema for AI Coding Agent Sessions

How agents make nondeterministic decisions requires a fundamentally different audit trail.

Staff Writer · · 10 min read
Cover illustration for “Structured Logging Schema for AI Coding Agent Sessions”
Observability · October 4, 2026 · 10 min read · 2,223 words

Standard application logs fail AI coding agent sessions for a structural reason: they were built to record deterministic execution paths, and agents do not execute deterministically. A developer debugging a failed deployment on a Saturday morning pulls up the session log expecting the same thing an application log has always given: a sequence of calls, a status code, a stack trace if something broke. What the log shows instead is a service account, a tool name, and the word "completed," with no trace of why the agent picked the migration path it picked or whether anyone had authorized it to touch that database. The tooling isn't missing fields so much as missing a premise: it assumes that running the same input twice produces the same output, which is the one guarantee an agent session cannot make.

Why standard application logs fail for agentic workflows

The guide Auditing and Logging AI Agent Activity names three specific ways this breaks down in practice. There's no link between the human intent that started a session and the sub-tasks the agent generated to pursue it. There's no record of the reasoning behind a given tool call, only the call itself. And there's identity dilution: a generic service account fires every action, so the log cannot distinguish an agent decision from a human override issued through the same credentials. These three gaps compound an existing problem. Industry analysis already puts SRE incident triage time at up to 50% spent manually filtering and correlating distributed text logs, before agents even enter the picture. Add non-deterministic decision-making on top of that baseline, and triage becomes closer to archaeology than to a search problem.

This forensic cost appears most clearly during replay. Re-running the same prompt against the same agent does not guarantee the same decision path, a property best named replay failure, because it breaks the standard debugging move of reproducing the bug to isolate it. And the practical cost occurs in exactly the scenario above: a log entry reading "status": "completed" carries no information when the thing that went wrong was the agent's judgment, not its execution. The question a session log has to answer is why the agent chose a given path, not merely whether the step it took finished without error. That question matters more now because agents no longer make single bounded calls to a model and stop. They run extended chains of actions, delegate pieces of the task to subagents, and take actions with real consequences in production systems. An audit trail built for one model output has to become a record of a whole decision-and-action chain instead, which is a different kind of record, built to a different schema.

The four layers of information a session log must contain

Diagram: Four Layers Every Agent Session Log Must Contain. Visualizes: Visualize the four mandatory layers of an agent session log as a vertical stack, each layer paired with the core question it answers and its key fields.

A session log adequate to that chain has to capture four distinct layers: intent, tool calls, state deltas, and cost attribution. Each layer answers a question none of the others can answer, which is what makes all four mandatory rather than optional refinements on top of a basic log.

The intent layer answers why the session existed in the first place: what the human asked for, who authorized the request, and under what delegated scope the agent was permitted to act. The tool call layer answers what happened during execution: which actions the agent took, with what parameters, in what order, and with what result. The state-delta layer answers what is different in the world now compared to before the session ran, a question execution logs alone cannot answer because a tool call recorded as "success" says nothing about what it actually changed. The cost attribution layer answers what the session cost and who is accountable for that spend, down to the token and the triggering identity.

Leaving any one of these out creates a specific, predictable blind spot. A schema that logs only tool calls and cost can tell you that a call failed and what it cost to fail, but not whether the agent was operating inside its authorized scope when it made the attempt. None of these partial schemas are broken in the sense of containing errors. They're incomplete in a way that only becomes visible at the exact moment someone needs the missing piece, usually during an incident.

This points to a broader claim about what a session log is actually for. Audit, in the framing of the five-plane reference architecture built for agentic systems, is evidence production, built from the start for the auditor, the regulator, and the incident responder. That reframing means the log itself has to serve as the governance instrument, not a feed piped into some separate governance system later. The four layers below are what that instrument looks like in practice.

Layer one: the intent fields that anchor every session to a human and a scope

Every tool call record in a session is forensically unanchored until something establishes who authorized the session and what it was permitted to do. A log can show, in perfect detail, that an agent ran a destructive database command, and still leave the one question that matters unanswered: did anyone grant it the authority to do that. Intent fields exist to close that gap, and they have to be written at session creation, not reconstructed afterward from whatever trace happens to survive.

A session record needs a session_id that correlates every downstream event back to one origin point, so the full chain can be reassembled later regardless of how many subagents or tool calls branch off it. It needs a parent_identity field recording the identifier of the human or system that authorized the session, which establishes the chain of responsibility a compliance review will eventually ask for. It needs delegation_scope, the specific OIDC permissions granted for that session, which is the record proving the agent stayed inside its security sandbox rather than simply acting as if it had. It needs session_goal, the plain-language statement of what the human actually wanted, captured before the agent starts planning its own sub-tasks. And it needs policy_version, recording which governance policy was active at the moment the session started, since policy changes over time and the version in force when the agent acted is the one that matters for judging what happened.

Auditing and Logging AI Agent Activity is direct about why identity fields carry this much weight: every agent has to function as a distinct non-human identity, with its own lifecycle governance, scoped permissions, and verifiable authentication. The five-plane architecture adds a related concept for schema design: composite principals, where authority passes down a delegation chain and narrows at each step, a process called capability attenuation. The delegation_scope field is the record of that narrowing. It's what lets an auditor confirm that an agent's authority never exceeded what was actually granted to it.

The practical stakes become clear in a simple comparison. A session triggered by a CI pipeline holding database-write credentials represents a fundamentally different risk surface than a session triggered from a developer's terminal holding read-only credentials. Nothing about a tool call log by itself communicates which of those two situations produced a given action. The intent layer makes that distinction visible in the record.

Layer two: the tool call fields that reconstruct the execution chain

Most agent failures happen at the moment of tool invocation, which makes this the layer where schema granularity matters most. A flat log line noting that a tool ran and returned a result is not a record of what happened: it's a label. A tool call entry built as a stringified blob flattens each invocation into a single line of prose, while one built as a proper structured record gives each invocation its own complete set of fields.

Each tool invocation needs a tool_name identifying the specific tool or MCP server endpoint called. It needs tool_params, the full JSON of arguments passed to that tool, where a hallucinated parameter, a wrong table name, a malformed file path, first appears. It needs parent_call_id, so that when an agent delegates work to a subagent that itself makes further calls, the causal chain between them survives in the record. It needs reasoning_context, the agent's stated justification at that specific decision point, not its full chain of thought but the narrower record of which options it weighed and which one it picked. And it needs policy_decision, the outcome of whatever governance check ran before the call reached the wire.

Consider a call like postgres_query running against a production database. That last field carries more weight than it might appear to at first glance. Auditing and Logging AI Agent Activity treats prompt injection as a live risk precisely because standard logs treat tool output as routine data retrieval with nothing to flag. The tool_params and raw_output fields are the only mechanism available for catching an injected instruction that arrived disguised as an ordinary tool response. And the record that a governance check ran before a call, separate from the record that the call itself happened, is what separates a log from an audit trail. A team that logs the tool call but not the policy decision that preceded it can show that something executed. It cannot show that the execution was compliant.

Layer three: state-delta fields that show what the agent changed, not just what it did

Recording that an agent called a file-write tool is a different fact than recording that a file actually changed. The gap between those two facts is where state-delta logging does its work: it captures what a resource looked like immediately before an action and immediately after, which is the information that makes replay, rollback, and forensic reconstruction possible after the fact. Early implementations are likely to skip this layer, because it takes more storage and more discipline than recording the call alone, and because its value only becomes obvious during an incident serious enough to need rollback.

A state-delta record needs a pre_state_snapshot, capturing the relevant state of the affected resource immediately before the action ran. It needs a delta_type, classifying the change as a create, modify, delete, or read, since reads matter as much as writes for any audit focused on data access rather than data modification alone. It needs affected_resource, naming the specific file, database record, API endpoint, or system entity the action touched. Irreversible actions, by virtue of that flag, should route to a human approval step before they execute rather than after, which is the direct operational payoff of tracking reversibility.

The paper on Agent-Native Telemetry formalizes this idea under the name verifiable state-delta evidence, structuring operational facts into four primitives: Transitions, Observations, Relations, and State Checkpoints. The same paper formalizes something called a verified negative theorem: a properly structured state-delta ledger can prove that an event did not occur, not only that it did. That capability marks the real line between a log and evidence. A log can tell you what was recorded. A ledger built this way can tell you, with proof, what never happened.

Memory operations deserve the same treatment as file or database changes, not a lighter one. Memory deletes carry particular weight under GDPR, where the ability to prove a specific piece of data was actually erased is itself a compliance requirement. The state-delta layer is the only layer in this schema built to catch that kind of compound, cumulative change, because it's the only layer tracking the actual before-and-after state of the world rather than the individual actions taken against it. Reconciling the reasoning record against the change record is what settles that question.

Layer four: cost attribution fields that make per-session spend controllable, not just visible

A token count and a latency figure logged once per session, as an aggregate, answer a billing question well enough but answer almost no governance question. Turning spend into something controllable, rather than something merely visible after the fact, requires attributing cost to the specific person or trigger that authorized the session, at a level of detail fine enough to let a cap actually stop a session before it does damage.

A cost-attribution record needs input_tokens and output_tokens counted per model call and summed per session, since a single session can span many calls with very different cost profiles. It needs model_id and model_params, recording which model ran at what temperature and token limit, because the same task run on a different model or a different setting produces a materially different cost. And it needs latency_ms per call along with total session latency, both for SLA monitoring and because a sudden latency spike is often the first visible signal of a runaway loop before the token count confirms it.

A billing alert tells a team, after the fact, that spending has crossed some threshold. A token budget enforced before execution can block a request, reroute it, or cap it outright, which makes pre-execution budget fields considerably more useful operationally than a spend total computed after the session has already finished. The risk this guards against is specific to agentic systems: a single task can trigger repeated model calls, repeated tool attempts, and an expanding context window, all before the session itself ever reaches completion. A session-level token total logged only at the end gives no warning while any of that is happening. It only confirms, once the session is over, how much damage already occurred.

Sources

  1. Agent-Native Telemetry: Verifiable State-Delta Evidence for Autonomous Operations
  2. A Five-Plane Reference Architecture for Runtime Governance of Production AI Agents
Filed underObservability

More in Observability