Token Cost Attribution by Developer, Team, and Agent Task
Agentic AI workflows expose the hidden costs that flat-rate models masked.

A chat interaction with an AI coding assistant has a predictable shape: a developer asks a question, the model answers, the exchange ends. An agentic session has no such edge. It runs in a loop, reading files, writing code, running tests, reading the failures, and trying again, sometimes for dozens of steps without a human in the middle of it. That structural difference is why token spend has become hard to forecast, and it arrived at the same moment every major AI coding platform dropped flat-rate subscriptions in favor of billing by token consumption. Spend is now a function of how much work an agent actually does, not how many seats a company has bought. A 20-step agent loop does not cost twenty times what a single chat turn costs. It costs far more than that, because each step in the loop re-sends everything the agent has already seen, so the context grows with every turn and the cost grows faster than the step count does. That growth pattern rarely appears in how engineering teams build their budgets, which is part of why actual bills have started to surprise them. Research from Zylos traces this shift directly to the move from single-session coding tools to agents that operate as participants in CI/CD pipelines, touching dozens of repositories at once. Once an agent is doing that, "who pays for what" is no longer a finance question that can be answered by looking at a seat count. It is an engineering problem, and it needs an engineering answer.
How token spend compounds inside a single agentic session
Most of what an agentic session spends tokens on has nothing to do with the code it eventually writes. A session opens by mapping the repository, so the agent can understand which files relate to which, and it follows that by pulling the relevant source files into its active context window. The system prompt and whatever project-specific instructions exist get sent again in full on every single turn, not just once at the start. Before any code gets written, the agent typically produces a chain of reasoning output, working through the plan step by step, and that reasoning consumes tokens the same way any other output does. The actual code edit, the part a developer would call the point of the exercise, is often a small fraction of the total tokens the session consumes. Controlling token spend means controlling how much context gets re-sent on each turn, how many retries a failing test triggers, and how long a session is allowed to run before someone decides it has gone far enough. Zylos's research finds that sessions touching more than one repository in a single turn add a complication: shared-library reads, context injected to support reasoning across projects, and the synthesis step that pulls everything together all burn tokens that cannot be cleanly handed to any one project or owner. A further wrinkle comes from caching. A cached system prompt written by one project's CI pipeline costs more to write the first time and less to read on every use after that, so a naive attribution scheme charges the project that happened to write the cache for the full cost while a later project that merely reads it gets a windfall it never earned.
Why spend is wildly uneven across a team
Token spend does not spread itself evenly across a team, and looking only at the total obscures the information that unevenness actually carries. A small group of developers can account for a disproportionate share of a team's total token volume, and that concentration can mean two very different things. It might mean those developers have built genuinely productive agentic workflows that the rest of the team should learn from and adopt. Or it might mean they are running expensive models on tasks a cheaper model would have handled just as well, burning budget without producing commensurate value. There is a known edge case in which a single developer's agent usage dwarfs an entire team's, illustrating how extreme the spread can get, but the more important pattern is the ordinary one: most teams have a wide, unremarkable spread between their highest and lowest consumers, and that spread goes unexamined because nobody is looking below the aggregate number. Without visibility at the level of the individual developer and the individual task, an engineering leader cannot tell a cost spike caused by one team's heavy agent-mode usage apart from a spike caused by a single repository's unusually large context, or from a model-tier mismatch where a task that needed a lighter model got routed to a heavier one. Aggregate figures cannot answer any of those questions, because only the developer-level and task-level detail can distinguish productive heavy usage from a model-tier mismatch or a context-heavy repository.
Why naive attribution approaches break down
Two attribution approaches tend to occur to teams first, and agentic workflows break both of them. The first charges whichever repository the agent happens to be working in at the moment of inference. The second splits the session's total cost evenly across every repository the session touched. Both fail for structural reasons, not edge-case ones. Consider a single agentic task inside a monorepo refactor that touches a core utilities package, three feature services, and an infrastructure module in the course of one run: the repository that happens to be "active" at any given moment is simply an artifact of the order in which the agent chose to work, and has no real relationship to which repository actually caused the cost. Splitting the bill evenly across repositories fails in the opposite direction. It undercharges the complex codebase that generated most of the context the agent had to process, and overcharges the simple one that barely contributed anything, which erases the very distinction that attribution exists to preserve. A gap in the tooling itself causes both failures. The OpenTelemetry GenAI semantic conventions, the standard many teams rely on for instrumenting model calls, do not define attributes for tenant, project, or cost center. Teams have to extend the schema themselves with custom attributes, and then make sure those attributes propagate through W3C trace context headers as a request moves through the system. Skip that propagation step, and a multi-hop agent chain loses its attribution context at every hop, so by the time the session finishes, nobody can reconstruct who was responsible for which part of the cost. Caching produces its own version of the same failure. The project whose pipeline writes a cached system prompt absorbs the write cost, while the project that later benefits from the cache hit gets to use it for less. Direct attribution hands one project the whole bill and the other a free ride, and averaging the cost across both simply hides that the dynamic exists.
The three architectural patterns that production systems use instead
Zylos's research finds that the industry has converged on three patterns that, together, solve what the naive approaches cannot: gateway-level direct attribution, hierarchical budget enforcement, and cost-aware model routing. Most production systems run a hybrid of all three. Gateway-level direct attribution works by attaching metadata, team ID, repository, pipeline step, cost center, to every request at the gateway layer, before the request ever reaches the model. A request arriving without complete attribution headers gets rejected outright with a hard error rather than quietly falling into a default bucket, which treats missing attribution as a structural failure that needs fixing, not a minor gap to shrug off. This is the most accurate of the three patterns, and it's the one production agentic CI/CD platforms have adopted as the backbone of their cost architecture. The accuracy comes at a price: every call path has to carry the metadata. Enforcing the contract has to happen at the API layer itself rather than hoping individual developers remember to tag their requests. Direct attribution handles the inference calls that can be cleanly tagged at the moment they're dispatched, but shared infrastructure costs, a vector database query used by several projects at once, a system prompt loaded once and reused across many sessions, need proportional allocation instead. The most accurate version of proportional allocation weights costs by complexity and task type, factoring in both how much was used and what kind of work it supported, but computing those weights carries its own overhead, and that overhead only pays for itself at real scale. Most production systems therefore mix the two: direct attribution wherever tagging is clean, proportional allocation wherever it isn't. The third pattern, cost-aware model routing, decides which model handles a given task based on remaining budget headroom before the task runs, not after the bill arrives. A task that would push a team over its remaining budget gets rerouted to a cheaper model or held in a queue at a reduced price. Combined with prompt caching and tighter context management, this routing layer produces real reductions in total production cost by stopping expensive calls before they happen.
Why budget controls need to act before execution
A billing alert tells a team it has already crossed a spending threshold. A token budget stops the spend before it happens, and in agentic workflows that distinction carries real weight, because a single task can trigger repeated model calls, repeated tool attempts, and a steadily expanding context window well before any monthly alert has a chance to fire. The production standard for handling this is hierarchical budget composition: organization-level caps roll up into team-level budgets, which roll up into budgets tied to individual virtual keys, which roll up further into provider-level configuration budgets inside a single key. That structure lets a team set real caps at the level of the session, the developer, the team, and the time period, rather than settling for watching an aggregate number after the fact. Where the enforcement sits in the pipeline matters as much as what level the budget is set at. Governance built into the execution loop catches a violation before it happens. Governance applied as a post-hoc audit only catches it once the money is already spent. Hard caps at the session level work as a necessary companion to per-developer attribution: knowing after the fact that one developer drove a cost spike last month has some value, but stopping that spike from happening in the first place has considerably more. The risk runs in one direction far more than the other, so the constraint belongs before execution rather than after it.
The compliance and audit requirements that token attribution now has to satisfy
Token attribution has moved from a cost-management nicety to something closer to a prerequisite for compliance. The stakes behind this shift came into sharp relief with the Cline incident in February 2026, when an unauthorized version of the Cline CLI package was published to npm and remained live for roughly eight hours, exploiting a prompt injection vulnerability in an AI-powered issue triage workflow that let an attacker with any GitHub account compromise production releases. Publishing credentials for npm were exfiltrated through a poisoned GitHub Actions cache. The incident is a concrete reminder that an AI agent with privileged access, if it isn't tracked inside a governed registry, becomes an attack surface a company didn't know it had. Zylos's research lays out what a credible production agent compliance program actually requires: five layers working together. A versioned registry of every agent lists the underlying models, prompts, tools, retrieval sources, owners, risk classification, and approval signature behind each one, with every entry carrying a unique ID and a timestamp. Alongside the registry sits a history of prompt and policy versions, a record of per-request OTel traces, a log of incidents, and periodic compliance reporting that ties the other four together. Shadow AI agents, the ones running with persistent privileged access outside this sanctioned registry, represent one of the most dangerous risks the industry has yet to deal with properly, for a simple reason: a company cannot govern an agent it has no way of attributing to a project, a team, or an owner.
What a complete attribution system looks like
Attribution by developer, by team, and by agent task are not three separate dashboards sitting side by side. They are one system, and the data has to hold together consistently at every level or none of it can be trusted. Per-developer attribution surfaces how consumption actually spreads across a team, who is driving volume, through which modes of interaction, against which repositories, and it's what makes per-developer budget caps and session-level enforcement possible in the first place, so that one developer's retry loop can't quietly consume a whole team's monthly allocation. That attribution only becomes genuinely useful once it connects to outcomes: token spend tied to a developer matters when it can be joined to the commits, pull requests, and shipped features that spend actually produced. Without that join, the numbers answer a cost question but leave the value unmeasured. Per-team attribution does different work. It supports chargeback and showback across cost centers, which matters for any organization splitting AI spend across multiple business units, and it reveals the shared-context costs that cross team boundaries, costs that no single team's view would show on its own. Per-task attribution is the smallest and arguably the most honest unit of the three, because cost per solved task, not simply cost per task attempted, captures the retries and failures a session absorbed along the way to a result. Put the three together, built on gateway-level attribution, enforced through hierarchical budgets, and backed by the compliance layers a registry and audit trail provide, and the result is a system that can finally answer the question that flat-rate pricing never had to ask: not just how much was spent, but who spent it, on what, and whether it was worth it.


