Introduction
Every token an AI agent spends traces back to a decision. What to read. Which tool to call. Whether to loop one more time. Yet most enterprises are trying to control agent costs by governing the bill, not the decisions that produce it.
At DAXA, we see runaway spend as the most visible symptom of an ungoverned agent. The same gaps that let an agent read data it should not, or take actions nobody authorised, are the gaps that let its costs climb unchecked. Govern the decisions, and the cost follows.
The research now points the same way. Gartner names ungoverned autonomy and bloated context windows as common drivers of token overspending, and Forcepoint's X-Labs research shows agents running up costs through attacks, misconfiguration and ordinary use, often without any single request looking unusual. AI agent cost control is no longer a budgeting exercise. It is a security control.
Why Are AI Agent Inference Costs Rising?
Tokens are getting cheaper, but agents use far more of them, on more expensive models. Gartner calls this the Inference Paradox.
Gartner predicts that AI inference costs per agentic workflow will rise more than fivefold through 2028. It identifies three trends behind this:
- Foundational model cost economics are improving quickly.
- That efficiency is unlocking more powerful, and more expensive, models for more sophisticated applications.
- Sophisticated AI workflows use far more tokens than simple chatbot interactions.
The result is that tokens are becoming more cost-efficient, but not as fast as the capabilities, and costs, of the workflows built on them.
The difference between a chatbot and an agent explains much of the gap. A chatbot reads a query and responds. An agent plans, calls tools, checks its own work and tries again. According to Gartner, routing a task to an agentic reasoning model increases provider inference costs by at least five times compared with a basic chatbot interaction, and often much more as complexity grows.
Pricing is shifting too. In separate research on AI coding costs, Gartner notes that coding agent vendors are moving from seat-based to consumption-based pricing, often with little transparency into how tokens are calculated and billed. It predicts that by 2028, AI coding costs will overtake the average developer's salary.
The effect is already visible. Uber reportedly exhausted its entire 2026 AI coding tools budget within four months as Claude Code spread across around 5,000 engineers. We covered the budgeting side in enterprise AI cost governance.
What Causes Runaway AI Agent Spend?
Structural trends raise the baseline. Runaway spend comes from ungoverned agent decisions: what an agent retrieves, what it does, how long it keeps going and whose identity it spends under.
Gartner is direct about this. It links token overspending to how usage is governed, naming ungoverned autonomy in agent-driven workflows, bloated context windows and the absence of structured feedback among the common failure modes. As analyst Nitish Tyagi puts it: "Token discipline will not emerge through developer choice alone."
This is the lens DAXA applies. Instead of asking how much an agent spent, ask which decision produced the spend, and whether it should have been allowed.
This is also why cost is now a security issue. Forcepoint's research shows that in these scenarios, no single request looks malicious or breaks usage limits, even as the total climbs. The same gap produces accidental overspend and deliberate abuse, and from the invoice alone, the two look identical. OWASP now ranks this risk, unbounded consumption, sixth in its 2026 Top 10 for LLM Applications, alongside the others we mapped in the OWASP LLM and Agentic Top 10.
Budget caps cannot close this gap on their own. They act after the spend happens, cannot tell legitimate work from abuse, and halt every valid task sharing the same key.
How DAXA Controls AI Agent Spend at the Source?
DAXA governs the decisions that create spend: what agents retrieve and what they do, evaluated in context at runtime.
Traditional IAM answers who can access which system. It was not built to reason about what an autonomous agent does once inside. DAXA extends it into Identity and Action Management: existing access permissions are enforced first, then the platform reasons over whether an agent should retrieve, write, delete or share specific data in that context, with enforcement at AI runtime.
Applied to agent cost, that means:
- Governed retrieval. Pebblo classifies data by confidentiality, compliance and semantic context, and evaluates each retrieval against user context, document intent and application context. The agent pulls in what its task needs, not everything it can reach, which limits context growth and data exposure together.
- Governed actions. Context-aware enforcement stops unauthorised actions before they happen, regardless of which agent, application or model is behind the request.
- Behavioural reasoning. DAXA's guardian agents maintain awareness of business workflows and data relationships, so the platform can reason about subtle signs of compromise or misuse rather than matching patterns.
- Real-time visibility. A data bill of materials shows what AI agents and applications access, modify, generate and share, giving teams the attribution that cost anomalies need.
DAXA does not replace spend limits or inference controls. It covers the layer they cannot see.
Tool fan-out shows why the third layer matters. As we covered in indirect prompt injection in AI coding agents, untrusted content shaping an agent's next step is a governance problem. We apply the same model to homegrown AI agents, AI coding assistants and across Pebblo.

How to Control AI Agent Costs: 8 Steps for Enterprises
Start with the decisions, then cap the totals.
- Match the model to the task. Gartner recommends aligning model selection with task complexity: route simpler, high-frequency tasks to smaller models and reserve frontier models for complex, high-value work.
- Set autonomy levels by use case. Define when agents should be used and how much autonomy each task gets. Gartner suggests classifying work as developer-led, developer-with-agent or fully agent-led.
- Govern what agents retrieve. Gartner recommends context engineering: including only relevant information and removing unnecessary data. Enforcing this at runtime keeps context lean and stops agents reading data they have no reason to see.
- Govern what agents do. Check each action against the agent's task and identity before it runs. This separates legitimate heavy work from tool fan-out, which spend caps alone cannot do.
- Bound loops and apply least privilege. Forcepoint recommends capping agent steps and loop-backs, detecting repetitive loops early, and using sandboxing and least-privilege access to limit how far a runaway process can spread.
- Attribute every model call. Trace each call to a known identity and workload, and keep API keys out of code repositories. Without attribution, a leaked key and legitimate growth look the same.
- Set thresholds and hard limits. Gartner recommends token thresholds, escalation policies and automated monitoring. Forcepoint adds hard caps per user, key and team, rather than alerts that fire too late.
- Review high-token workflows regularly. Gartner suggests reviewing high-consumption workflows in sprint retrospectives. Route unexplained spend to security review, not only to finance.
For the broader model behind these steps, see the agent governance stack.
Conclusion
AI agent costs are rising for structural reasons that will not reverse: more capable models, more complex workflows and consumption-based pricing. But runaway spend is different. It comes from agent decisions nobody is governing.
That is why AI agent cost control belongs in the security programme. Model routing and budget caps manage the baseline. Governing what agents retrieve and do, at runtime, stops the spend that should never have happened, which is also the autonomy we described in the coding agent paradox.


