// Generative Ai Tools

AI Agent Cost Control: Why Runaway Token Spend Is Now a Security Risk

September 9, 2026
min read

Introduction

Every token an AI agent spends traces back to a decision. What to read. Which tool to call. Whether to loop one more time. Yet most enterprises are trying to control agent costs by governing the bill, not the decisions that produce it.

At DAXA, we see runaway spend as the most visible symptom of an ungoverned agent. The same gaps that let an agent read data it should not, or take actions nobody authorised, are the gaps that let its costs climb unchecked. Govern the decisions, and the cost follows.

The research now points the same way. Gartner names ungoverned autonomy and bloated context windows as common drivers of token overspending, and Forcepoint's X-Labs research shows agents running up costs through attacks, misconfiguration and ordinary use, often without any single request looking unusual. AI agent cost control is no longer a budgeting exercise. It is a security control.

Why Are AI Agent Inference Costs Rising?

Tokens are getting cheaper, but agents use far more of them, on more expensive models. Gartner calls this the Inference Paradox.

Gartner predicts that AI inference costs per agentic workflow will rise more than fivefold through 2028. It identifies three trends behind this:

  1. Foundational model cost economics are improving quickly.
  2. That efficiency is unlocking more powerful, and more expensive, models for more sophisticated applications.
  3. Sophisticated AI workflows use far more tokens than simple chatbot interactions.

The result is that tokens are becoming more cost-efficient, but not as fast as the capabilities, and costs, of the workflows built on them.

The difference between a chatbot and an agent explains much of the gap. A chatbot reads a query and responds. An agent plans, calls tools, checks its own work and tries again. According to Gartner, routing a task to an agentic reasoning model increases provider inference costs by at least five times compared with a basic chatbot interaction, and often much more as complexity grows.

Pricing is shifting too. In separate research on AI coding costs, Gartner notes that coding agent vendors are moving from seat-based to consumption-based pricing, often with little transparency into how tokens are calculated and billed. It predicts that by 2028, AI coding costs will overtake the average developer's salary.

The effect is already visible. Uber reportedly exhausted its entire 2026 AI coding tools budget within four months as Claude Code spread across around 5,000 engineers. We covered the budgeting side in enterprise AI cost governance.

What Causes Runaway AI Agent Spend?

Structural trends raise the baseline. Runaway spend comes from ungoverned agent decisions: what an agent retrieves, what it does, how long it keeps going and whose identity it spends under.

Gartner is direct about this. It links token overspending to how usage is governed, naming ungoverned autonomy in agent-driven workflows, bloated context windows and the absence of structured feedback among the common failure modes. As analyst Nitish Tyagi puts it: "Token discipline will not emerge through developer choice alone."

This is the lens DAXA applies. Instead of asking how much an agent spent, ask which decision produced the spend, and whether it should have been allowed.
‍

Agent decision How it drives cost Evidence
What it retrieves Everything pulled into context is reprocessed on later turns, so cost grows with each step Gartner cites bloated context windows.
Forcepoint found per-message cost in a long support session rose roughly 100x by turn 100
What it follows Retrieved content steers the agent to fetch more, turning one task into a runaway chain Forcepoint describes pages seeded with hundreds of fake links that an agent follows during a legitimate task
What it does Tool calls and actions each carry cost, and ungoverned agents take steps nobody scoped Gartner cites ungoverned autonomy in agent-driven workflows
How long it keeps going Reasoning models can be pushed to re-check their own answers repeatedly, without a traffic spike Forcepoint's reasoning loop exhaustion scenario
Whose identity it spends under Leaked or unattributed keys let anyone spend on your account Forcepoint's example: a support chatbot's staging key leaked to a public code repository


This is also why cost is now a security issue. Forcepoint's research shows that in these scenarios, no single request looks malicious or breaks usage limits, even as the total climbs. The same gap produces accidental overspend and deliberate abuse, and from the invoice alone, the two look identical. OWASP now ranks this risk, unbounded consumption, sixth in its 2026 Top 10 for LLM Applications, alongside the others we mapped in the OWASP LLM and Agentic Top 10.

Budget caps cannot close this gap on their own. They act after the spend happens, cannot tell legitimate work from abuse, and halt every valid task sharing the same key.

How DAXA Controls AI Agent Spend at the Source?

DAXA governs the decisions that create spend: what agents retrieve and what they do, evaluated in context at runtime.

Traditional IAM answers who can access which system. It was not built to reason about what an autonomous agent does once inside. DAXA extends it into Identity and Action Management: existing access permissions are enforced first, then the platform reasons over whether an agent should retrieve, write, delete or share specific data in that context, with enforcement at AI runtime.

Applied to agent cost, that means:

  1. Governed retrieval. Pebblo classifies data by confidentiality, compliance and semantic context, and evaluates each retrieval against user context, document intent and application context. The agent pulls in what its task needs, not everything it can reach, which limits context growth and data exposure together.
  2. Governed actions. Context-aware enforcement stops unauthorised actions before they happen, regardless of which agent, application or model is behind the request.
  3. Behavioural reasoning. DAXA's guardian agents maintain awareness of business workflows and data relationships, so the platform can reason about subtle signs of compromise or misuse rather than matching patterns.
  4. Real-time visibility. A data bill of materials shows what AI agents and applications access, modify, generate and share, giving teams the attribution that cost anomalies need.

DAXA does not replace spend limits or inference controls. It covers the layer they cannot see.

‍

Control layer What it governs Where it stops short
Spend limits Total spend per user, key or team Acts after spend happens, cannot judge intent
Model and inference controls Model routing, loop counts, thinking tokens Cannot judge whether a retrieval or action should happen
Data and action governance (DAXA) What agents retrieve and do, at runtime Works alongside caps and inference controls, not instead of them

‍
Tool fan-out shows why the third layer matters. As we covered in indirect prompt injection in AI coding agents, untrusted content shaping an agent's next step is a governance problem. We apply the same model to homegrown AI agents, AI coding assistants and across Pebblo.

‍How to Control AI Agent Costs: 8 Steps for Enterprises

Start with the decisions, then cap the totals.

  1. Match the model to the task. Gartner recommends aligning model selection with task complexity: route simpler, high-frequency tasks to smaller models and reserve frontier models for complex, high-value work.

  2. Set autonomy levels by use case. Define when agents should be used and how much autonomy each task gets. Gartner suggests classifying work as developer-led, developer-with-agent or fully agent-led.

  3. Govern what agents retrieve. Gartner recommends context engineering: including only relevant information and removing unnecessary data. Enforcing this at runtime keeps context lean and stops agents reading data they have no reason to see.

  4. Govern what agents do. Check each action against the agent's task and identity before it runs. This separates legitimate heavy work from tool fan-out, which spend caps alone cannot do.

  5. Bound loops and apply least privilege. Forcepoint recommends capping agent steps and loop-backs, detecting repetitive loops early, and using sandboxing and least-privilege access to limit how far a runaway process can spread.

  6. Attribute every model call. Trace each call to a known identity and workload, and keep API keys out of code repositories. Without attribution, a leaked key and legitimate growth look the same.

  7. Set thresholds and hard limits. Gartner recommends token thresholds, escalation policies and automated monitoring. Forcepoint adds hard caps per user, key and team, rather than alerts that fire too late.

  8. Review high-token workflows regularly. Gartner suggests reviewing high-consumption workflows in sprint retrospectives. Route unexplained spend to security review, not only to finance.

For the broader model behind these steps, see the agent governance stack.

Conclusion

AI agent costs are rising for structural reasons that will not reverse: more capable models, more complex workflows and consumption-based pricing. But runaway spend is different. It comes from agent decisions nobody is governing.

That is why AI agent cost control belongs in the security programme. Model routing and budget caps manage the baseline. Governing what agents retrieve and do, at runtime, stops the spend that should never have happened, which is also the autonomy we described in the coding agent paradox.

from langchain.document_loaders.csv_loader import CSVLoader    
from langchain_community.document_loaders.pebblo import PebbloSafeLoader
‍

loader = PebbloSafeLoader(
          CSVLoader(file_path),
          name="acme-corp-rag-1", # App name (Mandatory)
          owner="Joe Smith", # Owner (Optional)
          description="Support RAG app",# Description(Optional)
)
‍
documents = loader.load()
vectordb = Chroma.from_documents(documents, OpenAIEmbeddings())
// FAQ’s

We’re here to answer your questions

What is AI agent cost control?

The controls that limit how much compute, token and financial resource AI agents consume, spanning model routing, spend limits and governance over what agents retrieve and do.

Why are AI agent inference costs rising?

Agents use far more tokens than chatbots and often run on more expensive reasoning models. Gartner predicts inference costs per agentic workflow will rise more than fivefold through 2028, even as per-token prices fall.

What causes runaway AI agent spend?

Ungoverned agent decisions: bloated context, uncontrolled autonomy, runaway loops, and leaked or unattributed API keys.

How can enterprises control AI agent costs?

Match models to task complexity, set autonomy levels, govern what agents retrieve and do at runtime, bound loops, attribute every model call, and use token thresholds and hard limits as a backstop.

Why is runaway AI agent spend a security risk?

The same control gap that causes accidental overspend can be triggered deliberately through leaked keys, planted content or crafted prompts, and individual requests often look normal.

Are budget caps enough to control AI agent spend?

No. Caps act after spending happens and cannot tell legitimate work from abuse. They work best as a backstop behind runtime governance.

Related Blogs