The next serious fight over AI budgets will look dull on the surface. It starts with a line item for token consumption, not a debate over what intelligence means.
That shift is useful. For two years many executives treated tokens as a technical detail buried in a vendor bill or smoothed over by a flat license. Inference is where AI becomes an operating cost. Training creates capability; inference consumes it. Each prompt, tool call, retrieved document, planning loop, and agent retry turns strategy into spend.
The infrastructure debate therefore matters to operators. McKinsey's June 2026 analysis shows cost per token and energy per token becoming central measures of AI economics. Goldman Sachs Research projects agentic AI will drive a 24-fold rise in token consumption by 2030. Goldman Sachs Asset Management notes that enterprise token use is already diverging between heavy adopters and median firms. The National Institute of Standards and Technology (NIST) and the Center for AI Standards and Innovation are advancing agent standards and identity work because agents need access to real systems before they produce real value.
Management sits where these signals meet. Working agents raise token volume. Rising volume makes cost and energy discipline material. Agents that touch real systems make identity and authorization material. Engineering can no longer own the budget alone.
Token budgets are where AI ambition meets operating discipline.
The Old Budget Was Easy
Old AI budgets were mostly procurement questions: the model, the copilot, the seat count, the vendor contract. Spend looked like software, so it was governed like software.
That approach held while usage stayed light, experimental, or limited by human prompting. A team could buy a tool, run pilots, and accept a loose link between activity and value. The main question was whether people used it.
Agentic systems change the cost structure. A human requests an outcome, and behind that request the system may plan, call tools, search files, write drafts, check its work, query another model, and retry a failed path. Users see one answer. Budgets see a chain of inference events.
That gap creates managerial risk. An employee requested a task. Systems consumed budget. Vendors record usage. Finance receives a bill. Security sees a new actor touching systems. The business still has to decide whether the work was worth doing.
The Unit Has Changed
Seat pricing taught managers to think in users. Token pricing forces them to think in work.
Users can be cheap and wasteful or expensive and valuable. Tokens might buy a hallucinated summary, a useful code review, a low-value slide rewrite, fraud triage, or an agent loop that should have stopped five calls earlier. The unit that matters is the priced unit of judgment, coordination, generation, or action that the token makes possible.
Blanket AI rationing is therefore crude. Forcing every workflow onto a small model can erase the economic value that justified AI in the first place. Running frontier inference on work a smaller model, script, template, or human decision would handle wastes money the other way.
The goal is tokens that buy the right kind of advantage.
The Budget Map
Executives need a map before they need a dashboard. A useful token budget has five lines.
| Budget item | Question | Failure mode |
|---|---|---|
| Model tier | Which tasks require frontier inference, and which can run on cheaper models or deterministic tools? | Using premium intelligence as a default setting. |
| Context budget | How much retrieved or pasted context is worth paying for before the marginal document adds noise? | Long prompts that feel thorough but inflate cost and error surface. |
| Agent loop | How many planning, tool-use, retry, and verification steps are allowed before escalation? | Autonomy that spends silently because nobody set a stop rule. |
| Value class | Is this inference buying revenue, risk reduction, cycle-time compression, learning speed, or convenience? | Counting every AI-assisted artifact as equally valuable. |
| Owner | Who owns the bill, the workflow outcome, the identity permissions, and the rollback trigger? | A shared budget with no responsible operator. |
This map turns token spend into a management artifact. Model tier separates expensive cognition from routine work. Context budget stops the organization from shoveling documents into prompts without relevance discipline. Agent loop makes autonomy accountable. Value class keeps convenience from counting as strategy. Owner keeps compute from becoming an orphan cost.
The Cheap Token Trap
Falling token costs raise the need for budgeting. They do not remove it.
McKinsey identifies model optimization, advanced packaging, custom silicon, and co-packaged optics as major levers for reducing inference cost over time. If those improvements arrive, more workflows become economically plausible. The familiar trap remains: when the unit gets cheaper, usage expands faster than discipline.
Cheap tokens still get wasted. Agent loops still touch the wrong system. Summaries still mislead decisions. Retrieval still exposes sensitive context. Workflows still dump review burden on humans who must inspect more machine output.
The lesson is discipline, not austerity. Cost decline should widen the test set. Judgment stays. As tokens get cheaper, the question shifts from whether a workflow can run to whether it should consume intelligence automatically.
The Security Line
The token budget also needs a security line. Agents act; they do more than answer.
The National Institute of Standards and Technology (NIST) AI Agent Standards Initiative frames adoption around trust, interoperability, security, and identity. Its Request for Information summary says respondents widely agreed that AI agents create novel security threats and that those concerns block adoption. That belongs in the budget discussion, not only in a security appendix.
Once an agent has permissions, spending tokens can mean spending authority. The agent may read a customer record, update a Customer Relationship Management field, trigger a support workflow, call an internal API, or query a privileged database. Each step has an inference cost and an authorization cost. Optimize both together.
For serious workflows, pair the token budget with identity controls: named human owner, scoped permissions, just-enough access, audit trail, step limit, exception path, and rollback rule. Otherwise the company is giving a probabilistic worker a company card.
The Founder Version
A founder needs no enterprise FinOps team to start. Three rules are enough.
- Put every AI workflow into a value class before scaling it: revenue, delivery, risk, learning, or convenience.
- Set a maximum loop depth for agentic work: number of tool calls, retries, documents, or minutes before human review.
- Review one weekly token spend review: top workflows by spend, output produced, human review burden, and stop-rule exceptions.
Those rules block two founder mistakes: starving a valuable workflow because the bill feels new, and subsidizing an impressive workflow because the output feels magical. The test is simple. If a workflow cannot name the business metric that token spend improves, it stays experimental.
The CFO Question
A single AI cost number hides the decision. The CFO should refuse that summary.
Separate token spend into strategic work, operating throughput, employee convenience, and waste. Coding assistants that reduce review bottlenecks may deserve more expensive inference than presentation helpers. Fraud agents may justify heavy context retrieval if they reduce loss or escalation. Meeting-summary tools often need a hard ceiling because the upside is modest and the volume is large.
The CFO's job is to price autonomy correctly. Suppressing consumption is the wrong goal. The budget should reward high-value inference and force low-value workflows to simplify, downshift, cache, template, or stop.
The Executive Test
Before approving a large AI workflow, require six answers.
- Name the default model tier and the event that allows escalation.
- State which context may enter the prompt or retrieval path.
- Set the maximum agent steps before a human owns the decision.
- Name the business value class that justifies the token spend.
- Name the human who owns the bill, permissions, and rollback.
- Define the weekly report that would expand, downshift, or kill the workflow.
These questions make AI spending legible and strategy less theatrical. A company can approve more AI while still refusing unmanaged inference.
The Better Argument
The market argument will keep swinging between abundance and scarcity. One side will point to falling cost per token. The other will point to infrastructure strain, energy limits, chip shortages, and rising agent usage. Both can be right.
The operator's argument is narrower. AI will be cheap enough to enter many workflows and expensive enough to punish lazy design. That combination rewards companies that allocate tokens like capital: deliberately, observably, and with a theory of return.
Winners in the next phase will ask where intelligence should be spent, not only whether AI works.
Source Notes
- McKinsey, "Frontiers of compute: The technologies to reduce AI inference costs"
- Goldman Sachs, "AI Agents Forecast to Boost Tech Cash Flow as Usage Soars"
- Goldman Sachs, "AI Investment Is Shifting as Inference, Enterprise Adoption Accelerate"
- National Institute of Standards and Technology (NIST), "Announcing the AI Agent Standards Initiative"
- National Institute of Standards and Technology (NIST), "Summary Analysis of Responses to the Request for Information Regarding Security Considerations for AI Agents"