The AI Hub · 2026-03-17 · AI Boutique Team · Tokens

Token use is through the roof. Your bill knows.

Prices per token keep collapsing. Total spend keeps climbing. Both things are true, and the gap between them is agents.

THE ROOF13× SINCE JAN 2025

Gartner’s analysis this March found enterprise token consumption has grown roughly thirteen-fold since January 2025. Goldman Sachs Research expects agentic AI to multiply it another 24 times by 2030 — to about 120 quadrillion tokens a month. If you budgeted your AI spend from early chatbot usage patterns, your projection isn’t slightly off. It’s a different order of magnitude.

checked 2026-07-06 — Gartner 13× figure; Goldman Sachs 24× / 120 quadrillion projection

The reason is structural. A chat answer costs a few thousand tokens. An agent that plans, retrieves, calls tools, retries and checks its own work consumes five to thirty times more per task. Point one at a real workflow and a single job can burn through what a whole team’s chat usage cost last year.

checked 2026-07-06 — 5–30× agent token multiplier (industry range estimate)

The paradox that matters

Token prices have fallen by orders of magnitude in two years. Enterprise AI spend rose over 300% in the same period. Cheaper units. Bigger bills.

checked 2026-07-06

Jevons, again

Economists have a name for this: when something useful gets cheaper, you don’t spend less on it — you use vastly more of it. It happened with coal, with electricity, with cloud storage. It’s happening with intelligence. Deloitte reported in January that AI is now the fastest-growing expense in corporate technology budgets, in some firms approaching half of total IT spend. Output tokens still cost four to five times input tokens, so verbose, multi-step reasoning workloads are exactly the expensive kind — and exactly the kind everyone is adopting.

checked 2026-07-06 — Deloitte January report; “half of total IT spend” qualifier is “in some firms”; 4–5× output/input ratio is approximate industry figure

None of this is an argument for using less AI. Companies making heavy, productive use of tokens are growing faster than those that aren’t. It’s an argument for knowing what a token earns you.

The question to ask

Not “how much are we spending on tokens?” but “what did each workload return per token?” The first question produces a cap. The second produces a portfolio — keep the workloads that pay, fix the ones that nearly do, kill the rest. That’s the arithmetic our Reality Check runs, and token bills are where we usually start.

Book an AI Reality Check

Ready to make the AI you've already paid for pay you back?