We built this because our own AI bill needed adult supervision. It now produces the fortnightly numbers behind everything we say about making AI pay.
Token-Spend Toolkit
Know what your AI actually costs — before the invoice tells you.
The Problem
Most businesses adopting AI have no idea what they spend on it. Usage is scattered across providers, models and teams; pricing changes without ceremony; and the first sign of trouble is usually a bill. The standard "solution" — a spreadsheet someone updates when they remember — fails precisely when spend spikes, which is the moment it matters.
- —No single view: Claude here, ChatGPT there, different export formats, different currencies.
- —No early warning: an expensive week looks identical to a normal one until month-end.
- —No accountability: nobody can say which model, team or workload drove the increase.
What We Built
A governed, tested pipeline — not a dashboard bolted onto hope. Seven tools that run in sequence or on their own:
- —Ingest — normalises Claude and OpenAI usage into one canonical store. Duplicate records with different content refuse to load silently; re-ingested data follows a latest-wins restatement rule; an unknown model raises an error rather than costing £0.
- —Forecast & budget tracker — per-client month-to-date spend, projection to month-end, budget percentage and a pace flag that says "you'll blow this budget on the 22nd" while there's time to act.
- —Spend explorer + Command Centre dashboard — a self-contained HTML dashboard: warnings, recommendations, provider split, forecast back-testing, run-context banner.
- —Anomaly detector — daily z-score plus an absolute floor, so it catches both "weird for you" and "expensive by any standard".
- —Automated report — TL;DR first, then drivers, anomalies and costed recommendations, formatted properly ($1,202.50 (£949.98)), with provenance on every figure.
- —Actions layer — turns report signals into a ranked advisory list: what to change, what it saves.
Runs weekly, unattended, every Monday morning — the report is on the desk before the coffee is.
The Discipline
This wasn't vibe-coded. It was delivered under a governed milestone plan (M0–M7), each milestone independently signed off against evidence, then put through a formal gap analysis whose findings were fixed and re-gated — including a circular cost-reconciliation defect and a silent data-drop path that most tools would never have looked for.
- —204 automated tests across the whole suite (2 skipped — they require live provider admin keys).
- —Board-signed v1 — 14 June 2026.
- —Honest limits, stated: live provider-API auto-pull is implemented but not yet verified against a real endpoint, so we don't claim it. Data arrives via provider exports.
Delivery timeline
governance skeleton, canonical schema, Claude ingest with validation gates.
budget tracker, dashboard aggregation core, z-score + floor detection.
automated narrative report, OpenAI adapter reconciling a mixed store.
gap-analysis fixes (R1–R8), integration suite, Actions layer, board sign-off.
What it proves
If we can meter, forecast and challenge our own AI spend to board-report standard, we can do it for yours. This is the working machinery behind our "make AI pay" positioning — not a slide about it.
We build cost-control pipelines like this for clients — or run yours as a service.
Book a Reality Check →