China’s models grew up. Your AI bill should notice.
GLM-5 and DeepSeek V4 now sit within touching distance of the Western frontier at a fraction of the token price. The twist of 2026: the cheap models are raising prices — and it still barely matters.
Two releases define the year so far. In February, Zhipu (Z.ai) shipped GLM-5 — a 744-billion-parameter open-weight model, trained entirely on Chinese silicon, that benchmarks alongside Anthropic’s frontier coding models. In April, DeepSeek previewed V4: a 1.6-trillion-parameter system with a 1-million-token context window as standard, weights on Hugging Face under MIT, and the official release landing this month. Both are serious production models, not curiosities. And both are reshaping what a token should cost.
The twist: cheap is putting its prices up
The story used to be simple — Chinese labs racing each other to zero. That race is over. Zhipu has raised prices twice in five months and its stock jumped on both announcements; investors read the hikes as proof the demand is real. DeepSeek’s peak-hour doubling is the same signal in a different shape: when a lab starts charging by the clock, inference demand — not training — has become its binding constraint. The era of flat, infinitely scalable, loss-leader tokens is closing.
And yet the gap barely moved. Even after two increases, GLM-5.2 lists at roughly a sixth of comparable US output pricing; DeepSeek V4 Pro, even at its doubled peak rate, stays far below any Western frontier list price. A Deutsche Bank analysis this June put V4 at roughly 1.5% of the cost of Anthropic’s flagship for the large majority of everyday tasks. Cheap got slightly less cheap. It did not stop being cheap.
Where the gap actually lives
Price is only half the ledger; the other half is what a model gets right first time. On SWE-bench Verified — the closest thing coding has to a common exam — DeepSeek’s best V4 configuration resolves 80.6% of tasks, against roughly 88–95% for the Western frontier. Eight to fifteen points doesn’t sound like much until you translate it: on a hundred real tickets, that’s eight to fifteen extra failures to catch, re-prompt or hand to a human. The dollar arithmetic still favours the cheap model — even three retries of V4 cost pennies against one frontier pass — but retries are paid in waiting, and a senior engineer’s afternoon is the most expensive token there is. That’s why the winning pattern in every stack we’ve audited this year is a mix, not a swap.
A day in the life of a surge-priced token
DeepSeek’s peak windows follow the Beijing working day — which is quietly good news for everyone west of it. For a London team, the 2× hours land between one and four in the morning and before mid-morning; for the US East Coast they sit overnight. Ordinary Western business-hours traffic mostly rides the cheap side of the clock, and scheduled work — evals, batch classification, data generation — can be pinned there deliberately. Timezone arbitrage is now a line in the runbook.
Six months, four price moves
Put the year’s announcements on one line and the direction is unmistakable: the two labs that taught the world to expect ever-cheaper tokens both spent 2026 repricing upwards — in opposite styles. Zhipu raised the sticker; DeepSeek kept the sticker and charged for the busy hours.
Coding Plan +30% next day — Zhipu’s first-ever rise
API +8–17% — second rise in eight weeks
DeepSeek makes its 75% cut permanent, 1M context standard
Surge pricing begins: 2× at Beijing peak, legacy endpoints retire
Whether any of this belongs in your stack is a measurement question — the kind a two-week Reality Check answers with your workloads, your data rules and your numbers, not a benchmark chart.
Route the boring 90% to cheap tokens. Classification, extraction, summarisation, high-volume agent loops — workloads with an existing quality bar are exactly where a 30–100× price gap turns into real money. Keep the frontier models for the hard 10% where the last benchmark points earn their premium.
The clock is now a cost lever. Surge pricing means the same call costs double at 10am Beijing and half that at 8pm. Nightly batch runs, evals and data generation belong off-peak. Expect Western providers to copy this within the year.
Cheap models can cost you time instead of money. The independently measured coding gap is real (V4 trails the frontier by 8+ points on SWE-bench Verified). A model that retries its way to an answer burns the hours you thought you’d saved. Benchmark on your own tasks before you switch.
Check your governance before your wallet. Routing customer data through a Chinese API is a data-protection and compliance question, not just a procurement one. The MIT-licensed weights are the clean answer — self-host or use a hosted instance in your jurisdiction — but that’s datacenter-class hardware, not a laptop.
Use the gap as leverage even if you never switch. The credible threat of a 6–100× cheaper alternative is now part of every enterprise AI negotiation. Your vendors know these numbers. You should too.