The AI Hub · 2026-07-06 · AI Boutique Team · Models · China

China’s models grew up. Your AI bill should notice.

GLM-5 and DeepSeek V4 now sit within touching distance of the Western frontier at a fraction of the token price. The twist of 2026: the cheap models are raising prices — and it still barely matters.

Two releases define the year so far. In February, Zhipu (Z.ai) shipped GLM-5 — a 744-billion-parameter open-weight model, trained entirely on Chinese silicon, that benchmarks alongside Anthropic’s frontier coding models. In April, DeepSeek previewed V4: a 1.6-trillion-parameter system with a 1-million-token context window as standard, weights on Hugging Face under MIT, and the official release landing this month. Both are serious production models, not curiosities. And both are reshaping what a token should cost.

What a million output tokens costs — July 2026 list pricesUSD · bar length = price
DeepSeek V4 Flash
$0.28
DeepSeek V4 Pro
$0.87
GLM-5-Turbo
$4.00
GLM-5.2 flagship
$4.40
US frontier (list)
~$25
Official API list prices, checked 2026-07-06 · US frontier = Anthropic Opus-class list ($25/M output, per SCMP) · DeepSeek peak hours bill at 2× from mid-July
DeepSeek V4
Official release this monthThe April “preview” graduates mid-July; legacy endpoints retire July 24.
1M-token context standardAcross the lineup, at no context premium. Pro: 1.6T parameters (49B active). Flash: 284B (13B active).
First surge pricing on a frontier API2× list during Beijing business hours (9–12, 14–18), off-peak unchanged.
Open weights, MIT licencePlus a free web chat and 5M free API tokens per new account.
GLM-5 → 5.2 (Zhipu / Z.ai)
Frontier-adjacent codingGLM-5.1 places third globally on a composite of major coding benchmarks and runs autonomously for 8-hour stretches.
Trained on Huawei siliconNo Nvidia in the stack. Open weights, MIT, on Hugging Face.
Coding Plan from ~$10/monthFlat-rate tiers (~$10/$30/$80) that plug into Claude Code and Cursor as a drop-in backend.
Two price rises this year+30% in February, +8–17% in April. Demand is outrunning their compute too.

The twist: cheap is putting its prices up

The story used to be simple — Chinese labs racing each other to zero. That race is over. Zhipu has raised prices twice in five months and its stock jumped on both announcements; investors read the hikes as proof the demand is real. DeepSeek’s peak-hour doubling is the same signal in a different shape: when a lab starts charging by the clock, inference demand — not training — has become its binding constraint. The era of flat, infinitely scalable, loss-leader tokens is closing.

And yet the gap barely moved. Even after two increases, GLM-5.2 lists at roughly a sixth of comparable US output pricing; DeepSeek V4 Pro, even at its doubled peak rate, stays far below any Western frontier list price. A Deutsche Bank analysis this June put V4 at roughly 1.5% of the cost of Anthropic’s flagship for the large majority of everyday tasks. Cheap got slightly less cheap. It did not stop being cheap.

Where the gap actually lives

Price is only half the ledger; the other half is what a model gets right first time. On SWE-bench Verified — the closest thing coding has to a common exam — DeepSeek’s best V4 configuration resolves 80.6% of tasks, against roughly 88–95% for the Western frontier. Eight to fifteen points doesn’t sound like much until you translate it: on a hundred real tickets, that’s eight to fifteen extra failures to catch, re-prompt or hand to a human. The dollar arithmetic still favours the cheap model — even three retries of V4 cost pennies against one frontier pass — but retries are paid in waiting, and a senior engineer’s afternoon is the most expensive token there is. That’s why the winning pattern in every stack we’ve audited this year is a mix, not a swap.

The quality gap — SWE-bench Verified, % of tasks resolved
DeepSeek V4 Pro-Max
80.6%
US frontier, low end
~88%
US frontier, high end
~95%
Independent evaluations, mid-2026 · checked 2026-07-06 · the gap = tasks you finish yourself

A day in the life of a surge-priced token

DeepSeek’s peak windows follow the Beijing working day — which is quietly good news for everyone west of it. For a London team, the 2× hours land between one and four in the morning and before mid-morning; for the US East Coast they sit overnight. Ordinary Western business-hours traffic mostly rides the cheap side of the clock, and scheduled work — evals, batch classification, data generation — can be pinned there deliberately. Timezone arbitrage is now a line in the runbook.

DeepSeek V4 pricing by hour — Beijing time (UTC+8)
000609–12 peak14–18 peak23
2× list price1× list price· Beijing 9–12 = London 1–4am · New York overnight
Checked 2026-07-06

Six months, four price moves

Put the year’s announcements on one line and the direction is unmistakable: the two labs that taught the world to expect ever-cheaper tokens both spent 2026 repricing upwards — in opposite styles. Zhipu raised the sticker; DeepSeek kept the sticker and charged for the busy hours.

The 2026 repricing timeline
Feb 11
GLM-5 ships

Coding Plan +30% next day — Zhipu’s first-ever rise

Apr 8
GLM-5.1

API +8–17% — second rise in eight weeks

Apr 24
V4 preview

DeepSeek makes its 75% cut permanent, 1M context standard

Mid-Jul
V4 official

Surge pricing begins: 2× at Beijing peak, legacy endpoints retire

Checked 2026-07-06

Whether any of this belongs in your stack is a measurement question — the kind a two-week Reality Check answers with your workloads, your data rules and your numbers, not a benchmark chart.

Key conclusions
What it means for your pricing — bluntly
01

Route the boring 90% to cheap tokens. Classification, extraction, summarisation, high-volume agent loops — workloads with an existing quality bar are exactly where a 30–100× price gap turns into real money. Keep the frontier models for the hard 10% where the last benchmark points earn their premium.

02

The clock is now a cost lever. Surge pricing means the same call costs double at 10am Beijing and half that at 8pm. Nightly batch runs, evals and data generation belong off-peak. Expect Western providers to copy this within the year.

03

Cheap models can cost you time instead of money. The independently measured coding gap is real (V4 trails the frontier by 8+ points on SWE-bench Verified). A model that retries its way to an answer burns the hours you thought you’d saved. Benchmark on your own tasks before you switch.

04

Check your governance before your wallet. Routing customer data through a Chinese API is a data-protection and compliance question, not just a procurement one. The MIT-licensed weights are the clean answer — self-host or use a hosted instance in your jurisdiction — but that’s datacenter-class hardware, not a laptop.

05

Use the gap as leverage even if you never switch. The credible threat of a 6–100× cheaper alternative is now part of every enterprise AI negotiation. Your vendors know these numbers. You should too.

Book an AI Reality Check

Ready to make the AI you've already paid for pay you back?