In a five-day window, three of China's top AI labs made moves that look contradictory on the surface. DeepSeek raised API prices by as much as 1,100% on certain tiers. Alibaba, on the same week, open-weighted a 2.4-trillion-parameter flagship model it had never released before. Zhipu AI shipped GLM-5.3, boosting coding benchmarks by roughly 6x using the exact same base model as its predecessor — no retraining involved. Together, these three moves signal that China's AI labs are shifting from competing on price alone to competing on pricing power itself. Pair this piece with our earlier Qwen3.8-Max launch note and DeepSeek V4 Flash review; those covered the product drops, this one covers the repricing.
00Timeline: What Happened, and When
| Date | Event |
|---|---|
| Jul 16, 2026 | Moonshot AI open-weights Kimi K3 (2.8T parameters), drawing US security scrutiny |
| Aug 2–3, 2026 | Alibaba previews, then launches, Qwen3.8-Max as a hosted API |
| Aug 10, 2026 | Meta releases Muse Glimmer (30B, Apache 2.0), teases open weights for flagship Muse Spark 1.2 |
| Aug 12, 2026 | Alibaba publishes Qwen3.8-2.4T-A95B open weights on Hugging Face/ModelScope; xAI ships Grok 4.6 |
| Aug 13, 2026 | DeepSeek-V4-Pro goes GA and announces a price increase effective Aug 17; Google ships discounted Gemini 3.7 Flash |
| Aug 14, 2026 | Zhipu ships GLM-5.3, reusing GLM-5.2's 743B base |
| Aug 17, 2026, 00:00 Beijing time | DeepSeek's new pricing takes effect |
Zoom out further and the picture gets more interesting: on Jul 30, OpenAI cut prices on its cheapest tier (GPT-5.6 Luna, down 80%), then on Aug 6–7 made Luna the free default with unlimited text chats. In other words, while Chinese labs were raising prices and opening up flagship weights, US labs were cutting prices and going free at the consumer layer — at the exact same time. That is not a coincidence; it is two sides of the same pricing fight. The Luna cut is covered in our GPT-5.6 price-cut brief.
PainWhere procurement teams get this week wrong
- Treating "1,100%" as the whole invoice. Outlets quoted 11x, 1,100%, and 350%. All three can be true. They describe different line items: peak cache-hit input, output, and cache-miss input.
- Assuming the official API is always cheapest. At peak hours, DeepSeek's own list price now sits above several resellers, including GMI Cloud and Novita.
- Reading "open weights" as Apache 2.0. Qwen3.8-Max ships under a custom Qwen3.8-Max License. Large MaaS and AI-assistant businesses still need a separate commercial deal.
- Repeating the geo-ban rumor. The published license has no US, EU, UK, or South Korea download ban. The constraints are revenue and attribution, not territory.
- Calling GLM-5.3 a new foundation model. It reuses the 743B GLM-5.2 base. The jump is post-training, and the scores are vendor-reported.
- Budgeting as if "Chinese model = cheapest model". Off-peak DeepSeek still undercuts Claude Opus 5, but Qwen's international API and OpenAI Luna now beat it on at least one dimension.
01The Numbers: What Actually Changed
DeepSeek's price hike, tier by tier (effective Aug 17, 00:00 Beijing time; peak hours are 9am–12pm and 2pm–6pm Beijing time). Prices are per 1M tokens.
| Billing item | Old price | New off-peak | New peak | Peak increase |
|---|---|---|---|---|
| V4-Flash cache hit (input) | ¥0.02 | ¥0.05 | ¥0.10 | ~400% |
| V4-Flash cache miss (input) | ¥1.0 | ¥1.5 | ¥3.0 | 200% |
| V4-Flash output | ¥2.0 | ¥4.5 | ¥9.0 | 350% |
| V4-Pro cache hit (input) | ¥0.025 | ¥0.15 | ¥0.30 | ~1,100% |
| V4-Pro cache miss (input) | ¥3.0 | ¥4.5 | ¥9.0 | 200% |
| V4-Pro output | ¥6.0 | ¥13.5 | ¥27.0 | 350% |
Qwen3.8-2.4T-A95B (Qwen3.8-Max open weights): key specs
| Spec | Detail |
|---|---|
| Parameters | 2.4T total, 95B active per token (MoE, 512 experts, 10 routed + 1 shared) |
| Context window | 262,144 tokens native (open checkpoint), extendable to ~1.01M; hosted Max version defaults to 1M |
| Release cadence | Preview Aug 2 → API live Aug 3 → open weights Aug 12 |
| API pricing (international) | $2/M input, $6/M output |
| License | Not Apache 2.0 — a custom Qwen3.8-Max License |
| Why it matters | First time Alibaba has open-weighted a Max-tier flagship; Qwen3.5/3.6/3.7 Max stayed API-only |
GLM-5.3 vs GLM-5.2: same base model, post-training only. These are Zhipu's own reported numbers — no independent third-party re-run has been published yet.
| Benchmark | GLM-5.2 | GLM-5.3 | Change |
|---|---|---|---|
| Terminal-Bench 3.0 | 4.6% | 28.3% | +23.7 pts |
| DeepSWE v1.1 | 46.2% | 66.9% | +20.7 pts |
| Agents' Last Exam (CLI) | 23.8% | 28.5% | +4.7 pts |
| CyberGym | 77.2% | 84.5% | +7.3 pts |
| AutomationBench | 26.2% | 48.2% | +22.0 pts |
02Breaking Down the Three Strategies
DeepSeek: from flat-rate pricing to time-of-day pricing — this is a capacity problem, not a strategy pivot. The easiest misread is "China's cheapest model finally caved to margin pressure." Look closer at the structure and it reads more like the opposite: a company making its compute constraints visible in the price sheet for the first time. DeepSeek's old flat, always-cheap pricing worked as a customer-acquisition tool as long as GPU capacity kept pace with demand. Once usage grew exponentially and capacity did not, something had to become explicit — and "encouraging more flexible workload scheduling" in the official announcement is corporate-speak for "our peak-hour compute is now scarce, please shift your load yourself."
One detail international coverage mostly missed: at peak hours, DeepSeek's own official API price is now higher than several third-party resellers (GMI Cloud, Novita, and others currently list V4 Pro below DeepSeek's new peak rate). The assumption that "the official API is always the cheapest way to run DeepSeek" — a core part of its reputation — has been broken for the first time.
Alibaba: open weights buy ecosystem goodwill; a custom license protects the revenue ceiling. Qwen3.8-Max's open-weighting is not a straightforward act of generosity. Alibaba did two things simultaneously: it published the full 2.4T-parameter checkpoint for free download, and it attached a custom license — not the permissive Apache 2.0 used for smaller Qwen models — that requires any "Model-as-a-Service" or "AI Work Assistant" business earning over $50 million in any 12-month period to negotiate a separate commercial license, and requires products with 100M+ monthly active users or $20M+ in monthly revenue to prominently display the model's name.
The logic: give away the weights to win developer mindshare (especially internationally, where "made-in-China model" still carries some hesitation among enterprise buyers), while keeping pricing leverage over the handful of companies actually capable of building a competing inference business on top of it. That is a materially different bet than Meta's Muse Glimmer, which ships under unrestricted Apache 2.0 — "open weights" does not mean the same thing across these two releases. One rumor worth killing explicitly: claims circulated online that Alibaba's license bans downloads from the US, EU, UK, and South Korea. That is false. The published license text contains no geographic or territorial clause of any kind.
GLM-5.3: no new base model, just a bigger post-training bet — and that is the real story. The most interesting fact about GLM-5.3 is not the score, it is the method: same 743B-parameter base as GLM-5.2, no retraining, and a roughly 6x jump on Terminal-Bench 3.0 (4.6% → 28.3%) purely from scaling up reinforcement learning environments in post-training. This confirms a trend that has been building industry-wide for months — as pretraining scaling laws show diminishing returns, post-training RL scale is becoming an independent performance lever with a much lower cost floor than retraining a new foundation model. That is a meaningfully lower barrier to entry, and it is why mid-tier labs without OpenAI-scale compute budgets can still close the gap on agentic and coding benchmarks.
03Head-to-Head: Is DeepSeek Still the Cheapest Frontier-Class Model?
| Model | Input (per 1M tokens) | Output (per 1M tokens) | Open weights? |
|---|---|---|---|
| DeepSeek V4-Pro (peak) | ¥9.0 (~$1.26) | ¥27.0 (~$3.78) | No |
| DeepSeek V4-Pro (off-peak) | ¥4.5 (~$0.63) | ¥13.5 (~$1.89) | No |
| Qwen3.8-Max (international API) | $2.00 | $6.00 | Yes (custom license) |
| OpenAI GPT-5.6 Luna | $0.20 | $1.20 | No |
| Claude Opus 5 (implied, per Alibaba's own comparison ratio) | ~$5.00 | ~$25.00 | No |
RMB-to-USD conversion at ~¥7.15/$1, approximate. The short answer: no. Even after accounting for the hike, DeepSeek V4-Pro's off-peak rate is still well below Claude Opus 5, but it is no longer the outright cheapest option on the table — both Qwen3.8-Max's international pricing and OpenAI's Luna now undercut DeepSeek's off-peak rate. "Chinese model = cheapest model" was true for most of 2025 and early 2026; it is not a safe assumption anymore.
04What's Disputed or Unverified
- The "1,100%" headline is technically accurate but misleading without context. It applies only to peak-hour cache-hit input pricing, the tier that started nearest to zero. Output pricing — the cost that dominates most real bills — rose 350%.
- Claims that Qwen3.8-Max runs on Alibaba's in-house Zhenwu M890 chips (reported by several Chinese financial outlets as evidence of a fully domestic-silicon inference stack) have not been independently confirmed by Alibaba's own technical documentation or third-party benchmarks. Treat this as vendor-adjacent, unverified reporting until confirmed.
- GLM-5.3's reported discovery of a "serious vulnerability" in Cursor comes from VentureBeat's reporting and Zhipu's own disclosure; specific technical details of the vulnerability have not been made public, so the claim should be read as a vendor-sourced, not independently audited, security finding.
- Reports that China's Ministry of Commerce may be preparing retaliatory export controls on AI/semiconductor technology are speculative and sourced to unconfirmed media reports, not an official announcement. Treat as background context, not established fact.
05Why This Matters: Two Price Wars Running in Parallel
Place this in the bigger frame and a pattern emerges. Over roughly the past month, China's top labs have shipped major releases at a pace domestic financial media has started calling "three model updates a week" — DeepSeek, Alibaba, and Zhipu, plus Moonshot's Kimi K3 (open-weighted Jul 16, 2.8T parameters) and MiniMax H3 before them. Chinese coverage broadly frames this as Chinese open-weight releases "forcing a global repricing of the AI industry" — a framing that is more assertive than most English-language coverage of the same events.
Meanwhile, US labs are running the opposite play at the consumer layer: OpenAI cut prices 80% on its cheapest tier (Jul 30) then made that model free and unlimited for all users a week later (Aug 6–7); Google shipped a coding-focused model at half the price of its three-week-old predecessor (Aug 13). So while Chinese labs open-weight flagships and introduce tiered, higher pricing on the compute-constrained top end, US labs are racing toward free and cheap at the consumer end. Both are real strategies; they are just optimizing for different parts of the funnel.
There is also a geopolitical layer worth naming carefully. Moonshot's Kimi K3 open-weighting in July already drew US security scrutiny; Alibaba choosing this specific window to open-weight a 2.4T flagship has been read by some analysts as a move to lock in international mindshare and a "technological parity" narrative before any potential regulatory tightening. That is an informed interpretation, not a confirmed fact — but it is part of the context that is hard to see if you are only reading English-language tech press, which has largely covered these releases as isolated product news rather than as a coordinated national pattern.
06Six steps: reprice the stack before the next invoice
Turn the week's headlines into an internal runbook. Do not let a single percentage rewrite the budget.
-
01
Split traffic by Beijing peak windows. Bucket existing calls into 09:00–12:00 and 14:00–18:00 Beijing time versus everything else. US daytime work mostly lands off-peak; China weekday daytime lands on-peak.
-
02
Measure cache-hit rate as its own KPI. The 1,100% headline sits on cache-hit input, which is still cheap in absolute yuan. Output (+350%) and cache-miss input (+200%) move the invoice. Heavy-usage models around 84M tokens/month, mostly off-peak, half hits, land nearer 1.8x.
-
03
Put official, reseller, Qwen, and Luna on one sheet. At peak, compare DeepSeek official, GMI Cloud / Novita, Qwen international ($2/$6), and GPT-5.6 Luna ($0.20/$1.20). Stop assuming official is cheapest.
-
04
Read the LICENSE before you download. Personal and internal use are largely unaffected. MaaS or AI Work Assistant revenue above $50M in any 12 months needs a separate Alibaba license. 100M MAU or $20M monthly revenue requires prominent model-name display. There is no geographic ban.
-
05
Treat GLM-5.3 as a post-training upgrade. Same 743B base. Vendor-reported scores. Terminal-Bench 3.0 still trails Sol at 34.6% and Fable 5 at 33.7%. Wait for an independent re-run before writing "open-weight #1" into a purchase memo.
-
06
Host self-serve evals on a dedicated plane. Pulling a 2.4T checkpoint, shifting batch jobs off-peak, and reproducing GLM post-training benches all need stable processes and a tenant boundary. Price dedicated Apple Silicon / cloud Mac nodes on the pricing page, then trial offline inference from the order page and decide on interrupt rate, not marketing copy.
[ ] Traffic split by Beijing peak / off-peak
[ ] Cache-hit rate and output tokens billed separately
[ ] Official peak vs reseller vs Qwen vs Luna
[ ] Qwen3.8-Max License revenue triggers read
[ ] GLM-5.3: vendor scores, wait for third party
[ ] Self-host plane isolated from production secrets
07Close and FAQ
DeepSeek wrote a capacity shortage into the price sheet. Alibaba traded flagship weights for mindshare and kept a custom license over the cash-generating layer. Zhipu showed that post-training scale can move coding benches without a new base. Shared-minute pools, oversold VPS instances, and a Mac under someone's desk still fail the same way: jittery bandwidth, noisy neighbors, and dropped long sessions wipe out off-peak batching and local eval audit trails. For a more stable production and evaluation plane, NUKCLOUD multi-region bare-metal Mac / cloud Mac nodes give dedicated Apple Silicon and a clear tenant boundary. Compare SKUs on the pricing page and start a trial from the order page.
Sources: DeepSeek's official pricing announcement, cross-checked against Wall Street CN, IT Home, AIGC.cn, and V2EX community discussion; Alibaba's official Qwen model repositories (Hugging Face / ModelScope) and South China Morning Post's reporting on license terms; Zhipu (Z.ai)'s official GLM-5.3 technical page, plus VentureBeat and StableLearn coverage; Meta AI Research's official blog and VentureBeat's coverage of Muse Glimmer; Chinese financial outlets (Yicai, Sohu Finance) on the pacing and framing of China's open-weight release cycle. Pricing, license terms, and benchmark figures reflect publicly available information as of publication. Verify the latest official pricing and license terms before republishing, and note that details flagged above as unverified (domestic chip claims, the Cursor vulnerability report, and export-control rumors) have not been independently confirmed.