Chinese AI model providers have triggered an API price war, with DeepSeek, Kimi, GLM, and other labs pricing flagship models at 5–12 times below equivalent OpenAI and Anthropic tiers. On OpenRouter, a platform aggregating model access, token share from U.S. companies for Chinese models has hovered above 30% each week since February 2026, peaking at 46%—compared to an average of 11% over the previous 12 months. Independent benchmarks show that typical production workloads migrate to Chinese models with an 87% cost reduction while output quality drops by only ~4% on average tasks.
The cost gap is stark at the premium end. DeepSeek V4 Pro runs a standard workload for roughly $11 while GPT-5.6 Sol costs ~$788 on the same job—a 70x delta in practice. Claude Opus 5 ($5/$25 per million input/output tokens) compares to DeepSeek at $0.30/$1.50. Baidu's Ernie, Zhipu AI's GLM-5.2 (trained on fraction of compute, still rivals GPT on coding), and MiniMax M3 all undercut flagship pricing. This follows aggressive price cuts by ByteDance, Alibaba, and Tencent: Alibaba reduced Qwen pricing by as much as 97%, and Baidu made Ernie models free for business users.
Developers are reacting in production. Lindy AI moved 100% of traffic from Claude to DeepSeek in June and projects savings of millions per month. Z.ai's GLM-5.2 saw the fastest adoption of any model on Vercel in 2026, with daily token volume growing 27x and customer count growing 80x in its first full week post-launch. For architects, the strategic question shifts from 'can we afford frontier models?' to 'where do we accept the 4-18% quality gap and self-host vs. cloud-host to preserve data sovereignty?' Open weights for DeepSeek, GLM, and Kimi K3 (MIT/Apache licenses) make self-hosting legal and feasible, but Chinese-hosted APIs raise data-governance questions that require case-by-case legal audit.