China-US AI gap drops to a record low

In June the LiveBench performance gap fell to 6% despite a 23x investment gap — $285.9B from the US versus $12.4B from China.

Author: Michael Kokin ·

According to Bloomberg Intelligence, in June the performance gap between Chinese and American models on the LiveBench benchmark narrowed to 6%, down from 9% in May. A year ago, the gap held at 10–15%.

GLM-5.2, from Zhipu AI (a Chinese AI lab), climbed to first place in agentic coding task rankings, with Moonshot's Kimi close behind. Bloomberg Intelligence analyst Robert Lee says Zhipu's success was no fluke.

The money math

The same picture shows up from another angle. In April, the Stanford AI Index put the gap between the top models from each country at 2.7%. Claude Opus 4.6 scored 1503 on the Arena leaderboard; ByteDance's Dola-Seed-2.0 scored 1464. For context, in May 2023 the gap between the two countries on major benchmarks ran between 17.5 and 31.6 percentage points.

The most surprising part of the Stanford report is the money math. The US poured $285.9 billion into private AI development in 2025; China put in $12.4 billion. A 23x spending gap, with a three-percentage-point quality gap in the top models. The report does include one important caveat: some Chinese investment flows through state funds that don't appear in private investment statistics, so the real spending gap is probably smaller than advertised.

The market votes with its wallet

According to a CNBC investigation from July 7, Chinese models' share of enterprise traffic on OpenRouter has stayed above 30% every week since February 8, peaking at 46%. A year ago the average was 11%; at the start of 2025 it was just 4.5%.

DeepSeek already outpaces every American lab in traffic volume, Anthropic included. It comes down to price: DeepSeek V4 Flash runs $0.14 per million tokens versus $5 for GPT-5.5 (and LLM prices keep falling).

Airbnb officially confirmed it uses the Chinese Qwen model in production, and two House committees promptly launched national security risk reviews.

What comes next

The US and China are building two increasingly incompatible tech stacks, Boston Consulting Group predicts, and the window for running both ecosystems in parallel is closing.

Broadly speaking, America holds the expensive proprietary models for well-funded corporations, while China keeps releasing models open-source for everyone. But there's a crack in that pattern too. ASPI (Australian Strategic Policy Institute) analysts estimate some Chinese labs are about four months away from building a model on par with Anthropic Mythos — the company's most powerful — and at that capability level, open-sourcing stops being the obvious move.