China AI Bench

China's AI models, tested by hand.

We run the models, publish the logs, and explain what the scores actually mean.

Best Chinese AI Models: Ranked & Sourced (2026)

2 formally tested in-house with raw logs, 5 smoke-tested (N=1), 3 compiled rows with every number sourced, 4 international reference — 6 tracked without scores. Last updated: 2026-08-08.

See the ranking →

Deep dives

Updated 2026-08-08
ReviewBenchmarks Desk10 min read

China vs US frontier: 9 models, 11 tasks, same battery

We compiled the same 11-task battery on 9 frontier models: GLM-5.2, Seed 2.1, Step 3.5 Flash, Hunyuan A13B, Ling 2.0 vs GPT-5.6 Luna Pro, Claude Opus 5, Gemini 3.6 Flash, Grok 4.5. Six scored 11/11; the only US-side failure was a quota issue, not capability. Costs spanned ~100x, from $0.0004 to $0.04. Raw logs public.

CommentaryField Notes3 min read

NIST CAISI vs DeepSeek V4: the 8-month gap

The US government's independent eval shop ran DeepSeek V4 Pro on closed benchmarks. Result: still the most capable Chinese model CAISI has tested, about 8 months behind the US frontier — and ahead on capability per dollar. What the gap actually means.