Benchmarks Desk
Hands-on tests, scores and head-to-heads — the desk behind Best Chinese AI Models.
ReviewComparison8 articlesMonthly full retest; 48h spot test after major releases.
About this desk
Benchmarks Desk is where we run and publish hands-on tests: task sets written in advance, fixed sampling parameters, versions snapshotted, raw logs kept. Vendor tables get cross-checked against independent evals where they exist. Nothing here is scraped from a leaderboard.
Featured
All articles
- China vs US frontier: 9 models, 11 tasks, same batteryAug 8, 2026 · Review
- V4 Flash on ARC-AGI: 89% for $0.02 a taskAug 8, 2026 · Commentary
- China vs US AI models, mid-2026: a scorecardAug 6, 2026 · Commentary
- GLM-5.2: the open-source SOTA, warts and allAug 6, 2026 · Review
- Kimi K3: 180K downloads, 340 fine-tunes, 72 hoursAug 6, 2026 · Review
- MiniMax H3: video generation at 0.8 yuan a secondAug 6, 2026 · Review
- Qwen 3.8: Alibaba's 2.4T-parameter launch, one week inAug 6, 2026 · Review
- DeepSeek V4 review 2026: benchmarks, pricing, my test runAug 3, 2026 · Review
Subscribe to this column: RSS feed