Benchmarks Desk
Testing & rankings desk · China AI Bench
The Benchmarks Desk is the collective byline behind the China AI Bench leaderboard (Ranked & Sourced) and the model deep-dives and rankings on this site. It is an organizational byline, not a single person: the desk runs the in-house test harness, keeps the raw logs, and compiles the sourced numbers.
Two models are tested in-house so far, DeepSeek V4 Pro and V4 Flash, run on a 22-task harness in August 2026; task sets, versions, and raw logs are published alongside them. Everything else on the board is compiled: vendor documentation, independent evals (Artificial Analysis, LMArena, NIST CAISI), and developer-reported tests, every number linked to the source we took it from, vendor-reported figures labeled as such. Rows we can't source yet sit in a tracking queue with no scores. We expand in-house testing model by model. When you see "Benchmarks Desk" on a byline, it means the data behind the article is either our raw logs or someone else's that we can point you to, never a number we can't back up.
Articles
- V4 Flash on ARC-AGI: 89% for $0.02 a taskAug 8, 2026
- China vs US AI models, mid-2026: a scorecardAug 6, 2026
- GLM-5.2: the open-source SOTA, warts and allAug 6, 2026
- Kimi K3: 180K downloads, 340 fine-tunes, 72 hoursAug 6, 2026
- MiniMax H3: video generation at 0.8 yuan a secondAug 6, 2026
- Qwen 3.8: Alibaba's 2.4T-parameter launch, one week inAug 6, 2026
News Desk and Benchmarks Desk are organizational bylines (collective editorial desks — see About); the editor is Eli Chen, a real person with a verifiable GitHub profile. Future individual contributors reuse this page pattern.