China AI Bench

Blog

The full archive: every article in one place — hands-on reviews with reproducible task batteries, benchmark head-to-heads and analysis. Newest first.

ReviewBenchmarks Deskby Eli Chen, Editor-in-chief10 min read

China vs US frontier: 9 models, 11 tasks, same battery

We compiled the same 11-task battery on 9 frontier models: GLM-5.2, Seed 2.1, Step 3.5 Flash, Hunyuan A13B, Ling 2.0 vs GPT-5.6 Luna Pro, Claude Opus 5, Gemini 3.6 Flash, Grok 4.5. Six scored 11/11; the only US-side failure was a quota issue, not capability. Costs spanned ~100x, from $0.0004 to $0.04. Raw logs public.

CommentaryBenchmarks Deskby Benchmarks Desk, Testing & rankings desk4 min read

V4 Flash on ARC-AGI: 89% for $0.02 a task

ARC Prize verified DeepSeek V4 Flash-0731: 89.0% on ARC-AGI-1 Semi-Private at $0.02 per task, 61.4% on ARC-AGI-2. What the verified scores mean, why the non-monotonic failures echo our own test run, and what $0.02 reasoning does to agent economics.

CommentaryNews Deskby News Desk, Daily briefing desk4 min read

MiniMax H3 Weights Are Up, License Bars 4 Markets

MiniMax H3's weights are on Hugging Face, and the license excludes the EU, UK, South Korea, and the US. The license text, verbatim, and what it means for developers.

CommentaryBenchmarks Deskby Benchmarks Desk, Testing & rankings desk8 min read

China vs US AI models, mid-2026: a scorecard

At mid-2026 the China-US model race has a scorecard with two columns: China wins open weights, volume, and price; the US still holds the frontier lead. 13 straight weeks of Chinese token volume, an 80% OpenAI price cut, and the numbers in between.

ReviewBenchmarks Deskby Benchmarks Desk, Testing & rankings desk6 min read

GLM-5.2: the open-source SOTA, warts and all

Zhipu's GLM-5.2 leads the open-source field on Artificial Analysis at 51 points, tops the Code Arena for available models, and publishes its weak spots: 45-minute reasoning runs and a 5% knowledge lag. We compiled the numbers.

ReviewBenchmarks Deskby Benchmarks Desk, Testing & rankings desk7 min read

Kimi K3: 180K downloads, 340 fine-tunes, 72 hours

Moonshot's Kimi K3 is the largest open-weight model ever shipped: 2.8T params, 180K Hugging Face downloads in 72 hours, 340+ fine-tunes, and a demand crush that paused new subscriptions. We compiled the numbers and the caveats.

ReviewBenchmarks Deskby Benchmarks Desk, Testing & rankings desk6 min read

MiniMax H3: video generation at 0.8 yuan a second

MiniMax H3 tops the Artificial Analysis video editing chart, hit #1 on Hugging Face trending, and prices video generation at 0.8 yuan per second, about a third of flagship rivals. We compiled the specs, the price math, and the open-weights promise.

ReviewBenchmarks Deskby Benchmarks Desk, Testing & rankings desk5 min read

Qwen 3.8: Alibaba's 2.4T-parameter launch, one week in

One week after Alibaba's Qwen 3.8 launch: what we can verify (2.4T params, ~95B active, IDC Leaders quadrant, eight day-one integrations) and what we can't (API pricing, independent benchmarks, the weights).

CommentaryField Notesby Eli Chen, Editor-in-chief3 min read

NIST CAISI vs DeepSeek V4: the 8-month gap

The US government's independent eval shop ran DeepSeek V4 Pro on closed benchmarks. Result: still the most capable Chinese model CAISI has tested, about 8 months behind the US frontier — and ahead on capability per dollar. What the gap actually means.