Long-form work: hands-on reviews with reproducible task batteries, benchmark head-to-heads and analysis. Every piece links its sources — task sets, raw logs, vendor tables and independent evals.
·ReviewBenchmarks Desk·by Eli Chen, Editor-in-chief·10 min read
We compiled the same 11-task battery on 9 frontier models: GLM-5.2, Seed 2.1, Step 3.5 Flash, Hunyuan A13B, Ling 2.0 vs GPT-5.6 Luna Pro, Claude Opus 5, Gemini 3.6 Flash, Grok 4.5. Six scored 11/11; the only US-side failure was a quota issue, not capability. Costs spanned ~100x, from $0.0004 to $0.04. Raw logs public.
·CommentaryBenchmarks Desk·by Benchmarks Desk, Testing & rankings desk·4 min read
ARC Prize verified DeepSeek V4 Flash-0731: 89.0% on ARC-AGI-1 Semi-Private at $0.02 per task, 61.4% on ARC-AGI-2. What the verified scores mean, why the non-monotonic failures echo our own test run, and what $0.02 reasoning does to agent economics.
·CommentaryNews Desk·by News Desk, Daily briefing desk·4 min read
MiniMax H3's weights are on Hugging Face, and the license excludes the EU, UK, South Korea, and the US. The license text, verbatim, and what it means for developers.
·CommentaryBenchmarks Desk·by Benchmarks Desk, Testing & rankings desk·8 min read
At mid-2026 the China-US model race has a scorecard with two columns: China wins open weights, volume, and price; the US still holds the frontier lead. 13 straight weeks of Chinese token volume, an 80% OpenAI price cut, and the numbers in between.
·ReviewBenchmarks Desk·by Benchmarks Desk, Testing & rankings desk·6 min read
Zhipu's GLM-5.2 leads the open-source field on Artificial Analysis at 51 points, tops the Code Arena for available models, and publishes its weak spots: 45-minute reasoning runs and a 5% knowledge lag. We compiled the numbers.
·ReviewBenchmarks Desk·by Benchmarks Desk, Testing & rankings desk·7 min read
Moonshot's Kimi K3 is the largest open-weight model ever shipped: 2.8T params, 180K Hugging Face downloads in 72 hours, 340+ fine-tunes, and a demand crush that paused new subscriptions. We compiled the numbers and the caveats.
·ReviewBenchmarks Desk·by Benchmarks Desk, Testing & rankings desk·6 min read
MiniMax H3 tops the Artificial Analysis video editing chart, hit #1 on Hugging Face trending, and prices video generation at 0.8 yuan per second, about a third of flagship rivals. We compiled the specs, the price math, and the open-weights promise.
·ReviewBenchmarks Desk·by Benchmarks Desk, Testing & rankings desk·5 min read
One week after Alibaba's Qwen 3.8 launch: what we can verify (2.4T params, ~95B active, IDC Leaders quadrant, eight day-one integrations) and what we can't (API pricing, independent benchmarks, the weights).
·CommentaryField Notes·by Eli Chen, Editor-in-chief·3 min read
The US government's independent eval shop ran DeepSeek V4 Pro on closed benchmarks. Result: still the most capable Chinese model CAISI has tested, about 8 months behind the US frontier — and ahead on capability per dollar. What the gap actually means.
·ReviewBenchmarks Desk·by Eli Chen, Editor-in-chief·13 min read
Hands-on review of DeepSeek V4 Pro and Flash-0731: vendor benchmarks vs NIST CAISI's independent eval, an 11-task API battery (11/11 and 10/11 passed, $0.0013 total), and the real pricing math.