About
China AI Bench is the independent intelligence layer between China's AI ecosystem and the global AI community. We run the models, check the claims, and publish the receipts in English. We evaluate models developed in China on their merits. We are not affiliated with any Chinese lab, vendor, or government agency.
Why this exists
Most English-language coverage of Chinese AI is translated headlines and vendor press releases. Nobody answers the questions developers ask: Is DeepSeek really that good? Can I build on Qwen, Kimi, or GLM? Do the benchmark numbers survive contact with real tasks? We built this site to answer them with evidence, not marketing.
What we cover
Three things:
- Benchmarks — hands-on evaluations of Chinese models, tasks and results published (Benchmarks).
- News — a daily briefing, source-backed, with a one-line take (News).
- Deep dives — technical reviews of what changed and why it matters for your stack (Deep Dives).
How we test
As of August 2026, two models carry our formal in-house badge — DeepSeek V4 Pro and V4 Flash, run on a 22-call harness in early August and re-confirmed on Aug 7. Five Chinese models — GLM-5.2, Doubao Seed 2.1, Step 3.5 Flash, Hunyuan A13B and Ling Flash 2.0 — passed a one-pass, N=1 smoke battery on Aug 7, 2026, published separately as "Smoke-tested" with raw logs, not as formal tests. Four US/EU models sit in an International reference zone (GPT-5.6 Luna Pro, Claude Opus 5, Gemini 3.6 Flash, Grok 4.5), same battery, reference only. The rest of the board is compiled with sources. Raw logs are public for every tested and smoke-tested row.
Editorial standards
Every claim sourced, every number verifiable. Articles are published under organizational bylines for desk reporting, with named editors responsible for analysis pieces. AI tools assist drafting and translation; nothing publishes without human review. We don't take money from model vendors, and rankings are not for sale. Vendors get a right of reply: their response to our findings runs alongside the original piece. Our methodology and raw data are public. Re-run our numbers if you like.
Who runs it
The editor, Eli Chen, is a full-stack developer in Beijing who spent over a decade writing production code — and changed his mind about Chinese models only after testing them hands-on. News Desk handles the briefings; Benchmarks Desk runs the evaluations.
Business model & privacy
Independent: no sponsored rankings, no paid placements. Ads fund the site. If we ever run affiliate links, they'll be disclosed inline. Analytics are aggregate and consent-based; we don't collect accounts or emails beyond an optional newsletter.
Full policy: Editorial Policy · Privacy (page ships at launch).
Contact
Email address goes live with the launch.