By Benchmarks Desk · updated
Best Chinese AI Models: Ranked & Sourced (2026)
Ranked by data verifiability, not raw performance. The rows we tested in-house anchor the board; smoke-tested rows sit in their own group (N=1, single-run); compiled rows follow, ordered by data completeness and ecosystem heat. Row order is a statement about evidence, not about which model is "better".
Last full retest: · Next: Sep 6, 2026
Two rows carry our formal in-house badge — DeepSeek V4 Pro (11/11) and V4 Flash-0731 (10/11), each measured across 22 API calls in August 2026 and re-confirmed on Aug 7. Task sets, versions and raw logs are public.
Five more Chinese models — GLM-5.2, Doubao Seed 2.1, Step 3.5 Flash, Hunyuan A13B and Ling Flash 2.0 — passed 10-11 of 11 tasks on a single N=1 pass each (Aug 7, 2026). They sit in a separate Smoke-tested group with public raw logs, not promoted to the formal badge.
Four US/EU models — GPT-5.6 Luna Pro, Claude Opus 5, Gemini 3.6 Flash and Grok 4.5 — ran the same 11-task battery through OpenRouter's free tier in an International reference zone. Reference only: nothing there is ranked against the Chinese models.
Kimi K3, Qwen 3.8 and MiniMax H3 are compiled rows: every number on them links to the source we took it from, with the capture date.
V4 Pro still trails the US frontier by about 8 months, per NIST CAISI's independent evaluation.
Six more models sit in a separate tracking queue — we're watching releases and official pages, but there's no verified data to rank them yet, so they get no scores.
How this board is ordered
Row order tracks data verifiability, not raw performance. Three kinds of rows sit on this board. "Tested by us": our formal in-house runs — two models (DeepSeek V4 Pro, V4 Flash), each measured across 22 API calls in August 2026 and re-confirmed, raw logs public. "Smoke-tested": five Chinese models run once through our 11-task battery on Aug 7, 2026 (N=1, single-run) — a smoke check with public raw logs, kept in its own group, not promoted to the formal badge. "Compiled": rows built from vendor docs, third-party evals and community tests, every number linked to its source.
Four US/EU models sit in a separate International reference zone below (reference only, not ranked against the Chinese models).
We don't blend scores from different pipelines into a composite: our task-set results, smoke runs and third-party indexes use different scales and are not directly comparable.
Every number on this board links to the source we took it from, with the capture date.
Where the numbers come from
Rows marked "Tested by us" are our formal in-house runs — two models, 22-call harness, re-confirmed on Aug 7, 2026. Task sets, sampling parameters, model versions, and raw logs are published alongside them.
Rows marked "Smoke-tested" are one-pass runs: each of the five Chinese models went through the same 11-task battery exactly once on Aug 7, 2026 (N=1). Single runs catch gross failures and shape a first impression; they are not statistical tests, and we don't call them that. Raw logs are public for every smoke row (GitHub + Hugging Face), so you can read the exact prompt, output and verdict yourself.
The International reference zone ran the same battery on four US/EU models through OpenRouter's free tier, N=1: same prompts, same judge. Reference only: not ranked against the Chinese models.
Everything else is compiled: from vendor documentation, independent evaluations (Artificial Analysis, LMArena, NIST CAISI), and tests reported by developers and Chinese-language tech media. Every number links to the source we took it from, with the date we captured it. Vendor-reported numbers are labeled vendor-reported. Unverified numbers say "unverified". We'd rather show you a sourced number we didn't run than a score we can't back up.
Rows we can't source yet sit in a separate tracking queue, clearly labeled, with no scores. We add tests as we run them.
Main board
How to read this board
For performance questions, read across, not down. The AA Intelligence Index column puts Kimi K3 at 57.1 (#3 of 189) and V4 Flash at 50; Frontend Code Arena puts K3 at 1679, #1. Those are third-party pipelines with their own scales.
Our in-house scores (11/11, 10/11) come from a different harness and are not convertible to them. The Smoke-tested rows are N=1 single runs — they tell you whether a model handled a task class at all, not how reliably. Row order is not a performance ranking.
- Spec
- 1.6T total / 49B activated, MIT, 1M ctx, text-only
- Price (snapshot)
- $0.435/$0.87; cache-hit $0.003625 (75% permanent discount)[3]2026-07
- Context
- 1M token context
- Spec
- 284B total / 13B activated, MIT, 1M ctx, Flash-0731 stable build
- Price (snapshot)
- $0.14/$0.28; cache-hit $0.0028[2]2026-06
- Context
- 1M token context
| Group | Model | Our run / Key numbers | Spec | Price (snapshot) | Context | Sources | Updated |
|---|---|---|---|---|---|---|---|
| Tested by us · T1 | preview (2026-07) | 1.6T total / 49B activated, MIT, 1M ctx, text-only | $0.435/$0.87; cache-hit $0.003625 (75% permanent discount)[3] snapshot: 2026-07 | 1M token context | raw logs, 2026-08-07 → | 2026-08-07 | |
| Tested by us · T2 | stable (2026-07-31) | 284B total / 13B activated, MIT, 1M ctx, Flash-0731 stable build | $0.14/$0.28; cache-hit $0.0028[2] snapshot: 2026-06 | 1M token context | raw logs, 2026-08-07 → | 2026-08-07 |
- AA Intelligence Index 57.1, #3 of 189[6](third-party eval)
- Frontend Code Arena 1679, #1 (preliminary; Fable 5 1631, Sol 1618)[5](third-party eval)
- SWE Marathon 42.0[5](third-party eval)
- Terminal-Bench 2.1 88.3[20](media report)
- HF downloads 180K in 72h; 340+ fine-tunes[9](media report)
- Specs: 2.8T/104B activated, 1M ctx, native vision[7](official)
K3's numbers are third-party evals and media reports we compile and link, not our in-house runs. Different pipelines: don't read them against our task-set scores.
- Spec
- 2.8T total / 104B activated (896 experts), Modified MIT, 1M ctx, native vision
- Price (snapshot)
- $3/$15; cache-hit $0.30[8]2026-08-06
- Context
- 1M token context, native visionLargest open-weight model to date; 4-bit weights ≈1.4TB, needs 64+ accelerators.
- Vendor self-report: 'next to Fable 5' (no independent eval yet — labeled vendor-reported)[12](vendor-reported)
- IDC MarketScape: only Chinese vendor in Leaders quadrant[15](third-party eval)
- Day-0 integration: OpenRouter, OpenCode, Hermes Agent, Vercel[12](media report)
- Specs: 2.4T/95B activated, ~1M ctx, released 8/3[12](official)
- Spec
- 2.4T total / 95B activated MoE, ~1M ctx (983,616 tokens), max output 131K, thinking-only mode
- Context
- ~1M token context, first trillion-parameter multimodal releaseWeights promised 'next week' as of 8/6, not yet published; Qwen3.8-27B dense model coming with them.
- Spec
- Unified text/image/video/audio generation, dual-channel A/V output, up to 15s 2K; H3-Base 33B / 42.5GB open weights (8/3, license excludes US/EU/UK/KR)
- Price (snapshot)
- ¥0.8/sec (2K video, ~1/3 of flagship tier)[17]2026-08-06
- Context
- Video-first flagship; H3-Base open for self-hostingOpen-weights promise from 7/31 landed 8/3 with H3-Base — region-restricted license, so not fully open.
| Group | Model | Our run / Key numbers | Spec | Price (snapshot) | Context | Sources | Updated |
|---|---|---|---|---|---|---|---|
| Compiled from public evals · C1 | Kimi K3 2.8T/104B activated, open weights 7/27 |
K3's numbers are third-party evals and media reports we compile and link, not our in-house runs. Different pipelines: don't read them against our task-set scores. | 2.8T total / 104B activated (896 experts), Modified MIT, 1M ctx, native vision | $3/$15; cache-hit $0.30[8] snapshot: 2026-08-06 | 1M token context, native vision Largest open-weight model to date; 4-bit weights ≈1.4TB, needs 64+ accelerators. | [6][5][20][9][7] | 2026-08-06 |
| Compiled from public evals · C2 | Qwen 3.8 2.4T/95B activated MoE, released 8/3 |
| 2.4T total / 95B activated MoE, ~1M ctx (983,616 tokens), max output 131K, thinking-only mode | snapshot: 2026-08-06 | ~1M token context, first trillion-parameter multimodal release Weights promised 'next week' as of 8/6, not yet published; Qwen3.8-27B dense model coming with them. | [12][15] | 2026-08-06 |
| Compiled from public evals · C3 | MiniMax H3 full-modality generation, H3-Base open 8/3 | Unified text/image/video/audio generation, dual-channel A/V output, up to 15s 2K; H3-Base 33B / 42.5GB open weights (8/3, license excludes US/EU/UK/KR) | ¥0.8/sec (2K video, ~1/3 of flagship tier)[17] snapshot: 2026-08-06 | Video-first flagship; H3-Base open for self-hosting Open-weights promise from 7/31 landed 8/3 with H3-Base — region-restricted license, so not fully open. | [16][17][18][19] | 2026-08-06 |
All 11 tasks passed with thinking disabled on the direct API.
All 11 tasks passed.
Step-3.5-Flash forces reasoning on SiliconFlow (ignores the thinking parameter). ipv4-validator failed: first run exhausted 1024 max_tokens (empty content); the 4096 retry burned ~15.5K reasoning tokens without converging. Behavioral/config trait, not a judged capability miss.
Hunyuan-A13B answered "3, -4" for the quadratic task (roots are 4 and -3) — the sign check was skipped (3²-3-12=-6≠0). Plain capability error, reported as-is.
| Group | Model (version locked) | Passed (N=1) | Battery cost | Note | Raw logs |
|---|---|---|---|---|---|
| Smoke-tested · S1 | GLM-5.2 glm-5.2 (thinking disabled) | 11/11 tasks N=1 · 2026-08-07 | $0.003323 | All 11 tasks passed with thinking disabled on the direct API. | HF →GitHub → |
| Smoke-tested · S2 | Doubao Seed 2.1 Pro doubao-seed-2-1-pro-260628 | 11/11 tasks N=1 · 2026-08-07 | $0.003306 | All 11 tasks passed. | HF →GitHub → |
| Smoke-tested · S3 | Step 3.5 Flash step-3.5-flash (SiliconFlow) | 10/11 tasks N=1 · 2026-08-07 | $0.002910 | Step-3.5-Flash forces reasoning on SiliconFlow (ignores the thinking parameter). ipv4-validator failed: first run exhausted 1024 max_tokens (empty content); the 4096 retry burned ~15.5K reasoning tokens without converging. Behavioral/config trait, not a judged capability miss. | HF →GitHub → |
| Smoke-tested · S4 | Hunyuan A13B hunyuan-a13b (SiliconFlow) | 10/11 tasks N=1 · 2026-08-07 | $0.000535 | Hunyuan-A13B answered "3, -4" for the quadratic task (roots are 4 and -3) — the sign check was skipped (3²-3-12=-6≠0). Plain capability error, reported as-is. | HF →GitHub → |
| Smoke-tested · S5 | Ling Flash 2.0 ling-flash-2.0 (SiliconFlow) | 11/11 tasks N=1 · 2026-08-07 | $0.000444 | All 11 tasks passed — cheapest battery in the batch. | HF →GitHub → |
Data layers: 2 models formally tested in-house (22-call harness, raw logs public) · 5 smoke-tested (11-task battery, N=1, Aug 7 2026, raw logs public) · 4 US/EU models in an International reference zone (reference only, not ranked) · 3 compiled from vendor docs and third-party evals, each number linked · 6 in the tracking queue.
International reference
Why are US/EU models on a Chinese-model board? These rows answer the question this board gets most: "how do these compare to GPT-5 / Claude?" Same 11-task battery, same judge, same day (Aug 7, 2026), N=1 — run through OpenRouter's free tier. Reference only: nothing here is ranked against the Chinese models above, and nothing here changes where a Chinese model sits on the board.
International reference — how to read this zone
These four rows ran on the same 11-task battery and judge as the rows above, through OpenRouter (free tier, N=1). Same prompts, same scoring. What differs: the channel (OpenRouter gateway, so latencies here are zone-comparable only), reasoning behavior (Claude and Gemini arrive with reasoning enabled by default; flagged per row), and free-tier quotas: after the $0.15 grant was spent, OpenRouter rejected further calls with HTTP 402, which is why Claude's retry didn't complete. We report the 10/11 as it ran, not as we'd like it to look. Cost to re-run Claude with reasoning off: about $0.013 (it's on the retest list). This zone is reference, not rank: nothing here is ranked against the Chinese models.
All 11 passed — cheapest battery in the international zone.
Claude Opus 5 failed only writing-product-copy, and that's a quota failure, not a capability one. With OpenRouter's default reasoning enabled, the model spent its 256 max_tokens on the reasoning trace and returned empty content. Our retry with thinking disabled (512 tokens) was rejected by the free-tier balance (HTTP 402), so we keep the honest 10/11. Re-run cost is about $0.013 — it's on the retest list.
Gemini 3.6 Flash arrives with reasoning enabled by default. writing-product-copy first ran truncated at 1024 max_tokens (FAIL) and passed on the 2048 retry (first_run_truncated kept in the log).
| Zone | Model (OpenRouter ID) | Passed (N=1) | Battery cost | List price $/1M | Note | Raw logs |
|---|---|---|---|---|---|---|
| Reference · I1 | GPT-5.6 Luna Pro openai/gpt-5.6-luna-pro · OpenRouter free tier | 11/11 tasks N=1 · 2026-08-07 | $0.004060 | $0.1/$0.6 | All 11 passed — cheapest battery in the international zone. | HF →GitHub → |
| Reference · I2 | Claude Opus 5 anthropic/claude-opus-5 · OpenRouter free tier | 10/11 tasks N=1 · 2026-08-07 | $0.031625 | $5/$25 | Claude Opus 5 failed only writing-product-copy, and that's a quota failure, not a capability one. With OpenRouter's default reasoning enabled, the model spent its 256 max_tokens on the reasoning trace and returned empty content. Our retry with thinking disabled (512 tokens) was rejected by the free-tier balance (HTTP 402), so we keep the honest 10/11. Re-run cost is about $0.013 — it's on the retest list. | HF →GitHub → |
| Reference · I3 | Gemini 3.6 Flash google/gemini-3.6-flash · OpenRouter free tier | 11/11 tasks N=1 · 2026-08-07 | $0.040614 | $1.5/$7.5 | Gemini 3.6 Flash arrives with reasoning enabled by default. writing-product-copy first ran truncated at 1024 max_tokens (FAIL) and passed on the 2048 retry (first_run_truncated kept in the log). | HF →GitHub → |
| Reference · I4 | Grok 4.5 x-ai/grok-4.5 · OpenRouter free tier | 11/11 tasks N=1 · 2026-08-07 | $0.019824 | $2/$6 | All 11 passed. | HF →GitHub → |
International rows tested via OpenRouter free tier, N=1, Aug 7 2026 · same 11-task battery and judge as the main board · reference only, not ranked against Chinese models · latencies include gateway overhead (zone-comparable only) · Claude Opus 5's 10/11 is a quota failure (max_tokens spent on reasoning; retry blocked by 402), not a capability failure.
In our tracking queue
These models are on our radar but we don't have enough verified data to rank them yet. We're watching release notes, official pricing pages, and independent evals. Rows graduate to the main board when a data point we can link appears.
- Why we're tracking
- Previous flagship, superseded by Qwen 3.8.
- What's needed to rank it
- None — we keep 3.8 as the only Qwen row on the main board.
- Why we're tracking
- Superseded by K3; useful as a historical comparison anchor.
- What's needed to rank it
- None — deep-dive comparisons can cite it without a board row.
- Why we're tracking
- Text-line flagship; M2.5 is already API-accessible on SiliconFlow.
- What's needed to rank it
- M2.5 marked testable — pending a task-card decision on a text-line smoke run (H3 stays Compiled: video flagship, 11-task text battery can't cover it).
- Why we're tracking
- Material gap closed; but vendor claims and third-party score disagree sharply, needs care.
- What's needed to rank it
- Collected 2026-08-06 (specs, tiered pricing; third-party score last vs vendor 'top tier' claim) — review required before promotion.
- Why we're tracking
- U-line is an image/multimodal line, not a text flagship; no AA-style third-party score.
- What's needed to rank it
- Partially collected (flagship pricing; U1.5-Lite-Preview open-sourced 8/3) — stays tracked until content-editor review.
- Why we're tracking
- Fully open-source, no API pricing; benchmarks are vendor-reported only.
- What's needed to rank it
- Partially collected (vendor-reported benchmark only, no third-party) — stays tracked until content-editor review.
Aug 7 battery, task by task
Every cell below is one real API call from the Aug 7, 2026 run — 9 models × 11 tasks, N=1 per cell. The formal "Tested by us" rows (DeepSeek V4 Pro, V4 Flash) are not in this matrix: they were re-confirmed on the same battery but carry the 22-call harness, a different batch.
| Model | IPv4 validatorcoding | Bug fix: second-largest distinct elementcoding | Math: quadratic roots with b+c = -13math | Math: consecutive integers summing to 575math | Logic: syllogismmath | Knowledge: 36th US presidentmultilingual | Knowledge: capital of Burkina Fasomultilingual | Writing: 3-sentence product copy, constrainedmultilingual | Extraction: dates and names as JSONcoding | Summary as 3-item JSON arraycoding | Needle in a 426-token documentlongContext | Total |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| GLM-5.2 | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | 11/11 |
| Doubao Seed 2.1 Pro | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | 11/11 |
| Step 3.5 Flash | ❌a | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | 10/11 |
| Hunyuan A13B | ✅ | ✅ | ❌b | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | 10/11 |
| Ling Flash 2.0 | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | 11/11 |
| GPT-5.6 Luna Proref | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | 11/11 |
| Claude Opus 5ref | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ❌c | ✅ | ✅ | ✅ | 10/11 |
| Gemini 3.6 Flashref | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | 11/11 |
| Grok 4.5ref | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | 11/11 |
Failure notes
- [a]Step 3.5 Flash — IPv4 validator: [retry max_tokens=4096] function not found -- initial FAIL was config-truncation, retry also FAILED
- [b]Hunyuan A13B — Math: quadratic roots with b+c = -13: contains '4'=True, '-3'=False
- [c]Claude Opus 5 — Writing: 3-sentence product copy, constrained: sentences=0 (need 3) — first run FAIL, OpenRouter default reasoning consumed 256/256 max_tokens (finish=length), final content empty. Config-fix retries (thinking disabled 512 / 2048 max_tokens) blocked by free-tier credit limit (HTTP 402, remaining balance <=449 output tokens at claude-opus-5 pricing). Recorded as real behavior, not retried on paid credits per budget rule.
Raw logs for every cell: raw logs page · mirrored on GitHub + Hugging Face.
6 of 9 models passed all 11 tasks on the Aug 7 battery; the only US-side failure was a quota issue, not a capability one.
By task dimension (tested rows only)
Coding
Pick: DeepSeek V4 Pro. 11/11 in our battery, and the only model that rejected '1.2.3.04' as an IPv4 address — strictness matters.
| Model | Score | Note |
|---|---|---|
| DeepSeek V4 Pro | 10 | Both code tasks passed first try |
| DeepSeek V4 Flash-0731 | 9.5 | returned sorted(set(xs)) — the fix we'd write |
| Compiled rows (not our runs) | — | K3 #1 Frontend Code Arena, GLM-5.2 #1 DesignArena — see main board |
Math & Reasoning
Pick: DeepSeek V4 Pro. Passed both math tasks. Flash failed 23×25 with thinking off — but fixed it with one flag (see review P.S.).
| Model | Score | Note |
|---|---|---|
| DeepSeek V4 Pro | 10 | Quadratic roots + consecutive integers |
| DeepSeek V4 Flash-0731 | 8 | 1/2 with thinking off; 2/2 with thinking on |
| — | — |
Chinese language
Pick: Untested yet. Our first battery was English-only. Chinese-language tasks are queued for the next retest — a Chinese-model site without a Chinese test would be unserious.
| Model | Score | Note |
|---|---|---|
| Queued | — | Next battery: zh task set v1 |
Multilingual
Pick: DeepSeek V4 Flash-0731. Both models handled English extraction/writing tasks at parity; Flash did it at 1/3 the price.
| Model | Score | Note |
|---|---|---|
| DeepSeek V4 Flash-0731 | 10 | English tasks at $0.28/1M out |
| DeepSeek V4 Pro | 10 | Same pass rate, 3x the price |
Long context
Pick: Both — with a caveat. Needle-in-a-426-token test passed on both. The full 1M token context is on our queue; at $0.14/1M in, DeepSeek is 18x cheaper than GPT-5.4 for a 300K doc.
| Model | Score | Note |
|---|---|---|
| DeepSeek V4 Flash-0731 | 10 | 426-token needle: PASS |
| DeepSeek V4 Pro | 10 | 426-token needle: PASS |
Cost efficiency
Pick: DeepSeek V4 Flash-0731. Our whole 22-call battery cost $0.0013. Artificial Analysis puts Flash-0731 at ~$0.03 per index task — the cheapest board-wide.
| Model | Score | Note |
|---|---|---|
| DeepSeek V4 Flash-0731 | 10 | $0.14/$0.28 per 1M tokens |
| DeepSeek V4 Pro | 7 | $0.435/$0.87 — still cheap, 3.1x Flash |
| GPT-5.4 (reference) | 3 | $2.50/$15 — 54x Flash output price |
Data sources
Every compiled number on this board resolves to one of the entries below — URL, capture date, and nature. Vendor-reported numbers are labeled as such; we don't launder vendor claims into third-party ones.
- [1] https://news.qq.com/rain/a/20260801A048U500captured 2026-08-06 · vendor-reported — DeepSeek V4 9-item agent benchmark release (Tencent News recap)
- [2] https://developer.puter.com/tutorials/deepseek-api-pricing/captured 2026-08-06 · official — DeepSeek API pricing page (Puter tutorial mirror)
- [3] https://www.36kr.com/p/3919319636290946captured 2026-08-06 · official — DeepSeek V4 Pro & GLM-5.2 pricing (36Kr)
- [4] https://www.oschina.net/news/486803captured 2026-08-06 · community — V4 Flash 8T-token single-day usage report (OSChina)
- [5] https://www.jiemian.com/article/14786633.htmlcaptured 2026-08-06 · third-party eval — K3 Frontend Code Arena 1679 #1, SWE Marathon 42.0 (Jiemian, preliminary)
- [6] https://macgpu.com/zh/blog/2026-kimi-k3-kaiyuan-quanzhong-fabu.htmlcaptured 2026-08-06 · third-party eval — K3 AA Intelligence Index 57.1, #3/189 (MACGPU)
- [7] https://www.ithome.com/0/982/259.htmcaptured 2026-08-06 · official — K3 specs: 2.8T/104B, 1M ctx, native vision (ITHome)
- [8] https://ollama.com/library/kimi-k3captured 2026-08-06 · official — K3 pricing $3/$15, cache $0.30 (Ollama model page)
- [9] https://blog.csdn.net/xyghehehehe/article/details/163353310captured 2026-08-06 · media report — K3 HF 180K downloads in 72h, 340+ fine-tunes (CSDN)
- [10] https://docs.bigmodel.cn/cn/guide/models/text/glm-5.2captured 2026-08-06 · official — GLM-5.2 specs, AA 51 open-source SOTA, FrontierSWE (Zhipu docs)
- [11] https://juejin.cn/post/7654122741078671402captured 2026-08-06 · community — GLM-5.2 DesignArena 1360 ELO #1, Code Arena #1 (Juejin)
- [12] https://news.qq.com/rain/a/20260804A0ACRT00captured 2026-08-06 · official — Qwen 3.8 release: 2.4T/95B, 8/3 launch, weights 'next week' (Tencent News)
- [13] https://www.datalearner.com/ai-models/pretrained-models/qwen3-8-maxcaptured 2026-08-06 · official — Qwen 3.8 CN pricing ¥12/¥36, cache ¥1.5 (DataLearner)
- [14] https://www.aipricedb.com/ai-news/qwen-launches-qwen3-8-max-with-2-4-trillion-parameters-c98759b1captured 2026-08-06 · official (converted) — Qwen 3.8 intl ~$2/$6, cache $0.25 — converted from vendor percentages (AIPriceDB)
- [15] https://www.leiphone.com/category/industrynews/1gS0E0CcUDxEgbQR.htmlcaptured 2026-08-06 · third-party eval — IDC MarketScape: Qwen only Chinese vendor in Leaders (Leiphone)
- [16] https://www.leiphone.com/category/industrynews/aiUMBoeUYbi8fX4x.htmlcaptured 2026-08-06 · third-party eval — MiniMax H3 AA video-editing #1, Arena image-to-video #1, HF trending #1 (Leiphone)
- [17] https://www.163.com/dy/article/L3623H0E05198NMR.htmlcaptured 2026-08-06 · media report — H3 ¥0.8/sec at 2K, ~1/3 of flagship (NetEase)
- [18] https://www.atlascloud.ai/zh/blog/guides/minimax-h3-open-source-weightscaptured 2026-08-06 · official — H3-Base 33B / 42.5GB, dual checkpoints, license excludes US/EU/UK/KR (AtlasCloud + HuggingFace MiniMaxAI/MiniMax-H3)
- [19] https://www.ithome.com/0/984/379.htmcaptured 2026-08-06 · media report — H3 open-weights landing 8/3 (ITHome)
- [20] https://explore.n1n.ai/zh/blog/kimi-k3-moonshot-ai-kaiyuan-moxing-2026-07-17captured 2026-08-06 · media report — K3 Terminal-Bench 2.1 88.3 (n1n.ai blog)
- [25] /blog/deepseek-v4-review-2026captured 2026-08-03 · our review — In-house 22-call battery + NIST CAISI cross-check (site review)
Methodology
- Three kinds of rows sit on this board. "Tested by us": our formal in-house runs — two models (DeepSeek V4 Pro, V4 Flash), each measured across 22 API calls in August 2026 and re-confirmed, raw logs public. "Smoke-tested": five Chinese models run once through our 11-task battery on Aug 7, 2026 (N=1, single-run) — a smoke check with public raw logs, kept in its own group, not promoted to the formal badge. "Compiled": rows built from vendor docs, third-party evals and community tests, every number linked to its source. Four US/EU models sit in a separate International reference zone below (reference only, not ranked against the Chinese models). We don't blend scores from different pipelines into a composite: our task-set results, smoke runs and third-party indexes use different scales and are not directly comparable.
- Formal in-house runs (T rows): 11 tasks written before the run — IPv4 validator, bug fix, 2 math, syllogism, 2 knowledge, constrained writing, JSON extraction, JSON summary, long-doc needle. Battery run twice (22 API calls per model); the single failure reproduced.
- Smoke runs (S rows, Aug 7, 2026): same 11-task battery, exactly one pass per model (N=1). Sampling: official vendor APIs (SiliconFlow / Volcano / Zhipu / MiniMax direct), temperature default, max_tokens 64-512 per task, thinking settings disclosed per row (GLM/Seed disabled; SiliconFlow trio at vendor default; Step forces reasoning). Single runs catch gross failures; they are not statistical tests.
- International zone (I rows): the same battery through OpenRouter free tier, N=1, same judge. Latencies include gateway overhead and are zone-comparable only. Claude and Gemini arrive with reasoning enabled by default; the Claude retry was blocked by a free-tier balance (HTTP 402) — reported as it ran.
- Versions snapshotted: deepseek-v4-flash serves Flash-0731 (stable, 2026-07-31); deepseek-v4-pro is the July preview; smoke/international rows carry their exact API model IDs.
- Compiled rows (C): each quantitative data point carries a footnote [n] resolving to a URL, capture date and nature (official / vendor-reported / third-party eval / media report / community). Vendor-reported numbers are labeled as such; unverified numbers say unverified.
- Cross-check: our results × vendor tables (arXiv:2606.19348) × NIST CAISI independent eval (2026-05-01). Where they disagree, we say so.
- Limits: no full-length 1M context test yet, no self-host run, no multimodal (V4 is text-only). Tracking rows get no scores by design.
Full task prompts, raw responses and rubrics for our in-house runs are published with the review: DeepSeek V4 review 2026. Every tested, smoke-tested and reference row links to its raw logs (GitHub + Hugging Face) from the tables above, and the raw logs page lists all 11 models with direct links.
PK: vendor-reported benchmarks (cross-checked, not trusted)
Every number below is vendor-reported (DeepSeek technical report arXiv:2606.19348 and competitor release materials). "n/r" = not reported. Our independent cross-check is in the CAISI section of the review.
| Benchmark | V4-Pro | V4-Flash | GPT-5.4 | Opus 4.6 | Gemini 3.1 Pro |
|---|---|---|---|---|---|
| SWE-bench Verified | 80.6 | 79.0 | 77.2 | 80.8 | 80.6 |
| SWE-bench Pro | 55.4 | 52.6 | n/r | n/r | n/r |
| LiveCodeBench Pass@1 | 93.5 | 91.6 | n/r | 88.8 | 91.7 |
| Terminal-Bench 2.0 | 67.9 | n/r | 75.1 | 65.4 | 68.5 |
| Codeforces rating | 3206 | 3052 | 3168 | n/r | 3052 |
| MMLU-Pro | 87.5 | 86.2 | n/r | n/r | n/r |
| GPQA Diamond | 90.1 | 88.1 | n/r | n/r | n/r |
| Humanity's Last Exam | 37.7 | n/r | 39.8 | 40.0 | 44.4 |
| SimpleQA-Verified | 57.9 | n/r | n/r | n/r | 75.6 |
PK: pricing snapshot (2026-08-06)
API prices in USD per 1M tokens (input/output), except where noted. Prices change — every row carries its snapshot date on the main board.
| Model | Input $/1M | Output $/1M | Note |
|---|---|---|---|
| DeepSeek V4-Flash | $0.14 | $0.28 | Stable build (Flash-0731) |
| DeepSeek V4-Pro | $0.435 | $0.87 | 75% cut July 2026 |
| Kimi K3 | $3 | $15 | Cache-hit $0.30 |
| GLM-5.2 | $1.4 | $4.4 | Coding Plan from $18/mo |
| Qwen 3.8 (intl) | $2 | $6 | Converted; CN ¥12/¥36 |
| GPT-5.4 | $2.5 | $15 | Long context >272K doubles input |
| Claude Opus 4.6 | $5 | $25 | |
| Gemini 3.1 Pro | $2 | $12 |
Deep dives
FAQ
Did you really test these models yourself?
Two of them formally: DeepSeek V4 Pro and V4 Flash, run on a 22-call harness in early August 2026 and re-confirmed on Aug 7. Task sets, versions, and raw logs are published. Five more Chinese models — GLM-5.2, Doubao Seed 2.1, Step 3.5 Flash, Hunyuan A13B, Ling Flash 2.0 — got a one-pass smoke battery on Aug 7, 2026 (N=1, raw logs public). We keep those in a separate "Smoke-tested" group rather than calling them formal tests, because a single run isn't one. Four US/EU models ran the same battery via OpenRouter in an International reference zone — reference only. The remaining rows are compiled, every number linked to its source. Models without verifiable data sit in the tracking queue, with no scores.
Why is the Smoke-tested group separate from Tested by us?
Because N=1 is a smoke check, not a benchmark. Each model ran the 11-task battery exactly once on Aug 7, 2026. A single pass tells you whether the model handles the task class at all and where it visibly fails; it can't measure reliability or rank models against each other. Formal "Tested by us" rows require a 22-call harness plus a re-confirmation pass. Smoke rows stay in their own group until they earn the formal badge.
Can I compare the International reference zone with the main board?
Same prompts, same judge, same 11-task battery: that's the point of the zone. Score comparisons between the zone and the Smoke-tested rows work, with N=1 caveats; latency comparisons don't (direct API vs OpenRouter gateway overhead). The zone is reference, not rank: nothing in it changes where a Chinese model sits on the board.
Are Chinese AI models safe to use?
Safety is a supply-chain question, not a yes/no. V4 is MIT-licensed open weight, so you can self-host and audit it. For API use, weigh data residency and the lack of western SLAs — our review covers the trade-offs, and we don't whitewash the concerns. Note that MiniMax H3's open weights exclude the US, EU, UK and South Korea, which matters if you plan to self-host it.
How do Chinese models compare to GPT-5?
Per NIST CAISI's independent May 2026 evaluation, DeepSeek V4 Pro sits at roughly GPT-5 level — about 8 months behind the US frontier. It wins decisively on capability per dollar, and loses on novel reasoning and knowledge recall. We also ran both sides on the same 11-task battery — see the International reference zone below. The compiled rows cite their own sources with dates.
What is the best Chinese AI model for coding?
Two pipelines answer this, and we keep them separate. On our formal harness, DeepSeek V4 Pro passed 11/11 tasks and V4 Flash-0731 passed 10/11. On our N=1 smoke battery (Aug 7, 2026), GLM-5.2, Doubao Seed 2.1 and Ling Flash 2.0 also passed 11/11 — single runs, not formal scores, so we don't rank from them. On third-party coding evals we compile, Kimi K3 tops Frontend Code Arena at 1679 (preliminary) and GLM-5.2 leads DesignArena at 1360 ELO — different pipelines from our task sets, each number linked and labeled. For agentic coding on a budget, Flash-0731 at $0.14/$0.28 per 1M tokens is the value pick on the board.
Is DeepSeek better than Qwen?
We tested DeepSeek V4 (Pro and Flash) in-house; Qwen 3.8 is on the board as a compiled row because we haven't run it yet. Its numbers come from Alibaba's own reporting (labeled vendor-reported) and IDC's MarketScape, both linked with capture dates. A direct head-to-head needs our own battery — Qwen joins the in-house queue.
How much do Chinese AI models cost?
Snapshot 2026-08-06: DeepSeek V4-Flash $0.14/$0.28 per 1M tokens (in/out), V4-Pro $0.435/$0.87; Kimi K3 $3/$15; GLM-5.2 $1.4/$4.4; Qwen 3.8 ¥12/¥36 in China and about $2/$6 internationally (converted). That's 1/10th to 1/54th of US frontier output pricing. Prices change; every price row carries its snapshot date.
Do Chinese models work well in English?
Our whole first battery ran in English: 21/22 task passes across both formally tested models, including strict JSON extraction and constrained product copy. English is not a weakness for V4 — knowledge recall is (SimpleQA 57.9 vs Gemini 3.1 Pro's 75.6, per our cross-check).
Update history
Last updated . Next full sweep: 2026-09-06. Price snapshots refresh weekly; in-house retests run quarterly or after a major release. International zone retests run quarterly or after a US/EU major release, subject to OpenRouter free-tier balance. Smoke-tested rows get a fresh N=1 pass on the same schedule; two passes in a row start the conversation about a formal badge, never one. Every change lands in the update history below.
v1: initial in-house battery — 11 tasks, 6 dimensions, 22 API calls; raw responses and rubrics published in the review.
Source: in-house run (site review)
v2 rebuild: main board reworked to 6 rows (2 tested in-house + 4 compiled with per-number source links); separate 10-row tracking queue with no scores; data-sources statement, layered FAQ and update-history structure added; Qwen 3.8 price filled in (CN ¥12/¥36, intl ~$2/$6); MiniMax H3 open-weights status updated (H3-Base landed 8/3, region-restricted license).
Source: 20260806-hotspot R1-R3 + 20260806-topup (23 records) + leaderboard spec v2
Plan-A revision: medal ranks (🥇🥈🥉) replaced by two evidence groups (Tested by us / Compiled from public evals) with in-group numbering; ordering methodology statement moved to the top; "How to read this board" added; K3 row gained Terminal-Bench 2.1 88.3 [20]; V4 Flash row gained AA Intelligence Index 50 [1].
Source: leaderboard spec v2 + plan-A revision brief (2026-08-07)
New "Smoke-tested · Aug 7, 2026" group — GLM-5.2, Doubao Seed 2.1, Step 3.5 Flash, Hunyuan A13B and Ling Flash 2.0 passed 10-11 of 11 tasks on a single N=1 run each; raw logs public (GitHub + Hugging Face). New "International reference" zone — GPT-5.6 Luna Pro, Claude Opus 5, Gemini 3.6 Flash and Grok 4.5, same battery via OpenRouter free tier, reference only, not ranked against Chinese models. Tested by us (2 rows) and Compiled (3 rows) unchanged. Raw logs page expanded to 11 models with direct links.
Source: benchmark runs 2026-08-07, data/raw-logs/ + GitHub EliChen-ai/china-ai-bench + HF EliChen-ai/china-ai-bench-benchmarks