China AI Bench

By Benchmarks Desk · updated

Best Chinese AI Models: Ranked & Sourced (2026)

Ranked by data verifiability, not raw performance. The rows we tested in-house anchor the board; smoke-tested rows sit in their own group (N=1, single-run); compiled rows follow, ordered by data completeness and ecosystem heat. Row order is a statement about evidence, not about which model is "better".

Last full retest: · Next: Sep 6, 2026

Two rows carry our formal in-house badge — DeepSeek V4 Pro (11/11) and V4 Flash-0731 (10/11), each measured across 22 API calls in August 2026 and re-confirmed on Aug 7. Task sets, versions and raw logs are public.

Five more Chinese models — GLM-5.2, Doubao Seed 2.1, Step 3.5 Flash, Hunyuan A13B and Ling Flash 2.0 — passed 10-11 of 11 tasks on a single N=1 pass each (Aug 7, 2026). They sit in a separate Smoke-tested group with public raw logs, not promoted to the formal badge.

Four US/EU models — GPT-5.6 Luna Pro, Claude Opus 5, Gemini 3.6 Flash and Grok 4.5 — ran the same 11-task battery through OpenRouter's free tier in an International reference zone. Reference only: nothing there is ranked against the Chinese models.

Kimi K3, Qwen 3.8 and MiniMax H3 are compiled rows: every number on them links to the source we took it from, with the capture date.

V4 Pro still trails the US frontier by about 8 months, per NIST CAISI's independent evaluation.

Six more models sit in a separate tracking queue — we're watching releases and official pages, but there's no verified data to rank them yet, so they get no scores.

How this board is ordered

Row order tracks data verifiability, not raw performance. Three kinds of rows sit on this board. "Tested by us": our formal in-house runs — two models (DeepSeek V4 Pro, V4 Flash), each measured across 22 API calls in August 2026 and re-confirmed, raw logs public. "Smoke-tested": five Chinese models run once through our 11-task battery on Aug 7, 2026 (N=1, single-run) — a smoke check with public raw logs, kept in its own group, not promoted to the formal badge. "Compiled": rows built from vendor docs, third-party evals and community tests, every number linked to its source.

Four US/EU models sit in a separate International reference zone below (reference only, not ranked against the Chinese models).

We don't blend scores from different pipelines into a composite: our task-set results, smoke runs and third-party indexes use different scales and are not directly comparable.

Every number on this board links to the source we took it from, with the capture date.

2
formally tested in-house
5
smoke-tested (N=1)
3
compiled rows
4
international reference
6
in the tracking queue
Where the numbers come from

Rows marked "Tested by us" are our formal in-house runs — two models, 22-call harness, re-confirmed on Aug 7, 2026. Task sets, sampling parameters, model versions, and raw logs are published alongside them.

Rows marked "Smoke-tested" are one-pass runs: each of the five Chinese models went through the same 11-task battery exactly once on Aug 7, 2026 (N=1). Single runs catch gross failures and shape a first impression; they are not statistical tests, and we don't call them that. Raw logs are public for every smoke row (GitHub + Hugging Face), so you can read the exact prompt, output and verdict yourself.

The International reference zone ran the same battery on four US/EU models through OpenRouter's free tier, N=1: same prompts, same judge. Reference only: not ranked against the Chinese models.

Everything else is compiled: from vendor documentation, independent evaluations (Artificial Analysis, LMArena, NIST CAISI), and tests reported by developers and Chinese-language tech media. Every number links to the source we took it from, with the date we captured it. Vendor-reported numbers are labeled vendor-reported. Unverified numbers say "unverified". We'd rather show you a sourced number we didn't run than a score we can't back up.

Rows we can't source yet sit in a separate tracking queue, clearly labeled, with no scores. We add tests as we run them.

Main board

How to read this board

For performance questions, read across, not down. The AA Intelligence Index column puts Kimi K3 at 57.1 (#3 of 189) and V4 Flash at 50; Frontend Code Arena puts K3 at 1679, #1. Those are third-party pipelines with their own scales.

Our in-house scores (11/11, 10/11) come from a different harness and are not convertible to them. The Smoke-tested rows are N=1 single runs — they tell you whether a model handled a task class at all, not how reliably. Row order is not a performance ranking.

Tested by us· 2 rows
our own runs — task sets, sampling parameters, model versions, raw logs published
preview (2026-07)
11/11 tasks(22 API calls, 2026-08-03)
C10 · M10 · Zh10 · ML10 · LC10 · $/perf7
raw logs →
  • SWE-bench Verified 80.6%[1](vendor-reported)
  • ~8 months behind the US frontier[25](per NIST CAISI)
Spec
1.6T total / 49B activated, MIT, 1M ctx, text-only
Price (snapshot)
$0.435/$0.87; cache-hit $0.003625 (75% permanent discount)[3]2026-07
Context
1M token context
raw logs, 2026-08-07 →Updated 2026-08-07
stable (2026-07-31)
10/11 tasks(22 API calls, 2026-08-03)
C10 · M8 · Zh10 · ML10 · LC10 · $/perf10
raw logs →
  • Terminal-Bench 2.1 82.7 / DeepSWE 54.4 / Toolathlon 70.3 / Agent Last Exam 25.2[1](vendor-reported)
  • AA Intelligence Index 50[1](vendor-reported)
  • 8/1 single-day usage: 8 trillion tokens[4](community test / usage report)
Spec
284B total / 13B activated, MIT, 1M ctx, Flash-0731 stable build
Price (snapshot)
$0.14/$0.28; cache-hit $0.0028[2]2026-06
Context
1M token context
raw logs, 2026-08-07 →Updated 2026-08-07
Compiled from public evals· 3 rows
every number links to the source we took it from, with the capture date
C1Kimi K3
2.8T/104B activated, open weights 7/27
  • AA Intelligence Index 57.1, #3 of 189[6](third-party eval)
  • Frontend Code Arena 1679, #1 (preliminary; Fable 5 1631, Sol 1618)[5](third-party eval)
  • SWE Marathon 42.0[5](third-party eval)
  • Terminal-Bench 2.1 88.3[20](media report)
  • HF downloads 180K in 72h; 340+ fine-tunes[9](media report)
  • Specs: 2.8T/104B activated, 1M ctx, native vision[7](official)

K3's numbers are third-party evals and media reports we compile and link, not our in-house runs. Different pipelines: don't read them against our task-set scores.

Spec
2.8T total / 104B activated (896 experts), Modified MIT, 1M ctx, native vision
Price (snapshot)
$3/$15; cache-hit $0.30[8]2026-08-06
Context
1M token context, native visionLargest open-weight model to date; 4-bit weights ≈1.4TB, needs 64+ accelerators.
Sources: [6][5][20][9][7]Updated 2026-08-06
C2Qwen 3.8
2.4T/95B activated MoE, released 8/3
  • Vendor self-report: 'next to Fable 5' (no independent eval yet — labeled vendor-reported)[12](vendor-reported)
  • IDC MarketScape: only Chinese vendor in Leaders quadrant[15](third-party eval)
  • Day-0 integration: OpenRouter, OpenCode, Hermes Agent, Vercel[12](media report)
  • Specs: 2.4T/95B activated, ~1M ctx, released 8/3[12](official)
Spec
2.4T total / 95B activated MoE, ~1M ctx (983,616 tokens), max output 131K, thinking-only mode
Price (snapshot)
CN ¥12/¥36 (cache ¥1.5); intl ~$2/$6 (converted; cache $0.25)[13][14]2026-08-06
Context
~1M token context, first trillion-parameter multimodal releaseWeights promised 'next week' as of 8/6, not yet published; Qwen3.8-27B dense model coming with them.
Sources: [12][15]Updated 2026-08-06
C3MiniMax H3
full-modality generation, H3-Base open 8/3
  • AA video-editing eval #1 globally; Arena image-to-video #1; HF trending #1[16](third-party eval)
  • ¥0.8/sec at 2K — about 1/3 of comparable flagships[17](media report)
  • H3-Base open weights landed 8/3 (33B, 42.5GB, dual checkpoints)[18][19](official)
Spec
Unified text/image/video/audio generation, dual-channel A/V output, up to 15s 2K; H3-Base 33B / 42.5GB open weights (8/3, license excludes US/EU/UK/KR)
Price (snapshot)
¥0.8/sec (2K video, ~1/3 of flagship tier)[17]2026-08-06
Context
Video-first flagship; H3-Base open for self-hostingOpen-weights promise from 7/31 landed 8/3 with H3-Base — region-restricted license, so not fully open.
Sources: [16][17][18][19]Updated 2026-08-06
Smoke-tested · Aug 7, 2026· 5 rows
One 11-task pass per model (N=1) on a single day — a smoke check, not a formal benchmark. These rows are separate from "Tested by us" for a reason: N=1 single runs don't get the formal badge. Raw logs for every row are public (GitHub + Hugging Face).
S1GLM-5.2
glm-5.2 (thinking disabled)
11/11 tasks N=1 · 2026-08-07

All 11 tasks passed with thinking disabled on the direct API.

Battery cost: $0.003323HF →GitHub →
S2Doubao Seed 2.1 Pro
doubao-seed-2-1-pro-260628
11/11 tasks N=1 · 2026-08-07

All 11 tasks passed.

Battery cost: $0.003306HF →GitHub →
S3Step 3.5 Flash
step-3.5-flash (SiliconFlow)
10/11 tasks N=1 · 2026-08-07

Step-3.5-Flash forces reasoning on SiliconFlow (ignores the thinking parameter). ipv4-validator failed: first run exhausted 1024 max_tokens (empty content); the 4096 retry burned ~15.5K reasoning tokens without converging. Behavioral/config trait, not a judged capability miss.

Battery cost: $0.002910HF →GitHub →
S4Hunyuan A13B
hunyuan-a13b (SiliconFlow)
10/11 tasks N=1 · 2026-08-07

Hunyuan-A13B answered "3, -4" for the quadratic task (roots are 4 and -3) — the sign check was skipped (3²-3-12=-6≠0). Plain capability error, reported as-is.

Battery cost: $0.000535HF →GitHub →
S5Ling Flash 2.0
ling-flash-2.0 (SiliconFlow)
11/11 tasks N=1 · 2026-08-07

All 11 tasks passed — cheapest battery in the batch.

Battery cost: $0.000444HF →GitHub →

Data layers: 2 models formally tested in-house (22-call harness, raw logs public) · 5 smoke-tested (11-task battery, N=1, Aug 7 2026, raw logs public) · 4 US/EU models in an International reference zone (reference only, not ranked) · 3 compiled from vendor docs and third-party evals, each number linked · 6 in the tracking queue.

International reference

Why are US/EU models on a Chinese-model board? These rows answer the question this board gets most: "how do these compare to GPT-5 / Claude?" Same 11-task battery, same judge, same day (Aug 7, 2026), N=1 — run through OpenRouter's free tier. Reference only: nothing here is ranked against the Chinese models above, and nothing here changes where a Chinese model sits on the board.

International reference — how to read this zone

These four rows ran on the same 11-task battery and judge as the rows above, through OpenRouter (free tier, N=1). Same prompts, same scoring. What differs: the channel (OpenRouter gateway, so latencies here are zone-comparable only), reasoning behavior (Claude and Gemini arrive with reasoning enabled by default; flagged per row), and free-tier quotas: after the $0.15 grant was spent, OpenRouter rejected further calls with HTTP 402, which is why Claude's retry didn't complete. We report the 10/11 as it ran, not as we'd like it to look. Cost to re-run Claude with reasoning off: about $0.013 (it's on the retest list). This zone is reference, not rank: nothing here is ranked against the Chinese models.

I1GPT-5.6 Luna Pro
openai/gpt-5.6-luna-pro · OpenRouter free tier
11/11 tasks N=1 · 2026-08-07

All 11 passed — cheapest battery in the international zone.

Battery cost: $0.004060 · List $$0.1/$0.6HF →GitHub →
I2Claude Opus 5
anthropic/claude-opus-5 · OpenRouter free tier
10/11 tasks N=1 · 2026-08-07

Claude Opus 5 failed only writing-product-copy, and that's a quota failure, not a capability one. With OpenRouter's default reasoning enabled, the model spent its 256 max_tokens on the reasoning trace and returned empty content. Our retry with thinking disabled (512 tokens) was rejected by the free-tier balance (HTTP 402), so we keep the honest 10/11. Re-run cost is about $0.013 — it's on the retest list.

Battery cost: $0.031625 · List $$5/$25HF →GitHub →
I3Gemini 3.6 Flash
google/gemini-3.6-flash · OpenRouter free tier
11/11 tasks N=1 · 2026-08-07

Gemini 3.6 Flash arrives with reasoning enabled by default. writing-product-copy first ran truncated at 1024 max_tokens (FAIL) and passed on the 2048 retry (first_run_truncated kept in the log).

Battery cost: $0.040614 · List $$1.5/$7.5HF →GitHub →
I4Grok 4.5
x-ai/grok-4.5 · OpenRouter free tier
11/11 tasks N=1 · 2026-08-07

All 11 passed.

Battery cost: $0.019824 · List $$2/$6HF →GitHub →

International rows tested via OpenRouter free tier, N=1, Aug 7 2026 · same 11-task battery and judge as the main board · reference only, not ranked against Chinese models · latencies include gateway overhead (zone-comparable only) · Claude Opus 5's 10/11 is a quota failure (max_tokens spent on reasoning; retry blocked by 402), not a capability failure.

In our tracking queue

These models are on our radar but we don't have enough verified data to rank them yet. We're watching release notes, official pricing pages, and independent evals. Rows graduate to the main board when a data point we can link appears.

Qwen 3.5
Alibaba
Why we're tracking
Previous flagship, superseded by Qwen 3.8.
What's needed to rank it
None — we keep 3.8 as the only Qwen row on the main board.
Kimi K2
Moonshot AI
Why we're tracking
Superseded by K3; useful as a historical comparison anchor.
What's needed to rank it
None — deep-dive comparisons can cite it without a board row.
MiniMax M2
MiniMax
Why we're tracking
Text-line flagship; M2.5 is already API-accessible on SiliconFlow.
What's needed to rank it
M2.5 marked testable — pending a task-card decision on a text-line smoke run (H3 stays Compiled: video flagship, 11-task text battery can't cover it).
ERNIE 5.0
Baidu
Why we're tracking
Material gap closed; but vendor claims and third-party score disagree sharply, needs care.
What's needed to rank it
Collected 2026-08-06 (specs, tiered pricing; third-party score last vs vendor 'top tier' claim) — review required before promotion.
SenseNova U1 / U1.5
SenseTime
Why we're tracking
U-line is an image/multimodal line, not a text flagship; no AA-style third-party score.
What's needed to rank it
Partially collected (flagship pricing; U1.5-Lite-Preview open-sourced 8/3) — stays tracked until content-editor review.
InternLM
Shanghai AI Lab
Why we're tracking
Fully open-source, no API pricing; benchmarks are vendor-reported only.
What's needed to rank it
Partially collected (vendor-reported benchmark only, no third-party) — stays tracked until content-editor review.

Aug 7 battery, task by task

Every cell below is one real API call from the Aug 7, 2026 run — 9 models × 11 tasks, N=1 per cell. The formal "Tested by us" rows (DeepSeek V4 Pro, V4 Flash) are not in this matrix: they were re-confirmed on the same battery but carry the 22-call harness, a different batch.

GLM-5.2
11/11
Doubao Seed 2.1 Pro
11/11
Step 3.5 Flash
10/11
✗a
Hunyuan A13B
10/11
✗b
Ling Flash 2.0
11/11
GPT-5.6 Luna Proref
11/11
Claude Opus 5ref
10/11
✗c
Gemini 3.6 Flashref
11/11
Grok 4.5ref
11/11

Failure notes

  • [a]Step 3.5 FlashIPv4 validator: [retry max_tokens=4096] function not found -- initial FAIL was config-truncation, retry also FAILED
  • [b]Hunyuan A13BMath: quadratic roots with b+c = -13: contains '4'=True, '-3'=False
  • [c]Claude Opus 5Writing: 3-sentence product copy, constrained: sentences=0 (need 3) — first run FAIL, OpenRouter default reasoning consumed 256/256 max_tokens (finish=length), final content empty. Config-fix retries (thinking disabled 512 / 2048 max_tokens) blocked by free-tier credit limit (HTTP 402, remaining balance <=449 output tokens at claude-opus-5 pricing). Recorded as real behavior, not retried on paid credits per budget rule.

Raw logs for every cell: raw logs page · mirrored on GitHub + Hugging Face.

6 of 9 models passed all 11 tasks on the Aug 7 battery; the only US-side failure was a quota issue, not a capability one.

By task dimension (tested rows only)

Coding

Pick: DeepSeek V4 Pro. 11/11 in our battery, and the only model that rejected '1.2.3.04' as an IPv4 address — strictness matters.

ModelScoreNote
DeepSeek V4 Pro10Both code tasks passed first try
DeepSeek V4 Flash-07319.5returned sorted(set(xs)) — the fix we'd write
Compiled rows (not our runs)K3 #1 Frontend Code Arena, GLM-5.2 #1 DesignArena — see main board

Math & Reasoning

Pick: DeepSeek V4 Pro. Passed both math tasks. Flash failed 23×25 with thinking off — but fixed it with one flag (see review P.S.).

ModelScoreNote
DeepSeek V4 Pro10Quadratic roots + consecutive integers
DeepSeek V4 Flash-073181/2 with thinking off; 2/2 with thinking on

Chinese language

Pick: Untested yet. Our first battery was English-only. Chinese-language tasks are queued for the next retest — a Chinese-model site without a Chinese test would be unserious.

ModelScoreNote
QueuedNext battery: zh task set v1

Multilingual

Pick: DeepSeek V4 Flash-0731. Both models handled English extraction/writing tasks at parity; Flash did it at 1/3 the price.

ModelScoreNote
DeepSeek V4 Flash-073110English tasks at $0.28/1M out
DeepSeek V4 Pro10Same pass rate, 3x the price

Long context

Pick: Both — with a caveat. Needle-in-a-426-token test passed on both. The full 1M token context is on our queue; at $0.14/1M in, DeepSeek is 18x cheaper than GPT-5.4 for a 300K doc.

ModelScoreNote
DeepSeek V4 Flash-073110426-token needle: PASS
DeepSeek V4 Pro10426-token needle: PASS

Cost efficiency

Pick: DeepSeek V4 Flash-0731. Our whole 22-call battery cost $0.0013. Artificial Analysis puts Flash-0731 at ~$0.03 per index task — the cheapest board-wide.

ModelScoreNote
DeepSeek V4 Flash-073110$0.14/$0.28 per 1M tokens
DeepSeek V4 Pro7$0.435/$0.87 — still cheap, 3.1x Flash
GPT-5.4 (reference)3$2.50/$15 — 54x Flash output price

Data sources

Every compiled number on this board resolves to one of the entries below — URL, capture date, and nature. Vendor-reported numbers are labeled as such; we don't launder vendor claims into third-party ones.

  1. [1] https://news.qq.com/rain/a/20260801A048U500captured 2026-08-06 · vendor-reported — DeepSeek V4 9-item agent benchmark release (Tencent News recap)
  2. [2] https://developer.puter.com/tutorials/deepseek-api-pricing/captured 2026-08-06 · official — DeepSeek API pricing page (Puter tutorial mirror)
  3. [3] https://www.36kr.com/p/3919319636290946captured 2026-08-06 · official — DeepSeek V4 Pro & GLM-5.2 pricing (36Kr)
  4. [4] https://www.oschina.net/news/486803captured 2026-08-06 · community — V4 Flash 8T-token single-day usage report (OSChina)
  5. [5] https://www.jiemian.com/article/14786633.htmlcaptured 2026-08-06 · third-party eval — K3 Frontend Code Arena 1679 #1, SWE Marathon 42.0 (Jiemian, preliminary)
  6. [6] https://macgpu.com/zh/blog/2026-kimi-k3-kaiyuan-quanzhong-fabu.htmlcaptured 2026-08-06 · third-party eval — K3 AA Intelligence Index 57.1, #3/189 (MACGPU)
  7. [7] https://www.ithome.com/0/982/259.htmcaptured 2026-08-06 · official — K3 specs: 2.8T/104B, 1M ctx, native vision (ITHome)
  8. [8] https://ollama.com/library/kimi-k3captured 2026-08-06 · official — K3 pricing $3/$15, cache $0.30 (Ollama model page)
  9. [9] https://blog.csdn.net/xyghehehehe/article/details/163353310captured 2026-08-06 · media report — K3 HF 180K downloads in 72h, 340+ fine-tunes (CSDN)
  10. [10] https://docs.bigmodel.cn/cn/guide/models/text/glm-5.2captured 2026-08-06 · official — GLM-5.2 specs, AA 51 open-source SOTA, FrontierSWE (Zhipu docs)
  11. [11] https://juejin.cn/post/7654122741078671402captured 2026-08-06 · community — GLM-5.2 DesignArena 1360 ELO #1, Code Arena #1 (Juejin)
  12. [12] https://news.qq.com/rain/a/20260804A0ACRT00captured 2026-08-06 · official — Qwen 3.8 release: 2.4T/95B, 8/3 launch, weights 'next week' (Tencent News)
  13. [13] https://www.datalearner.com/ai-models/pretrained-models/qwen3-8-maxcaptured 2026-08-06 · official — Qwen 3.8 CN pricing ¥12/¥36, cache ¥1.5 (DataLearner)
  14. [14] https://www.aipricedb.com/ai-news/qwen-launches-qwen3-8-max-with-2-4-trillion-parameters-c98759b1captured 2026-08-06 · official (converted) — Qwen 3.8 intl ~$2/$6, cache $0.25 — converted from vendor percentages (AIPriceDB)
  15. [15] https://www.leiphone.com/category/industrynews/1gS0E0CcUDxEgbQR.htmlcaptured 2026-08-06 · third-party eval — IDC MarketScape: Qwen only Chinese vendor in Leaders (Leiphone)
  16. [16] https://www.leiphone.com/category/industrynews/aiUMBoeUYbi8fX4x.htmlcaptured 2026-08-06 · third-party eval — MiniMax H3 AA video-editing #1, Arena image-to-video #1, HF trending #1 (Leiphone)
  17. [17] https://www.163.com/dy/article/L3623H0E05198NMR.htmlcaptured 2026-08-06 · media report — H3 ¥0.8/sec at 2K, ~1/3 of flagship (NetEase)
  18. [18] https://www.atlascloud.ai/zh/blog/guides/minimax-h3-open-source-weightscaptured 2026-08-06 · official — H3-Base 33B / 42.5GB, dual checkpoints, license excludes US/EU/UK/KR (AtlasCloud + HuggingFace MiniMaxAI/MiniMax-H3)
  19. [19] https://www.ithome.com/0/984/379.htmcaptured 2026-08-06 · media report — H3 open-weights landing 8/3 (ITHome)
  20. [20] https://explore.n1n.ai/zh/blog/kimi-k3-moonshot-ai-kaiyuan-moxing-2026-07-17captured 2026-08-06 · media report — K3 Terminal-Bench 2.1 88.3 (n1n.ai blog)
  21. [25] /blog/deepseek-v4-review-2026captured 2026-08-03 · our review — In-house 22-call battery + NIST CAISI cross-check (site review)

Methodology

  • Three kinds of rows sit on this board. "Tested by us": our formal in-house runs — two models (DeepSeek V4 Pro, V4 Flash), each measured across 22 API calls in August 2026 and re-confirmed, raw logs public. "Smoke-tested": five Chinese models run once through our 11-task battery on Aug 7, 2026 (N=1, single-run) — a smoke check with public raw logs, kept in its own group, not promoted to the formal badge. "Compiled": rows built from vendor docs, third-party evals and community tests, every number linked to its source. Four US/EU models sit in a separate International reference zone below (reference only, not ranked against the Chinese models). We don't blend scores from different pipelines into a composite: our task-set results, smoke runs and third-party indexes use different scales and are not directly comparable.
  • Formal in-house runs (T rows): 11 tasks written before the run — IPv4 validator, bug fix, 2 math, syllogism, 2 knowledge, constrained writing, JSON extraction, JSON summary, long-doc needle. Battery run twice (22 API calls per model); the single failure reproduced.
  • Smoke runs (S rows, Aug 7, 2026): same 11-task battery, exactly one pass per model (N=1). Sampling: official vendor APIs (SiliconFlow / Volcano / Zhipu / MiniMax direct), temperature default, max_tokens 64-512 per task, thinking settings disclosed per row (GLM/Seed disabled; SiliconFlow trio at vendor default; Step forces reasoning). Single runs catch gross failures; they are not statistical tests.
  • International zone (I rows): the same battery through OpenRouter free tier, N=1, same judge. Latencies include gateway overhead and are zone-comparable only. Claude and Gemini arrive with reasoning enabled by default; the Claude retry was blocked by a free-tier balance (HTTP 402) — reported as it ran.
  • Versions snapshotted: deepseek-v4-flash serves Flash-0731 (stable, 2026-07-31); deepseek-v4-pro is the July preview; smoke/international rows carry their exact API model IDs.
  • Compiled rows (C): each quantitative data point carries a footnote [n] resolving to a URL, capture date and nature (official / vendor-reported / third-party eval / media report / community). Vendor-reported numbers are labeled as such; unverified numbers say unverified.
  • Cross-check: our results × vendor tables (arXiv:2606.19348) × NIST CAISI independent eval (2026-05-01). Where they disagree, we say so.
  • Limits: no full-length 1M context test yet, no self-host run, no multimodal (V4 is text-only). Tracking rows get no scores by design.

Full task prompts, raw responses and rubrics for our in-house runs are published with the review: DeepSeek V4 review 2026. Every tested, smoke-tested and reference row links to its raw logs (GitHub + Hugging Face) from the tables above, and the raw logs page lists all 11 models with direct links.

PK: vendor-reported benchmarks (cross-checked, not trusted)

Every number below is vendor-reported (DeepSeek technical report arXiv:2606.19348 and competitor release materials). "n/r" = not reported. Our independent cross-check is in the CAISI section of the review.

BenchmarkV4-ProV4-FlashGPT-5.4Opus 4.6Gemini 3.1 Pro
SWE-bench Verified80.679.077.280.880.6
SWE-bench Pro55.452.6n/rn/rn/r
LiveCodeBench Pass@193.591.6n/r88.891.7
Terminal-Bench 2.067.9n/r75.165.468.5
Codeforces rating320630523168n/r3052
MMLU-Pro87.586.2n/rn/rn/r
GPQA Diamond90.188.1n/rn/rn/r
Humanity's Last Exam37.7n/r39.840.044.4
SimpleQA-Verified57.9n/rn/rn/r75.6

PK: pricing snapshot (2026-08-06)

API prices in USD per 1M tokens (input/output), except where noted. Prices change — every row carries its snapshot date on the main board.

ModelInput $/1MOutput $/1MNote
DeepSeek V4-Flash$0.14$0.28Stable build (Flash-0731)
DeepSeek V4-Pro$0.435$0.8775% cut July 2026
Kimi K3$3$15Cache-hit $0.30
GLM-5.2$1.4$4.4Coding Plan from $18/mo
Qwen 3.8 (intl)$2$6Converted; CN ¥12/¥36
GPT-5.4$2.5$15Long context >272K doubles input
Claude Opus 4.6$5$25
Gemini 3.1 Pro$2$12

Deep dives

FAQ

Did you really test these models yourself?

Two of them formally: DeepSeek V4 Pro and V4 Flash, run on a 22-call harness in early August 2026 and re-confirmed on Aug 7. Task sets, versions, and raw logs are published. Five more Chinese models — GLM-5.2, Doubao Seed 2.1, Step 3.5 Flash, Hunyuan A13B, Ling Flash 2.0 — got a one-pass smoke battery on Aug 7, 2026 (N=1, raw logs public). We keep those in a separate "Smoke-tested" group rather than calling them formal tests, because a single run isn't one. Four US/EU models ran the same battery via OpenRouter in an International reference zone — reference only. The remaining rows are compiled, every number linked to its source. Models without verifiable data sit in the tracking queue, with no scores.

Why is the Smoke-tested group separate from Tested by us?

Because N=1 is a smoke check, not a benchmark. Each model ran the 11-task battery exactly once on Aug 7, 2026. A single pass tells you whether the model handles the task class at all and where it visibly fails; it can't measure reliability or rank models against each other. Formal "Tested by us" rows require a 22-call harness plus a re-confirmation pass. Smoke rows stay in their own group until they earn the formal badge.

Can I compare the International reference zone with the main board?

Same prompts, same judge, same 11-task battery: that's the point of the zone. Score comparisons between the zone and the Smoke-tested rows work, with N=1 caveats; latency comparisons don't (direct API vs OpenRouter gateway overhead). The zone is reference, not rank: nothing in it changes where a Chinese model sits on the board.

Are Chinese AI models safe to use?

Safety is a supply-chain question, not a yes/no. V4 is MIT-licensed open weight, so you can self-host and audit it. For API use, weigh data residency and the lack of western SLAs — our review covers the trade-offs, and we don't whitewash the concerns. Note that MiniMax H3's open weights exclude the US, EU, UK and South Korea, which matters if you plan to self-host it.

How do Chinese models compare to GPT-5?

Per NIST CAISI's independent May 2026 evaluation, DeepSeek V4 Pro sits at roughly GPT-5 level — about 8 months behind the US frontier. It wins decisively on capability per dollar, and loses on novel reasoning and knowledge recall. We also ran both sides on the same 11-task battery — see the International reference zone below. The compiled rows cite their own sources with dates.

What is the best Chinese AI model for coding?

Two pipelines answer this, and we keep them separate. On our formal harness, DeepSeek V4 Pro passed 11/11 tasks and V4 Flash-0731 passed 10/11. On our N=1 smoke battery (Aug 7, 2026), GLM-5.2, Doubao Seed 2.1 and Ling Flash 2.0 also passed 11/11 — single runs, not formal scores, so we don't rank from them. On third-party coding evals we compile, Kimi K3 tops Frontend Code Arena at 1679 (preliminary) and GLM-5.2 leads DesignArena at 1360 ELO — different pipelines from our task sets, each number linked and labeled. For agentic coding on a budget, Flash-0731 at $0.14/$0.28 per 1M tokens is the value pick on the board.

Is DeepSeek better than Qwen?

We tested DeepSeek V4 (Pro and Flash) in-house; Qwen 3.8 is on the board as a compiled row because we haven't run it yet. Its numbers come from Alibaba's own reporting (labeled vendor-reported) and IDC's MarketScape, both linked with capture dates. A direct head-to-head needs our own battery — Qwen joins the in-house queue.

How much do Chinese AI models cost?

Snapshot 2026-08-06: DeepSeek V4-Flash $0.14/$0.28 per 1M tokens (in/out), V4-Pro $0.435/$0.87; Kimi K3 $3/$15; GLM-5.2 $1.4/$4.4; Qwen 3.8 ¥12/¥36 in China and about $2/$6 internationally (converted). That's 1/10th to 1/54th of US frontier output pricing. Prices change; every price row carries its snapshot date.

Do Chinese models work well in English?

Our whole first battery ran in English: 21/22 task passes across both formally tested models, including strict JSON extraction and constrained product copy. English is not a weakness for V4 — knowledge recall is (SimpleQA 57.9 vs Gemini 3.1 Pro's 75.6, per our cross-check).

Update history

Last updated . Next full sweep: 2026-09-06. Price snapshots refresh weekly; in-house retests run quarterly or after a major release. International zone retests run quarterly or after a US/EU major release, subject to OpenRouter free-tier balance. Smoke-tested rows get a fresh N=1 pass on the same schedule; two passes in a row start the conversation about a formal badge, never one. Every change lands in the update history below.

releaseDeepSeek V4 Pro, V4 Flash-0731

v1: initial in-house battery — 11 tasks, 6 dimensions, 22 API calls; raw responses and rubrics published in the review.

Source: in-house run (site review)

structuralAll

v2 rebuild: main board reworked to 6 rows (2 tested in-house + 4 compiled with per-number source links); separate 10-row tracking queue with no scores; data-sources statement, layered FAQ and update-history structure added; Qwen 3.8 price filled in (CN ¥12/¥36, intl ~$2/$6); MiniMax H3 open-weights status updated (H3-Base landed 8/3, region-restricted license).

Source: 20260806-hotspot R1-R3 + 20260806-topup (23 records) + leaderboard spec v2

structuralAll

Plan-A revision: medal ranks (🥇🥈🥉) replaced by two evidence groups (Tested by us / Compiled from public evals) with in-group numbering; ordering methodology statement moved to the top; "How to read this board" added; K3 row gained Terminal-Bench 2.1 88.3 [20]; V4 Flash row gained AA Intelligence Index 50 [1].

Source: leaderboard spec v2 + plan-A revision brief (2026-08-07)

expansionGLM-5.2, Doubao Seed 2.1, Step 3.5 Flash, Hunyuan A13B, Ling Flash 2.0, GPT-5.6 Luna Pro, Claude Opus 5, Gemini 3.6 Flash, Grok 4.5

New "Smoke-tested · Aug 7, 2026" group — GLM-5.2, Doubao Seed 2.1, Step 3.5 Flash, Hunyuan A13B and Ling Flash 2.0 passed 10-11 of 11 tasks on a single N=1 run each; raw logs public (GitHub + Hugging Face). New "International reference" zone — GPT-5.6 Luna Pro, Claude Opus 5, Gemini 3.6 Flash and Grok 4.5, same battery via OpenRouter free tier, reference only, not ranked against Chinese models. Tested by us (2 rows) and Compiled (3 rows) unchanged. Raw logs page expanded to 11 models with direct links.

Source: benchmark runs 2026-08-07, data/raw-logs/ + GitHub EliChen-ai/china-ai-bench + HF EliChen-ai/china-ai-bench-benchmarks