Skip to content
China's Best Open AI Model Scores 46 on an Independent Index Against 58 for the Top US Model, and Four Sources Put the Lag Between 2 and 8 Months
Artificial Intelligence

China's Best Open AI Model Scores 46 on an Independent Index Against 58 for the Top US Model, and Four Sources Put the Lag Between 2 and 8 Months

Illustration by tuput

English

Xiaomi's MiMo-V2.6-Pro leads open-weight models on Artificial Analysis's index with 46, ahead of Z.ai's GLM-5.3 and Moonshot's Kimi K3. Anthropic's Claude Opus 5.5 scores 58. DeepSeek's own report says it trails the US frontier by 3 to 6 months, and the US government's testing centre says about 8.

· · 7 min read

On 8 October 2026 the highest-placed Chinese model on Artificial Analysis’s Intelligence Index is Xiaomi’s MiMo-V2.6-Pro, which scores 46 and sits 18th overall. Anthropic’s Claude Opus 5.5 leads the same table at 58, and OpenAI’s GPT-6 Astra and Google’s Gemini 4 Argon score 53. The index combines ten evaluations of coding, agents, knowledge work and science into one score, according to Trending Topics.

Four sources have tried to turn that distance into months, and their answers run from 2 to 8. They use different tests and different US models as the yardstick, which is most of the reason they disagree.

The newest releases, in the labs’ own numbers

All of these are mixture-of-experts models. Each token is routed to a small set of specialist sub-networks, so the “active” parameter count, the adjustable numbers actually used per token, is far smaller than the total. A token is a word or word fragment.

ModelTotal / active parametersContext windowLicence
DeepSeek V4-Pro1.6 trillion / 49 billion1 million tokensMIT
DeepSeek V4.1 Flash552 billion / 16 billion1 million tokensMIT
Moonshot Kimi K32.8 trillion / 104 billion1 million tokensKimi K3 License
Alibaba Qwen3.8-2.4T-A95B2.4 trillion / 95 billion262,144 tokens, extendable to 1,010,000Qwen3.8-Max License
Z.ai GLM-5.3753 billion / 40 billion1 million tokensGLM-5.3 License
Xiaomi MiMo-V2.6-Pro1.0 trillion / 42 billion1 million tokensMIT

The DeepSeek V4 technical report, submitted to arXiv on 26 April 2026, describes a preview. It says V4-Pro was pre-trained on 33 trillion tokens and the smaller V4-Flash on 32 trillion. At a one-million-token context, V4-Pro needs 27% of the per-token computation and 10% of the memory cache of DeepSeek’s earlier V3.2, the report says. The V4.1 Flash and MiMo rows draw on Artificial Analysis’s table and Arena’s licence labels.

Kimi K3’s model card gives the 16-of-896 expert layout but no training token count and no training hardware. Qwen’s open checkpoint is text-only, and its card says thinking mode cannot be turned off. The hosted Qwen3.8-Max adds image input and a one-million-token default. GLM-5.3 reuses the base model of GLM-5.2, and Z.ai’s card says the gains come from post-training, the stage where a finished model is tuned with feedback.

Benchmark numbers on these cards are the labs’ own. DeepSeek’s card lists 80.6 for V4-Pro-Max on the coding test SWE-bench Verified. NIST’s CAISI measured 74 on the same test and attributes the difference to system prompts, scaffolding and token budgets.

What outside testers measured

Artificial Analysis, on version 4.3.2 of its index, lists these open-weight scores: MiMo-V2.6-Pro 46, GLM-5.3 45, Kimi K3 44, GLM-5.3-Flash 42, DeepSeek V4.1 Flash 39, Qwen3.8 27B (a 27-billion-parameter model) 34 and MiniMax-M3 29. Trending Topics reports 39.9 for the full Qwen3.8 2.4T. The best American open entries on the Artificial Analysis table are Thinking Machines’ Inkling at 25 and Nvidia’s Nemotron 3 Ultra at 23.

France’s Mistral Large 4 preview scored 38.4, which would rank eighth among open models, as tuput reported in the Mistral Large 4 article. Cost differs sharply. Trending Topics gives $0.13 per index task for MiMo-V2.6-Pro and about $2 each for GLM-5.3, Kimi K3 and Qwen3.8.

On Arena’s text leaderboard, Kimi K3 (max) is listed 16th at 1488, against 1525 for Google’s Gemini 4 Argon, which Arena marks as preliminary, and 1507 for Claude Opus 5.5. MiMo-V2.6-Pro is 26th, GLM-5.3 27th and DeepSeek V4.1 Flash 39th. None of the Indian models funded under the IndiaAI Mission appears in the Artificial Analysis open-weight table tuput read. The programme’s budget, GPUs and funded builders are set out in the IndiaAI Mission article.

Open weights, and what the licences require

“Open-weight” means anyone can download the trained numbers, and the terms of use differ by lab. DeepSeek and Xiaomi use the MIT licence, which has almost no conditions. Trending Topics reports that Artificial Analysis classes GLM-5.3, Kimi K3 and the large Qwen3.8 as restricted for commercial use, although the licence texts tuput read set narrower limits.

Kimi K3’s licence requires a company running a hosted-model business with more than $20 million in revenue over any 12 months to sign a separate agreement with Moonshot. A product with over 100 million monthly users, or $20 million in monthly revenue, must display “Kimi K3” on its interface. Internal use is exempt.

The Qwen3.8-Max licence has the same display rule and sets the separate-licence threshold at $50 million over 12 months for hosted-model or AI work-assistant businesses. Z.ai’s GLM-5.3 licence allows commercial use, with a Z.ai security review only for hosted-model businesses whose combined revenue exceeds $10 billion.

The training-cost claims and what they leave out

The figure most often quoted is $5.576 million for DeepSeek-V3. The V3 paper arrives at it from 2.788 million H800 GPU hours at an assumed rental price of $2 an hour. The paper states that the total covers only the official training run and excludes prior research and ablation experiments, the small tests used to pick architecture, algorithms and data.

The V4 report states token counts but, in the text tuput searched, no dollar cost or GPU-hour total. Kimi K3’s model card states no token count and no cost.

Which chips the labs say they use

DeepSeek’s report says its fused expert-communication kernel was validated on both Nvidia GPUs and Huawei Ascend NPUs, with a speed-up of 1.50 to 1.73 times on general inference and up to 1.96 times on latency-sensitive work such as reinforcement-learning rollouts. It does not say which chips trained the model.

ChinaTalk’s Irene Zhang, Aqib Zakaria and Jordan Schneider wrote on 27 April that V4 was “still trained on Nvidia chips”, which is their reading of the evidence and not a DeepSeek statement. The same piece quotes a footnote in DeepSeek’s launch announcement: “Due to constraints on high-end compute, V4-Pro’s service throughput is currently limited.” It adds that DeepSeek expects Pro’s price to fall once Huawei’s Ascend 950 supernodes ship in volume in the second half of 2026.

The GLM-5.3 card says deployment on Ascend is supported through vLLM-Ascend, xLLM and SGLang, and gives no training hardware. Kimi K3’s card says some of its evaluations ran on Nvidia H20 GPUs.

On the US side, The Next Web reported on 19 August, citing the Financial Times, that ByteDance and Tencent had each received about 10,000 Nvidia H200 chips. Beijing has told companies to keep them outside the mainland, and shipments go through Hong Kong. In July, the Next Web says, Commerce Under Secretary Jeffrey Kessler told Congress that “very few” had arrived. Diversion of restricted servers is the subject of tuput’s report on the Justice Department’s indictment over servers sent to China.

Four estimates of the gap

DeepSeek’s report says V4-Pro-Max beats GPT-5.2 and Gemini-3.0-Pro on standard reasoning tests but falls marginally short of GPT-5.4 and Gemini-3.1-Pro. It reads that as a lag of roughly 3 to 6 months behind the frontier.

NIST’s Center for AI Standards and Innovation tested V4 Pro in April and published on 1 May that its capabilities lag the frontier by about 8 months. On CAISI’s nine benchmarks, including two held back from the public, V4 Pro performed similarly to GPT-5, released about 8 months earlier. On ARC-AGI-2’s semi-private set it scored 46% against 79% for GPT-5.5. CAISI says its lag figure comes from a statistical model that approximates capability, not a direct measurement.

Nathan Lambert wrote on 21 September, in a piece expanded from testimony prepared for a Congressional briefing, that Chinese open-weight models are roughly 2 to 5 months behind the closed American frontier. American open models, he says, are 6 to 9 months behind OpenAI and Anthropic. He estimates that distillation, training a model on another model’s outputs, narrows the Chinese gap by 1 to 2 months.

Epoch AI’s Jack Edwards and Luke Emberson reported on 29 May that since January 2026 the most capable open-weight models have trailed the closed frontier by an average of four months on the Epoch Capabilities Index, or six months under a stricter test, an average gap of 8 index points. That covers open models generally, not only Chinese ones. Epoch adds that open models tend to score worse on private benchmarks, and that its estimate may understate the true gap.

Share
Copied!

Sources & further reading

  1. arXiv: DeepSeek-V4, Towards Highly Efficient Million-Token Context Intelligence (technical report, submitted 26 April 2026)
  2. Hugging Face: DeepSeek-V4-Pro model card
  3. arXiv: DeepSeek-V3 Technical Report
  4. Hugging Face: Kimi K3 model card (Moonshot AI)
  5. Hugging Face: Kimi K3 licence text
  6. Hugging Face: Qwen3.8-2.4T-A95B model card (Alibaba)
  7. Hugging Face: Qwen3.8-Max licence text
  8. Hugging Face: GLM-5.3 model card (Z.ai)
  9. Hugging Face: GLM-5.3 licence text
  10. Artificial Analysis: top open-weights models
  11. Artificial Analysis: leaderboard of all models by Intelligence Index
  12. Trending Topics: Mistral Large 4 on Artificial Analysis, Eighth Among Open Models (6 October 2026)
  13. Arena: text leaderboard
  14. NIST CAISI: Evaluation of DeepSeek V4 Pro (1 May 2026)
  15. Epoch AI: Open models lag state-of-the-art closed models by 4 months (29 May 2026)
  16. Interconnects: The current balance of power in open models (Nathan Lambert, 21 September 2026)
  17. ChinaTalk: DeepSeek V4 (27 April 2026)
  18. The Next Web: Nvidia's H200 finally reaches Chinese buyers (19 August 2026)

Researched and written with the help of AI tools and edited for accuracy. Provided for general information and discussion only, not professional advice. See our editorial standards and disclaimer. Spotted an error? Tell us.

#chinese ai#deepseek v4#kimi k3#qwen3.8#glm-5.3#open-weight models#artificial analysis#export controls

Enjoyed this? Get the next one.

One good read at a time, straight to your inbox. No spam, unsubscribe anytime.

More in Artificial Intelligence
Mistral's 1-Trillion-Parameter Le Chonk Scored 38 on an Independent AI Index, Eighth Among Open Models, Before Its Weights Have Even Shipped
France's Mistral has put a trillion-parameter model online and says the downloadable weights follow this month. Outside testers rank it the best open model outside China and eighth among open models, and Mistral's own figures for its size do not all agree.
US Prosecutors Say a California Businessman Routed More Than $300 Million of AI Servers to China via Malaysia and Singapore, an Allegation Not Yet Tested in Court
The owner of a City of Industry computer company is accused of moving servers full of restricted chips to Chinese buyers. The emails, the money and the rules behind the charge are public, and so are the gaps in the wider estimates.
Scale's SWE-Bench Pro Board Puts the Best Coding Agent at 61.5%, Maintainers Reject Many Fixes That Pass the Tests, and METR Says Its Latest Trial Cannot Show How Much Time Agents Save
Benchmarks say agents are strong. Maintainers, a randomised trial and a pile of incident reports say the picture is narrower. Here is what each number actually measures.
← all articles