Illustration by tuput
English
Xiaomi's MiMo-V2.6-Pro leads open-weight models on Artificial Analysis's index with 46, ahead of Z.ai's GLM-5.3 and Moonshot's Kimi K3. Anthropic's Claude Opus 5.5 scores 58. DeepSeek's own report says it trails the US frontier by 3 to 6 months, and the US government's testing centre says about 8.
On 8 October 2026 the highest-placed Chinese model on Artificial Analysis’s Intelligence Index is Xiaomi’s MiMo-V2.6-Pro, which scores 46 and sits 18th overall. Anthropic’s Claude Opus 5.5 leads the same table at 58, and OpenAI’s GPT-6 Astra and Google’s Gemini 4 Argon score 53. The index combines ten evaluations of coding, agents, knowledge work and science into one score, according to Trending Topics.
Four sources have tried to turn that distance into months, and their answers run from 2 to 8. They use different tests and different US models as the yardstick, which is most of the reason they disagree.
The newest releases, in the labs’ own numbers
All of these are mixture-of-experts models. Each token is routed to a small set of specialist sub-networks, so the “active” parameter count, the adjustable numbers actually used per token, is far smaller than the total. A token is a word or word fragment.
| Model | Total / active parameters | Context window | Licence |
|---|---|---|---|
| DeepSeek V4-Pro | 1.6 trillion / 49 billion | 1 million tokens | MIT |
| DeepSeek V4.1 Flash | 552 billion / 16 billion | 1 million tokens | MIT |
| Moonshot Kimi K3 | 2.8 trillion / 104 billion | 1 million tokens | Kimi K3 License |
| Alibaba Qwen3.8-2.4T-A95B | 2.4 trillion / 95 billion | 262,144 tokens, extendable to 1,010,000 | Qwen3.8-Max License |
| Z.ai GLM-5.3 | 753 billion / 40 billion | 1 million tokens | GLM-5.3 License |
| Xiaomi MiMo-V2.6-Pro | 1.0 trillion / 42 billion | 1 million tokens | MIT |
The DeepSeek V4 technical report, submitted to arXiv on 26 April 2026, describes a preview. It says V4-Pro was pre-trained on 33 trillion tokens and the smaller V4-Flash on 32 trillion. At a one-million-token context, V4-Pro needs 27% of the per-token computation and 10% of the memory cache of DeepSeek’s earlier V3.2, the report says. The V4.1 Flash and MiMo rows draw on Artificial Analysis’s table and Arena’s licence labels.
Kimi K3’s model card gives the 16-of-896 expert layout but no training token count and no training hardware. Qwen’s open checkpoint is text-only, and its card says thinking mode cannot be turned off. The hosted Qwen3.8-Max adds image input and a one-million-token default. GLM-5.3 reuses the base model of GLM-5.2, and Z.ai’s card says the gains come from post-training, the stage where a finished model is tuned with feedback.
Benchmark numbers on these cards are the labs’ own. DeepSeek’s card lists 80.6 for V4-Pro-Max on the coding test SWE-bench Verified. NIST’s CAISI measured 74 on the same test and attributes the difference to system prompts, scaffolding and token budgets.
What outside testers measured
Artificial Analysis, on version 4.3.2 of its index, lists these open-weight scores: MiMo-V2.6-Pro 46, GLM-5.3 45, Kimi K3 44, GLM-5.3-Flash 42, DeepSeek V4.1 Flash 39, Qwen3.8 27B (a 27-billion-parameter model) 34 and MiniMax-M3 29. Trending Topics reports 39.9 for the full Qwen3.8 2.4T. The best American open entries on the Artificial Analysis table are Thinking Machines’ Inkling at 25 and Nvidia’s Nemotron 3 Ultra at 23.
France’s Mistral Large 4 preview scored 38.4, which would rank eighth among open models, as tuput reported in the Mistral Large 4 article. Cost differs sharply. Trending Topics gives $0.13 per index task for MiMo-V2.6-Pro and about $2 each for GLM-5.3, Kimi K3 and Qwen3.8.
On Arena’s text leaderboard, Kimi K3 (max) is listed 16th at 1488, against 1525 for Google’s Gemini 4 Argon, which Arena marks as preliminary, and 1507 for Claude Opus 5.5. MiMo-V2.6-Pro is 26th, GLM-5.3 27th and DeepSeek V4.1 Flash 39th. None of the Indian models funded under the IndiaAI Mission appears in the Artificial Analysis open-weight table tuput read. The programme’s budget, GPUs and funded builders are set out in the IndiaAI Mission article.
Open weights, and what the licences require
“Open-weight” means anyone can download the trained numbers, and the terms of use differ by lab. DeepSeek and Xiaomi use the MIT licence, which has almost no conditions. Trending Topics reports that Artificial Analysis classes GLM-5.3, Kimi K3 and the large Qwen3.8 as restricted for commercial use, although the licence texts tuput read set narrower limits.
Kimi K3’s licence requires a company running a hosted-model business with more than $20 million in revenue over any 12 months to sign a separate agreement with Moonshot. A product with over 100 million monthly users, or $20 million in monthly revenue, must display “Kimi K3” on its interface. Internal use is exempt.
The Qwen3.8-Max licence has the same display rule and sets the separate-licence threshold at $50 million over 12 months for hosted-model or AI work-assistant businesses. Z.ai’s GLM-5.3 licence allows commercial use, with a Z.ai security review only for hosted-model businesses whose combined revenue exceeds $10 billion.
The training-cost claims and what they leave out
The figure most often quoted is $5.576 million for DeepSeek-V3. The V3 paper arrives at it from 2.788 million H800 GPU hours at an assumed rental price of $2 an hour. The paper states that the total covers only the official training run and excludes prior research and ablation experiments, the small tests used to pick architecture, algorithms and data.
The V4 report states token counts but, in the text tuput searched, no dollar cost or GPU-hour total. Kimi K3’s model card states no token count and no cost.
Which chips the labs say they use
DeepSeek’s report says its fused expert-communication kernel was validated on both Nvidia GPUs and Huawei Ascend NPUs, with a speed-up of 1.50 to 1.73 times on general inference and up to 1.96 times on latency-sensitive work such as reinforcement-learning rollouts. It does not say which chips trained the model.
ChinaTalk’s Irene Zhang, Aqib Zakaria and Jordan Schneider wrote on 27 April that V4 was “still trained on Nvidia chips”, which is their reading of the evidence and not a DeepSeek statement. The same piece quotes a footnote in DeepSeek’s launch announcement: “Due to constraints on high-end compute, V4-Pro’s service throughput is currently limited.” It adds that DeepSeek expects Pro’s price to fall once Huawei’s Ascend 950 supernodes ship in volume in the second half of 2026.
The GLM-5.3 card says deployment on Ascend is supported through vLLM-Ascend, xLLM and SGLang, and gives no training hardware. Kimi K3’s card says some of its evaluations ran on Nvidia H20 GPUs.
On the US side, The Next Web reported on 19 August, citing the Financial Times, that ByteDance and Tencent had each received about 10,000 Nvidia H200 chips. Beijing has told companies to keep them outside the mainland, and shipments go through Hong Kong. In July, the Next Web says, Commerce Under Secretary Jeffrey Kessler told Congress that “very few” had arrived. Diversion of restricted servers is the subject of tuput’s report on the Justice Department’s indictment over servers sent to China.
Four estimates of the gap
DeepSeek’s report says V4-Pro-Max beats GPT-5.2 and Gemini-3.0-Pro on standard reasoning tests but falls marginally short of GPT-5.4 and Gemini-3.1-Pro. It reads that as a lag of roughly 3 to 6 months behind the frontier.
NIST’s Center for AI Standards and Innovation tested V4 Pro in April and published on 1 May that its capabilities lag the frontier by about 8 months. On CAISI’s nine benchmarks, including two held back from the public, V4 Pro performed similarly to GPT-5, released about 8 months earlier. On ARC-AGI-2’s semi-private set it scored 46% against 79% for GPT-5.5. CAISI says its lag figure comes from a statistical model that approximates capability, not a direct measurement.
Nathan Lambert wrote on 21 September, in a piece expanded from testimony prepared for a Congressional briefing, that Chinese open-weight models are roughly 2 to 5 months behind the closed American frontier. American open models, he says, are 6 to 9 months behind OpenAI and Anthropic. He estimates that distillation, training a model on another model’s outputs, narrows the Chinese gap by 1 to 2 months.
Epoch AI’s Jack Edwards and Luke Emberson reported on 29 May that since January 2026 the most capable open-weight models have trailed the closed frontier by an average of four months on the Epoch Capabilities Index, or six months under a stricter test, an average gap of 8 index points. That covers open models generally, not only Chinese ones. Epoch adds that open models tend to score worse on private benchmarks, and that its estimate may understate the true gap.
Sources & further reading
- arXiv: DeepSeek-V4, Towards Highly Efficient Million-Token Context Intelligence (technical report, submitted 26 April 2026)
- Hugging Face: DeepSeek-V4-Pro model card
- arXiv: DeepSeek-V3 Technical Report
- Hugging Face: Kimi K3 model card (Moonshot AI)
- Hugging Face: Kimi K3 licence text
- Hugging Face: Qwen3.8-2.4T-A95B model card (Alibaba)
- Hugging Face: Qwen3.8-Max licence text
- Hugging Face: GLM-5.3 model card (Z.ai)
- Hugging Face: GLM-5.3 licence text
- Artificial Analysis: top open-weights models
- Artificial Analysis: leaderboard of all models by Intelligence Index
- Trending Topics: Mistral Large 4 on Artificial Analysis, Eighth Among Open Models (6 October 2026)
- Arena: text leaderboard
- NIST CAISI: Evaluation of DeepSeek V4 Pro (1 May 2026)
- Epoch AI: Open models lag state-of-the-art closed models by 4 months (29 May 2026)
- Interconnects: The current balance of power in open models (Nathan Lambert, 21 September 2026)
- ChinaTalk: DeepSeek V4 (27 April 2026)
- The Next Web: Nvidia's H200 finally reaches Chinese buyers (19 August 2026)
Researched and written with the help of AI tools and edited for accuracy. Provided for general information and discussion only, not professional advice. See our editorial standards and disclaimer. Spotted an error? Tell us.
Enjoyed this? Get the next one.
One good read at a time, straight to your inbox. No spam, unsubscribe anytime.