Illustration by tuput
English
Mistral Large 4 went into public preview on 6 October 2026 with a 1-trillion-parameter design and a promise to release its weights by the end of the month. Artificial Analysis scored it 38, the best open model from outside China but behind seven Chinese ones. Its licence has not been named.
Mistral Large 4, the model the French company nicknamed Le Chonk, scored 38 on the Artificial Analysis Intelligence Index when that independent testing firm ran the public preview on 6 October 2026. Trending Topics counts that as eighth among open-weight models, behind seven Chinese ones, although Artificial Analysis itself lists the model as proprietary because the weights are not yet public.
“Open-weight” means the company publishes the trained numbers that make up the model, so anyone with enough hardware can run it, change it and host it themselves. Mistral’s announcement says only: “We will release the weights by the end of the month.”
What was released on 6 October, and what was not
Mistral opened a public preview through its own API on 6 October 2026. Nothing can be downloaded yet, and the model card labels it “Open” without naming a licence. Forklog reports a full weights release on 27 October. Mistral’s own post gives no day.
The 7 October AI roundup covered the launch in a few lines. Two figures in it differ from Mistral’s own post: the roundup gave 49 billion active parameters and 4,000 chips, where Mistral says 52 billion and 3,800.
A trillion parameters, and three numbers that disagree
Large 4 is a mixture-of-experts model. Instead of using every parameter (a parameter is one adjustable number inside the network) for every word it produces, it routes each piece of text through a small share of specialist sub-networks. In its announcement Mistral describes it as a “1 trillion-parameter natively multimodal model with 52 billion active parameters,” meaning it reads images as well as text. The model card gives the total as 1.05 trillion, adds a 1.6 billion parameter vision encoder and lists a 1 million token context window, the amount of text the model can consider at once. Three of those figures are contested.
- Active parameters: Mistral’s post and model card say 52 billion. 49 billion is the figure given by the Register, IT Daily and Unite.AI’s account of the Artificial Analysis page.
- Context window: Mistral says 1 million tokens. Artificial Analysis measured about 524,000, and Vals.ai lists 512,000.
- Training chips: Mistral says it trained the model from scratch on 3,800 Nvidia Grace Blackwell GPUs in its own European datacentres. Forklog says about 4,000.
For scale, Mistral’s December 2025 Large 3 had 675 billion total and 41 billion active parameters, according to Mistral’s documentation. That makes the new model roughly 56% bigger in total size. Mistral gave no hardware guidance, the Register notes, but it adds that the model should still run on 8-way GPU servers such as Nvidia’s HGX B300 or AMD’s MI355X.
Training data covers more than 160 languages, Mistral says, including every official language of the European Union.
What Mistral says the model can do
Everything in this section is Mistral’s own reporting unless another name is attached, and the Register’s advice on Mistral’s charts is to “take them with a grain of salt.”
On coding, Mistral reports 61.7% on DeepSWE v1.1 and 49.8% on Artificial Analysis’s Coding Agent Index, ahead of DeepSeek V4 Pro 0813 and Qwen3.8 Max. In a blind test by the data firm Surge AI, annotators rated outputs from five models on a 1 to 5 scale. Mistral’s model came second of the five with 3.74, behind Claude Opus 5 at 4.22.
On agent tasks, it reports 59.9% on AutomationBench, which uses 657 business workflows. A figure Mistral’s summary leaves out comes from the Decoder: GLM-5.3 scored higher, at 62.2. On SciCode-Verified the Decoder reports 91.8%, behind GPT-6 Astra at 94.2% and Qwen3.8 at 93.8% but ahead of Claude Opus 5 at 91.3%.
On images, Mistral reports 42% on a visual grounding test called Dense 200, against 41% for GPT-6 Astra.
What the independent testers found
Artificial Analysis scored the preview 38 on version 4.3.2 of its Intelligence Index and ranked it 64th of 225 models. That equals GPT-6 Luna (max) at 38 and sits just below DeepSeek V4.1 Flash (max) at 39, per Unite.AI. Per the Register, early testing leaves it at a “major disadvantage” against the flagship models from OpenAI and Anthropic.
Trending Topics lists the open-weight leaders on the same index: Xiaomi’s MiMo-V2.6-Pro at 46.3, Z.ai’s GLM-5.3 at 44.8 and Moonshot AI’s Kimi K3 at 43.6. Among closed models it lists Claude Opus 5.5 at 57.6, GPT-6 Astra at 52.7 and Gemini 4 Argon at 52.6. South Korea’s Motif 3, at 33.6, was the previous best open model from outside China. ActuIA puts Thinking Machines’ Inkling Small, the best measured US open model, at 26.
On the same index Mistral’s earlier models scored far lower: Medium 3.5 at 14 and Large 3 at 9. ActuIA points out that Large 3 had no reasoning mode and the index favours reasoning.
Efficiency is the weak spot. Running the index cost $1.13 per task, against $0.27 for DeepSeek V4.1 Flash and $0.13 for MiMo-V2.6-Pro, according to Trending Topics. The model wrote about 200 million output tokens to finish the index, against a median of 81 million for comparable models. Speed measured 116 output tokens per second, with 1.46 seconds to the first token.
Vals.ai, which tests models on professional tasks, ranks Large 4 33rd of 45 on its overall index, at 48.05%. Harvey’s Legal Agent Benchmark is its best placing at 6th of 76, and Finance Agent v2 puts it 22nd of 76.
The cyber result needs a footnote
Mistral is selling Large 4 hardest on security work. It reports 82% on a test that asks the model to reproduce and patch a real open-source vulnerability, and says Claude Opus 5.5 and GPT-6 Astra score near zero on that test because they refuse the task.
Artificial Analysis’s own cyber measurements back that direction. ActuIA reports a Cyber Index score of 49.5, fifth of 18 models, and 81.7% on the index’s CyberGym-E2E test, the best of the 18. By that account GPT-6 Astra scored zero with a 100% refusal rate, and Claude Opus 5.5 scored 0.8% with 98.5% refusals.
The Decoder notes that this result reflects provider policy as well as ability, since a model that declines a task cannot complete it. Mistral has not explained how it separates legitimate research from attack preparation. Vetted partners and state authorities, Mistral says, will get a version with reduced moderation and expanded cyber capabilities during red-teaming. Trending Topics reports that the public version restricts those abilities.
Price and licence
The list price is $1.36 per million input tokens and $4.18 per million output tokens. Half those rates appeared on the model card, $0.68 and $2.09, as a sale price. Unite.AI’s account of Artificial Analysis’s page describes the discount as a launch offer for the first two weeks, while ActuIA found no end date stated.
The licence is the larger unknown. It decides whether companies can use the model commercially, and neither Mistral’s launch post nor its model card names one. Mistral’s documentation shows Large 3 shipped under Apache 2.0, a permissive licence. According to Trending Topics, Medium 3.5 carries a modified MIT licence that excludes companies with more than $20 million in monthly revenue, and VentureBeat reported a custom licence for Large 4 without details. tuput could not locate that VentureBeat report. ActuIA notes that Mistral’s list of details to come with the weights covers architecture, benchmarks and training method, and does not mention the licence.
Who it is for, and a rival due the same month
Mistral names finance, engineering, manufacturing, logistics, pharmaceuticals, science, shipping and the public sector as target customers, along with security teams that want to run the model on a private cloud or their own servers. Its post says: “Forged in Europe. Built for AI sovereignty.” The same argument sits behind India’s push for home-built models, which tuput covered in India Is Building AI in Its Own Languages. For a general explainer on what these systems do, see how a chatbot guesses the next word.
Another release is queued behind it. TechCrunch reported on 5 October that Reflection AI, a Brooklyn startup founded by two former Google DeepMind researchers, announced Beam, a text-only mixture-of-experts model with 501 billion total and 23 billion active parameters. Reflection says Beam matches Z.ai’s GLM-5.2 on advanced reasoning benchmarks and will release its weights in October. TechCrunch notes the claims are the company’s own and does not name a licence.
Mistral says the reinforcement learning run behind the preview is still going and expects large gains, so the 38 may not be the final score. Mistral’s deadline for the weights is the end of October.
Sources & further reading
- Mistral AI: Introducing Mistral Large 4
- Mistral AI docs: Mistral Large 4 model card
- Mistral AI docs: Mistral Large 3 model card
- Artificial Analysis: Mistral Large 4 Preview
- Vals.ai: Mistral Large 4 results
- Unite.AI: Benchmark Firm Gives Mistral Large 4 Preview a 38 Intelligence Score
- Trending Topics: Mistral Large 4 on Artificial Analysis, Eighth Among Open Models
- The Decoder: Mistral Large 4 is Europe's trillion-parameter answer to US models that refuse security work
- ActuIA: Mistral Large 4, how good is the best open model outside China?
- The Register: Mistral Large 4 coverage
- Forklog: Mistral Large 4 preview and weights date
- IT Daily: Mistral Large 4 (Le Chonk)
- TechCrunch: Reflection debuts Beam, an open-weight AI model to rival Chinese models
Researched and written with the help of AI tools and edited for accuracy. Provided for general information and discussion only, not professional advice. See our editorial standards and disclaimer. Spotted an error? Tell us.
Enjoyed this? Get the next one.
One good read at a time, straight to your inbox. No spam, unsubscribe anytime.