, , , ,

Mistral's Trillion-Parameter Model Is Built for the Work Claude Refuses

Close-up detail of a soldered server motherboard under a magnifying lamp, USB debug cable coiled on graph paper, flux residue on copper trac

Mistral put a 1-trillion-parameter model into public preview on October 6 and nicknamed it Le Chonk. The selling point is not the count. It is an 82 percent score on a cyber test that Claude Opus 5.5 and GPT-6 Astra refuse to take.

Mistral Large 4 is a natively multimodal mixture-of-experts system with 49 billion active parameters. It was trained from scratch on 3,800 Nvidia Grace Blackwell GPUs in the company's own European datacenters and is served from the same stack. The public preview is live on Mistral Studio. Open weights are promised by the end of October. Until then, Mistral says it is red-teaming the model with cybersecurity firms, vetted partners, and state authorities, some of whom will get reduced moderation and expanded cyber tools. Training data spanned more than 160 languages, including every official language of the European Union. The company is pitching a European deployment it operates end-to-end, under European law, independent of other cloud providers.

On the Artificial Analysis Cyber Index, Large 4 ranks among the top five models globally. On a test that asks a model to reproduce a real vulnerability in open-source software and then patch it, it scored 82 percent, the highest of any model Mistral cited. Claude Opus 5.5 and GPT-6 Astra scored near zero on the same task because they refuse it. Large 4 also solved 93 percent of Cybench's 40 competition-style challenges. Coding scores were more mixed: 61.7 percent on DeepSWE v1.1, 59.4 percent on SWE-Atlas-QnA, 28.3 percent on Terminal-Bench 4, and 49.8 percent on a combined Coding Agent Index, ahead of DeepSeek V4 Pro and Qwen3.8 Max in Mistral's comparison. In a blind human evaluation by Surge AI, professional annotators ranked it second of five models, behind Claude Opus 5.

The second-order product is a refusal gap, not a leaderboard. Closed labs have spent two years tightening safety filters that also block the first step of a patch: proving the hole is real. Mistral is selling that work as sovereignty — open weights, self-hosting, and a model that will do the job on a private rack. The weights are not out. The 82 percent is a preview score. Europe's trillion-parameter bet is that defenders will pay for a system Washington's frontier labs will not run, and that a license file will matter more than another closed API when an incident is already underway.

Image source: i.ibb.co