Mistral launched a public preview of Mistral Large 4 (ML4) on October 6, 2026, and announced the weights will drop by the end of the month. That two-step matters more than the benchmark chart: you can test through the preview API now, then deploy the same model on your own infrastructure weeks later. For anyone building on hosted-only frontier models, that’s a different procurement conversation.
The model itself
ML4 is a 1-trillion-parameter natively multimodal model with 49 billion active parameters, trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral’s own European datacenters. It was trained on data spanning more than 160 languages, including every official EU language. Mistral says it will publish architecture details, additional benchmarks, and its post-training methodology as the release approaches — those specifics are not in the announcement yet.
Cyber performance is the headline, for a structural reason
The most interesting claim isn’t a score; it’s why the score exists. On the Artificial Analysis Cyber Index, ML4 ranks among the top five models globally, and on a test requiring it to reproduce a real open-source vulnerability and then patch it, Mistral reports 82% — the highest of any model — while several closed frontier models score near zero. The gap, per Mistral, is refusal behavior: closed models decline the reproduce-a-flaw step, even though proving a flaw is real is how defense starts.
If you build security tooling, this is worth taking seriously but verifying yourself. The claim comes from Mistral’s own announcement and internal testing, and the preview API presumably still has moderation layers that a self-deployment would not. That difference between the preview and the eventual open weights is exactly what your evaluation should probe before you commit.
Coding, agents, and the usual benchmark spread
Beyond security, Mistral reports 61.7% on DeepSWE v1.1, 28.3% on Terminal-Bench 4, and a combined coding-agent index of 49.8%, placing it ahead of DeepSeek V4 Pro 0813 and Qwen3.8 Max. A blind human evaluation run with Surge AI ranked ML4 second of five models, behind Claude Opus 5 and ahead of Kimi K3 and GLM-5.3. On AutomationBench’s 657 business workflows it scores 59.9%, and on the security side it resists 93.3% of attacks on Lakera’s B3 benchmark, which Mistral says is the best score it sees among comparable open models.
Treat all of this as vendor-reported until independent numbers land. What’s genuinely useful is the shape: one model covering coding, agentic tool use, vision-grounded document work, and regulated-industry tasks, with weights you can audit.
What this changes for your stack decisions
The sovereignty angle is the part most likely to affect real roadmaps. Mistral is offering a European deployment it operates end-to-end under European law, plus self-deployment on private cloud or on-premise. If you’ve been weighing a frontier closed model against an open one and losing on capability, ML4’s pitch is that you no longer have to — at least on cyber, finance, and legal workloads.
That maps directly onto the tradeoffs I looked at earlier in Sovereign AI, One Year In: What Open Models Actually Changed for Builders: open weights buy you control and auditability, but you inherit the operational burden of hosting, scaling, and updating. ML4’s 49B active parameters keep inference costs plausible for a 1T model, but you should still cost a self-hosted deployment against the preview API before assuming sovereignty is worth it for your workload.
A practical next step: run the preview API against your own eval set this month, then re-run against the weights when they land. Refusal behavior and moderation thresholds can differ between the two, and that difference is the thing most likely to surprise you in production.
