跳到正文
原文
The Decoder· Jonathan Kemper·· 5 小時前AI 評分75

Mistral Large 4 主打美國模型拒絕處理的網絡安全任務,具備 1 trillion 參數

Mistral Large 4 is Europe's trillion-parameter answer to US models that refuse security work

AI 導讀

Mistral 發佈 Mistral Large 4 公開預覽版,這款 1 trillion 參數的多模態模型已可透過 Mistral Studio API 使用。

正文

Mistral has unveiled Mistral Large 4, its largest model to date. It has one trillion parameters and was trained in the company's own European data centers. In an independent index, the model makes a big leap forward but still falls well short of the leading closed models. Mistral's main pitch is that it handles security tasks that Claude and GPT-6 refuse to touch.

Mistral has released a public preview of Mistral Large 4 (ML4), nicknamed "le Chonk." The API is available now through Mistral Studio. Model weights are expected to follow at the end of October.

According to Mistral, ML4 is a natively multimodal model with one trillion parameters, 49 billion of which are active. On X, the company says ML4 is the best open-weight model from the US or Europe across aggregated benchmarks. It claims state-of-the-art performance in cyber defense, manufacturing, and finance, and says it beats even closed frontier models in visual grounding.

A big leap for Mistral, but Claude still leads by 20 points

An independent look comes from the Intelligence Index by Artificial Analysis, which aggregates ten benchmarks across different domains. ML4 scores 38 points there. That's a major step up for Mistral: its predecessor, Mistral Large 3, managed just 9 points in the index, and the most recent Mistral Medium 3.5 hit 14. At 38 points, ML4 also edges past GLM-5.2 from Z.ai, which Mistral offers on its own platform. But the gap to the top is still large. Claude Opus 5.5 (Max) leads with 58 points.

Balkendiagramm des Artificial Analysis Intelligence Index mit 25 Modellen, angeführt von Claude Opus 5.5 mit 58 Punkten vor Claude Sonnet 5.5 mit 56 sowie Claude Fable 5.1, GPT-6 Astra und Gemini 4 Argon mit je 53 Punkten; Mistral Large 4 Preview liegt mit 38 Punkten im hinteren Mittelfeld gleichauf mit GPT-6 Luna.
In the Artificial Analysis Intelligence Index, Mistral Large 4 trails the leading closed models from Anthropic, OpenAI, and Google, as well as several open-weight models from China. ML4 is still listed as proprietary because its weights haven't been released yet.

Until the weights ship, Mistral is red-teaming the model with security firms, vetted partners, and government agencies. Those groups get the same version with reduced moderation and expanded cyber capabilities.

Balkendiagramm zu SciCode-Verified pass@1 mit Mistral Large 4 Preview bei 91,8 Prozent im Vergleich zu elf Modellen, darunter GPT-6 Astra mit 94,2, Qwen3.8 mit 93,8 und Claude Opus 5 mit 91,3 Prozent.
Mistral Large 4 scores near the top in scientific coding, though GPT-6 Astra and Qwen3.8 still post higher numbers.

Mistral's cybersecurity edge partly comes from rivals refusing the task

Mistral devotes the bulk of its announcement to IT security. In the Artificial Analysis Cyber Index, ML4 ranks among the top five models worldwide, according to the company. Among open-weight models developed outside China, it leads by a wide margin.

One test asks a model to reproduce a real vulnerability in open-source software and then patch it. ML4 hits 82 percent, the highest score of any model. The reason is unusual, though. According to Mistral, Claude Opus 5.5 and GPT-6 Astra score near zero because they refuse the task entirely. So the test measures provider policies as much as model capabilities.

Mistral turns this into a selling point. Defending software often starts with proving that a vulnerability is real, the company argues. The safety filters on closed models block exactly that kind of work, while attackers jailbreak those same models anyway. Losing access to a provider mid-incident can itself become a security risk, Mistral adds. ML4 is designed to run in private clouds or on-premise as well.

At the same time, Mistral touts high refusal rates. On malicious cyber prompts from JailbreakBench, StrongREJECT, and AgentHarm, ML4 refuses more often than any other open model. How the model reliably tells legitimate vulnerability research apart from attack prep, Mistral doesn't explain. In Lakera's B3 AI Security Benchmark, ML4 blocks 93.3 percent of attacks, according to the company.

ML4 takes second place in coding, ahead of Deepseek and Qwen

In programming, ML4 scores 49.8 percent in the Artificial Analysis Coding Agent Index, putting it ahead of Deepseek V4 Pro and Qwen3.8 Max. Mistral also ran a blind evaluation with Surge AI, where professional annotators rated code quality without knowing which model produced it. ML4 landed second out of five models with 3.74 out of 5 points, ahead of GLM-5.3 and Kimi K3. Claude Opus 5 led with 4.22 points.

Balkendiagramm mit der gewichteten Gewinnrate von ML4 gegenüber GLM-5.3 nach Bereich, STEM 68, CAD 62, Finance 50 und Code 48.
Human evaluators prefer ML4 in STEM and CAD tasks, while both models perform roughly the same in finance and code.

On image understanding, Mistral claims a major improvement. ML4 is supposed to analyze documents, charts, gigapixel satellite imagery, and technical drawings, zooming in and checking details on its own. The lead over closed models is slim, though. On the Dense 200 benchmark, ML4 scores 42 percent while GPT-6 Astra hits 41 percent. According to an evaluation by Vals.ai, ML4 also beats GPT-6 Astra on legal and financial tasks.

Balkendiagramm zu AutomationBench von Artificial Analysis mit Mistral Large 4 Preview bei 59,9, GLM-5.3 bei 62,2, Kimi K3 bei 58,3, Qwen3.8 bei 57,2, DeepSeek V4 Pro bei 56,7, GLM-5.2 bei 28,4 und Mistral Medium 3.5 bei 6,3.
On automated business workflows, Mistral Large 4 trails GLM-5.3 slightly but beats Kimi K3 and Deepseek V4 Pro.

European sovereignty as a business model

Mistral says ML4 was trained from scratch on 3,800 Nvidia Grace Blackwell GPUs in its own European data centers. The preview runs on the same infrastructure. Alongside offerings in multiple regions, Mistral plans a European variant that it operates entirely on its own, independent of other service providers and under European law. Training data covers more than 160 languages, including all official EU languages.

Training happened in collaboration with companies in finance, manufacturing, logistics, pharma, shipping, and the public sector, according to Mistral. The company used the same training and RL environment it offers customers through Mistral Forge. For reinforcement learning post-training, Mistral uses about 3,000 GPUs. A single training run generates roughly 33 billion tokens per day.

Training isn't done yet

The RL run behind the preview is still ongoing and shows no signs of plateauing, Mistral says. The company expects significant improvements over the coming weeks. The compute expansion is funded by the Series D round of 3 billion euros, the largest equity round ever raised by a European tech company.

Drei Liniendiagramme zeigen Benchmark-Werte über die Trainingsschritte, gelb für Supervised Fine-Tuning und orange für Reinforcement Learning, mit Endwerten von 68,84 Prozent, 38,66 Prozent und 1.511.
Both supervised fine-tuning and the subsequent RL phase improve results on downstream benchmarks, according to Mistral.

A lot remains unclear for now. According to the model documentation, ML4 uses a fine-grained mixture-of-experts architecture with 1.05 trillion total parameters and a 1.6 billion parameter vision encoder. The context window spans one million tokens. Mistral plans to release architecture details, the license, and post-training methods along with the weights at the end of the month.

During the preview, Mistral charges $0.68 per million input tokens and $2.09 per million output tokens. Cached inputs cost $0.07. The documentation also lists prices at double those rates: $1.36 for input, $4.18 for output, and $0.14 for cached inputs. ML4 is also meant to serve as the foundation for a new generation of specialized Mistral models.

Mistral has been building its own infrastructure in Europe for months. In March, the company took out an $830 million loan for a data center near Paris, with 200 megawatts of compute capacity in Europe planned by the end of 2027. The company appears to be shifting its focus from consumers toward enterprise customers. In May, Mistral renamed its chatbot Le Chat to Vibe and rebuilt it as a work tool.

That same month, Mistral CEO Arthur Mensch warned a French parliamentary commission that Europe risks becoming dependent on US models for cybersecurity. The French military's codebases shouldn't be scanned by Anthropic's Mythos, Mensch said. Mistral's own models could find the same vulnerabilities that have been linked to Mythos, he added.

來源:The Decoder · the-decoder.com