EMPVResearch NotesIT
EMPV / RESEARCH NOTESNOTE / 010

NEWS NOTE / OPEN MODELS

Mistral Large 4: a trillion-parameter European model in preview

Mistral unveils a multimodal Mixture-of-Experts model trained on European infrastructure. The preview API is available, while weights are planned for late October. Independent tests find strengths in cyber and cost trade-offs.

Published2026-10-09 · 6 min
Mistral Large 4Open-weight AIMixture-of-ExpertsAI sovereignty
01 / The announcement

For now, it is an API preview. The weights cannot yet be downloaded.

On October 6, 2026, Mistral unveiled Mistral Large 4 (ML4), informally nicknamed “Le Chonk”, and opened a public preview through Mistral Studio. It is the French company's largest model yet, aimed at coding, agentic workflows and multimodal analysis.

Availability matters: as of October 9, developers can use the preview API, while Mistral says weights are due at the end of October. It would be premature to describe ML4 as already downloadable for production deployment.

02 / Architecture

One trillion total parameters does not mean activating them all for each token.

The technical documentation describes a multimodal granular Mixture-of-Experts (MoE) model with around 1.05 trillion total parameters and a 1.6-billion-parameter vision encoder. Mistral's model page lists 52 billion active parameters, while Artificial Analysis reports 49 billion. Both figures appear in public launch materials; until final architecture details arrive, they should not be treated as directly interchangeable.

MoE routing activates only a subset of experts for each generated token. Infrastructure still needs to host a model with roughly a trillion total parameters: the active-parameter count is not the memory footprint of a conventional dense 49- or 52-billion-parameter model.

Mistral says the training covers more than 160 languages. Its model page advertises a context window of up to one million tokens, while Artificial Analysis lists approximately 524,000 tokens for the API configuration it evaluated. Effective limits depend on the endpoint and should be checked before building a long-context application.

03 / Performance

The strongest results concern specific workloads, not every model benchmark.

In its launch announcement, Mistral reports 61.7% on DeepSWE v1.1, 28.3% on Terminal-Bench 4.0 and 59.9% on AutomationBench, which measures multi-application workflows. It also reports solving 93% of Cybench challenges. These are vendor-presented results: harnesses, settings and comparison conditions matter before applying them to an enterprise evaluation.

In Artificial Analysis's independent evaluation, the preview scores 38 on its Intelligence Index and 50 on its Cyber Index. The same source also highlights cost per task: $1.13 at standard API rates on its Intelligence Index, against $0.25 for GLM-5.3-Flash and $0.27 for DeepSeek V4.1 Flash, which have nearby intelligence scores. A temporary launch discount halves the ML4 figure to $0.57 per task. These are benchmark-specific costs, not estimates for all workloads.

Model policy also affects some cybersecurity tests: a model can refuse the requested action. Mistral highlights this constraint and is red-teaming ML4 with selected cybersecurity partners and government authorities. Benchmark scores reflect both technical capability and system behavior in a particular evaluation environment.

04 / European infrastructure

European training is documented. Operational autonomy depends on weights and deployment.

According to Mistral, ML4 was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in the company's European data centers. The public preview is served from that same infrastructure. This is a concrete infrastructure decision: for this version, model training and inference do not depend on consuming a US vendor's model API.

For an enterprise, however, “AI sovereignty” can refer to different boundaries: data location, inference operations, log custody, service continuity and portability. Releasing the weights would introduce a self-deployment option, but it would not make a model of this size straightforward to run on ordinary enterprise hardware.

The same decision boundary appears in our note on on-premise versus cloud AI: deployment location should follow data requirements, integration constraints and the team's ability to operate the system.

05 / Enterprise evaluation

A useful test covers an entire workflow, not an isolated answer.

Mistral's claimed capabilities across documents, images, tools and reasoning could be relevant to technical-document analysis, incident response support or preparing dossiers from mixed sources. The model is only one component, though: authorization, retrieval, APIs, human review and action controls sit outside its checkpoint.

A reproducible trial should use domain-specific documents and tasks, define successful completion in advance, log failures and human interventions, and measure end-to-end cost and latency. We outlined this approach in our note on evaluating a local LLM before production.

For example, when analysing industrial documents, correctly describing an engineering drawing is not enough. The system must identify the right revision, cite its source, avoid exposing unauthorized information and flag details that cannot be read reliably.

06 / What is not yet established

Downloadable weights still need to clarify licensing, deployment requirements and reproducibility.

Mistral says it will publish more architecture details, benchmark results and post-training information with the weights. At the date of this note, final license terms and self-hosted behavior cannot be evaluated against a public checkpoint because that checkpoint has not been released.

The next meaningful checks are whether the weights arrive, how much accelerator memory they require, which runtimes support them, what license applies and how self-hosted performance compares with the preview API. Artificial Analysis's early evaluation describes the current API, not every possible future deployment.

ML4 demonstrates that a European lab can design, train and serve a model at this scale on its own infrastructure. The degree to which enterprises can inherit that control will be clearer after the weights are released.

EMPV / TAKEAWAY

Mistral Large 4 is a significant European infrastructure milestone, available today through an API preview. Cyber results stand out, but private-deployment cost, hardware needs and autonomy still need testing after the weights ship.

EMPV / SHARE

Share this Research Note.

Research Notes