Updated October 8, 2026. Mistral Large 4 is the latest flagship AI model from Mistral AI. Announced on October 6 and nicknamed “Le Chonk,” it is available as a public preview, with downloadable weights planned for the end of the month.
The launch matters because it brings another European contender into the race for capable, deployable AI. Its one-trillion-parameter headline is striking, but the practical questions are more interesting: what can you test today, what does “open weight” actually promise, and when would running the model yourself make sense?
The quick answer: Mistral Large 4 combines text and image understanding with reasoning, coding and tool use. The preview gives developers a chance to evaluate it now. The planned weight release could give organizations more control over deployment, subject to the final license and technical requirements.

What changed with the Mistral Large 4 announcement?
Mistral says Large 4 is its largest and most capable model so far. It uses a mixture-of-experts architecture, with one trillion total parameters and 52 billion active parameters per token. The company also highlights native multimodal input and multilingual coverage across more than 160 languages. These are specifications reported in Mistral’s official launch announcement.
The distinction between a preview and a weight release is central to this story. You can experiment with the hosted model now; you cannot assume that the full downloadable package is already available. Organizations planning a private deployment should wait for that package before committing hardware or promising a launch date.
Mistral Large 4 at a glance
| Detail | Launch information |
|---|---|
| Availability | Public preview through Mistral’s API and Studio |
| Architecture | Hybrid instruction-and-reasoning mixture-of-experts model |
| Parameters | Approximately 1 trillion total (1.05T in the model documentation); 52 billion active per token |
| Inputs | Text and images |
| Language coverage | 160+ languages, including all official EU languages |
| Training infrastructure | 3,800 NVIDIA Grace Blackwell GPUs in European data centers |
| Announced list price | US$1.36 per million input tokens; US$4.18 per million output tokens |
| Preview sale shown October 8 | US$0.68 input; US$2.09 output; US$0.07 cached input per million tokens |
| Downloadable weights | Planned for the end of October 2026 |
Check the current model documentation before using the preview in a project. The documentation currently shows a discounted preview price alongside the original list price. Availability, serving limits and pricing can change. Also distinguish input tokens from output tokens when estimating a bill: a long generated report can cost more than the short question that prompted it.
Why one trillion parameters does not mean one trillion at work
A parameter is a learned value inside a neural network. In a dense model, most of the network participates in processing each token. A mixture-of-experts model uses a routing mechanism to select a smaller set of specialist components for a particular token.
Think of a large organization with several teams. The organization has broad expertise, but every employee does not join every task. The analogy is imperfect—experts are mathematical components rather than people—but it explains why total capacity and active computation are different numbers.

Sparse activation can make a very large model more economical to serve than its total size suggests. It does not remove the need to store the weights or move data between devices. GPU memory, interconnect bandwidth, quantization and batching remain important. Our AI inference guide explains how those choices affect speed and cost.
Parameter counts are therefore a poor shopping guide on their own. A model with fewer parameters may be more reliable on your task, easier to deploy or cheaper at your traffic level. Compare measured results rather than choosing the biggest number.
Where the new model could be useful
Code that requires repository context
A useful coding assistant needs to understand relationships across files, preserve existing behavior and explain the consequences of a change. Large 4 is positioned for this kind of agentic software work. A realistic trial would ask it to diagnose a bug in a controlled repository, make a small patch and run the relevant checks.
The acceptance test should include whether the patch fixes the actual problem, whether it creates a regression and whether a reviewer can understand it. A polished explanation is useful, but the working software is the evidence.
Documents, charts and visual interfaces
Native image understanding gives a model a way to reason over material that does not fit into plain text. Examples include charts, engineering diagrams and screenshots. Visual grounding adds a more precise task: identify the region in an image that corresponds to an instruction.
For document workflows, ask the model to identify the page and evidence behind an answer. A chart can contain ambiguous labels, a scanned table can lose a digit and a screenshot can be too small to read. Keep the original file available for verification. Our multimodal AI guide covers these input types and their limits.
Tool-based enterprise work
The model is also intended for agents that gather information, operate approved tools and produce documents or other deliverables. The application around it determines which systems it can reach and what actions it can take. A capable model does not automatically provide safe permissions or a reliable workflow.
For example, an internal reporting agent could retrieve approved records, draft a summary and prepare a spreadsheet. Sending the report or modifying a source record should be a separate, explicitly authorized step. See our AI agents guide for how models, tools and human oversight fit together.
Reading the benchmarks without overreading them
The launch post lists results across coding, cybersecurity, agentic work and vision. The table below summarizes selected published scores. These figures come from Mistral’s announcement; the Unlimited AI Editorial Team has not independently rerun the evaluations.
| Evaluation | Score listed by Mistral | What to examine before comparing |
|---|---|---|
| DeepSWE v1.1 | 61.7% | Repository tasks, tool scaffold and test conditions |
| Terminal-Bench 4 | 28.3% | Terminal environment and permitted tools |
| Cybench | 93% | Competition-style security tasks and refusal behavior |
| AutomationBench | 59.9% | App workflows, reliability and completion criteria |
| Dense 200 | 42% | Visual grounding setup and exact localization metric |
| B3 attack resistance | 93.3% | Attack set, application context and remaining failure cases |
A benchmark measures performance under a specific setup. Changing the instructions, available tools, sampling settings or time budget can change the score. Safety refusals can also affect comparisons, especially on cybersecurity tasks. A higher number does not always mean that a model has more technical knowledge.
Use a leaderboard to choose candidates for testing. Then use your own tasks to decide which candidate fits. The most useful question is whether the model produces correct, reviewable work at an acceptable cost with your actual inputs.
Open weights: the promise and the practical limits
Model weights are the learned values produced by training. Access to them can enable local hosting, adaptation, quantization and research. It can also reduce dependence on the availability or policy of one hosted API.
The final license determines what you can legally modify, redistribute and use commercially. Open weights alone do not establish that the training data, full training code or every supporting component is available. Read the actual release terms before describing a model as fully open source or building a commercial plan around it.

Self-hosting also transfers operational responsibility. Someone must provision hardware, monitor the service, restrict access, apply updates and investigate failures. A hosted API can be the more practical option for a small team or an irregular workload, even when downloadable weights exist.
For a regulated organization, deployment control may matter more than the lowest token price. Questions about data residency, logging, vendor access and service continuity belong in the decision. Our sovereign AI guide explains why sovereignty involves governance as well as geography.
Why the staged release deserves attention
Mistral says it is red-teaming the model with cybersecurity leaders, vetted partners and state authorities before releasing weights. That puts the model’s security capabilities and misuse risks at the center of the rollout.
Security assistance has legitimate uses, including analyzing a vulnerability and checking a patch. Those capabilities can also be misused. A controlled defensive evaluation should use authorized systems, isolated environments and limited credentials. An agent with access to production tools needs protections beyond the model’s own behavior.
A score is not a security guarantee: Even strong prompt-injection resistance leaves possible failures. Keep untrusted documents separate from trusted instructions, restrict tools to the minimum necessary permissions and require approval for consequential actions.
Our AI safety testing guide explains how capability testing, red-team exercises and deployment safeguards answer different questions.
A practical evaluation plan for the preview
Start with a small trial that can produce a decision. Choose representative tasks, define what a correct result looks like and compare Large 4 with the model you already use. Avoid moving a production workflow simply because a launch headline sounds impressive.
- Pick three real workflows. Include a text task, a visual task and a controlled tool task that reflect your users’ needs.
- Build a private test set. Include ordinary examples, ambiguous inputs, missing information and cases from each important language.
- Define success in advance. Record required facts, acceptable output structure and errors that would make a response unusable.
- Measure quality and cost together. Track completion rate, latency, token use, unsupported claims and human correction time.
- Test failures deliberately. Include unreadable images, conflicting instructions and tools that return an error.
- Reassess after the weights arrive. A quantized, fine-tuned or locally hosted version may behave differently from the preview.
For a concrete example, take ten anonymized internal reports and ask each candidate to extract the same five fields with page references. Have a reviewer check the fields against the originals. Record the time spent correcting each output. This makes the comparison about usable work instead of fluent prose.
Questions still open before the weight release
- What will the final license permit?
- What hardware will be needed for useful self-hosted performance?
- Which quantized variants and serving tools will be available?
- How closely will public weights match the hosted preview?
- How will independent evaluations compare with the launch results?
- What limitations appear in your own languages, documents and tool workflows?
These unknowns do not make the launch unimportant. They define what to verify next. A sensible deployment plan has a checkpoint after the technical package arrives, rather than assuming the announcement settles every requirement.
Mistral Large 4 FAQ
Can I try Mistral Large 4 now?
Yes. Mistral has made it available in public preview through its API and Studio. Use the current model documentation to check access and request details.
Are the downloadable weights available?
At the time of this October 8 review, Mistral’s announcement says weights are planned for the end of October 2026. Treat that as a release plan rather than completed availability.
Why are only 52 billion parameters active?
Large 4 uses mixture-of-experts routing. Selected components participate in processing each token, so active computation is smaller than the model’s total parameter count.
Does multimodal mean it generates images?
The announced multimodal capability concerns image input and understanding. Do not assume that understanding an image means the model provides a dedicated image-generation service.
Should a business self-host it?
That depends on privacy requirements, traffic, infrastructure and the final license. Compare the total operational cost with API access, and test the version you would actually deploy.
What this launch means next
Mistral Large 4 adds a significant new candidate for organizations evaluating coding, visual understanding and tool-based AI. Its planned weight release makes deployment control part of the conversation, alongside quality and price.
The useful next step is a measured preview trial. Keep the vendor’s reported results separate from your own findings, watch for the final release package and make the production decision from evidence that matches your workload.



