Mistral AI
Mistral Large 4
A trillion-parameter model opens in public preview

Key facts
- 1.05T
- Total parameters
- 1M tokens
- Context
- API preview
- Access today
- October 2026
- Weights planned
Mistral Large 4 is an AI model that reads words and pictures, writes answers and code, and works through software tools. Mistral opened its public preview on 6 October 2026.
Listen to the model guide
Podcast: Mistral's trillion-parameter preview
Mistral Large 4 is an AI model that reads text and images, writes answers and code, and takes actions through software tools. Mistral opened its public API preview on 6 October 2026, with downloadable weights planned by the end of October.
The preview gives developers a hosted version to try now. The planned weight release will provide the learned numerical settings needed to run the model on their own infrastructure.
What can Mistral Large 4 do?
Mistral Large 4 combines image understanding, coding and tool use in one model. Its model documentation lists a one-million-token context window, function calling, structured outputs, document questions and conversational agents. Tokens are the pieces of text and encoded content the model processes.
A tool-using assistant can inspect a document, ask an application to retrieve information and use the result in its answer. The application decides which tools it exposes and which actions require a person’s approval.
Mistral’s launch demonstrations include analysing financial filings, examining technical drawings and investigating software vulnerabilities. These examples describe the tasks the company is targeting. A useful evaluation uses the documents, codebase and tool permissions the intended deployment will actually handle.
How does its trillion-parameter network work?
Mistral’s model page lists 1.05 trillion total parameters and a mixture-of-experts architecture. Parameters are the numerical settings learned during training. A learned router selects expert networks for each token, as Nvidia’s explanation of the architecture describes.
That selection lets a model carry a large collection of learned knowledge while using selected parts of its network at each step. The expert outputs are combined before processing continues through later layers.
Hosting still requires memory for the weights and efficient communication between processors. Total model size, the work activated at each step and the length of the input all affect deployment requirements. A hosted API lets a team evaluate results before planning its own serving infrastructure.

How much does the preview cost?
Mistral’s pricing page marks the following promotional prices on 6 October 2026. Each figure is in US dollars per million tokens.
| Token category | Published promotional price |
|---|---|
| Input | $0.68 |
| Cached input | $0.07 |
| Output | $2.09 |
Input is the content sent to the model; output is the content it generates. Cached input is previously processed content reused under the service’s caching rules. The total charge depends on the quantities billed in each category.
The public preview is available through Mistral Studio. Mistral says its European data centres trained and serve the preview model, and it plans to publish the weights by the end of October. The release timetable gives teams a hosted testing period before they assess a self-hosted deployment.

How should you read its benchmark scores?
Mistral’s launch post reports the following selected results. They describe separate evaluations with different tasks and scoring rules.
| Evaluation | Reported score | Task |
|---|---|---|
| DeepSWE v1.1 | 61.7% | Software engineering |
| Terminal-Bench 4 | 28.3% | Terminal workflows |
| Cybench | 93% | Security challenges |
These are Mistral’s published launch figures. The company attributes its DeepSWE and terminal results to Artificial Analysis; Cybench covers a set of security exercises.
For a deployment decision, run the same representative tasks through the models being compared. Keep tool access, prompts and output budgets consistent, and record completed tasks, errors, elapsed time and token charges. That produces a comparison tied to the work the model will do.
More in Large Language Models
All LLMs →- Mistral AIMistral Large 3Europe's flagship
- Reflection AIBeamA sparse model with an October open-weight release planned
- GoogleGemma 4the laptop model
- Google DeepMindGemini 4 Argonlonger reasoning, coding and cyber defence
- OpenAIGPT-5.6the flagship from 9 July to 3 September 2026
- OpenAIGPT-6 Astrathe new generation, rolling out from 3 September 2026