France's Mistral AI began a public preview of its new flagship model, Mistral Large 4, on October 6, 2026. Nicknamed "Le Chonk," the model has 1.05 trillion total parameters, according to its model documentation. The API is already available, and Mistral plans to release the trained weights by the end of October. The company is heavily promoting the model's performance in cyber defense and professional tasks. But if we look separately at its overall ranking, its success rates on specific tasks, and its operating costs, the reasons to choose this giant model become more concrete. ## How to Read a 1-Trillion-Parameter Model Large 4 is larger than its predecessor, Mistral Large 3, which has 675 billion parameters. However, it does not use all of its parameters for every inference. It adopts a Mixture-of-Experts (MoE) architecture that activates only some specialist networks depending on the input, and Mistral's announcement puts the number of active parameters at 49 billion. The model documentation as of October 7, however, lists 52 billion, so the official materials do not agree on this figure. MoE reduces the computation required for each inference even in a very large model, but it does not shrink the weights that must be stored. If each of the 1.05 trillion parameters were stored in 4 bits, the total would be about 525 GB (1.05 trillion × 4 ÷ 8, with GB in decimal units). This is a simple calculation based on Mistral's published total parameter count. It does not mean a 4-bit quantized version will actually be distributed, and it excludes the extra information quantization requires as well as working memory during inference. The minimum GPU configuration needed for self-hosting should be judged after the weights are released, by checking the implementation and operating conditions. The model also accepts image input. The official model documentation lists a 1.6-billion-parameter vision encoder and a context length of 1 million tokens. Context length is the upper limit on how much text, conversation history, and other material can be handled at once. That is a major expansion from the stated 256k of Large 3 and Mistral Medium 3.5, but the independent evaluator Artificial Analysis lists the preview version at 524k. There is a gap between the stated figure and the conditions confirmed in the evaluation environment, and being able to accept 1 million tokens does not guarantee equally accurate answers across the entirety of a very long document. ## A Big Leap for Mistral, but Not No. 1 Globally On Artificial Analysis's Intelligence Index, Large 4 Preview scored 38, up sharply from Large 3's 9 and Medium 3.5's 14. For companies that use Mistral models, that means more options for handling complex work. Version 4.3.2 of the index combines 10 evaluations, including science, code generation, and task execution by agents.
| Model and evaluation setting as listed | Intelligence Index | | ----------------------------------------------- | ------------------------------- | | Mistral Large 3 | 9 | | Mistral Medium 3.5 | 14 | | Mistral Large 4 Preview | 38 | | DeepSeek V4 Pro 0813(max) | 36 | | Kimi K3(max) | 44 | | GLM-5.3(max) | 45 | | Qwen3.8 Max(0902) | 45 | | MiMo-V2.6-Pro | 46 | | GPT-6 Astra(max) | 53 | | Gemini 4 Argon(high) | 53 | | Claude Opus 5.5(max with fallback) | 58 |
The source is the [Artificial Analysis model comparison](https://artificialanalysis.ai/leaderboards/models), with values as listed on October 7, 2026; reasoning-effort and other evaluation settings follow the notation in each row. Because this was not an experiment giving every model the same compute or cost, score ratios cannot be read directly as multiples of capability. Large 3 was also released as a non-reasoning model, so the gap between generations cannot be attributed simply to the increase in parameters. A score of 38 shows that Large 4 has advanced significantly, but also that a gap remains to the top models. It beats DeepSeek V4 Pro 0813 but falls short of the top models from Kimi, GLM, and MiMo, while the leader, Claude Opus 5.5, scores 58. Mistral's claim that Large 4 is "the most capable open-weights model developed in the US and Europe" is also a comparison limited by region and release format. It does not mean the model is No. 1 worldwide when leading Chinese models and the flagships of companies that do not release weights are included. ## Cyber Defense Strength, Seen Alongside Refusal Policies and Scoring Conditions One area where Large 4 stands out is reproducing and fixing vulnerabilities in real software. On Artificial Analysis's "CyberGym-E2E-AA," Large 4 Preview took first place with a single-attempt pass rate of 81.7%, ahead of MiMo-V2.6-Pro at 78.6% and GPT-6 Luna on its max setting at 77.9%. The benchmark covers 131 tasks selected from real projects written in C/C++. Using a shared agent execution environment called Stirrup, each task is given 90 minutes. A task counts as a success if the model creates an input that crashes the unpatched program, the crash no longer occurs after the fix, and the existing functional tests still pass. However, the score does not include an additional check on whether the fix fully eliminated the vulnerability that the dataset treats as the ground truth. The 81.7% figure cannot be taken to mean that the model could completely fix that share of vulnerabilities in real incident response. Safety refusal policies also affect the scores. Mistral explains that Claude Opus 5.5 and GPT-6 Astra sometimes refuse such tasks for safety reasons, which leaves their scores close to zero. Artificial Analysis's Cyber Index also scores tasks as zero when a model or provider declines to answer for safety reasons, and it shows refusal rates separately. In other words, a low score can reflect not only cases where a model could not solve the task, but also cases where it followed a policy of not attempting it. Mistral's aim is to offer an option that does not uniformly block legitimate defensive work because of the provider's safety policy. Once the weights are released, organizations will be able to run the model in their own environments and under their own rules. However, the conditions differ between the public preview currently offered to general users and safety testing conducted with relaxed restrictions for vetted partners and authorities. It should not be assumed that the expanded cyber capabilities Mistral provides in its testing environment are already open to ordinary API users. ## Coding, Finance, and Legal Work Involve Different Comparisons In coding, Mistral announced that Large 4 scored 61.7% on DeepSWE v1.1 and 28.3% on Terminal-Bench 4, which evaluates complex terminal operations. In the published DeepSWE comparison chart, DeepSeek V4 Pro 0813 uses Codex, GLM-5.3 uses OpenCode, and Kimi K3 uses Kimi Code CLI. Because the tools and execution procedures given to each model differ, it is difficult to read the chart as a ranking of standalone model performance under identical conditions. A footnote to the chart also states that the results were evaluated by Artificial Analysis before the execution environment was made publicly available. The differences in execution environments cannot be ignored. The [public DeepSWE ranking](https://deepswe.datacurve.ai/) runs all models on mini-swe-agent and lists GLM-5.3 and Kimi K3 at 69% each, but Large 4 is not yet listed. To judge superiority by directly comparing the figures in Mistral's chart with the public ranking, the task sets and execution conditions would need to be aligned. In legal work, third-party public evaluations allow a more concrete comparison. On Vals.ai's Harvey's Legal Agent Benchmark, Large 4's task pass rate was 15.83%. That is higher than Kimi K3's 12.92%, GLM-5.3's 8.33%, and GPT-6 Astra's 5.42%. Gemini 4 Argon, however, scored higher at 19.58%, and Muse Spark 1.2 higher still at 25.42%. In this evaluation, models produce deliverables in a shared environment with internet access disabled, and a task passes only if it meets every condition set for it. The figure is the average of pass rates calculated by two AI graders, so 15.83% cannot be restated as the "correct-answer rate on legal questions." Even if a model satisfies many conditions, missing just one fails the whole task. In corporate practice, this ability to meet multiple conditions through to the end and complete the work matters separately from sheer knowledge. In finance and science too, progress over earlier Mistral models needs to be separated from global leadership. In the Finance Agent v2 chart Mistral published, Large 4 scored 54.7% and Medium 3.5 scored 32.1%, and Large 4 also beat GPT-6 Astra's 53.5%. GLM-5.3, however, is higher at 55.8%. In the company's published chart for SciCode-Verified, which tests implementing scientific processing as code, Large 4 scores 91.8%, but GLM-5.3 scored 92.5% and the 2.4T A95B version of Qwen3.8 scored 93.8%. All of these are figures from evaluation charts Mistral itself presented. In image understanding, the model is also intended for analyzing drawings and satellite imagery. According to Mistral, on Dense 200, which tests locating objects in images, Large 4 scored 42% and GPT-6 Astra scored 41%. The difference is only one point, though, and this result alone does not support a conclusion that Large 4 is better at image understanding in general or that the difference is statistically clear. When companies decide whether to adopt it, it is more practical to test different kinds of work separately for their own use cases, such as reading drawings, completing spreadsheets, and fixing code. ## Pricing Is Lower Than Medium 3.5 but Higher Than Large 3 Large 4's standard API pricing is $1.36 per million input tokens and $4.18 per million output tokens. For the first two weeks after launch, a 50% discount applies, bringing it to $0.68 for input and $2.09 for output. Within Mistral's own API, the standard price is higher than Large 3 and lower than Medium 3.5.
| Model / pricing tier | Per 1M input tokens | Per 1M output tokens | Total for 1M input + 1M output | | --------------------------------- | ----------------------- | ----------------------- | ---------------------------- | | Mistral Large 3 | $0.50 | $1.50 | $2.00 | | Mistral Medium 3.5 | $1.50 | $7.50 | $9.00 | | Mistral Large 4 (standard) | $1.36 | $4.18 | $5.54 | | Mistral Large 4 (launch discount) | $0.68 | $2.09 | $2.77 |
Calculated in standard US dollar unit prices based on the [official pricing table](https://docs.mistral.ai/inference/pricing) as of October 7, 2026. It excludes cached input and additional conditions such as region selection, and Batch and Priority pricing are treated separately. In a simple calculation using 1 million input tokens and 1 million output tokens at standard pricing, Large 4 comes to $5.54. That is about 38.4% cheaper than Medium 3.5's $9.00 but higher than Large 3's $2.00. The reduction is calculated as (9.00 − 5.54) ÷ 9.00 × 100. In other words, the arrival of the new flagship did not cut prices across Mistral's API. Instead, it added a candidate for more advanced work at a lower unit price than Medium 3.5. However, an estimate that assumes the same token volume does not match the actual cost of completing a given job. If a model generates long outputs for reasoning or fails and has to redo the work, the total cost rises even with a low unit price. When switching models, it is necessary to check how often the model can complete the work when given the same documents or code, and how many tokens that takes. According to Mistral, Large 4 was trained from scratch using 3,800 NVIDIA Grace Blackwell GPUs in its own data centers in Europe, and the current preview is served from the same infrastructure. Reinforcement learning is still continuing, and the company says it will keep improving the model's handling of long-running work involving code execution and external tools, using the model's own trial results. A mechanism that makes the training environment easy to expand could also become a foundation for developing industry-specific models in the future. However, further performance gains are Mistral's future plan and should be considered separate from results already demonstrated. The release of the model weights planned for the end of October will be the next major checkpoint for companies. Large 3 is distributed under Apache 2.0 and Medium 3.5 under a Modified MIT license, but it has not yet been decided that the same terms will apply to Large 4. Once the weights are released, companies will need to check the license terms and required hardware configuration, and verify whether the model can complete work through to the end on their own Japanese-language documents and actual tasks. If performance and operating costs fit a company's use case, and it can control where data is stored and how the model is run, Large 4 could become a strong option for companies that want to reduce their dependence on external APIs.