🔍 Read the full analysis: AI Frontier Progress Leaves Mistral Large 4 Trailing on ThorstenMeyerAI.com
Prime made for students and young adults
- Fast, free delivery for dorm and study essentials
- Prime Video and Amazon Music included
- Member-only deals
TL;DR
Mistral released Mistral Large 4 as an API preview on October 6, 2026. Artificial Analysis gave it an Intelligence Index score of 38, behind leading US models and several Chinese competitors; the score is a dated benchmark snapshot, not a direct measure of performance on every task. The model’s weights are scheduled for release later in October, and the source’s criticism of its reliability is based partly on personal, uncontrolled testing.
Mistral released Mistral Large 4 in public preview on October 6, but an October 7 benchmark snapshot places it below leading US models and several Chinese competitors. Artificial Analysis scored the preview 38 on its Intelligence Index, a result that raises questions about its fit for demanding agentic work while leaving performance on specific workloads to be tested.
Mistral describes Large 4 as its largest model to date: a mixture-of-experts system with one trillion total parameters and 49 billion active parameters, part of Mistral’s European AI ambitions. It accepts text and images and is currently available through a preview API. Mistral said it trained the model on its own infrastructure in Europe and continues to improve it. Its weights are scheduled for release later in October; they were not downloadable at the time of the source report.
In the Artificial Analysis comparison available October 7, Claude Opus 5.5 scored 58, Gemini 4 Argon scored 53, and GPT-6.1 Sol scored 52. China’s GLM-5.3 and Kimi K3 scored 45 and 44, respectively. DeepSeek V4.1 Flash scored 39, while Mistral Large 4 Preview and GPT-6 Luna each scored 38. Cohere Command A+ scored 13. These figures are index points, not percentages, and the models were assessed at different reasoning settings rather than identical compute budgets.
The source author, Thorsten Meyer, said the benchmark results and his own use led him not to choose the preview for demanding, long-running agentic work when higher-scoring alternatives were available. Meyer also reported encountering hallucinations during his testing. That observation is personal experience, not a controlled comparative study. Artificial Analysis reports a context capacity of roughly 512,000 tokens, but capacity alone does not establish accuracy across long inputs.
Benchmark Gaps Shape Model Choices
The result matters to developers weighing whether to use Mistral Large 4 for work that requires a model to plan, use tools, interpret results and carry decisions across multiple steps. In such workflows, an unsupported assumption early in a run can affect later actions, and a fluent final response does not by itself show that the process was reliable. The benchmark provides one comparative signal, but it does not directly measure success on a developer’s particular coding, research or business tasks.
The comparison puts Mistral behind several highly ranked alternatives, including models from the United States and China. That may influence evaluation and procurement decisions, but it does not prove that Mistral will fail a given task or that every competitor is better for every use. The source’s own comparison shows an exception: Cohere Command A+ scored below Mistral on this index. Cost, deployment needs, privacy, latency and workload-specific results can also affect model selection.
Mistral’s European training infrastructure is relevant to the region’s AI capacity, but it is separate from a claim about benchmark leadership. The preview’s current score and the company’s ongoing work should be treated as a snapshot, not a final assessment of the model or its eventual weights.
AI model performance benchmarking tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A Preview Before Weight Release
The release is not yet a public download of model weights. As of the source report, developers could access Mistral Large 4 through a preview API, while Mistral said weights would follow later in October. That distinction matters: API users can test the current preview, but cannot yet evaluate or run the scheduled weight release themselves.
The Intelligence Index figures are a dated comparison assembled by Artificial Analysis and reported as available on October 7, 2026. The source notes that the named models used different reasoning settings, so the scores do not represent tests conducted under identical compute budgets. Developer locations in the table identify the companies, not where individual API requests are processed.
Meyer’s recommendation is an assessment rather than an independent benchmark finding. He said he would begin with stronger-scoring alternatives for complex autonomous work, while acknowledging that aggregate scores do not establish reliability on his own workflows. Mistral, for its part, advertises strengths in agentic coding and specialized professional tasks; those claims require testing on the workloads in question.
“I would not choose it for demanding agentic work or long tasks when stronger models are available.”
— Thorsten Meyer, ThorstenMeyerAI.com
As an affiliate, we earn on qualifying purchases.
Reliability Beyond Index Scores
The available comparison does not show how the models perform on identical task sets, reasoning budgets or real-world agent workflows. An Intelligence Index score of 38 does not prove failure on a specific coding or research job, and it cannot establish how often the model hallucinates in a given deployment.
Meyer reported hallucinations in his own use, but did not describe a controlled comparative study or provide rates that could be compared across models. The source material also does not provide a complete cost comparison, despite discussing cost as a consideration. DeepSeek V4.1 Flash is described as having a much lower measured cost per task, but the underlying figures and measurement details are not supplied here.
It is also unclear how much the preview will change before the scheduled weight release, whether the released weights will match the API version, or how the model will score in later evaluations. The benchmark is a dated snapshot, and scores may change as models and evaluations are updated.
As an affiliate, we earn on qualifying purchases.
Weight Release and Further Testing
Mistral said the model’s weights were scheduled for release later in October 2026. Until that release, developers can assess the preview through the API, subject to its availability and terms. The source material does not provide a specific release date or details about the planned weight licence.
For developers considering the model, the next useful step is to test it against their own representative tasks, including multi-step tool use, constraint following and evidence checking. Comparisons should account for cost, latency, reasoning settings and supervision needs, rather than relying on a single aggregate score. Later benchmark updates and independent evaluations could clarify whether Mistral’s continued improvements change the current comparison.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is Mistral Large 4’s current benchmark score?
Artificial Analysis reported an Intelligence Index score of 38 for Mistral Large 4 Preview in data available on October 7, 2026. The score is an index result, not a percentage or a guarantee of performance on a particular task.
Is Mistral Large 4 available as downloadable weights?
Not at the time described in the source report. Mistral had opened a preview API and said the weights were scheduled for release later in October 2026. A specific release date was not given.
Does the score mean the model will fail at agentic tasks?
No. The benchmark provides an aggregate comparison, but it does not directly test every developer’s agent workflow. The source author recommends stronger-scoring alternatives for demanding, long tasks, while noting that workload-specific testing is still needed.
What did the source author report about hallucinations?
Thorsten Meyer said he encountered hallucinations while using the preview. He described this as his own experience, not a controlled comparative study; the source gives no measured hallucination rate.
How does Mistral compare with the models listed?
In the October 7 snapshot, Mistral’s score of 38 was below Claude Opus 5.5, Gemini 4 Argon, GPT-6.1 Sol, GLM-5.3, Kimi K3 and DeepSeek V4.1 Flash. It matched GPT-6 Luna and scored above Cohere Command A+. The tests used different reasoning settings, so the comparison is not under identical compute budgets.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
