AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Mistral Large 4: What Its Strength Outside The US And China Means For AI on ThorstenMeyerAI.com

Before you orderOffer from Amazon

Get the latest gadgets delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Mistral released Large 4 as a research preview, scoring 38.4 on Artificial Analysis’s Intelligence Index. The result makes it a leading model from outside the US and China, but the source’s benchmark data puts it behind current US and Chinese flagships and shows lower-cost models outperforming it on the same index.

Mistral AI released Mistral Large 4 as a research preview, scoring 38.4 on Artificial Analysis’s Intelligence Index v4.3.2. The result makes it the highest-scoring model in the source’s comparison from outside the United States and China, but it remains below major US and Chinese competitors, complicating claims that Europe has a frontier-level alternative.

The model has 1 trillion total parameters, with 49 billion active, and supports text and image input with text output. Mistral lists a 512,000-token context window. It is currently available through the company’s API as a research preview; Mistral has said it plans to release the weights at the end of October. The source says the model’s license had not been published at the time of writing.

On Artificial Analysis’s current index, Large 4’s 38.4 score is up from 9 for Mistral Large 3 and 14 for Medium 3.5 on the same index version, according to the source. But the listed US leaders score between 52.6 and 57.6, while Chinese models including GLM-5.3, Kimi K3 and DeepSeek V4.1 Flash also rank ahead. The source characterizes Large 4 as roughly level with OpenAI’s smaller GPT-6 Luna model.

Mistral’s listed API prices are $1.36 per million input tokens and $4.18 per million output tokens, with cached input at $0.14. The source reports a 50% discount for the first two weeks. It also says Artificial Analysis measured $1.13 per Intelligence Index task for Large 4, compared with $0.25 for GLM-5.3-Flash and $0.27 for DeepSeek V4.1 Flash; both of those models scored higher on the index. These task-cost estimates reflect the benchmark workload, not a fixed price for every customer use case.

At a glance
reportWhen: Released yesterday, according to the so…
The developmentMistral launched Large 4 in a research preview, with benchmark results positioning it as a strong European model but not a match for leading US and Chinese systems.
Mistral Large 4: Not a Frontier Model — Reality Check
AI Dispatch · Reality Check · 7 October 2026

Mistral Large 4: best outside the US and China — and still not a model to run your agents on

The headline is true: France has the most intelligent model outside the US and China. The independent data says the rest: every US and Chinese flagship scores higher, the best by 19 points. It costs 4× more per task than Chinese open models that outscore it, and it’s 2.5× as verbose as the median model.

Artificial Analysis Intelligence Index v4.3.2 — same version, like for like
Claude Opus 5.5 US57.6
Claude Sonnet 5.5 US56.0
Claude Fable 5.1 US53.4
GPT-6 Astra US52.7
Gemini 4 Argon US52.6
GPT-6.1 Sol US51.8
GLM-5.3 CN · open44.8
Kimi K3 CN · open43.6
GLM-5.3-Flash CN · open41.8
DeepSeek V4.1 Flash CN · open39.5
Mistral Large 4 (Preview) FR38.4
GPT-6 Luna US · small model~38
DeepSeek V4 Pro 0813 CN36.0
GLM-5.2 CN33.7
vs US frontier
−19.2 pts

~two-thirds of Opus 5.5. Level with OpenAI’s small model, Luna.

vs China open
8th

Eighth among open models once weights ship — behind seven Chinese ones. Beats GLM-5.2 and V4 Pro, loses to their successors.

vs Canada
n/a

Cohere doesn’t compete at this tier — reported ~14% hallucination at ~9% accuracy, because it declines most questions. A field of one.

The cost problem is worse than the intelligence problem — $ per Index task
Mistral Large 4
$1.13
Index 38.4 · $0.57 launch promo
GLM-5.3-Flash
$0.25
Index 41.8 · 4.5× cheaper
DeepSeek V4.1 Flash
$0.27
Index 39.5 · 4.2× cheaper
Gemini 4 Argon
~$1.99
Index 52.6 · +14 points
Per-token pricing looks competitive ($4.18/M output, well under the $10 median) — but it burns 200M output tokens on the Index vs an 81M median. Cheap tokens × 2.5 as many tokens is not a cheap model.
Why not for agentic or long-running work
The gap compounds
19 pts behind

The Index is now agentic-heavy — Briefcase, GDPval, AutomationBench, Terminal-Bench. Errors multiply across steps: tolerable in chat, fatal over a two-hour run.

AA v4.3.2
Verbosity
200M vs 81M

Output tokens to complete the Index. On an agent, verbosity is cost and latency on every step.

AA
Hallucination is back
observed

Confident false assertions in hands-on use. US frontier has largely moved past this — Gemini 4 Argon: 15%. In fairness Chinese open models are worse (Kimi K3 51%, DeepSeek V4 Pro 94%). In an agent, a fabrication is a wrong premise every later step builds on.

AUTHOR’S TESTING · not an AA figure
✓ What it’s genuinely good at
  • Cyber defence: 50 on the AA Cyber Index; 82% CyberGym-E2E (ahead of Luna’s 78%). Likely top-3 open model on cyber.
  • Documents & images: 19% GDP.pdf (+18 vs Large 3); 100 images per request.
  • Speed: 116 tok/s, 1.46s TTFT — well above median.
  • The jump: Large 3 scored 9 on this Index. 9 → 38 is real progress.
  • Jurisdiction: French parent, EU hosting, weights promised end of October.
▸ Who should actually use it
  • Legally bound buyers (defence, classified, DORA, health data): now the best European option by a wide margin. Wait for the weights, check the licence, pilot on cyber and documents.
  • Everyone else, for agentic or long tasks: don’t. A US frontier model is meaningfully more capable; GLM-5.3-Flash is more capable and 4× cheaper.
  • Note: Preview — Mistral says RL is still running, so scores may move. That changes next month’s decision, not today’s.
The take

Mistral says it has “essentially closed the gap.” It has closed the gap to where the Chinese open-weights field was a few months ago, while that field and the US frontier have both moved on. On every independent measure that matters for agents — intelligence, cost per task, verbosity and factual reliability — Large 4 is not a frontier model. “Most intelligent outside the US and China” is true mainly because almost nobody else outside those two countries is competing. Use it if you have to. Don’t use it because of the headline.

Sources: Artificial Analysis — Mistral Large 4 article & model/provider pages (6 Oct 2026), Index v4.3.2, comparison data; Trending Topics independent-ranking analysis; AA-derived reporting for frontier scores and AA-Omniscience rates (Argon 15%, Kimi K3 51%, DeepSeek V4 Pro 94%); Cohere profile as reported by Suprmind. Mistral Large 4’s AA-Omniscience result isn’t published in text — the hallucination point is the author’s own testing. Preview scores may change. Not investment advice.
thorstenmeyerai.com

Europe’s Model Faces a Cost Test

Large 4 matters because it offers a European-developed option in a market where the most prominent frontier models come from US and Chinese labs. For governments and businesses concerned about supplier concentration or where AI infrastructure is developed, that may make Mistral’s progress relevant even if the model is not the benchmark leader.

The data also sets limits on what the launch establishes. A score of 38.4 is a substantial rise from Mistral’s earlier results cited by the source, but it does not put Large 4 alongside the top US systems. Nor does being the highest-scoring model outside the US and China, within this particular comparison, mean it leads the global field. The source’s own figures show multiple Chinese models scoring above it.

Price and performance could shape purchasing decisions as much as geography. The source’s benchmark cost estimates put two lower-priced Chinese models ahead in score, while Large 4 generated substantially more output tokens than the median comparable model. That combination may matter for agentic tasks, where repeated model calls can add expense and delay. Buyers will need to test their own workloads rather than treat one index score or cost-per-task estimate as a universal measure.

Amazon

AI language model API access

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A Preview, Not a Final Release

The release is still an early-stage offering. Mistral Large 4 is available through the API as a research preview, and its weights have not yet been released. According to the source, Mistral plans to publish them at the end of October, but the license was still unpublished. That leaves open how developers will be permitted to use, adapt or distribute the model once the weights arrive.

The source attributes the benchmark comparison to Artificial Analysis Intelligence Index v4.3.2, which it says includes agentic work tests such as knowledge work, software workflows and coding. It also reports that Mistral said reinforcement learning was still running, meaning scores could change. The benchmark result should therefore be read as a measurement of the preview at this point, not a settled assessment of a final model.

Claims about hands-on reliability require a separate distinction. The source author reports seeing confident false statements during personal testing, but that is an observation rather than a published Artificial Analysis result. The source also cites hallucination figures for other models from AA-Omniscience, but those measures are not interchangeable with the Intelligence Index score and do not, by themselves, establish how a model will behave in every deployment.

“The ‘Western alternative outside the US’ race is a field of one, and Large 4 wins it by default.”

— Source article author

Amazon

text and image input AI models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Weights, Licensing and Reliability

Several points remain unresolved. The planned weight-release date is at the end of October, according to the source, but the release had not happened when the material was written. The eventual license and any restrictions on commercial use are also unknown in the supplied information.

It is also unclear how much the score will move as reinforcement learning continues, or whether the final model will differ materially from the preview. The source does not provide enough detail to establish how broadly its author’s hallucination observations apply, or to compare that anecdotal testing directly with standardized reliability measures. Benchmark rankings and task-cost estimates may not predict results for a particular company’s prompts, tools or workloads.

Amazon

large language model with 512k token window

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

October Weights and Updated Scores

The next stated milestone is Mistral’s planned release of Large 4’s weights at the end of October. The publication of those weights and their license will clarify whether developers can run the model independently and under what terms. Mistral may also update the preview as reinforcement learning continues; fresh benchmark results would show whether its index score and relative ranking change.

For now, prospective users can evaluate the API preview against their own requirements, including accuracy, latency, token use, cost and data-handling rules. Any comparison should separate the model’s regional significance from its measured performance: Large 4 expands Europe’s visible presence in advanced AI, while the available data does not show it matching the leading US or Chinese systems.

Amazon

AI development tools for researchers

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Mistral Large 4?

Mistral Large 4 is Mistral AI’s multimodal model, released as a research preview through its API. The source describes it as having 1 trillion total parameters, 49 billion active parameters and a 512,000-token context window.

How does it rank against other models?

It scored 38.4 on Artificial Analysis Intelligence Index v4.3.2, according to the source. That leads the models in the source’s comparison from outside the US and China, but several US and Chinese models score higher.

Are Mistral Large 4’s weights available?

Not yet, according to the source material. Mistral planned to release the weights at the end of October, and the license had not been published when the source was written.

Is Large 4 cheaper than competing models?

Not by the source’s benchmark task-cost estimates. It reports a cost of $1.13 per Intelligence Index task for Large 4, compared with $0.25 for GLM-5.3-Flash and $0.27 for DeepSeek V4.1 Flash. Actual costs depend on usage and workload.

Source: ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Advanced Technology and Scientific Equipment: A Practical Guide to Choosing the Right Tools

AIThis post was created with the assistance of artificial intelligence (AI).Advanced technology…

Navigating AI Challenges Post-SpaceX Acquisition Of Cursor

OpenAI publicly announces its decision on Cursor after SpaceX’s acquisition, affecting AI coding tools and developer workflows.

The 10 Best Mirrorless Cameras Of 2026 Featuring AI Power

Discover the 10 best mirrorless cameras of 2026, featuring advanced AI capabilities for photographers and content creators across all levels.

Microsoft Backs Away From Copilot PC Branding As OEMs Abandon The Label

Microsoft is shifting away from the Copilot PC branding after OEMs stopped using the label, signaling a strategic change in its Windows AI branding approach.