Mistral Large 4 For AI: A Standout Beyond The US And China, Not For Agents
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Mistral Large 4 For AI: A Standout Beyond The US And China, Not For Agents on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get hardware and tech essentials delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Mistral Large 4, released as a research public preview, scored 38.4 on Artificial Analysis Intelligence Index v4.3.2, below leading US and Chinese models but ahead of some older competitors. The source argues that its cost, high output volume and reported hallucinations make it a questionable choice for long-running AI agents; these concerns are not all independently measured by the index.

Mistral has released Large 4 as a research public preview, scoring 38.4 on the Artificial Analysis Intelligence Index v4.3.2. The result makes it a strong European entry in the source’s comparison, but leaves it below major US and Chinese models and raises questions about its price and use in multi-step AI agents.

The source describes Large 4 as a one-trillion-parameter model, with 49 billion active parameters, native text-and-image input, text output and a 512,000-token context window. It is available through Mistral’s API as a research public preview. Mistral has said its weights are expected at the end of October; the source says the licence has not been published.

Artificial Analysis’s current index puts Large 4 at 38.4 points. In the figures provided, that is below US models including Claude Opus 5.5 at 57.6 and GPT-6 Astra at 52.7, as well as Chinese models GLM-5.3 at 44.8, Kimi K3 at 43.6 and DeepSeek V4.1 Flash at 39.5. It scores above some earlier models, including DeepSeek V4 Pro 0813 at 36.0.

The source reports standard API pricing of $1.36 per million input tokens and $4.18 per million output tokens, with cached input priced at $0.14. A 50% discount is offered for the first two weeks. It also says Artificial Analysis measured $1.13 per Intelligence Index task for Large 4, compared with $0.25 for GLM-5.3-Flash and $0.27 for DeepSeek V4.1 Flash. Those task costs and index scores are figures reported in the source, not a guarantee of costs for every workload.

At a glance
reportWhen: Released yesterday, according to the so…
The developmentMistral released Large 4 as a research public preview, with an Artificial Analysis index score that places it behind current leading US and Chinese models.
Crypto market snapshot
Fear & Greed Index
71/100 — Greed
Bitcoin BTC$85,081▼ 0.6%
Ethereum ETH$2,681▼ 1.0%
Tether USDT$1▲ 0.0%
BNB BNB$775.21▼ 0.9%
XRP XRP$1.49▼ 0.8%
USDC USDC$1▲ 0.0%
Solana SOL$119.68▼ 0.8%
TRON TRX$0.3349▼ 0.4%
Live data · CoinGecko · alternative.me (24h change)
Mistral Large 4: Not a Frontier Model — Reality Check
AI Dispatch · Reality Check · 7 October 2026

Mistral Large 4: best outside the US and China — and still not a model to run your agents on

The headline is true: France has the most intelligent model outside the US and China. The independent data says the rest: every US and Chinese flagship scores higher, the best by 19 points. It costs 4× more per task than Chinese open models that outscore it, and it’s 2.5× as verbose as the median model.

Artificial Analysis Intelligence Index v4.3.2 — same version, like for like
Claude Opus 5.5 US57.6
Claude Sonnet 5.5 US56.0
Claude Fable 5.1 US53.4
GPT-6 Astra US52.7
Gemini 4 Argon US52.6
GPT-6.1 Sol US51.8
GLM-5.3 CN · open44.8
Kimi K3 CN · open43.6
GLM-5.3-Flash CN · open41.8
DeepSeek V4.1 Flash CN · open39.5
Mistral Large 4 (Preview) FR38.4
GPT-6 Luna US · small model~38
DeepSeek V4 Pro 0813 CN36.0
GLM-5.2 CN33.7
vs US frontier
−19.2 pts

~two-thirds of Opus 5.5. Level with OpenAI’s small model, Luna.

vs China open
8th

Eighth among open models once weights ship — behind seven Chinese ones. Beats GLM-5.2 and V4 Pro, loses to their successors.

vs Canada
n/a

Cohere doesn’t compete at this tier — reported ~14% hallucination at ~9% accuracy, because it declines most questions. A field of one.

The cost problem is worse than the intelligence problem — $ per Index task
Mistral Large 4
$1.13
Index 38.4 · $0.57 launch promo
GLM-5.3-Flash
$0.25
Index 41.8 · 4.5× cheaper
DeepSeek V4.1 Flash
$0.27
Index 39.5 · 4.2× cheaper
Gemini 4 Argon
~$1.99
Index 52.6 · +14 points
Per-token pricing looks competitive ($4.18/M output, well under the $10 median) — but it burns 200M output tokens on the Index vs an 81M median. Cheap tokens × 2.5 as many tokens is not a cheap model.
Why not for agentic or long-running work
The gap compounds
19 pts behind

The Index is now agentic-heavy — Briefcase, GDPval, AutomationBench, Terminal-Bench. Errors multiply across steps: tolerable in chat, fatal over a two-hour run.

AA v4.3.2
Verbosity
200M vs 81M

Output tokens to complete the Index. On an agent, verbosity is cost and latency on every step.

AA
Hallucination is back
observed

Confident false assertions in hands-on use. US frontier has largely moved past this — Gemini 4 Argon: 15%. In fairness Chinese open models are worse (Kimi K3 51%, DeepSeek V4 Pro 94%). In an agent, a fabrication is a wrong premise every later step builds on.

AUTHOR’S TESTING · not an AA figure
✓ What it’s genuinely good at
  • Cyber defence: 50 on the AA Cyber Index; 82% CyberGym-E2E (ahead of Luna’s 78%). Likely top-3 open model on cyber.
  • Documents & images: 19% GDP.pdf (+18 vs Large 3); 100 images per request.
  • Speed: 116 tok/s, 1.46s TTFT — well above median.
  • The jump: Large 3 scored 9 on this Index. 9 → 38 is real progress.
  • Jurisdiction: French parent, EU hosting, weights promised end of October.
▸ Who should actually use it
  • Legally bound buyers (defence, classified, DORA, health data): now the best European option by a wide margin. Wait for the weights, check the licence, pilot on cyber and documents.
  • Everyone else, for agentic or long tasks: don’t. A US frontier model is meaningfully more capable; GLM-5.3-Flash is more capable and 4× cheaper.
  • Note: Preview — Mistral says RL is still running, so scores may move. That changes next month’s decision, not today’s.
The take

Mistral says it has “essentially closed the gap.” It has closed the gap to where the Chinese open-weights field was a few months ago, while that field and the US frontier have both moved on. On every independent measure that matters for agents — intelligence, cost per task, verbosity and factual reliability — Large 4 is not a frontier model. “Most intelligent outside the US and China” is true mainly because almost nobody else outside those two countries is competing. Use it if you have to. Don’t use it because of the headline.

Sources: Artificial Analysis — Mistral Large 4 article & model/provider pages (6 Oct 2026), Index v4.3.2, comparison data; Trending Topics independent-ranking analysis; AA-derived reporting for frontier scores and AA-Omniscience rates (Argon 15%, Kimi K3 51%, DeepSeek V4 Pro 94%); Cohere profile as reported by Suprmind. Mistral Large 4’s AA-Omniscience result isn’t published in text — the hallucination point is the author’s own testing. Preview scores may change. Not investment advice.
thorstenmeyerai.com

A European Model With a Clear Gap

Large 4’s score marks a substantial improvement over Mistral’s previous models in the same index version: the source gives 9 points for Large 3 and 14 for Medium 3.5. The jump to 38.4 suggests Mistral has made significant progress, while the comparisons show it has not reached the performance of the leading US and Chinese systems assessed.

The result also matters for organisations seeking alternatives to US- and China-based AI providers. Mistral is a prominent European developer, but the source’s comparison indicates that geographic choice does not by itself establish competitive performance. Buyers will need to weigh capability, data-handling requirements, pricing, licensing and the tasks they plan to run.

The source raises a separate concern for agentic use. Artificial Analysis’s index includes benchmarks for knowledge work, software workflows and coding tasks. The author reports that Large 4 produced 200 million output tokens across index tasks, against a median of 81 million for comparable models. If that difference holds for a customer’s workload, higher token use could add cost and latency. The figure alone does not establish how the model will perform on every agent task.

Amazon

AI model API access

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Preview Status and Benchmark Basis

The score cited here comes from Artificial Analysis Intelligence Index v4.3.2, described by the source as its current version. Using the same index version helps make the listed model scores comparable, but benchmark results are not a complete measure of performance across different applications or operating conditions.

Large 4 remains a preview rather than a fully released open-weights model. Until Mistral publishes the weights and licence, customers cannot treat the planned release as an available option for self-hosting under known terms. Mistral has also said reinforcement learning is ongoing, according to the source, and that scores may change as development continues.

Amazon

large language model API

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Costs, Hallucinations and Final Scores

Several concerns in the source are not established by the benchmark table alone. The author says hands-on testing found confident hallucinations, but supplies no test protocol, sample size or independently verifiable results for that observation. The article also cites hallucination rates for other models from a separate measure; these figures do not establish Large 4’s rate on that measure.

The source’s index-task cost and output-token figures may not reflect every customer’s prompts, task mix or API usage. It is also unclear how much ongoing reinforcement learning will change the score, what the final weights’ licence will permit, and whether the temporary 50% discount applies to all users or usage patterns. The supplied source ends mid-comparison after introducing the cost of Gemini 4 Argon, so no complete cost comparison with that model can be reported.

Amazon

text and image AI input device

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Weights and Pricing After Preview

Mistral has said it plans to release Large 4’s weights at the end of October. The timing and licence terms will determine whether developers can run the model outside Mistral’s API and what uses will be permitted. The company’s ongoing reinforcement learning may also lead to revised benchmark results, according to the source.

For now, prospective users can assess the API preview against their own tasks and costs, while treating the reported index results as one comparison point. The source does not provide a confirmed date for a final release or further benchmark update.

Amazon

AI token counter tool

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Mistral Large 4?

It is Mistral’s one-trillion-parameter model, with 49 billion active parameters, text-and-image input, text output and a 512,000-token context window. The source says it is available as a research public preview through Mistral’s API.

How did Large 4 score against leading models?

It scored 38.4 on Artificial Analysis Intelligence Index v4.3.2. In the source’s table, leading US models and several Chinese models score higher, including DeepSeek V4.1 Flash at 39.5.

Is Large 4 open source or available to self-host now?

The source says Mistral plans to publish the weights at the end of October, but they were not yet available in the reported preview. It also says the licence had not been published.

Is Large 4 suitable for AI agents?

The source questions its suitability for long-running agents, citing its index score, reported output-token use and the author’s own observations of hallucinations. Those observations are not a complete independent evaluation, so teams would need to test the model on their specific workflows before drawing conclusions.

What does Large 4 cost to use?

The source reports standard API rates of $1.36 per million input tokens, $4.18 per million output tokens and $0.14 per million cached input tokens, plus a 50% discount for the first two weeks. Actual costs depend on usage and the applicable pricing terms.

Source: ThorstenMeyerAI.com

Nothing in this article is financial or investment advice. Cryptocurrency and precious-metal investments carry significant risk — do your own research and consider a licensed advisor.
HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Is The Future Of AI Less About Writing And More About Systemic Thinking?

TypeSafe AI introduces Jev, a decision-focused AI model emphasizing systemic thinking over traditional text generation, signaling a shift in enterprise AI development.

Is The $400 Million Public AI Initiative A Step Toward Sovereignty Or Subsidy Rhetoric?

A detailed analysis of France’s $400 million public AI program—does it foster AI sovereignty or serve as a subsidy with limited impact?

Why Only 34 Of 64 Changes Generalize In LLM-Engineered Agent Harnesses, ByteDance Seed Explains

ByteDance Seed’s HarnessDev project shows only 34 of 64 automated harness modifications by large language models transfer beyond initial conditions, highlighting limits in AI self-engineering.

The conversion. What turning the largest nonprofit into a company did to charity law.

OpenAI’s recent transformation kept control of its assets, challenging traditional charity laws and raising questions about future nonprofit conversions.