Mistral Large 4 Public Preview: 1 trillion parameters, weights open at the end of October—can it run locally?

Mistral Large 4 Public Preview: 1 trillion parameters, weights open at the end of October—can it run locally?

Oct 7, 20267 min read
Categories:Tech
Tags:#AI#Cutting-edge models

Info: Machine translation This post was machine-translated from the Chinese original. Wording may be rough in places — the Chinese version is authoritative.

What did Mistral Large 4 launch on October 6?

Mistral Large 4 opened its preview API on October 6, and the official plan is to release weights by the end of October 2026. Today, you can try remote calls on Mistral Studio, not download models to your own computer. Opening weights at the end of the month is just the beginning: licenses, inference software, and device resources must match, and you can’t guarantee that a regular Mac will run it at that time.

Officially, it is defined as a hybrid architecture with 1 trillion parameters and native multimodality, with about 49 billion activation parameters. The same announcement emphasized that 3,800 NVIDIA Grace Blackwell GPUs were used for training, with training and preview services running in Mistral’s own European data centers.

1 trillion total parameters versus about 49 billion activation parameters are most easily misunderstood as “only 49 billion parameters needed.” You can think of it as a small hospital with many specialties: a single visit only uses some departments, but other departments do not disappear. The hybrid expert model only calls some experts when processing a single input, reducing some computational load; Storage, memory, and data handling demands cannot be estimated solely from the activation volume of the current time. This difference is directly related to local deployment, not a negligible parameter detail.

Official announcement: total parameters 1 trillion, activation parameters 49 billion; Weighting planned to be released at the end of the month.

What are the strengths of the official self-reported areas?

Mistral promotes ML4 as “the strongest among open-weight models,” emphasizing cybersecurity, enterprise coding, and multimodal understanding. The following figures are from the original announcement and represent official report results, not independent verification.

Cybersecurity: Ranks among the top in open authority

  • The Artificial Analysis Cyber Index ranks among the top five globally and clearly leads in the “non-China openness weight” category.
  • On the index question “Reproducing and fixing real open-source software vulnerabilities,” it scored 82%, the highest among the announced models; On the same question, Claude Opus 5.5 and GPT-6 Astra scored “nearly zero” due to trigger rejection.
  • On Cybench (a benchmark of 40 security quizzes), the resolution rate is 93%.
  • Robustness to indirect prompt injection: Resisted 93.3% of attacks on Lakera’s public B3 AI Security Benchmark, with no higher results reported in the announcement.
  • On the average level of the three malicious request tests—JailbreakBench, StrongREJECT, and AgentHarm—ML4’s rejection rate is higher than that of the listed open-weighted control models.

Why should security rankings be based on testing methods? Fixing vulnerabilities usually requires first proving the problem exists in an authorized environment, then checking whether patches can block it. The official attributes some low scores for closed-source models to refusal to answer; This shows that the results mix task capability and security policy impact, and cannot be directly attributed to “all security tasks are stronger.” Mistral also said it is conducting red team testing with approved partners. Users should still limit the scope of authorization and not treat preview capability as an unrestricted attack tool.

Coding and proxy workflows

  • DeepSWE v1.1: 61.7%.

  • SWE-Atlas-QnA: 59.4%.

  • Terminal-Bench 4: 28.3%.

  • The three-component composite Code Agent Index is 49.8%, officially claiming it leads DeepSeek V4 Pro 0813 and Qwen 3.8 Max.

  • Surge AI blind review (5 models): ML4 preview score 3.74, second place; Comparison is Kimi K3 3.59, GLM-5.3 3.60, GLM-5.2 3.40, first place Claude Opus 5 4.22.

  • AutomationBench (657 business streams, covering Gmail, Google Sheets, Slack, Salesforce, and other applications): 59.9%.

  • AA-Briefcase (Long-Chain Knowledge Work Elo Score): 1,393.

Multimodal and visual localization

Official code review original text; For the complete benchmark chart, please check the source; this is not an independent retest.

Officially, ML4 can handle complex documents, charts, and natural images, combining visual and proxy capabilities. From gigapixel satellite image positioning, to magnifying engineering drawings for inspection, and PDF evidence retrieval, it is listed as a typical use case. For the Dense 200 visual positioning question, ML4 scored 42%, and GPT-6 Astra reported 41%.

Science and mathematics

The official statement states that ML4 is the best among open weights on SciCode-Verified; It also demonstrated multi-step chemistry tasks that generate a complete Hartree–Fock simulation in one go. In mathematics, the original announcement stated that ML4 is more precise and structured than GLM-5.3 inference, and can handle applied mathematics tasks related to theoretical physics for extended periods.

Knowledge work with legal affairs and finance

According to third-party evaluation agency vals.ai, ML4 outperforms GPT-6 Astra in both legal and financial representative tasks, and “surpasses all open-source models” on HarveyAI’s Legal Agent benchmark. FinWorkBench tests spreadsheet creation and editing capabilities, and officially released results for Finance Agent v2 and Finch (FinWorkBench).

Human preference evaluation

Original announcement: In coding, computer-aided design, and manual comparison of mathematics and physics, ML4 is preferred in CAD and STEM, while coding is “on par with or close to GLM-5.3” in finance. This aligns with the conclusion in the previous blind review section, where ML4 coding ranked second and Claude Opus 5 was first.

Training scale and the so-called “European sovereignty”

Officially, the model was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs, and the preview service ran in its own European data center. Post-training combined code execution, web retrieval, and various inspection methods to allow the model to receive feedback on task results. You could think of it as reviewing post-practice feedback: whether the file was written correctly or if the code passed checks, which is closer to real work than just judging whether the answer sounds correct. Training scale indicates R&D investment; running this model requires the same number of GPUs.

The announcement emphasized that ML4 will be available in multiple global regions, with European deployment operated by Mistral full-stack, independent from other digital service providers, governed by European law, and stated that the training corpora covers more than 160 languages, including all official EU languages.

It’s important to break this apart: the so-called ‘sovereign deployment’ currently refers to the European server operated by Mistral itself; After the weights are released at the end of the month, whether they can run in users’ own European data centers or local data centers will depend on the official announcement of architecture and license details. The announcement does not specify the requirements for VRAM, memory, or computing power for local deployment, so the local threshold cannot be verified at present.

API pricing and what can be done at this stage

The preview price for the announcement scraped on October 7 was: $1.36 per million input words, $4.18 per million output words. Word words are the unit charged after the model splits text and do not correspond fixedly to a single Chinese character or word. If a task accumulates 1 million characters input and 100,000 characters output, the unit price for these two items is about $1.778; This is a hypothetical example, not the actual payment for this bill, and does not include other service fees. Long tasks should look at the cumulative usage, not just the last short paragraph of answers.

Currently, what can be done by capability is: calling preview APIs in Mistral Studio, doing application integration and evaluation; and forming a cybersecurity red team with approved partners under a “more relaxed security policy.” Local self-deployment, enterprise intranet deployment, and secondary distribution after opening weights will all be discussed after the weighting and license window is released at the end of the month.

What can these achievements prove now, and what can’t be proven?

  • The percentages and rankings above are from Mistral announcements, some of which reference third-party reviews; This time, we did not review the original evaluation and reproduction conditions one by one, nor did we complete an independent horizontal review. Public announcements do not mean that the code, weights, and test records are all open source.
  • Comparisons like “the strongest among open weights” or “better than DeepSeek V4 Pro 0813 and Qwen 3.8 Max” only hold true on the benchmarks and competitors Mistral chooses itself; Different benchmarks, different prompts, and different inference parameters may lead to reversal of conclusions.
  • “Security models surpass Claude Opus 5.5 and GPT-6 Astra in cybersecurity” because some closed-source models refuse to perform such tasks, not because ML4 leads in all tasks; By the same logic, Claude Opus 5 still ranks first in the Surge AI blind review of coding problems.
  • “Sovereign deployment” and “European legal jurisdiction” currently cover Mistral’s self-operated European services; The secondary distribution boundary after the weight opening will await official license details.
  • “1 trillion parameters, 49 billion activations” cannot be estimated based solely on activation parameters. What weights to store, how to load experts, and what precision to operate must be combined with post-release architecture and software solutions; The current announcement does not provide a testing threshold for consumer-grade devices.
  • “Understanding Gigabit Satellite Images and Engineering Drawings at Once” is an official demo scenario. The announcement does not disclose specific visual resolution, prompt templates, or manual verification success rates. Do not treat demonstrations as stable capabilities before local testing.

What should we do next to view this update?

For those who want to use it today, the most direct way is to apply for a preview API trial at Mistral Studio. The price is estimated at $1.36 / $4.18 per million tokens, rather than making it a long-term commitment.

For those who only have one Mac on hand, now is not worth buying hardware just to “unlock weights at the end of the month.” Besides weights, runtime also consumes cache and image processing resources mentioned earlier; total parameters, accuracy, expert scheduling, and read speed all affect the experience. First, wait for the official architecture, license, and available running plan before testing startup, answering, and continuous tasks on your own device. What can be confirmed at this stage is that remote preview has been launched, but the local threshold has not been verified, so a 64GB Mac cannot be guaranteed.

For those concerned about cybersecurity workflows, the focus is on the red-teaming results Mistral performs under a “more relaxed security policy” with cybersecurity companies and European government agencies, which is the most differentiating and controversial aspect of ML4; Before making a selection, assess the compliance boundaries of your jurisdiction regarding capabilities like “reproducing and patching vulnerabilities.”

Tip At the end of the announcement, Mistral said it would release weights, complete architectural details, add evaluations, and post-training methods by the end of the month. Local feasibility, license boundaries, and whether it can truly replace the existing stack will be judged after the batch of materials is released.

Official information


🧰 Tools I build

I maintain all of these tools myself. Preview builds are clearly labeled; the release pages are the source of truth for downloads, updates and known limits.

Info: HyphenBox Status: Official releases

A radar for free LLM APIs: availability is re-tested continuously, one local interface for all of them, and keys stay on your machine

Downloads & updates

Info: LocalBrain Status: Official releases

A multimodal MCP toolbox for local models: TTS, Whisper and video generation in one place

Downloads & updates

Info: ScreenLex Status: Official releases

Learn new words while you watch shows. Free, for Mac and Windows

Downloads & updates

Info: HyphenScreen Status: Official releases

Screen recording and smart editing in one: a DaVinci-style timeline, automatic redaction and a check of the finished video before export. Free

Downloads & updates


Quote: HyphenTech Make AI your superpower Local deployment · Free resources · Self-made software https://hyphentech.top

Share:

Comments

Loading comments…

Back to home