Qwen4 preview open source: Even minimal quantization can't fit it all

Qwen4 preview open source: Even minimal quantization can't fit it all

Aug 26, 20268 min read
Categories:Thoughts
Tags:#AI#News

Info: Machine translation This post was machine-translated from the Chinese original. Wording may be rough in places — the Chinese version is authoritative.

Note This article was first published on the HyphenTech official WeChat account

Qwen4 preview open source: Even minimal quantization can’t fit it all

Qwen4 architecture preview open weights: each token activates 6B, but the current minimum GGUF is still 72.55GB. This article only analyzes official repositories, configurations, and file lists, and does not include native testing.

Note HyphenTech · 2026-08-26


🎯 ▍ 6B activation is tempting, but the hard drive first handed you a 360GB bill

# Qwen Official GitHub: Qwen3.8-Flash-Next is a multimodal MoE and an early preview of the Qwen4 architecture

After Qwen3.8-Flash-Next opened weights, the most eye-catching number isn’t 125B, but each token only activates 6B. This sounds like a big machine turning on only a few lights at a time, as if it could finally fit into a personal computer.

But the results from the File API are straightforward: 131 BF16 shards totaling 360,000,192,888 bytes, 360.00GB in decimal, 335.28GiB in binary.

The commonly seen “335GB” online is not another set of weights, but rather GiB written as GB. The two terms describe the same batch of files, with the difference coming from decimal and binary conversions. **The capacity unit is written incorrectly, and it looks like just a few letters, but when it comes to the hard drive, it can misjudge deployment. **

# Qwen Official Hugging Face Model Page: Page count about 180B storage parameters, pipeline for text-to-text

The size can actually be broken down by parameters: the 125B main model is about 250GB at BF16, the 51B N-gram embedding is about 102GB, and the 4B MTP is about 8GB, totaling about 360GB. There’s no mystical setup here, just a simple storage account. **Activation parameters determine how much weight each calculation carries, and storage parameters determine how much luggage the whole machine carries. **

After LLaMA’s weights leaked back then, quantization, fine-tuning, and on-device tools quickly grew into an ecosystem. It proved that open weights can accelerate the ability to sink downward, but it didn’t prove that any large model can immediately settle into ordinary computers. This time, Qwen proactively laid out the architecture, with the difference being that the ecosystem can be rapidly replenished, and physical memory doesn’t automatically expand due to community enthusiasm. **

**// The activation parameter is the compute ledger, the storage parameter is the in-memory ledger. **

$ 上手地址:Qwen3.8-Flash-Next 官方模型页面 https://huggingface.co/Qwen/Qwen3.8-Flash-Next

🔍 ▍ 125B, 51B, and 4B hold the true preview of Qwen4

This isn’t just a simple upgrade to the old architecture, but rather a Qwen4 preview car that has left the lab ahead of schedule.

The model_type in the configuration is qwen4_exp, and the main platform uses a multimodal MoE: 48 layers, 512 routed experts, 10 selected per token, plus shared experts. It puts the massive capacity in the background, only mobilizing a small portion of what is needed in the moment.

The attention layer is also not a neat replication. Every four layers contain 3 linear/GDN attention and 1 full/QSA attention.

Gated DeltaNet compresses history, while Qwen Sparse Attention uses a lightweight indexer to pick out important contexts by micro-block. The goal is clear: **Long contexts can’t rely on all tokens coming in pairs, or the cost will eventually drag you down. **

# Qwen Official Architecture Diagram: Three-layer GDN with one layer of QSA, plus gate control residual, MoE, and N-gram Embedding

Another major move is N-gram Embedding. It queries a table of 20,000,000 items based on local bigrams and trigrams, with the vocabulary itself containing 248,320 items, and is placed on the second layer.

The official design allows this large table to be left in host memory, where asynchronous prefetching covers part of the waiting time. This approach can reduce memory load, but it does not eliminate the system’s memory capacity constraints. **

Gated Residual expands the residual stream into four branches, with dynamic gates controlling read/write; On the optimization side, Muon and AdamW divide labor and refit scaling law.

The native context reaches 262,144 tokens, and the model card also provides a path to expand to 1,000,000 tokens. The direction is very clear: it targets long-range agents, office workflows, and multimodal interactions, not just a few more questions in the chat box. The focus of the Qwen4 preview is to redistribute computation, not to shrink storage.


⚡ ▍ 72.55GB is already very low, and 64GB still can’t fit it

The current smallest GGUF is UD-IQ1_S, divided into 3 files: 10,946,624 bytes, 49,990,818,368 bytes, and 22,544,696,352 bytes.

A total of 72,546,461,344 bytes, which is 72.55GB or 67.56GiB. It is already the smallest segment available; the regular IQ2, Q2, Q3, and Q4 will only be larger.

Based on about 180B storage parameters, this version averages only about 3.22 bits per parameter, still not pushing into 64GB. The reason is not mysterious: Dynamic Quant reserves higher precision for sensitive tensors, and 51B N-gram tables, 248,000-word tables, and non-uniform quantization tensors all take up space. **IQ1_S is the quantization level name, but it does not mean all parameters in the package use the same width. **

Equipment conditions Current borders Suggestion
64GB Mac The minimum file size is 72.55GB, which is lower than the runtime guide’s minimum of 78GB Continue using Qwen3.8-27B
96GB Mac Reach the current recommended starting point, only suitable for minimal quantization Waiting for official support and real speed
128GB Mac It allows for the system, KV cache, and higher precision Still not rushing to download the first release
RTX 4070 12GB Video memory is far from sufficient and still depends on the host system memory There is more than 80GB of memory that is partially offloaded after reassessment

※ Capacity boundaries are based on the current file size and runtime guide; Speed does not yet have local results.

What this table really wants to prevent is a common misjudgment: just because a file can be mmapped doesn’t mean it’s suitable for daily operation. Even if a 64GB machine boots via disk offload after the framework is perfected, the SSD may still keep switching pages; 512 expert visits will then expand the wait. **Being able to light up the boot screen and being able to work stably are two completely different things. **

# Actual warehouse file calculations: BF16 is 360.00GB, current minimum GGUF is 72.55GB; Unsloth recommends 78GB as minimum and 96GB as the minimum

Cramming large models into personal computers was originally the result of quantitative technology and inference frameworks being driven together. Only after billions of parameter levels became everyday options did cloud subscriptions begin to loosen. But this curve doesn’t just go downward: once the model swaps more capacity for expert, long-context, and multimodal support, **personal devices will crash back into memory walls. **64GB isn’t that you can’t tinker, but it’s not worth treating the hassle as productivity.


🧩 ▍ Having GGUF doesn’t mean stable performance; the first link is still one kilometer away

The file had been uploaded, but there was a noticeable time lag in the runlink. Qwen README states that llama.cpp supports text and visuals; On the same day, llama.cpp upstream support issue list remains open with #27741 and 0 comments. Unsloth’s guide also lists WIP and Coming very soon, pointing to specific PRs.

# Unsloth Launch Guide Original Page: At least 78GB RAM/unified memory, with clear indication that runtime support is still being developed

This doesn’t mean someone is wrong; it’s more like different levels of “support” are squeezed into the same word. The repository can recognize files, experimental branches can read architectures, but stable versions may not have already merged and overwritten text, visual, quantization, and offload. **Format recognition only gets the ticket; stable operation counts as real entry. **

The MLX route should also be judged with restraint. Currently, mlx-community does not have a corresponding conversion bin, and the model directory for MLX-LM main does not include qwen4_exp or qwen4 implementations.

The correct way is that it is “not yet ready,” and cannot be permanently deemed unsupported. Although Transformers, SGLang, vLLM, and TokenSpeed have boot commands, they are mainly aimed at multiple cards or servers and do not suddenly run out of consumer-grade memory.

Early factories each built their own generator sets, and as the unified grid matured, electricity gradually became a public capacity purchased on demand. Similar divisions of labor are emerging: large-capacity capacity is concentrated in the cloud, while small-capacity capacity remains in individual devices. The difference is that data and context are not ordinary currents; **some tasks are sent to the cloud, saving memory but giving away control. **

**// Just because the file arrives first doesn’t mean the toolchain has paved the way. **

$ 兼容状态:llama.cpp 问题单 #27741 https://github.com/ggml-org/llama.cpp/issues/27741

📉 ▍ The rankings are unevenly won; it resembles a specialized route for office intelligent agents

The most impressive official score is JobBench: 55.7, compared to Qwen 3.8-27B at 33.4, DeepSeek-V4-Flash at 41.3, and Claude Opus 4.6 at 36.6.

CoWorkBench also reached 73.9, higher than 27B’s 70.7, DeepSeek’s 45.1, and Claude’s 68.2. **Long-range tasks and office collaboration are indeed its most valuable areas. **

But the coding scores didn’t rank up in a complete crush. DeepSWE 1.1 was 58.7, 27B was 42.2, and DeepSeek was 54.4; In SWE-bench Pro, its 62.5 was only 0.8 higher than the 27B’s 61.7. NL2Repo-Bench even scored only 48.1, lower than DeepSeek’s 54.2. These counterexamples are more useful than the phrase “comprehensive is stronger.”

Multimodal and computer operations also show the same boundaries. OSWorld 2.0 binary is 19.4, exactly the same as 27B; Partial increases from 48.0 to 52.3. Vision2Web is 64.0, compared to 27B’s 62.9, with limited improvement. It’s more like focusing power on a specific workflow, rather than suddenly dropping into a single dimension.

# Qwen Official Model Card Representative Metrics Redrawn: Strengths Concentrated in Agent and Office, NL2Repo-Bench Still Trails DeepSeek

Gudhardt’s Law is a precise reminder: once a metric becomes a goal, people optimize it instead of focusing on what it originally measured. Here, you can’t just focus on the highest score; you need to look at whether your advantage is stable across tasks. Back to actual choices, **JobBench’s significant lead is worth noting, and NL2Repo’s lag shouldn’t be hidden in the footnotes either. **The most honest part of the rankings is often those scores that didn’t rise together.


💰 ▍ Who should download and who should wait? The answer is more realistic than the parameter list

The most time-saving option for 64GB Mac users now is to continue using Qwen 3.8-27B. When more power is needed, first use QwenCloud or QwenWork, and wait locally for official support and real speed data. Downloading the 72.55GB minimum quantization now only consumes hard drive, bandwidth, and troubleshooting time, and may not provide tools that work long-term.

96GB is the recommended starting point for current guidelines, but it’s only suitable for minimal quantization, and you should actively limit context to reserve space for KV cache and the system. 128GB is more relaxed, allowing you to consider higher precision or greater context, but you still can’t guarantee speed in advance. Don’t judge 12GB of graphics card solely by VRAM: if your Linux host has more than 80GB of system memory, you can still assess some offload.

Who benefits from this architecture? Cloud services can handle the large capacity that local storage can’t fit, while hardware vendors face stronger memory demands. Who bears the cost? Users who rush to launch have to pay for downloading, debugging, page changes, and waiting. The silent side is precisely existing subscription services: **They have no motivation to proactively remind you which tasks are already enough left on 27B. **

  • Use Now: 64GB Mac continues running Qwen3.8-27B

  •  ├ Requires big capability: Switch to QwenCloud or QwenWork

  •  ├ 96GB or 128GB: Wait for official support for llama.cpp and Unsloth

  •  └ Criteria for judgment: First, look at actual speed, memory usage, and long context stability

In the short term, ordinary GGUF quantization will struggle to reverse the 64GB conclusion, since IQ1_S is already the smallest tier. What could truly change the boundaries are official QAT or three-valued versions, low-bit solutions dedicated to handling N-gram tables, calibrated expert pruning, same-architecture small cups, and more mature asynchronous prefetch and expert scheduling. Each one requires real results; you can’t write your wishes as compatibility first.

After local operation shifted from geek toys to daily options, service providers’ pricing power did loosen; But this time, it was more like a re-drawing problem. Small models stayed on personal computers, large capacity was called on demand, and privacy-sensitive tasks were kept away as much as possible. **Open source returned the choice but didn’t promise every machine to swallow all ambitions for free. **Truly cost-effective deployment isn’t forced startup, but willingness to use every day.

$ 🔗 上手地址(直接点开就能用)
· 项目仓库:https://github.com/QwenLM/Qwen3.8-Flash-Next
$ 🎬 这个话题我做过视频
· 看美剧学英语,多半是「白看」——我写了个免费 Mac 和Win 软件治它|ScreenLex 光影词库
https://www.bilibili.com/video/BV1HgjU6UEe8?from=article_related_video
· 本地AI vs 云端API:一台Mac能干掉每月200块的订阅费吗?
https://www.bilibili.com/video/BV1bYGH6VEjH?from=article_related_video

Note Don’t download the 64GB version yet; for those above 96GB, the toolchain should be mature

Qwen3.8-Flash-Next opens up Qwen4’s GDN+QSA, gated residuals, N-gram Embedding, and Muon in advance. Each token activates 6B, but still stores about 180B parameters: BF16 weights are 360.00GB, and the current minimum GGUF is 72.55GB. 64GB machines are not suitable as daily deployment targets in the short term; 96GB is the recommended starting point for launch guides, and 128GB is more reasonable. What truly might redefine the boundaries are official QAT, reliable expert pruning, same-architecture small cups, and mature runtimes, rather than waiting for larger regular Q2/Q4 modes.

Note HyphenTech · Local AI / Free Purchase Guide / Tools I Make / New Product Express This article analyzes the official page, configuration files, and warehouse file lists; No download of the original 360GB weight, nor is the model running locally.


🧰 Tools I build

I maintain all of these tools myself. Preview builds are clearly labeled; the release pages are the source of truth for downloads, updates and known limits.

Info: HyphenBox Status: Official releases

A radar for free LLM APIs: availability is re-tested continuously, one local interface for all of them, and keys stay on your machine

Downloads & updates

Info: LocalBrain Status: Official releases

A multimodal MCP toolbox for local models: TTS, Whisper and video generation in one place

Downloads & updates

Info: ScreenLex Status: Official releases

Learn new words while you watch shows. Free, for Mac and Windows

Downloads & updates

Info: HyphenScreen Status: Official releases

Screen recording and smart editing in one: a DaVinci-style timeline, automatic redaction and a check of the finished video before export. Free

Downloads & updates


Quote: HyphenTech Make AI your superpower Local deployment · Free resources · Self-made software https://hyphentech.top

Share:

Comments

Loading comments…

Back to home