Open-weight model families now publish both styles side by side. Qwen ships Qwen3-32B (dense) and Qwen3-30B-A3B (MoE); Mistral ships Mixtral-8x7B (MoE) next to plain Llama-shaped dense checkpoints; IBM ships Granite 4 in MoE-only sizes. The “A3B” or “A22B” suffix is the key - it tells you how many parameters actually run per token, which is what sets VRAM, RAM and speed locally.
What the suffix means, exactly
A dense model uses every parameter on every token. Qwen3-32B’s config.json lists no num_experts field and an architectures entry of Qwen3ForCausalLM: each of its 64 layers runs the same MLP for every token (Qwen3-32B config). A Mixture-of-Experts model partitions the MLP into many small MLPs (“experts”) and routes each token to a small subset of them.
Qwen3-30B-A3B’s config carries num_experts: 128 and num_experts_per_tok: 8, with architectures: ["Qwen3MoeForCausalLM"] (Qwen3-30B-A3B config). The model card sums it up as “30.5B in total and 3.3B activated” (Qwen3-30B-A3B model card). Qwen3-235B-A22B-Instruct-2507 has the same expert count (128 routed, 8 active per token) but in a much larger container, summed as “235B in total and 22B activated” (Qwen3-235B-A22B-Instruct-2507 model card).
The classic third-party case is Mixtral-8x7B, where num_local_experts: 8 and num_experts_per_tok: 2 mean each token sees two of the eight experts (Mixtral config). DeepSeek-V3.2 is the most aggressive current example: n_routed_experts: 256, n_shared_experts: 1, num_experts_per_tok: 8, packed into a model whose safetensors.total is 685,355,329,792 bytes (~685 GB on disk at FP8). The model card only states “Model size: 685B params” (DeepSeek-V3.2 model card), but the eight-of-256-plus-one-shared pattern has been widely reported as activating roughly 37B parameters per token - similar to the DeepSeek-V3 baseline (DeepSeek-V3.2 config).
The point: total parameters predict storage and VRAM footprint; activated parameters predict compute per token and, roughly, quality. Those two numbers can differ by 5x to 10x, which is the whole reason MoE is interesting locally.
A side-by-side comparison
The table below uses only fields pulled from each model’s Hugging Face config.json and card on 2026-09-30. “Disk BF16” is the safetensors.total value when the model ships in BF16, rounded; “Disk FP8” applies where the published weight is FP8, which is the case for DeepSeek-V3.2. “Active per token” is the model’s own published activated-parameter figure (or num_experts_per_tok times the per-expert width where the card does not list a single number).
| Model | Architecture | Total | Active / token | Disk | Licence | HF gated |
|---|---|---|---|---|---|---|
| Llama-3.1-8B | Dense | 8B | 8B | 8.03 GB (BF16) | llama3.1 | manual |
| Qwen3-32B | Dense | 32B | 32B | 30.53 GB (BF16) | apache-2.0 | false |
| Mixtral-8x7B-Instruct-v0.1 | MoE (8 of 2) | ~46.7B | ~12.9B | 46.70 GB (BF16) | apache-2.0 | false |
| Qwen3-30B-A3B | MoE (128 of 8) | 30.5B | 3.3B | 30.52 GB (BF16) | apache-2.0 | false |
| granite-4.0-h-small | MoE (72 of 10) | 32B (per IBM) | 9B (per IBM) | varies by quant | apache-2.0 | false |
| Qwen3-235B-A22B-Instruct-2507 | MoE (128 of 8) | 235B | 22B | 235.09 GB (BF16) | apache-2.0 | false |
| DeepSeek-V3.2 | MoE (256 routed + 1 shared, 8 active) | 685B | ~37B (per third-party reporting) | 685.36 GB (FP8) | MIT | false |
Sources: Llama-3.1-8B API, Qwen3-32B API, Mixtral API, Qwen3-30B-A3B API, granite-4.0-h-small API, Qwen3-235B-A22B API, DeepSeek-V3.2 API. Disk figures are raw safetensors totals from those API responses.
Notice the column that does not move with total parameters: active per token. Qwen3-30B-A3B and Qwen3-32B live at almost the same on-disk weight size (30.5 GB vs 30.5 GB), but Qwen3-30B-A3B does roughly 3.3B-parameter math per token while Qwen3-32B does 32B-parameter math. On consumer hardware that often translates into the MoE model being faster and fitting in the same memory budget, because most of its experts stay on disk or in CPU memory and only the active subset is brought onto the GPU per layer (this is the “expert offload” pattern; current llama.cpp discussion threads describe it as an SSD-streaming target for MoE models that exceed system RAM: llama.cpp MoE issues).
When MoE is the right pick locally
MoE wins on three workloads in particular.
VRAM-bound setups that still want quality. A 24 GB consumer card cannot fit Qwen3-32B at usable quant and context, but it can fit Qwen3-30B-A3B at the same quant because only 8 of the 128 experts have to be resident per layer for any given token. The other 120 experts are still in the file, but the inference engine keeps them in CPU RAM and pages them in as the router’s choices land. The practical ceiling ends up being total-disk, not active-per-token.
Throughput on long prompts. MoE models’ per-token compute is set by the activated subset, not the total parameter count, so a 30B-active MoE tends to run faster than a 30B dense model at the same quant. Quality often tracks the larger total, not the smaller active number - the model card for Qwen3-235B-A22B-Instruct-2507 cites benchmarks comparable to much larger dense checkpoints while paying 22B-parameter compute per token.
Cost crossover. The “active compute” number is what determines the local-vs-API crossover. A 235B MoE whose per-token compute is 22B parameters is closer to a hosted 70B-class model on dollar-per-million-tokens than its total weight size would suggest, while the storage bill is closer to a 22B dense model on a single machine.
When dense is still the right pick locally
Three cases where MoE does not help.
Edge and tiny VRAM. Models under about 2B parameters ship dense in nearly every family because the routing overhead eats the savings. The Qwen 0.6B/1.7B/4B tier and the Llama 3.2 1B/3B tier are all dense; an MoE of the same activated size would carry a much bigger file for no quality win. See the without-a-GPU guide for the under-2B tier in detail.
Latency-sensitive first-token timing. The first token a MoE model emits requires loading every routed expert from off-device storage or CPU memory. On a fast NVMe that is a few hundred milliseconds; on a network share it can be seconds. A dense model with the same activated compute has a more predictable cold-start.
Inference engines that do not implement expert offload. llama.cpp has open issues for SSD streaming of routed expert weights and partial-offload scheduler crashes (llama.cpp MoE issues); Ollama has its own open MoE crash reports on Blackwell, ROCm mixed-arch systems and MLX Gemma 4 MoE imports (Ollama MoE issues). These are not deal-breakers for the listed models, but they are worth checking against the engine you actually run.
How to read a model card in 30 seconds
A four-field check tells you almost everything.
safetensors.totalon the HF API page - the on-disk weight size at the published quant.model_typeandarchitecturesinconfig.json- if it ends inMoeForCausalLMormixtralordeepseek_v32, it is MoE; otherwise it is dense.num_experts(orn_routed_experts) andnum_experts_per_tok- the storage vs active ratio. Qwen3-30B-A3B reads 128 and 8, so storage is the 128 number and active is roughly 8/128 of the per-layer MLP width.- The model card’s own “N total / N activated” sentence if present - usually the fastest read. Qwen’s cards always state it; DeepSeek’s card only states “Model size: 685B params” without an activated-parameter figure.
If fields 2 through 4 are missing on a card you want to run, treat it as unverified and either read config.json directly or pick a model whose card spells it out.
What this changes for the rest of the cluster
A few practical ripples for the other local-AI pages on this site. The VRAM arithmetic guide should be read with “total parameters = weight size” and “activated parameters = KV-cache pressure and per-token compute” treated as separate columns; MoE blurs the second number downward without changing the first. The inference-speed guide measures throughput in tokens per second, which is driven by the activated count plus memory bandwidth, not the total count. The family-picker hub is the right place to start when you have not decided whether MoE even applies to your workload.
Bottom line
MoE is a way to buy a larger model’s quality at a smaller model’s compute, paid for in disk and in memory bandwidth. On a single desktop with one consumer GPU and a fast NVMe, an MoE that activates 3B to 22B parameters per token is often the sweet spot for the 30B-to-235B total tier. A dense checkpoint is still simpler, still cold-starts faster, and still wins under about 8B parameters. Read the model’s config.json once before you commit a download.