Qwen3 30B-A3B
Alibaba Qwen
A mixture-of-experts model: 30B of weights in memory but only ~3B active per token, so it generates at roughly small-model speed.
Good fit
- Weights
- 57 GB
- KV cache
- 380 MB
- Overhead
- 1.0 GB
- 58 GB of 96 GB usable unified memory.
- Chosen as the best quality that still fits a 32,768-token context (59 GB at that length).
- Generates far faster than a dense 30B
- Excellent on high-memory, low-bandwidth machines
- Still needs the full 30B resident in memory
- MoE fine-tuning support is less mature
qwen3:30b-a3bRunning Qwen3 30B-A3B on MacBook Pro M4 Max (128 GB)
| Quantization | Quality | Weights | KV cache | Total | ~tok/s | Fit |
|---|---|---|---|---|---|---|
| F16Pick | lossless | 57 GB | 380 MB | 58 GB | 64 | Good fit |
| Q8_0 | near-lossless | 30 GB | 380 MB | 32 GB | 120 | Excellent fit |
| Q6_K | near-lossless | 23 GB | 380 MB | 25 GB | 156 | Excellent fit |
| Q5_K_M | high | 20 GB | 380 MB | 22 GB | 180 | Excellent fit |
| Q4_K_M | balanced | 17 GB | 380 MB | 19 GB | 212 | Excellent fit |
| Q3_K_M | degraded | 14 GB | 380 MB | 15 GB | 262 | Excellent fit |
KV cache is sized at 8,192 tokens. Longer contexts cost proportionally more — the recommendation above reserves room for a working context.
You can fine-tune this here
- Base weights
- 16 GB
- Optimizer
- 214 MB
- Activations
- 520 MB
- Peak
- 18 GB
- 18 GB peak against 96 GB usable — room to raise batch size or sequence length.
- Apple Silicon trains through MLX rather than CUDA kernels.
- Base weights
- 57 GB
- Optimizer
- 214 MB
- Activations
- 520 MB
- Peak
- 59 GB
- 59 GB peak against 96 GB usable — room to raise batch size or sequence length.
- Apple Silicon trains through MLX rather than CUDA kernels.
Where this model runs
VRAM 32 GB · Q5_K_M · 22 GB
VRAM 24 GB · Q4_K_M · 19 GB
VRAM 16 GB · Q4_K_M · 19 GB
VRAM 16 GB · Q4_K_M · 19 GB
VRAM 24 GB · Q4_K_M · 19 GB
VRAM 12 GB · Q4_K_M · 19 GB
VRAM 12 GB · Q4_K_M · 19 GB
VRAM 16 GB · Q4_K_M · 19 GB
VRAM 48 GB · Q8_0 · 32 GB
VRAM 80 GB · F16 · 58 GB
Unified 128 GB · F16 · 58 GB
Unified 48 GB · Q6_K · 25 GB
Unified 24 GB · Q3_K_M · 15 GB
Unified 192 GB · F16 · 58 GB
Unified 16 GB · Q4_K_M · 19 GB
VRAM 24 GB · Q4_K_M · 19 GB
VRAM 16 GB · Q4_K_M · 19 GB
VRAM 0 MB
VRAM 0 MB
Qwen2.5 32B Instruct
32.8B · Apache 2.0
Approaches 70B quality at half the memory. Runs on a single 24 GB card at Q4_K_M with a modest context window.
Qwen2.5 Coder 32B Instruct
32.8B · Apache 2.0
The strongest open code model that still fits one 24 GB card. The reason a lot of people buy a 4090.
Gemma 3 27B Instruct
27.4B · Gemma Terms of Use
The largest Gemma 3. Multimodal, long-context, and designed to run on a single high-memory accelerator.
DeepSeek-R1-Distill-Qwen-32B
32.8B · MIT
The strongest open reasoning model that fits a single 24 GB card. A genuinely different capability class on hard problems.
Catalogue figures come from each model’s published card. ModelLM has not independently measured them.