Qwen3 14B
Alibaba Qwen
The Qwen3 mid-size. Long context and a reasoning mode inside a footprint a 24 GB card handles comfortably.
Excellent fit
- Weights
- 15 GB
- KV cache
- 1.3 GB
- Overhead
- 720 MB
- 17 GB of 36 GB usable unified memory.
- Chosen as the best quality that still fits a 32,768-token context (20 GB at that length).
- 128k context at 14B
- Strong tool-call reliability
- Apache 2.0
- Long context costs a large KV cache
- Thinking mode is slow for interactive chat
qwen3:14bRunning Qwen3 14B on MacBook Pro M4 Pro (48 GB)
| Quantization | Quality | Weights | KV cache | Total | ~tok/s | Fit |
|---|---|---|---|---|---|---|
| F16 | lossless | 28 GB | 1.3 GB | 30 GB | 7 | Tight fit |
| Q8_0Pick | near-lossless | 15 GB | 1.3 GB | 17 GB | 13 | Excellent fit |
| Q6_K | near-lossless | 11 GB | 1.3 GB | 13 GB | 17 | Excellent fit |
| Q5_K_M | high | 9.8 GB | 1.3 GB | 12 GB | 20 | Excellent fit |
| Q4_K_M | balanced | 8.3 GB | 1.3 GB | 10 GB | 24 | Excellent fit |
| Q3_K_M | degraded | 6.7 GB | 1.3 GB | 8.7 GB | 29 | Excellent fit |
KV cache is sized at 8,192 tokens. Longer contexts cost proportionally more — the recommendation above reserves room for a working context.
You can fine-tune this here
- Base weights
- 7.8 GB
- Optimizer
- 957 MB
- Activations
- 1.3 GB
- Peak
- 12 GB
- 12 GB peak against 36 GB usable — room to raise batch size or sequence length.
- Apple Silicon trains through MLX rather than CUDA kernels.
- Base weights
- 28 GB
- Optimizer
- 957 MB
- Activations
- 1.3 GB
- Peak
- 32 GB
- 32 GB peak against 36 GB usable.
- Apple Silicon trains through MLX rather than CUDA kernels.
Where this model runs
VRAM 32 GB · Q8_0 · 17 GB
VRAM 24 GB · Q6_K · 13 GB
VRAM 16 GB · Q4_K_M · 10 GB
VRAM 16 GB · Q4_K_M · 10 GB
VRAM 24 GB · Q6_K · 13 GB
VRAM 12 GB · Q4_K_M · 10 GB
VRAM 12 GB · Q4_K_M · 10 GB
VRAM 16 GB · Q4_K_M · 10 GB
VRAM 48 GB · F16 · 30 GB
VRAM 80 GB · F16 · 30 GB
Unified 128 GB · F16 · 30 GB
Unified 48 GB · Q8_0 · 17 GB
Unified 24 GB · Q4_K_M · 10 GB
Unified 192 GB · F16 · 30 GB
Unified 16 GB · Q4_K_M · 10 GB
VRAM 24 GB · Q6_K · 13 GB
VRAM 16 GB · Q4_K_M · 10 GB
VRAM 0 MB
VRAM 0 MB
Qwen2.5 14B Instruct
14.8B · Apache 2.0
The sweet spot for 24 GB cards. Meaningfully stronger reasoning than 7B while still fine-tunable locally with QLoRA.
Qwen3 8B
8.2B · Apache 2.0
Switchable thinking mode: the same weights answer directly or reason step by step depending on the prompt. Long context for its size.
Mistral Nemo 12B Instruct
12.2B · Apache 2.0
A 12B with a 128k context and the Tekken tokenizer, which compresses non-English text far better than Llama’s.
Gemma 2 9B Instruct
9.2B · Gemma Terms of Use
Unusually good at natural, well-structured prose for its size. The short context is the catch.
Catalogue figures come from each model’s published card. ModelLM has not independently measured them.