DeepSeek-R1-Distill-Qwen-14B
DeepSeek
Qwen2.5 14B distilled on R1 reasoning traces. Writes out its thinking before answering, which costs tokens but wins on hard problems.
Excellent fit
- Weights
- 11 GB
- KV cache
- 1.5 GB
- Overhead
- 720 MB
- 14 GB of 23 GB usable VRAM.
- Chosen as the best quality that still fits a 32,768-token context (18 GB at that length).
- Outstanding maths and logic for 14B
- MIT licensed
- Runs on a 12 GB card at Q4_K_M
- Long chains of thought inflate latency
- Poor at short conversational replies
deepseek-r1:14bRunning DeepSeek-R1-Distill-Qwen-14B on GeForce RTX 4090
| Quantization | Quality | Weights | KV cache | Total | ~tok/s | Fit |
|---|---|---|---|---|---|---|
| F16 | lossless | 28 GB | 1.5 GB | 30 GB | 26 | Runs with CPU offload |
| Q8_0 | near-lossless | 15 GB | 1.5 GB | 17 GB | 50 | Good fit |
| Q6_KPick | near-lossless | 11 GB | 1.5 GB | 14 GB | 64 | Excellent fit |
| Q5_K_M | high | 9.8 GB | 1.5 GB | 12 GB | 74 | Excellent fit |
| Q4_K_M | balanced | 8.3 GB | 1.5 GB | 11 GB | 87 | Excellent fit |
| Q3_K_M | degraded | 6.7 GB | 1.5 GB | 9.0 GB | 108 | Excellent fit |
KV cache is sized at 8,192 tokens. Longer contexts cost proportionally more — the recommendation above reserves room for a working context.
You can fine-tune this here
- Base weights
- 7.8 GB
- Optimizer
- 1.0 GB
- Activations
- 1.3 GB
- Peak
- 12 GB
- 12 GB peak against 23 GB usable — room to raise batch size or sequence length.
- Base weights
- 28 GB
- Optimizer
- 1.0 GB
- Activations
- 1.3 GB
- Peak
- 32 GB
- Needs 32 GB — switch to QLoRA to cut the weight footprint.
Where this model runs
VRAM 32 GB · Q8_0 · 17 GB
VRAM 24 GB · Q6_K · 14 GB
VRAM 16 GB · Q5_K_M · 12 GB
VRAM 16 GB · Q5_K_M · 12 GB
VRAM 24 GB · Q6_K · 14 GB
VRAM 12 GB · Q4_K_M · 11 GB
VRAM 12 GB · Q4_K_M · 11 GB
VRAM 16 GB · Q5_K_M · 12 GB
VRAM 48 GB · F16 · 30 GB
VRAM 80 GB · F16 · 30 GB
Unified 128 GB · F16 · 30 GB
Unified 48 GB · Q8_0 · 17 GB
Unified 24 GB · Q5_K_M · 12 GB
Unified 192 GB · F16 · 30 GB
Unified 16 GB · Q4_K_M · 11 GB
VRAM 24 GB · Q6_K · 14 GB
VRAM 16 GB · Q5_K_M · 12 GB
VRAM 0 MB
VRAM 0 MB
Qwen2.5 7B Instruct
7.6B · Apache 2.0
The default starting point for local work on 8–12 GB cards. Strong instruction following and reliable tool-call formatting for its size.
Qwen2.5 14B Instruct
14.8B · Apache 2.0
The sweet spot for 24 GB cards. Meaningfully stronger reasoning than 7B while still fine-tunable locally with QLoRA.
Qwen2.5 32B Instruct
32.8B · Apache 2.0
Approaches 70B quality at half the memory. Runs on a single 24 GB card at Q4_K_M with a modest context window.
Qwen2.5 72B Instruct
72.7B · Qwen License
Frontier-adjacent open weights. Needs a workstation, a multi-GPU rig or a large unified-memory Mac.
Catalogue figures come from each model’s published card. ModelLM has not independently measured them.