DeepSeek-R1-Distill-Qwen-32B
DeepSeek
The strongest open reasoning model that fits a single 24 GB card. A genuinely different capability class on hard problems.
Runs with CPU offload
- Weights
- 18 GB
- KV cache
- 2.0 GB
- Overhead
- 1.0 GB
- Exceeds VRAM by 6.5 GB; layers spill to system RAM and generation slows sharply.
- Best local reasoning under 70B
- MIT licensed
- Excellent at maths and proofs
- Very verbose
- Tight on 24 GB
- Slow for interactive use
deepseek-r1:32bRunning DeepSeek-R1-Distill-Qwen-32B on Radeon RX 7800 XT
| Quantization | Quality | Weights | KV cache | Total | ~tok/s | Fit |
|---|---|---|---|---|---|---|
| F16 | lossless | 61 GB | 2.0 GB | 64 GB | 7 | Won't run |
| Q8_0 | near-lossless | 32 GB | 2.0 GB | 36 GB | 14 | Runs with CPU offload |
| Q6_K | near-lossless | 25 GB | 2.0 GB | 28 GB | 18 | Runs with CPU offload |
| Q5_K_M | high | 22 GB | 2.0 GB | 25 GB | 21 | Runs with CPU offload |
| Q4_K_MPick | balanced | 18 GB | 2.0 GB | 21 GB | 24 | Runs with CPU offload |
| Q3_K_M | degraded | 15 GB | 2.0 GB | 18 GB | 30 | Runs with CPU offload |
KV cache is sized at 8,192 tokens. Longer contexts cost proportionally more — the recommendation above reserves room for a working context.
Local fine-tuning on this machine
- Base weights
- 17 GB
- Optimizer
- 1.0 GB
- Activations
- 1.9 GB
- Peak
- 22 GB
- QLoRA still needs 22 GB; this machine has 15 GB. Choose a smaller base model.
- AMD requires a ROCm build of PyTorch; the Unsloth fast path is CUDA-only.
- Base weights
- 61 GB
- Optimizer
- 1.0 GB
- Activations
- 1.9 GB
- Peak
- 66 GB
- Needs 66 GB — switch to QLoRA to cut the weight footprint.
- AMD requires a ROCm build of PyTorch; the Unsloth fast path is CUDA-only.
Where this model runs
VRAM 32 GB · Q4_K_M · 21 GB
VRAM 24 GB · Q4_K_M · 21 GB
VRAM 16 GB · Q4_K_M · 21 GB
VRAM 16 GB · Q4_K_M · 21 GB
VRAM 24 GB · Q4_K_M · 21 GB
VRAM 12 GB · Q4_K_M · 21 GB
VRAM 12 GB · Q4_K_M · 21 GB
VRAM 16 GB · Q4_K_M · 21 GB
VRAM 48 GB · Q6_K · 28 GB
VRAM 80 GB · Q8_0 · 36 GB
Unified 128 GB · F16 · 64 GB
Unified 48 GB · Q4_K_M · 21 GB
Unified 24 GB · Q3_K_M · 18 GB
Unified 192 GB · F16 · 64 GB
Unified 16 GB · Q4_K_M · 21 GB
VRAM 24 GB · Q4_K_M · 21 GB
VRAM 16 GB · Q4_K_M · 21 GB
VRAM 0 MB
VRAM 0 MB
Qwen2.5 7B Instruct
7.6B · Apache 2.0
The default starting point for local work on 8–12 GB cards. Strong instruction following and reliable tool-call formatting for its size.
Qwen2.5 14B Instruct
14.8B · Apache 2.0
The sweet spot for 24 GB cards. Meaningfully stronger reasoning than 7B while still fine-tunable locally with QLoRA.
Qwen2.5 32B Instruct
32.8B · Apache 2.0
Approaches 70B quality at half the memory. Runs on a single 24 GB card at Q4_K_M with a modest context window.
Qwen2.5 72B Instruct
72.7B · Qwen License
Frontier-adjacent open weights. Needs a workstation, a multi-GPU rig or a large unified-memory Mac.
Catalogue figures come from each model’s published card. ModelLM has not independently measured them.