Qwen2.5 72B Instruct
Alibaba Qwen
Frontier-adjacent open weights. Needs a workstation, a multi-GPU rig or a large unified-memory Mac.
Good fit
- Weights
- 56 GB
- KV cache
- 2.5 GB
- Overhead
- 1.6 GB
- 60 GB of 96 GB usable unified memory.
- Chosen as the best quality that still fits a 32,768-token context (67 GB at that length).
- Strongest general quality in the Qwen 2.5 line
- Excellent multilingual coverage
- Out of reach for most single consumer GPUs
- Custom licence, not Apache
qwen2.5:72bRunning Qwen2.5 72B Instruct on MacBook Pro M4 Max (128 GB)
| Quantization | Quality | Weights | KV cache | Total | ~tok/s | Fit |
|---|---|---|---|---|---|---|
| F16 | lossless | 135 GB | 2.5 GB | 139 GB | 3 | Runs with CPU offload |
| Q8_0 | near-lossless | 72 GB | 2.5 GB | 76 GB | 5 | Good fit |
| Q6_KPick | near-lossless | 56 GB | 2.5 GB | 60 GB | 7 | Good fit |
| Q5_K_M | high | 48 GB | 2.5 GB | 52 GB | 8 | Excellent fit |
| Q4_K_M | balanced | 41 GB | 2.5 GB | 45 GB | 10 | Excellent fit |
| Q3_K_M | degraded | 33 GB | 2.5 GB | 37 GB | 12 | Excellent fit |
KV cache is sized at 8,192 tokens. Longer contexts cost proportionally more — the recommendation above reserves room for a working context.
You can fine-tune this here
- Base weights
- 38 GB
- Optimizer
- 1.6 GB
- Activations
- 2.9 GB
- Peak
- 45 GB
- 45 GB peak against 96 GB usable — room to raise batch size or sequence length.
- Apple Silicon trains through MLX rather than CUDA kernels.
- Base weights
- 135 GB
- Optimizer
- 1.6 GB
- Activations
- 2.9 GB
- Peak
- 143 GB
- Needs 143 GB — switch to QLoRA to cut the weight footprint.
- Apple Silicon trains through MLX rather than CUDA kernels.
Where this model runs
VRAM 32 GB · Q4_K_M · 45 GB
VRAM 24 GB · Q4_K_M · 45 GB
VRAM 16 GB · Q3_K_M · 37 GB
VRAM 16 GB · Q3_K_M · 37 GB
VRAM 24 GB · Q4_K_M · 45 GB
VRAM 12 GB
VRAM 12 GB
VRAM 16 GB · Q3_K_M · 37 GB
VRAM 48 GB · Q4_K_M · 45 GB
VRAM 80 GB · Q5_K_M · 52 GB
Unified 128 GB · Q6_K · 60 GB
Unified 48 GB · Q4_K_M · 45 GB
Unified 24 GB
Unified 192 GB · Q8_0 · 76 GB
Unified 16 GB
VRAM 24 GB · Q4_K_M · 45 GB
VRAM 16 GB · Q3_K_M · 37 GB
VRAM 0 MB
VRAM 0 MB
86.1%
Qwen2.5 model card
Reported by the model's author. ModelLM has not run these benchmarks and does not treat them as verified.
Qwen2.5 7B Instruct
7.6B · Apache 2.0
The default starting point for local work on 8–12 GB cards. Strong instruction following and reliable tool-call formatting for its size.
Qwen2.5 14B Instruct
14.8B · Apache 2.0
The sweet spot for 24 GB cards. Meaningfully stronger reasoning than 7B while still fine-tunable locally with QLoRA.
Qwen2.5 32B Instruct
32.8B · Apache 2.0
Approaches 70B quality at half the memory. Runs on a single 24 GB card at Q4_K_M with a modest context window.
Qwen2.5 Coder 7B Instruct
7.6B · Apache 2.0
The practical local copilot. Supports fill-in-the-middle, so it works as an inline completion model rather than only a chat assistant.
Catalogue figures come from each model’s published card. ModelLM has not independently measured them.