Qwen2.5 14B Instruct
The sweet spot for 24 GB cards. Meaningfully stronger reasoning than 7B while still fine-tunable locally with QLoRA.
- Params
- 14.8B
- Quant
- Q6_K
- Memory
- 14 GB
Loading…
Reasoning models write out intermediate steps before answering. They cost more tokens and more latency, and win on problems where the answer depends on the working.
The sweet spot for 24 GB cards. Meaningfully stronger reasoning than 7B while still fine-tunable locally with QLoRA.
Built deliberately for low latency on a single card — fewer layers, wider FFN. Apache 2.0 at a size that usually is not.
Trained largely on curated synthetic data. Punches far above its size on reasoning and maths; the short context limits what you can do with it.
| # | Model | Params | Licence | Quantization | Memory | ~tok/s | Fit | Fine-tune |
|---|---|---|---|---|---|---|---|---|
| 1 | Qwen2.5 14B Instruct | 14.8B | Apache 2.0 | Q6_K | 14 GB | 64 | Excellent fit | Yes |
| 2 | Mistral Small 24B Instruct | 23.6B | Apache 2.0 | Q4_K_M | 16 GB | 55 | Good fit | Yes |
| 3 | Phi-4 14B | 14.7B | MIT | Q8_0 | 17 GB | 50 | Good fit | Yes |
| 4 | Qwen3 14B | 14.8B | Apache 2.0 | Q6_K | 13 GB | 64 | Excellent fit | Yes |
| 5 | DeepSeek-R1-Distill-Qwen-14B | 14.8B | MIT | Q6_K | 14 GB | 64 | Excellent fit | Yes |
| 6 | Qwen2.5 Coder 32B Instruct | 32.8B | Apache 2.0 | Q4_K_M | 21 GB | 39 | Tight fit | Yes |
| 7 | Qwen2.5 32B Instruct | 32.8B | Apache 2.0 | Q4_K_M | 21 GB | 39 | Tight fit | Yes |
| 8 | Qwen3 8B | 8.2B | Apache 2.0 | Q8_0 | 9.8 GB | 89 | Excellent fit | Yes |
| 9 | DeepSeek-R1-Distill-Qwen-32B | 32.8B | MIT | Q4_K_M | 21 GB | 39 | Tight fit | Yes |
| 10 | Gemma 3 27B Instruct | 27.4B | Gemma Terms of Use | Q4_K_M | 21 GB | 47 | Tight fit | Yes |
| 11 | Qwen3 30B-A3B | 30.5B | Apache 2.0 | Q4_K_M | 19 GB | 391 | Tight fit | Yes |
| 12 | Phi-3.5 Mini Instruct | 3.8B | MIT | Q8_0 | 7.3 GB | 193 | Excellent fit | Yes |
| 13 | Qwen2.5 72B Instruct | 72.7B | Qwen License | Q4_K_M | 45 GB | 18 | Runs with CPU offload | No |
| 14 | Llama 3.3 70B Instruct | 70.6B | Llama 3.3 Community License | Q4_K_M | 44 GB | 18 | Runs with CPU offload | No |
The catalogue is filtered to models that declare this task, then ranked by capability — parameter count, published benchmarks, licence clarity — and derated by how comfortably each one fits the reference machine.
ModelLM can read your actual GPU, VRAM and RAM and size every model against it.
Detect my hardware