Qwen2.5 14B Instruct
The sweet spot for 24 GB cards. Meaningfully stronger reasoning than 7B while still fine-tunable locally with QLoRA.
- Params
- 14.8B
- Quant
- Q6_K
- Memory
- 14 GB
Loading…
Agent work lives or dies on reliable tool-call formatting. These are the models whose function-calling output is consistent enough to build on.
The sweet spot for 24 GB cards. Meaningfully stronger reasoning than 7B while still fine-tunable locally with QLoRA.
Built deliberately for low latency on a single card — fewer layers, wider FFN. Apache 2.0 at a size that usually is not.
A 12B with a 128k context and the Tekken tokenizer, which compresses non-English text far better than Llama’s.
| # | Model | Params | Licence | Quantization | Memory | ~tok/s | Fit | Fine-tune |
|---|---|---|---|---|---|---|---|---|
| 1 | Qwen2.5 14B Instruct | 14.8B | Apache 2.0 | Q6_K | 14 GB | 64 | Excellent fit | Yes |
| 2 | Mistral Small 24B Instruct | 23.6B | Apache 2.0 | Q4_K_M | 16 GB | 55 | Good fit | Yes |
| 3 | Mistral Nemo 12B Instruct | 12.2B | Apache 2.0 | Q6_K | 12 GB | 78 | Excellent fit | Yes |
| 4 | Qwen3 14B | 14.8B | Apache 2.0 | Q6_K | 13 GB | 64 | Excellent fit | Yes |
| 5 | Qwen2.5 32B Instruct | 32.8B | Apache 2.0 | Q4_K_M | 21 GB | 39 | Tight fit | Yes |
| 6 | Llama 3.1 8B Instruct | 8B | Llama 3.1 Community License | Q8_0 | 9.5 GB | 92 | Excellent fit | Yes |
| 7 | Qwen3 8B | 8.2B | Apache 2.0 | Q8_0 | 9.8 GB | 89 | Excellent fit | Yes |
| 8 | Qwen2.5 7B Instruct | 7.6B | Apache 2.0 | F16 | 15 GB | 51 | Good fit | Yes |
| 9 | Qwen3 30B-A3B | 30.5B | Apache 2.0 | Q4_K_M | 19 GB | 391 | Tight fit | Yes |
| 10 | Mistral 7B Instruct v0.3 | 7.25B | Apache 2.0 | F16 | 15 GB | 54 | Good fit | Yes |
| 11 | Qwen2.5 72B Instruct | 72.7B | Qwen License | Q4_K_M | 45 GB | 18 | Runs with CPU offload | No |
| 12 | Llama 3.3 70B Instruct | 70.6B | Llama 3.3 Community License | Q4_K_M | 44 GB | 18 | Runs with CPU offload | No |
The catalogue is filtered to models that declare this task, then ranked by capability — parameter count, published benchmarks, licence clarity — and derated by how comfortably each one fits the reference machine.
ModelLM can read your actual GPU, VRAM and RAM and size every model against it.
Detect my hardware