Mistral Nemo 12B Instruct
Mistral AI / NVIDIA
A 12B with a 128k context and the Tekken tokenizer, which compresses non-English text far better than Llama’s.
Mistral Nemo 12B Instruct does not fit on CPU only — 64 GB workstation.
CPU-only inference. Expect a few tokens per second at best.
- 128k context
- Efficient multilingual tokenizer
- Apache 2.0
- Middling reasoning for its size
- Long context needs a large KV cache
mistral-nemo:12bRunning Mistral Nemo 12B Instruct on CPU only — 64 GB workstation
| Quantization | Quality | Weights | KV cache | Total | ~tok/s | Fit |
|---|---|---|---|---|---|---|
| F16 | lossless | 23 GB | 1.6 GB | 25 GB | 4 | Not recommended |
| Q8_0 | near-lossless | 12 GB | 1.6 GB | 14 GB | 7 | Not recommended |
| Q6_K | near-lossless | 9.3 GB | 1.6 GB | 12 GB | 9 | Not recommended |
| Q5_K_M | high | 8.1 GB | 1.6 GB | 10 GB | 11 | Not recommended |
| Q4_K_M | balanced | 6.9 GB | 1.6 GB | 9.1 GB | 13 | Not recommended |
| Q3_K_M | degraded | 5.5 GB | 1.6 GB | 7.8 GB | 16 | Not recommended |
KV cache is sized at 8,192 tokens. Longer contexts cost proportionally more — the recommendation above reserves room for a working context.
Local fine-tuning on this machine
- Base weights
- 6.4 GB
- Optimizer
- 874 MB
- Activations
- 1.2 GB
- Peak
- 10 GB
- Local fine-tuning needs a GPU or Apple Silicon. CPU training is not practical.
- Base weights
- 23 GB
- Optimizer
- 874 MB
- Activations
- 1.2 GB
- Peak
- 26 GB
- Local fine-tuning needs a GPU or Apple Silicon. CPU training is not practical.
Where this model runs
VRAM 32 GB · Q8_0 · 14 GB
VRAM 24 GB · Q6_K · 12 GB
VRAM 16 GB · Q5_K_M · 10 GB
VRAM 16 GB · Q5_K_M · 10 GB
VRAM 24 GB · Q6_K · 12 GB
VRAM 12 GB · Q4_K_M · 9.1 GB
VRAM 12 GB · Q4_K_M · 9.1 GB
VRAM 16 GB · Q5_K_M · 10 GB
VRAM 48 GB · F16 · 25 GB
VRAM 80 GB · F16 · 25 GB
Unified 128 GB · F16 · 25 GB
Unified 48 GB · Q8_0 · 14 GB
Unified 24 GB · Q4_K_M · 9.1 GB
Unified 192 GB · F16 · 25 GB
Unified 16 GB · Q4_K_M · 9.1 GB
VRAM 24 GB · Q6_K · 12 GB
VRAM 16 GB · Q5_K_M · 10 GB
VRAM 0 MB
VRAM 0 MB
68%
Mistral Nemo announcement
Reported by the model's author. ModelLM has not run these benchmarks and does not treat them as verified.
Qwen2.5 7B Instruct
7.6B · Apache 2.0
The default starting point for local work on 8–12 GB cards. Strong instruction following and reliable tool-call formatting for its size.
Qwen2.5 14B Instruct
14.8B · Apache 2.0
The sweet spot for 24 GB cards. Meaningfully stronger reasoning than 7B while still fine-tunable locally with QLoRA.
Qwen2.5 Coder 7B Instruct
7.6B · Apache 2.0
The practical local copilot. Supports fill-in-the-middle, so it works as an inline completion model rather than only a chat assistant.
Qwen3 8B
8.2B · Apache 2.0
Switchable thinking mode: the same weights answer directly or reason step by step depending on the prompt. Long context for its size.
Catalogue figures come from each model’s published card. ModelLM has not independently measured them.