Gemma 3 12B Instruct
Multimodal and long-context in a 12B footprint. Interleaved local/global attention keeps the KV cache affordable.
Excellent fit
- Weights
- 23 GB
- KV cache
- 2.8 GB
- Overhead
- 670 MB
- 26 GB of 96 GB usable unified memory.
- Chosen as the best quality that still fits a 32,768-token context (35 GB at that length).
- Reads images as well as text
- 128k context
- Very broad language coverage
- Vision path needs runtime support
- Behind Qwen3 on pure reasoning
gemma3:12bRunning Gemma 3 12B Instruct on MacBook Pro M4 Max (128 GB)
| Quantization | Quality | Weights | KV cache | Total | ~tok/s | Fit |
|---|---|---|---|---|---|---|
| F16Pick | lossless | 23 GB | 2.8 GB | 26 GB | 17 | Excellent fit |
| Q8_0 | near-lossless | 12 GB | 2.8 GB | 16 GB | 33 | Excellent fit |
| Q6_K | near-lossless | 9.3 GB | 2.8 GB | 13 GB | 42 | Excellent fit |
| Q5_K_M | high | 8.1 GB | 2.8 GB | 12 GB | 49 | Excellent fit |
| Q4_K_M | balanced | 6.9 GB | 2.8 GB | 10 GB | 57 | Excellent fit |
| Q3_K_M | degraded | 5.5 GB | 2.8 GB | 9.0 GB | 71 | Excellent fit |
KV cache is sized at 8,192 tokens. Longer contexts cost proportionally more — the recommendation above reserves room for a working context.
You can fine-tune this here
- Base weights
- 6.4 GB
- Optimizer
- 967 MB
- Activations
- 1.1 GB
- Peak
- 10 GB
- 10 GB peak against 96 GB usable — room to raise batch size or sequence length.
- Apple Silicon trains through MLX rather than CUDA kernels.
- Base weights
- 23 GB
- Optimizer
- 967 MB
- Activations
- 1.1 GB
- Peak
- 27 GB
- 27 GB peak against 96 GB usable — room to raise batch size or sequence length.
- Apple Silicon trains through MLX rather than CUDA kernels.
Where this model runs
VRAM 32 GB · Q8_0 · 16 GB
VRAM 24 GB · Q8_0 · 16 GB
VRAM 16 GB · Q5_K_M · 12 GB
VRAM 16 GB · Q5_K_M · 12 GB
VRAM 24 GB · Q8_0 · 16 GB
VRAM 12 GB · Q4_K_M · 10 GB
VRAM 12 GB · Q4_K_M · 10 GB
VRAM 16 GB · Q5_K_M · 12 GB
VRAM 48 GB · F16 · 26 GB
VRAM 80 GB · F16 · 26 GB
Unified 128 GB · F16 · 26 GB
Unified 48 GB · Q8_0 · 16 GB
Unified 24 GB · Q5_K_M · 12 GB
Unified 192 GB · F16 · 26 GB
Unified 16 GB · Q4_K_M · 10 GB
VRAM 24 GB · Q8_0 · 16 GB
VRAM 16 GB · Q5_K_M · 12 GB
VRAM 0 MB
VRAM 0 MB
Qwen2.5 7B Instruct
7.6B · Apache 2.0
The default starting point for local work on 8–12 GB cards. Strong instruction following and reliable tool-call formatting for its size.
Qwen2.5 14B Instruct
14.8B · Apache 2.0
The sweet spot for 24 GB cards. Meaningfully stronger reasoning than 7B while still fine-tunable locally with QLoRA.
Qwen2.5 Coder 7B Instruct
7.6B · Apache 2.0
The practical local copilot. Supports fill-in-the-middle, so it works as an inline completion model rather than only a chat assistant.
Qwen3 8B
8.2B · Apache 2.0
Switchable thinking mode: the same weights answer directly or reason step by step depending on the prompt. Long context for its size.
Catalogue figures come from each model’s published card. ModelLM has not independently measured them.