Gemma 3 27B Instruct
The largest Gemma 3. Multimodal, long-context, and designed to run on a single high-memory accelerator.
Excellent fit
- Weights
- 15 GB
- KV cache
- 5.1 GB
- Overhead
- 940 MB
- 21 GB of 36 GB usable unified memory.
- Chosen as the best quality that still fits a 16,384-token context (27 GB at that length).
- Strong multimodal quality
- Excellent multilingual performance
- 128k context
- 16 KV heads make its cache larger than peers
- Needs 24 GB+ at Q4_K_M
gemma3:27bRunning Gemma 3 27B Instruct on MacBook Pro M4 Pro (48 GB)
| Quantization | Quality | Weights | KV cache | Total | ~tok/s | Fit |
|---|---|---|---|---|---|---|
| F16 | lossless | 51 GB | 5.1 GB | 57 GB | 4 | Runs with CPU offload |
| Q8_0 | near-lossless | 27 GB | 5.1 GB | 33 GB | 7 | Tight fit |
| Q6_K | near-lossless | 21 GB | 5.1 GB | 27 GB | 9 | Good fit |
| Q5_K_M | high | 18 GB | 5.1 GB | 24 GB | 11 | Good fit |
| Q4_K_MPick | balanced | 15 GB | 5.1 GB | 21 GB | 13 | Excellent fit |
| Q3_K_M | degraded | 12 GB | 5.1 GB | 19 GB | 16 | Excellent fit |
KV cache is sized at 8,192 tokens. Longer contexts cost proportionally more — the recommendation above reserves room for a working context.
You can fine-tune this here
- Base weights
- 14 GB
- Optimizer
- 874 MB
- Activations
- 1.7 GB
- Peak
- 19 GB
- 19 GB peak against 36 GB usable — room to raise batch size or sequence length.
- Apple Silicon trains through MLX rather than CUDA kernels.
- Base weights
- 51 GB
- Optimizer
- 874 MB
- Activations
- 1.7 GB
- Peak
- 56 GB
- Needs 56 GB — switch to QLoRA to cut the weight footprint.
- Apple Silicon trains through MLX rather than CUDA kernels.
Where this model runs
VRAM 32 GB · Q5_K_M · 24 GB
VRAM 24 GB · Q4_K_M · 21 GB
VRAM 16 GB · Q4_K_M · 21 GB
VRAM 16 GB · Q4_K_M · 21 GB
VRAM 24 GB · Q4_K_M · 21 GB
VRAM 12 GB · Q4_K_M · 21 GB
VRAM 12 GB · Q4_K_M · 21 GB
VRAM 16 GB · Q4_K_M · 21 GB
VRAM 48 GB · Q4_K_M · 21 GB
VRAM 80 GB · Q8_0 · 33 GB
Unified 128 GB · F16 · 57 GB
Unified 48 GB · Q4_K_M · 21 GB
Unified 24 GB · Q3_K_M · 19 GB
Unified 192 GB · F16 · 57 GB
Unified 16 GB · Q4_K_M · 21 GB
VRAM 24 GB · Q4_K_M · 21 GB
VRAM 16 GB · Q4_K_M · 21 GB
VRAM 0 MB
VRAM 0 MB
Qwen2.5 32B Instruct
32.8B · Apache 2.0
Approaches 70B quality at half the memory. Runs on a single 24 GB card at Q4_K_M with a modest context window.
Qwen2.5 Coder 32B Instruct
32.8B · Apache 2.0
The strongest open code model that still fits one 24 GB card. The reason a lot of people buy a 4090.
Qwen3 30B-A3B
30.5B · Apache 2.0
A mixture-of-experts model: 30B of weights in memory but only ~3B active per token, so it generates at roughly small-model speed.
Mistral Small 24B Instruct
23.6B · Apache 2.0
Built deliberately for low latency on a single card — fewer layers, wider FFN. Apache 2.0 at a size that usually is not.
Catalogue figures come from each model’s published card. ModelLM has not independently measured them.