StarCoder2 15B
BigCode
A base completion model trained on permissively-licensed source with full provenance. Not a chat model.
Excellent fit
- Weights
- 30 GB
- KV cache
- 630 MB
- Overhead
- 740 MB
- 31 GB of 79 GB usable VRAM.
- Chosen as the best quality that still fits a 16,384-token context (32 GB at that length).
- Fully traceable training data
- Excellent fill-in-the-middle
- 600+ languages
- Base model — needs instruction tuning for chat
- OpenRAIL-M carries use restrictions
starcoder2:15bRunning StarCoder2 15B on A100 80 GB
| Quantization | Quality | Weights | KV cache | Total | ~tok/s | Fit |
|---|---|---|---|---|---|---|
| F16Pick | lossless | 30 GB | 630 MB | 31 GB | 49 | Excellent fit |
| Q8_0 | near-lossless | 16 GB | 630 MB | 17 GB | 93 | Excellent fit |
| Q6_K | near-lossless | 12 GB | 630 MB | 14 GB | 120 | Excellent fit |
| Q5_K_M | high | 11 GB | 630 MB | 12 GB | 139 | Excellent fit |
| Q4_K_M | balanced | 9.0 GB | 630 MB | 10 GB | 163 | Excellent fit |
| Q3_K_M | degraded | 7.3 GB | 630 MB | 8.7 GB | 202 | Excellent fit |
KV cache is sized at 8,192 tokens. Longer contexts cost proportionally more — the recommendation above reserves room for a working context.
You can fine-tune this here
- Base weights
- 8.4 GB
- Optimizer
- 1.2 GB
- Activations
- 1.6 GB
- Peak
- 13 GB
- 13 GB peak against 79 GB usable — room to raise batch size or sequence length.
- Base weights
- 30 GB
- Optimizer
- 1.2 GB
- Activations
- 1.6 GB
- Peak
- 35 GB
- 35 GB peak against 79 GB usable — room to raise batch size or sequence length.
Where this model runs
VRAM 32 GB · Q8_0 · 17 GB
VRAM 24 GB · Q8_0 · 17 GB
VRAM 16 GB · Q4_K_M · 10 GB
VRAM 16 GB · Q4_K_M · 10 GB
VRAM 24 GB · Q8_0 · 17 GB
VRAM 12 GB · Q4_K_M · 10 GB
VRAM 12 GB · Q4_K_M · 10 GB
VRAM 16 GB · Q4_K_M · 10 GB
VRAM 48 GB · F16 · 31 GB
VRAM 80 GB · F16 · 31 GB
Unified 128 GB · F16 · 31 GB
Unified 48 GB · Q8_0 · 17 GB
Unified 24 GB · Q6_K · 14 GB
Unified 192 GB · F16 · 31 GB
Unified 16 GB · Q4_K_M · 10 GB
VRAM 24 GB · Q8_0 · 17 GB
VRAM 16 GB · Q4_K_M · 10 GB
VRAM 0 MB
VRAM 0 MB
Qwen2.5 14B Instruct
14.8B · Apache 2.0
The sweet spot for 24 GB cards. Meaningfully stronger reasoning than 7B while still fine-tunable locally with QLoRA.
Qwen3 14B
14.8B · Apache 2.0
The Qwen3 mid-size. Long context and a reasoning mode inside a footprint a 24 GB card handles comfortably.
Mistral Nemo 12B Instruct
12.2B · Apache 2.0
A 12B with a 128k context and the Tekken tokenizer, which compresses non-English text far better than Llama’s.
Gemma 3 12B Instruct
12.2B · Gemma Terms of Use
Multimodal and long-context in a 12B footprint. Interleaved local/global attention keeps the KV cache affordable.
Catalogue figures come from each model’s published card. ModelLM has not independently measured them.