Mixtral 8x7B Instruct
Mistral AI
The original open mixture-of-experts. 47B resident, ~13B active — fast generation if you have the memory to hold it.
Good fit
- Weights
- 31 GB
- KV cache
- 1.0 GB
- Overhead
- 1.3 GB
- 33 GB of 47 GB usable VRAM.
- Chosen as the best quality that still fits a 32,768-token context (36 GB at that length).
- Generates at ~13B speed
- Apache 2.0
- Well-supported MoE
- Needs ~28 GB even at Q4_K_M
- Superseded on quality by newer dense models
mixtral:8x7bRunning Mixtral 8x7B Instruct on RTX 6000 Ada Generation
| Quantization | Quality | Weights | KV cache | Total | ~tok/s | Fit |
|---|---|---|---|---|---|---|
| F16 | lossless | 87 GB | 1.0 GB | 89 GB | 29 | Runs with CPU offload |
| Q8_0 | near-lossless | 46 GB | 1.0 GB | 49 GB | 54 | Runs with CPU offload |
| Q6_K | near-lossless | 36 GB | 1.0 GB | 38 GB | 70 | Tight fit |
| Q5_K_MPick | high | 31 GB | 1.0 GB | 33 GB | 81 | Good fit |
| Q4_K_M | balanced | 26 GB | 1.0 GB | 29 GB | 95 | Good fit |
| Q3_K_M | degraded | 21 GB | 1.0 GB | 24 GB | 118 | Excellent fit |
KV cache is sized at 8,192 tokens. Longer contexts cost proportionally more — the recommendation above reserves room for a working context.
You can fine-tune this here
- Base weights
- 24 GB
- Optimizer
- 313 MB
- Activations
- 900 MB
- Peak
- 28 GB
- 28 GB peak against 47 GB usable — room to raise batch size or sequence length.
- Base weights
- 87 GB
- Optimizer
- 313 MB
- Activations
- 900 MB
- Peak
- 90 GB
- Needs 90 GB — switch to QLoRA to cut the weight footprint.
Where this model runs
VRAM 32 GB · Q4_K_M · 29 GB
VRAM 24 GB · Q4_K_M · 29 GB
VRAM 16 GB · Q4_K_M · 29 GB
VRAM 16 GB · Q4_K_M · 29 GB
VRAM 24 GB · Q4_K_M · 29 GB
VRAM 12 GB · Q4_K_M · 29 GB
VRAM 12 GB
VRAM 16 GB · Q4_K_M · 29 GB
VRAM 48 GB · Q5_K_M · 33 GB
VRAM 80 GB · Q8_0 · 49 GB
Unified 128 GB · Q8_0 · 49 GB
Unified 48 GB · Q4_K_M · 29 GB
Unified 24 GB · Q4_K_M · 29 GB
Unified 192 GB · F16 · 89 GB
Unified 16 GB · Q3_K_M · 24 GB
VRAM 24 GB · Q4_K_M · 29 GB
VRAM 16 GB · Q4_K_M · 29 GB
VRAM 0 MB
VRAM 0 MB
70.6%
Mixtral paper
Reported by the model's author. ModelLM has not run these benchmarks and does not treat them as verified.