Mixtral 8x7B Instruct
Mistral AI
The original open mixture-of-experts. 47B resident, ~13B active — fast generation if you have the memory to hold it.
Excellent fit
- Weights
- 46 GB
- KV cache
- 1.0 GB
- Overhead
- 1.3 GB
- 49 GB of 96 GB usable unified memory.
- Chosen as the best quality that still fits a 32,768-token context (52 GB at that length).
- Generates at ~13B speed
- Apache 2.0
- Well-supported MoE
- Needs ~28 GB even at Q4_K_M
- Superseded on quality by newer dense models
mixtral:8x7bRunning Mixtral 8x7B Instruct on MacBook Pro M4 Max (128 GB)
| Quantization | Quality | Weights | KV cache | Total | ~tok/s | Fit |
|---|---|---|---|---|---|---|
| F16 | lossless | 87 GB | 1.0 GB | 89 GB | 16 | Tight fit |
| Q8_0Pick | near-lossless | 46 GB | 1.0 GB | 49 GB | 31 | Excellent fit |
| Q6_K | near-lossless | 36 GB | 1.0 GB | 38 GB | 40 | Excellent fit |
| Q5_K_M | high | 31 GB | 1.0 GB | 33 GB | 46 | Excellent fit |
| Q4_K_M | balanced | 26 GB | 1.0 GB | 29 GB | 54 | Excellent fit |
| Q3_K_M | degraded | 21 GB | 1.0 GB | 24 GB | 67 | Excellent fit |
KV cache is sized at 8,192 tokens. Longer contexts cost proportionally more — the recommendation above reserves room for a working context.
You can fine-tune this here
- Base weights
- 24 GB
- Optimizer
- 313 MB
- Activations
- 900 MB
- Peak
- 28 GB
- 28 GB peak against 96 GB usable — room to raise batch size or sequence length.
- Apple Silicon trains through MLX rather than CUDA kernels.
- Base weights
- 87 GB
- Optimizer
- 313 MB
- Activations
- 900 MB
- Peak
- 90 GB
- Fits, but an OOM is likely if the sequence length or batch size rises.
- Apple Silicon trains through MLX rather than CUDA kernels.
Where this model runs
VRAM 32 GB · Q4_K_M · 29 GB
VRAM 24 GB · Q4_K_M · 29 GB
VRAM 16 GB · Q4_K_M · 29 GB
VRAM 16 GB · Q4_K_M · 29 GB
VRAM 24 GB · Q4_K_M · 29 GB
VRAM 12 GB · Q4_K_M · 29 GB
VRAM 12 GB
VRAM 16 GB · Q4_K_M · 29 GB
VRAM 48 GB · Q5_K_M · 33 GB
VRAM 80 GB · Q8_0 · 49 GB
Unified 128 GB · Q8_0 · 49 GB
Unified 48 GB · Q4_K_M · 29 GB
Unified 24 GB · Q4_K_M · 29 GB
Unified 192 GB · F16 · 89 GB
Unified 16 GB · Q3_K_M · 24 GB
VRAM 24 GB · Q4_K_M · 29 GB
VRAM 16 GB · Q4_K_M · 29 GB
VRAM 0 MB
VRAM 0 MB
70.6%
Mixtral paper
Reported by the model's author. ModelLM has not run these benchmarks and does not treat them as verified.