Best local AI models for GeForce RTX 4090
The reference local-AI card. 24 GB holds a 32B at Q4_K_M for inference or a comfortable 14B QLoRA fine-tune.
Specification
VRAM24 GB
Usable for models23 GB
System RAM64 GB
Bandwidth1008 GB/s
Local trainingSupported
Operating systemswindows, linux
Runs well
29 models fit this machine
Estimated
| Model | Params | Quantization | Memory | ~tok/s | Fit | Fine-tune |
|---|---|---|---|---|---|---|
| Qwen2.5 14B Instruct | 14.8B | Q6_K | 14 GB | 64 | Excellent fit | Yes |
| Mistral Small 24B Instruct | 23.6B | Q4_K_M | 16 GB | 55 | Good fit | Yes |
| Phi-4 14B | 14.7B | Q8_0 | 17 GB | 50 | Good fit | Yes |
| Mistral Nemo 12B Instruct | 12.2B | Q6_K | 12 GB | 78 | Excellent fit | Yes |
| Qwen3 14B | 14.8B | Q6_K | 13 GB | 64 | Excellent fit | Yes |
| DeepSeek-R1-Distill-Qwen-14B | 14.8B | Q6_K | 14 GB | 64 | Excellent fit | Yes |
| Qwen2.5 Coder 32B Instruct | 32.8B | Q4_K_M | 21 GB | 39 | Tight fit | Yes |
| Gemma 2 9B Instruct | 9.2B | Q8_0 | 12 GB | 80 | Excellent fit | Yes |
| DeepSeek-Coder-V2-Lite Instruct | 15.7B | Q5_K_M | 13 GB | 458 | Excellent fit | Yes |
| StarCoder2 15B | 16B | Q8_0 | 17 GB | 46 | Good fit | Yes |
| Qwen2.5 32B Instruct | 32.8B | Q4_K_M | 21 GB | 39 | Tight fit | Yes |
| Qwen2.5 Coder 7B Instruct | 7.6B | F16 | 15 GB | 51 | Good fit | Yes |
| Llama 3.1 8B Instruct | 8B | Q8_0 | 9.5 GB | 92 | Excellent fit | Yes |
| Gemma 3 12B Instruct | 12.2B | Q8_0 | 16 GB | 60 | Good fit | Yes |
| Qwen3 8B | 8.2B | Q8_0 | 9.8 GB | 89 | Excellent fit | Yes |
| Qwen2.5 7B Instruct | 7.6B | F16 | 15 GB | 51 | Good fit | Yes |
| DeepSeek-R1-Distill-Qwen-32B | 32.8B | Q4_K_M | 21 GB | 39 | Tight fit | Yes |
| Code Llama 7B Instruct | 6.7B | Q8_0 | 11 GB | 109 | Excellent fit | Yes |
Fine-tunable here
Models you can train on this machine
Qwen2.5 14B Instruct
QLORA · 12 GB peak
Mistral Small 24B Instruct
QLORA · 17 GB peak
Phi-4 14B
QLORA · 12 GB peak
Mistral Nemo 12B Instruct
QLORA · 10 GB peak
Qwen3 14B
QLORA · 12 GB peak
DeepSeek-R1-Distill-Qwen-14B
QLORA · 12 GB peak
Qwen2.5 Coder 32B Instruct
QLORA · 22 GB peak
Gemma 2 9B Instruct
QLORA · 8.2 GB peak