What can your machine actually run?
Capacity and bandwidth
| Machine | Type | Memory | Usable | Bandwidth | Training | Largest comfortable model |
|---|---|---|---|---|---|---|
| GeForce RTX 5090 | NVIDIA | 32 GB | 31 GB | 1792 GB/s | Yes | Qwen2.5 Coder 32B Instruct · Q4_K_M |
| GeForce RTX 4090 | NVIDIA | 24 GB | 23 GB | 1008 GB/s | Yes | Qwen2.5 14B Instruct · Q6_K |
| GeForce RTX 5080 | NVIDIA | 16 GB | 15 GB | 960 GB/s | Yes | Phi-4 14B · Q5_K_M |
| GeForce RTX 4080 Super | NVIDIA | 16 GB | 15 GB | 736 GB/s | Yes | Phi-4 14B · Q5_K_M |
| GeForce RTX 3090 | NVIDIA | 24 GB | 23 GB | 936 GB/s | Yes | Qwen2.5 14B Instruct · Q6_K |
| GeForce RTX 4070 | NVIDIA | 12 GB | 11 GB | 504 GB/s | Yes | Qwen2.5 Coder 7B Instruct · Q6_K |
| GeForce RTX 3060 12 GB | NVIDIA | 12 GB | 11 GB | 360 GB/s | Yes | Qwen2.5 Coder 7B Instruct · Q6_K |
| GeForce RTX 4060 Ti 16 GB | NVIDIA | 16 GB | 15 GB | 288 GB/s | Yes | Phi-4 14B · Q5_K_M |
| RTX 6000 Ada Generation | NVIDIA | 48 GB | 47 GB | 960 GB/s | Yes | Qwen2.5 Coder 32B Instruct · Q6_K |
| A100 80 GB | NVIDIA | 80 GB | 79 GB | 2039 GB/s | Yes | Qwen2.5 72B Instruct · Q5_K_M |
| MacBook Pro M4 Max (128 GB) | Apple | 128 GB | 96 GB | 546 GB/s | Yes | Qwen2.5 72B Instruct · Q6_K |
| MacBook Pro M4 Pro (48 GB) | Apple | 48 GB | 36 GB | 273 GB/s | Yes | Qwen2.5 Coder 32B Instruct · Q4_K_M |
| MacBook Air M4 (24 GB) | Apple | 24 GB | 18 GB | 120 GB/s | No | Phi-4 14B · Q5_K_M |
| Mac Studio M2 Ultra (192 GB) | Apple | 192 GB | 144 GB | 800 GB/s | Yes | Qwen2.5 72B Instruct · Q8_0 |
| MacBook Pro M1 Pro (16 GB) | Apple | 16 GB | 12 GB | 200 GB/s | No | Qwen2.5 Coder 7B Instruct · Q6_K |
| Radeon RX 7900 XTX | AMD | 24 GB | 23 GB | 960 GB/s | Yes | Qwen2.5 14B Instruct · Q6_K |
| Radeon RX 7800 XT | AMD | 16 GB | 15 GB | 624 GB/s | No | Phi-4 14B · Q5_K_M |
| CPU only — 32 GB DDR5 | Generic | 32 GB | 22 GB | 83 GB/s | No | Phi-3.5 Mini Instruct · Q4_K_M |
| CPU only — 64 GB workstation | Generic | 64 GB | 45 GB | 120 GB/s | No | Phi-3.5 Mini Instruct · Q4_K_M |
Usable memory reserves what the OS and the runtime allocator take. Apple Silicon budgets 75% of unified memory, matching the default wired-memory limit.
GeForce RTX 5090
32 GB VRAM · 1792 GB/s
32 GB of GDDR7 changes what fits: a 32B model at Q5_K_M with real context, or a 14B QLoRA run with room to spare.
31 GB usable
GeForce RTX 4090
24 GB VRAM · 1008 GB/s
The reference local-AI card. 24 GB holds a 32B at Q4_K_M for inference or a comfortable 14B QLoRA fine-tune.
23 GB usable
GeForce RTX 5080
16 GB VRAM · 960 GB/s
Very fast per gigabyte, but 16 GB puts a hard ceiling around 14B at Q4_K_M.
15 GB usable
GeForce RTX 4080 Super
16 GB VRAM · 736 GB/s
Comfortable with 7–14B models. QLoRA on a 14B fits; a 32B does not.
15 GB usable
GeForce RTX 3090
24 GB VRAM · 936 GB/s
The value pick: the same 24 GB as a 4090 at a fraction of the price, roughly 60–70% of the throughput.
23 GB usable
GeForce RTX 4070
12 GB VRAM · 504 GB/s
A solid 7–9B machine. 14B fits at Q4_K_M with a short context and no room for much else.
11 GB usable
GeForce RTX 3060 12 GB
12 GB VRAM · 360 GB/s
The cheapest sensible entry into local AI. 12 GB of VRAM matters more than raw speed.
11 GB usable
GeForce RTX 4060 Ti 16 GB
16 GB VRAM · 288 GB/s
16 GB on a narrow 128-bit bus: it will hold a 14B, but low bandwidth means slow generation. Capacity without speed.
15 GB usable
RTX 6000 Ada Generation
48 GB VRAM · 960 GB/s
48 GB in a single slot. Runs a 70B at Q4_K_M and fine-tunes 32B models locally.
47 GB usable
A100 80 GB
80 GB VRAM · 2039 GB/s
Datacentre class. Full fine-tunes of mid-size models and 70B inference at high precision.
79 GB usable
MacBook Pro M4 Max (128 GB)
128 GB unified · 546 GB/s
The largest model capacity you can carry in a bag. 128 GB unified memory runs a 70B comfortably; bandwidth caps throughput below a 4090.
96 GB usable
MacBook Pro M4 Pro (48 GB)
48 GB unified · 273 GB/s
Holds a 32B at Q4_K_M with room for context. Bandwidth, not capacity, is the limit here.
36 GB usable
MacBook Air M4 (24 GB)
24 GB unified · 120 GB/s
A fanless laptop that runs 7–14B models. Fine for chat and retrieval, not for training.
18 GB usable
Mac Studio M2 Ultra (192 GB)
192 GB unified · 800 GB/s
The most model-per-watt you can buy. 192 GB runs models no single consumer GPU can hold, at usable speed.
144 GB usable
MacBook Pro M1 Pro (16 GB)
16 GB unified · 200 GB/s
Runs 7B models well. 16 GB shared with the OS leaves about 12 GB for a model.
12 GB usable
CPU only — 32 GB DDR5
32 GB RAM · 83 GB/s
No GPU. Small models are usable for batch work; anything interactive will feel slow. Retrieval still works well.
22 GB usable
CPU only — 64 GB workstation
64 GB RAM · 120 GB/s
Enough memory to load large models; not enough bandwidth to generate with them at a usable rate.
45 GB usable