Nomic Embed Text v1.5
Nomic AI
The default local embedding model for Knowledge mode. 8k context and Matryoshka truncation, so you can trade index size for accuracy.
Excellent fit
- Weights
- 260 MB
- KV cache
- 280 MB
- Overhead
- 450 MB
- 990 MB of 23 GB usable VRAM.
- KV cache at 8,192 tokens is 280 MB — shorten the context to reclaim memory.
- Chosen as the best quality that still fits a 8,192-token context (990 MB at that length).
- 8192-token context
- Variable output dimensions
- Runs fast on CPU
- English-centric
- Not a generative model
nomic-embed-textRunning Nomic Embed Text v1.5 on GeForce RTX 4090
| Quantization | Quality | Weights | KV cache | Total | ~tok/s | Fit |
|---|---|---|---|---|---|---|
| F16Pick | lossless | 260 MB | 280 MB | 990 MB | 2844 | Excellent fit |
| Q8_0 | near-lossless | 140 MB | 280 MB | 870 MB | 5354 | Excellent fit |
KV cache is sized at 8,192 tokens. Longer contexts cost proportionally more — the recommendation above reserves room for a working context.
Local fine-tuning on this machine
This model does not publish weights that can be fine-tuned locally.
Where this model runs
VRAM 32 GB · F16 · 990 MB
VRAM 24 GB · F16 · 990 MB
VRAM 16 GB · F16 · 990 MB
VRAM 16 GB · F16 · 990 MB
VRAM 24 GB · F16 · 990 MB
VRAM 12 GB · F16 · 990 MB
VRAM 12 GB · F16 · 990 MB
VRAM 16 GB · F16 · 990 MB
VRAM 48 GB · F16 · 990 MB
VRAM 80 GB · F16 · 990 MB
Unified 128 GB · F16 · 990 MB
Unified 48 GB · F16 · 990 MB
Unified 24 GB · F16 · 990 MB
Unified 192 GB · F16 · 990 MB
Unified 16 GB · F16 · 990 MB
VRAM 24 GB · F16 · 990 MB
VRAM 16 GB · F16 · 990 MB
VRAM 0 MB · Q8_0 · 870 MB
VRAM 0 MB · Q8_0 · 870 MB
Llama 3.2 3B Instruct
3.2B · Llama 3.2 Community License
Small enough for a laptop CPU or a 6 GB card, and still coherent. A sensible target for edge deployment.
Llama 3.2 1B Instruct
1.24B · Llama 3.2 Community License
A genuinely tiny model for classification, routing and extraction rather than open-ended chat.
Gemma 3 4B Instruct
4.3B · Gemma Terms of Use
Vision plus 128k context on a laptop. The most capable genuinely small multimodal option.
Phi-3.5 Mini Instruct
3.8B · MIT
Strong reasoning at 3.8B with a 128k context — but full multi-head attention makes its KV cache expensive.
Catalogue figures come from each model’s published card. ModelLM has not independently measured them.