Local LLM memory calculator
Compare models and hardware to understand how much memory is worth planning for.
Your configuration
Narrow models to your hardware, and hardware to your model, using your current settings.
Memory needed
estimated use / available memory
- Model weights
- 5.4 GiB
- Conversation memory
- 0.3 GiB
- Runtime allowance
- 1.0 GiB
- Room left
- 16.3 GiB
1.0 GiB reserved for the system and other apps.
Calculation assumptions
Text inference estimate. Image, audio and video inputs need additional encoder and temporary memory.
Understand the tradeoffs before you buy.
Compare memory capacity, model size and precision to narrow your hardware shortlist.
- How much VRAM do you actually need?
Calculate model weights, KV cache and runtime overhead before choosing hardware for local AI.
- How much Mac memory do you need for local LLMs?
Compare 16GB, 24GB, 32GB and 64GB Mac memory for local LLMs. Budget unified memory, context and macOS headroom with worked model examples.
- Mac or GPU workstation for local AI?
Compare unified memory and discrete GPU memory, runtime compatibility and upgrade paths for local models.
- Two GPUs do not become one bigger GPU.
Understand the conditions under which multiple GPUs can run a larger local model, and what combined VRAM leaves out.
- Smaller Q8 or larger Q4: which local model should you run?
Choose between a smaller high-precision LLM and a larger quantized model with a practical test for quality, memory, context and response time.
- How context uses memory
See how context length, the KV cache and concurrent requests affect local LLM memory requirements.