What to plan for
- GPU
- A supported 12–16 GiB GPU; check the exact VRAM variant.
- System memory
- 32 GiB is a practical planning target for a single-user desktop.
- Storage
- An SSD with room for the model file, downloads and future variants.
- Power & enclosure
- Follow the exact board and CPU requirements; verify clearance and airflow.
Your first run
- Install a runtime compatible with your GPU, such as llama.cpp or an application built on it.
- Choose the exact instruct model and supported quantized artifact. Start with 4K–8K context.
- Check the runtime logs for offloading and actual memory use. Run several representative prompts.
The tradeoff
Smaller models make experimentation easier, but quality depends on the task. Increase model size only after evaluating your own examples.
These are component selection criteria, not a verified bill of materials. Check compatibility, current prices, power and cooling before purchasing. The examples assume Q4 weight precision, 8K context, one request and FP16 KV cache.