12–16 GiB GPU

Your first local model

Start small, learn the runtime and use hardware you already have if possible.

What to plan for

GPU
A supported 12–16 GiB GPU; check the exact VRAM variant.
System memory
32 GiB is a practical planning target for a single-user desktop.
Storage
An SSD with room for the model file, downloads and future variants.
Power & enclosure
Follow the exact board and CPU requirements; verify clearance and airflow.

Your first run

  1. Install a runtime compatible with your GPU, such as llama.cpp or an application built on it.
  2. Choose the exact instruct model and supported quantized artifact. Start with 4K–8K context.
  3. Check the runtime logs for offloading and actual memory use. Run several representative prompts.

The tradeoff

Smaller models make experimentation easier, but quality depends on the task. Increase model size only after evaluating your own examples.

These are component selection criteria, not a verified bill of materials. Check compatibility, current prices, power and cooling before purchasing. The examples assume Q4 weight precision, 8K context, one request and FP16 KV cache.