What to plan for
- GPU
- A 24 or 32 GiB GPU supported by the runtime. Compare used-card condition and new-card alternatives.
- System memory
- 64 GiB is a planning target for loading models while keeping desktop work available.
- Storage
- 1–2 TB NVMe as a starting capacity, adjusted to your existing projects and model library.
- Power & enclosure
- Size the supply for the chosen CPU and exact board. Plan for sustained heat and connector clearance.
Your first run
- Compare Q4 and Q5 memory requirements for your target model and context.
- Verify the model format, runtime version and GPU architecture are supported together.
- Measure response latency and output speed with your actual prompts before committing to a workflow.
The tradeoff
A faster 24 GiB card has essentially the same memory fit as another 24 GiB card. Paying for more compute does not solve a capacity limit.
These are component selection criteria, not a verified bill of materials. Check compatibility, current prices, power and cooling before purchasing. The examples assume Q4 weight precision, 8K context, one request and FP16 KV cache.