Unified memory

A high-memory Mac workspace

Room for large model weights in a compact workstation, with a shared system memory budget.

What to plan for

Machine
Choose an exact Apple silicon configuration with enough unified memory. Memory is not upgradeable later.
Memory policy
Budget for macOS and other apps. The planner reserves 25% of the configured capacity.
Storage
Leave space for several large model files and temporary downloads.
Runtime
Use a compatible Metal or MLX runtime and its supported model format.

Your first run

  1. Confirm the model’s architecture and quantization format are supported by your chosen runtime.
  2. Start with modest context, inspect peak memory and test alongside normal applications.
  3. Measure actual latency. Large capacity does not imply a particular generation speed.

The tradeoff

Choose the exact chip, GPU core count and memory size. These can change both capacity and bandwidth.

These are component selection criteria, not a verified bill of materials. Check compatibility, current prices, power and cooling before purchasing. The examples assume Q4 weight precision, 8K context, one request and FP16 KV cache.