Mac Studio M1 Ultra · 128 GB (48-core GPU)
Mac Studio, Apple M1 Ultra, 20-core CPU / 48-core GPU / 128 GB unified memory
Physical memory128GiB
Available in calculator 96.0GiB
Peak bandwidth 800GB/s
Use Metal or MLX with support for the chosen model. The OS and other apps share memory; Metal working-set limits vary.
Specifications and setup
- Memory
- 128 GiB Unified memory
- Connection
- On-chip unified memory; no discrete VRAM pool
- Runtime
- Metal / MLX. Actual usable memory is runtime- and OS-dependent; query recommendedMaxWorkingSetSize rather than treating all RAM as VRAM.
- Bandwidth basis
- Apple-published chip bandwidth; shared CPU/GPU traffic reduces bandwidth available to inference.
- System reserve
- Planner assumption: 25% of installed memory reserved for macOS, other apps and allocation headroom. This is adjustable, not an Apple GPU-allocation guarantee.
Model fit at Q4 and 8K
One request, FP16 conversation cache, default reserves. “Not verified” means the model’s cache formula needs more research.
| Model | Memory needed | Available | Estimate | Open plan |
|---|---|---|---|---|
| Qwen3.5 0.8B0.9B · Dense | 1.6 GiB | 96.0 GiB | Room to spare | |
| Qwen3.5 2B2.3B · Dense | 2.4 GiB | 96.0 GiB | Room to spare | |
| SmolLM3 3B3.1B · Dense | 3.3 GiB | 96.0 GiB | Room to spare | |
| Granite 4.2 3B3.7B · Dense | 3.7 GiB | 96.0 GiB | Room to spare | |
| Ministral 3 3B3.8B · Dense | 4.0 GiB | 96.0 GiB | Room to spare | |
| Qwen3.5 4B4.7B · Dense | 3.9 GiB | 96.0 GiB | Room to spare | |
| Gemma 4 E2B5.1B · Dense | 3.9 GiB | 96.0 GiB | Conditional · Room to spare | |
| OLMo 3 7B Instruct7.3B · Dense | 7.6 GiB | 96.0 GiB | Conditional · Room to spare | |
| Gemma 4 E4B8.0B · Dense | 5.6 GiB | 96.0 GiB | Conditional · Room to spare | |
| Granite 4.2 8B8.8B · Dense | 7.2 GiB | 96.0 GiB | Room to spare | |
| Nemotron Nano 9B v28.9B · Dense | 6.2 GiB | 96.0 GiB | Room to spare | |
| Ministral 3 8B8.9B · Dense | 7.0 GiB | 96.0 GiB | Room to spare | |
| Qwen3.5 9B9.7B · Dense | 6.7 GiB | 96.0 GiB | Room to spare | |
| GLM-4.6V Flash 9B10.3B · Dense | 7.1 GiB | 96.0 GiB | Room to spare | |
| Gemma 4 12B12.0B · Dense | 8.1 GiB | 96.0 GiB | Conditional · Room to spare | |
| Ministral 3 14B13.9B · Dense | 10.0 GiB | 96.0 GiB | Room to spare | |
| Phi-4 Reasoning Plus 14B14.7B · Dense | 10.8 GiB | 96.0 GiB | Room to spare | |
| gpt-oss 20B20.9B · Mixture of experts | 12.9 GiB | 96.0 GiB | Conditional · Room to spare | |
| ERNIE 4.5 21B-A3B Thinking21.8B · Mixture of experts | 13.6 GiB | 96.0 GiB | Room to spare | |
| Devstral Small 2 24B24.0B · Dense | 15.7 GiB | 96.0 GiB | Room to spare | |
| Gemma 4 26B-A4B25.8B · Mixture of experts | 15.9 GiB | 96.0 GiB | Conditional · Room to spare | |
| Qwen3.8 27B27.8B · Dense | 17.4 GiB | 96.0 GiB | Room to spare | |
| Granite 4.2 30B29.3B · Dense | 19.7 GiB | 96.0 GiB | Room to spare | |
| Qwen3 Coder 30B-A3B30.5B · Mixture of experts | 19.2 GiB | 96.0 GiB | Room to spare | |
| GLM-4.7 Flash 30B-A3B31.2B · Mixture of experts | 19.3 GiB | 96.0 GiB | Conditional · Room to spare | |
| Gemma 4 31B31.3B · Dense | 20.3 GiB | 96.0 GiB | Conditional · Room to spare | |
| Nemotron 3 Nano 30B-A3B31.6B · Mixture of experts | 19.2 GiB | 96.0 GiB | Room to spare | |
| OLMo 3.1 32B Think32.2B · Dense | 20.7 GiB | 96.0 GiB | Conditional · Room to spare | |
| Qwen3.6 35B-A3B36.0B · Mixture of experts | 21.9 GiB | 96.0 GiB | Room to spare | |
| Kimi Linear 48B-A3B49.1B · Mixture of experts | 30.8 GiB | 96.0 GiB | Conditional · Room to spare | |
| Llama 3.3 70B70.6B · Dense | 45.1 GiB | 96.0 GiB | Room to spare | |
| Qwen3 Coder Next 80B-A3B79.7B · Mixture of experts | 48.3 GiB | 96.0 GiB | Room to spare | |
| Llama 4 Scout 109B-A17B108.6B · Mixture of experts | 67.1 GiB | 96.0 GiB | Conditional · Room to spare | |
| gpt-oss 120B116.8B · Mixture of experts | 70.8 GiB | 96.0 GiB | Conditional · Room to spare | |
| Mistral Small 4 119B-A6.5B119.4B · Mixture of experts | 72.2 GiB | 96.0 GiB | Conditional · Room to spare | |
| Nemotron 3 Super 120B-A12B123.6B · Mixture of experts | 74.8 GiB | 96.0 GiB | Room to spare | |
| Devstral 2 123B125.0B · Dense | 78.2 GiB | 96.0 GiB | Room to spare | |
| Qwen3.5 122B-A10B125.1B · Mixture of experts | 75.8 GiB | 96.0 GiB | Room to spare | |
| Mistral Medium 3.5 128B127.7B · Dense | 79.8 GiB | 96.0 GiB | Room to spare | |
| Qwen3.8 Flash Next180.0B · Mixture of experts | 181.0 GiB | 96.0 GiB | Conditional · Over capacity | |
| MiniMax M2.7228.7B · Mixture of experts | 140.0 GiB | 96.0 GiB | Over capacity | |
| DeepSeek V4 Flash 0731304.2B · Mixture of experts | 183.6 GiB | 96.0 GiB | Conditional · Over capacity | |
| DeepSeek V4 Flash Vision304.6B · Mixture of experts | 183.9 GiB | 96.0 GiB | Conditional · Over capacity | |
| GLM-5.3 Flash321.3B · Mixture of experts | 194.1 GiB | 96.0 GiB | Conditional · Over capacity | |
| Qwen3.5 397B-A17B403.4B · Mixture of experts | 243.9 GiB | 96.0 GiB | Over capacity | |
| MiniMax M3427.0B · Mixture of experts | 258.8 GiB | 96.0 GiB | Conditional · Over capacity | |
| Nemotron 3 Ultra 550B-A55B560.5B · Mixture of experts | 338.8 GiB | 96.0 GiB | Over capacity | |
| Mistral Large 3 675B-A41B675.0B · Mixture of experts | 407.9 GiB | 96.0 GiB | Conditional · Over capacity | |
| GLM-5.3753.3B · Mixture of experts | 455.3 GiB | 96.0 GiB | Conditional · Over capacity | |
| Kimi K2.7 Code1,026.9B · Mixture of experts | 620.3 GiB | 96.0 GiB | Conditional · Over capacity | |
| DeepSeek V4 Pro 08131,650.5B · Mixture of experts | 996.2 GiB | 96.0 GiB | Conditional · Over capacity | |
| Qwen3.8 2.4T-A95B2,446.2B · Mixture of experts | 1,477.5 GiB | 96.0 GiB | Over capacity | |
| Kimi K32,779.9B · Mixture of experts | 1,678.3 GiB | 96.0 GiB | Conditional · Over capacity |