4 × H200 NVL 141 GB PCIe
4 identical H200 NVL 141 GB PCIe devices, 141 GiB nominal memory each; runtime-managed model splitting
Physical memory564GiB
Available in calculator 556.0GiB
Peak bandwidth / device 4800GB/s
Combined capacity is a planning upper bound. The runtime must split weights and KV cache; each device must fit its own layers and buffers. Layer splitting does not multiply single-request memory bandwidth. Host PCIe lanes, P2P support, cooling and power must be validated.
Specifications and setup
- Memory
- 141 GiB HBM3e ECC × 4 devices
- Connection
- Four-way NVLink bridge, up to 900 GB/s bidirectional per GPU; PCIe 5.0 host
- Runtime
- CUDA backend; use a runtime build that supports this GPU generation.
- Power
- 2400 W. Sum of device maximum/reference board power; host, cooling and power-supply losses excluded
- Bandwidth basis
- This is peak bandwidth of each device, not an aggregate rate. Do not multiply it by device count to predict sequential layer-split token latency.
- System reserve
- Planner assumption: 2 GiB reserve on each of 4 devices; placement and duplicated buffers can require additional memory.
Model fit at Q4 and 8K
One request, FP16 conversation cache, default reserves. “Not verified” means the model’s cache formula needs more research.
| Model | Memory needed | Available | Estimate | Open plan |
|---|---|---|---|---|
| Qwen3.5 0.8B0.9B · Dense | 4.6 GiB | 556.0 GiB | Conditional · Room to spare | |
| Qwen3.5 2B2.3B · Dense | 5.4 GiB | 556.0 GiB | Conditional · Room to spare | |
| SmolLM3 3B3.1B · Dense | 6.3 GiB | 556.0 GiB | Conditional · Room to spare | |
| Granite 4.2 3B3.7B · Dense | 6.7 GiB | 556.0 GiB | Conditional · Room to spare | |
| Ministral 3 3B3.8B · Dense | 7.0 GiB | 556.0 GiB | Conditional · Room to spare | |
| Qwen3.5 4B4.7B · Dense | 6.9 GiB | 556.0 GiB | Conditional · Room to spare | |
| Gemma 4 E2B5.1B · Dense | 6.9 GiB | 556.0 GiB | Conditional · Room to spare | |
| OLMo 3 7B Instruct7.3B · Dense | 10.6 GiB | 556.0 GiB | Conditional · Room to spare | |
| Gemma 4 E4B8.0B · Dense | 8.6 GiB | 556.0 GiB | Conditional · Room to spare | |
| Granite 4.2 8B8.8B · Dense | 10.2 GiB | 556.0 GiB | Conditional · Room to spare | |
| Nemotron Nano 9B v28.9B · Dense | 9.2 GiB | 556.0 GiB | Conditional · Room to spare | |
| Ministral 3 8B8.9B · Dense | 10.0 GiB | 556.0 GiB | Conditional · Room to spare | |
| Qwen3.5 9B9.7B · Dense | 9.7 GiB | 556.0 GiB | Conditional · Room to spare | |
| GLM-4.6V Flash 9B10.3B · Dense | 10.1 GiB | 556.0 GiB | Conditional · Room to spare | |
| Gemma 4 12B12.0B · Dense | 11.1 GiB | 556.0 GiB | Conditional · Room to spare | |
| Ministral 3 14B13.9B · Dense | 13.0 GiB | 556.0 GiB | Conditional · Room to spare | |
| Phi-4 Reasoning Plus 14B14.7B · Dense | 13.8 GiB | 556.0 GiB | Conditional · Room to spare | |
| gpt-oss 20B20.9B · Mixture of experts | 15.9 GiB | 556.0 GiB | Conditional · Room to spare | |
| ERNIE 4.5 21B-A3B Thinking21.8B · Mixture of experts | 16.6 GiB | 556.0 GiB | Conditional · Room to spare | |
| Devstral Small 2 24B24.0B · Dense | 18.7 GiB | 556.0 GiB | Conditional · Room to spare | |
| Gemma 4 26B-A4B25.8B · Mixture of experts | 18.8 GiB | 556.0 GiB | Conditional · Room to spare | |
| Qwen3.8 27B27.8B · Dense | 20.2 GiB | 556.0 GiB | Conditional · Room to spare | |
| Granite 4.2 30B29.3B · Dense | 22.4 GiB | 556.0 GiB | Conditional · Room to spare | |
| Qwen3 Coder 30B-A3B30.5B · Mixture of experts | 21.8 GiB | 556.0 GiB | Conditional · Room to spare | |
| GLM-4.7 Flash 30B-A3B31.2B · Mixture of experts | 21.9 GiB | 556.0 GiB | Conditional · Room to spare | |
| Gemma 4 31B31.3B · Dense | 22.9 GiB | 556.0 GiB | Conditional · Room to spare | |
| Nemotron 3 Nano 30B-A3B31.6B · Mixture of experts | 21.7 GiB | 556.0 GiB | Conditional · Room to spare | |
| OLMo 3.1 32B Think32.2B · Dense | 23.3 GiB | 556.0 GiB | Conditional · Room to spare | |
| Qwen3.6 35B-A3B36.0B · Mixture of experts | 24.3 GiB | 556.0 GiB | Conditional · Room to spare | |
| Kimi Linear 48B-A3B49.1B · Mixture of experts | 32.6 GiB | 556.0 GiB | Conditional · Room to spare | |
| Llama 3.3 70B70.6B · Dense | 45.9 GiB | 556.0 GiB | Conditional · Room to spare | |
| Qwen3 Coder Next 80B-A3B79.7B · Mixture of experts | 48.8 GiB | 556.0 GiB | Conditional · Room to spare | |
| Llama 4 Scout 109B-A17B108.6B · Mixture of experts | 67.1 GiB | 556.0 GiB | Conditional · Room to spare | |
| gpt-oss 120B116.8B · Mixture of experts | 70.8 GiB | 556.0 GiB | Conditional · Room to spare | |
| Mistral Small 4 119B-A6.5B119.4B · Mixture of experts | 72.2 GiB | 556.0 GiB | Conditional · Room to spare | |
| Nemotron 3 Super 120B-A12B123.6B · Mixture of experts | 74.8 GiB | 556.0 GiB | Conditional · Room to spare | |
| Devstral 2 123B125.0B · Dense | 78.2 GiB | 556.0 GiB | Conditional · Room to spare | |
| Qwen3.5 122B-A10B125.1B · Mixture of experts | 75.8 GiB | 556.0 GiB | Conditional · Room to spare | |
| Mistral Medium 3.5 128B127.7B · Dense | 79.8 GiB | 556.0 GiB | Conditional · Room to spare | |
| Qwen3.8 Flash Next180.0B · Mixture of experts | 181.0 GiB | 556.0 GiB | Conditional · Room to spare | |
| MiniMax M2.7228.7B · Mixture of experts | 140.0 GiB | 556.0 GiB | Conditional · Room to spare | |
| DeepSeek V4 Flash 0731304.2B · Mixture of experts | 183.7 GiB | 556.0 GiB | Conditional · Room to spare | |
| DeepSeek V4 Flash Vision304.6B · Mixture of experts | 183.9 GiB | 556.0 GiB | Conditional · Room to spare | |
| GLM-5.3 Flash321.3B · Mixture of experts | 194.1 GiB | 556.0 GiB | Conditional · Room to spare | |
| Qwen3.5 397B-A17B403.4B · Mixture of experts | 243.9 GiB | 556.0 GiB | Conditional · Room to spare | |
| MiniMax M3427.0B · Mixture of experts | 258.8 GiB | 556.0 GiB | Conditional · Room to spare | |
| Nemotron 3 Ultra 550B-A55B560.5B · Mixture of experts | 338.8 GiB | 556.0 GiB | Conditional · Room to spare | |
| Mistral Large 3 675B-A41B675.0B · Mixture of experts | 407.9 GiB | 556.0 GiB | Conditional · Room to spare | |
| GLM-5.3753.3B · Mixture of experts | 455.3 GiB | 556.0 GiB | Conditional · Room to spare | |
| Kimi K2.7 Code1,026.9B · Mixture of experts | 620.3 GiB | 556.0 GiB | Conditional · Over capacity | |
| DeepSeek V4 Pro 08131,650.5B · Mixture of experts | 996.2 GiB | 556.0 GiB | Conditional · Over capacity | |
| Qwen3.8 2.4T-A95B2,446.2B · Mixture of experts | 1,477.5 GiB | 556.0 GiB | Conditional · Over capacity | |
| Kimi K32,779.9B · Mixture of experts | 1,678.3 GiB | 556.0 GiB | Conditional · Over capacity |