2 × GeForce RTX 5090

2 identical GeForce RTX 5090 devices, 32 GiB nominal memory each; runtime-managed model splitting

Physical memory64GiB
Available in calculator 62.0GiB
Peak bandwidth / device 1792GB/s
Use in calculator

Combined capacity is a planning upper bound. The runtime must split weights and KV cache; each device must fit its own layers and buffers. Layer splitting does not multiply single-request memory bandwidth. Host PCIe lanes, P2P support, cooling and power must be validated.

Specifications and setup

Memory
32 GiB GDDR7 × 2 devices
Connection
Host PCIe 5.0; no NVLink
Runtime
CUDA backend; use a runtime build that supports this GPU generation.
Power
1150 W. Sum of device maximum/reference board power; host, cooling and power-supply losses excluded
Bandwidth basis
This is peak bandwidth of each device, not an aggregate rate. Do not multiply it by device count to predict sequential layer-split token latency.
System reserve
Planner assumption: 1 GiB reserve on each of 2 devices; placement and duplicated buffers can require additional memory.

Model fit at Q4 and 8K

One request, FP16 conversation cache, default reserves. “Not verified” means the model’s cache formula needs more research.

ModelMemory neededAvailableEstimateOpen plan
Qwen3.5 0.8B0.9B · Dense2.6 GiB62.0 GiBConditional · Room to spare
Qwen3.5 2B2.3B · Dense3.4 GiB62.0 GiBConditional · Room to spare
SmolLM3 3B3.1B · Dense4.3 GiB62.0 GiBConditional · Room to spare
Granite 4.2 3B3.7B · Dense4.7 GiB62.0 GiBConditional · Room to spare
Ministral 3 3B3.8B · Dense5.0 GiB62.0 GiBConditional · Room to spare
Qwen3.5 4B4.7B · Dense4.9 GiB62.0 GiBConditional · Room to spare
Gemma 4 E2B5.1B · Dense4.9 GiB62.0 GiBConditional · Room to spare
OLMo 3 7B Instruct7.3B · Dense8.6 GiB62.0 GiBConditional · Room to spare
Gemma 4 E4B8.0B · Dense6.6 GiB62.0 GiBConditional · Room to spare
Granite 4.2 8B8.8B · Dense8.2 GiB62.0 GiBConditional · Room to spare
Nemotron Nano 9B v28.9B · Dense7.2 GiB62.0 GiBConditional · Room to spare
Ministral 3 8B8.9B · Dense8.0 GiB62.0 GiBConditional · Room to spare
Qwen3.5 9B9.7B · Dense7.7 GiB62.0 GiBConditional · Room to spare
GLM-4.6V Flash 9B10.3B · Dense8.1 GiB62.0 GiBConditional · Room to spare
Gemma 4 12B12.0B · Dense9.1 GiB62.0 GiBConditional · Room to spare
Ministral 3 14B13.9B · Dense11.0 GiB62.0 GiBConditional · Room to spare
Phi-4 Reasoning Plus 14B14.7B · Dense11.8 GiB62.0 GiBConditional · Room to spare
gpt-oss 20B20.9B · Mixture of experts13.9 GiB62.0 GiBConditional · Room to spare
ERNIE 4.5 21B-A3B Thinking21.8B · Mixture of experts14.6 GiB62.0 GiBConditional · Room to spare
Devstral Small 2 24B24.0B · Dense16.7 GiB62.0 GiBConditional · Room to spare
Gemma 4 26B-A4B25.8B · Mixture of experts16.8 GiB62.0 GiBConditional · Room to spare
Qwen3.8 27B27.8B · Dense18.2 GiB62.0 GiBConditional · Room to spare
Granite 4.2 30B29.3B · Dense20.4 GiB62.0 GiBConditional · Room to spare
Qwen3 Coder 30B-A3B30.5B · Mixture of experts19.8 GiB62.0 GiBConditional · Room to spare
GLM-4.7 Flash 30B-A3B31.2B · Mixture of experts19.9 GiB62.0 GiBConditional · Room to spare
Gemma 4 31B31.3B · Dense20.9 GiB62.0 GiBConditional · Room to spare
Nemotron 3 Nano 30B-A3B31.6B · Mixture of experts19.7 GiB62.0 GiBConditional · Room to spare
OLMo 3.1 32B Think32.2B · Dense21.3 GiB62.0 GiBConditional · Room to spare
Qwen3.6 35B-A3B36.0B · Mixture of experts22.3 GiB62.0 GiBConditional · Room to spare
Kimi Linear 48B-A3B49.1B · Mixture of experts30.8 GiB62.0 GiBConditional · Room to spare
Llama 3.3 70B70.6B · Dense45.1 GiB62.0 GiBConditional · Room to spare
Qwen3 Coder Next 80B-A3B79.7B · Mixture of experts48.3 GiB62.0 GiBConditional · Room to spare
Llama 4 Scout 109B-A17B108.6B · Mixture of experts67.1 GiB62.0 GiBConditional · Over capacity
gpt-oss 120B116.8B · Mixture of experts70.8 GiB62.0 GiBConditional · Over capacity
Mistral Small 4 119B-A6.5B119.4B · Mixture of experts72.2 GiB62.0 GiBConditional · Over capacity
Nemotron 3 Super 120B-A12B123.6B · Mixture of experts74.8 GiB62.0 GiBConditional · Over capacity
Devstral 2 123B125.0B · Dense78.2 GiB62.0 GiBConditional · Over capacity
Qwen3.5 122B-A10B125.1B · Mixture of experts75.8 GiB62.0 GiBConditional · Over capacity
Mistral Medium 3.5 128B127.7B · Dense79.8 GiB62.0 GiBConditional · Over capacity
Qwen3.8 Flash Next180.0B · Mixture of experts181.0 GiB62.0 GiBConditional · Over capacity
MiniMax M2.7228.7B · Mixture of experts140.0 GiB62.0 GiBConditional · Over capacity
DeepSeek V4 Flash 0731304.2B · Mixture of experts183.6 GiB62.0 GiBConditional · Over capacity
DeepSeek V4 Flash Vision304.6B · Mixture of experts183.9 GiB62.0 GiBConditional · Over capacity
GLM-5.3 Flash321.3B · Mixture of experts194.1 GiB62.0 GiBConditional · Over capacity
Qwen3.5 397B-A17B403.4B · Mixture of experts243.9 GiB62.0 GiBConditional · Over capacity
MiniMax M3427.0B · Mixture of experts258.8 GiB62.0 GiBConditional · Over capacity
Nemotron 3 Ultra 550B-A55B560.5B · Mixture of experts338.8 GiB62.0 GiBConditional · Over capacity
Mistral Large 3 675B-A41B675.0B · Mixture of experts407.9 GiB62.0 GiBConditional · Over capacity
GLM-5.3753.3B · Mixture of experts455.3 GiB62.0 GiBConditional · Over capacity
Kimi K2.7 Code1,026.9B · Mixture of experts620.3 GiB62.0 GiBConditional · Over capacity
DeepSeek V4 Pro 08131,650.5B · Mixture of experts996.2 GiB62.0 GiBConditional · Over capacity
Qwen3.8 2.4T-A95B2,446.2B · Mixture of experts1,477.5 GiB62.0 GiBConditional · Over capacity
Kimi K32,779.9B · Mixture of experts1,678.3 GiB62.0 GiBConditional · Over capacity