GeForce RTX 5060

GeForce RTX 5060, desktop/reference specification, 8 GB GDDR7

Physical memory8GiB
Available in calculator 7.0GiB
Peak bandwidth 448GB/s
Use in calculator

Desktop reference card. Partner cooling and power limits vary. Use a CUDA/runtime build with Blackwell support; no NVLink.

Specifications and setup

Memory
8 GiB GDDR7
Connection
PCIe 5.0; no NVLink
Runtime
CUDA backend; use a runtime build that supports this GPU generation.
Power
145 W. Maximum/reference board power; excludes the host system and is not measured inference power
Bandwidth basis
Manufacturer peak memory bandwidth; not a measured sustained rate or a token/s estimate.
System reserve
Planner assumption: 1 GiB per device for driver/display and allocation headroom. Runtime buffers are modeled separately; actual free memory must be checked.

Model fit at Q4 and 8K

One request, FP16 conversation cache, default reserves. “Not verified” means the model’s cache formula needs more research.

ModelMemory neededAvailableEstimateOpen plan
Qwen3.5 0.8B0.9B · Dense1.6 GiB7.0 GiBRoom to spare
Qwen3.5 2B2.3B · Dense2.4 GiB7.0 GiBRoom to spare
SmolLM3 3B3.1B · Dense3.3 GiB7.0 GiBRoom to spare
Granite 4.2 3B3.7B · Dense3.7 GiB7.0 GiBRoom to spare
Ministral 3 3B3.8B · Dense4.0 GiB7.0 GiBRoom to spare
Qwen3.5 4B4.7B · Dense3.9 GiB7.0 GiBRoom to spare
Gemma 4 E2B5.1B · Dense3.9 GiB7.0 GiBConditional · Room to spare
OLMo 3 7B Instruct7.3B · Dense7.6 GiB7.0 GiBConditional · Over capacity
Gemma 4 E4B8.0B · Dense5.6 GiB7.0 GiBConditional · Tight fit
Granite 4.2 8B8.8B · Dense7.2 GiB7.0 GiBOver capacity
Nemotron Nano 9B v28.9B · Dense6.2 GiB7.0 GiBTight fit
Ministral 3 8B8.9B · Dense7.0 GiB7.0 GiBOver capacity
Qwen3.5 9B9.7B · Dense6.7 GiB7.0 GiBTight fit
GLM-4.6V Flash 9B10.3B · Dense7.1 GiB7.0 GiBOver capacity
Gemma 4 12B12.0B · Dense8.1 GiB7.0 GiBConditional · Over capacity
Ministral 3 14B13.9B · Dense10.0 GiB7.0 GiBOver capacity
Phi-4 Reasoning Plus 14B14.7B · Dense10.8 GiB7.0 GiBOver capacity
gpt-oss 20B20.9B · Mixture of experts12.9 GiB7.0 GiBConditional · Over capacity
ERNIE 4.5 21B-A3B Thinking21.8B · Mixture of experts13.6 GiB7.0 GiBOver capacity
Devstral Small 2 24B24.0B · Dense15.7 GiB7.0 GiBOver capacity
Gemma 4 26B-A4B25.8B · Mixture of experts15.9 GiB7.0 GiBConditional · Over capacity
Qwen3.8 27B27.8B · Dense17.4 GiB7.0 GiBOver capacity
Granite 4.2 30B29.3B · Dense19.7 GiB7.0 GiBOver capacity
Qwen3 Coder 30B-A3B30.5B · Mixture of experts19.2 GiB7.0 GiBOver capacity
GLM-4.7 Flash 30B-A3B31.2B · Mixture of experts19.3 GiB7.0 GiBConditional · Over capacity
Gemma 4 31B31.3B · Dense20.3 GiB7.0 GiBConditional · Over capacity
Nemotron 3 Nano 30B-A3B31.6B · Mixture of experts19.2 GiB7.0 GiBOver capacity
OLMo 3.1 32B Think32.2B · Dense20.7 GiB7.0 GiBConditional · Over capacity
Qwen3.6 35B-A3B36.0B · Mixture of experts21.9 GiB7.0 GiBOver capacity
Kimi Linear 48B-A3B49.1B · Mixture of experts30.8 GiB7.0 GiBConditional · Over capacity
Llama 3.3 70B70.6B · Dense45.1 GiB7.0 GiBOver capacity
Qwen3 Coder Next 80B-A3B79.7B · Mixture of experts48.3 GiB7.0 GiBOver capacity
Llama 4 Scout 109B-A17B108.6B · Mixture of experts67.1 GiB7.0 GiBConditional · Over capacity
gpt-oss 120B116.8B · Mixture of experts70.8 GiB7.0 GiBConditional · Over capacity
Mistral Small 4 119B-A6.5B119.4B · Mixture of experts72.2 GiB7.0 GiBConditional · Over capacity
Nemotron 3 Super 120B-A12B123.6B · Mixture of experts74.8 GiB7.0 GiBOver capacity
Devstral 2 123B125.0B · Dense78.2 GiB7.0 GiBOver capacity
Qwen3.5 122B-A10B125.1B · Mixture of experts75.8 GiB7.0 GiBOver capacity
Mistral Medium 3.5 128B127.7B · Dense79.8 GiB7.0 GiBOver capacity
Qwen3.8 Flash Next180.0B · Mixture of experts181.0 GiB7.0 GiBConditional · Over capacity
MiniMax M2.7228.7B · Mixture of experts140.0 GiB7.0 GiBOver capacity
DeepSeek V4 Flash 0731304.2B · Mixture of experts183.6 GiB7.0 GiBConditional · Over capacity
DeepSeek V4 Flash Vision304.6B · Mixture of experts183.9 GiB7.0 GiBConditional · Over capacity
GLM-5.3 Flash321.3B · Mixture of experts194.1 GiB7.0 GiBConditional · Over capacity
Qwen3.5 397B-A17B403.4B · Mixture of experts243.9 GiB7.0 GiBOver capacity
MiniMax M3427.0B · Mixture of experts258.8 GiB7.0 GiBConditional · Over capacity
Nemotron 3 Ultra 550B-A55B560.5B · Mixture of experts338.8 GiB7.0 GiBOver capacity
Mistral Large 3 675B-A41B675.0B · Mixture of experts407.9 GiB7.0 GiBConditional · Over capacity
GLM-5.3753.3B · Mixture of experts455.3 GiB7.0 GiBConditional · Over capacity
Kimi K2.7 Code1,026.9B · Mixture of experts620.3 GiB7.0 GiBConditional · Over capacity
DeepSeek V4 Pro 08131,650.5B · Mixture of experts996.2 GiB7.0 GiBConditional · Over capacity
Qwen3.8 2.4T-A95B2,446.2B · Mixture of experts1,477.5 GiB7.0 GiBOver capacity
Kimi K32,779.9B · Mixture of experts1,678.3 GiB7.0 GiBConditional · Over capacity