Local LLM models
Find a model by size and task, then check its memory requirements.
Choosing between model sizes? Compare a smaller Q8 model with a larger Q4 model, or learn what quantization means.
Qwen3.5 0.8B
Small text and image model for basic tasks on limited memory.
256K context
Qwen3.5 2B
Compact text and image model for local chat and simple tools.
256K context
SmolLM3 3B
Small text model with optional reasoning and a fully published training recipe.
64K context · 128K extended
Granite 4.2 3B
Small current Granite with optional reasoning and tool use.
128K context · 512K extended
Ministral 3 3B
Small text and image model with instruction following and tool use.
256K context
Qwen3.5 4B
Small text and image model with reasoning and tool use.
256K context · 1,010,000 extended
Gemma 4 E2B
Compact text, image and audio model; its full weights exceed 5B parameters.
128K context
OLMo 3 7B Instruct
Fully open small chat model with published training data and recipes.
64K context
Gemma 4 E4B
Text, image and audio model with roughly 8B total stored parameters.
128K context
Granite 4.2 8B
Current dense Granite for reasoning, code and multilingual tools.
128K context · 512K extended
Nemotron Nano 9B v2
Small dense hybrid model for reasoning and non-reasoning tasks.
128K context
Ministral 3 8B
Text and image model with a 256K supported context.
256K context
Qwen3.5 9B
Text, image and coding model with optional reasoning.
256K context · 1,010,000 extended
GLM-4.6V Flash 9B
Small vision-language model with native tool use.
128K context
Gemma 4 12B
Unified text, image and audio model with a 256K context.
256K context
Ministral 3 14B
Largest small dense Ministral for text and image tasks.
256K context
Phi-4 Reasoning Plus 14B
Dense reasoning specialist for mathematics, science and coding.
32K context
gpt-oss 20B
Small OpenAI reasoning model supplied with native MXFP4 expert weights.
128K context
ERNIE 4.5 21B-A3B Thinking
Compact sparse ERNIE with reasoning and a 128K context.
128K context
Devstral Small 2 24B
Coding agent model with image input for software development.
256K context
Gemma 4 26B-A4B
Sparse text and image model with about 3.8B active parameters.
256K context
Qwen3.8 27B
Current dense Qwen for coding, reasoning and image understanding.
256K context · 1,000,000 extended
Granite 4.2 30B
Largest current dense Granite for reasoning and agent tasks.
128K context · 512K extended
Qwen3 Coder 30B-A3B
Dedicated coding model with conventional attention and a 256K native context.
256K context · 1024K extended