Local LLM models

Find a model by size and task, then check its memory requirements.

Choosing between model sizes? Compare a smaller Q8 model with a larger Q4 model, or learn what quantization means.

53 models · Checked 2026-09-06
QwenGeneral

Qwen3.5 0.8B

0.9B stored parameters

Small text and image model for basic tasks on limited memory.

256K context

Q4 · 8K ≈ 1.6 GiB
QwenGeneral

Qwen3.5 2B

2.3B stored parameters

Compact text and image model for local chat and simple tools.

256K context

Q4 · 8K ≈ 2.4 GiB
Hugging FaceGeneral

SmolLM3 3B

3.1B stored parameters

Small text model with optional reasoning and a fully published training recipe.

64K context · 128K extended

Q4 · 8K ≈ 3.3 GiB
IBMReasoning

Granite 4.2 3B

3.7B stored parameters

Small current Granite with optional reasoning and tool use.

128K context · 512K extended

Q4 · 8K ≈ 3.7 GiB
Mistral AIGeneral

Ministral 3 3B

3.8B stored parameters

Small text and image model with instruction following and tool use.

256K context

Q4 · 8K ≈ 4.0 GiB
QwenGeneral

Qwen3.5 4B

4.7B stored parameters

Small text and image model with reasoning and tool use.

256K context · 1,010,000 extended

Q4 · 8K ≈ 3.9 GiB
GoogleGeneral

Gemma 4 E2B

5.1B stored parameters

Compact text, image and audio model; its full weights exceed 5B parameters.

128K context

Q4 · 8K ≈ 3.9 GiB
Ai2General

OLMo 3 7B Instruct

7.3B stored parameters

Fully open small chat model with published training data and recipes.

64K context

Q4 · 8K ≈ 7.6 GiB
GoogleGeneral

Gemma 4 E4B

8.0B stored parameters

Text, image and audio model with roughly 8B total stored parameters.

128K context

Q4 · 8K ≈ 5.6 GiB
IBMReasoning

Granite 4.2 8B

8.8B stored parameters

Current dense Granite for reasoning, code and multilingual tools.

128K context · 512K extended

Q4 · 8K ≈ 7.2 GiB
NVIDIAReasoning

Nemotron Nano 9B v2

8.9B stored parameters

Small dense hybrid model for reasoning and non-reasoning tasks.

128K context

Q4 · 8K ≈ 6.2 GiB
Mistral AIGeneral

Ministral 3 8B

8.9B stored parameters

Text and image model with a 256K supported context.

256K context

Q4 · 8K ≈ 7.0 GiB
QwenGeneral

Qwen3.5 9B

9.7B stored parameters

Text, image and coding model with optional reasoning.

256K context · 1,010,000 extended

Q4 · 8K ≈ 6.7 GiB
Z.aiGeneral

GLM-4.6V Flash 9B

10.3B stored parameters

Small vision-language model with native tool use.

128K context

Q4 · 8K ≈ 7.1 GiB
GoogleGeneral

Gemma 4 12B

12.0B stored parameters

Unified text, image and audio model with a 256K context.

256K context

Q4 · 8K ≈ 8.1 GiB
Mistral AIGeneral

Ministral 3 14B

13.9B stored parameters

Largest small dense Ministral for text and image tasks.

256K context

Q4 · 8K ≈ 10.0 GiB
MicrosoftReasoning

Phi-4 Reasoning Plus 14B

14.7B stored parameters

Dense reasoning specialist for mathematics, science and coding.

32K context

Q4 · 8K ≈ 10.8 GiB
OpenAIReasoning

gpt-oss 20B

20.9B stored parameters

Small OpenAI reasoning model supplied with native MXFP4 expert weights.

128K context

Q4 · 8K ≈ 12.9 GiB
BaiduReasoning

ERNIE 4.5 21B-A3B Thinking

21.8B stored parameters

Compact sparse ERNIE with reasoning and a 128K context.

128K context

Q4 · 8K ≈ 13.6 GiB
Mistral AICoding

Devstral Small 2 24B

24.0B stored parameters

Coding agent model with image input for software development.

256K context

Q4 · 8K ≈ 15.7 GiB
GoogleGeneral

Gemma 4 26B-A4B

25.8B stored parameters

Sparse text and image model with about 3.8B active parameters.

256K context

Q4 · 8K ≈ 15.9 GiB
QwenGeneral

Qwen3.8 27B

27.8B stored parameters

Current dense Qwen for coding, reasoning and image understanding.

256K context · 1,000,000 extended

Q4 · 8K ≈ 17.4 GiB
IBMReasoning

Granite 4.2 30B

29.3B stored parameters

Largest current dense Granite for reasoning and agent tasks.

128K context · 512K extended

Q4 · 8K ≈ 19.7 GiB
QwenCoding

Qwen3 Coder 30B-A3B

30.5B stored parameters

Dedicated coding model with conventional attention and a 256K native context.

256K context · 1024K extended

Q4 · 8K ≈ 19.2 GiB