Best Local AI Models for 8GB RAM Without a GPU
See which small local AI models safely fit an 8GB Windows laptop without a dedicated GPU.
Yes. A computer with 8GB RAM can run small local AI models. A safe first choice from the current MamiLens list is Qwen 2.5 1.5B. This is a memory-fit estimate, not a speed promise.
The device we checked
MamiLens keeps 3.5GB for the operating system and background work. That leaves a working model budget of about 4.5GB.
Best models to start with
Qwen 2.5 1.5B
Qwen 2.5 1.5B is a curated dense profile for general, writing, coding workloads.
- Calculated weights
- 1 GB
- Estimated total RAM
- 6 GB
- RAM left
- 2 GB
- Expected path
- CPU / RAM
- Speed
- Benchmark required
- Evidence
- Calculated estimate
ollama run qwen2.5:1.5bOpen evidence-labelled model profile →SmolLM2 1.7B
The largest SmolLM2 option for simple local chat and lightweight text tasks.
- Calculated weights
- 1.8 GB
- Estimated total RAM
- 6 GB
- RAM left
- 2 GB
- Expected path
- CPU / RAM
- Speed
- Benchmark required
- Evidence
- Calculated estimate
ollama run smollm2:1.7bOpen evidence-labelled model profile →Qwen 3 1.7B
Qwen 3 1.7B is a curated dense profile for general, coding, reasoning workloads.
- Calculated weights
- 1.4 GB
- Estimated total RAM
- 6 GB
- RAM left
- 2 GB
- Expected path
- CPU / RAM
- Speed
- Benchmark required
- Evidence
- Calculated estimate
ollama run qwen3:1.7bOpen evidence-labelled model profile →Gemma 2 2B
Gemma 2 2B is a curated dense profile for general, writing, reasoning workloads.
- Calculated weights
- 1.6 GB
- Estimated total RAM
- 6 GB
- RAM left
- 2 GB
- Expected path
- CPU / RAM
- Speed
- Benchmark required
- Evidence
- Calculated estimate
ollama run gemma2:2bOpen evidence-labelled model profile →Granite 3.2 2B
Granite 3.2 2B is a curated dense profile for general, reasoning, tools workloads.
- Calculated weights
- 1.5 GB
- Estimated total RAM
- 6 GB
- RAM left
- 2 GB
- Expected path
- CPU / RAM
- Speed
- Benchmark required
- Evidence
- Calculated estimate
ollama run granite3.2:2bOpen evidence-labelled model profile →Recommended inference engine
LM Studio or Ollama
A CPU-first setup is the safest assumption when dedicated GPU support is unknown.
CPU use is practical for small models, but speed must be tested.How accurate is this result?
This page uses the same calculator rules as the live MamiLens tool. It includes model weights, runtime overhead, operating-system reserve, context reserve and simultaneous-user cache when selected. A displayed speed range is calculated from published memory bandwidth. It is not a benchmark from this exact computer.
Start with the first model. Test a normal task, record response speed and watch memory use before using it for important work.