DT

Written and reviewed by DeviceterraDeviceterra editorial team · Updated September 9, 2026

KEY TAKEAWAY

You do not need a gaming GPU to start. You need a small model, free RAM, a narrow task, and patience.

TRY IT YOURSELF

Run a small model on a normal computer

You can try local AI without a dedicated graphics card. It will usually be slower, so start with a small model and a narrow task.

  1. 1

    Close heavy apps

    Save your work. Close games, video editors, and browser tabs you do not need.

  2. 2

    Install Ollama

    Visit ollama.com/download and run the installer for your operating system.

  3. 3

    Open PowerShell or Terminal

    Windows: click Start, type PowerShell, and open it. Mac: press Command and Space, type Terminal, and press Enter.

  4. 4

    Run a small model

    Type this command and press Enter. Wait for the first download to finish.

    ollama run gemma3:1b
  5. 5

    Test one short job

    Ask it to rewrite a short message. Do not paste a long document into the first test.

How to know it worked
  • The model answers without the computer freezing.
  • Normal memory use stays below the total available RAM.
  • The waiting time is acceptable for your task.
If something goes wrong
  • If the computer becomes unresponsive, close Ollama and choose a smaller model.
  • If answers become slower during a long chat, type /bye and start a new short chat.
UNDERSTAND THE DETAILS

Use the explanations below when you want to know why each step matters.

01

Choose a first model by available memory

On an 8 GB computer, begin with a 1B–2B quantized text model and one short task. A 16 GB computer has more room to test a 3B–8B model, but a larger model may respond too slowly on a laptop CPU. Installed RAM includes the memory Windows and your other apps already use.

For a reproducible small test, Qwen3 1.7B is available as an Ollama download of about 1.4 GB. That is the file size, not the total memory needed while it runs. Use the hardware calculator with your real free disk space before downloading.

A useful first task

Give it a short message and ask for a clearer rewrite without changing names, dates or amounts. Compare the result with your original. A fluent answer that changes the facts fails the test.

02

Measure usefulness instead of accepting a slow answer

Save three short examples from the work you actually do, with private details removed. Run each twice: the first attempt includes model loading, while the second better represents a model already in memory. Record how long you wait, whether the facts survived, and whether the computer remained usable.

Set your own limit before testing. For example, a draft that takes 30 seconds may be fine for occasional writing and unacceptable for live customer support. This is an example acceptance threshold, not a predicted speed.

  • If disk activity stays high and the computer freezes, reduce model size and close other apps. Do not assume adding virtual memory will make it fast.
  • If later replies slow down, start a new chat and reduce the context.
  • If answers are wrong, a faster processor will not repair them. Try a better model or narrow the task.
03

What a CPU can do

A CPU is the main processor in your computer. It can run small AI models for writing, summaries, sorting, simple coding help, and private document questions.

A CPU is usually slower than a good GPU. That may be fine if one person uses the model and the task does not need an instant answer.

04

Choose a small model

On an 8 GB computer, start with a model between 1B and 4B. On a 16 GB computer, you can often try models between 3B and 8B.

Choose a Q4 file first. Q4 is a smaller version of the model that uses less memory. These ranges are starting points, not guarantees.

Do not chase size

The best first model is the smallest one that can complete your task.

05

Keep the conversation short

Context is the text the model reads during one conversation. Long chats and large documents need more memory.

Start with short prompts. For many documents, use a search system that sends only the most useful passages to the model.

06

Simple setup steps

  • Close apps that use a lot of memory.
  • Install Ollama or LM Studio from its official website.
  • Use MamiLens to choose a small model.
  • Download one model.
  • Test five real prompts and record the waiting time.
  • Try a larger model only if the small one cannot do the job.
07

Tasks to leave for later

Large coding agents, very long documents, live web research, and many users need more memory and speed.

Start with one user and one short task. Buy better hardware only after your test shows clear value.

RESEARCH SOURCES

Official facts and real user evidence

Official documentation supports product and model facts. Community discussions show real setups, failures, and questions. A community result is supporting evidence, not a promise that another computer will perform the same way.

MAKE IT PRACTICAL

Find a model your computer can run.

MamiLens checks your hardware and shows a careful starting point.

Run the free compatibility check →

This guide is educational. Model software, licenses, and hardware support can change. Check official sources before an important deployment.