A DEVICETERRA PRODUCT

Technology for emerging-market businesses and institutions.

Explore Deviceterra ↗
LOCAL AI · YOUR HARDWARE · YOUR TASK

Find AI that fits your computer.

A practical shortlist, a compatible engine, and a setup you can test.

Estimates, explained.Memory fit is calculated. Speed and quality need a real test.How matching works →
Optional · saved only on this browser
YOUR MAMILENS JOURNEYHardware → model → software → a working setup.
STEP 1 · TELL US ABOUT YOUR COMPUTER

What computer will run the AI?

Answer the simple questions below. If you do not know a detail, MamiLens will use the safer option.

Open the dedicated hardware report →
START WITH THE CLOSEST DEVICE

You can correct every answer below. A starter never claims to know your exact hardware.

1 · YOUR COMPUTER
2 · GRAPHICS
3 · WHAT YOU WANT TO DO
4 · STORAGE SPACE
Can the model be downloaded safely?

Enter the free space on the drive where you will keep the model. MamiLens does not read your files or disk automatically.

System memory32 GBUsable verified VRAMnot assumedGPU catalogue151 profilesSpeedbenchmark required
Complete these four sections to receive a recommendation.

Enter your computer, graphics, main job and free disk space above. MamiLens does not detect your device automatically.

WANT MORE CHOICES?Explore all 98 modelsOpen a searchable list of every source-checked model
OPTIONAL MODEL EXPLORER

98 source-checked profiles

Use this only when you want to search beyond the three recommendations.

⌕
Q

Qwen 3.5

Alibaba Cloud · 9B

Recommended to test
Officially supported35/100 evidence coverage

Qwen 3.5 9B is a curated dense profile for general, coding, reasoning workloads.

GeneralCodingReasoningVision
MODEL FILE~6.6 GB
SAFE RAM FLOOR13 GB
GPU TARGET10 GB
ARCHITECTUREDense
The task and system-memory plan fit. Speed still needs a test on this computer.
L

Llama 3.1

Meta · 8B

Recommended to test
Officially supported35/100 evidence coverage

Llama 3.1 8B is a curated dense profile for general, writing, coding workloads.

GeneralWritingCodingTools
MODEL FILE~4.9 GB
SAFE RAM FLOOR11 GB
GPU TARGET8 GB
ARCHITECTUREDense
The task and system-memory plan fit. Speed still needs a test on this computer.
Q

Qwen 3

Alibaba Cloud · 8B

Recommended to test
Officially supported35/100 evidence coverage

Qwen 3 8B is a curated dense profile for general, coding, reasoning workloads.

GeneralCodingReasoningTools
MODEL FILE~5.2 GB
SAFE RAM FLOOR12 GB
GPU TARGET8 GB
ARCHITECTUREDense
The task and system-memory plan fit. Speed still needs a test on this computer.
G

Granite 3.2

IBM · 8B

Recommended to test
Officially supported35/100 evidence coverage

Granite 3.2 8B is a curated dense profile for general, reasoning, tools workloads.

GeneralReasoningToolsRAG
MODEL FILE~5 GB
SAFE RAM FLOOR11 GB
GPU TARGET8 GB
ARCHITECTUREDense
The task and system-memory plan fit. Speed still needs a test on this computer.
Q

Qwen 2.5

Alibaba Cloud · 7B

Recommended to test
Officially supported35/100 evidence coverage

Qwen 2.5 7B is a curated dense profile for general, writing, coding workloads.

GeneralWritingCodingReasoning
MODEL FILE~4.7 GB
SAFE RAM FLOOR11 GB
GPU TARGET7 GB
ARCHITECTUREDense
The task and system-memory plan fit. Speed still needs a test on this computer.
Q

Qwen 2.5 VL

Alibaba Cloud · 7B

Recommended to test
Officially supported35/100 evidence coverage

Qwen 2.5 VL 7B is a curated dense profile for vision, general, research workloads.

VisionGeneralResearchMultilingual
MODEL FILE~6 GB
SAFE RAM FLOOR12 GB
GPU TARGET9 GB
ARCHITECTUREDense
The task and system-memory plan fit. Speed still needs a test on this computer.
M

Mistral

Mistral AI · 7B

Recommended to test
Officially supported35/100 evidence coverage

Mistral 7B is a curated dense profile for general, writing, tools workloads.

GeneralWritingTools
MODEL FILE~4.1 GB
SAFE RAM FLOOR11 GB
GPU TARGET7 GB
ARCHITECTUREDense
The task and system-memory plan fit. Speed still needs a test on this computer.
Q

Qwen 3.5

Alibaba Cloud · 4B

Recommended to test
Officially supported35/100 evidence coverage

Qwen 3.5 4B is a curated dense profile for general, coding, reasoning workloads.

GeneralCodingReasoningVision
MODEL FILE~3.4 GB
SAFE RAM FLOOR10 GB
GPU TARGET6 GB
ARCHITECTUREDense
The task and system-memory plan fit. Speed still needs a test on this computer.
Q

Qwen 3

Alibaba Cloud · 4B

Recommended to test
Officially supported35/100 evidence coverage

Qwen 3 4B is a curated dense profile for general, coding, reasoning workloads.

GeneralCodingReasoningTools
MODEL FILE~2.5 GB
SAFE RAM FLOOR9 GB
GPU TARGET5 GB
ARCHITECTUREDense
The task and system-memory plan fit. Speed still needs a test on this computer.
G

Gemma 4

Google · E4B QAT

Recommended to test
Officially supported35/100 evidence coverage

Gemma 4 E4B QAT is a curated dense profile for general, writing, reasoning workloads.

GeneralWritingReasoningCoding
MODEL FILE~6.1 GB
SAFE RAM FLOOR13 GB
GPU TARGET8 GB
ARCHITECTUREDense
The task and system-memory plan fit. Speed still needs a test on this computer.
N

Nemotron 3 Nano

NVIDIA · 4B

Recommended to test
Officially supported35/100 evidence coverage

Nemotron 3 Nano 4B is a curated dense profile for general, reasoning, tools workloads.

GeneralReasoningToolsAgents
MODEL FILE~2.8 GB
SAFE RAM FLOOR9 GB
GPU TARGET6 GB
ARCHITECTUREDense
The task and system-memory plan fit. Speed still needs a test on this computer.
G

Gemma 3

Google · 4B

Recommended to test
Officially supported35/100 evidence coverage

Gemma 3 4B is a curated dense profile for general, writing, reasoning workloads.

GeneralWritingReasoningVision
MODEL FILE~3.3 GB
SAFE RAM FLOOR10 GB
GPU TARGET6 GB
ARCHITECTUREDense
The task and system-memory plan fit. Speed still needs a test on this computer.

Showing 12 of 98 matches

STEP 5 · CHOOSE THE SOFTWARE

Confirm a model to continue.

The engine and installation steps stay locked until you make your own model choice.

Return to the matches
READY FOR HUMAN REVIEW?

Get a checked setup before you spend time downloading.

During the public launch, Deviceterra can check the selected model, artifact, engine and hardware at no cost.

OPTIONAL PLANNING TOOLS

Go further after you have a model and engine.

Estimate memory, compare candidates or plan a first deployment. No sign-up and no hardware information leaves your browser.

Open all free tools →
KEEP YOUR LOCAL AI SHORTCUT

Come back when your models change.

Save your result on this device, install MamiLens for quick access, or receive occasional updates when new hardware guides and model profiles are published.

D
THE DEVICETERRA MISSION

Use technology today. Compete globally tomorrow.

Deviceterra helps small businesses and institutions in emerging countries increase profit using technology today, and build for the global market tomorrow. MamiLens delivers one part of that mission: practical, private and affordable AI adoption.

06 · LOCAL AI ANSWERS

Clear answers before you download.

Short, practical guidance for the questions people ask when starting with local AI. Each answer links to a complete, evidence-reviewed guide.

Can 8GB of RAM run local AI?

Yes. An 8GB computer can test small, quantized local models, usually in the 1B to 4B range. Close memory-heavy apps, keep context short and expect CPU generation to be slower than GPU inference.

Read the complete guide →

Do I need a GPU for local AI?

No. A dedicated GPU improves speed, but compact models can run on a CPU with system RAM. The useful test is whether the model completes your real task at an acceptable quality and waiting time.

Read the complete guide →

Ollama or LM Studio: which is better?

LM Studio is a strong visual starting point. Ollama is often better for command-line workflows, automation and applications that need a local API. The right choice depends on how you plan to use the model.

Read the complete guide →

Is local AI completely private?

Only when the complete workflow stays local. A local model can still leak data through cloud transcription, web search, plugins, remote storage or exposed APIs. Audit every component, not only the model.

Read the complete guide →

Which local AI model should I download?

Choose the smallest proven model that fits your available memory and passes a test for your exact job. Do not choose by parameter count or popularity alone.

Read the complete guide →

How much RAM does a local AI model need?

The model file, context cache, runtime, operating system and other applications all consume memory. Leave headroom above the download size and benchmark the exact quantization on your machine.

Read the complete guide →
PUBLIC LAUNCH ACCESS

Use MamiLens freely while we learn with the community.

The planning tools and configuration review are public during this launch phase. Clear paid services may come later, after MamiLens has earned trust and demonstrated consistent value.

MODEL ADVISOR

Free

Plan what to test first

  • 98 curated profiles
  • Conservative hardware matching
  • Exact verified Ollama commands
  • Official evidence links
SETUP SUPPORT

Request help

Tell Deviceterra where your setup is blocked

  • Engine and model troubleshooting
  • Safe first-task guidance
  • Clear next steps
  • No purchase obligation
DEVICETERRA PROJECT SOVEREIGN

Turn a compatible model
into a working system.

The model is only one component. Deviceterra installs the runtime, secures the data path, connects private knowledge and validates the workflow.

  • Worldwide remote installation
  • Private knowledge assistants
  • Local coding environments
  • School and SMB deployments
REVENUE PATH

One assessment. A clear outcome.

We first confirm that your hardware and use case are suitable. If local AI is the wrong answer, we say so before proposing a deployment.

Current recommended starting pointComplete the setup questions32 GB RAM · General · Windows
MAMILENS KNOWLEDGE BASE

Learn before you install.

Complete, evidence-first guides live here on MamiLens. Each one turns a technical decision into a practical next step.

Explore all guides →
LOCAL CONTENT · GLOBAL AUTHORITY

MamiLens owns the full guides and internal search journey. Relevant articles link to Deviceterra for the wider technology and deployment perspective.

EVIDENCE POLICY

What “works” means here.

1. Memory fit is not speed

A profile passes only when weights, runtime and selected context fit conservatively. Responsiveness still requires a benchmark on the actual machine.

2. Full GPU, hybrid and CPU are different

MamiLens labels full-VRAM fits, partial offload and CPU/RAM execution separately instead of treating every runnable model as equal.

3. Multi-GPU is runtime-dependent

Combined VRAM is useful only when the runtime supports the model, quantization and sharding topology. Interconnect and offload strategy matter.

4. Dense and MoE behave differently

MoE models load all weights but activate only some experts per token. Active parameters can improve compute efficiency without reducing storage.

5. Context consumes working memory

A published maximum is not a promise that your hardware can use it. The engine adds more headroom as the selected context grows.

6. Evidence has two layers

Official pages establish specifications. Community and creator tests inform practical fit, but MamiLens never copies speed results to untested hardware.