A DEVICETERRA PRODUCT

Practical technology for businesses and institutions.

Explore Deviceterra ↗
← RAM and hardware calculator
CALCULATED MAMILENS HARDWARE PROFILE

Best Local AI Models for RTX 4090 and 64GB RAM

Calculated local AI model fits for an RTX 4090 24GB workstation with 64GB RAM.

QUICK ANSWER · CALCULATED, NOT MEASURED

Yes. A computer with 64GB RAM can run small local AI models. A safe first choice from the current MamiLens list is GPT-OSS 20B. This is a memory-fit estimate, not a speed promise.

The device we checked

System RAM64 GB
GraphicsNVIDIA GeForce RTX 4090
WorkloadReasoning and maths
ContextAbout 32K

MamiLens keeps 8GB for the operating system and background work. That leaves a working model budget of about 56GB.

Best models to start with

#1 · Recommended to test

GPT-OSS 20B

GPT-OSS 20B is a curated mixture-of-experts profile for reasoning, coding, tools workloads.

Calculated weights
14 GB
Estimated total RAM
8 GB
RAM left
56 GB
Expected path
Verified GPU
Speed
25 to 47 tokens/s estimate
Evidence
Calculated estimate
ollama run gpt-oss:20bOpen evidence-labelled model profile →
#2 · Recommended to test

Qwen 2.5 14B

Qwen 2.5 14B is a curated dense profile for general, writing, coding workloads.

Calculated weights
9 GB
Estimated total RAM
9 GB
RAM left
55 GB
Expected path
Verified GPU
Speed
39 to 73 tokens/s estimate
Evidence
Calculated estimate
ollama run qwen2.5:14bOpen evidence-labelled model profile →
#3 · Recommended to test

Phi-4 Reasoning 14B

A reasoning-focused Phi-4 variant for difficult math, science and coding tasks.

Calculated weights
11 GB
Estimated total RAM
8 GB
RAM left
56 GB
Expected path
Verified GPU
Speed
32 to 60 tokens/s estimate
Evidence
Calculated estimate
ollama run phi4-reasoning:14bOpen evidence-labelled model profile →
#4 · Recommended to test

Qwen 3 14B

Qwen 3 14B is a curated dense profile for general, coding, reasoning workloads.

Calculated weights
9.3 GB
Estimated total RAM
8 GB
RAM left
56 GB
Expected path
Verified GPU
Speed
37 to 71 tokens/s estimate
Evidence
Calculated estimate
ollama run qwen3:14bOpen evidence-labelled model profile →
#5 · Recommended to test

DeepSeek R1 14B

DeepSeek R1 14B is a curated dense profile for reasoning, math, coding workloads.

Calculated weights
9 GB
Estimated total RAM
9 GB
RAM left
55 GB
Expected path
Verified GPU
Speed
39 to 73 tokens/s estimate
Evidence
Calculated estimate
ollama run deepseek-r1:14bOpen evidence-labelled model profile →

Recommended inference engine

Ollama or llama.cpp with CUDA

This is the simplest supported starting path for NVIDIA acceleration.

More CPU cores can help, but memory bandwidth and engine settings still matter.

How accurate is this result?

This page uses the same calculator rules as the live MamiLens tool. It includes model weights, runtime overhead, operating-system reserve, context reserve and simultaneous-user cache when selected. A displayed speed range is calculated from published memory bandwidth. It is not a benchmark from this exact computer.

Start with the first model. Test a normal task, record response speed and watch memory use before using it for important work.

Other common hardware checks