PUBLIC METHOD · VERSION 1.2 · SEPTEMBER 9, 2026

How MamiLens calculates and tests local AI.

A result is trustworthy only when people can see the inputs, understand the formula, repeat the test and find what remains unknown.

Submit a test result
Same setup

Save the model, engine, settings, and hardware.

Same task

Do not change the prompt between models.

Visible failure

Wrong answers and crashes stay in the record.

Clear evidence

Labels show what MamiLens has and has not checked.

THE CALCULATOR IN PLAIN ENGLISH

What every number means

1. Model weights

The listed Q4 artifact size is the starting point. Q5 is planned at about 1.18 times that size. Q8 is planned at about 1.82 times that size. The exact downloaded file remains the final authority.

2. Context and KV cache

Longer conversations need more cache memory. The estimate uses the selected context, model family, input type, cache precision and number of people active at once. Concurrent cache is included before deciding GPU fit. Q8 and Q4 cache factors are planning heuristics; the total never falls below the single-request uncompressed allowance. Exact runtime support still needs checking.

3. System reserve

MamiLens keeps 3.5 GB to 8 GB for the operating system, engine and background work. Larger computers receive a larger reserve.

4. GPU fit

VRAM is counted as verified acceleration only when the exact GPU and supported backend are known. Several GPUs are not treated as one simple memory pool.

5. Speed range

When the model fits a verified GPU, the calculator divides published memory bandwidth by calculated model size, then applies a conservative 35 to 65 percent efficiency range. This is a planning estimate, not a measured benchmark.

6. Final recommendation

A model must pass memory, input, task and published-context rules. Entered free disk space is also checked in the hardware calculator; when it is omitted, storage remains explicitly unchecked. MamiLens then ranks useful models near the strongest sensible size for the selected computer.

PRIMARY TECHNICAL SOURCES

Where the rules come from

Hugging FaceQuantization concepts and formatsllama.cppGGUF runtime and backend behaviorOllamaMemory, GPU loading and concurrencyNVIDIACUDA GPU-family supportAMD ROCmCurrent platform compatibilityIntel ArcOfficial GPU specifications
1
STEP 1

Record the computer

  • Operating system and version
  • Exact CPU
  • Exact GPU and dedicated VRAM
  • Total RAM
  • Free storage
2
STEP 2

Record the software

  • Engine and version
  • Exact model tag or file
  • Quantization
  • Context setting
  • Driver version when relevant
3
STEP 3

Prepare the test

  • Use a real task with private details removed
  • Save the exact prompt
  • Write the expected result before testing
  • Include a question the model should not answer when useful
4
STEP 4

Run the model

  • Run one warm-up prompt
  • Use the saved prompt without changing it
  • Record the time before the first word
  • Record speed and highest memory when possible
  • Save failures and error messages
5
STEP 5

Judge the result

  • Was the main answer correct?
  • Did it follow the requested format?
  • Did it invent facts?
  • Was the waiting time useful?
  • Could another person repeat the setup?
EVIDENCE LABELS

What each label means

Calculated estimate

A formula used official specifications and the hardware details entered. Nobody has measured this exact result yet.

Submitted report

A user sent the exact setup and outcome. MamiLens has not confirmed every detail.

Community confirmed

More than one independent report supports the same useful result.

MamiLens verified

The MamiLens team reviewed the evidence and repeated or directly checked the important parts.

Have a result another person can repeat?

Contribute your test