Speed numbers alone are not enough. A good test also checks whether the model completes your real task correctly.
Run a repeatable local benchmark
Online speed charts are useful clues, but they cannot tell you exactly how a model will behave on your computer. A fair local test records the full setup and uses the same prompts every time.
- 1
Record the computer
Write the operating system, CPU, RAM, GPU, VRAM, and free storage.
- 2
Record the software
Write the engine version, full model name, quantization, and context setting.
- 3
Prepare ten prompts
Use prompts from your real work. Save them in a text file so they never change between tests.
- 4
Run one warm-up prompt
The first answer may include model loading time. Run it once before measuring normal responses.
- 5
Run and score the test
Record first-word waiting time, total answer time, memory use, correctness, and instruction following.
- Someone else could repeat the test from your notes.
- You measured answer quality as well as speed.
- You can compare a future model against the saved result.
Use the explanations below when you want to know why each step matters.
Why online speed charts are not enough
Model speed changes with the computer, file type, AI app, settings, prompt length, drivers, and cooling. A result from another computer is only a clue.
Test the exact model file on the exact computer you plan to use.
Write down your setup
- Computer and operating system.
- CPU, RAM, GPU, and VRAM.
- AI app and version.
- Exact model name, file, and quantization.
- Context size and other important settings.
- Whether the model was loading for the first time.
Measure four things
First, measure how long it takes before the first word appears. Second, measure how fast the rest of the answer arrives. Third, watch the highest memory use. Fourth, score the answer itself.
A fast wrong answer is not a good result. A correct answer that takes too long may also be useless for the job.
Use the same test each time
- Create 20 real questions.
- Use the same prompt and settings.
- Keep the expected answer or scoring rule.
- Include difficult and should-refuse cases.
- Record errors, made-up facts, and broken formats.
Choose the winner
Pick the smallest model that passes your quality mark at an acceptable speed. Save the results so a new model must prove it is better.
Test again after changing the model, file, AI app, prompt, context, driver, or hardware.
Official facts and real user evidence
Official documentation supports product and model facts. Community discussions show real setups, failures, and questions. A community result is supporting evidence, not a promise that another computer will perform the same way.
Find a model your computer can run.
MamiLens checks your hardware and shows a careful starting point.
Run the free compatibility check →This guide is educational. Model software, licenses, and hardware support can change. Check official sources before an important deployment.
