DT

Written and reviewed by DeviceterraDeviceterra editorial team · Updated August 2026

KEY TAKEAWAY

Speed numbers alone are not enough. A good test also checks whether the model completes your real task correctly.

TRY IT YOURSELF

Run a repeatable local benchmark

Online speed charts are useful clues, but they cannot tell you exactly how a model will behave on your computer. A fair local test records the full setup and uses the same prompts every time.

  1. 1

    Record the computer

    Write the operating system, CPU, RAM, GPU, VRAM, and free storage.

  2. 2

    Record the software

    Write the engine version, full model name, quantization, and context setting.

  3. 3

    Prepare ten prompts

    Use prompts from your real work. Save them in a text file so they never change between tests.

  4. 4

    Run one warm-up prompt

    The first answer may include model loading time. Run it once before measuring normal responses.

  5. 5

    Run and score the test

    Record first-word waiting time, total answer time, memory use, correctness, and instruction following.

How to know it worked
  • Someone else could repeat the test from your notes.
  • You measured answer quality as well as speed.
  • You can compare a future model against the saved result.
UNDERSTAND THE DETAILS

Use the explanations below when you want to know why each step matters.

01

Why online speed charts are not enough

Model speed changes with the computer, file type, AI app, settings, prompt length, drivers, and cooling. A result from another computer is only a clue.

Test the exact model file on the exact computer you plan to use.

02

Write down your setup

  • Computer and operating system.
  • CPU, RAM, GPU, and VRAM.
  • AI app and version.
  • Exact model name, file, and quantization.
  • Context size and other important settings.
  • Whether the model was loading for the first time.
03

Measure four things

First, measure how long it takes before the first word appears. Second, measure how fast the rest of the answer arrives. Third, watch the highest memory use. Fourth, score the answer itself.

A fast wrong answer is not a good result. A correct answer that takes too long may also be useless for the job.

04

Use the same test each time

  • Create 20 real questions.
  • Use the same prompt and settings.
  • Keep the expected answer or scoring rule.
  • Include difficult and should-refuse cases.
  • Record errors, made-up facts, and broken formats.
05

Choose the winner

Pick the smallest model that passes your quality mark at an acceptable speed. Save the results so a new model must prove it is better.

Test again after changing the model, file, AI app, prompt, context, driver, or hardware.

RESEARCH SOURCES

Official facts and real user evidence

Official documentation supports product and model facts. Community discussions show real setups, failures, and questions. A community result is supporting evidence, not a promise that another computer will perform the same way.

MAKE IT PRACTICAL

Find a model your computer can run.

MamiLens checks your hardware and shows a careful starting point.

Run the free compatibility check →

This guide is educational. Model software, licenses, and hardware support can change. Check official sources before an important deployment.