DT

Written and reviewed by DeviceterraDeviceterra editorial team · Updated August 2026

KEY TAKEAWAY

A model is useful only when it gives good enough answers, at a useful speed, with an acceptable level of risk.

TRY IT YOURSELF

Create a small reliability test

A model opening successfully is not proof that it can do your job. Reliability means it keeps giving useful answers across many real examples.

  1. 1

    Collect twenty examples

    Remove names and secrets. Include normal tasks, hard tasks, and questions that should not be answered.

  2. 2

    Write the correct result

    Before running the AI, write the answer or rule you will use to score each example.

  3. 3

    Set a pass mark

    For example, decide that the model must pass at least 17 of 20 examples and must not fail any safety question.

  4. 4

    Run every example

    Use the same prompt, settings, and model version. Record the answer and waiting time.

  5. 5

    Review failures

    Group failures into wrong facts, broken format, missing evidence, slow answer, and unsafe answer.

How to know it worked
  • The pass mark was written before testing.
  • Every result can be checked again.
  • You can explain why the model passed or failed.
If something goes wrong
  • If the result changes often, run the same test more than once.
  • If the prompt changes, begin a new test version instead of mixing the results.
UNDERSTAND THE DETAILS

Use the explanations below when you want to know why each step matters.

01

Running is only the first test

A model may open and answer a question, but that does not prove it can do your job. It may be slow, ignore instructions, invent facts, or give a different answer each time.

Compatibility tells you the model fits the computer. Reliability tells you whether you can depend on it.

02

Build a small test set

Collect 20 to 30 examples from the real job. Remove private details. Include easy, hard, unclear, and should-refuse questions.

Write the correct answer or scoring rule before you test. This keeps you from choosing a model only because one answer looks impressive.

03

What to measure

  • Correctness: is the main answer right?
  • Evidence: can the answer be checked?
  • Instruction following: did it use the requested format?
  • Speed: is the waiting time acceptable?
  • Consistency: does it work across many examples?
  • Safety: does it refuse when it should?
04

Set a pass mark

Decide what the model must achieve before the test. A drafting tool may allow human editing. A system that affects money, health, safety, or legal work needs much stronger checks.

If the model does not pass, try a better prompt, a different model, or a smaller job. Do not hide the failure.

Good systems can stop

It is safer to say “I do not know” than to invent an answer.

05

Keep checking after launch

  • Save the model name and version.
  • Record important settings and prompts.
  • Collect failures from real users.
  • Test again after any model, prompt, or software update.
  • Keep a way to return to the last working version.
RESEARCH SOURCES

Official facts and real user evidence

Official documentation supports product and model facts. Community discussions show real setups, failures, and questions. A community result is supporting evidence, not a promise that another computer will perform the same way.

MAKE IT PRACTICAL

Find a model your computer can run.

MamiLens checks your hardware and shows a careful starting point.

Run the free compatibility check →

This guide is educational. Model software, licenses, and hardware support can change. Check official sources before an important deployment.