On this page
Why split text into pieces?A word is not always one tokenInput tokens and output tokensWhy tokens matter for local AIDoes a token count tell you quality?Try it: compare short and long textSourcesA token is a piece of text represented by a number for a model. It can be a word, part of a word, punctuation, or another text piece.
3Blue1Brown: Large Language Models explained briefly
This animation connects small text pieces with the way a language model builds a reply. We reuse the model lesson’s video here because the two ideas connect. The diagram below gives a simpler token example; it is not an exact tokenizer result.
Watch on YouTube or turn on captions ↗Here is what matters, why it matters and how to check it on your own setup.
Why split text into pieces?
People see sentences. A text model works with numbers. A tokenizer is the tool that breaks text into pieces and gives each piece a number. Those pieces are tokens.
Think of putting a sentence into small labelled boxes. The model works with the box labels. When it produces an answer, the system turns those labels back into readable text.
This does not mean the model sees the same pieces you would choose when reading. Different models can use different rules.
A word is not always one token
A common word may fit in one token. A less common word may split into several pieces. Spaces, punctuation, and parts of numbers can also affect the count.
Our pretend split is: “Local AI!” becomes “Local”, “ AI”, and “!”. It is an illustration of the idea only. A real tokenizer may split it differently.
Illustrative split only. The actual pieces and numbers depend on the tokenizer; the space before AI may be part of a token.
Input tokens and output tokens
Input tokens are the pieces sent to the model for a request. Output tokens are the pieces it produces in the answer. A long answer generally uses more output tokens than a short one.
Your typed question may not be the whole input. The app can also include earlier messages, instructions, and selected document passages. That is why a small new question can be part of a large request.
You do not need to count pieces by hand. If the app shows token usage, use that display. For an exact count, use the tokenizer for the exact model.
Why tokens matter for local AI
A model can handle only a limited amount of text at once. This is its context window, which we will cover next. Token counts help describe that limit.
More text can increase waiting time and working memory use. A reply also needs room. Downloading a model that fits on disk does not mean it can handle an unlimited conversation.
A label such as “4,000 tokens” does not mean exactly 4,000 words. The amount of readable text depends on the language, spelling, code, and tokenizer.
Does a token count tell you quality?
No. Tokens measure text pieces, not useful ideas. A long answer can be repetitive or wrong. A short answer can solve the problem well.
You may see speed listed as tokens per second. That measures how quickly output pieces are produced in a particular test. It is not the same as accuracy, and it does not capture all the time spent loading or reading the request.
Do not compare speed numbers without checking the model, computer, settings, and test. For learning, ask whether the waiting time is acceptable and the answer is useful.
Try it: compare short and long text
If your app displays token counts, send “Hello.” Then send a short paragraph you wrote yourself. Compare the input counts. Ask for a one-sentence answer, then a longer answer, and compare the output counts.
If your app does not display counts, you can still complete the main lesson: explain why a token is a text piece rather than always a whole word. Avoid guessing exact counts.
- Tokens are pieces, not points for intelligence.
- Different tokenizers can count the same text differently.
- Input includes more than your latest sentence in many chat apps.
- Longer text needs more room; lesson 6 will explain that room.
Is “100 tokens” exactly “100 words”? No. Some words take several tokens, and punctuation may take tokens too.
Official facts and real user evidence
Official documentation supports product and model facts. Community discussions show real setups, failures, and questions. A community result is supporting evidence, not a promise that another computer will perform the same way.
Watch a real test. Check what fits your computer.
See practical local AI tests from DeviceTerra, then use MamiLens to build a hardware-aware shortlist for your own setup.
This guide is educational. Model software, licenses, and hardware support can change. Check official sources before an important deployment.
