On this pageThe model has a limited working pageYour question is only part of the textSaved history is not the same as active contextA larger window does not promise perfect recallLonger text can need more computer memoryTry it: keep the important facts in viewSources
DT

Written and reviewed by DeviceterraPractical guidance · Updated October 7, 2026

THE SHORT ANSWER

The context window is the limit on the text pieces a model can use for a request. It is different from saved chat history and from what the model learned during training.

WATCH THE EXPLANATION

IBM Technology: What is a Context Window?

Watch the explanation of how much text fits into a request. The word token means a small text piece, as explained in lesson 5. A larger window is a capacity limit, not proof of perfect recall.

Watch on YouTube or turn on captions ↗
THE FULL GUIDE

Here is what matters, why it matters and how to check it on your own setup.

01

The model has a limited working page

Imagine answering a question while looking at one page of notes. Anything left off that page cannot help you answer. A language model also receives a limited amount of information for each request.

That limit is called its context window. The amount is counted in tokens, the small text pieces from lesson 5. A limit of 4,000 tokens is not a promise of 4,000 words.

The notes page is an analogy, not a picture of how the model thinks. The real system performs calculations using the text it receives.

02

Your question is only part of the text

A chat app may send instructions, earlier messages, document passages, and your latest question together. All of those can take up space. The model also needs room to produce a reply.

For an ordinary text chat, think about the request and reply together. Some models have separate reply limits too. Check the exact model and app rather than assuming every product uses the same limit.

Your question is only part of the text
Your question is only part of the textInstructions + earlier chatYour question + useful notesRoom for the reply

These pieces share a limited text budget. This diagram groups them for learning; it is not an exact token count or a picture of computer memory.

03

Saved history is not the same as active context

Your app can save a conversation on disk. That does not mean it sends every saved message back to the model for every answer.

When a conversation is too long, an app may shorten it, leave out messages, or refuse the request. The exact behaviour depends on the software. Check its documentation and any warning it shows.

If you add a fact to a chat, the model may use that fact in a later reply while it is included in the request. This does not, by itself, change the model’s learned values.

04

A larger window does not promise perfect recall

In the 2024 study Lost in the Middle, researchers tested questions about long collections of text. For the models and tasks they studied, moving the useful information into the middle could reduce answer accuracy.

That finding does not tell us the accuracy of every current model. It tells us why a published capacity limit should not be treated as a guarantee that every detail will be used correctly.

For a school notice, ask the model to quote the line containing the opening time. Then compare that line with the original notice yourself.

05

Longer text can need more computer memory

Ollama’s documentation says that increasing the context length increases the memory needed to run a model. More available text space can therefore make a local setup harder to fit.

The size printed on a model page and the amount your app actually uses may differ. Do not turn the setting up to its maximum just because the model advertises a large limit.

Begin with the text you need for the task. Remove repeated notes. For important facts, keep the original document and check the answer against it. This is practical guidance, not a measured memory result for your computer.

06

Try it: keep the important facts in view

Use this made-up note: “The library opens at 10 am on Tuesday. Returns go in the blue box.” Ask: “When does the library open? Quote the line you used.”

Now add a few harmless paragraphs around the note and try again. Compare the answer with the note. A correct answer on this small exercise is not proof that the model can handle a whole book.

If the app shows a context setting or usage meter, record it. Otherwise, do not guess the number. The purpose is to practise checking the source, not to measure a model’s maximum capacity.

Quick check: open the recap

Can a chat be saved but too long to send in full? Yes. Saved history and active context are different.

Finished this lesson?

Mark it complete when you have read the lesson and tried the exercise. This saves progress on this browser. It is your own assessment, not a test score.

RESEARCH SOURCES

Sources for this lesson

Official documentation supports product and model facts. Community discussions show real setups, failures, and questions. A community result is supporting evidence, not a promise that another computer will perform the same way.

YOUR NEXT STEP

Watch a real test. Check what fits your computer.

See practical local AI tests from DeviceTerra, then use MamiLens to build a hardware-aware shortlist for your own setup.

This guide is educational. Model software, licenses, and hardware support can change. Check official sources before an important deployment.