A language model has no memory between conversations. Everything it knows about your particular task has to be handed to it at the moment you ask. The context window is the size of that handover: how much text the model can hold in view at once, counting your instructions, any documents you paste in, and its own reply.
Text is measured in tokens rather than words. A token is roughly three quarters of a word in English, so a thousand tokens is about seven hundred and fifty words. A dense two-page document is somewhere near a thousand tokens. Your whole employee handbook might be thirty or forty thousand.
Specific numbers change every few months as vendors ship new versions, so the useful thing is not to memorise them but to understand what each order of magnitude opens up. Check the current documentation for whichever model you actually use before designing around a figure.
This was the standard when these tools first reached the public. It holds a long email thread or a short article. Anything longer had to be chopped into pieces and stitched back together, which is why early AI writing tools produced work that lost the thread halfway through.
The current mainstream. Roughly a short book. You can drop in a full contract, a quarter of transcripts, or an entire section of your website and ask questions across all of it. For nearly every small business use, this is more than enough and the practical constraints become cost and speed rather than capacity.
Now you are into whole codebases, a year of correspondence, or a stack of long reports analysed together. Genuinely useful for specific jobs like reviewing a large document set. Also considerably slower and more expensive per request, and this is where the caveat below starts to bite.
A large context window is a capacity, not a guarantee of attention. Models reliably struggle to use information buried in the middle of a very long input. Material at the beginning and the end gets weighted more heavily. If you paste four hundred pages in and ask about something on page two hundred, the answer is markedly less reliable than if you had pasted only the relevant twenty pages.
The practical rule follows directly: give the model the smallest amount of material that could possibly answer the question. A focused ten thousand tokens beats an unfocused five hundred thousand almost every time, and it costs a fraction as much.
You pay per token, both for what you send and what comes back. This has one consequence that catches people out. In a long chat, the entire conversation is resent with every message, so the tenth exchange costs far more than the first even though your question was just as short. If you are working through a tool that bills by usage rather than a flat subscription, starting a fresh conversation for a new topic is a real economy.
Several providers now offer caching, which lets you pay a reduced rate for material you send repeatedly, such as a standing set of instructions or a reference document. If you are building anything that reuses the same background text on every request, this is worth understanding before you build it.
Most owners never need to think about tokens at all. If you use a chat assistant a few times a week through a flat monthly subscription, capacity is not your constraint and the tier differences are academic.
It starts to matter in three situations. First, if you are feeding long documents in regularly, in which case a larger window saves you the tedium of splitting files. Second, if you are paying per use through an API, where context size is directly your bill. Third, if a vendor is selling you an automation and quoting a price, since their margin depends on how much context they push through per task and you are entitled to ask.
The useful question when comparing tools is therefore not "which has the biggest window" but "how much of my material actually needs to be in view at once". For most business tasks the honest answer is: much less than you would think. If you are weighing up what this costs in practice, our note on what AI marketing actually costs a small business covers the pricing side in more detail.
Somewhere around ninety to a hundred thousand words in English, so roughly a full-length book. That figure varies with the type of text: code, tables and unusual names use more tokens per word than ordinary prose. Treat any conversion as an estimate rather than a precise limit.
No. Capacity and capability are separate. A model with a very large window can still reason poorly, and a smaller-window model can be considerably better at the task you care about. Context size determines how much material you can work with in one go, not the quality of the thinking applied to it.
Either the conversation has exceeded the window and the earliest messages have been dropped, or the material is still present but buried in the middle where attention is weakest. Restating the key facts in your latest message usually fixes it, and is faster than trying to work out which of the two is happening.
Only if you regularly hit the limit of what you have. Most small business tasks, such as drafting, summarising a report or answering questions about one document, sit comfortably inside a mainstream window. Pay for larger capacity when you have a specific job that genuinely needs it, not pre-emptively.
Want this handled for you?
Get a free audit of your website, Google reviews, and local SEO — we’ll show you exactly where you’re losing customers. Delivered in 24 hours, no sales call.
Get my free audit → or book a 15-min callWe help local businesses in Stamford, Greenwich, Norwalk, and Fairfield County implement AI marketing that generates real results.
Get Your Free AI Marketing Audit →