Loading the context…
Loading the context…
The amount of information an AI model can process within a request, measured in tokens.
An application assembles the context before asking the model to respond. When it approaches the limit, it may remove older messages, select fewer documents or summarize earlier work. The provider can also impose a separate maximum response length. A larger window allows more material to fit, but fitting information inside it does not guarantee that the model will use every detail accurately.
Suppose a hypothetical model allows 12,000 total tokens. Instructions use 2,000, documents 6,000 and the conversation 1,000. That leaves at most 3,000 for a response, assuming those are the only counted items. Adding another 2,500-token document leaves only 500. The application must shorten the input, accept a shorter answer or choose a model with sufficient capacity; wishing for a 3,000-token answer cannot overcome the limit.