Context Window
The context window is the maximum number of tokens a language model can consider at once, covering the system prompt, conversation history, retrieved documents, tool results and the response it generates.
Everything the model uses to produce an answer must fit inside the context window, measured in tokens. When a conversation or document exceeds it, older or less relevant content must be truncated, summarised or retrieved selectively. Context windows have grown substantially, but processing long inputs increases latency and compute cost, since standard attention scales quadratically with sequence length.
Context size shapes application design: retrieval-augmented generation selects relevant chunks rather than inserting whole manuals, agents manage memory by summarising earlier steps, and long-context models can process entire contracts, codebases or log files in one request.
A large window does not guarantee the model uses all of it well; research has shown that models can use information in the middle of long inputs less reliably than information at the start or end. Prompt caching can reduce the cost of repeated content. Relevant, well-ordered context generally works better than simply adding more text, and anything placed in the context of a hosted model is sent to its provider.
Key points
- The maximum number of tokens a model can process in one request
- Includes prompts, history, retrieved text, tool output and the response
- Longer inputs increase latency and cost
- Models may use information in the middle of long inputs less reliably
Where AiVibe comes in
AiVibe delivers AI and machine learning services, chatbots and virtual assistants with RAG, MCP tools and voice, AI quality management including bias detection and model validation, and the AIMURUGA AI agent, and builds Intel-based edge AI devices using the Intel Distribution of OpenVINO toolkit.
Related terms
- Tokens and Tokenisation (LLMs)AI & Machine Learning
- Large Language Model (LLM)AI & Machine Learning
- Retrieval-Augmented Generation (RAG)AI & Machine Learning
- AI AgentAI & Machine Learning
- Prompt EngineeringAI & Machine Learning