Tokens and Tokenisation (LLMs)
In language models, tokenisation splits text into tokens, units such as words, word fragments or characters, each mapped to an integer ID. Models read and generate tokens, and context limits, speed and usage pricing are measured in them.
Most modern LLMs use subword tokenisers built with algorithms such as byte-pair encoding (BPE), WordPiece or Unigram, often through libraries such as SentencePiece. Frequent words become single tokens, while rare words, numbers and code are split into several pieces. Each token ID is mapped to an embedding vector that the model processes, and multimodal models also convert images and audio into tokens.
Token counts determine how much text fits in a model's context window, how long generation takes and, for hosted models, what usage costs. Engineers estimate token budgets when designing retrieval-augmented generation, size document chunks and prompts to fit, and monitor tokens per second as a key inference performance metric.
The same text can produce different token counts in different models, and languages that are less represented in tokeniser training data, including many Indian languages, often need more tokens per word, raising cost and latency. Tokenisation also explains some model weaknesses, such as difficulty counting letters. In data security, tokenisation means something different: replacing sensitive data with surrogate values.
Key points
- Tokens are words, subwords or characters mapped to integer IDs
- Common algorithms include byte-pair encoding, WordPiece and Unigram
- Context limits, latency and hosted-model pricing are measured in tokens
- Token counts differ between models and languages
Where AiVibe comes in
AiVibe delivers AI and machine learning services, chatbots and virtual assistants with RAG, MCP tools and voice, AI quality management including bias detection and model validation, and the AIMURUGA AI agent, and builds Intel-based edge AI devices using the Intel Distribution of OpenVINO toolkit.