Loading the context…
Loading the context…
The system that converts text into the token units used by an AI model.
Many modern text models use subword tokenization. Instead of needing a separate vocabulary entry for every possible word, they reuse smaller pieces to represent unfamiliar words. Different algorithms and learned vocabularies choose different boundaries. The tokenizer must match the model: sending the numerical identifiers from one vocabulary to a model trained on another would change what those numbers mean.
Consider an illustrative vocabulary containing “farm”, “ers” and a space. It could represent “farmers” using two tokens, while another vocabulary might contain “farmers” as one. This is a made-up demonstration of the principle, not the output of a particular commercial tokenizer. If 100 occurrences require two tokens each, they use 200 token positions; a one-token representation uses 100, before spaces or other formatting.