Loading the context…
Loading the context…
Units into which an AI model splits input and output, such as words, word fragments or punctuation.
A tokenizer converts the input into a sequence of token identifiers. The model uses that sequence to generate more tokens, which are decoded back into text. Providers commonly measure input and output usage in tokens. Instructions, earlier messages and tool results can consume input capacity too, even when the visible new question is short. The exact count depends on the tokenizer and request format.
Imagine a hypothetical request containing 2,000 tokens of instructions, 6,000 tokens of documents and a 500-token question. Its text input totals 8,500 tokens before any additional formatting overhead. If the answer contains 1,000 output tokens, input and output usage are different quantities: 8,500 and 1,000. To estimate cost, multiply each quantity by its applicable rate and then add them; a per-million rate requires dividing token counts by one million.
No published articles using this entry yet.