Loading the context…
Loading the context…
Reusing previously processed input in an AI service rather than treating all of it as newly processed input.
On a cache miss, the provider processes eligible input and may write a cache entry. A later request can reuse a matching entry while it remains valid. Changing text earlier in the prefix can prevent reuse of material after that change. Minimum lengths, expiry times, write charges and discounts vary by provider; stable instructions followed by changing questions are often easier to reuse than constantly reordered input.
Imagine four requests that each contain the same 12,000-token report followed by a different 1,000-token question. Without reuse, they contain 52,000 input tokens altogether. If the first report is cached and all three later requests hit that cache, 36,000 report tokens are reused, while the first 12,000 plus the four questions total 16,000 newly processed input tokens. Output and any cache-write charges must still be accounted for separately.