Prompt caching
Prompt caching
Many models can cache repeated prompt prefixes such as long system prompts, documents or tool definitions. Cached input is billed at a lower "cache read" price and responds faster.
By vendor
Claude example
Checking cache hits
Chat Completions: usage.prompt_tokens_details.cached_tokens Anthropic Messages: usage.cache_read_input_tokens and usage.cache_creation_input_tokens
Cache write and read prices are shown on each model page in the model catalog, and usage records in the dashboard list cached tokens separately.
Put stable content (system prompt, tool definitions, reference material) first and the parts that change last to maximize cache hits.