Prompt Caching
Reusing the processed form of a repeated prompt prefix so you do not pay full price for it again.
When many requests share a long, stable prefix — a system prompt, a document, a set of examples — caching lets the provider skip recomputing it. The saving is large and mostly free: put the stable content first and the variable content last.
In practice: A 20,000-token manual cached once, then queried a thousand times cheaply.
Where this comes up
- ChatGPT 5.6: Release Date, Models, Pricing & Access
- Claude Fable 5 vs Opus 4.8: Benchmarks, Pricing & When to Switch
- Claude Fable 5: Price, API, Access, Safeguards & Use Cases
- Claude Opus 4.8 Release Date, Pricing, API & Claude Code
- Claude Opus 5 vs Sonnet 5: Benchmarks, Pricing & Which to Use (July 2026)
- Claude Opus 5: Benchmarks, Pricing, and Full Guide (July 2026)