Cache Token Pricing For Model Apis
Captured source
source ↗Cache token pricing now available for Model APIs Announcing our Series F . Learn more
changelog / post
Cache token pricing now available for Model APIs
Apr 16, 2026 Go back
Cached input tokens are billed at a discounted rate on Model APIs for all models (excluding GPT-OSS), starting April 17, 2026. Cache token pricing is applied automatically to the portion of each request that hits the KV cache. The number of cached tokens will be visible in the API response for every request. We’re introducing cache token pricing to better fit the agentic workloads we serve on Model APIs. With that in mind, your bill should decrease in proportion to your cache hit rate. Cache token pricing is visible in-app and on our website. For those with high cache hit rates, you should expect substantial savings! For more information, see our pricing table or read our docs.
Explore Baseten today Start deploying Talk to an engineer
Popular models GLM 5.2
Kimi K2.7 Code
DeepSeek V4
GPT OSS 120B
Whisper Large V3
NVIDIA Nemotron 3 Ultra
Explore all
Popular models GLM 5.2
Kimi K2.7 Code
DeepSeek V4
GPT OSS 120B
Whisper Large V3
NVIDIA Nemotron 3 Ultra
Explore all
Notability
notability 5.0/10Substantive post on API caching, not a model release.