WritingBasetenBasetenpublished Apr 16, 2026seen Jun 26

Cache Token Pricing For Model Apis

Open original ↗

Captured source

source ↗
published Apr 16, 2026seen Jun 26captured Jun 27http 200method plain

Cache token pricing now available for Model APIs Announcing our Series F . Learn more

changelog / post

Cache token pricing now available for Model APIs

Apr 16, 2026 Go back

Cached input tokens are billed at a discounted rate on Model APIs for all models (excluding GPT-OSS), starting April 17, 2026. Cache token pricing is applied automatically to the portion of each request that hits the KV cache. The number of cached tokens will be visible in the API response for every request. We’re introducing cache token pricing to better fit the agentic workloads we serve on Model APIs. With that in mind, your bill should decrease in proportion to your cache hit rate. Cache token pricing is visible in-app and on our website. For those with high cache hit rates, you should expect substantial savings! For more information, see our pricing table or read our docs.

Explore Baseten today Start deploying Talk to an engineer

Popular models GLM 5.2

Kimi K2.7 Code

DeepSeek V4

GPT OSS 120B

Whisper Large V3

NVIDIA Nemotron 3 Ultra

Explore all

Popular models GLM 5.2

Kimi K2.7 Code

DeepSeek V4

GPT OSS 120B

Whisper Large V3

NVIDIA Nemotron 3 Ultra

Explore all

Notability

notability 5.0/10

Substantive post on API caching, not a model release.