WritingBasetenBasetenpublished Jul 23, 2026seen 3d

Glm 52 Fast Available On Baseten

Open original ↗

Captured source

source ↗
published Jul 23, 2026seen 3dcaptured 3dhttp 200method plain

GLM 5.2 Fast available on Baseten Announcing our Series F . Learn more

changelog / post

GLM 5.2 Fast available on Baseten

Jul 23, 2026 Go back

GLM 5.2 Fast is the first model in our new Fast tier on Model APIs: the same GLM 5.2 model weights served on dedicated capacity engineered for higher sustained per-user throughput, built for agentic coding and real-time conversational applications. It ships as its own model slug with its own pricing and rate limits, behind the same OpenAI-compatible API, so switching is a one-line change. If Fast capacity is temporarily saturated, requests keep serving on standard GLM 5.2 capacity: slower, not failed. curl https://inference.baseten.co/v1/chat/completions \ -H "Authorization: Bearer $BASETEN_API_KEY " \ -H "Content-Type: application/json" \ -d '{ "model": "zai-org/GLM-5.2-Fast", "messages": [{"role": "user", "content": "What makes an inference stack fast?"}] }' For more information, see our docs .

Explore Baseten today Start deploying Talk to an engineer

Popular models GLM-5.2 Fast

Inkling

GLM-5.2

Kimi K2.7 Code

Whisper Large V3

NVIDIA Nemotron 3 Ultra

Explore all

Popular models GLM-5.2 Fast

Inkling

GLM-5.2

Kimi K2.7 Code

Whisper Large V3

NVIDIA Nemotron 3 Ultra

Explore all