WritingBasetenBasetenpublished Jul 23, 2026seen 3d

Introducing Glm 52 Fast

Open original ↗

Captured source

source ↗
published Jul 23, 2026seen 3dcaptured 3dhttp 200method plain

Introducing GLM 5.2 Fast Announcing our Series F . Learn more

News

Introducing GLM 5.2 Fast

Introducing our new tier of Fast Model APIs.

Authors

Abu Qader

Philip Kiely

Alina Weinstein

Last updated July 23, 2026

Share

Today, we're introducing GLM-5.2 Fast: a separate Model API tier that serves the same GLM-5.2 weights on infrastructure tuned for per-user throughput for real-time applications. You can try it here now or switch over your GLM-5.2 standard to the endpoint zai-org/GLM-5.2-Fast .

Why agentic workloads need a different tier When GLM-5.2 shipped, it was strong enough at coding and tool use that teams started moving real-time agentic workloads onto it. Agentic workflows are systems: a main agent plans the work, breaks it into smaller tasks, delegates those tasks to specialized subagents, reviews their outputs, runs tools or code, and then merges the results back into the final answer or artifact. Each handoff creates another reasoning loop, another inference call, and another place where latency can compound. As these systems take on more important work, speed matters more because it determines whether the whole workflow keeps momentum. GLM-5.2 Fast transforms this workflow because it’s an open model we trust as the main agent, not just for the subagents around it. It’s smart enough to coordinate the workflow, cost-effective enough to scale across subagents, and fast enough to keep the entire system competitive. GLM-5.2 Fast preserves ease-of-use with tighter performance guardrails GLM 5.2 Fast serves the same weights as standard GLM-5.2, on infrastructure provisioned differently for real-time agentic use cases. It offers: Ease-of-use: Model APIs are just that; OpenAI-compatible API endpoints you point your code at. There’s no infrastructure to manage, and you pay only for what you use (per 1M tokens).

Tighter performance SLAs: our Fast tier is built to handle even the most variable, bursty workloads while maintaining a small variance in throughput and latency metrics.

Switching from standard GLM-5.2 Because Fast is just a model slug, switching is a one-line change: 1 import os 2 from openai import OpenAI 3 4 client = OpenAI( 5 api_key=os.environ[ "BASETEN_API_KEY" ], 6 base_url= "" , 7 ) 8 9 response = client.chat.completions.create( 10 model= "zai-org/GLM-5.2-Fast" , 11 messages=[{ "role" : "user" , "content" : "Refactor this function to be async." }], 12 stream= True , 13 ) You can run Fast alongside standard GLM-5.2, route by workload — agentic loops to Fast, batch summarization to standard — and see usage per tier separately in your dashboard. Most real deployments mix latency-sensitive and latency-tolerant traffic; routing each to the tier it needs cost-efficiently. Getting started Access is fully open at launch. Customers can access directly through the UI or through their account team. Switching over is a one-line model slug change. You can start using it here .

Subscribe to our newsletter Stay up to date on model performance, inference infrastructure, and more. ‌

Explore Baseten today Start deploying Talk to an engineer

Related posts View all News

News Announcing our Series F

Tuhin Srivastava 3 others

News Welcome, Gabe Stern!

Dannie Herzberg 1 other

News Welcome, Sameer Paranjpye!

Amir Haghighat

Popular models GLM-5.2 Fast

Inkling

GLM-5.2

Kimi K2.7 Code

Whisper Large V3

NVIDIA Nemotron 3 Ultra

Explore all

Popular models GLM-5.2 Fast

Inkling

GLM-5.2

Kimi K2.7 Code

Whisper Large V3

NVIDIA Nemotron 3 Ultra

Explore all