Monitor Concurrent Inference Requests
Captured source
source ↗Monitor concurrent inference requests Announcing our Series F . Learn more
changelog / post
Monitor concurrent inference requests
Feb 25, 2026 Go back
Track the number of in-progress inference requests across your deployments, including both requests currently being serviced and those waiting in the queue. This is the key indicator used to drive autoscaling decisions, and is now visible in the metrics dashboard and available through metrics export. For more information, see the supported metrics docs and the autoscaling documentation . ✕ Concurrent Requests Graph
Explore Baseten today Start deploying Talk to an engineer
Popular models GLM 5.2
Kimi K2.7 Code
DeepSeek V4
GPT OSS 120B
Whisper Large V3
NVIDIA Nemotron 3 Ultra
Explore all
Popular models GLM 5.2
Kimi K2.7 Code
DeepSeek V4
GPT OSS 120B
Whisper Large V3
NVIDIA Nemotron 3 Ultra
Explore all
Notability
notability 3.0/10Routine feature/tutorial post by Baseten, not a model release or major launch.