WritingBasetenBasetenpublished Feb 25, 2026seen Jun 26

Monitor Concurrent Inference Requests

Open original ↗

Captured source

source ↗
published Feb 25, 2026seen Jun 26captured Jun 27http 200method plain

Monitor concurrent inference requests Announcing our Series F . Learn more

changelog / post

Monitor concurrent inference requests

Feb 25, 2026 Go back

Track the number of in-progress inference requests across your deployments, including both requests currently being serviced and those waiting in the queue. This is the key indicator used to drive autoscaling decisions, and is now visible in the metrics dashboard and available through metrics export. For more information, see the supported metrics docs and the autoscaling documentation . ✕ Concurrent Requests Graph

Explore Baseten today Start deploying Talk to an engineer

Popular models GLM 5.2

Kimi K2.7 Code

DeepSeek V4

GPT OSS 120B

Whisper Large V3

NVIDIA Nemotron 3 Ultra

Explore all

Popular models GLM 5.2

Kimi K2.7 Code

DeepSeek V4

GPT OSS 120B

Whisper Large V3

NVIDIA Nemotron 3 Ultra

Explore all

Notability

notability 3.0/10

Routine feature/tutorial post by Baseten, not a model release or major launch.