Vllm And Sglang Metrics
Captured source
source ↗vLLM and SGLang metrics Announcing our Series F . Learn more
changelog / post
vLLM and SGLang metrics
Jun 10, 2026 Go back
Baseten now surfaces engine-native metrics for models served with vLLM or SGLang directly in the Metrics tab. Baseten automatically detects the engine through your container's /metrics endpoint, then graphs metrics such as tokens per second, time to first token, KV cache usage, and requests running or queued, no configuration or redeploy required. ✕ vLLM metrics You can also export these metrics to your own observability stack alongside Baseten's standard metrics.
For more information, see our docs .
Explore Baseten today Start deploying Talk to an engineer
Popular models GLM 5.2
Kimi K2.7 Code
DeepSeek V4
GPT OSS 120B
Whisper Large V3
NVIDIA Nemotron 3 Ultra
Explore all
Popular models GLM 5.2
Kimi K2.7 Code
DeepSeek V4
GPT OSS 120B
Whisper Large V3
NVIDIA Nemotron 3 Ultra
Explore all
Notability
notability 5.0/10Informative comparison post of two inference frameworks.