WritingBasetenBasetenpublished Jun 10, 2026seen Jun 26

Vllm And Sglang Metrics

Open original ↗

Captured source

source ↗
published Jun 10, 2026seen Jun 26captured Jun 27http 200method plain

vLLM and SGLang metrics Announcing our Series F . Learn more

changelog / post

vLLM and SGLang metrics

Jun 10, 2026 Go back

Baseten now surfaces engine-native metrics for models served with vLLM or SGLang directly in the Metrics tab. Baseten automatically detects the engine through your container's /metrics endpoint, then graphs metrics such as tokens per second, time to first token, KV cache usage, and requests running or queued, no configuration or redeploy required. ✕ vLLM metrics You can also export these metrics to your own observability stack alongside Baseten's standard metrics.

For more information, see our docs .

Explore Baseten today Start deploying Talk to an engineer

Popular models GLM 5.2

Kimi K2.7 Code

DeepSeek V4

GPT OSS 120B

Whisper Large V3

NVIDIA Nemotron 3 Ultra

Explore all

Popular models GLM 5.2

Kimi K2.7 Code

DeepSeek V4

GPT OSS 120B

Whisper Large V3

NVIDIA Nemotron 3 Ultra

Explore all

Notability

notability 5.0/10

Informative comparison post of two inference frameworks.