ReleaseNovita AINovita AIpublished Sep 2, 2026seen 1w

novitalabs/pegaflow v0.24.1

novitalabs/pegaflow

Open original ↗

Captured source

source ↗
published Sep 2, 2026seen 1wcaptured 1whttp 200method plain

v0.24.1

Repository: novitalabs/pegaflow

Tag: v0.24.1

Published: 2026-09-02T05:03:23Z

Prerelease: no

Release notes:

What's Changed

fix(connector): save every HMA recurrent boundary state vLLM hands off (#448)

Hybrid models (Kimi-K3, Kimi-Linear: KDA recurrent state + MLA) stored a single recurrent checkpoint per request at request_finished, so a shared prefix that did not end exactly there could not be resumed (HMA attention prefix of N blocks has no common recurrent checkpoint; recomputing instead) — a 0-block hit on agent traffic that shares a system prompt across sessions or branches mid-history.

  • Every recurrent boundary vLLM hands off (SchedulerOutput.kv_connector_block_state.boundary_state_offloads: chunk ends, internal prefill checkpoints, each block crossed while decoding) is now saved as a pinned boundary job and released once every TP worker reports it (PegaWorkerMetadata). Attention pages of hybrid requests save mid-flight at the normal per-block cadence.
  • When the external attention prefix runs past the last usable checkpoint, the connector sets Request.shared_prefix_boundary so vLLM commits the state at the end of the shared prefix; the first sharer recomputes once, later sharers resume.
  • KV cache registration accepts vLLM main's shared KV buffer views (per-layer [B, H, N, C] views at a nonzero storage offset).

HMA models now require a vLLM with the boundary-state hand-off (main since 2026-08-29). Non-hybrid models are unaffected.

Verified on Kimi-Linear-48B-A3B (single GPU): a sharer whose common prefix ends mid-way hits all shared blocks where 0.24.0 hit none; greedy warm output equals cold output.

Full Changelog: https://github.com/novitalabs/pegaflow/compare/v0.24.0...v0.24.1

Notability

notability 3.0/10

Routine minor release of pegaflow v0.24.1.