ForkFireworks AIFireworks AIpublished Jul 8, 2026seen 2w

fw-ai/tokenspeed

forked from lightseekorg/tokenspeed

Open original ↗

Captured source

source ↗
published Jul 8, 2026seen 2wcaptured 2whttp 200method plain

fw-ai/tokenspeed

Description: TokenSpeed is a speed-of-light LLM inference engine.

License: MIT

Stars: 0

Forks: 0

Open issues: 1

Created: 2026-07-08T19:51:48Z

Pushed: 2026-07-08T20:02:24Z

Default branch: main

Fork: yes

Parent repository: lightseekorg/tokenspeed

Archived: no

README:

TokenSpeed is a speed-of-light LLM inference engine designed for agentic workloads, with TensorRT-LLM-level performance and vLLM-level usability. Our goal is to be the most performant inference engine for production agentic workloads.

Core components:

  • Modeling layer: local-SPMD design with a static compiler that generates

collective communication from module-boundary placement annotations, so users do not hand-write parallelism logic.

  • Scheduler: C++ control plane and Python execution plane. Request

lifecycle, KV cache ownership, and overlap timing are encoded as a finite-state machine, with safe KV resource reuse enforced by the type system at compile time.

  • Kernels: pluggable, layered kernel system with a portable public API and

a centralized registry including one of the fastest MLA (Multi-head Latent Attention) implementations on Blackwell for agentic workload.

  • Entrypoint: SMG-integrated AsyncLLM for low-overhead CPU-side request

handling.

News

  • [2026/06] Deep dive into the design and optimization of TokenSpeed-Kernel. [blog]
  • [2026/05] 🚀 TokenSpeed hits 580 TPS on Qwen3.5-397B-A17B for agentic workloads. [blog]
  • [2026/05] TokenSpeed announced — a speed-of-light LLM inference engine for agentic workloads. [blog]

Blogs and Talks

For technical blogs, conference talks, and engineering articles from LightSeek Foundation, visit the LightSeek Blog.

Performance Comparison

Documentation

Start here: