
tokenspeed
PythonTokenSpeed combines a local-SPMD modeling layer with a static compiler that generates collective communication from annotations, a C++ control plane and Python execution plane scheduler modeled as a finite-state machine for safe KV cache reuse, a pluggable layered kernel system featuring one of the fastest MLA implementations on Blackwell, and an SMG-integrated AsyncLLM entrypoint for low-overhead CPU request handling. It achieves up to 580 TPS on Qwen3.5-397B-A17B for agentic workloads and offers comprehensive documentation and getting-started guides.