CUDA graphs expose the real decode bottleneck

The benchmark showed kernel-launch overhead, not model math, was what finally limited decode throughput

Started 1.41× slower than vLLM. Added continuous batching -> still behind. Added torch.compile -> somehow got worse. Added CUDA graphs -> 452 vs 460 tok/s. Nearly identical. kernel launch overhead is the real bottleneck at decode time, n
Ranked #17 on backlist 2026-06-26 (26 Jun 2026 UTC) · by (Jaydev) ·

How it ranks: Backlist reads my Twitter/X timeline, scores every tweet for substance with an LLM rubric (not engagement), and publishes the daily top picks with a one-line takeaway. Curated by Surya Dantuluri.