Building a Speculative Decoding Inference
Building a Speculative Decoding Inference speculative decoding (sds) is when a small "draft" model predicts multiple tokens fast, then a big "target" model verifies them all at once. if done right, you get ~2x faster generation without any