DFlash enters a production inference stack
A draft-model system moving into production signals that speculative decoding and related inference optimizations are becoming operational infrastructure
DFlash is now running in a production inference stack. More draft models coming soon. https:// github.com/z-lab/dflash