Abstract: LLMs face decoding bottlenecks. Speculative Decoding (SD) reduces latency via a small draft model for serial decoding and a large target model to verify in parallel. Despite this advantage, ...
Some results have been hidden because they may be inaccessible to you
Show inaccessible results