arXiv NLP
techCenter
Margins, Not Windows: Training-Free Per-Step Lossy Speculative Decodingtranslating…
arXiv:2609.02897v1 Announce Type: new
Abstract: Speculative decoding accelerates LLM inference by drafting candidate tokens and verifying them in parallel. Tree-attention drafters such as EAGLE-3 are widely adopted, yet typically hold two decisions fixed: (1) a strict…
Keywords#Speculative Decoding#Lossy#Lossy Speculative#Lossy Speculative Decoding#Margins
Comments
Sign in to join the discussion.
Loading comments…