papersSEP 10 04:00 UTC
Osprey: Target-Agnostic Pre-training Builds Stronger Draft Models for Speculative Decoding
Researchers introduce Osprey, a pre-training approach for draft models used in speculative decoding that is not tied to any specific target model. Draft models are typically tuned to a single target's output distribution, causing acceptance rates to fall when workloads change, and target-agnostic pre-training aims to keep them robust. The work focuses on achieving faster and more stable inference for large language models.