Accepted proposals allow several output positions to advance with fewer target-model steps, while rejected tokens are corrected according to the verification algorithm. Proper methods preserve the target distribution despite the draft approximation.
Speculative decoding accelerates generation by using a faster draft process to propose tokens that a target model verifies in groups.
Accepted proposals allow several output positions to advance with fewer target-model steps, while rejected tokens are corrected according to the verification algorithm. Proper methods preserve the target distribution despite the draft approximation.