Number of tokens to draft per iteration (default: draft model default). Recurrent (linear or sliding attention) models use more VRAM for longer drafts. This overhead multiplies with the max batch size, so for models with long drafts (e.g. DFlash with 15 tokens by default) shorter drafts may be preferable.
Declarations
Type
null or (positive integer, meaning >0)Default
nullExample
4