Extremely slow DSpark draft model performance (1-2 t/s) with DeepSeek-V4-Flash on llama-server compared to MTP?
1.95T2 sourcer/LocalLLaMA
Source record
Published by r/LocalLLaMA (T2 source). The original is at https://www.reddit.com/r/LocalLLaMA/comments/1vj8xoh/extremely_slow_dspark_draft_model_performance_12/.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryUser reports DSpark draft model achieves only 1-2 tokens/sec when paired with DeepSeek-V4-Flash on llama-server, noting poor performance relative to the model's native MTP speculative decoding.
Why it mattersConcrete throughput numbers comparing two speculative decoding paths for a specific local-inference stack. Useful if you run llama.cpp with DeepSeek drafts.
Cited by
No citations on record.
