I benchmark DFlash 2 (PR build) in llama.cpp on Qwen 3.8 27B against all speculative methods for 3 days. 2.26x on 100 real coding prompts, 4.68x with one n-gram drafter on top. Up to 8x on specific cases.
2.70T2 sourcer/LocalLLaMA
Source record
Published by r/LocalLLaMA (T2 source). The original is at https://www.reddit.com/r/LocalLLaMA/comments/1vvncyh/i_benchmark_dflash_2_pr_build_in_llamacpp_on_qwen/.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryA user benchmarked DFlash 2 (PR build) in llama.cpp on a Qwen 27B model against all speculative decoding methods over three days. DFlash alone yielded 2.26x speedup on 100 real coding prompts; pairing it with an n-gram drafter reached 4.68x overall and up to 8x on specific cases.
Why it mattersA side-by-side benchmark of every speculative method on a real coding workload is rare. The n-gram-on-top combination is a concrete, replicable finding for anyone running local inference.
Cited by
No citations on record.
