Qwen3.8-27B on 2x 3090 + vLLM + DFlash2: 218 tok/s single request
2.40T2 sourcer/LocalLLaMA
r/LocalLLaMAoriginal source ↗
Source record
Published by r/LocalLLaMA (T2 source). The original is at https://www.reddit.com/r/LocalLLaMA/comments/1vsccit/qwen3827b_on_2x_3090_vllm_dflash2_218_toks_single/.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryReddit user reports running a Qwen3 27B model on 2x RTX 3090 GPUs with vLLM and DFlash2 speculative decoding, reaching 218 tokens per second on a single request.
Why it mattersConcrete throughput number from a specific multi-GPU local inference stack; helps size 3090-class hardware against 27B-class models with speculative decoding.
Cited by
No citations on record.
