Qwen3.8-Flash-Next on 2x3090 + DDR4, part 4: 2.2-2.5x faster prefill by kicking the expert cache off the GPU while the prompt runs
3.15T2 sourcer/LocalLLaMA
Source record
Published by r/LocalLLaMA (T2 source). The original is at https://www.reddit.com/r/LocalLLaMA/comments/1wc6fsk/qwen38flashnext_on_2x3090_ddr4_part_4_2225x/.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryPart 4 of a user's investigation running Qwen3.8-Flash-Next on 2x RTX 3090 GPUs with DDR4, reporting 2.2-2.5x faster prefill by evicting the MoE expert cache off the GPU while the prompt is being processed.
Why it mattersSpecific benchmarked MoE prefill trick on consumer 3090s, with hard numbers across an ongoing series - directly usable for anyone running large MoE models on multi-GPU rigs.
Cited by
No citations on record.
