Qwen3.8-Flash-Next on 2x3090 + DDR4: 17 → 25-29 t/s decode with the expert cache PR
2.40T2 sourcer/LocalLLaMA
Source record
Published by r/LocalLLaMA (T2 source). The original is at https://www.reddit.com/r/LocalLLaMA/comments/1w5vjp6/qwen38flashnext_on_2x3090_ddr4_17_2529_ts_decode/.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryUser reports running Qwen3.8-Flash-Next on two RTX 3090s with DDR4, achieving decode speeds of 25–29 t/s (up from 17 t/s) after applying a specific expert cache pull request for the MoE model.
Why it mattersConcrete before/after benchmark of an MoE expert-cache PR on consumer GPUs, with reproducible hardware and token-rate numbers useful to local-inference tinkerers.
Cited by
No citations on record.
