Auto-fit vs tuned MoE offload: 564 → 1330 pp tok/s, unchanged decode (Qwen3.6-35B-A3B Q6 / RTX 3090)
2.85T2 sourcer/LocalLLaMA
Source record
Published by r/LocalLLaMA (T2 source). The original is at https://www.reddit.com/r/LocalLLaMA/comments/1vh22c8/autofit_vs_tuned_moe_offload_564_1330_pp_toks/.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryBenchmark on RTX 3090 running Qwen3.6-35B-A3B Q6 shows tuned MoE offload reaches 1330 prefill tokens/sec versus 564 with auto-fit, while decode throughput stays the same.
Why it mattersConcrete 2.4x prefill speedup from tuning MoE offload on a consumer GPU, with paired decode numbers so readers can judge the trade-off for their own setups.
Cited by
No citations on record.
