Could we have a --disk-moe or --n-disk-moe like --cpu-moe or --n-cpu-moe so we can use disk/cpu/gpu ?
1.50T2 sourcer/LocalLLaMA
Source record
Published by r/LocalLLaMA (T2 source). The original is at https://www.reddit.com/r/LocalLLaMA/comments/1vgldw6/could_we_have_a_diskmoe_or_ndiskmoe_like_cpumoe/.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryA user on r/LocalLLaMA proposes adding --disk-moe or --n-disk-moe flags to llama.cpp, analogous to existing CPU MoE offloading options, to allow mixture-of-experts layers to be stored and loaded from disk for systems with limited RAM and VRAM.
Why it mattersFeature request for llama.cpp disk-tier expert offloading; documents community demand for deeper storage hierarchies but provides no working implementation or benchmark.
Cited by
No citations on record.
