A llama.cpp PR caches “hot” MoE experts on the GPU — 33 → 56 tok/s reported with 8GB VRAM
2.70T2 sourcer/LocalLLaMA
r/LocalLLaMAoriginal source ↗
Source record
Published by r/LocalLLaMA (T2 source). The original is at https://www.reddit.com/r/LocalLLaMA/comments/1vfhns3/a_llamacpp_pr_caches_hot_moe_experts_on_the_gpu/.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryA pull request to llama.cpp adds caching of frequently-used Mixture of Experts experts on the GPU, with the reporter measuring a speed increase from 33 to 56 tokens per second on an 8GB VRAM setup.
Why it mattersUpstream PR with concrete benchmark on consumer hardware; useful for anyone running MoE models on a single GPU.
Cited by
No citations on record.
