llama: add Maple 20B-A1B ternary MoE architecture (CPU) by AlexGabbia · Pull Request #27000 · ggml-org/llama.cpp
2.55T2 sourcer/LocalLLaMA
r/LocalLLaMAoriginal source ↗
Source record
Published by r/LocalLLaMA (T2 source). The original is at https://www.reddit.com/r/LocalLLaMA/comments/1wg1o5b/llama_add_maple_20ba1b_ternary_moe_architecture/.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryPull request to llama.cpp adds inference support for the Maple 20B-A1B ternary MoE model architecture, including CPU execution path.
Why it mattersExtends llama.cpp with a new ternary-weight MoE model and CPU runtime path, broadening local inference options for less common architectures.
Cited by
No citations on record.
