llama : add --n-cpu-ffn option by John-194 · Pull Request #26622 · ggml-org/llama.cpp
2.70T2 sourcer/LocalLLaMA
r/LocalLLaMAoriginal source ↗
Source record
Published by r/LocalLLaMA (T2 source). The original is at https://www.reddit.com/r/LocalLLaMA/comments/1vzp4c9/llama_add_ncpuffn_option_by_john194_pull_request/.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryA pull request to llama.cpp adds a new --n-cpu-ffn command-line option to control how many feed-forward network layers are offloaded to the CPU during inference.
Why it mattersConcrete CLI flag for partial FFN offload in llama.cpp — relevant for users running local models on VRAM-constrained hardware.
Cited by
No citations on record.
