PSA: llama.cpp now loads MTP tensors by default for any draft-mtp arch, even with MTP disabled
2.40T2 sourcer/LocalLLaMA
Source record
Published by r/LocalLLaMA (T2 source). The original is at https://www.reddit.com/r/LocalLLaMA/comments/1va54em/psa_llamacpp_now_loads_mtp_tensors_by_default_for/.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryA PSA noting that llama.cpp now loads Multi-Token Prediction (MTP) tensors by default for any draft-mtp architecture, even when MTP functionality is disabled, which can silently increase memory usage.
Why it mattersFlags a default behavior change in llama.cpp that may inflate VRAM use for users of draft-mtp models, even with MTP turned off — worth a quick check of your local inference setup.
Cited by
No citations on record.
