It's unbelievable! I used the mmap function in llama.cpp to fit Qwen3.8-Flash-Next IQ3_XSS into 16G+64G RAM, and the speed still reached 26t/s.
2.55T2 sourcer/LocalLLaMA
Source record
Published by r/LocalLLaMA (T2 source). The original is at https://www.reddit.com/r/LocalLLaMA/comments/1w0t240/its_unbelievable_i_used_the_mmap_function_in/.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryA user reports using the mmap function in llama.cpp to load the quantized Qwen3.8-Flash-Next IQ3_XSS model into 16GB of primary RAM plus 64GB of secondary memory, achieving 26 tokens per second inference speed.
Why it mattersConcrete mmap technique in llama.cpp lets you exceed physical RAM on existing hardware — a specific working method for local LLM users constrained by memory.
Cited by
No citations on record.
