Qwen 3.8 Flash Next q4_k_m, 130k context, q8 cache on 16GB VRAM ann 64GB RAM, 15-20 t/s on 4080
2.40T2 sourcer/LocalLLaMA
Source record
Published by r/LocalLLaMA (T2 source). The original is at https://www.reddit.com/r/LocalLLaMA/comments/1wom3fe/qwen_38_flash_next_q4_k_m_130k_context_q8_cache/.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryA user reports running Qwen 3.8B Flash in q4_k_m quantization with 130k context and q8 KV cache on 16GB VRAM plus 64GB RAM, achieving 15-20 tokens per second on an RTX 4080.
Why it mattersConcrete hardware, quantization, and speed numbers for a current local model let readers judge fit for their own machine before downloading or buying parts.
Cited by
No citations on record.
