Running DeepSeek V4 Flash Q4_K_XL at ~100 tok/s prompt processing on 4× RTX 3060 12GB
2.25T2 sourcer/LocalLLaMA
r/LocalLLaMAoriginal source ↗
Source record
Published by r/LocalLLaMA (T2 source). The original is at https://www.reddit.com/r/LocalLLaMA/comments/1vrqf4f/running_deepseek_v4_flash_q4_k_xl_at_100_toks/.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryReddit user reports running DeepSeek V4 Flash in Q4_K_XL quantization, achieving approximately 100 tokens/second prompt processing on a 4× RTX 3060 12GB GPU setup.
Why it mattersConcrete consumer-GPU benchmark figure for a specific quantized model; useful reference point for readers planning multi-GPU local inference on budget hardware.
Cited by
No citations on record.
