DeepSeek V4 Flash 0731 IQ2_M benchmark for Dual 3060 and 96GB RAM ≈ 3.5 tok/s.
2.10T2 sourcer/LocalLLaMA
Source record
Published by r/LocalLLaMA (T2 source). The original is at https://www.reddit.com/r/LocalLLaMA/comments/1vcrd6d/deepseek_v4_flash_0731_iq2_m_benchmark_for_dual/.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryA Reddit user reports benchmarking DeepSeek V4 Flash at IQ2_M quantization on a dual RTX 3060 with 96GB system RAM, achieving approximately 3.5 tokens per second generation speed.
Why it mattersConcrete throughput number for a specific budget multi-GPU and high-RAM config running a heavily quantized model — useful planning data for similar local LLM builds.
Cited by
No citations on record.
