DeepSeek-V4-Flash-0731 UD-IQ3_S 12.5 tok/s on RTX 3090 +128GB DDR5
2.40T2 sourcer/LocalLLaMA
r/LocalLLaMAoriginal source ↗
Source record
Published by r/LocalLLaMA (T2 source). The original is at https://www.reddit.com/r/LocalLLaMA/comments/1vcz61x/deepseekv4flash0731_udiq3_s_125_toks_on_rtx_3090/.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryA user reports running DeepSeek-V4-Flash at UD-IQ3_S quantization, achieving 12.5 tokens/second on an RTX 3090 paired with 128GB DDR5 system RAM.
Why it mattersConcrete throughput number on consumer hardware at a specific quant level, useful for sizing local LLM rigs before committing to a build.
Cited by
No citations on record.
