DeepSeek V4-Flash (284B MoE) at 33 tok/s single / 68 tok/s aggregate on 2× RTX 3090 + a used quad-Xeon DDR4 server — full config
2.70T2 sourcer/LocalLLaMA
Source record
Published by r/LocalLLaMA (T2 source). The original is at https://www.reddit.com/r/LocalLLaMA/comments/1veow4b/deepseek_v4flash_284b_moe_at_33_toks_single_68/.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryReddit user reports running DeepSeek V4-Flash, a 284B-parameter mixture-of-experts model, at 33 tokens/sec single-stream and 68 tokens/sec aggregate on 2× RTX 3090 GPUs combined with a used quad-Xeon DDR4 server, sharing the full configuration.
Why it mattersConcrete throughput numbers and full hardware/software config for running a 284B MoE model on consumer GPUs plus a used DDR4 server, useful for anyone sizing local LLM rigs.
Cited by
No citations on record.
