[Deepseek-V4-Flash-0731] Full 1M context on a single RTX5090 + DDR5 Desktop Setup with VLLM CPU/Ram Offloading, ~800 tps pp & 15+ tps decode [Agentic Coding]
2.55T2 sourcer/LocalLLaMA
Source record
Published by r/LocalLLaMA (T2 source). The original is at https://www.reddit.com/r/LocalLLaMA/comments/1vfbcgx/deepseekv4flash0731_full_1m_context_on_a_single/.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryUser reports running Deepseek-V4-Flash-0731 with 1M token context on a single RTX 5090 plus DDR5 desktop memory using VLLM CPU/RAM offloading, achieving roughly 800 tokens/sec prompt processing and over 15 tokens/sec decode, framed for agentic coding workloads.
Why it mattersConcrete throughput numbers and an offloading config for a 1M-context model on one consumer GPU give readers a benchmark to compare against their own local-inference setups.
Cited by
No citations on record.
