The boring way to run Deepseek V4 Flash-0731 130-150 tks - 16x5060ti 16GB over 2 PLX88096 switches
2.55T2 sourcer/LocalLLaMA
Source record
Published by r/LocalLLaMA (T2 source). The original is at https://www.reddit.com/r/LocalLLaMA/comments/1vthcwk/the_boring_way_to_run_deepseek_v4_flash0731/.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryA user reports running Deepseek V4 Flash-0731 at 130-150 tokens per second on a setup of 16x 5060ti 16GB GPUs connected through 2 PLX88096 switches, presented as a straightforward local inference configuration.
Why it mattersConcrete multi-GPU topology and real tokens-per-second figures for a specific model on consumer-grade 5060ti hardware, useful for anyone planning similar local LLM rigs.
Cited by
No citations on record.
