Qwen3.8-Flash-Next 177B NVFP4(119GiB): SSD streaming at 9-10 tok/s on one 16 GB RTX 5060 Ti + 32 GB RAM
2.85T2 sourcer/LocalLLaMA
Source record
Published by r/LocalLLaMA (T2 source). The original is at https://www.reddit.com/r/LocalLLaMA/comments/1wrxap8/qwen38flashnext_177b_nvfp4119gib_ssd_streaming_at/.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryA Reddit post reports running Qwen3.8-Flash-Next 177B (NVFP4 quantized, 119 GiB) via SSD streaming on a single 16 GB RTX 5060 Ti with 32 GB system RAM, achieving 9-10 tokens/sec generation throughput.
Why it mattersConcrete NVFP4 plus SSD-offload benchmark on consumer hardware; specific GPU, RAM, and tok/s figures make it directly replicable on similar rigs.
Cited by
No citations on record.
