Qwen3.8-Flash-Next turns 4xR9700 into a local AI powerhouse! 120 t/s TG and 12k t/s PP single request with optimized vLLM
2.40T2 sourcer/LocalLLaMA
Source record
Published by r/LocalLLaMA (T2 source). The original is at https://www.reddit.com/r/LocalLLaMA/comments/1w2my8q/qwen38flashnext_turns_4xr9700_into_a_local_ai/.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryA Reddit user reports running Qwen3.8-Flash-Next on four AMD Radeon RX 9700 GPUs with an optimized vLLM build, achieving 120 tokens/s generation and 12,000 tokens/s prompt processing in a single request.
Why it mattersFirsthand RDNA 4 local LLM benchmarks with vLLM are uncommon; the throughput numbers give a concrete reference for AMD-based local inference rigs.
Cited by
No citations on record.
