153 tok/s on 1x AMD Radeon R9700 running Qwen3.8 27b NVFP4, 470 tok/s @ 8 conc requests, Prefill @ 3,619 tok/s
2.85T2 sourcer/LocalLLaMA
Source record
Published by r/LocalLLaMA (T2 source). The original is at https://www.reddit.com/r/LocalLLaMA/comments/1wiws8e/153_toks_on_1x_amd_radeon_r9700_running_qwen38/.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryUser reports inference throughput on a single AMD Radeon R9700 running Qwen3 27B at NVFP4 precision: 153 tok/s single-stream, 470 tok/s at 8 concurrent requests, with 3,619 tok/s prefill.
Why it mattersConcrete first-party benchmark on new AMD silicon with the emerging NVFP4 format; useful reference for anyone sizing local LLM builds around the R9700.
Cited by
No citations on record.
