You can now run Qwen 3.8 27B on AMD NPUs via FastFlowLM (at a killer 1 tps decode)
2.40T2 sourcer/LocalLLaMA
r/LocalLLaMAoriginal source ↗
Source record
Published by r/LocalLLaMA (T2 source). The original is at https://www.reddit.com/r/LocalLLaMA/comments/1wud77g/you_can_now_run_qwen_38_27b_on_amd_npus_via/.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryFastFlowLM runtime now supports running Qwen 27B models on AMD NPUs, achieving roughly 1 token per second decode speed on Ryzen AI hardware.
Why it mattersConcrete NPU-side local LLM stack for AMD hardware; throughput is low, but it documents an emerging on-device inference path beyond GPUs.
Cited by
No citations on record.
