How I got Qwen 3.8 27b running at ~75t/s decode on 16GB RTX 5080
2.85T2 sourcer/LocalLLaMA
Source record
Published by r/LocalLLaMA (T2 source). The original is at https://www.reddit.com/r/LocalLLaMA/comments/1w38s2d/how_i_got_qwen_38_27b_running_at_75ts_decode_on/.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryAuthor describes achieving ~75 tokens/second decode speed running Qwen 27B (likely Qwen 2.5 27B) on a single 16GB RTX 5080 GPU, with implicit details on quantization or inference backend used.
Why it mattersSpecific throughput numbers and hardware constraints for a current-gen consumer GPU running a 27B model are worth noting for local-inference readers.
Cited by
No citations on record.
