I pushed Qwen3.8-27B to 2.000 prefill per second and 132 decode per second on A RTX 3090.
2.40T2 sourcer/LocalLLaMA
r/LocalLLaMAoriginal source ↗
Source record
Published by r/LocalLLaMA (T2 source). The original is at https://www.reddit.com/r/LocalLLaMA/comments/1w49id7/i_pushed_qwen3827b_to_2000_prefill_per_second_and/.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryReddit post claims Qwen3.8-27B runs at 2,000 prefill tokens/sec and 132 decode tokens/sec on a single RTX 3090.
Why it mattersConcrete inference throughput on consumer hardware is useful for sizing local LLM setups against specific models.
Cited by
No citations on record.
