Qwen 3.8 27B at 50 tok/s with 100k Context on a 16GB GPU! (beellama.cpp)
2.55T2 sourcer/LocalLLaMA
Source record
Published by r/LocalLLaMA (T2 source). The original is at https://www.reddit.com/r/LocalLLaMA/comments/1w1lq7u/qwen_38_27b_at_50_toks_with_100k_context_on_a/.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryReddit post demonstrating a 27B-parameter Qwen model running at 50 tokens/second with 100k context window on a single 16GB GPU via a fork called beellama.cpp. Specific hardware and quantization details are not provided in the excerpt.
Why it mattersConcrete inference numbers on a single consumer GPU. Useful baseline for anyone sizing local LLM hardware, though the post lacks methodology details.
Cited by
No citations on record.
