Qwen3.8-27B at 262K context on a Strix Halo + RTX 3090 Ti: 9.5 -> 153 tok/s, and it beats a dual-3090 vLLM box on HumanEval
2.40T2 sourcer/LocalLLaMA
Source record
Published by r/LocalLLaMA (T2 source). The original is at https://www.reddit.com/r/LocalLLaMA/comments/1vu7yce/qwen3827b_at_262k_context_on_a_strix_halo_rtx/.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryA user on r/LocalLLaMA reports running Qwen3.8-27B at 262K context on a Strix Halo plus RTX 3090 Ti setup, achieving 9.5 to 153 tok/s, and claims it outperforms a dual-3090 vLLM configuration on HumanEval.
Why it mattersConcrete first-pass benchmark for Strix Halo as an AI inference platform, with specific tok/s and a HumanEval comparison against a familiar reference rig.
Cited by
No citations on record.
