Dear 24G owners, try VLLM you might be able to run Qwen3.8 27B INT4, 144K FP8 KV on RTX 3090 with better speed. (TLDR VLLM AOT)
2.25T2 sourcer/LocalLLaMA
Source record
Published by r/LocalLLaMA (T2 source). The original is at https://www.reddit.com/r/LocalLLaMA/comments/1wfdtm7/dear_24g_owners_try_vllm_you_might_be_able_to_run/.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryReddit user reports that owners of 24GB VRAM GPUs such as RTX 3090 can run Qwen3 27B at INT4 quantization with 144K FP8 KV cache using VLLM's AOT compilation, achieving better speed than typical setups.
Why it mattersConcrete config combination (VLLM AOT, INT4 weights, FP8 KV cache) that fits a 27B model onto a 24GB card, with claimed speed gains worth testing.
Cited by
No citations on record.
