Qwen 3.8 27B at ~3 BPW on an RTX 3060: GSQ vs ByteShape IQ3-XXS 2.88BPW
2.55T2 sourcer/LocalLLaMA
r/LocalLLaMAoriginal source ↗
Source record
Published by r/LocalLLaMA (T2 source). The original is at https://www.reddit.com/r/LocalLLaMA/comments/1wnechu/qwen_38_27b_at_3_bpw_on_an_rtx_3060_gsq_vs/.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryComparison of two quantization methods (GSQ and ByteShape IQ3-XXS at ~2.88 BPW) for running the Qwen 3 27B model on an RTX 3060 12GB GPU.
Why it mattersSide-by-side low-bit quantization benchmark on a common consumer GPU, directly relevant for anyone fitting a 27B model into 12GB of VRAM.
Cited by
No citations on record.
