Tuned/abliterated Qwen3.8-27b into a 24gb card 262k guff using the newest unreleased version of LexiPanel. It's fast with reliable draft acceptance. Made for 7900xtx but should work on whatever 24gb card with this setup and headless. Doesn't get dumber while coding like most of the other fine-tunes.
2.10T2 sourcer/LocalLLaMA
Source record
Published by r/LocalLLaMA (T2 source). The original is at https://www.reddit.com/r/LocalLLaMA/comments/1wvtxkx/tunedabliterated_qwen3827b_into_a_24gb_card_262k/.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryReddit user describes fitting a tuned/abliterated Qwen model into a 24GB GPU (tested on a 7900xtx, 262k context) using an unreleased LexiPanel build for speculative decoding. Claims fast inference, reliable draft acceptance, and preserved coding ability compared to other fine-tunes.
Why it mattersA concrete 24GB-GPU configuration with speculative decoding, context size, and a code-quality claim — a useful reference point for local LLM setup.
Cited by
No citations on record.
