R9V Update: Created and adopted KVA projections based on Deepseek V4.1 Flash + HySparse2/MiMo-V3 for Qwen3.8 Flash Next. This is a game changer for models that don't natively implement it. 1.45-1.85x speedup in prefill to 3k+ at a small deficit to perplexity. [2x R9700, 128GB DDR5]
2.55T2 sourcer/LocalLLaMA
Source record
Published by r/LocalLLaMA (T2 source). The original is at https://www.reddit.com/r/LocalLLaMA/comments/1wp2hqk/r9v_update_created_and_adopted_kva_projections/.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryA LocalLLaMA user reports creating and adopting KVA projections ported from Deepseek V4.1 Flash and HySparse2/MiMo-V3 to Qwen3.8 Flash Next, achieving 1.45-1.85x prefill speedup up to 3k tokens on 2x R9700 GPUs with 128GB DDR5, at a small perplexity cost.
Why it mattersFirsthand port of an attention optimization to a model that lacks it natively, with concrete hardware and speedup numbers attached.
Cited by
No citations on record.
