R9V Update: now ~100 tok/s in TG on Qwen3.8 Flash Next IQ4_XS on x2 R9700 + 128GB RAM. Fixed crashes with n-gram SSD streaming, improved diagnostics, plus pinned images. Q4_K_XL now supported, 50 tok/s TG.
2.85T2 sourcer/LocalLLaMA
Source record
Published by r/LocalLLaMA (T2 source). The original is at https://www.reddit.com/r/LocalLLaMA/comments/1wfroih/r9v_update_now_100_toks_in_tg_on_qwen38_flash/.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryUpdate to R9V, a local LLM inference tool, reporting ~100 tok/s on Qwen3.8 Flash Next IQ4_XS using dual R9700 GPUs with 128GB RAM, plus 50 tok/s for Q4_K_XL. Changes include n-gram SSD streaming crash fix, diagnostics improvements, and pinned images.
Why it mattersConcrete, measured performance figures for AMD R9700 inference with specific quantizations, plus a noted crash fix for n-gram SSD streaming — directly useful for local LLM operators on similar hardware.
Cited by
No citations on record.
