I forked Ninfer 3090 and converted it to run on the CMP170HX - doubled my Qwen3.6-35B from llama.cpp
2.85T2 sourcer/LocalLLaMA
Source record
Published by r/LocalLLaMA (T2 source). The original is at https://www.reddit.com/r/LocalLLaMA/comments/1vvjxg1/i_forked_ninfer_3090_and_converted_it_to_run_on/.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryUser forked the Ninfer 3090 llama.cpp backend and adapted it to run on a CMP170HX mining GPU, reporting a roughly 2x speedup for Qwen3-35B inference compared to the original port.
Why it mattersFirsthand CUDA backend port that unlocks a cheap second-hand mining card for usable 35B model throughput — a specific, reproducible datapoint few others have published.
Cited by
No citations on record.
