Nifer is insane. 700t/s with Qwen 3.6 35B (no thinking). Purpose build for RTX5090. Full 250k context too.
2.10T2 sourcer/LocalLLaMA
Source record
Published by r/LocalLLaMA (T2 source). The original is at https://www.reddit.com/r/LocalLLaMA/comments/1v8a7wb/nifer_is_insane_700ts_with_qwen_36_35b_no/.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryA Reddit post claims the Nifer inference engine reaches 700 tokens/sec on Qwen 3.6 35B in non-thinking mode, is purpose-built for the RTX 5090, and supports the full 250k context window. No benchmarks or configuration details are provided in the excerpt.
Why it mattersSingle throughput data point for a niche inference engine on the newest consumer GPU. Worth a glance for readers sizing local LLM rigs around the RTX 5090, but unverifiable from the title alone.
Cited by
No citations on record.
