The idea: on a CPU the decode speed depends on the active params per token, not the total. My objective is trying to run a 10B at 100tok/s on a mid level PC (No GPU).
1.50T2 sourcer/LocalLLaMA
Source record
Published by r/LocalLLaMA (T2 source). The original is at https://www.reddit.com/r/LocalLLaMA/comments/1v9vo75/the_idea_on_a_cpu_the_decode_speed_depends_on_the/.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryReddit post proposing that on CPU, decode speed depends on active parameters per token rather than total parameters, with a stated goal of running a 10B model at 100 tokens/sec on a mid-level GPU-less PC.
Why it mattersFrames a concrete CPU-inference target around the active-vs-total-params distinction, useful framing for anyone sizing local LLM hardware without a GPU.
Cited by
No citations on record.
