Running 104GB Qwen3.8-Flash-Next on 48GB Mac at ~12 tok/s
2.70T2 sourcer/LocalLLaMA
r/LocalLLaMAoriginal source ↗
Source record
Published by r/LocalLLaMA (T2 source). The original is at https://www.reddit.com/r/LocalLLaMA/comments/1w4z94f/running_104gb_qwen38flashnext_on_48gb_mac_at_12/.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryA user on r/LocalLLaMA reports running a 104GB Qwen3.8-Flash-Next model on a 48GB Mac, achieving roughly 12 tokens per second generation speed.
Why it mattersFirsthand throughput numbers for a 100B-class model on consumer Apple silicon, useful baseline for anyone planning local LLM hardware purchases.
Cited by
No citations on record.
