Qwen3.8-Flash-Next-oQ4e-mtp: 45 tok/s on M4 Max, 25 tok/s on M2 Ultra for local inference — llm-bench.io
2.40T2 sourcer/LocalLLaMA
r/LocalLLaMAoriginal source ↗
Source record
Published by r/LocalLLaMA (T2 source). The original is at https://www.reddit.com/r/LocalLLaMA/comments/1w8qo79/qwen38flashnextoq4emtp_45_toks_on_m4_max_25_toks/.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryBenchmark of a Qwen3.8-Flash-Next oQ4e mtp quant showing 45 tok/s on M4 Max and 25 tok/s on M2 Ultra for local inference, measured by llm-bench.io.
Why it mattersConcrete throughput numbers on two Apple Silicon chips for one specific quant, useful when sizing local LLM setups on Mac hardware.
Cited by
No citations on record.
