Ran Qwen3.8-Flash-Next (79 GB, 2-bit) at 350K ctx for 3.5 hours on a 128 GB M5 Max — speed vs context depth, 100 turns, one graph
2.85T2 sourcer/LocalLLaMA
Source record
Published by r/LocalLLaMA (T2 source). The original is at https://www.reddit.com/r/LocalLLaMA/comments/1w26y0w/ran_qwen38flashnext_79_gb_2bit_at_350k_ctx_for_35/.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryUser benchmarked the Qwen3.8-Flash-Next model (79 GB at 2-bit quantization) running 350K context over 100 turns for 3.5 hours on a 128 GB M5 Max Mac, with a graph showing the speed-versus-context-depth tradeoff.
Why it mattersFirsthand M5 Max benchmark of a 2-bit 79 GB model at extreme context length — concrete data for sizing Apple Silicon for long-context local inference.
Cited by
No citations on record.
