Two ~300B MoE models, each on ONE 128 GB mini PC (AMD Strix Halo): GLM-5.3-Flash at ~580 tok/s prefill, MiMo-V2.6-Flash up to 44 tok/s decode. EXL3 weights + open ROCm engine
3.15T2 sourcer/LocalLLaMA
Source record
Published by r/LocalLLaMA (T2 source). The original is at https://www.reddit.com/r/LocalLLaMA/comments/1wwocik/two_300b_moe_models_each_on_one_128_gb_mini_pc/.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryA report of running two ~300B parameter MoE models (GLM-5.3-Flash and MiMo-V2.6-Flash) on a single 128 GB AMD Strix Halo mini PC, achieving roughly 580 tok/s prefill and up to 44 tok/s decode, using EXL3 weights and an open ROCm inference engine.
Why it mattersConcrete inference benchmarks for large MoE models on a single consumer mini PC, useful for sizing local LLM and agent infrastructure.
Cited by
No citations on record.
