Speed-up Kimi K3(2.8T) on a 16x GB10 Cluster — 30 t/s coding throughput, 136 t/s concurrency peak.
2.25T2 sourcer/LocalLLaMA
Source record
Published by r/LocalLLaMA (T2 source). The original is at https://www.reddit.com/r/LocalLLaMA/comments/1wlt577/speedup_kimi_k328t_on_a_16x_gb10_cluster_30_ts/.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryReddit post on r/LocalLLaMA reports running Kimi K3 (2.8T-parameter model) on a 16x GB10 cluster, claiming 30 t/s coding throughput and 136 t/s peak concurrency. No body content available beyond the title.
Why it mattersConcrete throughput figures for a 2.8T model on Grace Blackwell hardware serve as a reference point for sizing local inference clusters.
Cited by
No citations on record.
