I hosted Kimi K3 (2.8T parameters) using 8 B300s. 92 tok/s, $190 per million tokens
2.70T2 sourcer/LocalLLaMA
r/LocalLLaMAoriginal source ↗
Source record
Published by r/LocalLLaMA (T2 source). The original is at https://www.reddit.com/r/LocalLLaMA/comments/1vw1j2p/i_hosted_kimi_k3_28t_parameters_using_8_b300s_92/.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryFirsthand report of self-hosting a 2.8T-parameter MoE model on 8 NVIDIA B300 GPUs, achieving 92 tokens/s inference and $190 per million tokens in cost.
Why it mattersConcrete performance and cost data for running a frontier-scale MoE model on new Blackwell hardware, useful for anyone weighing self-hosting against API pricing.
Cited by
No citations on record.
