Kimi K3 full model running on 16x GB10 cluster at 20+tps
2.10T2 sourcer/LocalLLaMA
Source record
Published by r/LocalLLaMA (T2 source). The original is at https://www.reddit.com/r/LocalLLaMA/comments/1vfl525/kimi_k3_full_model_running_on_16x_gb10_cluster_at/.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryA user reports running the full Kimi K3 model on a 16x GB10 (Grace Blackwell) cluster at over 20 tokens per second inference speed. No further details on configuration or context are provided.
Why it mattersConcrete tokens-per-second figure on a rare GB10 cluster setup; relevant to anyone sizing local infrastructure for very large models.
Cited by
No citations on record.
