Ling-3.0-flash quant ladder on one DGX Spark: the whole thing sits in a 32 to 40 tok/s band
2.40T2 sourcer/LocalLLaMA
r/LocalLLaMAoriginal source ↗
Source record
Published by r/LocalLLaMA (T2 source). The original is at https://www.reddit.com/r/LocalLLaMA/comments/1vlmun8/ling30flash_quant_ladder_on_one_dgx_spark_the/.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryA user reports running the Ling-3.0-flash model across different quantization levels on a single NVIDIA DGX Spark, with all variants producing throughput in the 32 to 40 tokens per second range.
Why it mattersHardware-specific throughput range for a current open-weight model on the new DGX Spark, useful as a buyer-side reference point.
Cited by
No citations on record.
