Two flags took the official Ling-3.0-flash INT4 from 20.8 to 38.7 tok/s on one DGX Spark
2.55T2 sourcer/LocalLLaMA
Source record
Published by r/LocalLLaMA (T2 source). The original is at https://www.reddit.com/r/LocalLLaMA/comments/1vjttcc/two_flags_took_the_official_ling30flash_int4_from/.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryA community post reports that adding two specific flags roughly doubled inference throughput of the Ling-3.0-flash INT4 model on a single DGX Spark, from 20.8 to 38.7 tokens per second.
Why it mattersConcrete, reproducible speedup on new local-inference hardware — useful for anyone tuning LLMs on a DGX Spark without waiting for official guides.
Cited by
No citations on record.
