Two 96 GB Ascend cards crun Qwen3.8-flash-next hardware notes, vLLM work, benchmarks, and what is next
2.85T2 sourcer/LocalLLaMA
Source record
Published by r/LocalLLaMA (T2 source). The original is at https://www.reddit.com/r/LocalLLaMA/comments/1wvt1m4/two_96_gb_ascend_cards_crun_qwen38flashnext/.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryReddit post documenting firsthand experience running a Qwen3 8B model on two 96 GB Huawei Ascend cards via vLLM, including hardware notes, benchmark results, and plans for further work.
Why it mattersFirsthand Ascend hardware benchmarks with vLLM in English are rare; useful baseline for anyone evaluating non-NVIDIA accelerators for local LLM serving.
Cited by
No citations on record.
