I turned an asymetric pair of Tesla V100s PCIe both (16 GB + 32 GB) into a surprisingly capable local LLM lab — 1.38k prompt tok/s, 40 decode tok/s with qwen3.8 27B Q6 and Q8...
2.40T2 sourcer/LocalLLaMA
Source record
Published by r/LocalLLaMA (T2 source). The original is at https://www.reddit.com/r/LocalLLaMA/comments/1wkwec5/i_turned_an_asymetric_pair_of_tesla_v100s_pcie/.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryA Reddit user reports building a local LLM inference rig from an asymmetric pair of Tesla V100 PCIe cards (16 GB + 32 GB), recording 1.38k prompt tok/s and 40 decode tok/s running Qwen 27B models at Q6 and Q8 quantization.
Why it mattersAsymmetric GPU pairings are uncommon; the concrete throughput numbers give a usable reference for anyone weighing older datacenter cards against newer consumer GPUs for local inference.
Cited by
No citations on record.
