NVFP4 on VOLTA! Despite being built for Blackwell, I made four 2017 V100s run Qwen 3.8 NVFP4 natively and match my $6000 RTX 5090.
2.55T2 sourcer/LocalLLaMA
Source record
Published by r/LocalLLaMA (T2 source). The original is at https://www.reddit.com/r/LocalLLaMA/comments/1vsq3zg/nvfp4_on_volta_despite_being_built_for_blackwell/.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryA Reddit user claims to have run Qwen 3.8B in NVFP4 quantization natively on four 2017 NVIDIA V100 (Volta) GPUs, matching the performance of an RTX 5090, despite NVFP4 being designed for Blackwell hardware.
Why it mattersDemonstrates that older datacenter GPUs can execute newer Blackwell-era quantization formats, which could extend the usable life of legacy hardware for local inference.
Cited by
No citations on record.
