Imbalanced VRAM usage between two GPUs in llama.cpp. Anyone successfully solve this?
1.65T2 sourcer/LocalLLaMA
r/LocalLLaMAoriginal source ↗
Source record
Published by r/LocalLLaMA (T2 source). The original is at https://www.reddit.com/r/LocalLLaMA/comments/1wruh65/imbalanced_vram_usage_between_two_gpus_in/.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryA user on r/LocalLLaMA asks whether anyone has resolved imbalanced VRAM usage across two GPUs when running inference with llama.cpp.
Why it mattersUseful only if the thread contains a confirmed fix; multi-GPU layer splitting in llama.cpp is a recurring pain point for local LLM setups.
Cited by
No citations on record.
