I made my iPhone a second GPU for my 24 GB MacBook: Qwen 3.8 27B prefills 29–44% faster & my holds part of the CTX window.
2.85T2 sourcer/LocalLLaMA
Source record
Published by r/LocalLLaMA (T2 source). The original is at https://www.reddit.com/r/LocalLLaMA/comments/1wvz1ex/i_made_my_iphone_a_second_gpu_for_my_24_gb/.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryReddit user reports using an iPhone as a secondary GPU for a 24GB MacBook to accelerate Qwen 3.8 27B inference, with the iPhone holding part of the context window and prefill running 29–44% faster.
Why it mattersConcrete measured setup for offloading LLM compute to an iPhone, with specific model and speedup numbers, not a vague announcement.
Cited by
No citations on record.
