Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
3.45T2 sourceHacker News · Show HN (50+ points)
Source record
Published by Hacker News · Show HN (50+ points) (T2 source). The original is at https://github.com/drumih/turbo-fieldfare.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryOpen-source Swift/Metal inference engine TurboFieldfare runs 4-bit Gemma 4 26B (MoE) on M-series Macs using ~2 GB RAM by keeping shared weights and KV cache in RAM and streaming routed experts from SSD with pread and a small expert cache. Achieves 5–6 tok/s on M2 Air, 31–35 tok/s on M5 Pro. Includes OpenAI-compatible local server.
Why it mattersConcrete open-source tool that runs a 26B MoE model on 8 GB Macs via SSD-streamed experts, with measured benchmarks and reproducible experiments documented in the repo.

Cited by
No citations on record.
