1-bit 27B in the browser: 25–30 tok/s on a 6 GB RTX 3060 Laptop (WebGPU, no install)
2.70T2 sourcer/LocalLLaMA
Source record
Published by r/LocalLLaMA (T2 source). The original is at https://www.reddit.com/r/LocalLLaMA/comments/1wbm50k/1bit_27b_in_the_browser_2530_toks_on_a_6_gb_rtx/.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryA Reddit post claims a 1-bit quantized 27B parameter language model runs in the browser via WebGPU, delivering 25–30 tokens per second on a laptop RTX 3060 with 6GB VRAM, requiring no installation.
Why it mattersA 27B model at usable speed in a browser tab on a mid-range laptop GPU lowers the setup bar for agentic work. The specific hardware and throughput numbers make the report checkable.
Cited by
No citations on record.
