Show HN: Needle2: 14MB agentic LLM for phones, wearables, smart home and robots
3.15T2 sourceHacker News · Show HN (50+ points)
Source record
Published by Hacker News · Show HN (50+ points) (T2 source). The original is at https://cactuscompute.com/needle.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryCactus released Needle 2, a 14MB agentic LLM (45M parameters at 2-bit compression) for edge devices including sub-$200 phones, wearables, and microcontrollers. It runs in 28MB RAM, reaches 300-1,500 tokens/sec across devices, and is based on Simple Attention Networks.
Why it mattersFirst-party release with concrete per-device benchmarks, MFLOP and power figures, and a linked paper. Relevant for anyone building on-device agentic workflows under tight size and energy budgets.

Cited by
No citations on record.
