A 2.6B model with tool calling and 128K context now runs at 30 tok/s on a phone
2.25T2 sourcer/LocalLLaMA
r/LocalLLaMAoriginal source ↗
Source record
Published by r/LocalLLaMA (T2 source). The original is at https://www.reddit.com/r/LocalLLaMA/comments/1vfn9vc/a_26b_model_with_tool_calling_and_128k_context/.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryA 2.6B parameter model supporting tool calling and a 128K context window reportedly runs at 30 tokens per second on a mobile phone.
Why it mattersOn-device small models with tool use and long context at usable speeds point toward fully local agents without cloud round-trips.
Cited by
No citations on record.
